Why AI Data Centres Are Breaking Their Own Power Gear
The core problem is that AI workloads do not draw electricity the way conventional data centres do. When models are being trained, hundreds of thousands of graphics processing units can ramp up and down within milliseconds. Experts interviewed for the Bloomberg report describe power use spiking as much as 50% above design capacity, meaning a 1-gigawatt facility may briefly demand 1.5 gigawatts.
That behaviour subjects batteries, generators, cooling systems and turbines to repeated shocks they were not designed to absorb. The article reports cracked gas turbines at xAI's Colossus site in Memphis, broken cranks on small natural gas engines, and batteries installed to smooth power flows that have had to be replaced within weeks or months because of the strain.
The consequences are direct and financial. Data-centre downtime can cost between thousands and hundreds of thousands of dollars per minute, while some facilities are said to be running at closer to 80% uptime instead of the assumed round-the-clock operation. A planned 2.67-gigawatt Texas campus backed by Chevron has slipped power delivery from 2027 to 2028, partly because extra engineering time had to be added to account for the interaction between load, generation and storage.
Regulators are also paying attention. The North American Electric Reliability Corp. found that about three-quarters of the load models for the operational US data centres it assessed could not adequately represent their dynamic behaviour. NERC issued a rare level-three reliability alert requiring large data centres to address immediate risks and submit responses by 3 August.
The Reliability Gap Behind Hyperscale AI Power Demand
Power architecture was not built for millisecond swings
The mismatch is structural. Traditional data-centre power systems expect fairly steady demand, but AI training mobilises large clusters of GPUs in unison. Drew Baglino of Heron Power Electronics Co. compares the result to a 1-gigawatt facility briefly drawing 1.5 gigawatts. Most batteries, transformers and generators were not specified for that kind of rapid, repeated swing, which is why cracks and premature wear are showing up from the US and UK to the Middle East and Africa.
Why the real cost is lost compute, not broken parts
Jason Hoffman of Switch highlights the core economics: when essential power equipment fails prematurely, the main financial consequence is not replacing a pump or breaker, but losing revenue from expensive compute capacity that has gone offline. Some facilities are reportedly seeing uptime close to 80%, which undermines the assumption that data centres operate 24/7 all year. Investors and lenders could feel the impact within the next 12 to 24 months if those reliability gaps are not closed.
The grid risk: sub-synchronous oscillations and NERC's warning
The wider power system is exposed too. Schneider Electric's Sreemant Roy warns that highly dynamic data-centre loads can create sub-synchronous oscillations that damage equipment elsewhere on the network. NERC has repeatedly warned that data centres are among the greatest risks to grid stability, and its assessment of more than 33 gigawatts of operational US data centres found roughly three-quarters of load models insufficient to represent their behaviour.
Industry response: Nvidia, power electronics and smoothing tricks
The AI industry is working on fixes, but they are not yet universal. Nvidia says it has worked more closely with power experts since its Blackwell GPUs began shipping, and Heron Power Electronics is developing equipment for Nvidia's even more energy-intensive next-generation servers due in 2027. Some operators run dummy computations to keep GPU demand steadier, though critics say this wastes electricity. A US Department of Energy test bed in Colorado is now available for developers to test whether their setups can handle AI's variability.
What Data-Centre, Utility and Investment Teams Need to Do
The story's real audience includes data-centre developers, power-equipment suppliers, utilities and the investors financing AI infrastructure. The steps below follow directly from the failures and deadlines the article identifies.
- Data-centre developers: Make power integration a first-phase engineering constraint, not a post-construction fix. Joulent's Chevron-backed Texas campus pushed first power from 2027 to 2028 after adding engineering time; projects that integrate load, generation, storage and grid connection early avoid that slippage.
- Operators: Specify stabilisation equipment — batteries, capacitors, transformers, flywheels and power electronics — sized for the 50% over-design spikes Heron Power Electronics describes, not just for nominal site capacity.
- Investors and lenders: Stress-test underwriting that assumes 24/7 uptime. The article reports some AI facilities closer to 80% uptime, so model downtime at thousands to hundreds of thousands of dollars per minute and review depreciation assumptions for GPU racks and power gear within 12–24 months.
- Utilities and grid operators: Act on NERC's level-three alert, replace the roughly three-quarters of data-centre load models NERC found insufficient, and test specifically for sub-synchronous oscillations that can damage equipment elsewhere on the network.
- Equipment and energy-storage suppliers: Target the 2027 Nvidia next-generation server build-out that Heron Power Electronics is already addressing; the article identifies unmet demand for power-smoothing equipment at new data centres.
Risk & Opportunity Assessment
| Commercial Risk | High | Downtime cost estimates range from thousands to hundreds of thousands of dollars per minute, some sites are seeing closer to 80% uptime, and investors could be affected within 12–24 months if reliability gaps persist. |
| Competitive Risk | Medium | Developers that integrate power systems early, such as the Joulent-Chevron Texas campus, show the cost of delaying power delivery; operators that fix the issue late may lose contracted uptime and revenue. |
| Regulatory Risk | Medium | NERC issued a rare level-three reliability alert with a 3 August response deadline and found about 75% of assessed data-centre load models were insufficient, signalling possible tighter reliability standards. |
| Reputation Risk | Medium | Reports of cracked turbines at xAI's Colossus facility and broader depreciation concerns around AI infrastructure add to existing investor scepticism about hyperscaler spending. |
| Technology Disruption | Transformational | Nvidia's next-generation servers due in 2027 will be even more energy-intensive, and current power equipment is not designed for the large millisecond-scale load swings already seen in AI workloads. |
| Commercial Opportunity | High | The article identifies unmet demand for power-stabilisation equipment — batteries, capacitors, transformers, flywheels, power electronics and hydrogen fuel cells — from providers such as Heron Power Electronics, Mainspring Energy, GeoPura and UL Solutions. |
Comments 0