On AI Infra Bottlenecks
Biting constraints on AI infra build-out, expressed through a tightness ratio
Whichever input is scarcest sets the ceiling on AI output — and the scarcest input is about to change for good.
The capacity of the AI industry to serve its users is, at any moment, set by whichever of three physical inputs is in shortest supply: semiconductors, high-bandwidth memory or electrical power. All three are tight. The question that matters for investors, operators and policymakers is not which constraint exists, but which one clears last — and on what timeline.
Our analysis, built bottom-up from physical unit flows rather than revenue figures, points to a clear sequence. Chips were the binding constraint in 2022-24. Memory has taken over in 2025-26. And from 2027, power infrastructure becomes the sole limiter — one that, unlike its predecessors, has no capital-cycle escape hatch.
From wafers to watts
The chip shortage that defined 2023 was never really about chipmaking. Nvidia's H100 occupied less than 5 per cent of TSMC's 5nm node. The true chokepoint was advanced packaging — TSMC's CoWoS process, which mounts GPU dies and memory stacks onto silicon interposers. As TSMC's then chairman Mark Liu put it in 2023: "It is not the shortage of AI chips, it is the shortage of our CoWoS capacity."
That constraint has responded to money. TSMC built four new CoWoS facilities in two years, taking monthly capacity from roughly 8,000 wafers in mid-2023 towards a targeted 120,000-130,000 by the end of 2026. Add maturing yields on the newer CoWoS-L process — rising from about 40 per cent to 70 per cent, the mathematical equivalent of a 75 per cent capacity expansion without laying a brick — and chip tightness falls steeply. By 2027, with 8-10mn H100-class GPUs still mid-life on depreciation schedules and custom silicon eroding Nvidia's share of demand, chips approach equilibrium.
Memory lags by roughly a year, and eases more slowly. The culprit is not unit volume but the per-chip escalation: the jump from H100 to B200 raised memory content per GPU by 140 per cent, the largest generational step-up on record, arriving precisely as chip volumes doubled. Demand compounded on demand. The result is a market sold out through 2026, with SK Hynix, Samsung and Micron — an oligopoly no entrant has cracked — fully booked and Micron able to meet only 55-60 per cent of core customer demand.
Relief comes in the second half of 2027, when three independent capacity additions — SK Hynix's Yongin cluster, Micron's Idaho fab and its Singapore packaging plant — converge just as the per-chip escalation plateaus: the coming Vera Rubin generation carries the same 288GB as its predecessor, the first flat transition since 2022.
The constraint with no capital cycle
Energy is categorically different. Chips and memory are capital-cycle constrained: given money and time, fabs get built. Power is regulatory and physically constrained. Of all US grid capacity that applied for interconnection between 2000 and 2019, just 13 per cent had reached commercial operation by the end of 2024. The median completed project takes five to seven years. AI data centre demand is growing at 30-35 per cent a year; usable new grid capacity in the key US markets is growing at 3-5 per cent.
Nor can efficiency ride to the rescue. The gains are real — Google measured a 33-fold reduction in energy per Gemini prompt in a single year — but they are consumed by demand before they touch the grid. Every time inference gets cheaper, workloads that were previously uneconomical become viable: longer contexts, agentic chains, more users. Google's per-prompt energy collapsed; its total AI energy spend still rose. Jevons Paradox, in silicon.
What the buyers are telling us
The most persuasive evidence is the transaction record — what rational, deep-pocketed buyers have actually paid for power they could not obtain through conventional channels.
Anthropic pays SpaceX $1.25bn a month for access to the Colossus campus in Memphis, an all-in rate of roughly $5,708 per megawatt-hour and $7.78 per GPU-hour against a median neocloud rate of $3.33. Six weeks later, Google signed for $11.46 per GPU-hour — a 47 per cent premium to the earlier deal. Microsoft's 20-year purchase of Three Mile Island's entire restarted output runs an estimated 150 per cent above wholesale spot. And Google's $4.75bn acquisition of Intersect Power bought not a single operating asset but permits, interconnection positions and the team that holds them — roughly $1,500-1,670 per kilowatt of capacity that exists only on paper. The premium is for queue position.
Every major hyperscaler has now signed at least one nuclear deal, with more than 9.8GW committed. These are not diversification plays. They are what a binding constraint looks like when it cannot be bought off through ordinary procurement.
The handoff, in short, is nearly complete. Chips cleared because capital was deployed. Memory will clear in 2027-28 for the same reason. Energy will not — because on any investment-relevant horizon, there is no such thing as doubling the grid.