AI Insiders reported yesterday that Nvidia is passing rising memory costs on to customers, pushing AI server prices up more than 15 percent. Venture investor Tomasz Tunguz argues that increase is the latest visible mark of a chain reaction that started with GPUs in 2023 and is still moving through the hardware supply chain, in a post on his blog, tomtunguz.com.
Tunguz frames the pattern with the Bullwhip Effect, a supply chain idea where a modest shift in demand gets amplified into much larger swings further upstream, because each link reacts late and overcorrects. Applied to AI hardware, a shortage in one component does not just raise its own price. Buyers switch to the next available part, and that switch shows up as a fresh shortage, months or years later, one step further back in the chain.
The first shock hit GPUs. Tunguz cites data showing on-demand Nvidia H100 rental rates pushed past $9 an hour as ChatGPT’s launch concentrated capital on chip procurement in early 2023. Conventional server shipments fell 22 percent that year, below 2018 levels, as buyers delayed refresh cycles to fund GPUs instead, a figure he draws from Omdia’s server tracker.
That drop in server demand hit memory makers already absorbing a post-pandemic glut, one Tunguz says cost the industry more than $20 billion and forced cuts to wafer output as steep as 40 percent. About eighteen months later, manufacturers began converting cleanroom capacity toward High Bandwidth Memory, the chip type built for AI accelerators. HBM needs close to three times the silicon per gigabyte that ordinary DDR5 does, which is Tunguz’s own explanation for why the shift squeezed conventional supply. He cites enterprise SSD contracts jumping 80 percent within one quarter and Micron reporting DRAM prices climbing near 60 percent quarter to quarter.
By late 2025, agentic software shifted the imbalance again. Training clusters have historically run roughly one CPU per eight GPUs, but Tunguz notes that autonomous agents, which spend cycles compiling code and calling tools, push that ratio toward one to one. Intel disclosed server CPU prices climbing 27 percent from a year earlier even as unit volumes fell. By 2026 the shortage reached storage: cloud architects, priced out of fast flash near $150 per terabyte, shifted toward slower hard drives for bulk training data, and Western Digital and Seagate have both said their nearline output for 2026 is fully booked.
The compounding cost lands hardest on the data centre itself. Tunguz puts new facilities at more than $20 billion per gigawatt, with electrical gear alone eating half the budget, and construction costs roughly tripling to $1,033 per square foot, land not included. Generator step-up transformers now carry roughly three-year lead times, and he reports that GE Vernova and Siemens Energy have already sold their turbine output through 2029, with orders booked into 2031. A Siemens executive described the resulting bind to Bloomberg: build too little and lose share, build too much and get stuck with fixed costs nobody wants to carry.
This is one analyst’s framing, not a company disclosure. The rental rates, shipment declines and price jumps come from earnings reports and industry trackers Tunguz cites directly. The three-to-one wafer substitution math and the Bullwhip label itself are his interpretation of why the shortage keeps reappearing one link further back.
A bullwhip amplifies the shock going up the chain, but it implies the correction overshoots on the way down too. Tunguz points to more than $2 billion in new transformer capacity, next-generation NAND fabs and additional turbine lines all scheduled to arrive in 2027 and 2028. Anyone signing a multi-year compute or data centre contract at today’s bottleneck prices should treat that wave as a live risk, not a footnote: locking in current rates assumes today’s scarcity holds, and the same framework that explains the price spike says scarcity in this chain never holds indefinitely.
Tomasz Tunguz, “The AI Bullwhip,” tomtunguz.com, published August 2026.