Server manufacturers that build systems under contract for the largest cloud operators — Microsoft, Alphabet’s Google and Oracle among them — have reportedly told those customers that prices on Nvidia’s next-generation AI systems will rise more than fifteen percent on shipments beginning early in 2027. The increases are said to apply to Vera Rubin and Grace Blackwell configurations, with the exact figure varying by chip generation and by how much memory a given rack carries.
This is reported, not confirmed. Bloomberg first published the account on 22 August 2026, citing people familiar with the matter; every subsequent write-up traces back to that reporting. Nvidia did not respond to requests for comment and has published nothing itself, and the contract builders and their customers are unnamed. The direction and cause are consistent across coverage, but the specific numbers rest on a single tier of anonymous sourcing and should be read as such.
Why it matters
For anyone budgeting AI infrastructure, a mid-teens percentage increase on flagship systems is a planning number, not a rounding error. TrendForce estimates the rise could add at least five billion dollars to the cost of building a one-gigawatt data centre. That lands directly on the buyers signing multi-year capacity leases now for delivery in 2027 and beyond, and it feeds through, eventually, to the price of inference that end users pay.
It also complicates the story of falling model prices. 18BYTE has documented how the price of frontier inference collapsed at the API level even as the hardware underneath it grows more expensive. Cheaper tokens sitting on pricier iron is a squeeze that has to resolve somewhere — in margins, in utilisation, or eventually back in price.
The cost centre moved from compute to memory
The reported cause is not the accelerator itself but the memory around it. Prices and shortages for DRAM, LPDDR and high-bandwidth memory have climbed as AI infrastructure spending outran the memory industry’s ability to supply it, even with capacity ramping. According to TrendForce, GPUs once accounted for more than 80 percent of an AI server’s cost; in next-generation systems that share has fallen toward half as memory-related costs rise.
That shift is the real signal inside the price story, and it is worth stating precisely: this is a characterisation of where cost sits in the bill of materials, reported by an industry analyst, not an audited breakdown from Nvidia. But if it holds, it reframes the bottleneck. The constraint on scaling is drifting from the logic die to the memory stack, and pricing power drifts with it — toward the HBM and DRAM suppliers. That is a different competitive map from the one in which the GPU vendor set the price alone.
This continues a pattern Nvidia’s own numbers already showed
The reported hike sits on top of a supply chain already stretched. When Nvidia last reported, 18BYTE noted that its purchase commitments grew faster than its revenue — the company pre-committing to components ahead of booked demand. Memory-driven price increases are what that forward buying looks like when it meets a shortage.
For model labs, the response has been to spread bets across chip families and to build their own — the logic behind Anthropic’s move into custom silicon. A rising memory bill does not spare custom accelerators, which need the same HBM, but it does raise the value of designs that use memory more efficiently per token served.
What this means for buyers
If you are negotiating compute for 2027 delivery, price in a mid-teens increase on Nvidia-based systems as a working assumption, and ask suppliers to be explicit about the memory configuration behind any quote — since the reported increase varies most with how much DRAM and HBM a rack carries, the configuration is where the money is. Do not treat the figure as fixed: it is unconfirmed, and Nvidia has said nothing.
For buyers of inference rather than hardware, the read is subtler. A more expensive supply chain is a reason to expect the recent, steep token-price declines to slow rather than continue indefinitely, and a reason to lock favourable rates where you can. It also strengthens the case for matching each workload to the cheapest adequate model rather than defaulting to the frontier, which is the discipline the falling-then-firming price environment rewards.
What would change our reading
We would move this from reported to confirmed if Nvidia, a named OEM, or one of the cloud buyers stated the increase on the record, or if it appeared in a filing or an official price schedule. We would revise the magnitude if later reporting narrowed the figure below fifteen percent or showed it applying to a subset of configurations only. And we would reconsider the memory-bottleneck framing if HBM and DRAM supply caught up faster than expected — new capacity from the major memory makers coming online ahead of schedule would relieve exactly the pressure this story rests on, and could unwind the increase before it reaches customers.
Sources
- Bloomberg, business news — first report that Nvidia AI server prices will rise more than 15 percent from early 2027, citing people familiar with the matter — finance.yahoo.com
- TrendForce, semiconductor market research — analysis of the reported price hikes, memory cost drivers and the cost-share shift from GPU to memory — trendforce.com
- Tom’s Hardware, PC hardware news — report on Nvidia warning large customers of >15 percent AI server price hikes on rising memory costs — tomshardware.com