At a Glance
- Nvidia told Microsoft, Google, and Oracle that AI server prices are going up 15%+ in 2027, while model inference prices keep dropping. Most enterprises are watching only one side of this.
- OpenAI’s Jalapeno chip outperformed Nvidia’s Blackwell chip in inference efficiency. Big labs are now designing their chips. This shows where Nvidia’s pricing is headed.
- If you are not tracking AI cost at the workload level, you are probably spending more than you think.
Last week we talked about how inference is getting cheaper. 80% price cuts, open-weight models flooding the market, everyone undercutting everyone. Great news, right?
This week, Nvidia quietly told a different story.
They informed Microsoft, Google, and Oracle that Vera Rubin and Grace Blackwell server prices are going up 15%+ starting early 2027. Why? Memory. DDR4 jumped 50% this quarter because everyone is chasing HBM (High Bandwidth Memory) for AI chips, and that cost has to go somewhere.
So tokens are cheaper. But the machines serving those tokens are getting more expensive. Both things are happening at the same time. And nobody in most enterprises is looking at both numbers together.
Meanwhile, OpenAI showed up at Hot Chips with benchmark results for Jalapeno, their custom inference chip built with Broadcom. 1.5 to 1.9x more efficient than Blackwell. Up to 4x faster on the kind of workloads ChatGPT actually runs. Designed in 16 months. Shipping late this year.
Google has TPUs. Amazon has Trainium. Now OpenAI has Jalapeno. When your biggest customers start building their own chips to get away from your pricing, that is a signal.
But here is the thing: most enterprises are not Google or OpenAI. They cannot design custom silicon. They pay through cloud providers, who will absorb these hikes and pass them along in ways that don’t show up as a line item labelled “Nvidia tax.”
Atgeir’s take:-
We keep seeing the same thing with clients. Teams look at the API price, see it dropping, generative ai service is getting cheaper. Nobody is adding up what the compute actually costs, how many new workloads got spun up last quarter because tokens felt cheap, or what the fully loaded cost per outcome looks like.
That is the gap. Not the model price. The total picture. For organizations investing in ai application development services, this becomes increasingly important as workloads scale.
If you are not tracking AI cost at the workload level, not the API level, you are probably already spending more than you realise. And with infrastructure costs headed up, that gap is only going to widen.