Software That Makes the Chip Cheap
NVIDIA put out a blog post at the end of June that I want to walk you through, because everyone is reading it wrong and I think the actual story underneath it is way more interesting than the headline.
The pitch on the surface is simple. NVIDIA’s software made AI tokens dramatically cheaper. Like, cut the cost by 5x in a single month on one model, cheaper. And a bunch of companies you’ve maybe heard of, Baseten, Together AI, Cognition, are all quoted saying yeah, this is working great for us.
If you read that as a consumer story, it sounds great. Cheaper tokens, cheaper AI products, more stuff gets built, everybody wins.
But if you read it as an infrastructure guy, which is what I do for a living, it’s actually a confession. NVIDIA is telling you that the thing everyone spent the last two years fighting over, who has the most GPUs, who’s chip constrained, who got allocation, isn’t the game anymore. The chip stopped being the bottleneck. The software running on top of the chip is the bottleneck now. And NVIDIA owns that too.
Let me explain why that distinction actually matters to your business, your job, or your portfolio.
Why “cost per token” is the tell
Here’s the line in NVIDIA’s post that I keep coming back to. They said companies are shifting away from evaluating chips by their peak specs and instead judging everything by cost per token, meaning how many useful tokens you can squeeze out per dollar, per watt, and within a certain speed requirement.
That’s not just a new marketing phrase. That’s an admission about where the money actually lives now.
When everyone cared about raw compute power, that was a hardware race. Theoretically anyone with enough cash could buy their way into competing. But cost per token isn’t a hardware number. It’s what happens when you take the GPUs, the networking, the memory, and a giant pile of orchestration software, and tune all of it together like an engine.
NVIDIA describes their own stack as three layers stacked on top of each other, one that handles the plumbing of serving requests at scale, one that squeezes performance out of the model itself, and one that just exposes the raw hardware so developers don’t have to think about it. That’s not a chip. That’s basically an operating system for manufacturing intelligence, and NVIDIA sits at every single layer of it.
These cost improvements are not one-time events. That’s the whole ballgame.
NVIDIA’s newest chips are apparently pushing out something like 2.7 times more tokens than they were six months ago, on the exact same hardware, which works out to cutting the cost of producing each token by more than 60 percent without buying anything new. Other benchmarks put their flagship system at around twelve cents per million tokens, something like 35 times cheaper than the previous generation, and roughly 50 times more efficient per watt. Stack a few of these software tricks together and NVIDIA claims you get up to 20 times more throughput out of hardware you already own.
None of that came from a new chip. It all came from software updates.
Think about what that actually means economically. A GPU is basically melting ice the second you buy it. It’s worth less every month as newer chips ship. But a software stack that keeps squeezing more tokens out of that same GPU is the opposite, it’s an asset that gets more valuable the longer you hold it, as long as you stay inside NVIDIA’s ecosystem. That’s not really a discount they’re giving you. That’s a really elegant way of making you dependent on staying put, because the second you leave, you lose access to gains that only exist inside their stack.
And here’s the twist that I think is genuinely clever. The fact that the models themselves, DeepSeek, Llama, all these open weight models, are free and open source isn’t a threat to NVIDIA at all. It’s actually helping them. Anyone can download the model. What they can’t download is the specific combination of NVIDIA’s networking fabric, their inference software, and years of tuning that determines whether that free model actually runs cheaply at scale. So the model gets commoditized, prices race to zero at that layer, while NVIDIA owns the only layer left where you can actually make money. It’s the exact same move cloud companies pulled with open source databases a decade ago. Give away the part that anyone can copy, keep the part that nobody can.
Who’s actually cashing in
Look at the list of partners NVIDIA name drops in their own post. Baseten builds on NVIDIA’s software. Cognition runs its infrastructure on NVIDIA’s framework instead of building their own. Deep Infra, same thing. Together AI, same thing.
None of these companies are competing with NVIDIA. They’re distribution channels.
Every one of them building their edge on top of NVIDIA’s tools is, whether they think about it this way or not, extending NVIDIA’s grip on the market while NVIDIA collects the profit from the one layer none of them can replicate on their own.
Fireworks helped a company called Sentient get 25 to 50 percent better efficiency.
Together AI helped a voice AI company cut their cost per conversation by about 6x. Those are real wins for those businesses. They’re also proof that the fastest way to get competitive costs is to go deeper into NVIDIA’s stack, not away from it.
We’ve seen this movie before. It happened with cloud computing, where “don’t get locked in” became the advice everyone gave and nobody followed. It’s happening again here, except this time the lock-in shows up directly in your gross margin, which makes it even harder to walk away from.
What this actually means for you
If you’re building something on top of all this, the short term read is genuinely good news. If token costs keep falling anywhere near this rate, and there’s research suggesting infrastructure efficiency is compounding by something like 10x a year, then stuff that didn’t make financial sense eighteen months ago might pencil out now.
That should absolutely change what you build.
But here’s the catch. If cheap tokens become available to literally everyone building on the same stack, then cheap tokens stop being your competitive advantage almost as fast as you get access to them. The moat moves somewhere else, up into whoever owns the distribution, the data, or the workflow lock-in, because the inference layer underneath it is turning into a commodity, ironically because NVIDIA is the one commoditizing it, in service of protecting the layer they still own.
If you’re a smaller company, this hits you harder. You can rent access to the same hardware everyone else uses and get the baseline numbers NVIDIA publishes. But the real gains, the kind Baseten and Together AI are bragging about, come from custom tuning on top of the base stack, and that takes engineering talent most small teams don’t have and can’t afford to hire. So you end up with two tiers, well funded companies who capture the full 20x gain, and everyone else paying closer to retail for a fraction of it.
That gap is basically why companies like Baseten exist as separate businesses instead of just being a checkbox inside NVIDIA’s own product. NVIDIA gets paid either way.
On the labor side, I don’t think this is really about jobs disappearing overnight. It’s about which tasks suddenly become cheap enough to automate without a second thought. When a reasoning task that used to cost real money now costs a fraction of a cent, the decision to replace a human step in a workflow with an AI agent stops being a big strategic bet and starts being a rounding error. That’s a much faster trigger for change inside a company than any leap in how smart the model actually is.
One thing that bugs me a little. A lot of the benchmarks getting cited here, InferenceMAX, SemiAnalysis, come from a research firm that has commercial relationships across the AI infrastructure world, and NVIDIA is the one funding and amplifying most of their headline numbers.
That doesn’t mean the numbers are fake. NVIDIA also posted record results on MLPerf, which actually is an independent industry benchmark, so the broader trend checks out. But when the company that benefits the most from a narrative is also the one paying for and distributing the evidence behind that narrative, you shouldn’t throw it out, you should just discount it a little.
Cost per token is becoming the metric everyone in the industry cares about at exactly the moment the company with the most to gain from that metric is also the one deciding how it gets measured.
That’s the kind of thing regulators eventually notice, usually years after the pricing power has already locked in, never before.
So what do you actually do with this
If you’re building an AI product, treat falling token costs as a nice tailwind, not a moat. Put the savings into whatever actually makes you hard to copy, your data, your workflow, your distribution, because cheap inference is available to your competitor too.
If you’re running infrastructure at scale, actually benchmark your own workloads instead of trusting the vendor slides, and start treating the ability to switch stacks as a real line item in your budget, because the lock-in feels soft today and gets a lot more expensive to escape in two years.
If you’re investing in this space, stop asking who has the best model and start asking who controls the layer that decides whether any model is affordable to run at scale. Be skeptical of any inference startup whose entire pitch is “we optimized on top of NVIDIA’s stack,” because whatever edge that gives them, it’s compounding for NVIDIA at least as much as it’s compounding for them.
And if you’re thinking about this from a national policy angle, the chip export controls everyone reached for first are already fighting the last war. If the real chokepoint moved from silicon to the software that makes silicon usable at scale, then staying competitive depends on software and systems talent just as much as GPU access, and that’s a much harder thing to control at a border.
The token got cheaper. Who controls what makes it cheap didn’t move an inch.

