The Chip Got Better. The Empire Got Bigger.
NVIDIA just shipped something called Vera Rubin. Ten times better performance per watt. One tenth the cost per token. 350 factory sites in 30 countries building the thing. CoreWeave ran it against DeepSeek-R1 and got ten times the tokens per megawatt compared to the old Blackwell racks.
On paper, that’s 800,000 tokens a second out of a rack that used to give you 80,000, same power draw.
Cool numbers.
But you know what’s important now.. the story is who owns the thing that was scarce before the chip ever mattered: the power to run it. And once you see that, the whole Vera Rubin launch reads differently.
Who gets to convert electricity into intelligence at the lowest cost, and what happens to everyone else.
Why “Performance Per Watt” Is the New Scoreboard
Here’s the thing that changed in the last two years.
For a long time, the bottleneck in AI was just getting your hands on chips. You wanted H100s, there was a line, NVIDIA controlled the line. That’s over.
TSMC can ramp chip production in twelve to eighteen months when demand shows up. What TSMC and NVIDIA cannot do is build you a new electrical substation or get you a grid connection, because that’s not a manufacturing problem, it’s a permitting and physics problem, and it runs on a completely different clock.
Interconnection queues in the biggest US markets, Northern Virginia, Phoenix, Dallas, are running four to seven years right now. There’s over two thousand gigawatts of generation and storage sitting in queues waiting to connect to grids that were never built for this.
So when NVIDIA says Vera Rubin gives you ten times the tokens per megawatt, they’re not selling you a faster chip. They’re selling you a way to get more intelligence out of the same electrical connection you already fought years to secure.
That’s the whole game now.
If you’re stuck with, say, 150 megawatts of power and no realistic path to more before 2030, the only lever left is squeezing more tokens out of every watt you’ve got.
Revenue basically becomes tokens-per-watt multiplied by however many gigawatts you actually control. Not GPUs you own. Watts you can plug into.
That reframes the question: it’s not “how much faster is this chip,” it’s “how does this change who wins in a world where the grid, not the wafer, is the ceiling.”
A cloud provider sitting on a fixed power allocation just found a way to serve ten times more customers or ten times more tokens without asking the utility for another megawatt.
That’s a direct, structural advantage over anyone still running Blackwell-era hardware on the same power budget. And because government-run energy governance, permitting boards, utility commissions, is what decides who gets new watts at all, the companies that already have power contracts locked in are about to become dramatically more valuable, because efficiency gains compound on top of an allocation nobody else can get.
That’s the answer to how energy governance evolves here: it doesn’t need new laws to matter more, it already controls who gets to play, and Vera Rubin just raised the stakes on every megawatt already spoken for.
The Empire Gets Bigger, Not More Crowded
Now follow the money on who actually benefits.
CoreWeave, Google Cloud, Microsoft Azure, and Oracle are the first four names getting Vera Rubin racks.
Not fifty companies. Four.
And NVIDIA’s own numbers say a rack that used to deliver one tenth the token throughput now costs one tenth as much per million tokens to run. If you’re one of those four, your cost structure for serving AI inference just fell off a cliff, and your competitors haven’t caught up yet, because Rubin allocation itself is scarce and follows existing NVIDIA relationships.
Smaller neo-clouds and marketplace providers are looking at 2027 before they see meaningful Rubin capacity.
That’s the mechanism by which this technology could concentrate market share and influence rather than spread it around.
The efficient frontier doesn’t get more crowded, it gets owned by the four or five players who already had the capital, the power contracts, and the NVIDIA relationship to get first access.
A smaller AI startup trying to compete on inference costs isn’t just fighting for market share, it’s fighting a cost curve that just moved ten times in the other direction for its biggest competitors, and it can’t buy its way onto that curve for another year or two even if it has the cash, because the hardware simply isn’t allocated to it yet.
This is where the smaller players actually have a move, though, and it’s not “wait for Rubin.” It’s picking your battles on the workload. The efficiency gains are concentrated in large mixture-of-experts models doing high-concurrency inference, that’s where the ten-times number comes from.
If you’re running smaller models, under maybe seventy billion parameters, on existing Blackwell or even Hopper hardware, you’re not compute-bound in the first place, so the Rubin gap barely touches you. The leverage for a startup is choosing not to compete where the giants just built a moat, and instead building on model sizes and use cases where raw rack efficiency isn’t the deciding factor yet. That’s a genuinely different competitive game than trying to out-infrastructure Google Cloud.
Five Fights Nobody’s Settled Yet
This is the part of the story that gets skipped in every breathless press write-up, because the people writing them have a stake in the “everything is fine” version. It isn’t fine, it’s contested, and here’s where the actual disagreement lives.
Fight one: does cheaper, more efficient infrastructure create monopolies or competition?
One camp says this consolidates power hard: whoever gets first access to Rubin racks locks in a cost advantage that compounds, and the four names getting first shipments become nearly impossible to displace on price.
The other camp points out that cheaper tokens usually mean more applications get built, because things that were too expensive to run suddenly pencil out, and historically cheaper infrastructure has expanded who can compete, not shrunk it, the same way cheaper cloud compute in the 2010s created a wave of new companies rather than just entrenching AWS.
Both things can be true at once, actually, concentration at the infrastructure layer and explosion at the application layer, which is exactly what happened with cloud computing.
Fight two: can the grid actually handle this, or is gigascale AI running into a wall?
Some people look at the water-saving cooling design and the ten-times efficiency and conclude the industry is solving its own energy problem just fast enough.
Others point to forty percent of AI data centers projected to be power-constrained by 2027 and multi-year interconnection queues in every major market and say efficiency gains at the chip level are getting eaten alive by the sheer scale of buildout, so the wall is coming regardless of how good the racks get.
This one isn’t really an opinion disagreement, it’s a math disagreement, whether efficiency gains grow faster than demand, and right now demand looks like it’s winning.
Fight three: does the “cheapest token wins” model eventually lock out smaller players?
One side says a race to the cheapest token per unit of compute becomes a scale game, and if you can’t afford the rack-scale hardware, you simply can’t compete on unit economics no matter how good your model is.
The other side argues token costs eventually become commoditized and cheap for everyone, the way storage and bandwidth did, and today’s frontier advantage is tomorrow’s baseline, so the barrier is temporary, not permanent.
History leans toward the second view eventually happening, but “eventually” has been known to take years that kill a startup’s runway first.
Fight four: is localized, distributed manufacturing across 350 sites a strength or a new fragility?
The optimistic read is that spreading production across thirty countries makes the supply chain more resilient to any single country’s disruption, unlike the old model where a shortage in one region stalled everything.
The skeptical read is that this many interdependent sites and seven co-designed chips built as a single system creates more points of failure, not fewer, because now you need every piece of a complex, tightly coupled chain to work in sync, and a disruption anywhere in thirty countries has more surface area to hit than a disruption in three.
Whether this reshapes trade dependencies in a good or bad direction really depends on whether any single node in that network becomes a chokepoint, which nobody knows yet because the network is brand new.
Fight five: does AI efficiency like this kill jobs or create new ones?
The standard worry is that if a rack can now do ten times the inference work per watt, you need proportionally fewer people managing that infrastructure, and automation absorbs roles that used to require humans watching dashboards and provisioning capacity.
The counter-argument is that cheaper, more available inference means way more AI gets deployed everywhere, and someone has to build, monitor, secure, and govern all of that new deployment, which is a bigger job market than the narrower one it replaces.
Both sides are usually right about different timeframes, contraction in the specific old roles happens fast, expansion in new AI-operations roles happens slower and unevenly, and the people who lose the first job aren’t always the ones who get the second one.
The Geography of Who Controls This
The 350-site, 30-country manufacturing footprint deserves its own beat.
Spreading fabrication and assembly across that many countries changes which governments have leverage over the AI buildout, and it changes what “supply chain risk” even means.
A single-country chokepoint used to be the nightmare scenario, think Taiwan and advanced chip fabrication. A thirty-country web is harder to disrupt on purpose, but it’s also harder to fully audit or secure, and it means more national governments now have a stake, and potential leverage, in how this technology gets built and deployed.
Watch which countries start attaching conditions, export requirements, local employment mandates, data sovereignty rules, to hosting a piece of that manufacturing chain. That’s where the second-order geopolitics of Vera Rubin will show up, not in NVIDIA’s press release, but in trade ministries over the next eighteen months.
There’s also a quieter effect worth naming: what does a ten-times jump in AI factory efficiency do to industries that aren’t AI at all.
Cheaper, denser compute per megawatt makes AI-driven automation viable in sectors that couldn’t previously afford the compute bill, logistics optimization, industrial quality control, scientific simulation.
That’s a slow-burn effect on traditional manufacturing sectors, not because robots show up on the floor tomorrow, but because the compute cost of running sophisticated optimization and simulation just dropped enough that projects which didn’t pencil out a year ago suddenly do.
What This Actually Means for You
If you’re a founder: don’t try to out-infrastructure the four companies who just got first access to this. Build on the model sizes and workloads where the efficiency gap doesn’t apply yet, and treat every year of “we’re not compute-bound” as a genuine head start, not a consolation prize.
If you’re an operator: the tokens-per-watt framework is now the real unit economics of your AI stack, not tokens-per-dollar. If your power allocation is fixed, that’s your actual ceiling on growth, and squeezing efficiency out of your existing infrastructure matters more than it did eighteen months ago.
If you’re an investor: the interesting bet isn’t NVIDIA, that’s priced in. It’s whoever controls power contracts and grid interconnection rights in constrained markets, because that scarcity doesn’t get solved by a better chip, and it’s the actual gate everyone else has to get through.
If you’re in government: the leverage you have isn’t the chip, it’s the grid connection and the manufacturing footprint sitting inside your borders. Whoever writes the rules for interconnection queues, energy allocation to data centers, and manufacturing conditions is deciding, right now, who gets to build the next generation of this stuff and who waits in line. That’s a bigger strategic lever than any AI policy paper you’re about to commission.
The chip got ten times better. The bottleneck just moved to the thing nobody can manufacture their way out of.

