I used to judge chip wars by teraflops. Then Google came up with a chip that loses almost every benchmark against NVIDIA. But it might still win the decade anyway.
That made me change the way I used to evaluate AI inference hardware.
So, here’s a shorter version. NVIDIA’s new Vera Rubin platform is faster than anything Google has built, chip for chip, no argument there. Google didn’t try to out-spec it. It split its flagship line into two separate chips, TPU 8t for training and TPU 8i for inference, and started competing on cost per query instead of peak throughput.
Key Takeaways
- NVIDIA’s Rubin R100 wins on every raw spec; nearly 5 times the FP4 throughput and almost triple the memory bandwidth of TPU 8i.
- Google isn’t chasing NVIDIA’s speed. It split its chip line into TPU 8t for training and TPU 8i for inference, betting on cost per token instead of peak FLOPS.
- AI inference now costs the industry more than training does, the first year that’s ever been true, which is why this fight is happening at all.
- Anthropic’s roughly 1 million TPU, 3.5 gigawatt commitment shows the real contest is about locking in power capacity for 2027, not winning a single benchmark.
- Most teams should still default to NVIDIA for day-one compatibility. TPU 8i only pays off at high, steady AI inference volume, and the JAX rewrite cost is real.
What’s Changed in AI Inference Economics This Year?
You can train once. But AI inference happens every single time you hit send. A model that costs tens of millions to train can add up to several times that in serving costs over its life, one API call at a time. That’s why inference is a much bigger factor for AI profitability in 2026.

Gartner just confirmed it with a number. 55 cents of every cloud AI dollar now goes to inference. And this is the first time that AI inference cost has beaten training. But most of the 2026 chip race coverage still leads with specs. The real story is what those specs cost to run at scale, and that’s the question reshaping AI computing budgets across the industry in 2026.
The training-to-inference cost flip
Training cost is just a one-time expense. But AI inference is like a non-cancellable subscription. Every chatbot reply, every coding suggestion, every agent double-checking its own work runs through inference.
A frontier model doesn’t make or lose money on the day it finishes training. It makes or loses money on every query, for years afterward, which is exactly why the hardware underneath it suddenly matters this much.
One check lets you see through most chip announcements. Does the number describe a lab benchmark, or what happens on the millionth query?
The first question is much easier to answer. But almost none answer the second. And that’s exactly the gap.
TPU vs GPU: Two Very Different Bets on Chip Design
NVIDIA and Google just made opposite bets on what AI inference hardware should look like. And neither are wrong, yet. NVIDIA is trying to build one extremely capable, general-purpose accelerator, keep it compatible with the CUDA software the whole industry already runs, and let scale do the rest.
In contrast, Google’s trying to stop trying to be everything to everyone, and split the chip in two instead. Calling this simply TPU vs GPU undersells it. The real split is specialization against ecosystem gravity, and both companies are aware of exactly what they’re betting on.

NVIDIA Vera Rubin: The ecosystem behind the chip
Here’s a very interesting detail. Vera Rubin is a platform, not a single chip. It pairs NVIDIA’s Vera CPU with the Rubin GPU. The inference workhorse inside it, Rubin R100, puts up 50 PFLOPS of NVFP4 compute and 22 TB/s of memory bandwidth, nearly 3X what Google’s inference chip can move.
Google TPU 8: Why one chip became two
Google’s answer wasn’t to catch up on raw speed. At Cloud Next 2026, it split its flagship TPU line into two dedicated chips. TPU 8t for training and TPU 8i built specifically for AI inference, packing 288GB of HBM and 384MB of on-chip SRAM, 3X the prior generation. That’s basically Google admitting that training and inference are different jobs and shouldn’t share a chip anymore.
TPU 8 vs. NVIDIA Vera Rubin: The Spec and Cost Comparison
Put the numbers side by side, and the gap is bigger than most headlines admit.
| Metric | Rubin R100 | TPU 8i |
| FP4 compute | 50 PFLOPS | 10.1 PFLOPS |
| Memory bandwidth | 22 TB/s | 8.6 TB/s |
| HBM capacity | 288GB | 288GB |
| On-chip SRAM | Standard | 384MB |
| Cost per token | ~$1.03/million (B200, on-demand) | Not yet public |
| MLPerf results | Published | None yet |
| Framework | CUDA, vLLM, TensorRT-LLM | JAX/MaxText only |
Per accelerator, this isn’t close. NVIDIA wins on nearly every row that measures raw hardware.
But the 384MB of on-chip SRAM on TPU 8i is triple the prior generation. That’s not a headline spec, but it’s Google’s actual answer to lower bandwidth. More on-chip memory means fewer trips out to HBM per AI inference step, which is exactly the kind of architectural choice a spec sheet buries and a cost-per-token number would expose.
The numbers where NVIDIA actually wins
Before making Google’s case, the numbers deserve honesty. Rubin R100 has close to 5X the FP4 throughput of TPU 8i and almost 3X the bandwidth. It has published MLPerf results; TPU 8 has none.
The New Economics of AI Inference: Cost Per Token Over Peak FLOPS
NVIDIA’s roadmap is built around winning every benchmark chart it enters. Google’s roadmap is built around making that chart optional for the buyer. Surprisingly, both goals can succeed at the same time. And that’s kind of a trap for anyone who still shops for AI infrastructure off a spec sheet in 2026.
Google claims TPU 8i delivers 80% better AI inference performance per dollar than its previous Ironwood chip. I’d suggest you consider it as a vendor benchmark.
Why cost per token beats a benchmark chart.
A chip that wins every FLOPS chart can still lose on the invoice. Utilization, batch size, and idle capacity eat into peak performance long before it reaches your bill. For NVIDIA, they’ve got some real, independently verifiable numbers. That’s roughly $1.03 per million tokens on-demand on B200. TPU 8i doesn’t have that yet.
AI Infrastructure Is Now a GW Race, Not a Chip Race
Anthropic went well past a pilot. It committed to roughly 1 million TPUs covering 3.5 GWs of 2027 capacity. That volume of commitment allows you to secure years of power capacity in advance. Google is locking in customers even before shipping the hardware at broader scale, and framing it as feeding an AI Hypercomputer stack built for agentic workloads at scale.
Power is the real scarce resource behind this entire story. Whoever controls gigawatts in 2027 controls pricing power. NVIDIA knows this too, which is part of why investor chatter around NVDA has started pricing in Google’s cloud-scale threat alongside its chip specs.
What This Means If You’re Actually Buying AI Computing Power
Here are some points:
- Pick NVIDIA Vera Rubin if you need day-one CUDA compatibility, multi-cloud flexibility, or spot pricing for bursty workloads.
- Consider TPU 8i only if you run high-volume, steady-state inference, are willing to rewrite in JAX or MaxText, and can live with single-vendor lock-in.
- Budget for the switch cost honestly. FlashAttention 3 is CUDA-only. There is no drop-in vLLM path to TPU today.
- Re-run this math every two quarters. AI computing pricing is moving fast enough that a decision from January can be stale by August.
The unglamorous truth is that most teams should default to NVIDIA unless their AI inference volume is large and predictable enough to justify the migration pain. Google only pays off at scale, and scale is the one thing a spec sheet can’t tell you whether you actually have.
Final Thoughts
Which chip wins today matters less than whether the best chip survives as a useful category once specialization and captive cloud capacity decide the market. NVIDIA still builds the faster part, and that’s not close.
But Google is betting that speed stops mattering once the buyer is locked into a power contract three years out. That bet looks right, even if TPU 8i isn’t ready to prove it.
For more info on AI and tech, visit Yaabot.
FAQs
No, not per chip. However, Rubin R100 performs significantly better than the others on raw FP4 throughput and memory bandwidth. The goal of the competition between TPU 8i and others is not speed, but cost per token and cloud economics.
The GPUs are general-purpose and CUDA-supported, allowing you to use existing tools as is. TPUs are specialized for Google Cloud and JAX, and offer a compromise between flexibility and possible cost efficiency when working at scale.
NVIDIA’s B200 runs about $1.03 per million tokens on-demand. TPU 8i has no public pricing yet, so any figure you see for it is an estimate, not a verified number.
It’s a platform, not one chip. Vera is the CPU, Rubin is the GPU, and Rubin R100 is the inference-focused die inside it, rated at 50 PFLOPS NVFP4.
While this is the case if you need to run inference in high volume, and if you can get comfortable with a JAX rewrite, CUDA portability is a much better option for most teams today.
Because AI inference now eats more cloud budget than training does. That’s where the recurring revenue and margin actually live, so that’s where NVIDIA and Google are now fighting hardest.

