The 2026 AI chip race looks like one contest on a spec sheet, but Nvidia's Rubin, Google's Ironwood, and Groq's LPU are actually built for three different jobs — and comparing their PFLOPS numbers head-to-head misses why each one wins where it does.
Every few months a new comparison chart makes the rounds: Nvidia's newest GPU against Google's newest TPU against AMD's newest accelerator, ranked by peak FLOPS like it's a drag race. The chart is usually accurate and usually misleading, because it implies these chips are fungible — that you'd pick whichever one wins the spec sheet and run your workload on it. In practice, the 2026 AI hardware landscape has split into three genuinely different jobs, and the chips built for each one make different tradeoffs that a single "PFLOPS per dollar" number can't capture.
The three jobs are: large-scale pretraining (where interconnect and memory capacity matter more than any single chip's peak throughput), general-purpose training-and-inference (Nvidia's territory, and increasingly AMD's), and high-throughput, low-latency inference at fixed model sizes (where Groq and Cerebras have built chips that would be terrible at training but are very good at exactly one thing). Comparing a Groq LPU to a Trainium3 pod is like comparing a drag racer to a delivery van — both have engines, but the design briefs don't overlap.
| Chip | Maker | Job | Memory | Bandwidth | Compute (per chip) | Power |
|---|---|---|---|---|---|---|
| Rubin GPU (NVL72) | Nvidia | Training + inference | 288 GB HBM4 | 22 TB/s | ~50 PFLOPS FP4 / ~25 PFLOPS FP8 | ~2,300 W |
| Instinct MI450 | AMD | Training + inference | 432 GB HBM4 | 19.6 TB/s | ~40 PFLOPS FP4 / ~20 PFLOPS FP8 | ~1,000–1,400 W |
| TPU Ironwood (v7) | Hyperscale pretraining | 192 GB HBM3E | 7.37 TB/s | ~4.6 PFLOPS FP8 | ~300–600 W | |
| Trainium3 | AWS | Hyperscale training | 144 GB HBM3E | 4.9 TB/s | 2.52 PFLOPS FP8 | — |
| Groq LPU | Groq | Fixed-model inference | 230 MB SRAM (on-chip) | — | 8-bit native | ~375 W |
| Cerebras WSE-3 | Cerebras | Fixed-model inference | 44 GB SRAM (on-die) | 21 PB/s (on-chip fabric) | 125 PFLOPS (system) | ~27 kW (system) |
A few things jump out before you even get to the prose. Nvidia's Rubin has more than four times the memory bandwidth of Google's Ironwood — but Ironwood isn't trying to feed one giant model through one chip as fast as possible, it's trying to keep 9,216 chips in a pod fed cheaply enough that a hyperscaler can afford to run it around the clock. And Groq and Cerebras don't report HBM bandwidth at all, because they don't have HBM — that's the whole point of the design, and it's worth unpacking why.
Every GPU-shaped chip — Rubin, MI450, even the TPU and Trainium ASICs — uses the same basic memory architecture: a relatively small pool of fast on-chip cache backed by a much larger pool of HBM sitting a few millimeters away on the same package. That HBM is what limits inference throughput in practice. Generating each token requires streaming the entire set of model weights (and the growing KV cache) through the compute units, and HBM bandwidth — not FLOPS — is usually the bottleneck. This is the "memory wall" that shapes on-device and edge inference too, just at datacenter scale instead of a phone's power budget.
Groq's answer was to delete the wall: the LPU holds the entire model in on-chip SRAM, which runs at roughly an order of magnitude higher bandwidth than HBM, with no round trip off-die at all. Cerebras went further and made the "chip" an entire silicon wafer — the WSE-3 is roughly 46,000 mm² of continuous silicon with 44 GB of SRAM and 900,000 cores wired together by an on-chip fabric running at 21 petabytes per second. Both designs trade capacity for speed: a Groq LPU can't hold a 400B-parameter model on one chip, so you shard across a rack (Groq's LPX racks scale to 256 chips), and a Cerebras wafer's 44 GB ceiling means you're choosing between a smaller model at full wafer speed or model-parallel sharding across wafers.
The payoff shows up in tokens per second, not FLOPS. On Llama 3.3 70B, Groq serves around 750 tokens/second and Cerebras around 2,100 tokens/second — both several times faster than a typical GPU-served endpoint at the same model size, because neither one is HBM-bound. The cost side splits differently: Groq is currently cheaper per token on most published benchmarks (roughly $0.15/$0.60 per million input/output tokens on GPT-OSS-120B, versus Cerebras around $0.35/$0.75), while Cerebras wins on raw throughput. Neither number tells you anything about how either chip would perform on a training run, because neither one is built to do backward passes at scale — Cerebras's own CS-4 roadmap still frames the wafer-scale line primarily as an inference and fine-tuning product, not a pretraining replacement for GPU clusters.
Google, Amazon, and (via Maia and MTIA) Microsoft and Meta didn't build custom silicon because Nvidia's chips are bad — Ironwood and Trainium3 are both objectively slower than Rubin on nearly every per-chip spec in the table above. They built them because at hyperscaler volume, the question isn't "which chip is fastest," it's "what does a token cost when you're running ten million of them a day, and can you get enough chips to matter." Nvidia's GPUs sell at GPU margins; TPUs and Trainium chips are built at cost by companies that also run the datacenters, so the comparison that matters internally is dollars-per-training-FLOP and dollars-per-inference-token at their own workloads, not peak throughput against Rubin.
That's also why Ironwood's per-chip numbers look modest next to Rubin's — 4.6 PFLOPS of FP8 versus Rubin's ~25 — but the pod-level number is the one Google actually cares about: 9,216 chips networked together for roughly 42.5 exaflops. The architecture optimizes for scale-out efficiency (SparseCores, a custom interconnect, tight integration with JAX and the XLA compiler) rather than any single chip winning a benchmark. The tradeoff is lock-in: Ironwood only runs well inside Google's own software stack, the same way Trainium is tied to AWS's Neuron SDK. None of the standard CUDA-based inference tooling most open-source model servers assume — vLLM, SGLang, TensorRT-LLM — runs natively on either, which is a real cost even when the per-token economics look good on paper.
Nvidia's counter to this dynamic is the CUDA moat itself: every framework, every kernel optimization, every inference server assumes a GPU-shaped chip first and back-fills TPU or Trainium support later, if at all. That software gravity is arguably a bigger competitive advantage than any single generation's FLOPS lead, and it's the reason AMD's MI450 — which on paper offers more memory capacity and better power efficiency than Rubin — still has to fight an uphill software battle, not a hardware one, to take share.
Power and cooling are the two variables that never make it into a marketing slide but end up determining what you can actually deploy. Rubin's ~2,300 W per chip and Cerebras's 27 kW full-system draw both assume liquid cooling and rack-level power delivery most colocation facilities weren't built for five years ago; that's part of why hyperscalers are retrofitting or building new datacenters specifically around Blackwell- and Rubin-class thermal envelopes rather than dropping new GPUs into old rooms. Ironwood's 300–600 W per chip looks unglamorous next to Rubin, but it's a deliberate design choice — Google has said openly that performance-per-watt, not peak performance, is the metric it optimizes pod design around, because at pod scale the power bill dominates the total cost of ownership.
The other thing a spec sheet hides is precision. Most of the eye-catching PFLOPS numbers — Rubin's 50, MI450's 40 — are FP4 numbers, a data type that barely existed in production inference two years ago and still requires careful quantization to avoid accuracy loss on smaller or more sensitive models. The FP8 numbers, roughly half those figures, are the more honest comparison point for workloads that haven't been tuned for 4-bit inference, and even FP8 accuracy loss is workload-dependent enough that "PFLOPS" alone tells you little about what you'll actually get in production without benchmarking your specific model and quantization scheme.
If you're picking hardware rather than just reading about it, the question to ask isn't "which chip wins the table" — it's which of the three jobs your workload actually is. Pretraining a frontier-scale model from scratch is a hyperscaler-cluster problem where TPU pods and Trainium UltraClusters are cost-competitive with GPU clusters if you're already inside that cloud's ecosystem. Fine-tuning and general-purpose serving across many model sizes is still Nvidia and AMD's game, because the software stack assumes GPU shapes and switching costs are real. And serving a fixed, well-known model at maximum tokens-per-second and minimum latency — the coding-agent and voice-agent use case where every millisecond of time-to-first-token matters — is exactly the niche Groq and Cerebras built their chips for, and exactly where a general-purpose GPU, sized for flexibility rather than throughput on one model, will lose on both speed and cost per token.