The three labs' cheap-tier models now sit within a few cents of each other per million tokens — so the real differences are in release cadence, cache economics, and how much of your "cheap" budget goes to invisible reasoning tokens.
Every frontier lab now ships a cheap tier, and for most production AI workloads — agent fan-out, batch classification, high-volume tool calling — the cheap tier is the workload, not a fallback for it. The flagship models get the keynote slide; the mini and Flash models get the API bill. As of October 2026, the three live options are Anthropic's Claude Haiku 4.5 (released October 15, 2025), OpenAI's GPT-5.4 Mini (released March 17, 2026), and Google's Gemini 3.8 Flash (released September 2, 2026). They've converged on strikingly similar per-token pricing, which makes the spec sheet almost useless as a decision tool. The differences that actually matter are structural, not numerical.
| Claude Haiku 4.5 | GPT-5.4 Mini | Gemini 3.8 Flash | |
|---|---|---|---|
| Released | Oct 15, 2025 | Mar 17, 2026 | Sep 2, 2026 |
| Input price (per 1M tokens) | $1.00 | $0.75 | $0.75 |
| Output price (per 1M tokens) | $5.00 | $4.50 | $3.75 |
| Cached input price (per 1M tokens) | $0.10 (90% off) | $0.075 (90% off) | ~$0.05–0.08 (similar discount tier) |
| Context window | 200K tokens | ~272K–400K tokens (sources disagree; OpenAI lists 400K) | 1,048,576 tokens (1M) |
| Max output tokens | 64K | 128K | 64K |
| Reasoning/thinking tokens | Optional, billed as output | Default on for "Thinking" routing; billed as output | Optional, billed as output |
| Positioning | "Near-frontier intelligence" at low latency | "Most capable small model yet," ChatGPT free-tier default fallback | "Pro-grade reasoning, Flash-level latency and cost" |
Pricing is directional and pulled from public API pricing pages and third-party trackers (OpenRouter, llm-stats, cloudprice.net) current as of this writing; check each provider's pricing page before budgeting against it, since all three labs have moved these numbers more than once in the last year.
Look at the table for ten seconds and the takeaway seems to be "they're all about a dollar in, a few dollars out, pick whichever fits your stack." That's not wrong, exactly, but it's missing the three things that actually separate these models in production.
The dates in that table understate how differently each lab treats the mini tier. Anthropic shipped Haiku 4.5 once, in October 2025, and it has been the same model since — a fixed target you can benchmark once and trust for months. OpenAI's GPT-5.4 Mini is one point release in a faster-moving line (GPT-5, 5.1, 5.2, 5.4, and GPT-5.5 already following behind it by the time you read this). Google is the extreme case: it shipped three Flash point releases in six weeks over July–September 2026 (3.6, 3.7, then 3.8), each retrained on top of the last rather than built from scratch — Gemini 3.8 Flash's own model card defers its architecture, training data, and safety sections to the 3.7 Flash card.
This matters for anyone running evals. A benchmark you ran against "Gemini Flash" in July is testing a materially different model than the one serving your traffic in October, and the model ID in your logs may have quietly moved underneath you if you're pinned to an alias rather than a dated snapshot. Google's gains have been real — Gemini 3.8 Flash jumped from 81.6% to 90.8% on Terminal-Bench 2.1 over the 3.7→3.8 step alone, a meaningful jump for agentic coding tasks — but it also means Flash-tier evals have a shelf life measured in weeks, not quarters. If your CI pipeline includes an LLM-as-judge or a regression eval against a mini-tier model, pin the dated model ID, not the alias, or your baseline will drift without anyone changing a line of code.
Gemini 3.8 Flash's 1M-token window looks like a 5x advantage over Haiku 4.5's 200K. In practice, the usable fraction of a stated context window — the point past which needle-in-a-haystack recall degrades — has historically been well short of the advertised ceiling for every vendor, Google included. A 1M window is genuinely useful for dumping a whole mid-sized repo or a long document corpus into context once, which none of Haiku 4.5's 200K or GPT-5.4 Mini's ~272–400K can match. But if your actual working set is 50K tokens of conversation history plus tool schemas, the headline number is marketing, not a constraint you're bumping into. Size the context window to your real payload, not to the biggest number on the page — and if you do need the long end, test recall at the depth you'll actually use, not just at the top of the range the vendor benchmarks.
The output price column hides a multiplier that doesn't show up anywhere else in the table. All three mini-tier models now support an internal "thinking" or extended-reasoning mode, and on OpenAI's side, GPT-5.4 Mini defaults into "Thinking" routing in several ChatGPT contexts rather than being opt-in. Reasoning tokens are billed as output tokens even though you never see them in the response — they're the model's scratch space, charged at the same $4.50/$3.75/$5.00-per-million rate as the answer itself.
This is the single biggest variable cost in a "cheap model" budget, and it's invisible until the invoice arrives. A classification task that should cost a few hundred output tokens can balloon to several thousand if the model decides to reason through the answer by default. If you're running one of these models for high-volume, low-complexity work — tagging, extraction, simple routing — check whether reasoning effort is configurable and turn it down or off. You bought the mini tier for cost control; a model quietly spending 3,000 tokens thinking about whether an email is spam defeats the point.
The cached-input discount is the line in the table that matters most for agentic workloads and barely matters for one-shot chat. All three vendors now discount cached input roughly 90% off the base input rate — Anthropic charges $0.10 per million cached tokens against a $1.00 base, OpenAI charges $0.075 against $0.75. That discount compounds hard in any agent loop where a large, mostly-static payload — a system prompt, a set of tool definitions, a long document — gets resent on every turn.
This is exactly the shape of a tool-calling agent talking to an MCP server. An agent that calls a server exposing a few dozen tool schemas — something like Utilix's own MCP server, which surfaces its developer-utility tools as callable functions — resends that schema block on effectively every turn unless the client caches it. At $1.00/M uncached versus $0.10/M cached, a long agent session where the tool definitions dominate the prompt can see the bulk of its input cost evaporate purely from cache hits, with no change to model choice at all. If you're comparing mini-tier models for an agent workload, model the cache-hit rate your actual tool-calling pattern will produce before trusting the sticker price — a model with slightly higher base pricing but a session shape that caches well can land cheaper in practice than one with a lower list price and poor cache locality.
The practical move either way is the same: don't budget off the list price. Model your actual token mix — including hidden reasoning tokens and your real cache-hit rate for repeated tool schemas or system prompts — because at this tier, those two factors move your real cost more than the quarter-cent difference between any two vendors' input price.