Claude Haiku 5.5: Pricing, Benchmarks, vs Sonnet 5.5 & Luna
Claude Haiku 5.5 launched Oct 7 at $0.10/$0.50 per Mtok with 1M context. Full pricing, official benchmarks, head-to-heads vs Sonnet 5.5 and GPT-6 Luna.
If you run any kind of agentic workload — coding agents, subagents, classification pipelines, extraction jobs — the model you route cheap tasks to just got a lot more interesting. On October 7, 2026, Anthropic shipped Claude Haiku 5.5, and the headline is not subtle: the same $0.10/$0.50 per million tokens OpenAI charges for GPT-6 Luna, with Anthropic's benchmark numbers in front of it, a 1M-token context window, and — a first for a Haiku-class model — an adjustable effort dial.
This article is for whoever has to make the routing decision: the developer wondering whether the haiku alias in Claude Code should stay default, the platform team pricing out a high-volume pipeline, the Copilot admin eyeing the new model picker entry. Every number below was checked against Anthropic's and OpenAI's official pages on the day of publication, and at the end there is a short, honest section on the router-plus-DeepSeek route for people whose real constraint is the bill itself.
Quick answer
- Released October 7, 2026. Anthropic calls it "our fastest model to date" and "our fastest, cheapest, and most capable small model yet."
- API pricing (prompts up to 100K tokens): $0.10 input / $0.50 output per MTok, with cache reads at $0.01. Above 100K prompt tokens it jumps to $0.50/$2.50 — the long-context premium.
- Context window: 1M tokens, 128K max output. Standard on the model — but Haiku 5.5 is the only current Claude that pays a surcharge for long prompts; every other Claude 4.6+ model includes 1M at flat pricing.
- Model ID:
claude-haiku-5-5, identical on the Anthropic API, Amazon Bedrock, Google Cloud, and Microsoft Foundry. - vs Sonnet 5.5: one-twentieth the input and output price ($0.10/$0.50 vs $2/$10). Sonnet 5.5 is meaningfully stronger on every published benchmark; Haiku 5.5 wins on latency and on cost per token by 20×.
- vs GPT-6 Luna: dead-even price match at the headline rate, right down to the $0.01 cache read. Luna's long-context surcharge starts higher (272K vs 100K) and is far gentler ($0.20/$0.75 vs $0.50/$2.50).
- First Haiku with adjustable effort: Low / Medium / High / Xhigh / Max, defaulting to medium.
- Available everywhere already: Claude Code 2.1.293 made it the default Haiku, GitHub Copilot added it on launch day, and Cline's catalog refreshed to it on October 8.
What Claude Haiku 5.5 is — and when it shipped
Haiku is Anthropic's fast, cheap tier, and the official model docs describe 5.5 as built "for high-volume, latency-sensitive tasks such as classification, extraction, and routing." The launch post goes further — "our fastest model to date" — with a footnote conceding it "runs less quickly than our Opus models in Fast Mode." That footnote matters: Opus 5.5's Fast Mode is an $8/$40 product; this is not it.
The release date is October 7, 2026, completing the generation Anthropic started teasing in September: Opus 5.5 shipped September 22, Sonnet 5.5 followed September 28, and Haiku 5.5 landed nine days later. All three share the June 2026 knowledge cutoff and the same shape — 1M context, 128K max output, adaptive thinking.
Two quieter announcements rode along: Sonnet 5.5 cache reads were halved from $0.20 to $0.10 per MTok — Anthropic says around 20% off agentic costs for cache-heavy loops — and monthly API credits arrived for subscription holders: $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team.
The first Haiku-class effort setting is a bigger deal than it sounds. Haiku 4.5 had no dial; 5.5 exposes the same Low/Med/High/Xhigh/Max ladder as the bigger models, defaulting to medium. One deployment can run at maximum cheapness for routing and push effort up where it counts, without switching model IDs.
Claude Haiku 5.5 pricing: every rate on the card
Here is the full price card from Anthropic's official pricing page, verified the day this was published. The one structural thing to understand before reading it: Haiku 5.5 is priced by prompt length. A prompt of up to 100,000 tokens pays the headline rate; a prompt over 100,000 tokens pays roughly 5× on every line. Anthropic's launch post says around 90% of requests fit under the threshold.
| Rate (per MTok) | ≤100K prompt | >100K prompt |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache write (5m) | $0.125 | $0.625 |
| Cache write (1h) | $0.20 | $1.00 |
| Cache read | $0.01 | $0.05 |
For orientation, the neighbors on the same card:
| Model | Input | Output | Cache read |
|---|---|---|---|
| Haiku 5.5 | $0.10 | $0.50 | $0.01 |
| Haiku 4.5 | $1.00 | $5.00 | $0.10 |
| Sonnet 5.5 | $2.00 | $10.00 | $0.10 |
| Sonnet 5 | $2.00 | $10.00 | $0.20 |
| Opus 5.5 | $4.00 | $20.00 | $0.20 |
That makes Haiku 5.5 one-tenth the price of its own predecessor — Haiku 4.5's $1/$5 card was already the cheap option, and 5.5 undercuts it 10× on input while claiming to be more capable. Against the rest of the current lineup it is 20× cheaper than Sonnet 5.5 and 40× cheaper than Opus 5.5 on input tokens.
One honest caveat, the same one we raised in our Claude Pro default model piece: sticker price is not your bill. The tokenizer changed — Anthropic concedes the model "uses slightly more tokens per task" — so naive per-token comparisons flatter 5.5. And the long-context tier is not a rounding error for coding agents: whole-repo prompts cross 100K routinely, and everything above the line bills at 5×. Run your own numbers before assuming the $0.10 headline.
The 1M context window — and the 100K catch
The spec sheet reads like its siblings: 1M-token context window, 128K-token max output. The official model overview lists 1M as the standard window for the whole current lineup, and Claude Code's changelog entry for 2.1.293 confirms it in the same breath as the price: "Added Claude Haiku 5.5 (claude-haiku-5-5) … 1M context, $0.10/$0.50 per Mtok ($0.50/$2.50 for prompts over 100K)."
But the pricing page contains a sentence that matters: Claude 4.6-and-later models — except Haiku 5.5 — "include the full 1M token context window at standard pricing." In other words, Sonnet 5.5 and Opus 5.5 charge the same per token whether your prompt is 9K or 900K. Haiku 5.5 switches to its long-context rate card the moment your prompt passes 100K tokens.
The consequence is easy to state. Haiku 5.5 has the biggest window per dollar for prompts under 100K — nothing else at $0.10 comes with 1M context. Between 100K and the ceiling you are paying $0.50/$2.50: still cheaper than Sonnet 5.5's flat $2/$10, but no longer a 20× blowout. A 300K-token prompt costs $0.15 of input on Haiku 5.5 versus $0.60 on Sonnet 5.5 — a 4× gap. Compute your prompt-size distribution before picking a winner.
Claude Haiku 5.5 vs Sonnet 5.5
Everything here comes from the two model cards and Anthropic's own launch benchmark table, so the comparison is at least run on the same track.
| Dimension | Claude Haiku 5.5 | Claude Sonnet 5.5 |
|---|---|---|
| Input / output | $0.10 / $0.50 (≤100K) | $2.00 / $10.00, flat |
| Cache read | $0.01 | $0.10 |
| Long-prompt prices | 5× above 100K | none — flat across 1M |
| Context / output | 1M / 128K | 1M / 128K |
| Default effort | medium | high |
| Effort range | Low → Max | Low → Max |
| Official one-liner | "For high-volume, latency-sensitive tasks" | "The best combination of speed and intelligence" |
| GDPval-AA v2.1 | 1620 | 1840 |
| OSWorld 2.1 | 72.4% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 70.6% |
| FrontierCode 1.1 | 46.4% | 52.1% (at Xhigh effort) |
The benchmark gaps are real but uneven, and that unevenness is the whole routing argument. On OSWorld's computer-use tasks Haiku 5.5 gets within about 11 points of Sonnet 5.5 — for click-and-extract work, a gap you can afford. On Terminal-Bench, Sonnet 5.5 nearly doubles it (70.6% vs 39.2%), and terminal-heavy agentic work is exactly where a failed run costs more than the tokens you saved. FrontierCode lands in between; note Sonnet 5.5's score was recorded at Xhigh effort, neither Haiku's default nor its price point.
Our read, clearly labelled as ours: Haiku 5.5 has replaced Sonnet 5.5 as the rational default for short, well-scoped tasks — classification, extraction, routing, subagent summaries, quick edits, first-draft anything — because a 20× price gap buys a lot of retries. Sonnet 5.5 keeps every job where a wrong answer is expensive: multi-file edits, terminal-heavy agentic runs, anything where the model has to hold a plan and execute it. GitHub's own Copilot changelog framing is a useful third-party anchor: in early testing Haiku 5.5 "matched Claude Sonnet 5 on many coding tasks while using significantly fewer tokens" — note that is the previous Sonnet, not 5.5.
Claude Haiku 5.5 vs GPT-6 Luna
This is the comparison the launch was designed to pick, with one sentence of naming housekeeping first: OpenAI's catalog lists GPT-6 Luna (gpt-6-luna) — there is no "6.1" Luna. The 6.1 refresh covered Sol; Luna remains the GPT-6-era efficiency model, described by OpenAI as "our most efficient model for focused, high-volume tasks."
| Dimension | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Input / output | $0.10 / $0.50 (≤100K) | $0.10 / $0.50 (≤272K) |
| Cached input read | $0.01 | $0.01 |
| Cache write | $0.125 | $0.125 |
| Long-context rates | $0.50 / $2.50 above 100K | $0.20 / $0.75 above 272K |
| Context window | 1M | 1.05M |
| Max output | 128K | 128K |
| Knowledge cutoff | Jun 2026 | May 18, 2026 |
| Batch / Flex pricing | — | half price ($0.05 / $0.25 short-context) |
| GDPval-AA v2.1 | 1620 | 1437 |
| OSWorld 2.1 | 72.4% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 16.4% |
| FrontierCode 1.1 | 46.4% | 42.4% |
At the headline rate the two models are priced to the penny identically — input, output, cache read, even the 5-minute cache write. That parity is itself the news: Anthropic explicitly priced its small model at Luna's number, and the pitch is that you get more model for the same money.
Where they part ways is the long-context fine print, and it cuts in Luna's favor. Luna's surcharge does not start until 272K input tokens, and when it does, it is 2× input / 1.5× output — $0.20/$0.75 per OpenAI's pricing page. Haiku 5.5 starts surcharging at 100K and jumps 5×. Run the numbers on a 500K-token prompt: 500K × $0.20 = $0.10 of input cost on Luna, versus 500K × $0.50 = $0.25 on Haiku 5.5. For giant-context workloads, Luna is 2.5× cheaper on input.
On benchmarks, the table above is Anthropic's own numbers, not OpenAI's — treat it as the vendor's claim and benchmark on your own tasks. But if it holds, Haiku 5.5 leads on every row, most dramatically on OSWorld (72.4% vs 48.9%) and Terminal-Bench (39.2% vs 16.4%). Luna's counterarguments are structural rather than benchmark-shaped: Batch/Flex at half price for throughput workloads, a slightly larger window with a much gentler surcharge curve, and OpenAI's full tool surface (web search, file search, computer use) attached at the same sticker price.
Benchmarks: what Anthropic actually published
Anthropic's launch post leads with a benchmark table. The full set, all run by Anthropic against Haiku 4.5, GPT-6 Luna, and Sonnet 5.5:
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 157 | 86 | 141 | 336 |
| OSWorld 2.1 (offline) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | — | 56.9% |
| Humanity's Last Exam (tools) | 57.4% | 18.7% | — | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1% |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
The step-change versus Haiku 4.5 is the story Anthropic wants told, and the numbers support it: OSWorld goes from 15.7% to 72.4%, Terminal-Bench from literally zero to 39.2%, GDPval-AA more than doubles. Whatever you believe about vendor tables, Haiku 4.5 — still the cheap tier in a lot of production code — is not in the same class.
Three honesty notes. First, these are vendor-run numbers; OpenAI has not published Luna scores against this exact suite, so read the Luna columns as Anthropic's claim. Second, the customer quotes on the launch page — Box reporting scores "11 points higher than Haiku 4.5 at about half the latency," Asana seeing "over a 30% reduction in latency" and "up to 2.5x faster inference per agent turn" — are directional, not reproducible. Third, Anthropic published no tokens-per-second figure, so any "Haiku 5.5 does X tok/s" you see elsewhere is a measurement, not a spec; the official language is just "our fastest model to date."
Where you can use Haiku 5.5 today
Availability was unusually good on day one — the launch post says it is "available now on all platforms," and the tooling catalogs confirm it:
- Claude Code. Version 2.1.293 added
claude-haiku-5-5and made it "the default Haiku model on the Anthropic API." Thehaikualias now resolves to 5.5, so anyone who pinned"model": "haiku"inherited it automatically — the alias-pin behavior from our default model guide cuts both ways. - GitHub Copilot. The changelog entry dated October 7, 2026 lists it across VS Code, Visual Studio, JetBrains IDEs, Xcode, Eclipse, Copilot CLI, the cloud agent, the Copilot app, github.com, and GitHub Mobile, for Pro, Pro+, Max, Business, and Enterprise plans — billed at provider list pricing under usage-based billing, with admin control via the model policy.
- Cline. SDK v0.0.92 and the matching desktop (v0.0.45) and CLI (v3.0.70) releases on October 8 refreshed the model catalog to add Haiku 5.5 — and made it the default model on Google Vertex AI among several providers, displacing Sonnet 5.5. That is third-party tooling voting with its defaults.
If you are watching token spend in any of these tools, our guide on Claude Code usage limits covers the metering side.
Bedrock, Vertex, and Foundry access
The access story for platform teams is clean: the model overview lists the same ID — claude-haiku-5-5 — on the Anthropic API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, so a multi-cloud deployment needs no per-provider ID mapping, and the launch post names AWS, Google Cloud, and Microsoft Azure explicitly ("available now on all platforms").
One operational note: on Bedrock and friends, cross-region inference profiles and per-region availability can lag a launch by days. If a region 404s on the model ID, check the inference-profile list before assuming the model is not there yet.
When to pick Haiku 5.5 over Sonnet or Opus
The decision is simpler than the benchmark tables make it look, because the price gaps are large enough that moderate quality differences rarely change the answer:
- Pick Haiku 5.5 when the task is short, bounded, and high-volume: classification, extraction, routing, summarization, subagent side-quests, lint-style code edits, or any pipeline where you can afford to retry. With effort adjustable down to Low, it is also the cheapest way to put a frontier-family model in front of a trivial task.
- Pick Sonnet 5.5 when the task is a real coding job — multi-file changes, plan-and-execute agentic runs, anything touching Terminal-Bench-shaped work — and when your prompts routinely exceed 100K tokens, because its flat pricing across the full 1M window quietly beats Haiku's 5× surcharge on giant contexts.
- Pick Opus 5.5 for the long-horizon autonomous runs that justify $4/$20, ideally with Fast Mode if the bottleneck is wall-clock time rather than budget.
The honest cost alternative: router plus DeepSeek
One more option belongs in a pricing piece, stated plainly: if your real constraint is the monthly bill rather than the model tier, stop paying frontier prices for cheap tasks at all. Keep Claude Code as the harness and point it at a compatible endpoint — the $0.10 input rate has a competitor. DeepSeek's Anthropic-compatible endpoint meters per token at published rates, with deepseek-flash from $0.15 per million input tokens off-peak; the pricing table and two-variable wiring are in our Codex usage dashboard piece, and the DeepSeek page covers the models themselves.
What you give up is real: these are different models, the benchmark table above is not their benchmark table, and on agentic coding the 5.5 generation does things cheap endpoints may not. What you buy is per-token billing you control and the ability to send trivial tasks to something priced like a utility. Haiku 5.5 narrows that argument — at $0.10 it is closer to utility pricing than any Claude before it — but for pure high-volume throughput where the ceiling does not matter, the cheaper-API route still exists. And if you want to feel the difference without configuring anything, the DeepSeek App Builder runs those models in a browser, one prompt to a working app.
FAQ
When was Claude Haiku 5.5 released? October 7, 2026 — the third model of the 5.5 generation, after Opus 5.5 (September 22) and Sonnet 5.5 (September 28).
How much does Claude Haiku 5.5 cost? $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, with cache reads at $0.01 per MTok. Prompts over 100K tokens bill at $0.50/$2.50, with cache reads at $0.05.
Does Haiku 5.5 have a 1M context window? Yes — 1M tokens of context and 128K max output, same as Sonnet 5.5 and Opus 5.5. The difference is billing: the other models include 1M at standard pricing, while Haiku 5.5 switches to its long-context rate card above 100K prompt tokens.
Is Claude Haiku 5.5 cheaper than GPT-6 Luna? At the headline rate they are identical — $0.10/$0.50 per MTok and a $0.01 cache read on both. Luna is cheaper for very large prompts: its surcharge starts at 272K input tokens and only doubles the rate ($0.20/$0.75), versus Haiku's 5× jump above 100K. Whether Haiku 5.5 is worth it anyway depends on the benchmark deltas — which, per Anthropic's own table, favor the Claude.
Is Haiku 5.5 available on Amazon Bedrock?
Yes. The model ID claude-haiku-5-5 is the same on the Anthropic API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, and the launch post confirms all-platform availability.
What is the haiku alias in Claude Code now?
Haiku 5.5, since version 2.1.293. If you pinned the alias, you are already on it; if you pinned the dated ID claude-haiku-4-5, nothing moved until you change it.
Related reading
- Claude Pro default model: Opus 5.5 in Claude Code — the same launch generation, from the subscription side
- GPT-6.1 Sol pricing and context window — OpenAI's $2/$10 tier, the rung above Luna
- Claude Code usage limits — how the plan meters actually work
- Codex usage dashboard — per-token billing with DeepSeek pricing, spelled out
- DeepSeek App Builder — the cheap-API models, zero setup, in a browser
