Anthropic · Model guide
Claude Haiku 5.5
The small model in Anthropic's Claude 5.5 family and the successor to Claude Haiku 4.5, built for high-volume, latency-sensitive work such as classification, extraction, routing, summarization and subagent tasks. Anthropic describes it as its fastest model at standard speed. It is the first Haiku with adaptive thinking and an effort setting (default medium), and it moves to a 1M-token context window with up to 128K output tokens. Manual thinking budgets, non-default sampling parameters and assistant prefill are rejected, and it uses the newer tokenizer, so the same text counts as roughly 30% more tokens than on Haiku 4.5.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorFor high-volume, latency-sensitive tasks such as classification, extraction, and routing.source
High-volume, latency-bound calls · Anthropic builds it for classification, routing, extraction, summaries and compaction, and calls it its fastest model at standard speed.(vendor claim)source
Subagent for coding agents · Anthropic pairs it with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work, and says the larger models remain the better choice for complex agentic coding itself.(vendor claim)source
Computer and browser use · Anthropic reports a 72.4% partial score on the OSWorld 2.1 offline subset against 15.7% for Claude Haiku 4.5, and highlights it for speed-sensitive browser use.(vendor claim)source
- anthropic/claude-sonnet-5-5stronger —Anthropic states Claude Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those in Terminal-Bench 4.0.(vendor claim)source
- anthropic/claude-haiku-4-5predecessor —Claude Haiku 4.5 still accepts manual thinking budgets, sampling parameters and assistant prefill, all of which return a 400 on this model.source
Pricing and billing
Tiered pricing: the rate changes once the input passes the threshold.
| Price / 1M tokens by input size | Input < 100.0K | Input ≥ 100.0K |
|---|---|---|
| Input | $ 0.10 | $ 0.50 |
| Output | $ 0.50 | $ 2.50 |
| Cache read | $ 0.010 | $ 0.050 |
| Cache write | $ 0.13 | $ 0.63 |
At this input size you are on the Input < 100.0K tier.
+ 2.0K × $ 0.50
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 128.0K tokens
- Input / output modalities
- Text + Image → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· can be turned off
- Reasoning effort
- low · medium · high · xhigh · max
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
- reasoning.toggle is unavailable when Reasoning effort=xhigh,max source
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- FrontierCode v1.1 (Main)score·Anthropic · 2026-10-07effort maxClaude Codesetup not fully disclosed
Run and reported by Cognition in Claude Code; mean of five runs per task. Max is this model's best effort level here; at xhigh Anthropic reports 45.8%. Same table: Claude Sonnet 5.5 46.2% at max (52.1% at xhigh), GPT-6 Luna 42.4%.
46.4% - Terminal-Bench 4.0pass_rate·Anthropic · 2026-10-07effort maxClaude Codesetup not fully disclosed
Claude Code in --bare mode, 10 trials per task (660 trials), no internet egress; standard error ±1.9 pts. Run with safeguards on and no fallback model: 1.8% of trials were stopped by a flagged request and counted as failures. Same table: Claude Sonnet 5.5 70.6%, Claude Haiku 4.5 0.0%, GPT-6 Luna 16.4%.
39.2%
- OSWorld 2.1 (offline subset, partial score)score·Anthropic · 2026-10-07effort maxsetup not fully disclosed
The benchmark's official offline subset: 82 of 108 tasks, VM without internet access. Mean per-task checkpoint credit, Pass@1 averaged over five attempts, 1080p, up to 500 action steps; strict pass rate 37.1%. Not comparable with the full-set OSWorld 2.1 partial score reported in earlier system cards. Re-evaluated under this configuration: Claude Sonnet 5.5 83.9%, Claude Opus 5.5 87.2%, Claude Haiku 4.5 15.7%.
72.4%
- AA-Briefcase v1.1score·Anthropic · 2026-10-07effort maxsetup not fully disclosed
Artificial Analysis long-horizon knowledge-work rating, not a percentage; run independently by Artificial Analysis. Same table: Claude Sonnet 5.5 1824, Claude Haiku 4.5 614, GPT-6 Luna 1336. At the default medium effort Anthropic reports 1372.
1578 - GDPval-AA v2.1elo·Anthropic · 2026-10-07effort maxsetup not fully disclosed
Run independently by Artificial Analysis. Elo is relative to the field, anchored to DeepSeek V4.1 Flash (max) at 1600; same table: Claude Sonnet 5.5 1840, Claude Haiku 4.5 735, GPT-6 Luna 1437. At the default medium effort Anthropic reports 1277.
1620
- Humanity's Last Examaccuracy·Anthropic · 2026-10-07effort maxsetup not fully disclosed
No tools, 980k-token task budget; graded by Claude Opus 4.6. Same table: Claude Sonnet 5.5 56.9%, Claude Haiku 4.5 10.2%.
45.9% - Humanity's Last Exam (with tools)accuracy·Anthropic · 2026-10-07effort maxsetup not fully disclosed
Tools were web search, web fetch, programmatic tool calling and code execution, with a 980k-token task budget and HLE-discussing sources blocklisted; graded by Claude Opus 4.6. Same table: Claude Sonnet 5.5 64.5%, Claude Haiku 4.5 18.7%.
57.4%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Claude Haiku 5.5?
How much does Claude Haiku 5.5 cost?
Does Claude Haiku 5.5 support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.