Anthropic · Model guide
Claude Sonnet 5.5
The second model in Anthropic's Claude 5.5 family, positioned as its best combination of speed and intelligence and a faster complement to Claude Opus 5.5. Anthropic describes it as strongest at well-scoped everyday tasks, bug fixing, and polished documents, slides and spreadsheets, and reports output more than 30% faster than Sonnet 5. Adaptive thinking is on by default at high effort; thinking can no longer be set to disabled, and forced tool choice is rejected.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorThe best combination of speed and intelligence.source
Agentic coding · Anthropic reports 70.6% on Terminal-Bench 4.0 against 10.3% for Claude Sonnet 5, and a CursorBench 4.0 score within about two points of Claude Opus 5.5.(vendor claim)source
Documents, slides and spreadsheets · Anthropic reports 1844 on GDPval-AA v2.1, two points below Claude Opus 5.5, and highlights decks and documents that need minimal editing.(vendor claim)source
Charts and computer use · Anthropic reports it close to Claude Opus 5.5 on Chartography and OSWorld 2.1, far ahead of Claude Sonnet 5 on both.(vendor claim)source
- anthropic/claude-opus-5-5stronger —Anthropic states Claude Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment.(vendor claim)source
- anthropic/claude-sonnet-5predecessor —Claude Sonnet 5 still accepts thinking disabled and forced tool choice, both of which return a 400 on this model.source
- anthropic/claude-haiku-5-5cheaper —Claude Haiku 5.5 is cheaper and faster again for short, latency-bound calls.
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 2.00 |
| Output | $ 10.00 |
| Cache read | $ 0.10 |
| Cache write | $ 2.50 |
+ 2.0K × $ 10.00
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 128.0K tokens
- Input / output modalities
- Text + Image → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· always on
- Reasoning effort
- low · medium · high · xhigh · max
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- CursorBench 4.0score·Anthropic · 2026-09-28effort maxsetup not fully disclosed
Measured and reported by Cursor in its production agent harness. Anthropic reports 53.1% at xhigh, 47.8% at high and 39.2% at medium. Same table: Claude Opus 5.5 57.8%, Claude Sonnet 5 34.1%.
55.5% - FrontierCode v1.1 (Main)score·Anthropic · 2026-09-28effort maxClaude Codesetup not fully disclosed
Run and reported by Cognition in Claude Code. At xhigh effort Anthropic reports 52.1%: the grader penalises out-of-scope edits, and at max the model more often ran a multi-subagent code review that led to timeouts or extra edits. Same table: Claude Opus 5.5 54.4%, Claude Sonnet 5 42.4%.
46.2% - Terminal-Bench 4.0pass_rate·Anthropic · 2026-09-28effort maxClaude Codesetup not fully disclosed
Claude Code in --bare mode, five trials per task, no internet egress; standard error ±2.5 pts. Run with safeguards on: 1.2% of requests were answered by a fallback model, affecting 1.5% of trials. Same table: Claude Opus 5.5 66.4% (at xhigh), Claude Sonnet 5 10.3%.
70.6%
- OSWorld 2.1 (partial credit)score·Anthropic · 2026-09-28effort maxsetup not fully disclosed
Mean per-task checkpoint credit, Pass@1 averaged over five runs, 1080p, up to 500 action steps, task files of 2026-09-10 (v2.1); strict pass rate 43.5%. Same configuration and table: Claude Opus 5.5 81.8%, Claude Sonnet 5 57.0%.
80.1%
- AA-Briefcase v1.1score·Anthropic · 2026-09-28effort maxsetup not fully disclosed
Artificial Analysis long-horizon knowledge-work rating, not a percentage. Run on the same pre-release deployment as GDPval-AA. Same table: Claude Opus 5.5 1822, Claude Sonnet 5 1359. At xhigh effort Anthropic reports 1746.
1811 - GDPval-AA v2.1elo·Anthropic · 2026-09-28effort maxsetup not fully disclosed
Run by Artificial Analysis on a pre-release deployment that Anthropic says had a since-fixed bug degrading structured-output requests. Elo is relative to the field; same table: Claude Opus 5.5 1846, Claude Sonnet 5 1449. At xhigh effort Anthropic reports 1725.
1844
- Humanity's Last Examaccuracy·Anthropic · 2026-09-28effort maxsetup not fully disclosed
No tools, 980k-token task budget. From the system card's capability summary; the same table rates Claude Opus 5.5 at 64.4% and Claude Sonnet 5 at 43.1%.
56.9% - Humanity's Last Exam (with tools)accuracy·Anthropic · 2026-09-28effort maxsetup not fully disclosed
Tools were web search, web fetch, programmatic tool calling and code execution, with a 980k-token task budget and HLE-discussing sources blocklisted. Same table: Claude Opus 5.5 67.7%, Claude Sonnet 5 54.9%.
64.5%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Claude Sonnet 5.5?
How much does Claude Sonnet 5.5 cost?
Does Claude Sonnet 5.5 support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.