Anthropic · Model guide
Claude Opus 5.5
The first model in Anthropic's Claude 5.5 family and the one Anthropic recommends starting with for most workloads, built for long-running agentic coding and knowledge work. Anthropic reports it performs at the level of Claude Fable 5.1 on most work and generates output noticeably faster than Opus 5. Thinking is always on, effort is the only depth control and now defaults to medium, and forced tool choice is rejected.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorFor long-running agentic coding and knowledge work.source
Agentic coding · Anthropic reports it ahead of Claude Fable 5.1 and Claude Opus 5 on Terminal-Bench 4.0, FrontierCode and CursorBench 4.0, and highlights codebase-wide migrations and audits.(vendor claim)source
Professional knowledge work · Anthropic reports 1846 on GDPval-AA v2.1, ahead of Claude Fable 5.1 and Claude Opus 5 in the same table.(vendor claim)source
Charts and screenshots · Anthropic reports much more precise reading of dense charts, diagrams and screenshots than Claude Opus 5, even without tools.(vendor claim)source
- anthropic/claude-fable-5-1stronger —Anthropic recommends Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when Opus 5.5 at higher effort still falls short.(vendor claim)source
- anthropic/claude-opus-5predecessor —Claude Opus 5 still accepts thinking disabled and forced tool choice, both of which return a 400 on this model.source
- anthropic/claude-sonnet-5cheaper —Claude Sonnet 5 is half the input price for agent work that does not need an Opus-class model.
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 4.00 |
| Output | $ 20.00 |
| Cache read | $ 0.20 |
| Cache write | $ 5.00 |
+ 2.0K × $ 20.00
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 128.0K tokens
- Input / output modalities
- Text + Image → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· always on
- Reasoning effort
- low · medium · high · xhigh · max
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- CursorBench 4.0score·Anthropic · 2026-09-22effort maxsetup not fully disclosed
At the default medium effort Anthropic reports 52.5%. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
57.8% - FrontierCode v1.1 (Main)score·Anthropic · 2026-09-22effort maxsetup not fully disclosed
At the default medium effort Anthropic reports 54.6%. This table rates Claude Opus 5 at 48.0% while the Claude Opus 5 launch table rated it 53.4%; compare within one table. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
54.4% - Terminal-Bench 4.0pass_rate·Anthropic · 2026-09-22effort xhighsetup not fully disclosed
Reported at xhigh effort, Anthropic's highest score for this model. Standard error ±2.6 pts. Anthropic's setup reproduces the public leaderboard (5 trials/task, Claude Code harness) for Claude Opus 5 at 52.3% against 51.8% published. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
66.4%
- OSWorld 2.1 (partial credit)score·Anthropic · 2026-09-22effort maxsetup not fully disclosed
Mean per-task checkpoint credit on OSWorld 2.1 (task files of 2026-09-10), not the OSWorld 2.0 release behind Claude Fable 5.1's own 77.9%. Strict pass rate 48.7% per the Claude Sonnet 5.5 system card, which reuses this configuration. Same table: Claude Fable 5.1 80.7%, Claude Opus 5 74.0%. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
81.8%
- GDPval-AA v2.1elo·Anthropic · 2026-09-22effort maxsetup not fully disclosed
Artificial Analysis's evaluation across 44 occupations. Elo is relative to the field; this table rates Claude Fable 5.1 at 1735 and Claude Opus 5 at 1708. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
1846
- Chartography (with tools)score·Anthropic · 2026-09-22effort maxsetup not fully disclosed
Visual chart recognition. Anthropic does not disclose which tools the run had. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
89%
- Humanity's Last Exam (with tools)accuracy·Anthropic · 2026-09-22effort maxsetup not fully disclosed
Anthropic does not disclose which tools the run had. This table rates Claude Fable 5.1 at 65.6% and Claude Opus 5 at 63.6%, slightly off the 65.0% and 64.7% stored from their own launch tables. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
67.7%
- Terminal-Bench-Science 0.1accuracy·Anthropic · 2026-09-22effort maxsetup not fully disclosed
Standard error ±3.5–5 pts per model. Anthropic's setup reproduces the public leaderboard (3 trials/task, Claude Code harness) for Claude Opus 5 at 29.0% against 30.0% published. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.
58.7%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Claude Opus 5.5?
How much does Claude Opus 5.5 cost?
Does Claude Opus 5.5 support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.