Claude Opus 5.5
anthropic/claude-opus-5-5
Back to model catalog

Anthropic · Model guide

Claude Opus 5.5

Compare
anthropic/claude-opus-5-5Released Sep 22, 2026

The first model in Anthropic's Claude 5.5 family and the one Anthropic recommends starting with for most workloads, built for long-running agentic coding and knowledge work. Anthropic reports it performs at the level of Claude Fable 5.1 on most work and generates output noticeably faster than Opus 5. Thinking is always on, effort is the only depth control and now defaults to medium, and forced tool choice is rejected.

Modalities
TextImageText
Context
1M
Max output
128K
Price / 1M tokens
$ 4.00in$ 20.00out
StreamingTool callingJSON outputImage inputThinking mode

What this model is good at

What the vendor positions it for, and which of our models to reach for instead.

VendorFor long-running agentic coding and knowledge work.source

Best for
  • Agentic coding · Anthropic reports it ahead of Claude Fable 5.1 and Claude Opus 5 on Terminal-Bench 4.0, FrontierCode and CursorBench 4.0, and highlights codebase-wide migrations and audits.(vendor claim)source

  • Professional knowledge work · Anthropic reports 1846 on GDPval-AA v2.1, ahead of Claude Fable 5.1 and Claude Opus 5 in the same table.(vendor claim)source

  • Charts and screenshots · Anthropic reports much more precise reading of dense charts, diagrams and screenshots than Claude Opus 5, even without tools.(vendor claim)source

Neighbouring models
  • anthropic/claude-fable-5-1stronger —Anthropic recommends Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when Opus 5.5 at higher effort still falls short.(vendor claim)source
  • anthropic/claude-opus-5predecessor —Claude Opus 5 still accepts thinking disabled and forced tool choice, both of which return a 400 on this model.source
  • anthropic/claude-sonnet-5cheaper —Claude Sonnet 5 is half the input price for agent work that does not need an Opus-class model.
Compare them side by side →

Pricing and billing

Billed per token. Cached input is charged at the cache-read rate.

Price / 1M tokensYou pay
Input$ 4.00
Output$ 20.00
Cache read$ 0.20
Cache write$ 5.00
Estimate a request
$ 0.0800
Estimated cost per request
10.0K × $ 4.00
+ 2.0K × $ 20.00
≈ $ 80.00 per 1,000 requests

Capabilities and limits

What CrossModel guarantees across every route this model can take right now.

Context window
1.0M tokens
Max output
128.0K tokens
Input / output modalities
Text + Image → Text
Streaming
Supported
Tool calling
Supported
Structured output (JSON)
Supported
Image input
Supported
Thinking mode
Supported· always on
Reasoning effort
low · medium · high · xhigh · max
Available endpoints
/v1/chat/completions · /v1/responses · /v1/messages
Conditional limitsDisclosed by the vendor. These do not make the capability unavailable — they say when it is not.
  • reasoning.toggle is unavailable when source
  • tool_use.forced_choice is unavailable when source

Published benchmarks

Scores the vendor reported, with the evaluation setup each one came from.

Agentic coding
  • CursorBench 4.0
    score·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    At the default medium effort Anthropic reports 52.5%. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    57.8%
  • FrontierCode v1.1 (Main)
    score·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    At the default medium effort Anthropic reports 54.6%. This table rates Claude Opus 5 at 48.0% while the Claude Opus 5 launch table rated it 53.4%; compare within one table. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    54.4%
  • Terminal-Bench 4.0
    pass_rate·Anthropic · 2026-09-22effort xhighsetup not fully disclosed

    Reported at xhigh effort, Anthropic's highest score for this model. Standard error ±2.6 pts. Anthropic's setup reproduces the public leaderboard (5 trials/task, Claude Code harness) for Claude Opus 5 at 52.3% against 51.8% published. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    66.4%
Agentic work
  • AutomationBench
    score·Anthropic · 2026-09-22setup not fully disclosed

    Run and reported by Zapier during early access, without fallback models, so safeguard interventions counted as failures; Anthropic states this lowers the score against practical use. Effort level not disclosed.

    40%
Computer use
  • OSWorld 2.1 (partial credit)
    score·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    Mean per-task checkpoint credit on OSWorld 2.1 (task files of 2026-09-10), not the OSWorld 2.0 release behind Claude Fable 5.1's own 77.9%. Strict pass rate 48.7% per the Claude Sonnet 5.5 system card, which reuses this configuration. Same table: Claude Fable 5.1 80.7%, Claude Opus 5 74.0%. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    81.8%
Knowledge work
  • GDPval-AA v2.1
    elo·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    Artificial Analysis's evaluation across 44 occupations. Elo is relative to the field; this table rates Claude Fable 5.1 at 1735 and Claude Opus 5 at 1708. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    1846
Multimodal reasoning
  • Chartography (with tools)
    score·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    Visual chart recognition. Anthropic does not disclose which tools the run had. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    89%
Reasoning
  • Humanity's Last Exam (with tools)
    accuracy·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    Anthropic does not disclose which tools the run had. This table rates Claude Fable 5.1 at 65.6% and Claude Opus 5 at 63.6%, slightly off the 65.0% and 64.7% stored from their own launch tables. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    67.7%
Scientific research
  • Terminal-Bench-Science 0.1
    accuracy·Anthropic · 2026-09-22effort maxsetup not fully disclosed

    Standard error ±3.5–5 pts per model. Anthropic's setup reproduces the public leaderboard (3 trials/task, Claude Code harness) for Claude Opus 5 at 29.0% against 30.0% published. Anthropic evaluated Claude Opus 5.5 with production safeguards enabled; when they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier LLM development tasks by Claude Opus 5, which Anthropic states likely lowers the score.

    58.7%

Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.

Use it in your tools

Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.

See all integration guides →

Frequently asked questions

What is Claude Opus 5.5?
Claude Opus 5.5 is available on CrossModel as anthropic/claude-opus-5-5. It has a 1M-token context window and can return up to 128K tokens per request.
How much does Claude Opus 5.5 cost?
Currently $ 4.00 per 1M input tokens and $ 20.00 per 1M output tokens. Cached input is billed at $ 0.20 per 1M.
Does Claude Opus 5.5 support tool calling and structured output?
Tool calling: Supported · Structured output (JSON): Supported · Image input: Supported · Streaming: Supported · Thinking mode: Supported
Which endpoint do I call?
For text models, OpenAI Chat Completions (/v1/chat/completions), OpenAI Responses (/v1/responses) and Anthropic Messages (/v1/messages) all work — pick whichever your SDK already speaks.
Can I try it without writing code?
Yes. Sign in and open the Playground on this page to chat with the model directly — no API key needed. Usage is billed from your wallet at the normal rate.

Chat first, integrate later

Chat first, then wire it in once you like the answers.