Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Claude Opus 4.8Claude Opus 5Chat with these 2
AttributeClaude Opus 4.8anthropic/claude-opus-4-8Claude Opus 5anthropic/claude-opus-5
Specs
Context window1M1M
Max output128K128K
Input modalitiesTextImageTextImage
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputSupportedSupported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 5.00$ 5.00
Output$ 25.00$ 25.00
Cache read$ 0.50$ 0.50
Cache write$ 6.25$ 6.25
What it is good at
Best for
  • Agentic coding
  • Security research with reduced guardrails
  • Agentic coding
  • Professional knowledge work
  • Scientific research
Published benchmarks
Frontier-Bench v0.1score
  • 43.3%(setup not fully disclosed)
FrontierCode v1.1 (Main)score
  • 53.4%(setup not fully disclosed)
Terminal-Bench 2.1pass_rate
  • 74.6%(setup not fully disclosed)
BrowseComp (with tools)accuracy
  • 90.8%(setup not fully disclosed)
AutomationBenchscore
  • 26%(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 69.2%(setup not fully disclosed)
OSWorld 2.0success_rate
  • 70.6%(setup not fully disclosed)
OSWorld-Verifiedsuccess_rate
  • 83.4%(setup not fully disclosed)
Finance Agent v2score
  • 53.9%(setup not fully disclosed)
GDPval-AAelo
  • 1890(setup not fully disclosed)
GDPval-AA v2elo
  • 1861(setup not fully disclosed)
ARC-AGI-3score
  • 30.2%(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 49.8%(setup not fully disclosed)
  • 56.3%(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracy
  • 57.9%(setup not fully disclosed)
  • 64.7%(setup not fully disclosed)