Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Claude Sonnet 4.6Claude Sonnet 5Claude Haiku 4.5Chat with these 3
AttributeClaude Sonnet 4.6anthropic/claude-sonnet-4-6Claude Sonnet 5anthropic/claude-sonnet-5Claude Haiku 4.5anthropic/claude-haiku-4-5
Specs
Context window1M1M200K
Max output128K128K64K
Input modalitiesTextImageTextImageTextImage
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 3.00$ 2.00$ 1.00
Output$ 15.00$ 10.00$ 5.00
Cache read$ 0.30$ 0.20$ 0.10
Cache write$ 3.75$ 2.50$ 1.25
What it is good at
Best for
  • Agentic coding
  • Long-context work
  • Agentic coding
  • High-volume agent workloads
  • Real-time, low-latency tasks
  • High-volume workloads
Published benchmarks
Terminal-Benchpass_rate
  • 41%(setup not fully disclosed)
Terminal-Bench 2.0pass_rate
  • 59.1%(setup not fully disclosed)
Terminal-Bench 2.1pass_rate
  • 80.4%(setup not fully disclosed)
BrowseComp (with tools)accuracy
  • 74.7%(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 63.2%(setup not fully disclosed)
SWE-Bench Verifiedresolved_rate
  • 79.6%(setup not fully disclosed)
  • 73.3%(setup not fully disclosed)
OSWorldsuccess_rate
  • 50.7%(setup not fully disclosed)
OSWorld-Verifiedsuccess_rate
  • 78.5%(setup not fully disclosed)
  • 81.2%(setup not fully disclosed)
GDPval-AAelo
  • 1633(setup not fully disclosed)
GDPval-AA v2elo
  • 1618(setup not fully disclosed)
AIME 2025accuracy
  • 80.7%(setup not fully disclosed)
MMMLUaccuracy
  • 89.3%(setup not fully disclosed)
  • 83%(setup not fully disclosed)
MMMUaccuracy
  • 73.2%(setup not fully disclosed)
MMMU Pro (no tools)accuracy
  • 74.5%(setup not fully disclosed)
ARC-AGI-2score
  • 58.3%(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 34.6%(setup not fully disclosed)
  • 43.2%(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracy
  • 46.8%(setup not fully disclosed)
  • 57.4%(setup not fully disclosed)
GPQA Diamondaccuracy
  • 89.9%(setup not fully disclosed)
  • 73%(setup not fully disclosed)
τ2-bench Retailscore
  • 83.2%(setup not fully disclosed)
τ2-bench Telecomscore
  • 83%(setup not fully disclosed)