Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Claude Haiku 4.5Claude Sonnet 5Chat with these 2
AttributeClaude Haiku 4.5anthropic/claude-haiku-4-5Claude Sonnet 5anthropic/claude-sonnet-5
Specs
Context window200K1M
Max output64K128K
Input modalitiesTextImageTextImage
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputSupportedSupported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 1.00$ 2.00
Output$ 5.00$ 10.00
Cache read$ 0.10$ 0.20
Cache write$ 1.25$ 2.50
What it is good at
Best for
  • Real-time, low-latency tasks
  • High-volume workloads
  • Agentic coding
  • High-volume agent workloads
Published benchmarks
Terminal-Benchpass_rate
  • 41%(setup not fully disclosed)
Terminal-Bench 2.1pass_rate
  • 80.4%(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 63.2%(setup not fully disclosed)
SWE-Bench Verifiedresolved_rate
  • 73.3%(setup not fully disclosed)
OSWorldsuccess_rate
  • 50.7%(setup not fully disclosed)
OSWorld-Verifiedsuccess_rate
  • 81.2%(setup not fully disclosed)
GDPval-AA v2elo
  • 1618(setup not fully disclosed)
AIME 2025accuracy
  • 80.7%(setup not fully disclosed)
MMMLUaccuracy
  • 83%(setup not fully disclosed)
MMMUaccuracy
  • 73.2%(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 43.2%(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracy
  • 57.4%(setup not fully disclosed)
GPQA Diamondaccuracy
  • 73%(setup not fully disclosed)
τ2-bench Retailscore
  • 83.2%(setup not fully disclosed)
τ2-bench Telecomscore
  • 83%(setup not fully disclosed)