Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

GPT-5.5 ProGPT-5.5GPT-5.6 SolChat with these 3
AttributeGPT-5.5 Proopenai/gpt-5.5-proGPT-5.5openai/gpt-5.5GPT-5.6 Solopenai/gpt-5.6-sol
Specs
Context window1M1M1M
Max output128K128K128K
Input modalitiesTextImageTextImageTextImage
Output modalitiesTextTextText
Capabilities
StreamingNot supportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 30.00>= 272K$ 60.00$ 5.00>= 272K$ 10.00$ 5.00>= 272K$ 10.00
Output$ 180.00>= 272K$ 270.00$ 30.00>= 272K$ 45.00$ 30.00>= 272K$ 45.00
Cache read$ 30.00>= 272K$ 60.00$ 0.50>= 272K$ 1.00$ 0.50>= 272K$ 1.00
Cache write$ 30.00>= 272K$ 60.00$ 5.00>= 272K$ 10.00$ 6.25>= 272K$ 12.50
What it is good at
Best for
  • Difficult research and analysis
  • High-stakes knowledge work
  • Agentic coding
  • Professional knowledge work
  • Frontier agentic coding
  • Complex professional work
Published benchmarks
Terminal-Bench 2.0pass_rate
  • 82.7%effort xhigh(setup not fully disclosed)
Terminal-Bench 2.1pass_rate
  • 88.8%effort max(setup not fully disclosed)
BrowseComp (with tools)accuracysetups differ by: reasoning
  • 90.1%effort xhigh(setup not fully disclosed)
  • 84.4%effort xhigh(setup not fully disclosed)
  • 90.4%effort max(setup not fully disclosed)
Agents' Last Examscore
  • 52.7%effort max(setup not fully disclosed)
SWE-Bench Pro (Public)pass_ratesetups differ by: reasoning
  • 58.6%effort xhigh(setup not fully disclosed)
  • 64.6%effort max(setup not fully disclosed)
OSWorld 2.0success_rate
  • 62.6%effort max(setup not fully disclosed)
OSWorld-Verifiedsuccess_rate
  • 78.7%effort xhigh(setup not fully disclosed)
GDPval (wins or ties)wins_or_ties
  • 82.3%effort xhigh(setup not fully disclosed)
  • 84.9%effort xhigh(setup not fully disclosed)
OpenAI MRCR v2 8-needle 512K–1Maccuracysetups differ by: reasoning
  • 74%effort xhigh(setup not fully disclosed)
  • 73.8%effort max(setup not fully disclosed)
FrontierMath Tier 4accuracy
  • 39.6%effort xhigh(setup not fully disclosed)
MMMU Pro (with tools)accuracysetups differ by: reasoning
  • 83.2%effort xhigh(setup not fully disclosed)
  • 84.6%effort max(setup not fully disclosed)
Humanity's Last Exam Full (with tools)accuracy
  • 57.2%effort xhigh(setup not fully disclosed)
  • 52.2%effort xhigh(setup not fully disclosed)
GPQA Diamondaccuracysetups differ by: reasoning
  • 93.6%effort xhigh(setup not fully disclosed)
  • 94.6%effort max(setup not fully disclosed)
GeneBenchaccuracy
  • 33.2%effort xhigh(setup not fully disclosed)
Toolathlonscoresetups differ by: reasoning
  • 55.6%effort xhigh(setup not fully disclosed)
  • 58%effort max(setup not fully disclosed)