Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

GPT-5.6 LunaGPT-5.6 TerraGPT-5.4 NanoChat with these 3
AttributeGPT-5.6 Lunaopenai/gpt-5.6-lunaGPT-5.6 Terraopenai/gpt-5.6-terraGPT-5.4 Nanoopenai/gpt-5.4-nano
Specs
Context window1M1M400K
Max output128K128K128K
Input modalitiesTextImageTextImageTextImage
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 0.20>= 272K$ 0.40$ 2.00>= 272K$ 4.00$ 0.20
Output$ 1.20>= 272K$ 1.80$ 12.00>= 272K$ 18.00$ 1.25
Cache read$ 0.020>= 272K$ 0.040$ 0.20>= 272K$ 0.40$ 0.020
Cache write$ 0.25>= 272K$ 0.50$ 2.50>= 272K$ 5.00$ 0.20
What it is good at
Best for
  • High-volume agent execution
  • Routine implementation and subagents
  • Everyday knowledge work
  • Cost-aware agentic coding
  • Classification and extraction
  • Ranking and simple subagents
Published benchmarks
Terminal-Bench 2.0pass_rate
  • 46.3%effort xhigh(setup not fully disclosed)
Terminal-Bench 2.1pass_rate
  • 84.7%effort max(setup not fully disclosed)
  • 87.4%effort max(setup not fully disclosed)
BrowseComp (with tools)accuracy
  • 83.3%effort max(setup not fully disclosed)
  • 87.5%effort max(setup not fully disclosed)
Agents' Last Examscore
  • 50.3%effort max(setup not fully disclosed)
  • 50.4%effort max(setup not fully disclosed)
SWE-Bench Pro (Public)pass_ratesetups differ by: reasoning
  • 62.7%effort max(setup not fully disclosed)
  • 63.4%effort max(setup not fully disclosed)
  • 52.4%effort xhigh(setup not fully disclosed)
OSWorld 2.0success_rate
  • 45.6%effort max(setup not fully disclosed)
  • 50.2%effort max(setup not fully disclosed)
OSWorld-Verifiedsuccess_rate
  • 39%effort xhigh(setup not fully disclosed)
OpenAI MRCR v2 8-needle 128K–256Kaccuracy
  • 33.1%effort xhigh(setup not fully disclosed)
OpenAI MRCR v2 8-needle 512K–1Maccuracy
  • 41.3%effort max(setup not fully disclosed)
  • 72.5%effort max(setup not fully disclosed)
MMMU Pro (no tools)accuracy
  • 66.1%effort xhigh(setup not fully disclosed)
MMMU Pro (with tools)accuracy
  • 79.5%effort max(setup not fully disclosed)
  • 82%effort max(setup not fully disclosed)
Humanity's Last Exam Full (with tools)accuracy
  • 37.7%effort xhigh(setup not fully disclosed)
GPQA Diamondaccuracysetups differ by: reasoning
  • 92.3%effort max(setup not fully disclosed)
  • 92.9%effort max(setup not fully disclosed)
  • 82.8%effort xhigh(setup not fully disclosed)
Toolathlonscoresetups differ by: reasoning
  • 53.4%effort max(setup not fully disclosed)
  • 53.1%effort max(setup not fully disclosed)
  • 35.5%effort xhigh(setup not fully disclosed)