Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Qwen3.7 MaxQwen3.7 PlusChat with these 2
AttributeQwen3.7 Maxqwen/qwen3.7-max20% offQwen3.7 Plusqwen/qwen3.7-plus10% off
Specs
Context window1M1M
Max output66K64K
Input modalitiesTextTextImageVideo
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputNot supportedSupported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 1.88$ 1.50$ 0.32$ 0.29>= 256K$ 0.96$ 0.86
Output$ 5.63$ 4.50$ 1.25$ 1.13>= 256K$ 3.75$ 3.38
Cache read$ 0.38$ 0.30$ 0.032$ 0.029>= 256K$ 0.096$ 0.086
Cache write$ 2.35$ 1.88$ 0.40$ 0.36>= 256K$ 1.20$ 1.08
What it is good at
Best for
  • Complex reasoning
  • Long-running coding agents
  • Enterprise knowledge work
  • Multimodal coding agents
Published benchmarks
Terminal-Bench 2.0pass_rate
  • 69.7%Terminus(setup not fully disclosed)
  • 70.3%Terminus(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 60.6%(setup not fully disclosed)
  • 57.6%(setup not fully disclosed)
SWE-Bench Verifiedresolved_rate
  • 80.4%(setup not fully disclosed)
  • 77.7%(setup not fully disclosed)
AndroidWorldsuccess_rate
  • 81%(setup not fully disclosed)
OSWorld Verifiedsuccess_rate
  • 73.3%(setup not fully disclosed)
IFBenchscore
  • 79.1%(setup not fully disclosed)
MRCR v2 128Kscore
  • 90.4%(setup not fully disclosed)
  • 91.7%(setup not fully disclosed)
MMMU Proaccuracy
  • 79%(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 41.4%(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracy
  • 53.5%(setup not fully disclosed)
GPQA Diamondaccuracy
  • 92.4%(setup not fully disclosed)
  • 90.3%(setup not fully disclosed)
MCPMarkscore
  • 60.8%(setup not fully disclosed)
  • 58.7%(setup not fully disclosed)
VideoMMMUaccuracy
  • 85.4%(setup not fully disclosed)