Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Grok 4.3Grok 4.5Chat with these 2
AttributeGrok 4.3x-ai/grok-4.3Grok 4.5x-ai/grok-4.5
Specs
Context window1M500K
Max output1M500K
Input modalitiesTextImageTextImage
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputSupportedSupported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 1.25>= 200K$ 2.50$ 2.00>= 200K$ 4.00
Output$ 2.50>= 200K$ 5.00$ 6.00>= 200K$ 12.00
Cache read$ 0.20>= 200K$ 0.40$ 0.30>= 200K$ 0.60
Cache write$ 1.25>= 200K$ 2.50$ 2.00>= 200K$ 4.00
What it is good at
Best for
  • Large-context workflows
  • Tool-assisted knowledge work
  • Agentic engineering
  • Technical knowledge work
Published benchmarks
SWE Marathonpass_at_1
  • 29%(setup not fully disclosed)
Terminal-Bench 2.1pass_ratesetups differ by: reasoning
  • 39.7%effort high(setup not fully disclosed)
  • 83.3%(setup not fully disclosed)
DeepSWE 1.0resolved_rate
  • 62%provider(setup not fully disclosed)
DeepSWE 1.1resolved_rate
  • 53%mini-swe-agent
SWE-Bench Proresolved_rate
  • 64.7%(setup not fully disclosed)
SciCodescore
  • 47.3%effort high(setup not fully disclosed)
GDPval-AA v2elo
  • 1087.5effort high(setup not fully disclosed)
AA-LCRscore
  • 66%effort high(setup not fully disclosed)
AA Intelligence Index v4.1.1score
  • 37.9effort high(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 37.2%effort high(setup not fully disclosed)
CritPtscore
  • 8%effort high(setup not fully disclosed)
GPQA Diamondaccuracy
  • 90.1%effort high(setup not fully disclosed)
tau3-Bankingscore
  • 12.4%effort high(setup not fully disclosed)