Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Grok 4.5Grok 4.3Grok Build 0.1Chat with these 3
AttributeGrok 4.5x-ai/grok-4.5Grok 4.3x-ai/grok-4.3Grok Build 0.1x-ai/grok-build-0.1
Specs
Context window500K1M256K
Max output500K1M256K
Input modalitiesTextImageTextImageTextImage
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 2.00>= 200K$ 4.00$ 1.25>= 200K$ 2.50$ 1.00>= 200K$ 2.00
Output$ 6.00>= 200K$ 12.00$ 2.50>= 200K$ 5.00$ 2.00>= 200K$ 4.00
Cache read$ 0.30>= 200K$ 0.60$ 0.20>= 200K$ 0.40$ 0.20>= 200K$ 0.40
Cache write$ 2.00>= 200K$ 4.00$ 1.25>= 200K$ 2.50$ 1.00>= 200K$ 2.00
What it is good at
Best for
  • Agentic engineering
  • Technical knowledge work
  • Large-context workflows
  • Tool-assisted knowledge work
  • Fast coding agents
  • Interactive edit loops
Published benchmarks
SWE Marathonpass_at_1
  • 29%(setup not fully disclosed)
Terminal-Bench 2.1pass_ratesetups differ by: reasoning
  • 83.3%(setup not fully disclosed)
  • 39.7%effort high(setup not fully disclosed)
  • 52.1%enabled(setup not fully disclosed)
DeepSWE 1.0resolved_rate
  • 62%provider(setup not fully disclosed)
DeepSWE 1.1resolved_rate
  • 53%mini-swe-agent
SWE-Bench Proresolved_rate
  • 64.7%(setup not fully disclosed)
SciCodescoresetups differ by: reasoning
  • 47.3%effort high(setup not fully disclosed)
  • 50.2%enabled(setup not fully disclosed)
GDPval-AA v2elosetups differ by: reasoning
  • 1087.5effort high(setup not fully disclosed)
  • 1215.1enabled(setup not fully disclosed)
AA-LCRscoresetups differ by: reasoning
  • 66%effort high(setup not fully disclosed)
  • 70%enabled(setup not fully disclosed)
AA Intelligence Index v4.1.1scoresetups differ by: reasoning
  • 37.9effort high(setup not fully disclosed)
  • 40.7enabled(setup not fully disclosed)
Humanity's Last Examaccuracysetups differ by: reasoning
  • 37.2%effort high(setup not fully disclosed)
  • 38.3%enabled(setup not fully disclosed)
CritPtscoresetups differ by: reasoning
  • 8%effort high(setup not fully disclosed)
  • 9.1%enabled(setup not fully disclosed)
GPQA Diamondaccuracysetups differ by: reasoning
  • 90.1%effort high(setup not fully disclosed)
  • 89.5%enabled(setup not fully disclosed)
tau3-Bankingscoresetups differ by: reasoning
  • 12.4%effort high(setup not fully disclosed)
  • 13.4%enabled(setup not fully disclosed)