Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Grok Build 0.1Grok 4.5Chat with these 2
AttributeGrok Build 0.1x-ai/grok-build-0.1Grok 4.5x-ai/grok-4.5
Specs
Context window256K500K
Max output256K500K
Input modalitiesTextImageTextImage
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputSupportedSupported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 1.00>= 200K$ 2.00$ 2.00>= 200K$ 4.00
Output$ 2.00>= 200K$ 4.00$ 6.00>= 200K$ 12.00
Cache read$ 0.20>= 200K$ 0.40$ 0.30>= 200K$ 0.60
Cache write$ 1.00>= 200K$ 2.00$ 2.00>= 200K$ 4.00
What it is good at
Best for
  • Fast coding agents
  • Interactive edit loops
  • Agentic engineering
  • Technical knowledge work
Published benchmarks
SWE Marathonpass_at_1
  • 29%(setup not fully disclosed)
Terminal-Bench 2.1pass_ratesetups differ by: reasoning
  • 52.1%enabled(setup not fully disclosed)
  • 83.3%(setup not fully disclosed)
DeepSWE 1.0resolved_rate
  • 62%provider(setup not fully disclosed)
DeepSWE 1.1resolved_rate
  • 53%mini-swe-agent
SWE-Bench Proresolved_rate
  • 64.7%(setup not fully disclosed)
SciCodescore
  • 50.2%enabled(setup not fully disclosed)
GDPval-AA v2elo
  • 1215.1enabled(setup not fully disclosed)
AA-LCRscore
  • 70%enabled(setup not fully disclosed)
AA Intelligence Index v4.1.1score
  • 40.7enabled(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 38.3%enabled(setup not fully disclosed)
CritPtscore
  • 9.1%enabled(setup not fully disclosed)
GPQA Diamondaccuracy
  • 89.5%enabled(setup not fully disclosed)
tau3-Bankingscore
  • 13.4%enabled(setup not fully disclosed)