Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

GLM-5.1GLM-5.2GLM-5 TurboChat with these 3
AttributeGLM-5.1z-ai/glm-5.1GLM-5.2z-ai/glm-5.2GLM-5 Turboz-ai/glm-5-turbo
Specs
Context window200K1M200K
Max output128K128K128K
Input modalitiesTextTextText
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputNot supportedNot supportedNot supported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 1.00>= 32K$ 1.20$ 1.20$ 0.90>= 32K$ 1.10
Output$ 3.80>= 32K$ 4.40$ 4.40$ 3.70>= 32K$ 4.30
Cache read$ 0.20>= 32K$ 0.30$ 0.30$ 0.18>= 32K$ 0.27
Cache write$ 1.00>= 32K$ 1.20$ 1.20$ 0.90>= 32K$ 1.10
What it is good at
Best for
  • Long-horizon engineering
  • Performance engineering
  • Million-token workflows
  • Agentic software engineering
  • Persistent agent workflows
Published benchmarks
Terminal-Bench 2.1pass_rate
  • 81%Terminus-2
SWE-Bench Proresolved_rate
  • 58.4%(setup not fully disclosed)
  • 62.1%(setup not fully disclosed)
SciCodescore
  • 43.6%enabled(setup not fully disclosed)
AA-LCRscore
  • 66.7%enabled(setup not fully disclosed)
KernelBench Level 3 Optimizationgeometric_mean_speedup
  • 3.6×(setup not fully disclosed)
AA Intelligence Index v4.1.1score
  • 39.1enabled(setup not fully disclosed)
GPQA Diamondaccuracy
  • 91.2%(setup not fully disclosed)
Humanity's Last Examaccuracysetups differ by: reasoning, tools
  • 40.5%
  • 27.8%enabled(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracy
  • 54.7%
CritPtscore
  • 0.3%enabled(setup not fully disclosed)
GPQA Diamondaccuracy
  • 84.7%enabled(setup not fully disclosed)
MCP Atlas Publicscore
  • 76.8%(setup not fully disclosed)
Tool Decathlonscore
  • 48.2%(setup not fully disclosed)