Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

GLM-5.3GLM-5.2Chat with these 2
AttributeGLM-5.3z-ai/glm-5.3GLM-5.2z-ai/glm-5.2
Specs
Context window1M1M
Max output128K128K
Input modalitiesTextText
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputNot supportedNot supported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 1.20$ 1.20
Output$ 4.40$ 4.40
Cache read$ 0.30$ 0.30
Cache write$ 1.20$ 1.20
What it is good at
Best for
  • Long-horizon software engineering
  • Security-focused code review
  • Million-token workflows
  • Agentic software engineering
Published benchmarks
Terminal-Bench 2.1pass_rate
  • 81%Terminus-2
Terminal-Bench 3.0pass_rate
  • 28.3%effort max · Claude Code 2.1.207 · 3 runs(setup not fully disclosed)
Agents' Last Exam (CLI)score
  • 28.5%effort max(setup not fully disclosed)
DeepSWE 1.1resolved_rate
  • 66.9%mini-swe-agent(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 62.1%(setup not fully disclosed)
CyberGymscore
  • 84.5%effort max · Claude Code 2.1.207(setup not fully disclosed)
ExploitBenchscore
  • 54.4%effort max · Claude Code 2.1.207 · 3 runs(setup not fully disclosed)
ExploitGym (6h)tasks_completed
  • 130effort max · Claude Code 2.1.207(setup not fully disclosed)
GDPval-AA v2elo
  • 1769(setup not fully disclosed)
GPQA Diamondaccuracy
  • 91.2%(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 40.5%
Humanity's Last Exam (with tools)accuracysetups differ by: reasoning, tools
  • 62.5%effort max(setup not fully disclosed)
  • 54.7%
MCP Atlas Publicscore
  • 76.8%(setup not fully disclosed)
Tool Decathlonscore
  • 48.2%(setup not fully disclosed)