Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Hy4 PreviewHy3Chat with these 2
AttributeHy4 Previewtencent/hy4-previewHy3tencent/hy3
Specs
Context window1M262K
Max output66K131K
Input modalitiesTextText
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputNot supportedNot supported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 0.96$ 0.16
Output$ 2.88$ 0.64
Cache read$ 0.048$ 0.040
Cache write$ 0.96$ 0.16
What it is good at
Best for
  • Agentic coding
  • Long-context work
  • Office and analysis
  • Scientific research
  • Agentic coding
  • Productivity agents
Published benchmarks
Terminal-Bench 2.1pass_ratesetups differ by: harness
  • 85.4%effort high · Claude Code(setup not fully disclosed)
  • 71.7%effort high · Terminus-2(setup not fully disclosed)
BrowseComp (with tools)accuracy
  • 84.2%effort high(setup not fully disclosed)
Claw-Eval Generalpass3_rate
  • 68.5%effort high(setup not fully disclosed)
Agents' Last Exam (CLI)score
  • 22.8%effort high · Claude Code(setup not fully disclosed)
SWE-Bench Multilingualresolved_rate
  • 82.9%effort high · SWE-agent(setup not fully disclosed)
  • 75.8%effort high · SWE-agent(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 65.7%effort high · SWE-agent(setup not fully disclosed)
  • 57.9%effort high · SWE-agent(setup not fully disclosed)
SWE-Bench Verifiedresolved_rate
  • 78%effort high(setup not fully disclosed)
OfficeQA Proscore
  • 66.2%effort high · Claude Code(setup not fully disclosed)
GDPval-AA v2elo
  • 1678effort high(setup not fully disclosed)
AA-LCRscore
  • 73.4%effort high(setup not fully disclosed)
OneMillionBench (with tools)score
  • 65.4%effort high(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracysetups differ by: tools
  • 55.4%effort high(setup not fully disclosed)
  • 53.2%effort high(setup not fully disclosed)
GPQA Diamondaccuracy
  • 92.3%effort high(setup not fully disclosed)
  • 90.4%effort high(setup not fully disclosed)
MCP Atlas Publicscore
  • 83.7%effort high(setup not fully disclosed)
  • 79.1%effort high(setup not fully disclosed)
Toolathlon Verifiedpass_rate
  • 74.1%effort high · 3 runs(setup not fully disclosed)