Tencent · Model guide
Hy4 Preview
Tencent's next-generation open flagship: a 770B-parameter mixture of experts that activates 49B per token across 78 layers, with a native multi-token-prediction layer for speculative decoding. Gated DeepSeek Sparse Attention with cross-layer index reuse carries a one-million-token context. This preview is text-only and defaults to deep chain-of-thought reasoning, with a direct-answer mode available, and Tencent built its training data around the work its own software, finance, game and security teams ship.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorTencent presents Hy4 preview as an early release at the open-source frontier, built for productivity work, and says it ships with known issues so that feedback can shape the full version.source
Agentic coding · Tencent reports 85.4 on Terminal-Bench 2.1 and 82.9 on SWE-bench Multilingual, and co-designs the model with its own CodeBuddy coding agent.source
Long-context work · The context window reaches one million tokens, four times the previous generation, and Tencent scores it on a with-tools million-token benchmark.source
Office and analysis · Tencent highlights turning scattered context into documents, spreadsheets and presentations, with data analysis and financial modelling.(vendor claim)source
Scientific research · Tencent claims gains on hard research questions across AI research, molecular dynamics, condensed matter physics and pure mathematics.(vendor claim)source
- tencent/hy3predecessor —Hy3 is the previous generation: a much smaller and cheaper model with a 256K context, and the stable release rather than a preview.
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 0.96 |
| Output | $ 2.88 |
| Cache read | $ 0.048 |
| Cache write | $ 0.96 |
+ 2.0K × $ 2.88
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 65.5K tokens
- Input / output modalities
- Text → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Not supported
- Thinking mode
- Supported· can be turned off
- Reasoning effort
- none · high
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- Terminal-Bench 2.1pass_rate·Tencent Hy · 2026-08-28effort highClaude Codesetup not fully disclosed
Claude Code harness, up to 500 turns and a 12-hour timeout per trial, capped at 16 CPUs and 32 GB RAM. The same table reports Hy3 at 70.8 rather than the 71.7 in this catalogue; Tencent notes the harness, judge model and anti-hacking mechanisms changed, so the two are not the same run.
85.4%
- 92.3%
- MCP Atlas Publicscore·Tencent Hy · 2026-08-28effort highsetup not fully disclosed
Scale's April 2026 methodology on the 500-task public set, 100 tool-call budget per task, with the judge model moved from Gemini 2.5 Pro to Gemini 3.1 Pro Preview.
83.7% - Toolathlon Verifiedpass_rate·Tencent Hy · 2026-08-28effort high3 runssetup not fully disclosed
Tencent's internal agent scaffold with some MCP tools reimplemented, 2 CPU cores and 10 GB per task and a 2-hour timeout instead of 5400 seconds; the 108-task Verified set, official evaluators, 100-step budget and 3-run Pass@1 protocol are unchanged.
74.1%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Hy4 Preview?
How much does Hy4 Preview cost?
Does Hy4 Preview support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.