Qwen · Model guide
Qwen3.8 Flash
The hosted production build of Qwen3.8-Flash-Next, a 125B-parameter mixture of experts activating 6B per token plus a 51B n-gram embedding table. A hybrid Gated DeltaNet and Qwen Sparse Attention stack keeps a 1M-token context affordable to serve, and the model takes text, image and video in with adjustable reasoning, function calling, context caching and built-in tools. Qwen aims it at high-concurrency applications, tool-driven workflows and coding or office assistants.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorQwen says the model strikes an excellent balance between capability, latency and cost, and is well suited to high-concurrency applications, tool-driven workflows and coding and cowork assistants.source
High-concurrency workloads · Only 6B parameters are active per token, and Qwen reports up to 7.6x prefill and 4.9x decode attention-kernel speedups at 1M tokens.(vendor claim)source
Coding and cowork agents · Qwen reports its largest gains over Qwen3.7-Plus in coding and office tasks, and documents the model against Claude Code, Codex, Qoder and Qwen Code.(vendor claim)source
Million-token context · The hosted build serves a 1M-token context by default, with up to 991K input and a 262K thinking budget.source
- qwen/qwen3.8-maxstronger —Choose Qwen3.8 Max when maximum capability matters more than throughput or cost.
- qwen/qwen3.7-flashpredecessor —Stay on Qwen3.7 Flash when an existing integration is tuned to the previous generation's tiered context pricing.
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 0.13 |
| Output | $ 0.43 |
| Cache read | $ 0.016 |
| Cache write | $ 0.20 |
+ 2.0K × $ 0.43
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 131.1K tokens
- Input / output modalities
- Text + Image + Video → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· can be turned off
- Reasoning effort
- low · medium · xhigh
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- DeepSWE 1.1resolved_rate·Qwen Team · 2026-08-26setup not fully disclosed
官方在 Claude Code 与 mini-SWE-agent 两个框架下评测并取较高分(Qwen 说明本模型在 mini-SWE-agent 上最好),故不写单一 harness;上下文窗口 256K。
58.7% - 91.9%
- SWE-Bench Multilingualresolved_rate·Qwen Team · 2026-08-26mini-swe-agentsetup not fully disclosed
上下文窗口 256K。
81% - SWE-Bench Proresolved_rate·Qwen Team · 2026-08-26Claude Codesetup not fully disclosed
已修正存在问题的任务,并在修正后的基准上重测全部基线;上下文窗口 256K。口径与 qwen/qwen3.8-max 那条相同。
62.5%
- 91.7%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Qwen3.8 Flash?
How much does Qwen3.8 Flash cost?
Does Qwen3.8 Flash support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.