Qwen3.8 Flash
qwen/qwen3.8-flash
Back to model catalog

Qwen · Model guide

Qwen3.8 Flash

Compare
qwen/qwen3.8-flashReleased Aug 26, 2026

The hosted production build of Qwen3.8-Flash-Next, a 125B-parameter mixture of experts activating 6B per token plus a 51B n-gram embedding table. A hybrid Gated DeltaNet and Qwen Sparse Attention stack keeps a 1M-token context affordable to serve, and the model takes text, image and video in with adjustable reasoning, function calling, context caching and built-in tools. Qwen aims it at high-concurrency applications, tool-driven workflows and coding or office assistants.

Modalities
TextImageVideoText
Context
1M
Max output
131K
Price / 1M tokens
$ 0.13in$ 0.43out
StreamingTool callingJSON outputImage inputThinking mode · can be turned off

What this model is good at

What the vendor positions it for, and which of our models to reach for instead.

VendorQwen says the model strikes an excellent balance between capability, latency and cost, and is well suited to high-concurrency applications, tool-driven workflows and coding and cowork assistants.source

Best for
  • High-concurrency workloads · Only 6B parameters are active per token, and Qwen reports up to 7.6x prefill and 4.9x decode attention-kernel speedups at 1M tokens.(vendor claim)source

  • Coding and cowork agents · Qwen reports its largest gains over Qwen3.7-Plus in coding and office tasks, and documents the model against Claude Code, Codex, Qoder and Qwen Code.(vendor claim)source

  • Million-token context · The hosted build serves a 1M-token context by default, with up to 991K input and a 262K thinking budget.source

Neighbouring models
  • qwen/qwen3.8-maxstronger —Choose Qwen3.8 Max when maximum capability matters more than throughput or cost.
  • qwen/qwen3.7-flashpredecessor —Stay on Qwen3.7 Flash when an existing integration is tuned to the previous generation's tiered context pricing.
Compare them side by side →

Pricing and billing

Billed per token. Cached input is charged at the cache-read rate.

Price / 1M tokensYou pay
Input$ 0.13
Output$ 0.43
Cache read$ 0.016
Cache write$ 0.20
Estimate a request
$ 0.0022
Estimated cost per request
10.0K × $ 0.13
+ 2.0K × $ 0.43
≈ $ 2.16 per 1,000 requests

Capabilities and limits

What CrossModel guarantees across every route this model can take right now.

Context window
1.0M tokens
Max output
131.1K tokens
Input / output modalities
Text + Image + Video → Text
Streaming
Supported
Tool calling
Supported
Structured output (JSON)
Supported
Image input
Supported
Thinking mode
Supported· can be turned off
Reasoning effort
low · medium · xhigh
Available endpoints
/v1/chat/completions · /v1/responses · /v1/messages

Published benchmarks

Scores the vendor reported, with the evaluation setup each one came from.

Agentic work
  • Agents' Last Exam
    score·Qwen Team · 2026-08-26setup not fully disclosed

    官方同时给出 Pass@1 24.3 与 Score 51.2,此处按目录内该基准统一的 score 口径记 51.2;来源页未披露评测设置。

    51.2%
Coding
  • DeepSWE 1.1
    resolved_rate·Qwen Team · 2026-08-26setup not fully disclosed

    官方在 Claude Code 与 mini-SWE-agent 两个框架下评测并取较高分(Qwen 说明本模型在 mini-SWE-agent 上最好),故不写单一 harness;上下文窗口 256K。

    58.7%
  • LiveCodeBench v6
    pass_rate·Qwen Team · 2026-08-26setup not fully disclosed

    官方评测表未给出该项的评测设置。

    91.9%
  • SWE-Bench Multilingual
    resolved_rate·Qwen Team · 2026-08-26mini-swe-agentsetup not fully disclosed

    上下文窗口 256K。

    81%
  • SWE-Bench Pro
    resolved_rate·Qwen Team · 2026-08-26Claude Codesetup not fully disclosed

    已修正存在问题的任务,并在修正后的基准上重测全部基线;上下文窗口 256K。口径与 qwen/qwen3.8-max 那条相同。

    62.5%
Computer use
  • AndroidWorld
    success_rate·Qwen Team · 2026-08-26setup not fully disclosed

    官方评测表未给出该项的评测设置。同表 Qwen3.7-Plus 记 81.0,与目录里那条一致。

    84.5%
Multimodal reasoning
  • CharXiv Reasoning (no tools)
    accuracy·Qwen Team · 2026-08-26setup not fully disclosed

    取官方「不含 CI」(不启用代码解释器)一列;启用代码解释器时为 90.6,属另一种设置,不并入本条。评测使用固定 prompt。

    84.6%
Reasoning
  • Humanity's Last Exam
    accuracy·Qwen Team · 2026-08-26setup not fully disclosed

    由 GPT-4o 判分;官方未说明是否使用工具,故按目录内不带工具的默认身份记录。

    35.9%
Scientific reasoning
  • GPQA Diamond
    accuracy·Qwen Team · 2026-08-26setup not fully disclosed

    官方评测表未给出该项的评测设置。

    91.7%
Tool use
  • Toolathlon Verified
    pass_rate·Qwen Team · 2026-08-26setup not fully disclosed

    Pass@1;口径与 qwen/qwen3.8-max 那条相同。

    73.5%

Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.

Use it in your tools

Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.

See all integration guides →

Frequently asked questions

What is Qwen3.8 Flash?
Qwen3.8 Flash is available on CrossModel as qwen/qwen3.8-flash. It has a 1M-token context window and can return up to 131K tokens per request.
How much does Qwen3.8 Flash cost?
Currently $ 0.13 per 1M input tokens and $ 0.43 per 1M output tokens. Cached input is billed at $ 0.016 per 1M.
Does Qwen3.8 Flash support tool calling and structured output?
Tool calling: Supported · Structured output (JSON): Supported · Image input: Supported · Streaming: Supported · Thinking mode: Supported
Which endpoint do I call?
For text models, OpenAI Chat Completions (/v1/chat/completions), OpenAI Responses (/v1/responses) and Anthropic Messages (/v1/messages) all work — pick whichever your SDK already speaks.
Can I try it without writing code?
Yes. Sign in and open the Playground on this page to chat with the model directly — no API key needed. Usage is billed from your wallet at the normal rate.

Chat first, integrate later

Chat first, then wire it in once you like the answers.