Qwen · Model guide
Qwen3.8 Omni Flash
Qwen's latest natively omni-modal model, built on the Qwen3.8-Flash-Next architecture. It takes text, image, audio and video into a single one-million-token context and answers in text, with adjustable reasoning, function calling and context caching. Qwen aims it at agentic work where sound and video are the raw material: video editing, music-video creation, film commentary, meetings transcribed with speaker attribution, and recordings turned into structured notes.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorQwen says the model's core goal is to strengthen agent ability in real productivity settings, pushing omni-modal models from understanding omni-modal content towards planning tasks, calling tools and producing finished work.source
Audio and video agent workflows · Qwen reports a 36.5-point gain on WildClawBench-MM and 22.3 on AgenticVBench over the previous generation, and ships plugins for video editing, dubbing and film commentary pipelines.(vendor claim)source
Hour-long recordings · A one-million-token context with up to 991K input, a 262K thinking budget, and native support for audio-visual input up to one hour long.source
Meetings and long-horizon office work · Joint audio and visual speaker identification cuts AliMeeting diarisation error from 88.1 to 3.4, and the model scores 75.3 on Qwen's internal CoWorkBench office benchmark.(vendor claim)source
- qwen/qwen3.8-flashsibling —Choose Qwen3.8 Flash when the inputs are only text, images and video, since it carries the same token prices without the omni-modal input path.
- qwen/qwen3.8-maxstronger —Choose Qwen3.8 Max when peak text capability matters more than omni-modal input.
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 0.13 |
| Output | $ 0.43 |
| Cache read | $ 0.016 |
| Cache write | $ 0.13 |
+ 2.0K × $ 0.43
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 131.1K tokens
- Input / output modalities
- Text + Image + Audio + Video → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· can be turned off
- Reasoning effort
- low · medium · xhigh
- Thinking budget
- 0 – 0 tokens
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
- Structured output (JSON) is unavailable when thinking mode is on source
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- 91.6%
- 81.8%
- DeepSWE 1.1resolved_rate·Qwen Team · 2026-09-18setup not fully disclosed
官方在 Claude Code 与 mini-SWE-agent 两个框架下评测并取较高分,故不写单一 harness;上下文窗口 256K。口径与 qwen/qwen3.8-flash 那条相同。
57.8% - 92.6%
- SWE-Bench Multilingualresolved_rate·Qwen Team · 2026-09-18mini-swe-agentsetup not fully disclosed
上下文窗口 256K。
80.5% - SWE-Bench Proresolved_rate·Qwen Team · 2026-09-18Claude Codesetup not fully disclosed
已修正存在问题的任务并在修正后的基准上重测全部基线;上下文窗口 256K。
63.3%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is Qwen3.8 Omni Flash?
How much does Qwen3.8 Omni Flash cost?
Does Qwen3.8 Omni Flash support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.