Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flash
Back to model catalog

Qwen · Model guide

Qwen3.8 Omni Flash

Compare
qwen/qwen3.8-omni-flashReleased Sep 18, 2026

Qwen's latest natively omni-modal model, built on the Qwen3.8-Flash-Next architecture. It takes text, image, audio and video into a single one-million-token context and answers in text, with adjustable reasoning, function calling and context caching. Qwen aims it at agentic work where sound and video are the raw material: video editing, music-video creation, film commentary, meetings transcribed with speaker attribution, and recordings turned into structured notes.

Modalities
TextImageAudioVideoText
Context
1M
Max output
131K
Price / 1M tokens
$ 0.13in$ 0.43out
StreamingTool callingJSON outputImage inputThinking mode · can be turned off

What this model is good at

What the vendor positions it for, and which of our models to reach for instead.

VendorQwen says the model's core goal is to strengthen agent ability in real productivity settings, pushing omni-modal models from understanding omni-modal content towards planning tasks, calling tools and producing finished work.source

Best for
  • Audio and video agent workflows · Qwen reports a 36.5-point gain on WildClawBench-MM and 22.3 on AgenticVBench over the previous generation, and ships plugins for video editing, dubbing and film commentary pipelines.(vendor claim)source

  • Hour-long recordings · A one-million-token context with up to 991K input, a 262K thinking budget, and native support for audio-visual input up to one hour long.source

  • Meetings and long-horizon office work · Joint audio and visual speaker identification cuts AliMeeting diarisation error from 88.1 to 3.4, and the model scores 75.3 on Qwen's internal CoWorkBench office benchmark.(vendor claim)source

Neighbouring models
  • qwen/qwen3.8-flashsibling —Choose Qwen3.8 Flash when the inputs are only text, images and video, since it carries the same token prices without the omni-modal input path.
  • qwen/qwen3.8-maxstronger —Choose Qwen3.8 Max when peak text capability matters more than omni-modal input.
Compare them side by side →

Pricing and billing

Billed per token. Cached input is charged at the cache-read rate.

Price / 1M tokensYou pay
Input$ 0.13
Output$ 0.43
Cache read$ 0.016
Cache write$ 0.13
Estimate a request
$ 0.0022
Estimated cost per request
10.0K × $ 0.13
+ 2.0K × $ 0.43
≈ $ 2.16 per 1,000 requests

Capabilities and limits

What CrossModel guarantees across every route this model can take right now.

Context window
1.0M tokens
Max output
131.1K tokens
Input / output modalities
Text + Image + Audio + Video → Text
Streaming
Supported
Tool calling
Supported
Structured output (JSON)
Supported
Image input
Supported
Thinking mode
Supported· can be turned off
Reasoning effort
low · medium · xhigh
Thinking budget
0 – 0 tokens
Available endpoints
/v1/chat/completions · /v1/responses · /v1/messages
Conditional limitsDisclosed by the vendor. These do not make the capability unavailable — they say when it is not.
  • Structured output (JSON) is unavailable when thinking mode is on source

Published benchmarks

Scores the vendor reported, with the evaluation setup each one came from.

Audio interaction
  • VoiceBench
    score·Qwen Team · 2026-09-18setup not fully disclosed

    官方评测表未给出该项的评测设置,也未公开 metric 定义。

    91.6%
Audio understanding
  • MMAU
    score·Qwen Team · 2026-09-18setup not fully disclosed

    官方评测表未给出该项的评测设置,也未公开 metric 定义。

    81.8%
Audio visual reasoning
  • OmniVideoBench (static)
    score·Qwen Team · 2026-09-18setup not fully disclosed

    取 Static 模式(直接理解输入内容);接入 Qwen Code 的 Agentic 模式为 67.8 并把单查询 token 从 145,736 降到 79,117,属另一种设置,不并入本条。官方未公开 metric 定义。

    63.4%
Coding
  • DeepSWE 1.1
    resolved_rate·Qwen Team · 2026-09-18setup not fully disclosed

    官方在 Claude Code 与 mini-SWE-agent 两个框架下评测并取较高分,故不写单一 harness;上下文窗口 256K。口径与 qwen/qwen3.8-flash 那条相同。

    57.8%
  • LiveCodeBench v6
    pass_rate·Qwen Team · 2026-09-18setup not fully disclosed

    官方评测表未给出该项的评测设置。

    92.6%
  • SWE-Bench Multilingual
    resolved_rate·Qwen Team · 2026-09-18mini-swe-agentsetup not fully disclosed

    上下文窗口 256K。

    80.5%
  • SWE-Bench Pro
    resolved_rate·Qwen Team · 2026-09-18Claude Codesetup not fully disclosed

    已修正存在问题的任务并在修正后的基准上重测全部基线;上下文窗口 256K。

    63.3%
Computer use
  • AndroidWorld
    success_rate·Qwen Team · 2026-09-18setup not fully disclosed

    官方评测表未给出该项的评测设置。同表 Qwen3.8-Flash 记 84.5,与目录里那条一致。

    87.1%
Multimodal
  • WildClawBench-MM
    score·Qwen Team · 2026-09-18Claude Codesetup not fully disclosed

    只评测 WildClawBench 中涉及图像、视频或音频的多模态任务;上一代 Qwen3.5-Omni-Plus 记 34.5。官方未公开 metric 定义,故按保守的 score 口径记录。

    71%
Multimodal reasoning
  • CharXiv Reasoning (no tools)
    accuracy·Qwen Team · 2026-09-18setup not fully disclosed

    取官方「不含 CI」(不启用代码解释器)一列;启用代码解释器时为 91.4,属另一种设置,不并入本条。评测使用固定 prompt。

    83.5%
Reasoning
  • Humanity's Last Exam
    accuracy·Qwen Team · 2026-09-18setup not fully disclosed

    由 GPT-4o 判分;官方未说明是否使用工具,故按目录内不带工具的默认身份记录。

    36.5%
Scientific reasoning
  • GPQA Diamond
    accuracy·Qwen Team · 2026-09-18setup not fully disclosed

    官方评测表未给出该项的评测设置。

    91%

Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.

Use it in your tools

Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.

See all integration guides →

Frequently asked questions

What is Qwen3.8 Omni Flash?
Qwen3.8 Omni Flash is available on CrossModel as qwen/qwen3.8-omni-flash. It has a 1M-token context window and can return up to 131K tokens per request.
How much does Qwen3.8 Omni Flash cost?
Currently $ 0.13 per 1M input tokens and $ 0.43 per 1M output tokens. Cached input is billed at $ 0.016 per 1M.
Does Qwen3.8 Omni Flash support tool calling and structured output?
Tool calling: Supported · Structured output (JSON): Supported · Image input: Supported · Streaming: Supported · Thinking mode: Supported
Which endpoint do I call?
For text models, OpenAI Chat Completions (/v1/chat/completions), OpenAI Responses (/v1/responses) and Anthropic Messages (/v1/messages) all work — pick whichever your SDK already speaks.
Can I try it without writing code?
Yes. Sign in and open the Playground on this page to chat with the model directly — no API key needed. Usage is billed from your wallet at the normal rate.

Chat first, integrate later

Chat first, then wire it in once you like the answers.