Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Qwen3.8 Omni FlashQwen3.8 FlashQwen3.8 MaxChat with these 3
AttributeQwen3.8 Omni Flashqwen/qwen3.8-omni-flashQwen3.8 Flashqwen/qwen3.8-flashQwen3.8 Maxqwen/qwen3.8-max
Specs
Context window1M1M1M
Max output131K131K131K
Input modalitiesTextImageAudioVideoTextImageVideoTextImageVideo
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 0.13$ 0.13$ 1.88
Output$ 0.43$ 0.43$ 5.63
Cache read$ 0.016$ 0.016$ 0.23
Cache write$ 0.13$ 0.20$ 2.35
What it is good at
Best for
  • Audio and video agent workflows
  • Hour-long recordings
  • Meetings and long-horizon office work
  • High-concurrency workloads
  • Coding and cowork agents
  • Million-token context
  • Maximum reasoning
  • Multimodal coding agents
Published benchmarks
Terminal-Bench 2.1pass_rate——
  • 86.6%Claude Code · 10 runs(setup not fully disclosed)
Agents' Last Examscore—
  • 51.2%(setup not fully disclosed)
—
VoiceBenchscore
  • 91.6%(setup not fully disclosed)
——
MMAUscore
  • 81.8%(setup not fully disclosed)
——
OmniVideoBench (static)score
  • 63.4%(setup not fully disclosed)
——
DeepSWE 1.1resolved_rate
  • 57.8%(setup not fully disclosed)
  • 58.7%(setup not fully disclosed)
—
LiveCodeBench v6pass_rate
  • 92.6%(setup not fully disclosed)
  • 91.9%(setup not fully disclosed)
—
SWE-Bench Multilingualresolved_rate
  • 80.5%mini-swe-agent(setup not fully disclosed)
  • 81%mini-swe-agent(setup not fully disclosed)
—
SWE-Bench Proresolved_rate
  • 63.3%Claude Code(setup not fully disclosed)
  • 62.5%Claude Code(setup not fully disclosed)
  • 67.7%Claude Code(setup not fully disclosed)
AndroidWorldsuccess_rate
  • 87.1%(setup not fully disclosed)
  • 84.5%(setup not fully disclosed)
—
OSWorld Verifiedsuccess_rate——
  • 86.1%(setup not fully disclosed)
MRCR v2 256K (8-needle)score——
  • 92.9%(setup not fully disclosed)
WildClawBench-MMscore
  • 71%Claude Code(setup not fully disclosed)
——
CharXiv Reasoning (no tools)accuracy
  • 83.5%(setup not fully disclosed)
  • 84.6%(setup not fully disclosed)
—
MMMU Proaccuracy——
  • 82.3%(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 36.5%(setup not fully disclosed)
  • 35.9%(setup not fully disclosed)
  • 43.6%(setup not fully disclosed)
PaperBench (BasicAgent)score——
  • 93%3 runs(setup not fully disclosed)
GPQA Diamondaccuracy
  • 91%(setup not fully disclosed)
  • 91.7%(setup not fully disclosed)
  • 92.6%(setup not fully disclosed)
Toolathlon Verifiedpass_ratesetups differ by: pass_k—
  • 73.5%(setup not fully disclosed)
  • 72.5%(setup not fully disclosed)
VideoMMMUaccuracy——
  • 88.7%(setup not fully disclosed)