DeepSeek · Model guide
DeepSeek V4.1 Flash
DeepSeek's newest Flash model and the first of its V4.1 architecture series: natively multimodal, a 1M-token context window, up to 384K output tokens, and a thinking mode that can be turned off or dialled between low, high and max. DeepSeek reports it has surpassed V4 Pro on capability, speed and end-to-end latency, and it is also the model now serving the retired V4 Flash model names.
What this model is good at
What the vendor positions it for, and which of our models to reach for instead.
VendorDeepSeek describes it as the smallest model in a new architecture series designed for a higher capability ceiling, faster inference and greater throughput, and says it has surpassed V4 Pro across performance, speed and total elapsed time.source
Terminal and coding agents · The vendor reports Terminal-Bench 2.1 at 90.6 and DeepSWE 1.1 at 74.2.source
Agents that read screenshots and charts · Image understanding is native to this model rather than a separate experimental variant, and one request can carry up to 600 images.source
Hard knowledge and research · The vendor reports GPQA Diamond at 90.9, and Humanity's Last Exam at 63.9 when tools are allowed.source
High-throughput production traffic · DeepSeek documents a concurrency limit of 2500 for the flash tier, five times the V4 Pro limit.source
- deepseek/deepseek-v4-propredecessor —V4 Pro is still a distinct model until 2026-09-14 12:00 Beijing time; after that DeepSeek routes deepseek-v4-pro requests to V4.1 Flash.source
- deepseek/deepseek-v4-flashpredecessor —V4 Flash and V4 Flash Vision Exp are retired upstream; those model ids are kept as compatibility aliases and are served by V4.1 Flash.source
Pricing and billing
Billed per token. Cached input is charged at the cache-read rate.
| Price / 1M tokens | You pay |
|---|---|
| Input | $ 0.30$ 0.27 |
| Output | $ 1.20$ 1.08 |
| Cache read | $ 0.0060$ 0.0054 |
| Cache write | $ 0.30$ 0.27 |
+ 2.0K × $ 1.08
Capabilities and limits
What CrossModel guarantees across every route this model can take right now.
- Context window
- 1.0M tokens
- Max output
- 384.0K tokens
- Input / output modalities
- Text + Image → Text
- Streaming
- Supported
- Tool calling
- Supported
- Structured output (JSON)
- Supported
- Image input
- Supported
- Thinking mode
- Supported· can be turned off
- Reasoning effort
- low · medium · high · xhigh · max
- Available endpoints
- /v1/chat/completions · /v1/responses · /v1/messages
Published benchmarks
Scores the vendor reported, with the evaluation setup each one came from.
- Terminal-Bench 2.1pass_rate·DeepSeek · 2026-09-10setup not fully disclosed
厂商自测。来源页这一轮未重述 harness 与采样设置(V4 系列此前统一用 DeepSeek Harness 极简模式、max 档位),所以 settings 留空;不可与第三方独立运行的同名成绩直接互换。
90.6% - Terminal-Bench 4.0pass_rate·DeepSeek · 2026-09-10setup not fully disclosed
厂商自测。来源页同时给出 Terminal-Bench 3.0 的 30.0;不同版本是不同身份,此处只录最新一版。来源页未披露该项的评测设置。
31.2%
Vendor-reported numbers, not CrossModel measurements. Scores are only comparable when the benchmark version, metric and evaluation setup match, so nothing here is averaged or ranked.
Use it in your tools
Point the base URL at CrossModel and paste this model ID — every tool below has a setup guide.
Frequently asked questions
What is DeepSeek V4.1 Flash?
How much does DeepSeek V4.1 Flash cost?
Does DeepSeek V4.1 Flash support tool calling and structured output?
Which endpoint do I call?
Can I try it without writing code?
Chat first, integrate later
Chat first, then wire it in once you like the answers.