Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

DeepSeek V4.1 FlashDeepSeek V4 ProDeepSeek V4 FlashChat with these 3
AttributeDeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash10% offDeepSeek V4 Prodeepseek/deepseek-v4-pro10% offDeepSeek V4 Flashdeepseek/deepseek-v4-flash10% off
Specs
Context window1M1M1M
Max output384K384K384K
Input modalitiesTextImageTextText
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedNot supportedNot supported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 0.30$ 0.27$ 1.35$ 1.22$ 0.30$ 0.27
Output$ 1.20$ 1.08$ 4.05$ 3.65$ 1.20$ 1.08
Cache read$ 0.0060$ 0.0054$ 0.045$ 0.041$ 0.0060$ 0.0054
Cache write$ 0.30$ 0.27$ 1.35$ 1.22$ 0.30$ 0.27
What it is good at
Best for
  • Terminal and coding agents
  • Agents that read screenshots and charts
  • Hard knowledge and research
  • High-throughput production traffic
  • Complex long-horizon agents
  • Research and hard reasoning
  • Cost-sensitive coding agents
  • High-volume tool workflows
Published benchmarks
Terminal-Bench 2.1pass_ratesetups differ by: reasoning
  • 90.6%(setup not fully disclosed)
  • 78.7%effort max(setup not fully disclosed)
Terminal-Bench 4.0pass_rate
  • 31.2%(setup not fully disclosed)
Agents' Last Examscore
  • 31.8%(setup not fully disclosed)
AutomationBenchscore
  • 54.8%(setup not fully disclosed)
DeepSWE 1.1resolved_rate
  • 74.2%(setup not fully disclosed)
NL2Reposcore
  • 65.4%(setup not fully disclosed)
SciCodescore
  • 49.2%effort max(setup not fully disclosed)
CyberGymscore
  • 88.1%(setup not fully disclosed)
GDPval-AA v2elo
  • 1590.3effort max(setup not fully disclosed)
AA-LCRscore
  • 75.3%effort max(setup not fully disclosed)
MathArena Apexscore
  • 65.6%(setup not fully disclosed)
Chartography (with tools)score
  • 78.9%(setup not fully disclosed)
AA Intelligence Index v4.1.1score
  • 53effort max(setup not fully disclosed)
Humanity's Last Examaccuracysetups differ by: reasoning
  • 36.8%(setup not fully disclosed)
  • 39.3%effort max(setup not fully disclosed)
Humanity's Last Exam (with tools)accuracy
  • 63.9%(setup not fully disclosed)
CritPtscore
  • 18%effort max(setup not fully disclosed)
GPQA Diamondaccuracysetups differ by: reasoning
  • 90.9%(setup not fully disclosed)
  • 92.8%effort max(setup not fully disclosed)
tau3-Bankingscore
  • 39.6%effort max(setup not fully disclosed)