Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Gemini 3 Flash PreviewGemini 3.1 Pro PreviewGemini 3.6 FlashChat with these 3
AttributeGemini 3 Flash Previewgemini/gemini-3-flash-previewGemini 3.1 Pro Previewgemini/gemini-3.1-pro-previewGemini 3.6 Flashgemini/gemini-3.6-flash
Specs
Context window1M1M1M
Max output66K66K66K
Input modalitiesTextImageAudioVideoTextImageAudioVideoTextImageAudioVideo
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 0.50$ 2.00>= 200K$ 4.00$ 1.50
Output$ 3.00$ 12.00>= 200K$ 18.00$ 7.50
Cache read$ 0.050$ 0.20>= 200K$ 0.40$ 0.15
Cache write$ 0.50$ 2.00>= 200K$ 4.00$ 1.50
What it is good at
Best for
  • Responsive coding agents
  • Interactive multimodal analysis
  • Advanced coding agents
  • Large multimodal corpora
  • Rapid agentic coding loops
  • Efficient knowledge work
  • Spatial and multimodal reasoning
Published benchmarks
Terminal-Bench 2.0pass_rate
  • 68.5%effort high · Terminus 2(setup not fully disclosed)
Terminal-Bench 2.1pass_ratesetups differ by: reasoning
  • 58%effort high · Terminus 2(setup not fully disclosed)
  • 78%Terminus 2(setup not fully disclosed)
BrowseComp (with tools)accuracy
  • 85.9%effort high · tools: browse, python, search(setup not fully disclosed)
SWE-Bench Pro (Public)pass_rate
  • 58.7%(setup not fully disclosed)
SWE-Bench Verifiedresolved_rate
  • 78%effort high(setup not fully disclosed)
  • 80.6%effort high(setup not fully disclosed)
OSWorld-Verifiedsuccess_rate
  • 83%(setup not fully disclosed)
GDPval-AA v2elo
  • 1421(setup not fully disclosed)
CharXiv Reasoning (no tools)accuracy
  • 85.2%(setup not fully disclosed)
MMMU Pro (no tools)accuracy
  • 81.2%effort high(setup not fully disclosed)
Humanity's Last Examaccuracy
  • 33.7%effort high(setup not fully disclosed)
  • 44.4%effort high(setup not fully disclosed)
GPQA Diamondaccuracy
  • 90.4%effort high(setup not fully disclosed)
  • 94.3%effort high(setup not fully disclosed)