Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Kimi K2.5Kimi K2.6Chat with these 2
AttributeKimi K2.5moonshot/kimi-k2.5Kimi K2.6moonshot/kimi-k2.6
Specs
Context window262K262K
Max output262K262K
Input modalitiesTextImageVideoTextImageVideo
Output modalitiesTextText
Capabilities
StreamingSupportedSupported
Tool callingSupportedSupported
Structured output (JSON)SupportedSupported
Image inputSupportedSupported
Thinking modeSupportedSupported
Price / 1M tokens
Input$ 0.62$ 1.00
Output$ 3.30$ 4.16
Cache read$ 0.11$ 0.18
Cache write$ 0.62$ 1.00
What it is good at
Best for
  • Visual-to-code
  • Agentic search
  • Agentic coding
  • Tool-assisted research
  • Multimodal workflows
Published benchmarks
Terminal-Bench 2.0pass_rate
  • 66.7%Terminus 2(setup not fully disclosed)
BrowseComp (with context management)accuracy
  • 74.9%thinking · tools: browser, code_interpreter, search · discard-all(setup not fully disclosed)
BrowseComp (with tools)accuracy
  • 83.2%thinking(setup not fully disclosed)
DeepSearchQAf1
  • 92.5%(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 58.6%(setup not fully disclosed)
SWE-Bench Verifiedresolved_ratesetups differ by: harness, reasoning, runs
  • 76.8%non-thinking · moonshot-internal-agent · 5 runs(setup not fully disclosed)
  • 80.2%(setup not fully disclosed)
OSWorld Verifiedsuccess_rate
  • 73.1%(setup not fully disclosed)
MMMU Proaccuracy
  • 79.4%(setup not fully disclosed)
Humanity's Last Exam Full (with tools)accuracy
  • 50.2%thinking · tools: browser, code_interpreter, search · retain-latest-tool-round(setup not fully disclosed)
Humanity's Last Exam Full (with tools)accuracy
  • 54%thinking · tools: browser, code_interpreter, search(setup not fully disclosed)
GPQA Diamondaccuracy
  • 90.5%thinking(setup not fully disclosed)