Back to model catalog

Compare models

Context, limits, capabilities, price and use cases side by side — up to 4 models at a time. The URL holds your selection, so you can share it.

Kimi K2.6Kimi K2.7 CodeKimi K3Chat with these 3
AttributeKimi K2.6moonshot/kimi-k2.6Kimi K2.7 Codemoonshot/kimi-k2.7-codeKimi K3moonshot/kimi-k3
Specs
Context window262K262K1M
Max output262K262K1M
Input modalitiesTextImageVideoTextImageVideoTextImageVideo
Output modalitiesTextTextText
Capabilities
StreamingSupportedSupportedSupported
Tool callingSupportedSupportedSupported
Structured output (JSON)SupportedSupportedSupported
Image inputSupportedSupportedSupported
Thinking modeSupportedSupportedSupported
Price / 1M tokens
Input$ 1.00$ 1.00$ 3.00
Output$ 4.16$ 4.16$ 15.00
Cache read$ 0.18$ 0.18$ 0.30
Cache write$ 1.00$ 1.00$ 3.00
What it is good at
Best for
  • Agentic coding
  • Tool-assisted research
  • Multimodal workflows
  • Repository-scale coding
  • Tool orchestration
  • Long-horizon engineering
  • Deep research
  • Computer use
Published benchmarks
Kimi Code Bench v2score
  • 62%Kimi Code CLI(setup not fully disclosed)
Terminal-Bench 2.0pass_rate
  • 66.7%Terminus 2(setup not fully disclosed)
Terminal-Bench 2.1pass_rate
  • 88.3%(setup not fully disclosed)
BrowseComp (with tools)accuracysetups differ by: reasoning
  • 83.2%thinking(setup not fully disclosed)
  • 91.2%(setup not fully disclosed)
DeepSearchQAf1
  • 92.5%(setup not fully disclosed)
  • 95%(setup not fully disclosed)
Kimi Claw 24/7score
  • 46.9%OpenClaw · 3 runs(setup not fully disclosed)
DeepSWEresolved_rate
  • 67.5%(setup not fully disclosed)
MLS Bench Litescore
  • 35.1%Kimi Code CLI(setup not fully disclosed)
ProgramBenchscoresetups differ by: harness
  • 53.6%Kimi Code CLI(setup not fully disclosed)
  • 77.8%(setup not fully disclosed)
SWE-Bench Proresolved_rate
  • 58.6%(setup not fully disclosed)
SWE-Bench Verifiedresolved_rate
  • 80.2%(setup not fully disclosed)
OSWorld Verifiedsuccess_rate
  • 73.1%(setup not fully disclosed)
  • 84.8%(setup not fully disclosed)
MMMU Proaccuracy
  • 79.4%(setup not fully disclosed)
Humanity's Last Exam Full (with tools)accuracysetups differ by: reasoning, tools
  • 54%thinking · tools: browser, code_interpreter, search(setup not fully disclosed)
  • 56%effort max(setup not fully disclosed)
GPQA Diamondaccuracysetups differ by: reasoning
  • 90.5%thinking(setup not fully disclosed)
  • 93.5%effort max(setup not fully disclosed)
MCP Atlasscore
  • 76%3 runs(setup not fully disclosed)
MCPMark Verifiedscoresetups differ by: max_output_tokens, runs
  • 81.1%3 runs(setup not fully disclosed)
  • 94.5%(setup not fully disclosed)