Kimi K2 Thinking
Moonshot AINov 6, 2025#3
Open weights
Late-2025 open-weight milestone that ran 200–300 sequential tool calls reliably and briefly held the open-model crown on HLE (~45%). Superseded within Moonshot's own lineup, still excellent as a self-hosted reasoner.
- Status
- Generally available
- License
- Open weights · Custom
- Context window
- 256K
- Max output
- 64K
- Input · $/M
- $0.6
- Output · $/M
- $2.50
- Cached input · $/M
- $0.15
- Blended · $/M
- $1.07
- Parameters
- ~1T MoE
- Inputs
- Text only
Benchmarks
Composite 70.2· mean of published scores
Reasoning & knowledge
GPQA Diamond
no public score
MMLU-Pro
no public score
Humanity's Last ExamLeader
44.9%
ARC-AGI-2
no public score
Coding
SWE-bench Verified
71.3%
dataset best 93.4SWE-bench Pro
no public score
LiveCodeBench
no public score
Agents & computer use
Terminal-Bench 2
no public score
OSWorld 2.0
no public score
BrowseComp
no public score
Math
AIME 2026Leader
94.5%
Known strengths
- Interleaved tool-use thinking
- Self-hostable at scale