Pick this when
The strongest model you can actually download. Long autonomous coding runs.
Skip when
You only care about the API invoice. GLM-5.3 is cheaper for similar hosted work.
Leads open-weight GPQA. License allows self-host and fine-tune. Reselling access at scale needs a read of Moonshot's terms.
Price
- Input / 1M
- $3.00
- Output / 1M
- $15
- Measured task
- $0.84
Self-host to dodge the API bill. Cache-hit input is $0.30 / 1M on Moonshot's API.
Shape
- Context
- 1M
- Max output
- 64K
- Released
- 2026-07
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index60
GPQA Diamond94
Terminal-Bench 2.184
Agentic index41
- Arena
- 1816
- Speed
- 77/s
- GPQA
- 93.5%
- Term.
- 84.1%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
Moonshot · Beijing. Kimi. Built for long agentic coding runs. K3 is the strongest model you can download today. The license is Moonshot's own, not MIT.