Skip to catalog
FbFable
CatalogLabsScoresCompare
Catalog
K

Moonshot

Kimi K3

kimi-k3

60Index
Other · restrictedGenerally available1M context2.8TTextImage

Pick this when

The strongest model you can actually download. Long autonomous coding runs.

Skip when

You only care about the API invoice. GLM-5.3 is cheaper for similar hosted work.

Leads open-weight GPQA. License allows self-host and fine-tune. Reselling access at scale needs a read of Moonshot's terms.

Price

Input / 1M
$3.00
Output / 1M
$15
Measured task
$0.84

Self-host to dodge the API bill. Cache-hit input is $0.30 / 1M on Moonshot's API.

Shape

Context
1M
Max output
64K
Released
2026-07
Tools
Yes
Reasoning
Yes

Scores

Intelligence index60
GPQA Diamond94
Terminal-Bench 2.184
Agentic index41
Arena
1816
Speed
77/s
GPQA
93.5%
Term.
84.1%
  • AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
  • GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
  • Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
  • Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.

Moonshot · Beijing. Kimi. Built for long agentic coding runs. K3 is the strongest model you can download today. The license is Moonshot's own, not MIT.

Compiled 2026-08-28 from public Artificial Analysis, LMSYS-style arena, and first-party price sheets. Scores move. Recheck before you spend.

A two-point gap on the intelligence index is noise for most work. Cost per finished task is the number that shows up on the invoice.