Pick this when
Frontier-class coding at a mid-tier API bill.
Skip when
You need multimodal input or weights you can host.
Ties Kimi K3 on the composite. Zhipu reports a lead on vulnerability spotting. Treat vendor security benches as claims until replicated.
Price
- Input / 1M
- $1.40
- Output / 1M
- $4.40
- Measured task
- $0.68
API only. Weights not released as of 28 Aug 2026.
Shape
- Context
- 1M
- Max output
- 64K
- Released
- 2026-08-18
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index60
GPQA Diamond90
Terminal-Bench 2.182
Agentic index41
- Arena
- —
- Speed
- 95/s
- GPQA
- 90.1%
- Term.
- 82.4%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
Z.AI · Beijing. GLM. 5.3 matches Kimi on the composite at a lower API price. Weights are not public, so this is still a rental.