Pick this when
Near-frontier reasoning you can download under MIT. Best open permissive option.
Skip when
You need multimodal input or the last few intelligence points.
The first open-weight model that looks like it belongs on the frontier table. Slower than the closed leaders. The license is the point.
Price
- Input / 1M
- $1.32
- Output / 1M
- $3.96
- Measured task
- $0.25
First-party API. Other hosts and self-host will bill differently. MIT weights.
Shape
- Context
- 1M
- Max output
- 64K
- Released
- 2026-08-13
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index53
GPQA Diamond87
Terminal-Bench 2.179
Agentic index50
- Arena
- —
- Speed
- 78/s
- GPQA
- 86.8%
- Term.
- 78.7%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
DeepSeek · Hangzhou. Open weights under MIT. V4 Pro is the first downloadable model that actually sits in the frontier cluster. Flash is the volume lane.