Pick this when
Efficient Qwen coding when Max is more model than you want to run.
Skip when
You need Max's multimodal stack or a million-token window.
High efficiency, smaller window. A workhorse, not a flagship.
Price
- Input / 1M
- $0.15
- Output / 1M
- $0.60
- Measured task
- $0.14
New efficiency tier. Hosted rates still settling.
Shape
- Context
- 256K
- Max output
- 32K
- Released
- 2026-08
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index50
GPQA Diamond81
Terminal-Bench 2.158
Agentic index37
- Arena
- —
- Speed
- 160/s
- GPQA
- 80.6%
- Term.
- 58.4%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
Qwen · Hangzhou. Two flagships: Qwen3.8-Max is a proprietary multimodal API. The 2.4T checkpoint is downloadable, text-only, and commercially restricted.