Pick this when
Multimodal and multilingual work on a budget, especially Chinese and mixed media.
Skip when
Speed is the constraint. 47 tok/s is slow for a chat UI.
Capable, not fast. The API is closed even though Qwen is known for open releases.
Price
- Input / 1M
- $2.00
- Output / 1M
- $6.00
- Measured task
- $1.13
Proprietary API. Separate open checkpoint exists under a different license.
Shape
- Context
- 1M
- Max output
- 64K
- Released
- 2026-08
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index58
GPQA Diamond89
Terminal-Bench 2.175
Agentic index40
- Arena
- —
- Speed
- 47/s
- GPQA
- 89.4%
- Term.
- 75.2%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
Qwen · Hangzhou. Two flagships: Qwen3.8-Max is a proprietary multimodal API. The 2.4T checkpoint is downloadable, text-only, and commercially restricted.