Pick this when
The Claude default. Drafting, analysis, coding, and anything you will actually read.
Skip when
You need the top of the board, or you are optimizing purely for token cost.
Most teams should start here, not at Opus. Style and caution are the reason people stay on Claude after the scores get close.
Price
- Input / 1M
- $2.00
- Output / 1M
- $10
- Measured task
- $1.72
Promotional $2/$10 was made permanent on 10 Aug 2026. Flat through 1M context.
Shape
- Context
- 1M
- Max output
- 64K
- Released
- 2026-05
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index55
GPQA Diamond84
Terminal-Bench 2.172
Agentic index34
- Arena
- 1588
- Speed
- 75/s
- GPQA
- 84.2%
- Term.
- 72.4%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
Anthropic · San Francisco. Claude. Opus 5 leads public intelligence rankings. Fable 5 is the named flagship and the expensive one. Sonnet 5 is the daily driver.