Pick this when
Flagship agents and coding. Dead heat with Opus on hard jobs, fewer tokens to finish them.
Skip when
Huge documents past 272K, or a budget that cannot eat $30 output.
Leads GPQA Diamond. On DeepSWE it ties Opus while using about half the output tokens and fewer agent steps. Broadest native tool surface.
Price
- Input / 1M
- $5.00
- Output / 1M
- $30
- Measured task
- $1.23
Above 272K input the whole request doubles: $10 / $45.
Shape
- Context
- 1.1M
- Max output
- 128K
- Released
- 2026-07
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index61
GPQA Diamond95
Terminal-Bench 2.188
Agentic index58
- Arena
- 2134
- Speed
- 62/s
- GPQA
- 94.6%
- Term.
- 88.0%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
OpenAI · San Francisco. ChatGPT's parent. The GPT-5.6 line is a three-rung ladder: Sol for hard work, Terra for most production traffic, Luna for volume.