Pick this when
Teams already on 4.5 who have not moved. Still cheap for the score.
Skip when
You can call 4.6. Same price, clearly stronger.
Previous Grok flagship. Keep it only if a pin or a region has not caught up.
Price
- Input / 1M
- $2.00
- Output / 1M
- $6.00
- Measured task
- $0.92
Same 200K pricing cliff as 4.6.
Shape
- Context
- 500K
- Max output
- 64K
- Released
- 2026-07
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index56
GPQA Diamond86
Terminal-Bench 2.171
Agentic index44
- Arena
- —
- Speed
- 72/s
- GPQA
- 86.4%
- Term.
- 71.2%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
SpaceXAI · Palo Alto / Memphis. Grok. Wired into X and the live web. Grok 4.6 sits on the frontier scoreboard at a mid-tier price. Context is 500K, half of the pack.