Pick this when
Frontier score at a mid-tier bill. Live X and web. Knowledge-work tasks like GDPval.
Skip when
Prompts past 200K, or terminal agents if your own eval disagrees with the public split.
Ties Sol on the Intelligence Index at a fifth of Sol's output price. Leads the current Agentic Index snapshot. Consumer Grok apps were still rolling out at publish time.
Price
- Input / 1M
- $2.00
- Output / 1M
- $6.00
- Measured task
- $0.84
Rate doubles above ~200K tokens to $4 / $12 for the whole request.
Shape
- Context
- 500K
- Max output
- 64K
- Released
- 2026-08-12
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index61
GPQA Diamond91
Terminal-Bench 2.188
Agentic index59
- Arena
- —
- Speed
- 66/s
- GPQA
- 91.2%
- Term.
- 88.4%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
SpaceXAI · Palo Alto / Memphis. Grok. Wired into X and the live web. Grok 4.6 sits on the frontier scoreboard at a mid-tier price. Context is 500K, half of the pack.