Pick this when
Two-million-token prompts: dumps, logs, multi-doc review, cheap long context.
Skip when
You need frontier reasoning. Fast is a window, not a brain.
The context number is real. Effective use of the far end of a 2M window still varies. Test on your own documents.
Price
- Input / 1M
- $0.20
- Output / 1M
- $0.50
- Measured task
- $0.18
Largest practical context window in this catalog.
Shape
- Context
- 2M
- Max output
- 32K
- Released
- 2026-06
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index48
GPQA Diamond72
Terminal-Bench 2.155
Agentic index25
- Arena
- —
- Speed
- 210/s
- GPQA
- 72.1%
- Term.
- 54.8%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
SpaceXAI · Palo Alto / Memphis. Grok. Wired into X and the live web. Grok 4.6 sits on the frontier scoreboard at a mid-tier price. Context is 500K, half of the pack.