Grok 4.6
Aug 2026 · Proprietary · Current · supported evidence
Ties GPT-5.6 Sol on intelligence at $2 / $6. Strong agents and GPQA. No public SWE-bench; 500K context is the ceiling.
Intelligence
60.9
AA Index
API price
$2 / $6
Input / output per 1M
Context
500K
58 tok/s
Capability
Intelligence Index60.9
Coding index76.8
Agentic index58.7
Arena Elo (offset)164
Use it when
- ▸Frontier quality at mid price
- ▸Agents
- ▸Coding assistants
Skip if
- –You must have 1M context or a published SWE-bench Verified
Benchmarks
GPQA Diamond Graduate-level science questions designed so Google search is not enough. Still one of the cleanest knowledge/reasoning splits. | 94.9% |
|---|---|
Humanity’s Last Exam Expert-written questions across fields. Harder than MMLU; the current differentiator for “does this model actually know things.” | 42.9% |
MMLU-Pro Harder, less-saturated successor to MMLU. Classic MMLU is above 90% for every flagship and no longer ranks the field. | — |
SWE-bench Verified 500 human-validated GitHub issues. Score swings 5–15 points by harness — treat vendor numbers as an upper bound. | — |
SWE-bench Pro Harder, contamination-resistant coding eval. Currently the best public split between “can code” and “can maintain a repo.” | — |
Terminal-Bench 2.1 End-to-end tasks in a real terminal. Better proxy for coding agents than HumanEval, which is fully saturated. | 88.4% |
ARC-AGI-2 Abstract visual puzzles. Rewards generalization over memorization. GPT-5.6 Sol currently leads the published set. | 67.1% |
AIME 2025 American Invitational Mathematics Examination. Contest math; reasoning-mode models dominate. | — |
MMMU College-level multimodal understanding across diagrams, charts, and exam figures. | — |
Served via the xAI / SpaceXAI API. Reasoning effort (low/medium/high/xhigh) moves the score by several points.
Modalities
text · vision · tools
Reasoning mode
reasoning
Hybrid and reasoning models spend tokens thinking. That raises GPQA and agents, and also raises latency and bill.