Skip to catalog
FbFable
CatalogLabsScoresCompare
Catalog
X

SpaceXAI

Grok 4.6

grok-4.6

61Index
ClosedGenerally available500K contextTextImage

Pick this when

Frontier score at a mid-tier bill. Live X and web. Knowledge-work tasks like GDPval.

Skip when

Prompts past 200K, or terminal agents if your own eval disagrees with the public split.

Ties Sol on the Intelligence Index at a fifth of Sol's output price. Leads the current Agentic Index snapshot. Consumer Grok apps were still rolling out at publish time.

Price

Input / 1M
$2.00
Output / 1M
$6.00
Measured task
$0.84

Rate doubles above ~200K tokens to $4 / $12 for the whole request.

Shape

Context
500K
Max output
64K
Released
2026-08-12
Tools
Yes
Reasoning
Yes

Scores

Intelligence index61
GPQA Diamond91
Terminal-Bench 2.188
Agentic index59
Arena
—
Speed
66/s
GPQA
91.2%
Term.
88.4%
  • AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
  • GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
  • Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
  • Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.

SpaceXAI · Palo Alto / Memphis. Grok. Wired into X and the live web. Grok 4.6 sits on the frontier scoreboard at a mid-tier price. Context is 500K, half of the pack.

Compiled 2026-08-28 from public Artificial Analysis, LMSYS-style arena, and first-party price sheets. Scores move. Recheck before you spend.

A two-point gap on the intelligence index is noise for most work. Cost per finished task is the number that shows up on the invoice.