Skip to catalog
FbFable
CatalogLabsScoresCompare
Catalog
X

SpaceXAI

Grok 4.5

grok-4.5

56Index
ClosedGenerally available500K contextTextImage

Pick this when

Teams already on 4.5 who have not moved. Still cheap for the score.

Skip when

You can call 4.6. Same price, clearly stronger.

Previous Grok flagship. Keep it only if a pin or a region has not caught up.

Price

Input / 1M
$2.00
Output / 1M
$6.00
Measured task
$0.92

Same 200K pricing cliff as 4.6.

Shape

Context
500K
Max output
64K
Released
2026-07
Tools
Yes
Reasoning
Yes

Scores

Intelligence index56
GPQA Diamond86
Terminal-Bench 2.171
Agentic index44
Arena
—
Speed
72/s
GPQA
86.4%
Term.
71.2%
  • AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
  • GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
  • Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
  • Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.

SpaceXAI · Palo Alto / Memphis. Grok. Wired into X and the live web. Grok 4.6 sits on the frontier scoreboard at a mid-tier price. Context is 500K, half of the pack.

Compiled 2026-08-28 from public Artificial Analysis, LMSYS-style arena, and first-party price sheets. Scores move. Recheck before you spend.

A two-point gap on the intelligence index is noise for most work. Cost per finished task is the number that shows up on the invoice.