← Back to explorer
SpaceXAIReleased 2026-07

Cursor Grok 4.5

SpaceXAI’s frontier model, trained jointly with Cursor for coding, agents, and knowledge work. Scored 54 and ranked #4 on the Artificial Analysis Intelligence Index in the source app’s dated launch snapshot — the biggest Grok generation jump yet — with standout cost efficiency and token thrift.

Context
500K
Speed
Price /M
$2 / $6
Access
api

When to use it

Coding agents in Cursor / Grok Build, long-running engineering tasks, and near-frontier intelligence at a fraction of Opus/GPT Pro spend.

Watch-outs

500K context (down from Grok 4.3’s 1M). Hallucination rate rose with knowledge gains on AA-Omniscience. Prefer max-effort Claude/OpenAI when absolute peak coding ceiling matters more than $/task.

CodingAgentsWritingBest valueScience

Benchmarks

Arena Elo

Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.

AA Intelligence Index54

Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).

SWE-bench Verified

Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.

GPQA Diamond

PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.

MMLU-Pro

Harder multi-choice knowledge/reasoning suite than classic MMLU.

Humanity's Last Exam

Frontier closed-ended academic difficulty across many domains.

Model Gauge · built by Cursor Grok 4.5 High · curated snapshot, not a live leaderboard

Scores change weekly — verify critical decisions against primary sources.