Cursor Grok 4.5
SpaceXAI’s frontier model, trained jointly with Cursor for coding, agents, and knowledge work. Scored 54 and ranked #4 on the Artificial Analysis Intelligence Index in the source app’s dated launch snapshot — the biggest Grok generation jump yet — with standout cost efficiency and token thrift.
When to use it
Coding agents in Cursor / Grok Build, long-running engineering tasks, and near-frontier intelligence at a fraction of Opus/GPT Pro spend.
Watch-outs
500K context (down from Grok 4.3’s 1M). Hallucination rate rose with knowledge gains on AA-Omniscience. Prefer max-effort Claude/OpenAI when absolute peak coding ceiling matters more than $/task.
Benchmarks
Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.
Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).
Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.
PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.
Harder multi-choice knowledge/reasoning suite than classic MMLU.
Frontier closed-ended academic difficulty across many domains.