Grok 4.5
xAIJul 16, 2026N/A
Trained with Cursor on tens of thousands of GB300s for multi-step software engineering. Best published score on SWE Marathon (29.0%), strong Terminal-Bench (83.3%), and remarkable token efficiency — ~4× fewer tokens per task than Opus 4.8 max.
- Status
- Generally available
- License
- Proprietary API
- Context window
- 500K
- Max output
- 64K
- Input · $/M
- $2
- Output · $/M
- $6
- Cached input · $/M
- $0.5
- Blended · $/M
- $3.00
- Parameters
- Undisclosed
- Inputs
- Text + images
- Speed
- ≈80 tok/s serving
Benchmarks
Composite 74.0· mean of published scores
Reasoning & knowledge
GPQA Diamond
no public score
MMLU-Pro
no public score
Humanity's Last Exam
no public score
ARC-AGI-2
no public score
Coding
SWE-bench Verified
no public score
SWE-bench Pro
64.7%
dataset best 80LiveCodeBench
no public score
Agents & computer use
Terminal-Bench 2Leader
83.3%
OSWorld 2.0
no public score
BrowseComp
no public score
Math
AIME 2026
no public score
Known strengths
- SWE Marathon leader
- ~16k tokens/task average (highly efficient)
- Cheap flat $2/$6 pricing
Caveats
- Trails Fable 5 on most max-effort coding suites