Previous flagship — now value-priced, still excellent.
Released
Oct 20, 2025
Context
400K
Pricing
$1.50/$6.00
per 1M tok
Best at
Reliable reasoning
Performance
Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.
Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.
American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.
Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.
Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.
Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).
Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.
Value score: 66/100
Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.
More from
Data last checked Jun 1, 2026 (2mo ago)