Leads on reasoning with a native 2M-token context.
Released
May 20, 2026
Context
2M
Pricing
$2.50/$10
per 1M tok
Best at
Long-context king
Performance
Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.
Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.
American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.
Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.
Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.
Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).
Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.
Value score: 77/100
Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.
More from
Data last checked Jun 14, 2026 (1mo ago)