Code-specialist tuned for fast IDE autocomplete.
Released
Jan 30, 2026
Context
256K
Pricing
Open
per 1M tok
Best at
Code specialist
Performance
Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.
Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.
American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.
Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.
Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.
Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).
Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.
Value score: 48/100
Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.
More from
Data last checked May 12, 2026 (2mo ago)