Meta's largest open-weights model, near-frontier quality.
Released
Apr 29, 2026
Context
1M
Pricing
Open
per 1M tok
Best at
Open weights
Performance
Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.
Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.
American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.
Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.
Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.
Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).
Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.
Value score: 82/100
Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.
More from
Data last checked Jun 9, 2026 (1mo ago)