About
ModelLab is an independent dashboard. Here's what the numbers mean, where they come from, and what they don't tell you.
Benchmark scores are the best publicly reported figures for each model, drawn from official model cards and papers, plus independent leaderboards: LMArena (Arena Elo), swebench.com (SWE-bench), and LiveBench. Pricing reflects public list API pricing per 1 million tokens in USD.
Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.
Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.
American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.
Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.
Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.
Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).
Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.
17 models from 8 labs.
OpenAI
San Francisco, USA
Anthropic
San Francisco, USA
Mountain View, USA
xAI
Bay Area, USA
Meta
Menlo Park, USA
DeepSeek
Hangzhou, China
Qwen
Hangzhou, China
Mistral
Paris, France