ModelLab
ModelsCompare

About

How this data works

ModelLab is an independent dashboard. Here's what the numbers mean, where they come from, and what they don't tell you.

Where the numbers come from

Benchmark scores are the best publicly reported figures for each model, drawn from official model cards and papers, plus independent leaderboards: LMArena (Arena Elo), swebench.com (SWE-bench), and LiveBench. Pricing reflects public list API pricing per 1 million tokens in USD.

The benchmarks

Chatbot Arena

Elo

Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.

SWE-bench Verified

0–100

Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.

AIME 2025

0–100

American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.

GPQA Diamond

0–100

Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.

MMLU-Pro

0–100

Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.

Aider Polyglot

0–100

Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).

OSWorld

0–100

Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.

Labs covered

17 models from 8 labs.

OA

OpenAI

San Francisco, USA

A

Anthropic

San Francisco, USA

G

Google

Mountain View, USA

X

xAI

Bay Area, USA

∞

Meta

Menlo Park, USA

DS

DeepSeek

Hangzhou, China

Q

Qwen

Hangzhou, China

M

Mistral

Paris, France

Read benchmark scores with healthy skepticism

  • • Labs self-report many scores and can select favorable conditions. Independent, third-party runs are more trustworthy.
  • • Benchmarks saturate — once models score near the ceiling, the metric stops differentiating them.
  • • A leaderboard rank says nothing about your workload. The best way to choose is to test on your own data.
  • • “Value score” is a heuristic, not a published metric. Use it as a rough guide.
Back to leaderboardCompare models

ModelLab

An independent dashboard for comparing frontier AI models. Benchmark scores are drawn from public model cards, papers, and leaderboards. Always verify against the original source before making decisions.

Navigate

  • All models
  • Compare
  • About the data

Sources

  • LMArena ↗
  • SWE-bench ↗
  • LiveBench ↗

Data snapshot last checked June 17, 2026.

Not affiliated with any AI lab. Scores are approximations.