Methodology
Model Gauge is a decision surface, not a live scrape. The dataset is a curated snapshot dated 2026-07-15, covering 19 models across 11 labs.
What we optimize for
- Usefulness first. Presets answer “best for coding / science / value / open weights,” not only “who is #1 today.”
- Multiple signals. Human preference (Arena), composite intelligence (AA Index), coding (SWE-bench), and science (GPQA) rarely agree — that disagreement is the point.
- Practical constraints. Price, speed, context, and access type sit next to quality so developers can route traffic.
Benchmarks explained
Arena Elo
Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.
AA Intelligence Index
Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).
SWE-bench Verified
Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.
GPQA Diamond
PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.
MMLU-Pro
Harder multi-choice knowledge/reasoning suite than classic MMLU.
Humanity's Last Exam
Frontier closed-ended academic difficulty across many domains.
Sources
- LMArena / Chatbot Arena — Crowdsourced human preference Elo
- Artificial Analysis Intelligence Index — Composite intelligence + speed/price metrics
- SWE-bench Verified — Real-world software engineering tasks
- GPQA Diamond — Graduate-level science reasoning
Limits
Vendor-reported numbers can differ from independent runs. Some models lack public scores on every bench — those cells show an em dash rather than inventing data. Arena Elo moves daily; treat ranks as bands, not gospel. For procurement, re-check primary sources on the day you decide.