← Back to explorer

Methodology

Model Gauge is a decision surface, not a live scrape. The dataset is a curated snapshot dated 2026-07-15, covering 19 models across 11 labs.

What we optimize for

  • Usefulness first. Presets answer “best for coding / science / value / open weights,” not only “who is #1 today.”
  • Multiple signals. Human preference (Arena), composite intelligence (AA Index), coding (SWE-bench), and science (GPQA) rarely agree — that disagreement is the point.
  • Practical constraints. Price, speed, context, and access type sit next to quality so developers can route traffic.

Benchmarks explained

Arena Elo

Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.

AA Intelligence Index

Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).

SWE-bench Verified

Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.

GPQA Diamond

PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.

MMLU-Pro

Harder multi-choice knowledge/reasoning suite than classic MMLU.

Humanity's Last Exam

Frontier closed-ended academic difficulty across many domains.

Sources

Limits

Vendor-reported numbers can differ from independent runs. Some models lack public scores on every bench — those cells show an em dash rather than inventing data. Arena Elo moves daily; treat ranks as bands, not gospel. For procurement, re-check primary sources on the day you decide.

Model Gauge · built by Cursor Grok 4.5 High · curated snapshot, not a live leaderboard

Scores change weekly — verify critical decisions against primary sources.