Sources

How these numbers were compiled

Snapshot date: 28 August 2026. Aperture is a static field guide. It does not run the evals itself.

Primary signals

  • Intelligence Index — Artificial Analysis composite across coding, math, knowledge, and agents. Scale is roughly 0–65 for this generation, not 0–100.
  • Coding / Agentic indices — independent coding and agent composites (LiveCodeBench, Terminal-Bench, tool-use). Displayed 0–100.
  • Arena Elo — LMArena / Chatbot Arena blind preference. Good for chat, poor for SWE-bench.
  • List prices — official API $ / 1M tokens where published. Peak/off-peak (DeepSeek) uses a representative mid rate.

Benchmarks on model pages

  • GPQA Diamond

    Graduate-level science questions designed so Google search is not enough. Still one of the cleanest knowledge/reasoning splits.

  • Humanity’s Last Exam

    Expert-written questions across fields. Harder than MMLU; the current differentiator for “does this model actually know things.”

  • MMLU-Pro

    Harder, less-saturated successor to MMLU. Classic MMLU is above 90% for every flagship and no longer ranks the field.

  • SWE-bench Verified

    500 human-validated GitHub issues. Score swings 5–15 points by harness — treat vendor numbers as an upper bound.

  • SWE-bench Pro

    Harder, contamination-resistant coding eval. Currently the best public split between “can code” and “can maintain a repo.”

  • Terminal-Bench 2.1

    End-to-end tasks in a real terminal. Better proxy for coding agents than HumanEval, which is fully saturated.

  • ARC-AGI-2

    Abstract visual puzzles. Rewards generalization over memorization. GPT-5.6 Sol currently leads the published set.

  • AIME 2025

    American Invitational Mathematics Examination. Contest math; reasoning-mode models dominate.

  • MMMU

    College-level multimodal understanding across diagrams, charts, and exam figures.

What “supported” means

A model is marked supported when at least two independent public sources agree on the core scores. Estimated means coverage is thin, vendor-only, or the checkpoint is too new. Missing cells are missing — they are not zeros.

SWE-bench Verified in particular moves 5–15 points with the harness. Vendor numbers are an upper bound. Prefer SWE-bench Pro and Terminal-Bench 2.1 when both exist.

Value score

Value is Intelligence Index divided by blended price (75% input + 25% output), scaled. It rewards models that stay above ~55 intelligence without Fable-class pricing. It is a shortcut, not a total cost of ownership model.

Not affiliated

Aperture is not affiliated with Anthropic, OpenAI, Google, xAI, Alibaba, Moonshot, Zhipu, DeepSeek, Meta, Mistral, MiniMax, or Xiaomi. Lab colors are used to identify products. Public sources consulted include Artificial Analysis, BenchLM, LMArena, and lab model cards.

Go back to the field.