How to read the board · 2026-08-28
A leaderboard is a compressed story. These are the axes Fable uses, and the jobs they actually speak to.
01 · 0–100
Artificial Analysis composite of nine evals: agents, coding, scientific reasoning, and long-context work.
The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
02 · %
198 PhD-level science questions written so Google does not help a smart non-expert.
A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
03 · %
The model has to finish real tasks in a shell: install, debug, patch, verify.
Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
04 · Elo
Blind pairwise human votes on coding answers.
Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
05 · 0–100
Independent composite of multi-step tool use: plan, act, check, retry.
Chat quality and agent quality have split. Some models think well and stall in a loop. Others grind through jobs.
06 · tok/s
Tokens generated per second on a hosted API, averaged over a rolling window.
Matters for streaming chat and agent loops. Irrelevant for overnight batch jobs. Price per task usually matters more.
07 · USD
Measured spend to finish a standard independent task, not the sticker price per million tokens.
Verbose reasoning, retries, and long-context surcharges eat cheap rates. This is the number that shows up on the invoice.