Graph the top current models by task, cost, coverage, and confidence. Archived models and agent systems are available, but they stay out of the default ranking.
376
scores
25
files
0
unmatched
View
Current raw leader: Claude Fable 5
Treat this as a benchmark reference only. It should not be the default production choice.
Claude Opus 4.8
Anthropic
Expensive at output-heavy scale, but broad current coverage makes it one of the safest top comparisons.
GPT-5.5
OpenAI
Multiple submitted harness variants are preserved; compare the condition labels before declaring a winner.
Claude Opus 4.7
Anthropic
Older than Opus 4.8/Fable 5 but still a top coding and agentic comparison point.
GLM-5.2
Z.ai
Best open-weight value story in the submitted set, but not the best absolute closed-model score.
Green zone marks high score with lower blended token cost. 4 crowded labels are kept in the legend and hover instead of overlapping.
Top-right means strong on both code and agent tasks. 5 crowded labels are kept in the legend and hover instead of overlapping.
| Metric | Coding | Agentic | Reasoning | Math | Long Context | Value | Reliability |
|---|---|---|---|---|---|---|---|
| Coding | 1.00 | 0.98 | 0.49 | 0.63 | -0.32 | -0.61 | 0.81 |
| Agentic | 0.98 | 1.00 | 0.58 | 0.67 | -0.25 | -0.63 | 0.85 |
| Reasoning | 0.49 | 0.58 | 1.00 | 0.65 | 0.10 | -0.85 | 0.67 |
| Math | 0.63 | 0.67 | 0.65 | 1.00 | -0.37 | -0.78 | 0.68 |
| Long Context | -0.32 | -0.25 | 0.10 | -0.37 | 1.00 | 0.27 | -0.02 |
| Value | -0.61 | -0.63 | -0.85 | -0.78 | 0.27 | 1.00 | -0.57 |
| Reliability | 0.81 | 0.85 | 0.67 | 0.68 | -0.02 | -0.57 | 1.00 |
Agreement is computed only over visible models with both scores. It is a rank signal for overlap, not proof of benchmark validity.
Dark cells have score evidence. Pale cells are missing or unknown, not bad scores.
| Model | Arena | SWE | Terminal | GPQA | HLE | Math | Context | Reliability |
|---|---|---|---|---|---|---|---|---|
Claude Fable 5 Mixed confidence | 1535 Elo | 78.0% | 85.0% | 93.0% | 40.0% | 99.7% | 1M tokens | Mixed |
Claude Opus 4.8 Strong confidence | 1552 Elo | 2x | 88.5% | 94.2% | 38.5% | 98.3% | 1M tokens | Strong |
Claude Opus 4.7 Strong confidence | 1567 Elo | 83.5% | 90.2% | 91.8% | 35.2% | 97.8% | 500K tokens | Strong |
GPT-5.5 Strong confidence | 2x | 2x | 3x | 93.6% | 36.2% | 99.2% | 1M tokens | Strong |
GLM-5.2 Mixed confidence | 1468 Elo | 62.1% | 81.0% | 84.0% | Missing | Missing | 1M tokens | Mixed |
Grok 4.3 Limited confidence | 1498 Elo | 75.8% | 80.0% | 87.3% | Missing | 92.0% | 256K tokens | Limited |
DeepSeek V4 Pro Mixed confidence | 1470 Elo | 72.0% | Missing | 90.0% | Missing | 96.5% | 1M tokens | Mixed |
Qwen3.6 Max Preview Limited confidence | 1541 Elo | 70.5% | Missing | 87.0% | Missing | 95.5% | 256K tokens | Limited |
Anthropic · Reasoning / Research / Frontier math
Not generally available right now; use as a top-score reference, not an automatic production default.
Unavailable: Not generally available right now. Treat as a benchmark reference, not a deployable choice.
Score
94
Context
1M
Price
$20/$100
Anthropic · Agents / Coding / Knowledge work
Expensive at output-heavy scale, but broad current coverage makes it one of the safest top comparisons.
Score
93
Context
1M
Price
$5/$25
Anthropic · SWE-bench / Web dev / Agents
Older than Opus 4.8/Fable 5 but still a top coding and agentic comparison point.
Score
92
Context
500K
Price
$15/$75
OpenAI · General / Coding / Agents
Multiple submitted harness variants are preserved; compare the condition labels before declaring a winner.
Score
90
Context
1M
Price
$5/$30
Z.ai · Open weights / Coding / Cost
Best open-weight value story in the submitted set, but not the best absolute closed-model score.
Score
84
Context
1M
Price
$1.4/$4.4
xAI · General / Fast chat / Agents
Good top-set alternate, but public benchmark breadth is thinner than Anthropic/OpenAI.
Score
82
Context
256K
Price
$3/$15
DeepSeek · Open weights / Reasoning / Cost
Strong value and open-deployment option; benchmark breadth is less complete than the top closed models.
Score
79
Context
1M
Price
$0.55/$2.19
Alibaba Qwen · Coding / Multilingual / General
Worth watching, especially for coding/value, but preview status lowers confidence.
Restricted: Preview access and serving behavior should be validated before production use.
Score
78
Context
256K
Price
$0.8/$2.4
Google DeepMind · Multimodal / Long context / Research
Searchable for completeness, but not recommended by default in this explorer.
Restricted: Searchable for completeness, but not recommended by default in this explorer.
Score
76
Context
2M
Price
$2.5/$10
Export scope
Markdown
Best for notes, docs, and AI chat context.
HTML
Portable table snippet with semantic markup.
JSON
Structured payload for agents and crawlers.
| Model | Lab | Overall | Availability | Confidence | Coverage |
|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | 94 | Unavailable | Mixed | 8 |
| Claude Opus 4.8 | Anthropic | 93 | Available | Strong | 8 |
| Claude Opus 4.7 | Anthropic | 92 | Available | Strong | 8 |
| GPT-5.5 | OpenAI | 90 | Available | Strong | 8 |
| GLM-5.2 | Z.ai | 84 | Available | Mixed | 6 |
| Grok 4.3 | xAI | 82 | Available | Limited | 7 |
| DeepSeek V4 Pro | DeepSeek | 79 | Available | Mixed | 6 |
| Qwen3.6 Max Preview | Alibaba Qwen | 78 | Restricted | Limited | 6 |
| Gemini 3.1 Pro | Google DeepMind | 76 | Restricted | Limited | 7 |