Benchmark results
Top AI Model Explorer

Latest models first. Compare what actually matters.

Graph the top current models by task, cost, coverage, and confidence. Archived models and agent systems are available, but they stay out of the default ranking.

376

scores

25

files

0

unmatched

View

Recommended usable set
What to compare if you need to use a model now.
This view excludes unavailable benchmark-only models and downshifts weaker evidence before ranking.
Benchmark-only warning
Raw score is not the same as usable.
Raw leaders remain visible for transparency, but the practical recommendation layer excludes unavailable models.

Current raw leader: Claude Fable 5

Treat this as a benchmark reference only. It should not be the default production choice.

94
raw
0
usable
Unavailable
access
3
Strong
3
Mixed
3
Limited
2
Restricted
1
Unavailable
Top ranking
Overall leaders
Raw benchmark ranking. Missing scores are excluded, and unavailable models can still appear here.
Claude Fable 5AnthropicClaude Opus 4.8AnthropicClaude Opus 4.7AnthropicGPT-5.5OpenAIGLM-5.2Z.aiGrok 4.3xAIDeepSeek V4 ProDeepSeekQwen3.6 Max PreviewAlibaba QwenGemini 3.1 ProGoogle DeepMind
Active compare
Selected models
Defaults to the usable recommendation set. Use raw leaders when you want benchmark-only comparison.

Claude Opus 4.8

Anthropic

93

Expensive at output-heavy scale, but broad current coverage makes it one of the safest top comparisons.

GPT-5.5

OpenAI

90

Multiple submitted harness variants are preserved; compare the condition labels before declaring a winner.

Claude Opus 4.7

Anthropic

92

Older than Opus 4.8/Fable 5 but still a top coding and agentic comparison point.

GLM-5.2

Z.ai

84

Best open-weight value story in the submitted set, but not the best absolute closed-model score.

Quality vs cost
Best value is not always the top score.

Green zone marks high score with lower blended token cost. 4 crowded labels are kept in the legend and hover instead of overlapping.

Claude Fable 5AnthropicClaude Opus 4.8AnthropicGPT-5.5OpenAIClaude Opus 4.7AnthropicGrok 4.3xAIGLM-5.2Z.aiDeepSeek V4 ProDeepSeekQwen3.6 Max PreviewAlibaba QwenGemini 3.1 ProGoogle DeepMind
Coding vs agentic
Code strength and task execution diverge.

Top-right means strong on both code and agent tasks. 5 crowded labels are kept in the legend and hover instead of overlapping.

Claude Fable 5AnthropicClaude Opus 4.8AnthropicGPT-5.5OpenAIClaude Opus 4.7AnthropicGrok 4.3xAIGLM-5.2Z.aiDeepSeek V4 ProDeepSeekQwen3.6 Max PreviewAlibaba QwenGemini 3.1 ProGoogle DeepMind
Benchmark agreement
Which scores are actually saying the same thing?
Spearman rank correlation across the visible model set. Strong agreement can mean redundancy; weak agreement can reveal a real tradeoff.
MetricCodingAgenticReasoningMathLong ContextValueReliability
Coding1.000.980.490.63-0.32-0.610.81
Agentic0.981.000.580.67-0.25-0.630.85
Reasoning0.490.581.000.650.10-0.850.67
Math0.630.670.651.00-0.37-0.780.68
Long Context-0.32-0.250.10-0.371.000.27-0.02
Value-0.61-0.63-0.85-0.780.271.00-0.57
Reliability0.810.850.670.68-0.02-0.571.00

Agreement is computed only over visible models with both scores. It is a rank signal for overlap, not proof of benchmark validity.

Benchmark arbitrage
Where the leaderboard hides useful alternatives.
Models whose value rank diverges from the overall rank are often the most interesting practical options.
Coverage heatmap

Know when the data is missing.

Dark cells have score evidence. Pale cells are missing or unknown, not bad scores.

ModelArenaSWETerminalGPQAHLEMathContextReliability

Claude Fable 5

Mixed confidence

1535 Elo78.0%85.0%93.0%40.0%99.7%1M tokensMixed

Claude Opus 4.8

Strong confidence

1552 Elo2x88.5%94.2%38.5%98.3%1M tokensStrong

Claude Opus 4.7

Strong confidence

1567 Elo83.5%90.2%91.8%35.2%97.8%500K tokensStrong

GPT-5.5

Strong confidence

2x2x3x93.6%36.2%99.2%1M tokensStrong

GLM-5.2

Mixed confidence

1468 Elo62.1%81.0%84.0%MissingMissing1M tokensMixed

Grok 4.3

Limited confidence

1498 Elo75.8%80.0%87.3%Missing92.0%256K tokensLimited

DeepSeek V4 Pro

Mixed confidence

1470 Elo72.0%Missing90.0%Missing96.5%1M tokensMixed

Qwen3.6 Max Preview

Limited confidence

1541 Elo70.5%Missing87.0%Missing95.5%256K tokensLimited
Model table
Top current models
Search can reveal non-promoted models, but old references stay hidden unless enabled.

Claude Fable 5

latestUnavailable

Anthropic · Reasoning / Research / Frontier math

Not generally available right now; use as a top-score reference, not an automatic production default.

Unavailable: Not generally available right now. Treat as a benchmark reference, not a deployable choice.

Score

94

Context

1M

Price

$20/$100

Claude Opus 4.8

latest

Anthropic · Agents / Coding / Knowledge work

Expensive at output-heavy scale, but broad current coverage makes it one of the safest top comparisons.

Score

93

Context

1M

Price

$5/$25

Claude Opus 4.7

current

Anthropic · SWE-bench / Web dev / Agents

Older than Opus 4.8/Fable 5 but still a top coding and agentic comparison point.

Score

92

Context

500K

Price

$15/$75

GPT-5.5

latest

OpenAI · General / Coding / Agents

Multiple submitted harness variants are preserved; compare the condition labels before declaring a winner.

Score

90

Context

1M

Price

$5/$30

GLM-5.2

latestOpen weights

Z.ai · Open weights / Coding / Cost

Best open-weight value story in the submitted set, but not the best absolute closed-model score.

Score

84

Context

1M

Price

$1.4/$4.4

Grok 4.3

latest

xAI · General / Fast chat / Agents

Good top-set alternate, but public benchmark breadth is thinner than Anthropic/OpenAI.

Score

82

Context

256K

Price

$3/$15

DeepSeek V4 Pro

latestOpen weights

DeepSeek · Open weights / Reasoning / Cost

Strong value and open-deployment option; benchmark breadth is less complete than the top closed models.

Score

79

Context

1M

Price

$0.55/$2.19

Qwen3.6 Max Preview

latestRestricted

Alibaba Qwen · Coding / Multilingual / General

Worth watching, especially for coding/value, but preview status lowers confidence.

Restricted: Preview access and serving behavior should be validated before production use.

Score

78

Context

256K

Price

$0.8/$2.4

Gemini 3.1 Pro

latestRestrictedNot promoted

Google DeepMind · Multimodal / Long context / Research

Searchable for completeness, but not recommended by default in this explorer.

Restricted: Searchable for completeness, but not recommended by default in this explorer.

Score

76

Context

2M

Price

$2.5/$10

Selected compare
Side-by-side score shape
Bars omit missing values instead of treating them as zero.
Claude Opus 4.8AnthropicGPT-5.5OpenAIClaude Opus 4.7AnthropicGLM-5.2Z.ai
Claude Opus 4.88 benchmarks · Strong
GPT-5.58 benchmarks · Strong
Claude Opus 4.78 benchmarks · Strong
GLM-5.26 benchmarks · Mixed
AI export
Make the benchmark view easy to cite, copy, and parse.
Exports include the current task, view, filter result, selected models, caveats, and the artifact-based disclaimer.
This is an artifact-based explorer built from submitted benchmark-app data. It is not a certified live benchmark authority; verify model availability, pricing, and source benchmark claims before making production decisions.

Export scope

Markdown

Best for notes, docs, and AI chat context.

HTML

Portable table snippet with semantic markup.

JSON

Structured payload for agents and crawlers.

Public machine-readable JSON is available at /ai-benchmark-results/model-centric-data/export.json.
Export preview
9 models in this export
The preview shows the same model set sent to Markdown, HTML, and JSON.
AI benchmark export preview
ModelLabOverallAvailabilityConfidenceCoverage
Claude Fable 5Anthropic94UnavailableMixed8
Claude Opus 4.8Anthropic93AvailableStrong8
Claude Opus 4.7Anthropic92AvailableStrong8
GPT-5.5OpenAI90AvailableStrong8
GLM-5.2Z.ai84AvailableMixed6
Grok 4.3xAI82AvailableLimited7
DeepSeek V4 ProDeepSeek79AvailableMixed6
Qwen3.6 Max PreviewAlibaba Qwen78RestrictedLimited6
Gemini 3.1 ProGoogle DeepMind76RestrictedLimited7