ModelBench

20 frontier models · Jun 2026 benchmarks

Top SWE

Claude Opus 4.7

AAnthropic83.5%

Top GPQA

Gemini 3.1 Pro

GGoogle94.1%

Top HLE

Gemini 3.1 Pro

GGoogle46.4%

Top Arena

Claude Opus 4.7

AAnthropic1567

Filters

Lab

Tier

Modality

Access

Min context

Showing 20 models

Benchmark data sourced from independently-run evaluations (Epoch AI, Scale AI, OpenAI GDPval). Scores may differ from vendor self-reports.

Updated June 2026 · Not affiliated with any model provider.