AI

Frontier Model Index

Snapshot: June 18, 2026

Sources

Filterable benchmark atlas

Pick the AI model that fits the job.

Compare leading AI labs across intelligence, coding, knowledge work, context, cost, access, and deployment control. Scores mix public provider tables with independent benchmark context where available.

Models tracked13
Major labs10
Open-weight options6
Current top visibleClaude Opus 4.8
OpenAIAnthropicGooglexAIMetaMistralDeepSeekQwenCohereZ.ai

13 models

Sorted by best fit for general.

ANT

Anthropic

Claude Opus 4.8

Frontier API model · May 2026

APIConsumerEnterpriseClosed1M context

Fit

92

Intelligence

61

Coding

72

Price in/out

$5.00 / $25.00

Benchmark details and guidance

Score profile

AA Intelligence
61
SWE / SWE-Pro
69.2
Terminal-Bench
85
GPQA Diamond
94.2
FrontierMath
43.8
GDPval / work
80.3
OSWorld
78
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

AgentsCodingKnowledge workResearch

Why it matters

  • Very strong agentic reliability and long-context work.
  • Often a safer choice than newer restricted models because API behavior is production-exposed.

Watchouts

  • Expensive at high output volume.
  • May be slower on simple high-throughput workloads.
OAI

OpenAI

GPT-5.5

Frontier API model · May 2026

APIConsumerClosed1M context

Fit

91

Intelligence

60

Coding

70

Price in/out

$5.00 / $30.00

Benchmark details and guidance

Score profile

AA Intelligence
60
SWE / SWE-Pro
58.6
Terminal-Bench
82.7
GPQA Diamond
93.6
FrontierMath
51.7
GDPval / work
84.9
OSWorld
78.7
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
83.2

Best use

GeneralCodingAgentsKnowledge workResearch

Why it matters

  • Broad frontier baseline with strong coding, tools, and knowledge-work results.
  • 1M context and predictable API pricing make it easier to budget than many top reasoning models.

Watchouts

  • Not open weight.
  • Specialized coding leaders can beat it on some software-engineering benchmarks.
GDM

Google DeepMind

Gemini 3.1 Pro Preview

Preview API model · March 2026

APIConsumerEnterpriseClosed1M context

Fit

83

Intelligence

57

Coding

58

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
57
SWE / SWE-Pro
54.2
Terminal-Bench
68.5
GPQA Diamond
94.3
FrontierMath
36.9
GDPval / work
67.3
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
80.5

Best use

MultimodalResearchGeneralAgents

Why it matters

  • Strong multimodal platform coverage and Google ecosystem integration.
  • Competitive scientific reasoning scores with preview-era room to improve.

Watchouts

  • Preview status can mean behavior changes.
  • Coding and knowledge-work scores trail the top closed models in the cited tables.
ZAI

Z.ai

GLM-5.2

Open-weight coding challenger · June 2026

APIOpen weightsOpen weight1M context

Fit

83

Intelligence

51

Coding

69

Price in/out

$1.40 / $4.40

Benchmark details and guidance

Score profile

AA Intelligence
51
SWE / SWE-Pro
62.1
Terminal-Bench
81
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
74
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

CodingAgentsOpen deploymentCost sensitive

Why it matters

  • Leading open-weight score on Artificial Analysis at this snapshot.
  • 1M context and MIT licensing make it attractive for private coding-agent stacks.

Watchouts

  • Newest entry in the set; expect more independent validation to arrive.
  • Reasoning and multimodal breadth still need closer comparison with closed frontier models.
DS

DeepSeek

DeepSeek V4 Pro

Open-source preview · April 2026

APIOpen weightsOpen weight1M context

Fit

81

Intelligence

52

Coding

0

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
52
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
73.5
LiveCodeBench
n/a
MMMU
n/a

Best use

Open deploymentCodingCost sensitiveAgents

Why it matters

  • Large sparse model with 49B active parameters and a 1M context window.
  • Strong option when self-hosting control matters more than top proprietary scores.

Watchouts

  • Preview model, so operational polish and serving behavior need validation.
  • Published benchmark coverage is less uniform across frontier tasks.
xAI

xAI

Grok 4.3

Current chat flagship · May 2026

APIConsumerClosedUnknown context

Fit

79

Intelligence

53

Coding

0

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
53
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

GeneralAgentsCost sensitive

Why it matters

  • xAI positions it as the default general model for chat.
  • Artificial Analysis reports a meaningful intelligence gain alongside lower pricing versus prior Grok 4.20.

Watchouts

  • Less public benchmark detail than OpenAI, Anthropic, or Google releases.
  • Use task-specific evaluation before adopting for regulated work.
QWN

Alibaba Qwen

Qwen3.6 Max Preview

Preview flagship · April 2026

APIClosedUnknown context

Fit

77

Intelligence

52

Coding

0

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
52
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

CodingGeneralAgents

Why it matters

  • Qwen3.6 focuses on stability, coding productivity, and real-world utility.
  • A serious China-lab option for teams already using Alibaba/Qwen tooling.

Watchouts

  • Preview branding means API and score movement are likely.
  • Compare against GLM and DeepSeek for open deployment needs.
META

Meta

Llama 4 Maverick

Open-source multimodal · April 2025

Open weightsOpen weightUnknown context

Fit

69

Intelligence

n/a

Coding

56

Price in/out

$0.19 / $0.49

Benchmark details and guidance

Score profile

AA Intelligence
n/a
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
69.8
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
80.5
LiveCodeBench
43.4
MMMU
73.4

Best use

Open deploymentMultimodalCost sensitive

Why it matters

  • Mature open ecosystem with strong multimodal and low-cost deployment story.
  • Good baseline for fine-tuning, distillation, and private hosting.

Watchouts

  • No longer the newest open-weight frontier entrant.
  • Reasoning scores trail 2026 closed frontier models.
META

Meta

Llama 4 Scout

Efficient open-source multimodal · April 2025

Open weightsOpen weightUnknown context

Fit

69

Intelligence

n/a

Coding

43

Price in/out

$0.19 / $0.49

Benchmark details and guidance

Score profile

AA Intelligence
n/a
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
57.2
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
74.3
LiveCodeBench
32.8
MMMU
69.4

Best use

Open deploymentCost sensitiveMultimodal

Why it matters

  • Cheaper and easier to operate than frontier-class closed systems.
  • Useful for constrained deployments where open weights matter.

Watchouts

  • Not a first choice for difficult reasoning or agentic coding.
  • Needs task-specific tuning to compete with larger open models.
ANT

Anthropic

Claude Fable 5

Restricted / watchlist · June 2026

RestrictedClosed1M context

Fit

69

Intelligence

65

Coding

80

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
65
SWE / SWE-Pro
80.3
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
90
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

AgentsKnowledge workCodingResearch

Why it matters

  • Current top score on the Artificial Analysis Intelligence Index when fallback is counted.
  • Provider-reported gains are strongest on complex analytical and long-running agent tasks.

Watchouts

  • Access is volatile and not a safe default for production planning.
  • Some headline numbers are provider-reported and need independent replication.
MST

Mistral AI

Mistral Medium 3.5

Open frontier-class model · April 2026

APIOpen weightsOpen weightUnknown context

Fit

66

Intelligence

n/a

Coding

0

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
n/a
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

Open deploymentCodingAgentsMultimodal

Why it matters

  • Mistral positions Medium 3.5 for agentic and coding workloads.
  • European open-model option for teams with sovereignty requirements.

Watchouts

  • Public cross-benchmark numbers are less standardized in the sources used here.
  • Run an internal bake-off against GLM, DeepSeek, and Llama before committing.
MST

Mistral AI

Mistral Large 3

Open-weight general model · December 2025

Open weightsEnterpriseOpen weight256K context

Fit

66

Intelligence

n/a

Coding

0

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
n/a
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

Open deploymentMultimodalEnterprise

Why it matters

  • State-of-the-art open-weight general multimodal model in Mistral's catalog.
  • A practical choice when permissive deployment and regional control matter.

Watchouts

  • Older than the newest 2026 coding-specialized challengers.
  • Use Medium 3.5 or newer coding models for agent-heavy workflows.
COH

Cohere

Command A+

Enterprise MoE · May 2026

APIEnterpriseClosed128K context

Fit

63

Intelligence

n/a

Coding

0

Price in/out

Not public

Benchmark details and guidance

Score profile

AA Intelligence
n/a
SWE / SWE-Pro
n/a
Terminal-Bench
n/a
GPQA Diamond
n/a
FrontierMath
n/a
GDPval / work
n/a
OSWorld
n/a
MMLU Pro
n/a
LiveCodeBench
n/a
MMMU
n/a

Best use

EnterpriseKnowledge workCost sensitiveMultimodal

Why it matters

  • Designed for enterprise deployment, reasoning, translation, and better tokenizer efficiency.
  • Can fit on a small high-end GPU footprint for private enterprise serving.

Watchouts

  • Not generally positioned as the strongest raw benchmark model.
  • Evaluate for sovereign deployment, latency, and multilingual ROI rather than headline rank.

Compare

Select up to four models.

OpenAI

GPT-5.5

Fit
91
AA index
60
Coding
70
Context
1M
Price
$5.00 / $30.00

Anthropic

Claude Opus 4.8

Fit
92
AA index
61
Coding
72
Context
1M
Price
$5.00 / $25.00

Reading the scores

A blank benchmark means the source set did not publish a directly comparable number. The fit score is an editorial composite that changes with the selected use case and penalizes restricted availability.

Sources and Method Notes

OpenAI

Introducing GPT-5.5

Provider release details, pricing, context window, and published benchmark tables.

OpenAI API

All models

Current OpenAI model catalog.

Claude API Docs

Models overview

Current Claude family and API availability notes.

Anthropic

Introducing Claude Opus 4.8

Provider release context for Claude Opus 4.8.

Anthropic

Claude Fable 5 and Claude Mythos 5

Provider release context for Anthropic's Mythos-class models.

Google AI for Developers

Gemini 3.1 Pro Preview

Google's developer-facing model description and positioning.

xAI Docs

Models

xAI model selection guidance for Grok 4.3 and modality-specific APIs.

Meta

Llama: industry leading, open-source AI

Llama 4 model positioning, benchmark table, and cost range.

Mistral Docs

Models overview

Mistral model catalog, release versions, and open model positioning.

DeepSeek API Docs

DeepSeek V4 Preview Release

DeepSeek V4 Pro and Flash size, context, and availability.

QwenLM

Qwen3.6

Qwen3.6 family release notes and goals.

Cohere Documentation

An overview of Cohere's models

Command A+ model details, context, output limit, and deployment positioning.

Artificial Analysis

GLM-5.2 is the new leading open weights model

Independent GLM-5.2 intelligence, cost, context, license, and open-weight comparison.

Artificial Analysis

Comparison of AI models across intelligence, performance, and price

Independent Intelligence Index and speed context across hundreds of models.

SWE-bench

SWE-bench leaderboards

Independent coding benchmark leaderboard context.