Frontier BenchAI model benchmarks · Jul 2026
ModelsCompareLabsMethodology

Labs

The organizations behind the models. Tap a model to drill in.

A

Anthropic

www.anthropic.com

Maker of the Claude family — focused on safety and agentic coding.

Founded 2021 · HQ San Francisco, USA · 4 models tracked

  • Claude Opus 4.8

    Current agentic-coding and computer-use leader. Best default for engineering teams.

    ≈$10.00/M

    1M ctx

  • Claude Opus 4.7

    Previous flagship. Still a strong repository-reasoning pick; superseded by 4.8.

    ≈$10.00/M

    1M ctx

  • Claude Sonnet 4.7

    Best price-to-capability ratio in the Claude lineup for everyday coding.

    ≈$6.00/M

    1M ctx

  • Claude Haiku 4.7

    Sub-second latency tier for chat and classification at scale.

    ≈$1.60/M

    250K ctx

O

OpenAI

openai.com

Maker of GPT and ChatGPT — broad multimodal frontier models.

Founded 2015 · HQ San Francisco, USA · 3 models tracked

  • GPT-5.5

    Best terminal/CLI autonomy and dense academic reasoning. Effort variants up to max.

    ≈$15.00/M

    400K ctx

  • GPT-5.5 Pro

    Higher-effort Pro tier. Wins on HLE with tools and FrontierMath.

    ≈$26.25/M

    400K ctx

  • GPT-5.1 Mini

    Workhorse small model — cheap, fast, multimodal, ubiquitous.

    ≈$0.70/M

    200K ctx

GD

Google DeepMind

deepmind.google

Maker of Gemini — massive context windows and tight Google ecosystem.

Founded 2010 · HQ London, UK · 2 models tracked

  • Gemini 3.1 Pro

    Undisputed long-context leader at 2M tokens. Best $/token for batch reasoning.

    ≈$4.50/M

    2M ctx

  • Gemini 3.5 Flash

    Beats the prior Pro flagship on coding at 4× the speed. Volume-pipeline darling.

    ≈$3.38/M

    1M ctx

X

xAI

x.ai

Maker of Grok — large context, integrated with X.

Founded 2023 · HQ San Francisco, USA · 2 models tracked

  • Grok 4.20

    2M context at frontier-cheap output pricing. Strong on raw math and HumanEval.

    ≈$3.00/M

    2M ctx

  • Grok 4.1 Fast

    Budget-tier coding leader. Cheapest credible agent for high-volume workloads.

    ≈$0.28/M

    2M ctx

D

DeepSeek

www.deepseek.com

Open-weight MoE leader — frontier-grade coding at a fraction of the cost.

Founded 2023 · HQ Hangzhou, China · 2 models tracked

  • DeepSeek V4 Pro

    Best open-weight coder. MIT license, 1M context, 49B active MoE.

    ≈$0.54/M

    1M ctx

  • DeepSeek V4 Flash

    Cheapest credible API on the market. Matches reasoning with a larger thinking budget.

    ≈$0.18/M

    1M ctx

AQ

Alibaba (Qwen)

www.alibaba.com

Maker of Qwen — Apache-2.0 enterprise and multilingual leader.

Founded 1999 · HQ Hangzhou, China · 2 models tracked

  • Qwen 3.7 Max

    Apache-2.0 hosted frontier model. Top multilingual and agentic-coding open option.

    ≈$3.75/M

    1M ctx

  • Qwen 3.5 Coder

    Code-tuned open model with a huge fine-tune ecosystem. IDE autocomplete favorite.

    ≈$0.30/M

    256K ctx

MA

Meta AI

ai.meta.com

Maker of Llama — open-weight ecosystem with the largest context windows.

Founded 2013 · HQ Menlo Park, USA · 2 models tracked

  • Llama 4 Maverick

    MoE flagship with deep ecosystem tooling. 17B active, 400B total.

    ≈$0.30/M

    1M ctx

  • Llama 4 Scout

    10M-token context — the only realistic choice for whole-repo or whole-corpus input.

    ≈$0.22/M

    10M ctx

MA

Mistral AI

mistral.ai

European open-weight lab — Apache-2.0 dense models, deployable anywhere.

Founded 2023 · HQ Paris, France · 2 models tracked

  • Mistral Large 3

    675B dense flagship. Cleanest license for EU data-sovereign deployment.

    ≈$0.60/M

    256K ctx

  • Mistral Small 4

    119B dense. Runs on a single 8×H100 node — easy self-host generalist.

    ≈$0.15/M

    128K ctx

ZA

Zhipu AI (Z.ai)

z.ai

Maker of GLM — cost-effective agentic and coding models.

Founded 2019 · HQ Beijing, China · 1 models tracked

  • GLM-4.6

    Cost-effective coding & agentic model. Strong math for the price.

    ≈$0.76/M

    200K ctx

Built by GLM-5.2 · Scores are directional, drawn from public mid-2026 benchmarks.

SWE-bench, GPQA, MMLU, AIME, HLE etc. are trademarks of their respective owners.