MFMODEL FIELD NOTES
Snapshot · 15 Jul 2026/7 labs · 7 models

FRONTIER MODEL INDEX ↗

AI models, compared without the theatre.

A practical field guide to the models shipping now — benchmark signals, context, modality, and where each one actually fits.

↓

HOW TO READ THIS

Use the benchmark columns as signals, then use the fit notes to make a decision. No single score tells the whole story.

01START HERE
BUILD AGENTSGPT-5.5 · Opus 4.8long-horizon + tool use
LONG CONTEXTGemini 3.1 · DeepSeek V41M token class
OPEN WEIGHTSDeepSeek · Mistral · Llamaself-host / fine-tune

02 / THE WORKSPACE

Find the right model for the job.

7 of 7 models
LAB
USE CASE
SORT BY
#MODEL ↕RELEASECONTEXTSWE-ProcodingTerminal 2.0agenticGPQAscienceBrowseCompresearchARC-AGI 2abstractBEST FIT
01
G
Gemini 3.1 ProGoogle · Closed
Feb 20261M54.2%68.5%94.3%85.9%77.1%Strongest long-context + multimodal blend
02
O
GPT-5.5OpenAI · Closed
Apr 2026400K58.6%82.7%93.6%84.4%85.0%Best all-rounder for long-running work
03
A
Claude Opus 4.8Anthropic · Closed
May 20261M——93.6%——Careful collaborator for complex agent work
04
∞
Llama 4 MaverickMeta · Open weights
Apr 20251M—————Flexible multimodal base for custom products
05
M
Mistral Medium 3.5Mistral · Open weights
May 2026——————Self-hostable coding + agent workhorse
06
x
Grok 4.5xAI · Closed
Jul 2026—64.7%83.3%———Fast path for real-world engineering tasks
07
D
DeepSeek-V4 ProDeepSeek · Open weights
Apr 20261M—————Open-weight frontier for self-hosted agents
01
G
Gemini 3.1 ProGoogle · Feb 2026

Strongest long-context + multimodal blend

CONTEXT1MGPQA94.3%SWE-PRO54.2%WEIGHTSCLOSED
02
O
GPT-5.5OpenAI · Apr 2026

Best all-rounder for long-running work

CONTEXT400KGPQA93.6%SWE-PRO58.6%WEIGHTSCLOSED
03
A
Claude Opus 4.8Anthropic · May 2026

Careful collaborator for complex agent work

CONTEXT1MGPQA93.6%SWE-PRO—WEIGHTSCLOSED
04
∞
Llama 4 MaverickMeta · Apr 2025

Flexible multimodal base for custom products

CONTEXT1MGPQA—SWE-PRO—WEIGHTSOPEN
05
M
Mistral Medium 3.5Mistral · May 2026

Self-hostable coding + agent workhorse

CONTEXT—GPQA—SWE-PRO—WEIGHTSOPEN
06
x
Grok 4.5xAI · Jul 2026

Fast path for real-world engineering tasks

CONTEXT—GPQA—SWE-PRO64.7%WEIGHTSCLOSED
07
D
DeepSeek-V4 ProDeepSeek · Apr 2026

Open-weight frontier for self-hosted agents

CONTEXT1MGPQA—SWE-PRO—WEIGHTSOPEN

03 / SIGNALS, NOT SCORES

Benchmarks have a point of view.

01CODINGpass@1 / %
Grok 4.564.7
GPT-5.558.6
Gemini 3.154.2

Public SWE-Pro results currently favor models tuned for autonomous software work.

02SCIENCEGPQA Diamond / %
Gemini 3.194.3
GPT-5.593.6
Opus 4.893.6

At the frontier, the spread is tiny. Method and reasoning budget matter more than the decimal.

03CONTEXTmaximum reported
1M
400K
128K+

Context is a capability, not a guarantee. Retrieval quality and cost still decide production fit.

MF

Model Field Notes is a research-style UI prototype. Data reflects provider-published pages available on 15 Jul 2026.

Sources ↗Back to top ↑
2/3models queued for compare
OGPT-5.5GGemini 3.1 Pro