ApertureAug 2026
FieldCompareLabsGuideMethod
  • Field
  • Compare
  • Labs
  • Guide

Field notes

How to pick a model in August 2026

Leaderboards answer “who scored highest on this test.” You still have to answer “what is the job, what does it cost, and can I actually buy it.”

If you only need one

Why these

Frontier

A

Claude Opus 5

Anthropic

Highest measured intelligence that you can actually buy.

63.1 · $5 / $25

Coding

A

Claude Fable 5

Anthropic

Repo work, SWE-bench Pro, Terminal-Bench.

62.1 · $10 / $50

Agents

Z

GLM-5.3

Z.ai (Zhipu)

Tool loops, computer use, long jobs.

59.5 · $1.40 / $4.40

Value

Z

GLM-5.3 Flash

Z.ai (Zhipu)

At least 55 intelligence, best points per dollar.

57.5 · $0.15 / $0.50

Open weight

K

Kimi K3

Moonshot AI

Highest open model with real evidence.

59.7 · $3 / $15

Speed

G

Gemini 3.7 Flash

Google DeepMind

Fastest model that still clears 50 intelligence.

56 · $0.75 / $3.75

Multimodal

G

Gemini 3.7 Flash

Google DeepMind

Vision plus video, ranked by intelligence.

56 · $0.75 / $3.75

Volume

Z

GLM-5.3 Flash

Z.ai (Zhipu)

Best model at or under $0.30 / 1M input.

57.5 · $0.15 / $0.50

Start with the job, not the lab

  1. Hard repo work. Prefer SWE-bench Pro and Terminal-Bench over HumanEval. Claude Fable 5 and Opus 5 still lead the hardest coding evals; GPT-5.6 Sol wins Terminal-Bench and ARC-AGI-2.
  2. Agents. Look at the agentic index and cost per task, not input price. GLM-5.3, Opus 5, and Grok 4.6 cluster here.
  3. Chat quality. Arena Elo is preference, not homework. Fable 5 and Gemini 3.7 Flash poll well; that does not mean they win SWE-Pro.
  4. Volume. GPT-5.6 Luna, DeepSeek V4 Flash, and Qwen3.8 Flash are the actual production defaults for high QPS.

Price is not the list price

Sol looks expensive at $5 / $30 and is often cheaper per completed task than Fable 5 because it spends fewer tokens. Grok 4.6 looks mid-priced and is near-Sol intelligence. Luna looks like a toy and is not. Always check cost per task when a lab publishes it.

Open weight is a product decision

Qwen3.8 Max is the evidence-backed open flagship. Kimi K3 is close and larger. GLM-5.3 is the agentic open pick if you can live without vision. DeepSeek V4 is what you self-host when the bill is the product. Llama 4 Maverick is for fine-tunes, not for beating Claude.

Things that no longer discriminate

MMLU above 90%, HumanEval, and “1M context” as a headline. Most flagships now have a million-token window. Grok 4.6 is the notable exception at 500K. Use ARC-AGI-2, HLE, SWE-Pro, and Terminal-Bench when the old exams have flattened.

Numbers and caveats: methodology.

Aperture is a field guide, not a vendor. Scores compiled 28 Aug 2026.

MethodologyHow to pick