Field notes
How to pick a model in August 2026
Leaderboards answer “who scored highest on this test.” You still have to answer “what is the job, what does it cost, and can I actually buy it.”
If you only need one
Why theseStart with the job, not the lab
- Hard repo work. Prefer SWE-bench Pro and Terminal-Bench over HumanEval. Claude Fable 5 and Opus 5 still lead the hardest coding evals; GPT-5.6 Sol wins Terminal-Bench and ARC-AGI-2.
- Agents. Look at the agentic index and cost per task, not input price. GLM-5.3, Opus 5, and Grok 4.6 cluster here.
- Chat quality. Arena Elo is preference, not homework. Fable 5 and Gemini 3.7 Flash poll well; that does not mean they win SWE-Pro.
- Volume. GPT-5.6 Luna, DeepSeek V4 Flash, and Qwen3.8 Flash are the actual production defaults for high QPS.
Price is not the list price
Sol looks expensive at $5 / $30 and is often cheaper per completed task than Fable 5 because it spends fewer tokens. Grok 4.6 looks mid-priced and is near-Sol intelligence. Luna looks like a toy and is not. Always check cost per task when a lab publishes it.
Open weight is a product decision
Qwen3.8 Max is the evidence-backed open flagship. Kimi K3 is close and larger. GLM-5.3 is the agentic open pick if you can live without vision. DeepSeek V4 is what you self-host when the bill is the product. Llama 4 Maverick is for fine-tunes, not for beating Claude.
Things that no longer discriminate
MMLU above 90%, HumanEval, and “1M context” as a headline. Most flagships now have a million-token window. Grok 4.6 is the notable exception at 500K. Use ARC-AGI-2, HLE, SWE-Pro, and Terminal-Bench when the old exams have flattened.
Numbers and caveats: methodology.