Anthropic’s limited-access peak. Tops composite boards, but you cannot buy it on the public API. Treat Fable 5 as the available twin.
- Price /1M
- $10 / $50
- Context
- 1M
- Speed
- —
- Status
- Restricted
Field guide · 28 Aug 2026
Aperture ranks 34 current and recent models from 12 labs on intelligence, coding, agents, price, and speed. Filter the field, sort by the job, compare up to four, and leave with a pick you can defend.
Intelligence Index is the Artificial Analysis composite (approx. 0–65). Cards are the default on a phone; switch to table or the price map on a larger screen.
25 models
Anthropic’s limited-access peak. Tops composite boards, but you cannot buy it on the public API. Treat Fable 5 as the available twin.
Highest generally-available Intelligence Index. Half of Fable’s price, first-choice for agents and knowledge work.
The available Claude that wins the hardest coding evals. Expensive and slower than Opus 5. Default when the repo is the product.
OpenAI’s flagship. Leads ARC-AGI-2 and Terminal-Bench. Token-efficient per task even though list price looks high.
Ties GPT-5.6 Sol on intelligence at $2 / $6. Strong agents and GPQA. No public SWE-bench; 500K context is the ceiling.
Largest open-weight model that is actually good. 97% of the lead score at a lower output price. Native multimodal.
One point off Opus 5 on the Intelligence Index. Best agentic score in the open-weight set. Text-only; weights were staged late August.
Best open-weight flagship on evidence-backed boards. Slow to sample. Strong multilingual and agentic for the license.
Almost GLM-5.3 intelligence at Flash prices. If it holds under independent harnesses, it is the value model of August.
Meta’s closed frontier. Fast on 1.1, strong HLE, still thinner independent coding evidence than Claude or GPT.
The middle GPT-5.6. Keeps most of Sol’s coding, drops some agentic depth, much faster and cheaper.
Fastest serious reasoning model in the field. GPQA leader on several boards. Agentic depth still behind Claude and Grok.
Anthropic’s daily driver. Not the crown, but the one you actually leave on in production if Opus is too rich.
The default cheap-and-good open model for coding. Peak/off-peak API pricing. Not a 60-index flagship.
The punchline of the 5.6 stack: near-Terra coding at Luna prices. Default OpenAI pick for volume.
Dense 27B that hangs with models 50× its size. The local / workstation Qwen, not the datacenter Max.
Flash sibling of V4 Pro. Absurdly cheap, still 50+ intelligence. The volume open-weight pick.
Qwen’s cheap Fast lane. Use it as the open-weight Luna, not as a Max replacement.
Google’s Pro that actually shipped. Knowledge and multimodal are real; agentic scores are the hole in the card.
Xiaomi’s first model that belongs on a frontier page. Arena is real; published eval coverage is still thin.
Vendor SWE-bench is high; independent intelligence is mid. A specialist, not a general flagship.
Best generally available European API in this set. Context is 128K — a generation behind the 1M club.
Anthropic’s fast cheap tier. Useful for classification and routing, not for the frontier jobs on this board.
OpenAI’s open-weight line. Fine to self-host; do not confuse it with the 5.6 API stack.
The Llama you can actually fine-tune. Not on the 2026 intelligence frontier. Still the default open Meta line.