A decision-first index of the frontier: benchmark evidence, real API cost, context, openness, and what each model is actually useful for.
Higher is better. Further left costs less. Only models with a published Terminal-Bench result and a listed API rate are plotted.
9 of 9 models · ranked for all-rounder
High-stakes engineering, tool-use workflows, and professional knowledge work.
Text · Vision · Tools
Premium API price; reported scores use xhigh reasoning effort in a research environment.
Benchmark scores are transcribed from first-party launch posts, model cards, and API documentation. A dash means no sufficiently comparable figure was found—not a zero.
Fit scores are an editorial decision aid based on reported capability, price, modalities, openness, and deployment tradeoffs. They are intentionally separate from benchmark values.
Prices are USD per million tokens at the standard API rate where publicly listed. “Self-host” means weights are available; your infrastructure still has a cost. “Not listed” is intentionally not estimated.