Still widely deployed prior GPT-5 generation — strong computer-use and coding history, often cheaper via existing contracts.
Teams already on GPT-5.x infra who need stability over bleeding edge.
Superseded by 5.5/5.6 on preference and newest composites.
Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.
Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).
Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.
PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.
Harder multi-choice knowledge/reasoning suite than classic MMLU.
Frontier closed-ended academic difficulty across many domains.
General frontier reasoning, mixed multimodal apps, and OpenAI ecosystem tooling.
Hard reasoning tasks, research assistants, and multimodal product surfaces.
Default production chat, tools, and mixed workloads on OpenAI.