ApertureAug 2026
FieldCompareLabsGuideMethod
  • Field
  • Compare
  • Labs
  • Guide
OpenAI
O

GPT-5.6 Sol

Jul 2026 · Proprietary · Current · supported evidence

OpenAI’s flagship. Leads ARC-AGI-2 and Terminal-Bench. Token-efficient per task even though list price looks high.

Open compareLab site

Intelligence

60.9

AA Index

API price

$5 / $30

Input / output per 1M

Context

1.05M

74 tok/s

Capability

Intelligence Index60.9
Coding index77.4
Agentic index57.8
Arena Elo (offset)182

Use it when

  • ▸Abstract reasoning
  • ▸Codex-style agents
  • ▸Tool-heavy apps

Skip if

  • –You need the cheapest intelligence point

Benchmarks

GPQA Diamond
Graduate-level science questions designed so Google search is not enough. Still one of the cleanest knowledge/reasoning splits.
94.6%
Humanity’s Last Exam
Expert-written questions across fields. Harder than MMLU; the current differentiator for “does this model actually know things.”
49.5%
MMLU-Pro
Harder, less-saturated successor to MMLU. Classic MMLU is above 90% for every flagship and no longer ranks the field.
—
SWE-bench Verified
500 human-validated GitHub issues. Score swings 5–15 points by harness — treat vendor numbers as an upper bound.
—
SWE-bench Pro
Harder, contamination-resistant coding eval. Currently the best public split between “can code” and “can maintain a repo.”
64.6%
Terminal-Bench 2.1
End-to-end tasks in a real terminal. Better proxy for coding agents than HumanEval, which is fully saturated.
88.8%
ARC-AGI-2
Abstract visual puzzles. Rewards generalization over memorization. GPT-5.6 Sol currently leads the published set.
92.5%
AIME 2025
American Invitational Mathematics Examination. Contest math; reasoning-mode models dominate.
—
MMMU
College-level multimodal understanding across diagrams, charts, and exam figures.
—

List price cut in late August 2026; some boards still show $4/$20.

Modalities

text · vision · audio · tools

Reasoning mode

reasoning

Hybrid and reasoning models spend tokens thinking. That raises GPQA and agents, and also raises latency and bill.

Close on intelligence

  • AClaude Fable 562.1
  • AClaude Opus 563.1
  • xGrok 4.660.9

Same lab

  • GPT-5.6 Terra
  • GPT-5.6 Luna
  • GPT-5.5
  • GPT-5.4

Aperture is a field guide, not a vendor. Scores compiled 28 Aug 2026.

MethodologyHow to pick