Pick this when
High volume, classification, fan-out, and the cheap first pass of an agent stack.
Skip when
The task is actually hard. Luna will spend tokens retrying and lose the discount.
Lowest measured task cost in the closed frontier families. Scores with good open models while remaining a first-party OpenAI endpoint.
Price
- Input / 1M
- $0.20
- Output / 1M
- $1.20
- Measured task
- $0.05
80% price cut on 30 Jul 2026. Same 272K surcharge pattern.
Shape
- Context
- 1.1M
- Max output
- 128K
- Released
- 2026-07
- Tools
- Yes
- Reasoning
- Yes
Scores
Intelligence index52
GPQA Diamond78
Terminal-Bench 2.161
Agentic index28
- Arena
- —
- Speed
- 180/s
- GPQA
- 78.4%
- Term.
- 61.2%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
OpenAI · San Francisco. ChatGPT's parent. The GPT-5.6 line is a three-rung ladder: Sol for hard work, Terra for most production traffic, Luna for volume.