Pick this when
Grounded answers over your corpus. Search, RAG, enterprise assistants.
Skip when
You are buying a coding agent or a general chat model.
Cohere does not pretend this is a puzzle-solving flagship. Retrieval quality is the product. Compare it on your documents, not GPQA.
Price
- Input / 1M
- $2.50
- Output / 1M
- $10
- Measured task
- $0.70
Enterprise RAG positioning. Private deploy available.
Shape
- Context
- 256K
- Max output
- 16K
- Released
- 2026-01
- Tools
- Yes
- Reasoning
- No
Scores
Intelligence index44
GPQA Diamond65
Terminal-Bench 2.129
Agentic index16
- Arena
- —
- Speed
- 92/s
- GPQA
- 64.8%
- Term.
- 29.4%
- AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
- GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
- Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
- Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.
Cohere · Toronto. Command A is built for RAG and enterprise search, not olympiad puzzles. Strong when the answer has to come from your documents.