Skip to catalog
FbFable
CatalogLabsScoresCompare
Catalog
Mi

Mistral

Ministral 3 14B

ministral-3-14b-reasoning

36Index
Apache-2.0Generally available256K context14BText

Pick this when

On-device and air-gapped work where 14B is the budget.

Skip when

Cloud quality is an option. This is an edge model.

Honest about its size. That is rarer than it should be.

Price

Input / 1M
$0.20
Output / 1M
$0.20
Measured task
$0.06

Edge / on-device. Apache 2.0.

Shape

Context
256K
Max output
16K
Released
2026-02
Tools
Yes
Reasoning
Yes

Scores

Intelligence index36
GPQA Diamond50
Terminal-Bench 2.122
Agentic index11
Arena
—
Speed
70/s
GPQA
50.0%
Term.
22.1%
  • AA Index. The single number most people mean by “how smart.” It is a blend, so a specialist can lose here and still win the job you care about.
  • GPQA. A clean test of hard reasoning. PhD experts sit around 65%. It says little about writing, tools, or taste.
  • Term.. Closest public proxy for coding agents that live in a terminal. A high index score with a weak terminal score is a warning.
  • Arena. Captures taste and usefulness that unit tests miss. Sample size varies by model, so treat gaps under ~100 Elo as noise.

Mistral · Paris. Europe's lab. Smaller, cheaper models with EU data residency. Large 3 is the hosted flagship. Ministral is the on-device one.

Compiled 2026-08-28 from public Artificial Analysis, LMSYS-style arena, and first-party price sheets. Scores move. Recheck before you spend.

A two-point gap on the intelligence index is noise for most work. Cost per finished task is the number that shows up on the invoice.