Frontier BenchAI model benchmarks · Jul 2026
ModelsCompareLabsMethodology
All models
AQ

Qwen 3.5 Coder

Alibaba (Qwen)

Code-tuned open model with a huge fine-tune ecosystem. IDE autocomplete favorite.

Open-weightApache 2.0256K contextReleased Feb 2026textcode

API price · per 1M tokens

$0.2/$0.6

input / output

≈ $0.30/M blended

Benchmarks

Bar shows this model's score; tick shows the best score across all tracked models for context.

SWE-bench VerifiedReal GitHub issues, human-verified patches. The standard coding-agent benchmark.
75.0/ best 90.1
SWE-bench Pro2,294 harder real-world GitHub issues. Where the frontier separates now.
48.0/ best 69.2
HumanEval+Functional code generation from docstrings.
93.0/ best 96.0
MATH-500Competition-style math problems.
84.0/ best 97.8

Other Alibaba (Qwen) models

AQ

Qwen 3.7 Max

Frontier

≈$3.75/M

Built by GLM-5.2 · Scores are directional, drawn from public mid-2026 benchmarks.

SWE-bench, GPQA, MMLU, AIME, HLE etc. are trademarks of their respective owners.