Claude Opus 4.7AAnthropicFlagshipLeads SWE-bench Verified and Text Arena coding. Best for software engineering.Context500KInput$15.00/1MReleasedFeb 2026SWE83.5%GPQA91.8%MMLU90.5%HLE35.2%SWE-bench leaderWeb devAgentsAdd to compare
Claude Opus 4.8AAnthropicFlagshipAnthropic's top model. Excels at agentic terminal tasks and web development.Context500KInput$15.00/1MReleasedApr 2026SWE81.0%GPQA92.5%MMLU91.8%HLE38.5%AgentsCodingLong documentsAdd to compare
GPT-5.5◯OpenAIFlagshipOpenAI's latest flagship with extended thinking. Top-tier on GPQA and competitive on SWE-bench.Context400KInput$2.50/1MReleasedMay 2026SWE80.6%GPQA94.0%MMLU92.1%HLE36.2%ReasoningCodingMathAdd to compare
Gemini 3.5 FlashGGoogleFastFast, cheap model that punches above its weight on coding benchmarks.Context1MInput$0.15/1MReleasedMay 2026SWE79.3%GPQA88.5%MMLU86.0%HLE30.0%SpeedCostSWE-bench valueAdd to compare
GPT-5.5 Pro◯OpenAIReasoningMaximum compute reasoning variant. Leads FrontierMath tiers and OTIS Mock AIME.Context400KInput$15.00/1MReleasedMay 2026SWE79.1%GPQA93.9%MMLU93.4%HLE44.3%Deep reasoningFrontier mathResearchAdd to compare
Claude Fable 5AAnthropicReasoningLeads SimpleBench and FrontierMath Tier 4. Anthropic's reasoning specialist.Context500KInput$20.00/1MReleasedMay 2026SWE78.0%GPQA93.0%MMLU92.5%HLE40.0%Common senseFrontier mathResearchAdd to compare
Gemini 3.1 ProGGoogleFlagshipGoogle's flagship with industry-leading context and Humanity's Last Exam score.Context2MInput$1.25/1MReleasedApr 2026SWE77.5%GPQA94.1%MMLU91.5%HLE46.4%MultimodalHLE leader2M contextAdd to compare
GPT-5.4◯OpenAIBalancedStrong all-rounder at a lower price point than 5.5 series.Context256KInput$1.75/1MReleasedMar 2026SWE76.9%GPQA93.3%MMLU91.0%HLE34.0%General purposeCost-efficientCodingAdd to compare
GPT-5.3 Codex◯OpenAICodingSpecialized coding model tuned for IDE and agentic development workflows.Context256KInput$2.00/1MReleasedJan 2026SWE75.5%GPQA82.0%MMLU85.0%HLE22.0%Code generationIDE integrationSWEAdd to compare
Claude Sonnet 4.6AAnthropicBalancedBest price-performance ratio in the Claude family.Context500KInput$3.00/1MReleasedMar 2026SWE74.5%GPQA88.0%MMLU88.2%HLE28.0%ValueSpeedCodingAdd to compare
DeepSeek V4 ProDDeepSeekFlagshipTop open-weight contender with API access at aggressive pricing.Context128KInput$0.55/1MReleasedFeb 2026SWE72.0%GPQA90.0%MMLU89.5%HLE30.0%Open weightsValueReasoningAdd to compare
o3◯OpenAIReasoningReasoning-focused model with strong math and science performance.Context200KInput$10.00/1MReleasedApr 2025SWE71.2%GPQA87.7%MMLU88.5%HLE24.8%STEM reasoningLong context recallAdd to compare
Qwen 3.7 MaxQQwenFlagshipStrong coding arena performer with open-weight availability.Context256KInput$0.80/1MReleasedMar 2026SWE70.5%GPQA87.0%MMLU88.8%HLE25.0%Coding arenaMultilingualOpen weightsAdd to compare
Grok 4𝕏xAIFlagshipxAI flagship with strong long-context comprehension and real-time X integration.Context256KInput$3.00/1MReleasedDec 2025SWE68.5%GPQA86.5%MMLU88.0%HLE26.0%Real-time dataLong fiction recallAdd to compare
Grok 4.1 Fast𝕏xAIFastOptimized for latency-sensitive applications with live data access.Context256KInput$0.50/1MReleasedMar 2026SWE65.0%GPQA82.0%MMLU84.0%HLE20.0%SpeedReal-timeLow costAdd to compare
Gemini 2.5 ProGGoogleBalancedPrevious-gen workhorse still widely deployed for production workloads.Context1MInput$1.25/1MReleasedJun 2025SWE63.2%GPQA84.0%MMLU87.5%HLE22.0%Long contextMultimodalStableAdd to compare
Mistral Large 3MMistralFlagshipEuropean flagship with strong multilingual capabilities.Context128KInput$2.00/1MReleasedNov 2025SWE55.0%GPQA78.0%MMLU85.0%HLE15.0%EU hostingMultilingualOpen weightsAdd to compare
DeepSeek R1DDeepSeekReasoningPioneering open reasoning model. Strong math, weaker on agentic tasks.Context128KInput$0.55/1MReleasedJan 2025SWE49.2%GPQA71.5%MMLU84.0%HLE8.5%Open weightsMathSelf-hostAdd to compare
Llama 4 Maverick∞MetaBalancedMeta's open multimodal model for on-premise and fine-tuning workflows.Context1MInputFree (OSS)ReleasedApr 2025SWE45.0%GPQA72.0%MMLU82.5%HLE12.0%Open weights1M contextSelf-hostAdd to compare
Command R+CCohereBalancedEnterprise-focused model optimized for retrieval-augmented generation.Context128KInput$2.50/1MReleasedAug 2024SWE38.0%GPQA58.0%MMLU75.5%HLE8.0%RAGEnterpriseRetrievalAdd to compare