This page compares the submitted Next.js benchmark apps as build outputs. The previews preserve the model builders' own interfaces, while the review flags stale data, weak sourcing, and likely hallucination risk without certifying the benchmark claims.
Review mode
Findings are based on the submitted artifacts, visible app behavior, and peer comparison. This is not a full source-backed audit of every benchmark number.
20
Submissions
20
Live source previews
6/20
Build times
The benchmark explorer leads with current top models, cost/quality scatter plots, coding vs agentic charts, coverage heatmaps, and a sticky compare tray.
Open explorerEach link opens the submitted app's own interface, with Abundance chrome removed except for a small back button.
The gallery's best-sourced single page: per-benchmark source links, explicitly versioned tracks, six shipped unit tests, and a dated Sep 4, 2026 snapshot, held back only by reliance on publisher comparison tables (including competitors' cards) and the recurring SpaceXAI relabel.
pass
3
watch
2
fail
0
Source
Benchmarks/Astra
Features
7/7
The gallery's most complete build: five working routes, job-based recommendations, URL-shareable filter state, and a dated Aug 28, 2026 snapshot across 31 models, with aggregate sourcing and an unpublished cost-per-task harness as the residual caveats.
pass
3
watch
2
fail
0
Source
Benchmarks/Fable/fable
Features
7/7
The most transparent single-page board in the gallery: current launch-week coverage, an explicit Sep 2, 2026 snapshot, and honest per-record caveats, held back only by vendor-only figures and a blended score that quietly rewards thin rows.
pass
2
watch
3
fail
0
Source
Benchmarks/Musespark 1.3
Features
7/7
A tidy three-route dashboard with solid compare and detail flows, failing the freshness bar on an 18-model dataset that stops at early-2025 releases and cites no sources.
pass
2
watch
0
fail
2
Source
Benchmarks/Ling 3.0 Flash/ai-models-benchmarks
Features
6/7
A clean, fast leaderboard with unusually good operational columns, undermined by a mid-2025 dataset presented under a 'Live' title and by sourcing that is never stated.
pass
2
watch
0
fail
2
Source
Benchmarks/Muse Spark 1.2/bench
Features
6/7
The strongest multi-route submission in the gallery: Aperture pairs a 34-model, 12-lab August 2026 snapshot with honest access labeling, skip-if guidance, a Pareto comparison, and a stated missing-data policy.
pass
3
watch
1
fail
0
Source
Benchmarks/Grok 4.6 Grok Build
Features
7/7
The most feature-complete single-page submission, but it fails the freshness bar: a Feb-2025-era dataset is framed as today's leaderboard, so the useful interactions sit on top of stale evidence.
pass
2
watch
1
fail
2
Source
Benchmarks/Gemini Flash 3.7
Features
7/7
The most editorially disciplined submission in the gallery: ModelPulse pairs 38 curated records with explicit null handling, rank gating, primary-source links, and shareable URL state, at the cost of many empty cells for the newest releases.
pass
3
watch
2
fail
0
Source
Benchmarks/Zcode 5.3 Flash/modelpulse
Features
6/7
A dense, polished comparison dashboard: Zcode 5.3 shipped 32 typed records across 13 labs with dual views, four-way compare, and a working QA script—the main caveat is aggregate acknowledgment of estimated values instead of per-cell flags.
pass
2
watch
2
fail
0
Source
Benchmarks/Zcode 5.3/modeldex
Features
6/7
Gemini 3.6 Flash High delivered the most complete interaction surface of the Gemini submissions and now builds on Next.js 16.2.11, but its data remains an older static benchmark snapshot that needs sourcing and modernization.
pass
2
watch
2
fail
0
Source
E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash
Features
7/7
Fast and approachable, but it fails the latest-model requirement: the submitted dataset looks stale next to the other entries.
pass
2
watch
1
fail
2
Source
project root
Features
7/7
The broadest app structure, with real routes and strong comparison surfaces; the main weakness is that the confident benchmark data still needs verification.
pass
2
watch
2
fail
0
Source
ai-bench/
Features
6/7
The most evidence-aware submission: it still needs external verification, but it did the best job separating useful UI from caveats and source notes.
pass
2
watch
1
fail
0
Source
src/components/model-explorer.tsx and src/data/models.ts
Features
7/7
A clean, usable explorer with strong dark UI fidelity, but its benchmark claims are still mostly asserted rather than evidenced.
pass
2
watch
2
fail
0
Source
Archive/Composer/Grok Composer
Features
6/7
Best for interface completeness, but the data layer is openly approximate and carries real hallucination risk.
pass
2
watch
1
fail
1
Source
Archive/Grok
Features
6/7
Terra is visually polished and fast to scan, but its main comparison and detail affordances stop short of a result. The normalized numbers should be treated as editorial fit signals until a formula and claim-level sources are published.
pass
2
watch
2
fail
0
Source
C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas
Features
6/7
Luna is the more rigorous of the two additions: its benchmark caveats, blank-value handling, mobile table alternative, and functioning comparison modal make the artifact useful without pretending the provider results are normalized.
pass
2
watch
2
fail
0
Source
C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5
Features
6/7
A genuinely different and useful decision surface, especially for task presets and practical constraints. Its self-featured Grok result is candid about missing cells but still needs independent claim-level verification.
pass
2
watch
2
fail
0
Source
E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5
Features
7/7
The most structurally ambitious result in this round. GLM built a real multi-route benchmark product with excellent comparison ergonomics, while its exact scores still require source-by-source verification.
pass
3
watch
1
fail
0
Source
E:/Projects/Benchmark Tests/GLM/glm-5.2-coder
Features
6/7
The strongest complete newcomer: Sol shipped a restrained, mobile-first decision tool with unusually clear benchmark caveats, useful interactions, and executable source tests.
pass
3
watch
1
fail
0
Source
components/model-explorer.tsx and data/models.ts
Features
7/7
These findings review the submitted apps as artifacts. They flag obvious stale data, weak sourcing, and likely hallucination risk, but they are not a complete fact-check of every AI benchmark claim.
Search, lab multi-select, access and use-case filters, benchmark and sort selects, table and chart views, mobile cards, a detail dialog, a methodology dialog, a three-model compare, CSV export, and a saved shortlist all respond client-side; dialogs are native <dialog> elements with Escape dismissal and focus containment.
Evidence: The explorer state pipeline in components/explorer.tsx and the native-dialog Modal in components/ui.tsx.
Score cells open the publisher table they came from (with scoreSources overrides such as GPT-5.6 Sol's Terminal 2.1), and model dialogs add documentation and pricing links; the CSV export carries benchmark, terminal-override, details, and price source columns.
Evidence: The SourceLink usage in components/model-dialog.tsx and the toCSV headers in lib/explorer.ts.
Filtering, null-last two-way sorting, the three-model cap, and CSV serialization live in a dependency-free lib/explorer.ts covered by six node:test cases; the data module enforces bounded scores and https sources, and the QA report records axe 4.12.1 with zero violations.
Evidence: tests/explorer.test.ts and docs/qa-report.md in the Astra source.
Qwen3.8-Max's figures come from Z.ai's GLM-5.3 model card and several Claude and Gemini rows cite OpenAI's GPT-6 comparison table, so cross-lab numbers pass through a rival's curation; the app discloses this in each note, but the source link a user opens is the competitor's page, not the subject lab's.
Evidence: The sourceLabel and source fields on the qwen-3-8-max, claude-opus-5, and gemini-3-8-flash records in lib/models.ts.
xAI is relabeled 'SpaceXAI' in the lab list and release notes, repeating a naming quirk other gallery submissions were dinged for, and Gemini 3.1 Pro keeps a preview-table DeepSWE of 11.8 in the main ranking behind a status tag that is easy to miss.
Evidence: The labs entry for id 'xai' and the gemini-3-1-pro record in lib/models.ts.
How Astra did it
Selecting any of nine job presets re-computes a weighted blend of benchmark, cost, and availability factors, re-ranks the catalog, and labels each card Best match, Strong fit, Usable, or Weak fit.
Evidence: jobScore and fitFromScore in lib/catalog.ts driving JobPicker and the card badges.
Catalog, models, labs, benchmarks, and compare routes all render statically; filters serialize to the URL, compare persists to localStorage, and an eight-test suite covers ranking, filtering, search, and delayed-availability behavior.
Evidence: lib/url-query.ts, the compare store, and lib/catalog.test.ts.
A single DATA_AS_OF date (2026-08-28) appears on the catalog, methodology, and footer; releases run through 2026-08-18, and the one unavailable model is listed as delayed with an explicit warning rather than dropped.
Evidence: DATA_AS_OF in lib/benchmarks.ts and the Gemini 3.5 Pro record in lib/models.ts.
The footer attributes scores to Artificial Analysis, arena standings, and first-party price sheets, but no model or benchmark links to its source, and the site URL collected for every one of the 14 labs is never rendered anywhere in the app.
Evidence: The footer copy and the unused site field in lib/labs.ts.
Cost per task drives the budget-weighted job scores, but the app states only that it is measured spend on a standard task; the harness, task set, and date are unpublished, so the numbers cannot be reproduced.
Evidence: The benchmark description in lib/benchmarks.ts and the pricing.note fields in lib/models.ts.
How Fable did it
Search, use-case chips, nine-lab multi-select, three toggles, eleven sort modes, cards and table views, a detail modal, and a three-model compare with per-row stars all respond client-side with no dead controls.
Evidence: The single client state pipeline in app/page.tsx over data/models.ts.
The data module, methodology section, and footer all carry the same compiled-on date, name the figures as vendor-reported, and quote independent re-run deltas such as Epoch AI's V4-Pro SWE result.
Evidence: The models.ts header comment, the on-page methodology copy, and the footer snapshot line.
Unreported benchmark cells count as a neutral 50 of the best-in-lineup ratio, so Gemini 3.8 Flash (one reported figure) still lands around 61 overall and Astra (zero reported) scores exactly 50. The method is disclosed, but the ranking still flatters launch-day records.
Evidence: The blendedRank function in data/models.ts and the methodology paragraph that defends it.
No benchmark number links to a source: Epoch AI, Terminal-Bench, and WebDev Arena are named in prose without URLs, and xAI is relabeled 'SpaceXAI' in the lab list, footer, and metadata, contradicting the lab naming used elsewhere in the gallery.
Evidence: The methodology fine print, the footer disclaimer, and the LABS entry for id 'xai'.
The cards/table toggle is hidden below the sm breakpoint except inside the expanded filters sheet, so mobile users must open filters to change views.
Evidence: The hidden sm:flex and sm:hidden classes on the view-mode controls in app/page.tsx.
How Muse Spark 1.3 did it
The board filters and sorts client-side, compare opens as a modal and a dedicated page for up to four models, and each model renders a detail page with full benchmark bars and cost efficiency.
Evidence: app/page.tsx state, CompareModal, app/compare/page.tsx, and app/model/[id]/page.tsx.
Records are typed with pricing (including cached rates), architecture, params, and modalities, and the model page renders straight from the typed module.
Evidence: lib/models.ts and the model detail route in the submitted source.
The layout metadata promises 'the latest AI models', but the dataset stops at GPT-4.5, Claude 3.5, and Gemini 2.0-era releases with no 2025-late or 2026 coverage.
Evidence: The 18-record dataset in lib/models.ts versus the metadata description.
Ten benchmark families are populated per model with no citations, methodology, or compiled-on date anywhere in the app.
Evidence: Benchmark entries in lib/models.ts carry descriptions but no source fields.
How Ling 3.0 Flash did it
Search, lab, modality, and access filters, sorting, cards and table views, a benchmark-focus selector, and a detail modal all respond client-side with no dead controls.
Evidence: The single-page client state pipeline in app/page.tsx.
Filters collapse into a sheet on small screens and the compare tray stays reachable while scrolling.
Evidence: showFilters state and the compare tray markup in app/page.tsx.
The app titles itself 'BENCHMARK — Live AI Model Leaderboard' but the newest records are Claude 4, Gemini 2.5 Pro, and Grok 3, with no GPT-5.x, Claude 4.5+/5, Gemini 3.x, GLM-5.x, or Grok 4.x coverage.
Evidence: The 20-record dataset in data/models.ts versus the title in the submitted layout metadata.
The app presents benchmark figures with no methodology, source list, or compiled-on date, so visitors cannot tell how current or verifiable the numbers are.
Evidence: The page copy and data module contain no sourcing or snapshot note.
How Muse Spark 1.2 did it
Records distinguish restricted from generally available models, mark evidence quality, and state that scores are compiled from public leaderboards, not live eval runs.
Evidence: status and evidence fields in lib/models.ts plus the README and methodology copy.
Visitors can search and filter the field board, switch cards, table, and price-map views, open use-case picks, compare up to four models with a Pareto chart, and read per-model, per-lab, guide, and methodology pages.
Evidence: FieldBoard, UseCasePicks, CompareView with ParetoChart, and the route files under app/.
Model and lab pages render statically from typed records, compare state is URL-driven, and the compare view is Suspense-wrapped for the searchParams read.
Evidence: generateStaticParams in the submitted models/[slug] and labs/[slug] pages and the Suspense boundary in app/compare/page.tsx.
The methodology names public sources (Artificial Analysis, BenchLM, LMArena, lab cards) but individual cells do not link to the specific claim they came from.
Evidence: Methodology page copy versus the unlinked numeric cells in lib/models.ts.
How Grok 4.6 did it
Visitors can search, filter by lab, category, access type, price ceiling, and context floor, switch between cards and a leaderboard table, sort every metric, open a detail modal, run a recommendation wizard, and compare up to four models with radar and bars.
Evidence: page.tsx state pipeline plus FilterBar, LeaderboardTable, ModelWizard, ComparisonDock, ComparisonModal, ModelDetailModal, and BenchmarkGlossary components.
Seventeen media queries across the component styles adapt the grid, table, dock, and modals, and buttons, selects, and inputs keep a 40px minimum touch height.
Evidence: Media queries in the CSS modules and the touch-target rules in the app styles.
The app presents a late-2024/early-2025 snapshot as the current frontier: Claude 3.7 Sonnet wears the 'King' badge and o1 is framed as OpenAI's premier reasoning model, with no GPT-5.x, Claude 4.x/5, Gemini 3.x, GLM-5.x, or Grok 4.x coverage.
Evidence: The 16-record dataset in data/models.ts and the badges and overview copy attached to those records.
Recommendations and leader labels describe year-old releases as today's best, so a visitor following the wizard would be pointed at superseded models.
Evidence: ModelWizard outputs and HighlightsBanner copy drawn from the same stale dataset.
Ten benchmark families are populated with plausible values, but records carry no source links a visitor could verify; the footer links only to lab homepages.
Evidence: BenchmarkScores fields in data/models.ts and the Footer lab links.
How Gemini 3.7 Flash did it
Unpublished scores render as em dashes, ranks require at least three published benchmarks, and the module header states that numbers are never invented to fill gaps.
Evidence: data.ts header comment, the rank gating in stats.ts, and the on-page methodology list.
All-round, Reasoning, Coding, and Agents tabs re-frame default sorts, and every filter, tab, sort, and selection serializes into the query string and rehydrates on load.
Evidence: view.ts serializeFilters/parseFilters plus the history.replaceState sync in explorer.tsx.
The board is a client surface, compare reads its model IDs from searchParams, and each of the 38 models renders as a static page with sources, caveats, and same-lab siblings.
Evidence: app/page.tsx, app/compare/page.tsx, and app/models/[id]/page.tsx with generateStaticParams in the submitted source.
Many headline numbers are vendor-reported on vendor-chosen harnesses; model pages flag material disagreements (for example Kimi K3 and DeepSeek V4 SWE-bench runs) but no independent evaluation exists.
Evidence: Per-model caveats arrays and the 'Read the fine print' methodology block.
GLM-5.3, Grok 4.6, and Claude Opus 5 ship with empty score sets, so benchmark columns show dashes and comparisons fall back to price and context takeaways.
Evidence: Empty scores objects for those records in data.ts.
How Zcode 5.3 Flash did it
Visitors can search, filter by lab, open weights, and capability, switch between card and table views, sort every column, open a detail sheet, and compare up to four models with radar, bars, and win counts.
Evidence: Explorer.tsx holds one shared state pipeline; CompareModal and DetailModal render as bottom sheets on mobile with Escape and scroll-lock handling.
The submission ships a headless QA script covering search, lab filtering, sorting, the compare flow, the table, and modals, plus desktop and mobile screenshots of each state.
Evidence: qa/qa.mjs and the submitted QA screenshots; Abundance reran lint, typecheck, build, and preview checks before publishing.
The data note states that some recent models carry estimated values, but individual cells do not distinguish estimates from measured results.
Evidence: README 'About the data' section versus the unmarked numeric cells in data/models.ts.
Scores compile from official announcements and public leaderboards, but model records do not attach source links a visitor can follow.
Evidence: data/models.ts records and the README sourcing note; contrast with submissions that attach per-model URLs.
How Zcode 5.3 did it
The submitted package pins Next.js and eslint-config-next to 16.2.11, and the supplied verification report records a clean production build with no compilation or type errors.
Evidence: Submitted package.json and July 22, 2026 build verification report.
Visitors can filter, sort, switch between grid and table layouts, inspect details, compare up to four models, use the recommendation wizard, and explore a price-versus-performance chart.
Evidence: Submitted Home, FilterBar, ModelGrid, ModelTable, ComparisonDrawer, ModelDetailModal, UseCaseWizard, and PerformanceChart components.
Its 18 records foreground DeepSeek R1, o3-mini, Claude 3.5, Gemini 2.0, GPT-4o, Llama 3.x, and other 2024/early-2025 entries rather than the latest model families represented elsewhere in the benchmark collection.
Evidence: src/data/modelsData.ts in the submitted project.
The typed records expose precise ELO, MMLU-Pro, GPQA, HumanEval, MATH-500, SWE-bench, speed, and pricing values, but users cannot trace individual claims to primary sources or a shared evaluation harness.
Evidence: The submitted model data contains values and descriptions but no per-claim source links or methodology module.
How Gemini 3.6 Flash High did it
The card UI, filters, compare drawer, modal detail, and recommendation wizard make the app easy to try quickly.
Evidence: project.zip includes FilterPanel, ModelCard, CompareDrawer, RecommendationWizard, and RadarChart components.
The compare drawer and wizard are well-suited to a small-screen model picker.
Evidence: The preview uses a sticky comparison drawer and modal-style recommendation flow.
The prompt asked for the latest AI models, but this dataset centers older families such as GPT-4o, o1, Claude 3.5, Gemini 1.5/2.0, Llama 3.x, and Mistral Large 2.
Evidence: Gemini data lacks the GPT-5.5, Claude 4.x/5-style, Gemini 3.x, GLM-5.2, and Grok 4.x-style coverage seen elsewhere.
Because older models are framed as current leaders, the app risks giving users outdated or hallucinated recommendations.
Evidence: Submitted model notes present older model families as the useful current benchmark set.
The app gives benchmark values, but does not provide enough visible source trail for users to trust the numbers.
Evidence: The submitted preview exposes benchmark fields without claim-level citations.
How Gemini 3.5 Flash High did it
The app has leaderboard, compare, about/methodology, and per-model detail pages rather than only a single dashboard.
Evidence: ai-bench.zip source includes /, /about, /compare, and /models/[id] routes.
The compare tray, provider marks, radar chart, sortable leaderboard, and detail pages make it feel like a real tool.
Evidence: Leaderboard, CompareTray, RadarChart, BenchmarkBars, and ProviderMark components.
The model lineup and benchmark numbers are presented confidently and should be checked before being treated as facts.
Evidence: Source data contains current-looking frontier names and exact benchmark scores.
The app is strong for users who know how to compare metrics, but less guided for users who want a step-by-step model recommendation.
Evidence: No recommendation wizard route or component was present in the inspected source.
How GLM 5.2 did it
The source data has source records, caution fields, and partial-data handling instead of presenting every number as equally certain.
Evidence: ai-model-benchmarks-source.zip data module includes sources and caution metadata.
The app supports search, filters, task-weighted sorting, compare selections, benchmark drilldowns, and method notes.
Evidence: Submitted completion report and preserved model-explorer component.
The app is more careful than the others, but the underlying GPT-5.5-era model names and figures still need checking.
Evidence: Dataset includes current-looking frontier model names and benchmark fields.
How GPT 5.5 High did it
The source separates data, filtering, cards, details, stats, lab badges, and compare UI into readable modules.
Evidence: Archive/Composer/Grok Composer/src/components, data, and lib folders.
Search, sorting, filters, model detail, and compare all work as expected for a model-selection dashboard.
Evidence: ModelExplorer wires FilterPanel, ModelCard, ModelDetail, ComparePanel, and StatsOverview.
The app includes benchmark-heavy model data but does not visibly attach claim-level citations to the numbers.
Evidence: The preview footer has a general sourcing note; the cards and detail views do not expose per-claim sources.
It is a strong explorer, but it does not guide a non-technical user through a recommendation flow.
Evidence: No wizard component or explicit step-by-step recommendation route exists in the Composer source.
How Composer 2.5 did it
The submitted output gives users cards, table view, filters, quick presets, compare, details, cost estimates, and CSV export.
Evidence: Archive/Grok/app/page.tsx implements the full ModelBench single-page dashboard.
The preview now uses the submitted black background, white text, zinc panels, and compact mobile-first controls.
Evidence: Archive/Grok/app/layout.tsx and globals.css set bg-zinc-950/text-zinc-200 styling.
The source comment says the data is real-ish and synthesized from leaderboards, Artificial Analysis, and provider announcements.
Evidence: Archive/Grok/app/page.tsx dataset comment above ALL_MODELS.
The app names current-looking models and exact benchmark scores, but does not attach source links or evidence to individual claims.
Evidence: Footer contains a general verification warning instead of per-model sources.
How Grok Build did it
Seven rows combine provider identity, status, context, four meters, price, and comparison selection without turning the page into a card mosaic.
Evidence: Model list, at-a-glance rail, and responsive row CSS in the submitted source.
The navigation, hero, filters, model meters, insights, briefing, and fixed selection tray all reflow for narrow screens.
Evidence: The source includes a dedicated max-width 760px layout and passed the submitted production build.
Up to three models can be queued and removed, but clicking Compare does not reveal a table, drawer, modal, or new route.
Evidence: compare-cta is rendered without an onClick handler or link target.
Reasoning, coding, vision, and speed are stored as exact 0-100 numbers, while the page only says they are normalized across available evaluations.
Evidence: Typed model array, scoreLabels metadata, and list-heading disclosure; no calculation or source fields exist in the source app.
How Codex GPT-5 did it
The methodology disclosure states that lab-reported values are not normalized and that benchmark families and tool settings cannot be treated as interchangeable.
Evidence: Read methodology panel and comparison-modal footnote in the submitted page.
Visitors can narrow seven models, select up to three, and open a side-by-side modal covering lab, GPQA, SWE-Pro, context, weights, and best fit.
Evidence: Filter state, toggleCompare cap, compare dock, and compare modal in app/page.tsx.
Null results remain visually blank but are coerced to zero for score sorting, which is a practical ordering rule rather than a benchmark result.
Evidence: compareValue returns zero for null before the GPQA and SWE-Pro sort comparators run.
Provider pages are attached at model level, not to each benchmark cell, so exact values still require checking against their original evaluation conditions.
Evidence: Each typed model record contains one source and sourceLabel alongside up to five metric fields.
How Codex GPT-5 did it
Six one-tap presets reshape filters and sorting for frontier, coding, science, value, open-weight, and speed decisions.
Evidence: PRESETS metadata, filter reducer, FilterBar, and ExplorerProvider.
The featured Cursor Grok 4.5 record publishes AA Index 54 and price/context facts while leaving five benchmark families and speed null.
Evidence: cursor-grok-4-5 entry in src/data/models.ts.
The source calls Cursor Grok 4.5 a jointly trained SpaceXAI/Cursor frontier model and ranks it #4 from AA Index 54 without an attached claim-level source.
Evidence: README, featured model record, and methodology source list.
Methodology links Arena, Artificial Analysis, SWE-bench, and GPQA, but model records do not identify which source supports each exact number.
Evidence: SOURCES metadata and model records.
How Cursor Grok 4.5 High did it
The source prerenders the explorer, comparison, lab, methodology, and model-detail surfaces from shared typed data.
Evidence: README route inventory plus app, component, data, and query modules in glm-5.2-coder.
Pinned models persist in localStorage, feed a sticky mobile tray, and appear in a winner-highlighted table and hand-rolled radar chart.
Evidence: PinProvider, CompareTray, CompareView, and RadarChart source modules.
The toolbar becomes a bottom sheet, the first comparison column stays sticky, and the compare action remains thumb-reachable.
Evidence: Toolbar, CompareView, CompareTray, responsive class contracts, and README UX notes.
The methodology names vendor tables and independent leaderboards, but exact cells are not connected to individual source URLs.
Evidence: Dataset header, README data-source paragraph, and methodology route.
How GLM 5.2 did it
Visitors can search, filter by lab and openness, rank by job, sort four ways, inspect benchmark detail, plot cost against capability, and compare up to three models.
Evidence: The submitted model-explorer component implements every interaction directly and keeps mobile controls touch-sized.
Every model carries a first-party source, missing values remain missing, and the page labels fit scores as an editorial decision aid rather than a published benchmark.
Evidence: Typed model records, source links, price states, tradeoff copy, and the visible methodology section.
The source includes three automated dataset checks and the submitted report records lint, type, build, responsive browser, gateway, and production dependency-audit verification.
Evidence: tests/model-data.test.ts plus the submitted completion report; Abundance reruns its own release checks before publishing.
The app correctly warns that provider benchmark versions, reasoning effort, and run configurations differ, so the values should be read as directional evidence.
Evidence: Dataset header comment, model tradeoffs, README data notes, and on-page methodology copy.
How GPT 5.6 Sol High did it
Strong means the artifact handled that area well. Weak means users should treat the submitted claims with extra skepticism before using the app for model selection.
| Build | Latest coverage | Sourcing | Hallucination risk | Benchmark verification | Interaction quality |
|---|---|---|---|---|---|
| Astra Strong | strong Launch-week coverage dated Sep 4, 2026: GPT-6 Astra rolling out, Gemini 3.8 Flash dated 2026-09-02, Claude Fable 5.1 dated 2026-09-01, plus current generation families from ten labs. | strong Per-benchmark source links with per-result overrides, plus documentation and pricing sources on every record — the only submission in the gallery where each reported figure opens its own source. | mixed Unverified values are left blank and one unverifiable DeepSWE result is omitted outright, but the underlying figures are publisher-reported, sometimes via competitor comparison tables. | mixed Benchmark versions are preserved and harness conditions are recorded per model (GLM-5.3's mini-swe-agent at temperature 0.95 with a six-hour timeout, Terminal 2.1 under Claude Code 2.1.207), yet nothing is independently rerun. | strong Filters, two-way sorting, chart and table views, mobile cards, native-dialog compare, CSV export, URL-shareable state, Cmd+K search, and a localStorage shortlist all work, with axe-clean QA and reduced-motion support. |
| Fable Strong | strong Releases run through 2026-08-18 (GLM-5.3 and GLM-5.3 Flash) under an Aug 28 snapshot; the one delayed model is labeled rather than hidden. | mixed Aggregate attribution with a dated footer line, but neither per-model citations nor the collected lab URLs appear in the UI. | mixed Editorial voice is consistently hedged and gap states are explicit, while cost-per-task and arena figures ride on unpublished harnesses. | mixed Seven metrics with full what/why metadata, including an explicit treat-gaps-under-100-Elo-as-noise caution, but no re-run or verification trail. | strong Search, eight sorts, license/context/modality/lab filters, job re-ranking, four-way compare, and per-model pages all work, with state synced to the URL. |
| Muse Spark 1.3 Strong | strong Newest records are launch-week releases dated 2026-09-02 (Gemini 3.8 Flash, Muse Spark 1.3, Qwen3.8-Max); every current-generation family the other submissions cover is present. | mixed A compiled-on date and a methodology section exist, but attribution is prose-only and no figure links to a source. | mixed Per-record caveats and re-run deltas are quoted honestly, but launch-week claims such as 'released today' rest entirely on vendor posts. | mixed All eight columns are vendor-reported; the app quotes independent deltas (V4-Pro SWE 77.6 vs 80.6 claimed) rather than normalizing them into the board. | strong Search, filters, toggles, eleven sorts, dual views, modal detail, and capped compare all work; the mobile view switcher is the one reachable-but-buried control. |
| Ling 3.0 Flash Needs Review | weak Newest records are GPT-4.5 and Gemini 2.0 Flash from early 2025; GPT-5.x, Claude 4.x+, and Gemini 3.x are absent. | weak Benchmark values carry descriptions but no sources or snapshot date. | weak Stale releases are presented as the latest models in the page metadata and hero. | weak Includes retired benchmarks (GSM8K, HellaSwag) without dates or harness notes. | strong Search, filters, five sort modes, benchmark bars, four-way compare, and detail pages all work. |
| Muse Spark 1.2 Needs Review | weak Newest records are mid-2025 releases (Claude 4, Gemini 2.5 Pro, Grok 3); every 2026 generation is absent. | weak No sources, methodology, or snapshot date appear anywhere in the app. | weak A 'Live' title over a stale dataset invites readers to trust outdated standings. | weak Ten benchmark columns are populated without any stated harness, date, or verification trail. | strong Search, filters, sorting, dual views, benchmark focus, detail modal, and compare all work. |
| Grok 4.6 Grok Build Strong | strong 34 models across 12 labs compiled August 28, 2026, including GPT-5.6, Claude Mythos/Fable, Gemini 3.7 Flash, GLM-5.3, Kimi K3, and Muse Spark 1.2. | mixed Public leaderboard sources are named in the methodology but cells carry no per-claim links. | strong Missing cells stay null and restricted-access models are not presented as purchasable options. | mixed Scores compile public leaderboards with stated caveats, but nothing is independently rerun. | strong Search, filters, three view modes, use-case picks, docked four-way compare, and per-model guidance all work across routes. |
| Gemini 3.7 Flash High Needs Review | weak Snapshot centers Claude 3.7 Sonnet, o1, and Gemini 2.0 from late 2024/early 2025 and misses every 2026 generation. | weak No per-claim citations; benchmark values are unverifiable from the app. | weak Stale models are framed as current leaders with badges like 'King' and 'Best Value'. | weak Ten benchmark families are reported without a stated harness, date, or verification trail. | strong Wizard, filters, dual views, scatterplot, detail modal, comparison dock, and glossary all work with no dead controls. |
| Zcode 5.3 Flash Strong | strong 38 models across 13 labs compiled August 26, 2026, including GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, and GLM-5.3. | strong Model pages attach primary source links and name the independent leaderboards consulted. | strong Null preservation is explicit and enforced—dashes for unpublished scores, no estimated fills, prices absent where none are published. | mixed Vendor-reported figures on vendor harnesses with disputes noted, but no independent rerun normalizes them. | strong Tabs, filters, sorting, four-way compare with takeaways, and model pages all work with keyboard focus states and touch-sized targets. |
| Zcode 5.3 Strong | strong 32 models across 13 labs compiled August 2026, including GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro. | mixed Aggregates official announcements and public leaderboards but omits per-record source links. | mixed Estimated values are admitted in aggregate rather than marked per cell, so readers cannot tell which numbers are estimates. | mixed No independent rerun; the README warns that harness, sampling, and eval-date differences move the figures. | strong Search, filters, dual views, column sorting, medals, detail sheets, and four-way comparison all run client-side with no dead controls. |
| Gemini 3.6 Flash High Needs Review | weak A polished but static 2024/early-2025 model snapshot rather than a current frontier roster. | weak Exact values have no visible claim-level citations. | mixed The app is explicit about its fields, but unsourced precision and stale positioning need verification before model-selection use. | weak No visible cross-provider normalization or independent benchmark rerun is supplied. | strong Filtering, sorting, comparison, recommendations, analytics, details, and mobile navigation are all implemented. |
| Gemini 3.5 Flash High Needs Review | weak Fails the latest-model requirement; dataset appears stale. | weak No strong source trail for exact benchmark claims. | weak Older models are framed as current recommendations. | weak Benchmark claims need verification and likely updating. | strong Wizard, filters, detail modal, and compare drawer are useful. |
| GLM 5.2 Strong | strong Broad current-looking frontier model coverage. | mixed Includes methodology/about surface, but claim-level evidence is limited. | mixed Confident exact numbers and model names need verification. | mixed Methodology page helps but does not make this a certified benchmark feed. | strong Best multi-route comparison experience. |
| GPT 5.5 High Strong | strong Broad model set with GPT, Claude, Gemini, Grok, GLM, DeepSeek, Qwen, Meta, Mistral, and Cohere. | strong Best source and caution scaffolding among the submissions. | mixed Lower than others, but current-looking model claims still need verification. | mixed Method notes help; independent verification is still out of scope here. | strong Search, filter, sort, compare, drilldowns, and recommendation-style weighting are present. |
| Composer 2.5 Mixed | strong Includes many 2026-style model names across major labs. | weak Generic sourcing note, not claim-level citations. | mixed Current-looking names and scores require verification. | weak Benchmark fields are asserted in data files without evidence records. | strong Search, filter, sort, details, compare, and mobile controls are present. |
| Grok Build Mixed | mixed Covers current-looking 2026 model names, but several claims need verification. | weak General source note only; no claim-level citations. | weak The submitted source calls the data real-ish and synthesized. | weak Exact benchmark scores are not backed by visible source records. | strong Filtering, sorting, compare, details, export, and cost estimates all exist. |
| Terra Codex GPT-5 Mixed | mixed Seven-model snapshot includes 2026 flagships but also older 2025 reference models. | weak The completion report names official lab pages, but the executable source contains no claim-level URLs or score provenance. | mixed The benchmark warning helps, but unexplained normalized scores can look more objective than the evidence supports. | weak No normalization formula or reproducible evaluation harness is included. | mixed Search, chips, sorting, and selection work; compare, detail, and secondary mobile filter buttons are incomplete. |
| Luna Codex GPT-5 Strong | strong Seven-provider July 15, 2026 snapshot including GPT-5.5, Opus 4.8, Gemini 3.1 Pro, and Grok 4.5. | mixed Every model links an official provider page, but exact benchmark cells are not individually cited. | strong Null fields remain null and the app repeatedly warns against apples-to-apples interpretation. | mixed The source explains evaluation mismatch; Abundance publishes the artifact rather than reproducing the lab runs. | strong Responsive filtering, sorting, mobile cards, methodology reveal, and a working comparison modal. |
| Cursor Grok 4.5 High Mixed | strong Nineteen models across 11 labs, including a separate SpaceXAI identity for Cursor Grok 4.5. | mixed Four authoritative benchmark-family links are visible, but exact model claims are not mapped to those sources. | mixed Null handling is good; the self-featured joint-training and rank claims still need independent verification. | mixed AA Index, Arena, SWE-bench, and GPQA are explained, but Abundance did not reproduce the submitted numbers. | strong Task presets, deep filters, nine sorts, model details, and capped comparison make the snapshot useful. |
| GLM 5.2 Cursor Strong | strong Twenty models across nine labs, spanning proprietary, open-weight, fast, and frontier tiers. | mixed Source families and harness caveats are named, but individual model-score records lack claim-level links. | mixed Missing data is tolerated, yet many precise mid-2026 values are asserted rather than independently reproduced. | mixed The app explains harness variance and saturation; Abundance did not rerun the 10 benchmark suites. | strong Filters, metric sorting, persistent four-model comparison, radar visualization, detail routes, and mobile states are all implemented. |
| GPT 5.6 Sol High Strong | mixed Final snapshot covers nine selected models across seven labs rather than claiming exhaustive market coverage. | strong Every model has a first-party HTTPS source and a human-readable source label. | strong Missing prices and scores stay explicit; self-hosted and not-listed states are not converted into guesses. | mixed Provider claims are source-linked and caveated, but no independent rerun normalizes the different harnesses. | strong Search, filters, use-case ranking, sorting, detail, scatterplot, and capped comparison all work in one surface. |
This is a product-surface comparison only. It checks whether each submitted app exposed the interaction pattern in its source or build output.
| Build | Filters | Sorting | Compare | Model detail | Recommendation | Lab branding | Mobile-first UI |
|---|---|---|---|---|---|---|---|
| Astra Benchmarks/Astra | |||||||
| Fable Benchmarks/Fable/fable | |||||||
| Muse Spark 1.3 Benchmarks/Musespark 1.3 | |||||||
| Ling 3.0 Flash Benchmarks/Ling 3.0 Flash/ai-models-benchmarks | |||||||
| Muse Spark 1.2 Benchmarks/Muse Spark 1.2/bench | |||||||
| Grok 4.6 Grok Build Benchmarks/Grok 4.6 Grok Build | |||||||
| Gemini 3.7 Flash High Benchmarks/Gemini Flash 3.7 | |||||||
| Zcode 5.3 Flash Benchmarks/Zcode 5.3 Flash/modelpulse | |||||||
| Zcode 5.3 Benchmarks/Zcode 5.3/modeldex | |||||||
| Gemini 3.6 Flash High E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash | |||||||
| Gemini 3.5 Flash High project root | |||||||
| GLM 5.2 ai-bench/ | |||||||
| GPT 5.5 High src/components/model-explorer.tsx and src/data/models.ts | |||||||
| Composer 2.5 Archive/Composer/Grok Composer | |||||||
| Grok Build Archive/Grok | |||||||
| Terra Codex GPT-5 C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas | |||||||
| Luna Codex GPT-5 C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5 | |||||||
| Cursor Grok 4.5 High E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5 | |||||||
| GLM 5.2 Cursor E:/Projects/Benchmark Tests/GLM/glm-5.2-coder | |||||||
| GPT 5.6 Sol High components/model-explorer.tsx and data/models.ts |
These previews are a way to inspect the work produced by each builder. The report card calls out obvious artifact issues, but it is not a live benchmark feed or a certified source for model-selection advice.