This page compares the submitted Next.js benchmark apps as build outputs. The previews preserve the model builders' own interfaces, while the review flags stale data, weak sourcing, and likely hallucination risk without certifying the benchmark claims.
Review mode
Findings are based on the submitted artifacts, visible app behavior, and peer comparison. This is not a full source-backed audit of every benchmark number.
11
Submissions
11
Live source previews
5/11
Build times
The benchmark explorer leads with current top models, cost/quality scatter plots, coding vs agentic charts, coverage heatmaps, and a sticky compare tray.
Open explorerEach link opens the submitted app's own interface, with Abundance chrome removed except for a small back button.
The strongest complete newcomer: Sol shipped a restrained, mobile-first decision tool with unusually clear benchmark caveats, useful interactions, and executable source tests.
pass
3
watch
1
fail
0
Source
components/model-explorer.tsx and data/models.ts
Features
7/7
The most structurally ambitious result in this round. GLM built a real multi-route benchmark product with excellent comparison ergonomics, while its exact scores still require source-by-source verification.
pass
3
watch
1
fail
0
Source
E:/Projects/Benchmark Tests/GLM/glm-5.2-coder
Features
6/7
A genuinely different and useful decision surface, especially for task presets and practical constraints. Its self-featured Grok result is candid about missing cells but still needs independent claim-level verification.
pass
2
watch
2
fail
0
Source
E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5
Features
7/7
Luna is the more rigorous of the two additions: its benchmark caveats, blank-value handling, mobile table alternative, and functioning comparison modal make the artifact useful without pretending the provider results are normalized.
pass
2
watch
2
fail
0
Source
C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5
Features
6/7
Terra is visually polished and fast to scan, but its main comparison and detail affordances stop short of a result. The normalized numbers should be treated as editorial fit signals until a formula and claim-level sources are published.
pass
2
watch
2
fail
0
Source
C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas
Features
6/7
Best for interface completeness, but the data layer is openly approximate and carries real hallucination risk.
pass
2
watch
1
fail
1
Source
Archive/Grok
Features
6/7
A clean, usable explorer with strong dark UI fidelity, but its benchmark claims are still mostly asserted rather than evidenced.
pass
2
watch
2
fail
0
Source
Archive/Composer/Grok Composer
Features
6/7
The most evidence-aware submission: it still needs external verification, but it did the best job separating useful UI from caveats and source notes.
pass
2
watch
1
fail
0
Source
src/components/model-explorer.tsx and src/data/models.ts
Features
7/7
The broadest app structure, with real routes and strong comparison surfaces; the main weakness is that the confident benchmark data still needs verification.
pass
2
watch
2
fail
0
Source
ai-bench/
Features
6/7
Fast and approachable, but it fails the latest-model requirement: the submitted dataset looks stale next to the other entries.
pass
2
watch
1
fail
2
Source
project root
Features
7/7
Gemini 3.6 Flash High delivered the most complete interaction surface of the Gemini submissions and now builds on Next.js 16.2.11, but its data remains an older static benchmark snapshot that needs sourcing and modernization.
pass
2
watch
2
fail
0
Source
E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash
Features
7/7
These findings review the submitted apps as artifacts. They flag obvious stale data, weak sourcing, and likely hallucination risk, but they are not a complete fact-check of every AI benchmark claim.
Visitors can search, filter by lab and openness, rank by job, sort four ways, inspect benchmark detail, plot cost against capability, and compare up to three models.
Evidence: The submitted model-explorer component implements every interaction directly and keeps mobile controls touch-sized.
Every model carries a first-party source, missing values remain missing, and the page labels fit scores as an editorial decision aid rather than a published benchmark.
Evidence: Typed model records, source links, price states, tradeoff copy, and the visible methodology section.
The source includes three automated dataset checks and the submitted report records lint, type, build, responsive browser, gateway, and production dependency-audit verification.
Evidence: tests/model-data.test.ts plus the submitted completion report; Abundance reruns its own release checks before publishing.
The app correctly warns that provider benchmark versions, reasoning effort, and run configurations differ, so the values should be read as directional evidence.
Evidence: Dataset header comment, model tradeoffs, README data notes, and on-page methodology copy.
How GPT 5.6 Sol High did it
The source prerenders the explorer, comparison, lab, methodology, and model-detail surfaces from shared typed data.
Evidence: README route inventory plus app, component, data, and query modules in glm-5.2-coder.
Pinned models persist in localStorage, feed a sticky mobile tray, and appear in a winner-highlighted table and hand-rolled radar chart.
Evidence: PinProvider, CompareTray, CompareView, and RadarChart source modules.
The toolbar becomes a bottom sheet, the first comparison column stays sticky, and the compare action remains thumb-reachable.
Evidence: Toolbar, CompareView, CompareTray, responsive class contracts, and README UX notes.
The methodology names vendor tables and independent leaderboards, but exact cells are not connected to individual source URLs.
Evidence: Dataset header, README data-source paragraph, and methodology route.
How GLM 5.2 did it
Six one-tap presets reshape filters and sorting for frontier, coding, science, value, open-weight, and speed decisions.
Evidence: PRESETS metadata, filter reducer, FilterBar, and ExplorerProvider.
The featured Cursor Grok 4.5 record publishes AA Index 54 and price/context facts while leaving five benchmark families and speed null.
Evidence: cursor-grok-4-5 entry in src/data/models.ts.
The source calls Cursor Grok 4.5 a jointly trained SpaceXAI/Cursor frontier model and ranks it #4 from AA Index 54 without an attached claim-level source.
Evidence: README, featured model record, and methodology source list.
Methodology links Arena, Artificial Analysis, SWE-bench, and GPQA, but model records do not identify which source supports each exact number.
Evidence: SOURCES metadata and model records.
How Cursor Grok 4.5 High did it
The methodology disclosure states that lab-reported values are not normalized and that benchmark families and tool settings cannot be treated as interchangeable.
Evidence: Read methodology panel and comparison-modal footnote in the submitted page.
Visitors can narrow seven models, select up to three, and open a side-by-side modal covering lab, GPQA, SWE-Pro, context, weights, and best fit.
Evidence: Filter state, toggleCompare cap, compare dock, and compare modal in app/page.tsx.
Null results remain visually blank but are coerced to zero for score sorting, which is a practical ordering rule rather than a benchmark result.
Evidence: compareValue returns zero for null before the GPQA and SWE-Pro sort comparators run.
Provider pages are attached at model level, not to each benchmark cell, so exact values still require checking against their original evaluation conditions.
Evidence: Each typed model record contains one source and sourceLabel alongside up to five metric fields.
How Codex GPT-5 did it
Seven rows combine provider identity, status, context, four meters, price, and comparison selection without turning the page into a card mosaic.
Evidence: Model list, at-a-glance rail, and responsive row CSS in the submitted source.
The navigation, hero, filters, model meters, insights, briefing, and fixed selection tray all reflow for narrow screens.
Evidence: The source includes a dedicated max-width 760px layout and passed the submitted production build.
Up to three models can be queued and removed, but clicking Compare does not reveal a table, drawer, modal, or new route.
Evidence: compare-cta is rendered without an onClick handler or link target.
Reasoning, coding, vision, and speed are stored as exact 0-100 numbers, while the page only says they are normalized across available evaluations.
Evidence: Typed model array, scoreLabels metadata, and list-heading disclosure; no calculation or source fields exist in the source app.
How Codex GPT-5 did it
The submitted output gives users cards, table view, filters, quick presets, compare, details, cost estimates, and CSV export.
Evidence: Archive/Grok/app/page.tsx implements the full ModelBench single-page dashboard.
The preview now uses the submitted black background, white text, zinc panels, and compact mobile-first controls.
Evidence: Archive/Grok/app/layout.tsx and globals.css set bg-zinc-950/text-zinc-200 styling.
The source comment says the data is real-ish and synthesized from leaderboards, Artificial Analysis, and provider announcements.
Evidence: Archive/Grok/app/page.tsx dataset comment above ALL_MODELS.
The app names current-looking models and exact benchmark scores, but does not attach source links or evidence to individual claims.
Evidence: Footer contains a general verification warning instead of per-model sources.
How Grok Build did it
The source separates data, filtering, cards, details, stats, lab badges, and compare UI into readable modules.
Evidence: Archive/Composer/Grok Composer/src/components, data, and lib folders.
Search, sorting, filters, model detail, and compare all work as expected for a model-selection dashboard.
Evidence: ModelExplorer wires FilterPanel, ModelCard, ModelDetail, ComparePanel, and StatsOverview.
The app includes benchmark-heavy model data but does not visibly attach claim-level citations to the numbers.
Evidence: The preview footer has a general sourcing note; the cards and detail views do not expose per-claim sources.
It is a strong explorer, but it does not guide a non-technical user through a recommendation flow.
Evidence: No wizard component or explicit step-by-step recommendation route exists in the Composer source.
How Composer 2.5 did it
The source data has source records, caution fields, and partial-data handling instead of presenting every number as equally certain.
Evidence: ai-model-benchmarks-source.zip data module includes sources and caution metadata.
The app supports search, filters, task-weighted sorting, compare selections, benchmark drilldowns, and method notes.
Evidence: Submitted completion report and preserved model-explorer component.
The app is more careful than the others, but the underlying GPT-5.5-era model names and figures still need checking.
Evidence: Dataset includes current-looking frontier model names and benchmark fields.
How GPT 5.5 High did it
The app has leaderboard, compare, about/methodology, and per-model detail pages rather than only a single dashboard.
Evidence: ai-bench.zip source includes /, /about, /compare, and /models/[id] routes.
The compare tray, provider marks, radar chart, sortable leaderboard, and detail pages make it feel like a real tool.
Evidence: Leaderboard, CompareTray, RadarChart, BenchmarkBars, and ProviderMark components.
The model lineup and benchmark numbers are presented confidently and should be checked before being treated as facts.
Evidence: Source data contains current-looking frontier names and exact benchmark scores.
The app is strong for users who know how to compare metrics, but less guided for users who want a step-by-step model recommendation.
Evidence: No recommendation wizard route or component was present in the inspected source.
How GLM 5.2 did it
The card UI, filters, compare drawer, modal detail, and recommendation wizard make the app easy to try quickly.
Evidence: project.zip includes FilterPanel, ModelCard, CompareDrawer, RecommendationWizard, and RadarChart components.
The compare drawer and wizard are well-suited to a small-screen model picker.
Evidence: The preview uses a sticky comparison drawer and modal-style recommendation flow.
The prompt asked for the latest AI models, but this dataset centers older families such as GPT-4o, o1, Claude 3.5, Gemini 1.5/2.0, Llama 3.x, and Mistral Large 2.
Evidence: Gemini data lacks the GPT-5.5, Claude 4.x/5-style, Gemini 3.x, GLM-5.2, and Grok 4.x-style coverage seen elsewhere.
Because older models are framed as current leaders, the app risks giving users outdated or hallucinated recommendations.
Evidence: Submitted model notes present older model families as the useful current benchmark set.
The app gives benchmark values, but does not provide enough visible source trail for users to trust the numbers.
Evidence: The submitted preview exposes benchmark fields without claim-level citations.
How Gemini 3.5 Flash High did it
The submitted package pins Next.js and eslint-config-next to 16.2.11, and the supplied verification report records a clean production build with no compilation or type errors.
Evidence: Submitted package.json and July 22, 2026 build verification report.
Visitors can filter, sort, switch between grid and table layouts, inspect details, compare up to four models, use the recommendation wizard, and explore a price-versus-performance chart.
Evidence: Submitted Home, FilterBar, ModelGrid, ModelTable, ComparisonDrawer, ModelDetailModal, UseCaseWizard, and PerformanceChart components.
Its 18 records foreground DeepSeek R1, o3-mini, Claude 3.5, Gemini 2.0, GPT-4o, Llama 3.x, and other 2024/early-2025 entries rather than the latest model families represented elsewhere in the benchmark collection.
Evidence: src/data/modelsData.ts in the submitted project.
The typed records expose precise ELO, MMLU-Pro, GPQA, HumanEval, MATH-500, SWE-bench, speed, and pricing values, but users cannot trace individual claims to primary sources or a shared evaluation harness.
Evidence: The submitted model data contains values and descriptions but no per-claim source links or methodology module.
How Gemini 3.6 Flash High did it
Strong means the artifact handled that area well. Weak means users should treat the submitted claims with extra skepticism before using the app for model selection.
| Build | Latest coverage | Sourcing | Hallucination risk | Benchmark verification | Interaction quality |
|---|---|---|---|---|---|
| GPT 5.6 Sol High Strong | mixed Final snapshot covers nine selected models across seven labs rather than claiming exhaustive market coverage. | strong Every model has a first-party HTTPS source and a human-readable source label. | strong Missing prices and scores stay explicit; self-hosted and not-listed states are not converted into guesses. | mixed Provider claims are source-linked and caveated, but no independent rerun normalizes the different harnesses. | strong Search, filters, use-case ranking, sorting, detail, scatterplot, and capped comparison all work in one surface. |
| GLM 5.2 Cursor Strong | strong Twenty models across nine labs, spanning proprietary, open-weight, fast, and frontier tiers. | mixed Source families and harness caveats are named, but individual model-score records lack claim-level links. | mixed Missing data is tolerated, yet many precise mid-2026 values are asserted rather than independently reproduced. | mixed The app explains harness variance and saturation; Abundance did not rerun the 10 benchmark suites. | strong Filters, metric sorting, persistent four-model comparison, radar visualization, detail routes, and mobile states are all implemented. |
| Cursor Grok 4.5 High Mixed | strong Nineteen models across 11 labs, including a separate SpaceXAI identity for Cursor Grok 4.5. | mixed Four authoritative benchmark-family links are visible, but exact model claims are not mapped to those sources. | mixed Null handling is good; the self-featured joint-training and rank claims still need independent verification. | mixed AA Index, Arena, SWE-bench, and GPQA are explained, but Abundance did not reproduce the submitted numbers. | strong Task presets, deep filters, nine sorts, model details, and capped comparison make the snapshot useful. |
| Luna Codex GPT-5 Strong | strong Seven-provider July 15, 2026 snapshot including GPT-5.5, Opus 4.8, Gemini 3.1 Pro, and Grok 4.5. | mixed Every model links an official provider page, but exact benchmark cells are not individually cited. | strong Null fields remain null and the app repeatedly warns against apples-to-apples interpretation. | mixed The source explains evaluation mismatch; Abundance publishes the artifact rather than reproducing the lab runs. | strong Responsive filtering, sorting, mobile cards, methodology reveal, and a working comparison modal. |
| Terra Codex GPT-5 Mixed | mixed Seven-model snapshot includes 2026 flagships but also older 2025 reference models. | weak The completion report names official lab pages, but the executable source contains no claim-level URLs or score provenance. | mixed The benchmark warning helps, but unexplained normalized scores can look more objective than the evidence supports. | weak No normalization formula or reproducible evaluation harness is included. | mixed Search, chips, sorting, and selection work; compare, detail, and secondary mobile filter buttons are incomplete. |
| Grok Build Mixed | mixed Covers current-looking 2026 model names, but several claims need verification. | weak General source note only; no claim-level citations. | weak The submitted source calls the data real-ish and synthesized. | weak Exact benchmark scores are not backed by visible source records. | strong Filtering, sorting, compare, details, export, and cost estimates all exist. |
| Composer 2.5 Mixed | strong Includes many 2026-style model names across major labs. | weak Generic sourcing note, not claim-level citations. | mixed Current-looking names and scores require verification. | weak Benchmark fields are asserted in data files without evidence records. | strong Search, filter, sort, details, compare, and mobile controls are present. |
| GPT 5.5 High Strong | strong Broad model set with GPT, Claude, Gemini, Grok, GLM, DeepSeek, Qwen, Meta, Mistral, and Cohere. | strong Best source and caution scaffolding among the submissions. | mixed Lower than others, but current-looking model claims still need verification. | mixed Method notes help; independent verification is still out of scope here. | strong Search, filter, sort, compare, drilldowns, and recommendation-style weighting are present. |
| GLM 5.2 Strong | strong Broad current-looking frontier model coverage. | mixed Includes methodology/about surface, but claim-level evidence is limited. | mixed Confident exact numbers and model names need verification. | mixed Methodology page helps but does not make this a certified benchmark feed. | strong Best multi-route comparison experience. |
| Gemini 3.5 Flash High Needs Review | weak Fails the latest-model requirement; dataset appears stale. | weak No strong source trail for exact benchmark claims. | weak Older models are framed as current recommendations. | weak Benchmark claims need verification and likely updating. | strong Wizard, filters, detail modal, and compare drawer are useful. |
| Gemini 3.6 Flash High Needs Review | weak A polished but static 2024/early-2025 model snapshot rather than a current frontier roster. | weak Exact values have no visible claim-level citations. | mixed The app is explicit about its fields, but unsourced precision and stale positioning need verification before model-selection use. | weak No visible cross-provider normalization or independent benchmark rerun is supplied. | strong Filtering, sorting, comparison, recommendations, analytics, details, and mobile navigation are all implemented. |
This is a product-surface comparison only. It checks whether each submitted app exposed the interaction pattern in its source or build output.
| Build | Filters | Sorting | Compare | Model detail | Recommendation | Lab branding | Mobile-first UI |
|---|---|---|---|---|---|---|---|
| GPT 5.6 Sol High components/model-explorer.tsx and data/models.ts | |||||||
| GLM 5.2 Cursor E:/Projects/Benchmark Tests/GLM/glm-5.2-coder | |||||||
| Cursor Grok 4.5 High E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5 | |||||||
| Luna Codex GPT-5 C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5 | |||||||
| Terra Codex GPT-5 C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas | |||||||
| Grok Build Archive/Grok | |||||||
| Composer 2.5 Archive/Composer/Grok Composer | |||||||
| GPT 5.5 High src/components/model-explorer.tsx and src/data/models.ts | |||||||
| GLM 5.2 ai-bench/ | |||||||
| Gemini 3.5 Flash High project root | |||||||
| Gemini 3.6 Flash High E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash |
These previews are a way to inspect the work produced by each builder. The report card calls out obvious artifact issues, but it is not a live benchmark feed or a certified source for model-selection advice.