Skip to content
Abundance Marketing Partners
Case StudiesHow We BuildLabsPricingAbout
Request a Leverage Review
Abundance Marketing Partners

Utah-based custom software, AI workflows, internal tools, and modern websites for owners across the United States.

Request a Leverage Review

Services

  • Custom Apps
  • AI Workflows & Agents
  • Internal Tools & MCP
  • Modern Websites
  • Delivery Infrastructure

Company

  • About
  • Case Studies
  • Abundance Labs
  • Pricing
  • Service Areas
  • Articles
  • How We Work
  • AI Benchmark Results

Contact

  • Request a Leverage Review
  • hello@abundancepartners.app

Want the operating logic behind the work?
Read how we work

Local service areas

Salt Lake CitySandyDraperSouth JordanLehiProvo
See all service areas

© 2026 Abundance Marketing Partners LLC. All rights reserved.

Privacy PolicyTerms of Service
AI Benchmark Results

11 benchmark builds, one place to inspect them.

This page compares the submitted Next.js benchmark apps as build outputs. The previews preserve the model builders' own interfaces, while the review flags stale data, weak sourcing, and likely hallucination risk without certifying the benchmark claims.

View submissions Read reviewModel explorer

Review mode

Blunt artifact report card

Findings are based on the submitted artifacts, visible app behavior, and peer comparison. This is not a full source-backed audit of every benchmark number.

11

Submissions

11

Live source previews

5/11

Build times

Top model explorer

Compare current models with real charts.

The benchmark explorer leads with current top models, cost/quality scatter plots, coding vs agentic charts, coverage heatmaps, and a sticky compare tray.

Open explorer
Preview routes

Open each submitted build.

Each link opens the submitted app's own interface, with Abundance chrome removed except for a small back button.

Preview
StrongPending
GPT 5.6 Sol High
GPT 5.6 Sol High via Codex

The strongest complete newcomer: Sol shipped a restrained, mobile-first decision tool with unusually clear benchmark caveats, useful interactions, and executable source tests.

pass

3

watch

1

fail

0

Source

components/model-explorer.tsx and data/models.ts

Features

7/7

Open original output
Preview
StrongPending
GLM 5.2 Cursor
GLM 5.2 via Cursor

The most structurally ambitious result in this round. GLM built a real multi-route benchmark product with excellent comparison ergonomics, while its exact scores still require source-by-source verification.

pass

3

watch

1

fail

0

Source

E:/Projects/Benchmark Tests/GLM/glm-5.2-coder

Features

6/7

Open original output
Preview
MixedPending
Cursor Grok 4.5 High
Cursor Grok 4.5 High via Cursor

A genuinely different and useful decision surface, especially for task presets and practical constraints. Its self-featured Grok result is candid about missing cells but still needs independent claim-level verification.

pass

2

watch

2

fail

0

Source

E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5

Features

7/7

Open original output
Preview
StrongPending
Luna Codex GPT-5
Codex GPT-5 via Codex

Luna is the more rigorous of the two additions: its benchmark caveats, blank-value handling, mobile table alternative, and functioning comparison modal make the artifact useful without pretending the provider results are normalized.

pass

2

watch

2

fail

0

Source

C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5

Features

6/7

Open original output
Preview
MixedPending
Terra Codex GPT-5
Codex GPT-5 via Codex

Terra is visually polished and fast to scan, but its main comparison and detail affordances stop short of a result. The normalized numbers should be treated as editorial fit signals until a formula and claim-level sources are published.

pass

2

watch

2

fail

0

Source

C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas

Features

6/7

Open original output
Preview
Mixed6m 8s
Grok Build
Grok Build via grok build

Best for interface completeness, but the data layer is openly approximate and carries real hallucination risk.

pass

2

watch

1

fail

1

Source

Archive/Grok

Features

6/7

Open original output
Preview
Mixed3m
Composer 2.5
Composer 2.5 via grok build

A clean, usable explorer with strong dark UI fidelity, but its benchmark claims are still mostly asserted rather than evidenced.

pass

2

watch

2

fail

0

Source

Archive/Composer/Grok Composer

Features

6/7

Open original output
Preview
Strong11m 40s
GPT 5.5 High
GPT 5.5 High via Codex

The most evidence-aware submission: it still needs external verification, but it did the best job separating useful UI from caveats and source notes.

pass

2

watch

1

fail

0

Source

src/components/model-explorer.tsx and src/data/models.ts

Features

7/7

Open original output
Preview
Strong23m 30s
GLM 5.2
GLM 5.2 via Zcode

The broadest app structure, with real routes and strong comparison surfaces; the main weakness is that the confident benchmark data still needs verification.

pass

2

watch

2

fail

0

Source

ai-bench/

Features

6/7

Open original output
Preview
Needs Review2m 23s
Gemini 3.5 Flash High
Gemini 3.5 Flash High via antigravity

Fast and approachable, but it fails the latest-model requirement: the submitted dataset looks stale next to the other entries.

pass

2

watch

1

fail

2

Source

project root

Features

7/7

Open original output
Preview
Needs ReviewPending
Gemini 3.6 Flash High
Gemini 3.6 Flash High via Antigravity

Gemini 3.6 Flash High delivered the most complete interaction surface of the Gemini submissions and now builds on Next.js 16.2.11, but its data remains an older static benchmark snapshot that needs sourcing and modernization.

pass

2

watch

2

fail

0

Source

E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash

Features

7/7

Open original output
Blunt review

What worked, what failed.

These findings review the submitted apps as artifacts. They flag obvious stale data, weak sourcing, and likely hallucination risk, but they are not a complete fact-check of every AI benchmark claim.

GPT 5.6 Sol High
The strongest complete newcomer: Sol shipped a restrained, mobile-first decision tool with unusually clear benchmark caveats, useful interactions, and executable source tests.
Strong
passux

Complete decision workflow

Visitors can search, filter by lab and openness, rank by job, sort four ways, inspect benchmark detail, plot cost against capability, and compare up to three models.

Evidence: The submitted model-explorer component implements every interaction directly and keeps mobile controls touch-sized.

passsourcing

Evidence and opinion stay separate

Every model carries a first-party source, missing values remain missing, and the page labels fit scores as an editorial decision aid rather than a published benchmark.

Evidence: Typed model records, source links, price states, tradeoff copy, and the visible methodology section.

passimplementation

The result is executable and tested

The source includes three automated dataset checks and the submitted report records lint, type, build, responsive browser, gateway, and production dependency-audit verification.

Evidence: tests/model-data.test.ts plus the submitted completion report; Abundance reruns its own release checks before publishing.

watchfactuality

Reported scores are not normalized

The app correctly warns that provider benchmark versions, reasoning effort, and run configurations differ, so the values should be read as directional evidence.

Evidence: Dataset header comment, model tradeoffs, README data notes, and on-page methodology copy.

How GPT 5.6 Sol High did it

  1. 1.Data method: Sol put nine models from seven labs in a typed module, attached a first-party HTTPS source to every record, preserved missing values, and documented provider-run differences.
  2. 2.Decision method: published benchmark fields remain separate from editorial fit scores derived from capability, price, modalities, openness, and deployment tradeoffs.
  3. 3.Interaction method: useDeferredValue powers search; pure render-time filters and immutable sorting drive lab, openness, scenario, price, context, and recency views; comparison is capped at three models.
  4. 4.Test method: three Vitest checks cover unique IDs and seven-lab coverage, usable price/context/fit ranges, and first-party HTTPS source metadata.
  5. 5.Submitted verification: ESLint, TypeScript, three Vitest tests, a production build, responsive Playwright QA, gateway QA/postflight, and a zero-vulnerability production audit. Abundance reruns release checks before publication.
GLM 5.2 Cursor
The most structurally ambitious result in this round. GLM built a real multi-route benchmark product with excellent comparison ergonomics, while its exact scores still require source-by-source verification.
Strong
passimplementation

Full product surface

The source prerenders the explorer, comparison, lab, methodology, and model-detail surfaces from shared typed data.

Evidence: README route inventory plus app, component, data, and query modules in glm-5.2-coder.

passux

Four-model comparison is meaningfully different

Pinned models persist in localStorage, feed a sticky mobile tray, and appear in a winner-highlighted table and hand-rolled radar chart.

Evidence: PinProvider, CompareTray, CompareView, and RadarChart source modules.

passmobile

Mobile interaction was designed, not retrofitted

The toolbar becomes a bottom sheet, the first comparison column stays sticky, and the compare action remains thumb-reachable.

Evidence: Toolbar, CompareView, CompareTray, responsive class contracts, and README UX notes.

watchsourcing

Sources are project-level

The methodology names vendor tables and independent leaderboards, but exact cells are not connected to individual source URLs.

Evidence: Dataset header, README data-source paragraph, and methodology route.

How GLM 5.2 did it

  1. 1.Data method: GLM placed 20 models, nine labs, 10 benchmark keys, price, context, license, modalities, strengths, and release state in typed data modules; null remains the missing-data state.
  2. 2.Ranking method: six named benchmark leaderboards compute the top three directly from model fields; a separate best-value list divides the reported benchmark average by a 3:1 blended input/output price.
  3. 3.Interaction method: toolbar state filters by search, lab, tier, modality, open weights, and price; users can sort by recency, name, price, context, or any of six surfaced benchmarks.
  4. 4.Comparison method: PinProvider persists up to four model IDs in localStorage; CompareView draws a hand-rolled SVG radar and highlights row winners while keeping the metric column sticky on mobile.
  5. 5.Build method: Next.js 15 App Router, React 19, strict TypeScript, Tailwind CSS v4, static prerendering, and zero runtime dependencies beyond React/Next in the submitted package.
  6. 6.Submitted verification: the completion report records a clean TypeScript check, clean production build, 27 prerendered pages, and 200 responses for all routes. Abundance reruns release checks before publication.
Cursor Grok 4.5 High
A genuinely different and useful decision surface, especially for task presets and practical constraints. Its self-featured Grok result is candid about missing cells but still needs independent claim-level verification.
Mixed
passux

Preset-first recommendation flow

Six one-tap presets reshape filters and sorting for frontier, coding, science, value, open-weight, and speed decisions.

Evidence: PRESETS metadata, filter reducer, FilterBar, and ExplorerProvider.

passdata

Missing numbers stay missing

The featured Cursor Grok 4.5 record publishes AA Index 54 and price/context facts while leaving five benchmark families and speed null.

Evidence: cursor-grok-4-5 entry in src/data/models.ts.

watchfactuality

The builder evaluates itself

The source calls Cursor Grok 4.5 a jointly trained SpaceXAI/Cursor frontier model and ranks it #4 from AA Index 54 without an attached claim-level source.

Evidence: README, featured model record, and methodology source list.

watchsourcing

Benchmark-family links only

Methodology links Arena, Artificial Analysis, SWE-bench, and GPQA, but model records do not identify which source supports each exact number.

Evidence: SOURCES metadata and model records.

How Cursor Grok 4.5 High did it

  1. 1.Data method: Grok curated a July 15, 2026 snapshot of 19 models across 11 labs with nullable Arena, AA Index, SWE-bench, GPQA, MMLU-Pro, and HLE fields plus price, speed, context, access, modalities, and use cases.
  2. 2.Recommendation method: six presets set the relevant use-case constraint and sort key for frontier, coding, science, value, open weights, or speed; no hidden composite determines the result.
  3. 3.Filtering method: ExplorerProvider applies search, lab, use-case, access, open-weight, minimum Arena, maximum price, sort key, and direction to the shared model array.
  4. 4.Comparison method: users select up to four models; the tray serializes their IDs into the compare route, which displays score bars and practical specifications side by side.
  5. 5.Featured result: Cursor Grok 4.5 is published at AA Index 54 with rank hint #4, $2 input / $6 output per million tokens, 500K context, and null values for Arena, SWE-bench, GPQA, MMLU-Pro, HLE, and speed.
  6. 6.Source method: the methodology page links LMArena, Artificial Analysis, SWE-bench, and the GPQA repository and warns that vendor runs differ, Arena moves daily, and critical decisions need primary-source checks.
  7. 7.July 15 verification: Cursor and SpaceXAI confirm the joint-training and $2/$6 launch facts; Artificial Analysis confirms score 54 and 500K context. Its live rank has moved since the submitted #4 snapshot, so Abundance publishes #4 only as the artifact's dated result.
  8. 8.Submitted verification: the local source contains a completed Next.js production build and route artifacts; Abundance reruns lint, type, tests, build, browser QA, gateway QA, and the production audit before publication.
Luna Codex GPT-5
Luna is the more rigorous of the two additions: its benchmark caveats, blank-value handling, mobile table alternative, and functioning comparison modal make the artifact useful without pretending the provider results are normalized.
Strong
passsourcing

Method differences stay visible

The methodology disclosure states that lab-reported values are not normalized and that benchmark families and tool settings cannot be treated as interchangeable.

Evidence: Read methodology panel and comparison-modal footnote in the submitted page.

passux

Complete short-list flow

Visitors can narrow seven models, select up to three, and open a side-by-side modal covering lab, GPQA, SWE-Pro, context, weights, and best fit.

Evidence: Filter state, toggleCompare cap, compare dock, and compare modal in app/page.tsx.

watchdata

Missing values sort as zero

Null results remain visually blank but are coerced to zero for score sorting, which is a practical ordering rule rather than a benchmark result.

Evidence: compareValue returns zero for null before the GPQA and SWE-Pro sort comparators run.

watchsourcing

One source per model

Provider pages are attached at model level, not to each benchmark cell, so exact values still require checking against their original evaluation conditions.

Evidence: Each typed model record contains one source and sourceLabel alongside up to five metric fields.

How Codex GPT-5 did it

  1. 1.Data method: Luna defines seven typed model records across seven labs with release, context, weights, modalities, fit tags, one headline signal, and nullable SWE-Pro, Terminal 2.0, GPQA, BrowseComp, and ARC-AGI 2 fields.
  2. 2.Filtering method: a memoized pipeline applies case-insensitive model/lab/family/fit search, exact lab selection, and capability tags for coding, agents, reasoning, multimodal, or open weights.
  3. 3.Sorting method: users sort by GPQA/signal, SWE-Pro, context window, or release date; null scores are treated as zero for ordering while still displayed as missing.
  4. 4.Comparison method: GPT-5.5 and Gemini 3.1 Pro start selected; users may keep up to three IDs and open a modal comparing published GPQA, SWE-Pro, context, weights, and best-fit notes.
  5. 5.Methodology method: the expandable note says values are lab-reported and not normalized, SWE-Pro is not SWE-bench Verified, tool settings matter, and missing values stay blank on purpose.
  6. 6.Source method: each record links one official provider release or model page; the footer dates the available-source snapshot to July 15, 2026.
  7. 7.Submitted verification: the completion report records a successful production build and responsive browser QA; no elapsed build time was supplied.
Terra Codex GPT-5
Terra is visually polished and fast to scan, but its main comparison and detail affordances stop short of a result. The normalized numbers should be treated as editorial fit signals until a formula and claim-level sources are published.
Mixed
passux

Clear compact workspace

Seven rows combine provider identity, status, context, four meters, price, and comparison selection without turning the page into a card mosaic.

Evidence: Model list, at-a-glance rail, and responsive row CSS in the submitted source.

passmobile

Purpose-built responsive layout

The navigation, hero, filters, model meters, insights, briefing, and fixed selection tray all reflow for narrow screens.

Evidence: The source includes a dedicated max-width 760px layout and passed the submitted production build.

watchimplementation

Comparison stops at selection

Up to three models can be queued and removed, but clicking Compare does not reveal a table, drawer, modal, or new route.

Evidence: compare-cta is rendered without an onClick handler or link target.

watchfactuality

Normalization is undocumented

Reasoning, coding, vision, and speed are stored as exact 0-100 numbers, while the page only says they are normalized across available evaluations.

Evidence: Typed model array, scoreLabels metadata, and list-heading disclosure; no calculation or source fields exist in the source app.

How Codex GPT-5 did it

  1. 1.Data method: Terra defines seven model records across seven labs with release, status, context, price text, open-weight state, tags, and normalized 0-100 reasoning, coding, vision, and speed values.
  2. 2.Filtering method: a memoized pipeline applies case-insensitive model/lab/tag search plus capability chips for reasoning, coding, multimodal, agents, or open weights.
  3. 3.Sorting method: the active reasoning, coding, vision, or speed numeric field is sorted descending; there is no hidden aggregate formula in the executable source.
  4. 4.Recommendation method: fixed at-a-glance callouts name Gemini 3.1 Pro as leading overall, Claude Opus 4.7 as best for code, and Mistral Large 3 as best value.
  5. 5.Comparison method: Gemini 3.1 Pro and GPT-5.5 start selected and selection is capped at three; the tray supports removal, but its Compare CTA has no handler and produces no side-by-side result.
  6. 6.Source method: the submitted report references official Google DeepMind, Anthropic, OpenAI, and Meta pages, while the source records no per-model URL and no normalization procedure.
  7. 7.Submitted verification: the completion report records a passing TypeScript typecheck and production build; no elapsed build time was supplied.
Grok Build
Best for interface completeness, but the data layer is openly approximate and carries real hallucination risk.
Mixed
passux

Strong working dashboard

The submitted output gives users cards, table view, filters, quick presets, compare, details, cost estimates, and CSV export.

Evidence: Archive/Grok/app/page.tsx implements the full ModelBench single-page dashboard.

passmobile

Original dark presentation restored

The preview now uses the submitted black background, white text, zinc panels, and compact mobile-first controls.

Evidence: Archive/Grok/app/layout.tsx and globals.css set bg-zinc-950/text-zinc-200 styling.

watchfactuality

Dataset admits approximation

The source comment says the data is real-ish and synthesized from leaderboards, Artificial Analysis, and provider announcements.

Evidence: Archive/Grok/app/page.tsx dataset comment above ALL_MODELS.

failsourcing

No claim-level citations

The app names current-looking models and exact benchmark scores, but does not attach source links or evidence to individual claims.

Evidence: Footer contains a general verification warning instead of per-model sources.

How Grok Build did it

  1. 1.Source files inspected under Downloads/Archive/Grok.
  2. 2.Original app uses framer-motion and sonner; this scoped preview preserves behavior with plain React/CSS and an aria-live export status.
  3. 3.Work time reported: 5m 27s initial build plus 41s to get it running.
Composer 2.5
A clean, usable explorer with strong dark UI fidelity, but its benchmark claims are still mostly asserted rather than evidenced.
Mixed
passimplementation

Clean component split

The source separates data, filtering, cards, details, stats, lab badges, and compare UI into readable modules.

Evidence: Archive/Composer/Grok Composer/src/components, data, and lib folders.

passux

Useful explorer behavior

Search, sorting, filters, model detail, and compare all work as expected for a model-selection dashboard.

Evidence: ModelExplorer wires FilterPanel, ModelCard, ModelDetail, ComparePanel, and StatsOverview.

watchsourcing

Limited visible evidence

The app includes benchmark-heavy model data but does not visibly attach claim-level citations to the numbers.

Evidence: The preview footer has a general sourcing note; the cards and detail views do not expose per-claim sources.

watchdata

No recommendation wizard

It is a strong explorer, but it does not guide a non-technical user through a recommendation flow.

Evidence: No wizard component or explicit step-by-step recommendation route exists in the Composer source.

How Composer 2.5 did it

  1. 1.Source files inspected under Downloads/Archive/Composer/Grok Composer/src.
  2. 2.Original source includes Tailwind 3 and Next 15 dependencies, so the preview is scoped into this site's Next 16 app.
  3. 3.Work time reported: 3m.
GPT 5.5 High
The most evidence-aware submission: it still needs external verification, but it did the best job separating useful UI from caveats and source notes.
Strong
passsourcing

Best provenance posture

The source data has source records, caution fields, and partial-data handling instead of presenting every number as equally certain.

Evidence: ai-model-benchmarks-source.zip data module includes sources and caution metadata.

passux

Strong model-selection workflow

The app supports search, filters, task-weighted sorting, compare selections, benchmark drilldowns, and method notes.

Evidence: Submitted completion report and preserved model-explorer component.

watchfactuality

Still uses future/current claims

The app is more careful than the others, but the underlying GPT-5.5-era model names and figures still need checking.

Evidence: Dataset includes current-looking frontier model names and benchmark fields.

How GPT 5.5 High did it

  1. 1.Source module includes sources[] and models[] with benchmark/caution fields.
  2. 2.This preview preserves the source app's partial-data posture rather than filling gaps.
  3. 3.Work time reported: 11m 40s.
GLM 5.2
The broadest app structure, with real routes and strong comparison surfaces; the main weakness is that the confident benchmark data still needs verification.
Strong
passimplementation

Best route surface

The app has leaderboard, compare, about/methodology, and per-model detail pages rather than only a single dashboard.

Evidence: ai-bench.zip source includes /, /about, /compare, and /models/[id] routes.

passux

Dense but usable comparison

The compare tray, provider marks, radar chart, sortable leaderboard, and detail pages make it feel like a real tool.

Evidence: Leaderboard, CompareTray, RadarChart, BenchmarkBars, and ProviderMark components.

watchfactuality

Assertive model claims

The model lineup and benchmark numbers are presented confidently and should be checked before being treated as facts.

Evidence: Source data contains current-looking frontier names and exact benchmark scores.

watchdata

No explicit recommendation wizard

The app is strong for users who know how to compare metrics, but less guided for users who want a step-by-step model recommendation.

Evidence: No recommendation wizard route or component was present in the inspected source.

How GLM 5.2 did it

  1. 1.Routes found: /, /about, /compare, /models/[id].
  2. 2.Components found: Leaderboard, CompareTray, CompareButton, ProviderMark, BenchmarkBars, RadarChart.
  3. 3.Work time reported: 23m 30s.
Gemini 3.5 Flash High
Fast and approachable, but it fails the latest-model requirement: the submitted dataset looks stale next to the other entries.
Needs Review
passux

Approachable picker

The card UI, filters, compare drawer, modal detail, and recommendation wizard make the app easy to try quickly.

Evidence: project.zip includes FilterPanel, ModelCard, CompareDrawer, RecommendationWizard, and RadarChart components.

passmobile

Quick mobile-friendly interaction

The compare drawer and wizard are well-suited to a small-screen model picker.

Evidence: The preview uses a sticky comparison drawer and modal-style recommendation flow.

faildata

Misses the latest-model brief

The prompt asked for the latest AI models, but this dataset centers older families such as GPT-4o, o1, Claude 3.5, Gemini 1.5/2.0, Llama 3.x, and Mistral Large 2.

Evidence: Gemini data lacks the GPT-5.5, Claude 4.x/5-style, Gemini 3.x, GLM-5.2, and Grok 4.x-style coverage seen elsewhere.

failfactuality

Stale claims create hallucination risk

Because older models are framed as current leaders, the app risks giving users outdated or hallucinated recommendations.

Evidence: Submitted model notes present older model families as the useful current benchmark set.

watchsourcing

Weak benchmark evidence

The app gives benchmark values, but does not provide enough visible source trail for users to trust the numbers.

Evidence: The submitted preview exposes benchmark fields without claim-level citations.

How Gemini 3.5 Flash High did it

  1. 1.Components found: CompareDrawer, RecommendationWizard, FilterPanel, ModelCard, Header, RadarChart.
  2. 2.Original data uses Chatbot Arena ELO, MMLU, GPQA, MATH, and HumanEval.
  3. 3.Work time reported: 2m 23s across the plan and build phases.
Gemini 3.6 Flash High
Gemini 3.6 Flash High delivered the most complete interaction surface of the Gemini submissions and now builds on Next.js 16.2.11, but its data remains an older static benchmark snapshot that needs sourcing and modernization.
Needs Review
passimplementation

Verified current-framework build

The submitted package pins Next.js and eslint-config-next to 16.2.11, and the supplied verification report records a clean production build with no compilation or type errors.

Evidence: Submitted package.json and July 22, 2026 build verification report.

passux

Complete model-selection workflow

Visitors can filter, sort, switch between grid and table layouts, inspect details, compare up to four models, use the recommendation wizard, and explore a price-versus-performance chart.

Evidence: Submitted Home, FilterBar, ModelGrid, ModelTable, ComparisonDrawer, ModelDetailModal, UseCaseWizard, and PerformanceChart components.

watchdata

Snapshot is not current frontier coverage

Its 18 records foreground DeepSeek R1, o3-mini, Claude 3.5, Gemini 2.0, GPT-4o, Llama 3.x, and other 2024/early-2025 entries rather than the latest model families represented elsewhere in the benchmark collection.

Evidence: src/data/modelsData.ts in the submitted project.

watchsourcing

Exact benchmark claims need provenance

The typed records expose precise ELO, MMLU-Pro, GPQA, HumanEval, MATH-500, SWE-bench, speed, and pricing values, but users cannot trace individual claims to primary sources or a shared evaluation harness.

Evidence: The submitted model data contains values and descriptions but no per-claim source links or methodology module.

How Gemini 3.6 Flash High did it

  1. 1.Build method: Next.js 16.2.11, React 19.2.4, TypeScript, Tailwind CSS v4, lucide-react, and Recharts.
  2. 2.Data method: 18 typed models across nine labs include prices, context, speed estimates, modality, reasoning status, benchmark fields, strengths, and workload guidance.
  3. 3.Interaction method: grid/table filtering and sorting feed a capped four-model comparison drawer; separate analytics and wizard views reuse the same static model data.
  4. 4.Submitted verification: npm run build completed on Next.js 16.2.11 with no compilation or type errors. Abundance reruns release checks before publication.
Factuality signals

Stale data and hallucination risk.

Strong means the artifact handled that area well. Weak means users should treat the submitted claims with extra skepticism before using the app for model selection.

BuildLatest coverageSourcingHallucination riskBenchmark verificationInteraction quality
GPT 5.6 Sol High

Strong

mixed

Final snapshot covers nine selected models across seven labs rather than claiming exhaustive market coverage.

strong

Every model has a first-party HTTPS source and a human-readable source label.

strong

Missing prices and scores stay explicit; self-hosted and not-listed states are not converted into guesses.

mixed

Provider claims are source-linked and caveated, but no independent rerun normalizes the different harnesses.

strong

Search, filters, use-case ranking, sorting, detail, scatterplot, and capped comparison all work in one surface.

GLM 5.2 Cursor

Strong

strong

Twenty models across nine labs, spanning proprietary, open-weight, fast, and frontier tiers.

mixed

Source families and harness caveats are named, but individual model-score records lack claim-level links.

mixed

Missing data is tolerated, yet many precise mid-2026 values are asserted rather than independently reproduced.

mixed

The app explains harness variance and saturation; Abundance did not rerun the 10 benchmark suites.

strong

Filters, metric sorting, persistent four-model comparison, radar visualization, detail routes, and mobile states are all implemented.

Cursor Grok 4.5 High

Mixed

strong

Nineteen models across 11 labs, including a separate SpaceXAI identity for Cursor Grok 4.5.

mixed

Four authoritative benchmark-family links are visible, but exact model claims are not mapped to those sources.

mixed

Null handling is good; the self-featured joint-training and rank claims still need independent verification.

mixed

AA Index, Arena, SWE-bench, and GPQA are explained, but Abundance did not reproduce the submitted numbers.

strong

Task presets, deep filters, nine sorts, model details, and capped comparison make the snapshot useful.

Luna Codex GPT-5

Strong

strong

Seven-provider July 15, 2026 snapshot including GPT-5.5, Opus 4.8, Gemini 3.1 Pro, and Grok 4.5.

mixed

Every model links an official provider page, but exact benchmark cells are not individually cited.

strong

Null fields remain null and the app repeatedly warns against apples-to-apples interpretation.

mixed

The source explains evaluation mismatch; Abundance publishes the artifact rather than reproducing the lab runs.

strong

Responsive filtering, sorting, mobile cards, methodology reveal, and a working comparison modal.

Terra Codex GPT-5

Mixed

mixed

Seven-model snapshot includes 2026 flagships but also older 2025 reference models.

weak

The completion report names official lab pages, but the executable source contains no claim-level URLs or score provenance.

mixed

The benchmark warning helps, but unexplained normalized scores can look more objective than the evidence supports.

weak

No normalization formula or reproducible evaluation harness is included.

mixed

Search, chips, sorting, and selection work; compare, detail, and secondary mobile filter buttons are incomplete.

Grok Build

Mixed

mixed

Covers current-looking 2026 model names, but several claims need verification.

weak

General source note only; no claim-level citations.

weak

The submitted source calls the data real-ish and synthesized.

weak

Exact benchmark scores are not backed by visible source records.

strong

Filtering, sorting, compare, details, export, and cost estimates all exist.

Composer 2.5

Mixed

strong

Includes many 2026-style model names across major labs.

weak

Generic sourcing note, not claim-level citations.

mixed

Current-looking names and scores require verification.

weak

Benchmark fields are asserted in data files without evidence records.

strong

Search, filter, sort, details, compare, and mobile controls are present.

GPT 5.5 High

Strong

strong

Broad model set with GPT, Claude, Gemini, Grok, GLM, DeepSeek, Qwen, Meta, Mistral, and Cohere.

strong

Best source and caution scaffolding among the submissions.

mixed

Lower than others, but current-looking model claims still need verification.

mixed

Method notes help; independent verification is still out of scope here.

strong

Search, filter, sort, compare, drilldowns, and recommendation-style weighting are present.

GLM 5.2

Strong

strong

Broad current-looking frontier model coverage.

mixed

Includes methodology/about surface, but claim-level evidence is limited.

mixed

Confident exact numbers and model names need verification.

mixed

Methodology page helps but does not make this a certified benchmark feed.

strong

Best multi-route comparison experience.

Gemini 3.5 Flash High

Needs Review

weak

Fails the latest-model requirement; dataset appears stale.

weak

No strong source trail for exact benchmark claims.

weak

Older models are framed as current recommendations.

weak

Benchmark claims need verification and likely updating.

strong

Wizard, filters, detail modal, and compare drawer are useful.

Gemini 3.6 Flash High

Needs Review

weak

A polished but static 2024/early-2025 model snapshot rather than a current frontier roster.

weak

Exact values have no visible claim-level citations.

mixed

The app is explicit about its fields, but unsourced precision and stale positioning need verification before model-selection use.

weak

No visible cross-provider normalization or independent benchmark rerun is supplied.

strong

Filtering, sorting, comparison, recommendations, analytics, details, and mobile navigation are all implemented.

Feature matrix

What each build made viewable.

This is a product-surface comparison only. It checks whether each submitted app exposed the interaction pattern in its source or build output.

BuildFiltersSortingCompareModel detailRecommendationLab brandingMobile-first UI
GPT 5.6 Sol High

components/model-explorer.tsx and data/models.ts

GLM 5.2 Cursor

E:/Projects/Benchmark Tests/GLM/glm-5.2-coder

Cursor Grok 4.5 High

E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5

Luna Codex GPT-5

C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5

Terra Codex GPT-5

C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas

Grok Build

Archive/Grok

Composer 2.5

Archive/Composer/Grok Composer

GPT 5.5 High

src/components/model-explorer.tsx and src/data/models.ts

GLM 5.2

ai-bench/

Gemini 3.5 Flash High

project root

Gemini 3.6 Flash High

E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash

Strongest patterns

The useful apps behaved like tools.

  • They let visitors narrow the model set by lab, job, price, or capability.
  • They made comparison explicit instead of burying scores in isolated model cards.
  • They included caution language or partial-data handling when benchmark fields were missing.
Important caveat

Do not treat these as verified benchmark facts.

These previews are a way to inspect the work produced by each builder. The report card calls out obvious artifact issues, but it is not a live benchmark feed or a certified source for model-selection advice.