Skip to content
Abundance Marketing Partners
Case StudiesPricingHow We BuildLabsAbout
Request a Project Review
Abundance Marketing Partners

Utah-based custom software, AI workflows, internal tools, and modern websites for owners across the United States.

Request a Project Review

Services

  • Custom Apps
  • AI Workflows & Agents
  • Internal Tools & MCP
  • Modern Websites
  • Delivery Infrastructure

Company

  • About
  • Case Studies
  • Abundance Labs
  • Pricing
  • Service Areas
  • Articles
  • How We Build
  • AI Benchmark Results

Contact

  • Request a Project Review
  • hello@abundancepartners.app

What happens after you get in touch?
Read how we work

Prefer proof first?
Explore our client work · Browse Labs

Local service areas

Salt Lake CitySandyDraperSouth JordanLehiProvo
See all service areas

© 2026 Abundance Marketing Partners LLC. All rights reserved.

Privacy PolicyTerms of Service
AI Benchmark Results

20 benchmark builds, one place to inspect them.

This page compares the submitted Next.js benchmark apps as build outputs. The previews preserve the model builders' own interfaces, while the review flags stale data, weak sourcing, and likely hallucination risk without certifying the benchmark claims.

View submissions Read reviewModel explorer

Review mode

Blunt artifact report card

Findings are based on the submitted artifacts, visible app behavior, and peer comparison. This is not a full source-backed audit of every benchmark number.

20

Submissions

20

Live source previews

6/20

Build times

Top model explorer

Compare current models with real charts.

The benchmark explorer leads with current top models, cost/quality scatter plots, coding vs agentic charts, coverage heatmaps, and a sticky compare tray.

Open explorer
Preview routes

Open each submitted build.

Each link opens the submitted app's own interface, with Abundance chrome removed except for a small back button.

Preview
Strong29m
Astra
Astra via Astra

The gallery's best-sourced single page: per-benchmark source links, explicitly versioned tracks, six shipped unit tests, and a dated Sep 4, 2026 snapshot, held back only by reliance on publisher comparison tables (including competitors' cards) and the recurring SpaceXAI relabel.

pass

3

watch

2

fail

0

Source

Benchmarks/Astra

Features

7/7

Open original output
Preview
StrongPending
Fable
Fable via Fable

The gallery's most complete build: five working routes, job-based recommendations, URL-shareable filter state, and a dated Aug 28, 2026 snapshot across 31 models, with aggregate sourcing and an unpublished cost-per-task harness as the residual caveats.

pass

3

watch

2

fail

0

Source

Benchmarks/Fable/fable

Features

7/7

Open original output
Preview
StrongPending
Muse Spark 1.3
Muse Spark 1.3 via Muse

The most transparent single-page board in the gallery: current launch-week coverage, an explicit Sep 2, 2026 snapshot, and honest per-record caveats, held back only by vendor-only figures and a blended score that quietly rewards thin rows.

pass

2

watch

3

fail

0

Source

Benchmarks/Musespark 1.3

Features

7/7

Open original output
Preview
Needs ReviewPending
Ling 3.0 Flash
Ling 3.0 Flash via Ling

A tidy three-route dashboard with solid compare and detail flows, failing the freshness bar on an 18-model dataset that stops at early-2025 releases and cites no sources.

pass

2

watch

0

fail

2

Source

Benchmarks/Ling 3.0 Flash/ai-models-benchmarks

Features

6/7

Open original output
Preview
Needs ReviewPending
Muse Spark 1.2
Muse Spark 1.2 via Muse

A clean, fast leaderboard with unusually good operational columns, undermined by a mid-2025 dataset presented under a 'Live' title and by sourcing that is never stated.

pass

2

watch

0

fail

2

Source

Benchmarks/Muse Spark 1.2/bench

Features

6/7

Open original output
Preview
StrongPending
Grok 4.6 Grok Build
Grok 4.6 via Grok Build

The strongest multi-route submission in the gallery: Aperture pairs a 34-model, 12-lab August 2026 snapshot with honest access labeling, skip-if guidance, a Pareto comparison, and a stated missing-data policy.

pass

3

watch

1

fail

0

Source

Benchmarks/Grok 4.6 Grok Build

Features

7/7

Open original output
Preview
Needs ReviewPending
Gemini 3.7 Flash High
Gemini 3.7 Flash via Antigravity

The most feature-complete single-page submission, but it fails the freshness bar: a Feb-2025-era dataset is framed as today's leaderboard, so the useful interactions sit on top of stale evidence.

pass

2

watch

1

fail

2

Source

Benchmarks/Gemini Flash 3.7

Features

7/7

Open original output
Preview
StrongPending
Zcode 5.3 Flash
Zcode 5.3 Flash via Zcode

The most editorially disciplined submission in the gallery: ModelPulse pairs 38 curated records with explicit null handling, rank gating, primary-source links, and shareable URL state, at the cost of many empty cells for the newest releases.

pass

3

watch

2

fail

0

Source

Benchmarks/Zcode 5.3 Flash/modelpulse

Features

6/7

Open original output
Preview
StrongPending
Zcode 5.3
Zcode 5.3 via Zcode

A dense, polished comparison dashboard: Zcode 5.3 shipped 32 typed records across 13 labs with dual views, four-way compare, and a working QA script—the main caveat is aggregate acknowledgment of estimated values instead of per-cell flags.

pass

2

watch

2

fail

0

Source

Benchmarks/Zcode 5.3/modeldex

Features

6/7

Open original output
Preview
Needs ReviewPending
Gemini 3.6 Flash High
Gemini 3.6 Flash High via Antigravity

Gemini 3.6 Flash High delivered the most complete interaction surface of the Gemini submissions and now builds on Next.js 16.2.11, but its data remains an older static benchmark snapshot that needs sourcing and modernization.

pass

2

watch

2

fail

0

Source

E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash

Features

7/7

Open original output
Preview
Needs Review2m 23s
Gemini 3.5 Flash High
Gemini 3.5 Flash High via antigravity

Fast and approachable, but it fails the latest-model requirement: the submitted dataset looks stale next to the other entries.

pass

2

watch

1

fail

2

Source

project root

Features

7/7

Open original output
Preview
Strong23m 30s
GLM 5.2
GLM 5.2 via Zcode

The broadest app structure, with real routes and strong comparison surfaces; the main weakness is that the confident benchmark data still needs verification.

pass

2

watch

2

fail

0

Source

ai-bench/

Features

6/7

Open original output
Preview
Strong11m 40s
GPT 5.5 High
GPT 5.5 High via Codex

The most evidence-aware submission: it still needs external verification, but it did the best job separating useful UI from caveats and source notes.

pass

2

watch

1

fail

0

Source

src/components/model-explorer.tsx and src/data/models.ts

Features

7/7

Open original output
Preview
Mixed3m
Composer 2.5
Composer 2.5 via grok build

A clean, usable explorer with strong dark UI fidelity, but its benchmark claims are still mostly asserted rather than evidenced.

pass

2

watch

2

fail

0

Source

Archive/Composer/Grok Composer

Features

6/7

Open original output
Preview
Mixed6m 8s
Grok Build
Grok Build via grok build

Best for interface completeness, but the data layer is openly approximate and carries real hallucination risk.

pass

2

watch

1

fail

1

Source

Archive/Grok

Features

6/7

Open original output
Preview
MixedPending
Terra Codex GPT-5
Codex GPT-5 via Codex

Terra is visually polished and fast to scan, but its main comparison and detail affordances stop short of a result. The normalized numbers should be treated as editorial fit signals until a formula and claim-level sources are published.

pass

2

watch

2

fail

0

Source

C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas

Features

6/7

Open original output
Preview
StrongPending
Luna Codex GPT-5
Codex GPT-5 via Codex

Luna is the more rigorous of the two additions: its benchmark caveats, blank-value handling, mobile table alternative, and functioning comparison modal make the artifact useful without pretending the provider results are normalized.

pass

2

watch

2

fail

0

Source

C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5

Features

6/7

Open original output
Preview
MixedPending
Cursor Grok 4.5 High
Cursor Grok 4.5 High via Cursor

A genuinely different and useful decision surface, especially for task presets and practical constraints. Its self-featured Grok result is candid about missing cells but still needs independent claim-level verification.

pass

2

watch

2

fail

0

Source

E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5

Features

7/7

Open original output
Preview
StrongPending
GLM 5.2 Cursor
GLM 5.2 via Cursor

The most structurally ambitious result in this round. GLM built a real multi-route benchmark product with excellent comparison ergonomics, while its exact scores still require source-by-source verification.

pass

3

watch

1

fail

0

Source

E:/Projects/Benchmark Tests/GLM/glm-5.2-coder

Features

6/7

Open original output
Preview
StrongPending
GPT 5.6 Sol High
GPT 5.6 Sol High via Codex

The strongest complete newcomer: Sol shipped a restrained, mobile-first decision tool with unusually clear benchmark caveats, useful interactions, and executable source tests.

pass

3

watch

1

fail

0

Source

components/model-explorer.tsx and data/models.ts

Features

7/7

Open original output
Blunt review

What worked, what failed.

These findings review the submitted apps as artifacts. They flag obvious stale data, weak sourcing, and likely hallucination risk, but they are not a complete fact-check of every AI benchmark claim.

Astra
The gallery's best-sourced single page: per-benchmark source links, explicitly versioned tracks, six shipped unit tests, and a dated Sep 4, 2026 snapshot, held back only by reliance on publisher comparison tables (including competitors' cards) and the recurring SpaceXAI relabel.
Strong
passux

Complete catalog surface with accessible dialogs

Search, lab multi-select, access and use-case filters, benchmark and sort selects, table and chart views, mobile cards, a detail dialog, a methodology dialog, a three-model compare, CSV export, and a saved shortlist all respond client-side; dialogs are native <dialog> elements with Escape dismissal and focus containment.

Evidence: The explorer state pipeline in components/explorer.tsx and the native-dialog Modal in components/ui.tsx.

passsourcing

Every reported figure links to a source

Score cells open the publisher table they came from (with scoreSources overrides such as GPT-5.6 Sol's Terminal 2.1), and model dialogs add documentation and pricing links; the CSV export carries benchmark, terminal-override, details, and price source columns.

Evidence: The SourceLink usage in components/model-dialog.tsx and the toCSV headers in lib/explorer.ts.

passimplementation

Pure logic shipped with its own test suite

Filtering, null-last two-way sorting, the three-model cap, and CSV serialization live in a dependency-free lib/explorer.ts covered by six node:test cases; the data module enforces bounded scores and https sources, and the QA report records axe 4.12.1 with zero violations.

Evidence: tests/explorer.test.ts and docs/qa-report.md in the Astra source.

watchdata

Competitor cards as primary sources

Qwen3.8-Max's figures come from Z.ai's GLM-5.3 model card and several Claude and Gemini rows cite OpenAI's GPT-6 comparison table, so cross-lab numbers pass through a rival's curation; the app discloses this in each note, but the source link a user opens is the competitor's page, not the subject lab's.

Evidence: The sourceLabel and source fields on the qwen-3-8-max, claude-opus-5, and gemini-3-8-flash records in lib/models.ts.

watchfactuality

SpaceXAI relabel and preview-generation scores

xAI is relabeled 'SpaceXAI' in the lab list and release notes, repeating a naming quirk other gallery submissions were dinged for, and Gemini 3.1 Pro keeps a preview-table DeepSWE of 11.8 in the main ranking behind a status tag that is easy to miss.

Evidence: The labs entry for id 'xai' and the gemini-3-1-pro record in lib/models.ts.

How Astra did it

  1. 1.Data method: 18 typed records across 10 labs with four versioned benchmark tracks, per-benchmark scoreSources overrides, documentation and pricing sources, license notes, release dates, and evaluation conditions in each notes field; unverified values are absent, never zeroed.
  2. 2.Interaction method: a dependency-free lib/explorer.ts handles conjunctive filters, null-last two-way sorting, the three-model selection cap, and CSV serialization, covered by six node:test cases shipped with the source.
  3. 3.Dataset era: SNAPSHOT is 2026-09-04, carried consistently in the metadata, the coverage line, the model dialogs, and the CSV's verified-date column.
  4. 4.Verification: the source ships its own QA report (typecheck, lint, build, six passing tests, axe 4.12.1 zero violations, 320-1440px responsive checks); Abundance reran the test suite plus its own lint, typecheck, and build before publishing.
Fable
The gallery's most complete build: five working routes, job-based recommendations, URL-shareable filter state, and a dated Aug 28, 2026 snapshot across 31 models, with aggregate sourcing and an unpublished cost-per-task harness as the residual caveats.
Strong
passux

Job-based re-ranking with visible fit

Selecting any of nine job presets re-computes a weighted blend of benchmark, cost, and availability factors, re-ranks the catalog, and labels each card Best match, Strong fit, Usable, or Weak fit.

Evidence: jobScore and fitFromScore in lib/catalog.ts driving JobPicker and the card badges.

passimplementation

Five routes with shareable state and a tested core

Catalog, models, labs, benchmarks, and compare routes all render statically; filters serialize to the URL, compare persists to localStorage, and an eight-test suite covers ranking, filtering, search, and delayed-availability behavior.

Evidence: lib/url-query.ts, the compare store, and lib/catalog.test.ts.

passdata

Dated snapshot with honest gaps

A single DATA_AS_OF date (2026-08-28) appears on the catalog, methodology, and footer; releases run through 2026-08-18, and the one unavailable model is listed as delayed with an explicit warning rather than dropped.

Evidence: DATA_AS_OF in lib/benchmarks.ts and the Gemini 3.5 Pro record in lib/models.ts.

watchsourcing

Aggregate attribution only, and unused lab links

The footer attributes scores to Artificial Analysis, arena standings, and first-party price sheets, but no model or benchmark links to its source, and the site URL collected for every one of the 14 labs is never rendered anywhere in the app.

Evidence: The footer copy and the unused site field in lib/labs.ts.

watchfactuality

Cost per task has no published harness

Cost per task drives the budget-weighted job scores, but the app states only that it is measured spend on a standard task; the harness, task set, and date are unpublished, so the numbers cannot be reproduced.

Evidence: The benchmark description in lib/benchmarks.ts and the pricing.note fields in lib/models.ts.

How Fable did it

  1. 1.Data method: 31 typed records across 14 labs with release dates, context and output tokens, license objects, availability states, per-million pricing plus a measured cost per task, and seven score fields with full benchmark metadata.
  2. 2.Interaction method: the catalog is a server-rendered rank over URL-parsed query state; the job picker, filters, and search all serialize through lib/url-query.ts, and compare selection persists to localStorage via a store with an SSR fallback.
  3. 3.Route method: five routes (catalog, models/[id], labs, benchmarks, compare) with per-model pages statically generated from the typed module and unknown ids falling through to a Missing page.
  4. 4.Dataset era: a single snapshot constant, DATA_AS_OF = 2026-08-28, is rendered on the catalog header, methodology page, and footer; release dates run through 2026-08-18 and month-precision dates appear on three Qwen records.
  5. 5.Verification: the source ships an eight-test catalog suite covering ranking, filtering, search, availability, and context sorting; Abundance reran lint, typecheck, build, and preview checks before publishing.
Muse Spark 1.3
The most transparent single-page board in the gallery: current launch-week coverage, an explicit Sep 2, 2026 snapshot, and honest per-record caveats, held back only by vendor-only figures and a blended score that quietly rewards thin rows.
Strong
passux

Complete interaction surface on one page

Search, use-case chips, nine-lab multi-select, three toggles, eleven sort modes, cards and table views, a detail modal, and a three-model compare with per-row stars all respond client-side with no dead controls.

Evidence: The single client state pipeline in app/page.tsx over data/models.ts.

passsourcing

Snapshot date and lab-vs-independent caveats stated

The data module, methodology section, and footer all carry the same compiled-on date, name the figures as vendor-reported, and quote independent re-run deltas such as Epoch AI's V4-Pro SWE result.

Evidence: The models.ts header comment, the on-page methodology copy, and the footer snapshot line.

watchdata

Blended overall score rewards thin rows

Unreported benchmark cells count as a neutral 50 of the best-in-lineup ratio, so Gemini 3.8 Flash (one reported figure) still lands around 61 overall and Astra (zero reported) scores exactly 50. The method is disclosed, but the ranking still flatters launch-day records.

Evidence: The blendedRank function in data/models.ts and the methodology paragraph that defends it.

watchfactuality

Prose-only attribution and a renamed lab

No benchmark number links to a source: Epoch AI, Terminal-Bench, and WebDev Arena are named in prose without URLs, and xAI is relabeled 'SpaceXAI' in the lab list, footer, and metadata, contradicting the lab naming used elsewhere in the gallery.

Evidence: The methodology fine print, the footer disclaimer, and the LABS entry for id 'xai'.

watchmobile

View switcher tucked into the filter panel

The cards/table toggle is hidden below the sm breakpoint except inside the expanded filters sheet, so mobile users must open filters to change views.

Evidence: The hidden sm:flex and sm:hidden classes on the view-mode controls in app/page.tsx.

How Muse Spark 1.3 did it

  1. 1.Data method: 13 typed records across 9 labs with release dates, context, pricing (null when undisclosed), open-weights and availability flags, access notes, modalities, use-case tags, and eight benchmark fields including SWE-bench Pro and AA-AnalystAgent.
  2. 2.Interaction method: one client pipeline drives search, use-case, lab, and toggle filters, eleven sort modes, view mode, the detail modal, and a three-model compare; the blended overall score is a transparent mean of best-in-lineup ratios with a neutral fill for unreported cells.
  3. 3.Dataset era: explicitly dated — 'All figures vendor/lab-reported launch numbers compiled Sep 2, 2026' — the only submission besides Aperture to state a snapshot date.
  4. 4.Verification: the source ships with the default create-next-app README and no QA script; Abundance reran lint, typecheck, build, and preview checks before publishing.
Ling 3.0 Flash
A tidy three-route dashboard with solid compare and detail flows, failing the freshness bar on an 18-model dataset that stops at early-2025 releases and cites no sources.
Needs Review
passux

Three usable routes with a working compare flow

The board filters and sorts client-side, compare opens as a modal and a dedicated page for up to four models, and each model renders a detail page with full benchmark bars and cost efficiency.

Evidence: app/page.tsx state, CompareModal, app/compare/page.tsx, and app/model/[id]/page.tsx.

passimplementation

Typed dataset with per-model detail pages

Records are typed with pricing (including cached rates), architecture, params, and modalities, and the model page renders straight from the typed module.

Evidence: lib/models.ts and the model detail route in the submitted source.

faildata

Misses the latest-model brief

The layout metadata promises 'the latest AI models', but the dataset stops at GPT-4.5, Claude 3.5, and Gemini 2.0-era releases with no 2025-late or 2026 coverage.

Evidence: The 18-record dataset in lib/models.ts versus the metadata description.

failfactuality

No sources or snapshot date

Ten benchmark families are populated per model with no citations, methodology, or compiled-on date anywhere in the app.

Evidence: Benchmark entries in lib/models.ts carry descriptions but no source fields.

How Ling 3.0 Flash did it

  1. 1.Data method: 18 typed records across 8 labs with release dates, architecture, params, context, modalities, cached-input pricing, and per-benchmark entries with category and full-mark metadata.
  2. 2.Interaction method: one client state pipeline drives search, lab/category filters, and five sort modes; compare is capped at four IDs and available as both a modal and a route.
  3. 3.Route method: the board and compare are client surfaces; the per-model page renders from the typed module and guards unknown IDs with a not-found state.
  4. 4.Verification: the source ships with the default create-next-app README and no QA script; Abundance reran lint, typecheck, build, and preview checks before publishing.
Muse Spark 1.2
A clean, fast leaderboard with unusually good operational columns, undermined by a mid-2025 dataset presented under a 'Live' title and by sourcing that is never stated.
Needs Review
passux

Fast, readable board with operational columns

Search, lab, modality, and access filters, sorting, cards and table views, a benchmark-focus selector, and a detail modal all respond client-side with no dead controls.

Evidence: The single-page client state pipeline in app/page.tsx.

passmobile

Mobile filter sheet and touch targets

Filters collapse into a sheet on small screens and the compare tray stays reachable while scrolling.

Evidence: showFilters state and the compare tray markup in app/page.tsx.

faildata

Mid-2025 dataset under a live-leaderboard title

The app titles itself 'BENCHMARK — Live AI Model Leaderboard' but the newest records are Claude 4, Gemini 2.5 Pro, and Grok 3, with no GPT-5.x, Claude 4.5+/5, Gemini 3.x, GLM-5.x, or Grok 4.x coverage.

Evidence: The 20-record dataset in data/models.ts versus the title in the submitted layout metadata.

failfactuality

No stated sources or snapshot date

The app presents benchmark figures with no methodology, source list, or compiled-on date, so visitors cannot tell how current or verifiable the numbers are.

Evidence: The page copy and data module contain no sourcing or snapshot note.

How Muse Spark 1.2 did it

  1. 1.Data method: 20 typed records across 10 labs with family, release month, params, context, modalities, access channels, pricing, arena Elo, latency, and ten benchmark fields.
  2. 2.Interaction method: one client pipeline drives search, filters, sorting, view mode, benchmark focus, detail modal, and the compare tray; no routing beyond the single page.
  3. 3.Dataset era: newest record is Gemini 2.5 Pro; the module carries no compiled-on date, so the gallery dates it from the release set.
  4. 4.Verification: the source ships with the default create-next-app README and no QA script; Abundance reran lint, typecheck, build, and preview checks before publishing.
Grok 4.6 Grok Build
The strongest multi-route submission in the gallery: Aperture pairs a 34-model, 12-lab August 2026 snapshot with honest access labeling, skip-if guidance, a Pareto comparison, and a stated missing-data policy.
Strong
passdata

Access and evidence stay explicit

Records distinguish restricted from generally available models, mark evidence quality, and state that scores are compiled from public leaderboards, not live eval runs.

Evidence: status and evidence fields in lib/models.ts plus the README and methodology copy.

passux

Complete decision workflow across six routes

Visitors can search and filter the field board, switch cards, table, and price-map views, open use-case picks, compare up to four models with a Pareto chart, and read per-model, per-lab, guide, and methodology pages.

Evidence: FieldBoard, UseCasePicks, CompareView with ParetoChart, and the route files under app/.

passimplementation

Real routes with statically generated details

Model and lab pages render statically from typed records, compare state is URL-driven, and the compare view is Suspense-wrapped for the searchParams read.

Evidence: generateStaticParams in the submitted models/[slug] and labs/[slug] pages and the Suspense boundary in app/compare/page.tsx.

watchsourcing

Compiled figures without per-claim links

The methodology names public sources (Artificial Analysis, BenchLM, LMArena, lab cards) but individual cells do not link to the specific claim they came from.

Evidence: Methodology page copy versus the unlinked numeric cells in lib/models.ts.

How Grok 4.6 did it

  1. 1.Data method: 34 typed records across 12 labs with release month, status, license, evidence, reasoning mode, context, pricing, speed, composite intelligence/coding/agentic scores, per-benchmark fields, modalities, summary, bestFor, and skipIf guidance.
  2. 2.Interaction method: the field board drives search, filters, sorting, and view modes; compare state lives in a client provider mirrored to the URL; the Pareto chart plots composite intelligence against task cost.
  3. 3.Route method: the field board, guide, and methodology are single surfaces; compare reads up to four IDs from searchParams; per-model and per-lab pages render statically.
  4. 4.Verification: the README documents the route map, sources, and missing-data policy; Abundance reran lint, typecheck, build, and preview checks before publishing.
Gemini 3.7 Flash High
The most feature-complete single-page submission, but it fails the freshness bar: a Feb-2025-era dataset is framed as today's leaderboard, so the useful interactions sit on top of stale evidence.
Needs Review
passux

Complete decision workflow

Visitors can search, filter by lab, category, access type, price ceiling, and context floor, switch between cards and a leaderboard table, sort every metric, open a detail modal, run a recommendation wizard, and compare up to four models with radar and bars.

Evidence: page.tsx state pipeline plus FilterBar, LeaderboardTable, ModelWizard, ComparisonDock, ComparisonModal, ModelDetailModal, and BenchmarkGlossary components.

passmobile

Touch-sized and responsive

Seventeen media queries across the component styles adapt the grid, table, dock, and modals, and buttons, selects, and inputs keep a 40px minimum touch height.

Evidence: Media queries in the CSS modules and the touch-target rules in the app styles.

faildata

Misses the latest-model brief

The app presents a late-2024/early-2025 snapshot as the current frontier: Claude 3.7 Sonnet wears the 'King' badge and o1 is framed as OpenAI's premier reasoning model, with no GPT-5.x, Claude 4.x/5, Gemini 3.x, GLM-5.x, or Grok 4.x coverage.

Evidence: The 16-record dataset in data/models.ts and the badges and overview copy attached to those records.

failfactuality

Stale framing creates hallucination risk

Recommendations and leader labels describe year-old releases as today's best, so a visitor following the wizard would be pointed at superseded models.

Evidence: ModelWizard outputs and HighlightsBanner copy drawn from the same stale dataset.

watchsourcing

No per-claim benchmark trail

Ten benchmark families are populated with plausible values, but records carry no source links a visitor could verify; the footer links only to lab homepages.

Evidence: BenchmarkScores fields in data/models.ts and the Footer lab links.

How Gemini 3.7 Flash did it

  1. 1.Data method: 16 typed records across 8 labs in one module, each with tagline, overview copy, category, access type, badge, pricing (including cached-input rates), context window, modalities, and ten benchmark fields.
  2. 2.Benchmark set: LMArena overall and coding Elo, SWE-bench Verified, GPQA Diamond, MMLU-Pro, MATH-500, AIME 2024, LiveCodeBench, MMMU, and Tau-bench.
  3. 3.Interaction method: one client-side state pipeline drives search, filters, sorting, the wizard, the comparison dock, and modals; the value-frontier scatterplot plots composite quality against blended price.
  4. 4.Verification: the source ships with the default create-next-app README and no QA script; Abundance reran lint, typecheck, build, and preview checks before publishing.
Zcode 5.3 Flash
The most editorially disciplined submission in the gallery: ModelPulse pairs 38 curated records with explicit null handling, rank gating, primary-source links, and shareable URL state, at the cost of many empty cells for the newest releases.
Strong
passdata

Missing data stays missing

Unpublished scores render as em dashes, ranks require at least three published benchmarks, and the module header states that numbers are never invented to fill gaps.

Evidence: data.ts header comment, the rank gating in stats.ts, and the on-page methodology list.

passux

Workload tabs and shareable state

All-round, Reasoning, Coding, and Agents tabs re-frame default sorts, and every filter, tab, sort, and selection serializes into the query string and rehydrates on load.

Evidence: view.ts serializeFilters/parseFilters plus the history.replaceState sync in explorer.tsx.

passimplementation

Real routes instead of one screen

The board is a client surface, compare reads its model IDs from searchParams, and each of the 38 models renders as a static page with sources, caveats, and same-lab siblings.

Evidence: app/page.tsx, app/compare/page.tsx, and app/models/[id]/page.tsx with generateStaticParams in the submitted source.

watchfactuality

Vendor-reported figures with noted disputes

Many headline numbers are vendor-reported on vendor-chosen harnesses; model pages flag material disagreements (for example Kimi K3 and DeepSeek V4 SWE-bench runs) but no independent evaluation exists.

Evidence: Per-model caveats arrays and the 'Read the fine print' methodology block.

watchdata

Sparse public data for the newest releases

GLM-5.3, Grok 4.6, and Claude Opus 5 ship with empty score sets, so benchmark columns show dashes and comparisons fall back to price and context takeaways.

Evidence: Empty scores objects for those records in data.ts.

How Zcode 5.3 Flash did it

  1. 1.Data method: 38 records compiled August 26, 2026 from vendor announcements, API docs, and independent leaderboards; the module header states null scores are never invented.
  2. 2.State method: every filter, tab, sort, direction, and selection serializes through history.replaceState and rehydrates from the URL on load, so any board view is shareable.
  3. 3.Rank method: overall ranks publish only for models with at least three published benchmarks; sparse sorts fall back to composite with recency tie-breaking.
  4. 4.Route method: the board is one client surface; compare reads up to four model IDs from searchParams; per-model pages render statically with primary-source links, caveats, and same-lab siblings.
  5. 5.Submitted verification: the source ships a headless QA script and desktop/mobile screenshots; Abundance reran lint, typecheck, build, and preview checks before publishing.
Zcode 5.3
A dense, polished comparison dashboard: Zcode 5.3 shipped 32 typed records across 13 labs with dual views, four-way compare, and a working QA script—the main caveat is aggregate acknowledgment of estimated values instead of per-cell flags.
Strong
passux

Complete exploration workflow

Visitors can search, filter by lab, open weights, and capability, switch between card and table views, sort every column, open a detail sheet, and compare up to four models with radar, bars, and win counts.

Evidence: Explorer.tsx holds one shared state pipeline; CompareModal and DetailModal render as bottom sheets on mobile with Escape and scroll-lock handling.

passimplementation

Executable source with QA coverage

The submission ships a headless QA script covering search, lab filtering, sorting, the compare flow, the table, and modals, plus desktop and mobile screenshots of each state.

Evidence: qa/qa.mjs and the submitted QA screenshots; Abundance reran lint, typecheck, build, and preview checks before publishing.

watchfactuality

Estimated values are only acknowledged in aggregate

The data note states that some recent models carry estimated values, but individual cells do not distinguish estimates from measured results.

Evidence: README 'About the data' section versus the unmarked numeric cells in data/models.ts.

watchsourcing

Compiled figures without per-claim links

Scores compile from official announcements and public leaderboards, but model records do not attach source links a visitor can follow.

Evidence: data/models.ts records and the README sourcing note; contrast with submissions that attach per-model URLs.

How Zcode 5.3 did it

  1. 1.Data method: 32 typed records across 13 labs live in one module with release dates, context windows, in/out pricing, open-weights flags, capability chips, six benchmark fields, and LMArena Elo.
  2. 2.Computation method: the overall index is the mean of the four core benchmarks; value score divides the index by blended 3:1 price raised to the 0.25 power so ultra-cheap weak models cannot automatically win.
  3. 3.Interaction method: cards and table views consume one filter pipeline; comparison is capped at four IDs; detail and compare render in modals with Escape-to-close and scroll lock.
  4. 4.Verification: qa/qa.mjs headless checks (filters, sorting, compare flow, table, modals) plus submitted desktop/mobile screenshots; Abundance reran lint, typecheck, build, and preview checks.
Gemini 3.6 Flash High
Gemini 3.6 Flash High delivered the most complete interaction surface of the Gemini submissions and now builds on Next.js 16.2.11, but its data remains an older static benchmark snapshot that needs sourcing and modernization.
Needs Review
passimplementation

Verified current-framework build

The submitted package pins Next.js and eslint-config-next to 16.2.11, and the supplied verification report records a clean production build with no compilation or type errors.

Evidence: Submitted package.json and July 22, 2026 build verification report.

passux

Complete model-selection workflow

Visitors can filter, sort, switch between grid and table layouts, inspect details, compare up to four models, use the recommendation wizard, and explore a price-versus-performance chart.

Evidence: Submitted Home, FilterBar, ModelGrid, ModelTable, ComparisonDrawer, ModelDetailModal, UseCaseWizard, and PerformanceChart components.

watchdata

Snapshot is not current frontier coverage

Its 18 records foreground DeepSeek R1, o3-mini, Claude 3.5, Gemini 2.0, GPT-4o, Llama 3.x, and other 2024/early-2025 entries rather than the latest model families represented elsewhere in the benchmark collection.

Evidence: src/data/modelsData.ts in the submitted project.

watchsourcing

Exact benchmark claims need provenance

The typed records expose precise ELO, MMLU-Pro, GPQA, HumanEval, MATH-500, SWE-bench, speed, and pricing values, but users cannot trace individual claims to primary sources or a shared evaluation harness.

Evidence: The submitted model data contains values and descriptions but no per-claim source links or methodology module.

How Gemini 3.6 Flash High did it

  1. 1.Build method: Next.js 16.2.11, React 19.2.4, TypeScript, Tailwind CSS v4, lucide-react, and Recharts.
  2. 2.Data method: 18 typed models across nine labs include prices, context, speed estimates, modality, reasoning status, benchmark fields, strengths, and workload guidance.
  3. 3.Interaction method: grid/table filtering and sorting feed a capped four-model comparison drawer; separate analytics and wizard views reuse the same static model data.
  4. 4.Submitted verification: npm run build completed on Next.js 16.2.11 with no compilation or type errors. Abundance reruns release checks before publication.
Gemini 3.5 Flash High
Fast and approachable, but it fails the latest-model requirement: the submitted dataset looks stale next to the other entries.
Needs Review
passux

Approachable picker

The card UI, filters, compare drawer, modal detail, and recommendation wizard make the app easy to try quickly.

Evidence: project.zip includes FilterPanel, ModelCard, CompareDrawer, RecommendationWizard, and RadarChart components.

passmobile

Quick mobile-friendly interaction

The compare drawer and wizard are well-suited to a small-screen model picker.

Evidence: The preview uses a sticky comparison drawer and modal-style recommendation flow.

faildata

Misses the latest-model brief

The prompt asked for the latest AI models, but this dataset centers older families such as GPT-4o, o1, Claude 3.5, Gemini 1.5/2.0, Llama 3.x, and Mistral Large 2.

Evidence: Gemini data lacks the GPT-5.5, Claude 4.x/5-style, Gemini 3.x, GLM-5.2, and Grok 4.x-style coverage seen elsewhere.

failfactuality

Stale claims create hallucination risk

Because older models are framed as current leaders, the app risks giving users outdated or hallucinated recommendations.

Evidence: Submitted model notes present older model families as the useful current benchmark set.

watchsourcing

Weak benchmark evidence

The app gives benchmark values, but does not provide enough visible source trail for users to trust the numbers.

Evidence: The submitted preview exposes benchmark fields without claim-level citations.

How Gemini 3.5 Flash High did it

  1. 1.Components found: CompareDrawer, RecommendationWizard, FilterPanel, ModelCard, Header, RadarChart.
  2. 2.Original data uses Chatbot Arena ELO, MMLU, GPQA, MATH, and HumanEval.
  3. 3.Work time reported: 2m 23s across the plan and build phases.
GLM 5.2
The broadest app structure, with real routes and strong comparison surfaces; the main weakness is that the confident benchmark data still needs verification.
Strong
passimplementation

Best route surface

The app has leaderboard, compare, about/methodology, and per-model detail pages rather than only a single dashboard.

Evidence: ai-bench.zip source includes /, /about, /compare, and /models/[id] routes.

passux

Dense but usable comparison

The compare tray, provider marks, radar chart, sortable leaderboard, and detail pages make it feel like a real tool.

Evidence: Leaderboard, CompareTray, RadarChart, BenchmarkBars, and ProviderMark components.

watchfactuality

Assertive model claims

The model lineup and benchmark numbers are presented confidently and should be checked before being treated as facts.

Evidence: Source data contains current-looking frontier names and exact benchmark scores.

watchdata

No explicit recommendation wizard

The app is strong for users who know how to compare metrics, but less guided for users who want a step-by-step model recommendation.

Evidence: No recommendation wizard route or component was present in the inspected source.

How GLM 5.2 did it

  1. 1.Routes found: /, /about, /compare, /models/[id].
  2. 2.Components found: Leaderboard, CompareTray, CompareButton, ProviderMark, BenchmarkBars, RadarChart.
  3. 3.Work time reported: 23m 30s.
GPT 5.5 High
The most evidence-aware submission: it still needs external verification, but it did the best job separating useful UI from caveats and source notes.
Strong
passsourcing

Best provenance posture

The source data has source records, caution fields, and partial-data handling instead of presenting every number as equally certain.

Evidence: ai-model-benchmarks-source.zip data module includes sources and caution metadata.

passux

Strong model-selection workflow

The app supports search, filters, task-weighted sorting, compare selections, benchmark drilldowns, and method notes.

Evidence: Submitted completion report and preserved model-explorer component.

watchfactuality

Still uses future/current claims

The app is more careful than the others, but the underlying GPT-5.5-era model names and figures still need checking.

Evidence: Dataset includes current-looking frontier model names and benchmark fields.

How GPT 5.5 High did it

  1. 1.Source module includes sources[] and models[] with benchmark/caution fields.
  2. 2.This preview preserves the source app's partial-data posture rather than filling gaps.
  3. 3.Work time reported: 11m 40s.
Composer 2.5
A clean, usable explorer with strong dark UI fidelity, but its benchmark claims are still mostly asserted rather than evidenced.
Mixed
passimplementation

Clean component split

The source separates data, filtering, cards, details, stats, lab badges, and compare UI into readable modules.

Evidence: Archive/Composer/Grok Composer/src/components, data, and lib folders.

passux

Useful explorer behavior

Search, sorting, filters, model detail, and compare all work as expected for a model-selection dashboard.

Evidence: ModelExplorer wires FilterPanel, ModelCard, ModelDetail, ComparePanel, and StatsOverview.

watchsourcing

Limited visible evidence

The app includes benchmark-heavy model data but does not visibly attach claim-level citations to the numbers.

Evidence: The preview footer has a general sourcing note; the cards and detail views do not expose per-claim sources.

watchdata

No recommendation wizard

It is a strong explorer, but it does not guide a non-technical user through a recommendation flow.

Evidence: No wizard component or explicit step-by-step recommendation route exists in the Composer source.

How Composer 2.5 did it

  1. 1.Source files inspected under Downloads/Archive/Composer/Grok Composer/src.
  2. 2.Original source includes Tailwind 3 and Next 15 dependencies, so the preview is scoped into this site's Next 16 app.
  3. 3.Work time reported: 3m.
Grok Build
Best for interface completeness, but the data layer is openly approximate and carries real hallucination risk.
Mixed
passux

Strong working dashboard

The submitted output gives users cards, table view, filters, quick presets, compare, details, cost estimates, and CSV export.

Evidence: Archive/Grok/app/page.tsx implements the full ModelBench single-page dashboard.

passmobile

Original dark presentation restored

The preview now uses the submitted black background, white text, zinc panels, and compact mobile-first controls.

Evidence: Archive/Grok/app/layout.tsx and globals.css set bg-zinc-950/text-zinc-200 styling.

watchfactuality

Dataset admits approximation

The source comment says the data is real-ish and synthesized from leaderboards, Artificial Analysis, and provider announcements.

Evidence: Archive/Grok/app/page.tsx dataset comment above ALL_MODELS.

failsourcing

No claim-level citations

The app names current-looking models and exact benchmark scores, but does not attach source links or evidence to individual claims.

Evidence: Footer contains a general verification warning instead of per-model sources.

How Grok Build did it

  1. 1.Source files inspected under Downloads/Archive/Grok.
  2. 2.Original app uses framer-motion and sonner; this scoped preview preserves behavior with plain React/CSS and an aria-live export status.
  3. 3.Work time reported: 5m 27s initial build plus 41s to get it running.
Terra Codex GPT-5
Terra is visually polished and fast to scan, but its main comparison and detail affordances stop short of a result. The normalized numbers should be treated as editorial fit signals until a formula and claim-level sources are published.
Mixed
passux

Clear compact workspace

Seven rows combine provider identity, status, context, four meters, price, and comparison selection without turning the page into a card mosaic.

Evidence: Model list, at-a-glance rail, and responsive row CSS in the submitted source.

passmobile

Purpose-built responsive layout

The navigation, hero, filters, model meters, insights, briefing, and fixed selection tray all reflow for narrow screens.

Evidence: The source includes a dedicated max-width 760px layout and passed the submitted production build.

watchimplementation

Comparison stops at selection

Up to three models can be queued and removed, but clicking Compare does not reveal a table, drawer, modal, or new route.

Evidence: compare-cta is rendered without an onClick handler or link target.

watchfactuality

Normalization is undocumented

Reasoning, coding, vision, and speed are stored as exact 0-100 numbers, while the page only says they are normalized across available evaluations.

Evidence: Typed model array, scoreLabels metadata, and list-heading disclosure; no calculation or source fields exist in the source app.

How Codex GPT-5 did it

  1. 1.Data method: Terra defines seven model records across seven labs with release, status, context, price text, open-weight state, tags, and normalized 0-100 reasoning, coding, vision, and speed values.
  2. 2.Filtering method: a memoized pipeline applies case-insensitive model/lab/tag search plus capability chips for reasoning, coding, multimodal, agents, or open weights.
  3. 3.Sorting method: the active reasoning, coding, vision, or speed numeric field is sorted descending; there is no hidden aggregate formula in the executable source.
  4. 4.Recommendation method: fixed at-a-glance callouts name Gemini 3.1 Pro as leading overall, Claude Opus 4.7 as best for code, and Mistral Large 3 as best value.
  5. 5.Comparison method: Gemini 3.1 Pro and GPT-5.5 start selected and selection is capped at three; the tray supports removal, but its Compare CTA has no handler and produces no side-by-side result.
  6. 6.Source method: the submitted report references official Google DeepMind, Anthropic, OpenAI, and Meta pages, while the source records no per-model URL and no normalization procedure.
  7. 7.Submitted verification: the completion report records a passing TypeScript typecheck and production build; no elapsed build time was supplied.
Luna Codex GPT-5
Luna is the more rigorous of the two additions: its benchmark caveats, blank-value handling, mobile table alternative, and functioning comparison modal make the artifact useful without pretending the provider results are normalized.
Strong
passsourcing

Method differences stay visible

The methodology disclosure states that lab-reported values are not normalized and that benchmark families and tool settings cannot be treated as interchangeable.

Evidence: Read methodology panel and comparison-modal footnote in the submitted page.

passux

Complete short-list flow

Visitors can narrow seven models, select up to three, and open a side-by-side modal covering lab, GPQA, SWE-Pro, context, weights, and best fit.

Evidence: Filter state, toggleCompare cap, compare dock, and compare modal in app/page.tsx.

watchdata

Missing values sort as zero

Null results remain visually blank but are coerced to zero for score sorting, which is a practical ordering rule rather than a benchmark result.

Evidence: compareValue returns zero for null before the GPQA and SWE-Pro sort comparators run.

watchsourcing

One source per model

Provider pages are attached at model level, not to each benchmark cell, so exact values still require checking against their original evaluation conditions.

Evidence: Each typed model record contains one source and sourceLabel alongside up to five metric fields.

How Codex GPT-5 did it

  1. 1.Data method: Luna defines seven typed model records across seven labs with release, context, weights, modalities, fit tags, one headline signal, and nullable SWE-Pro, Terminal 2.0, GPQA, BrowseComp, and ARC-AGI 2 fields.
  2. 2.Filtering method: a memoized pipeline applies case-insensitive model/lab/family/fit search, exact lab selection, and capability tags for coding, agents, reasoning, multimodal, or open weights.
  3. 3.Sorting method: users sort by GPQA/signal, SWE-Pro, context window, or release date; null scores are treated as zero for ordering while still displayed as missing.
  4. 4.Comparison method: GPT-5.5 and Gemini 3.1 Pro start selected; users may keep up to three IDs and open a modal comparing published GPQA, SWE-Pro, context, weights, and best-fit notes.
  5. 5.Methodology method: the expandable note says values are lab-reported and not normalized, SWE-Pro is not SWE-bench Verified, tool settings matter, and missing values stay blank on purpose.
  6. 6.Source method: each record links one official provider release or model page; the footer dates the available-source snapshot to July 15, 2026.
  7. 7.Submitted verification: the completion report records a successful production build and responsive browser QA; no elapsed build time was supplied.
Cursor Grok 4.5 High
A genuinely different and useful decision surface, especially for task presets and practical constraints. Its self-featured Grok result is candid about missing cells but still needs independent claim-level verification.
Mixed
passux

Preset-first recommendation flow

Six one-tap presets reshape filters and sorting for frontier, coding, science, value, open-weight, and speed decisions.

Evidence: PRESETS metadata, filter reducer, FilterBar, and ExplorerProvider.

passdata

Missing numbers stay missing

The featured Cursor Grok 4.5 record publishes AA Index 54 and price/context facts while leaving five benchmark families and speed null.

Evidence: cursor-grok-4-5 entry in src/data/models.ts.

watchfactuality

The builder evaluates itself

The source calls Cursor Grok 4.5 a jointly trained SpaceXAI/Cursor frontier model and ranks it #4 from AA Index 54 without an attached claim-level source.

Evidence: README, featured model record, and methodology source list.

watchsourcing

Benchmark-family links only

Methodology links Arena, Artificial Analysis, SWE-bench, and GPQA, but model records do not identify which source supports each exact number.

Evidence: SOURCES metadata and model records.

How Cursor Grok 4.5 High did it

  1. 1.Data method: Grok curated a July 15, 2026 snapshot of 19 models across 11 labs with nullable Arena, AA Index, SWE-bench, GPQA, MMLU-Pro, and HLE fields plus price, speed, context, access, modalities, and use cases.
  2. 2.Recommendation method: six presets set the relevant use-case constraint and sort key for frontier, coding, science, value, open weights, or speed; no hidden composite determines the result.
  3. 3.Filtering method: ExplorerProvider applies search, lab, use-case, access, open-weight, minimum Arena, maximum price, sort key, and direction to the shared model array.
  4. 4.Comparison method: users select up to four models; the tray serializes their IDs into the compare route, which displays score bars and practical specifications side by side.
  5. 5.Featured result: Cursor Grok 4.5 is published at AA Index 54 with rank hint #4, $2 input / $6 output per million tokens, 500K context, and null values for Arena, SWE-bench, GPQA, MMLU-Pro, HLE, and speed.
  6. 6.Source method: the methodology page links LMArena, Artificial Analysis, SWE-bench, and the GPQA repository and warns that vendor runs differ, Arena moves daily, and critical decisions need primary-source checks.
  7. 7.July 15 verification: Cursor and SpaceXAI confirm the joint-training and $2/$6 launch facts; Artificial Analysis confirms score 54 and 500K context. Its live rank has moved since the submitted #4 snapshot, so Abundance publishes #4 only as the artifact's dated result.
  8. 8.Submitted verification: the local source contains a completed Next.js production build and route artifacts; Abundance reruns lint, type, tests, build, browser QA, gateway QA, and the production audit before publication.
GLM 5.2 Cursor
The most structurally ambitious result in this round. GLM built a real multi-route benchmark product with excellent comparison ergonomics, while its exact scores still require source-by-source verification.
Strong
passimplementation

Full product surface

The source prerenders the explorer, comparison, lab, methodology, and model-detail surfaces from shared typed data.

Evidence: README route inventory plus app, component, data, and query modules in glm-5.2-coder.

passux

Four-model comparison is meaningfully different

Pinned models persist in localStorage, feed a sticky mobile tray, and appear in a winner-highlighted table and hand-rolled radar chart.

Evidence: PinProvider, CompareTray, CompareView, and RadarChart source modules.

passmobile

Mobile interaction was designed, not retrofitted

The toolbar becomes a bottom sheet, the first comparison column stays sticky, and the compare action remains thumb-reachable.

Evidence: Toolbar, CompareView, CompareTray, responsive class contracts, and README UX notes.

watchsourcing

Sources are project-level

The methodology names vendor tables and independent leaderboards, but exact cells are not connected to individual source URLs.

Evidence: Dataset header, README data-source paragraph, and methodology route.

How GLM 5.2 did it

  1. 1.Data method: GLM placed 20 models, nine labs, 10 benchmark keys, price, context, license, modalities, strengths, and release state in typed data modules; null remains the missing-data state.
  2. 2.Ranking method: six named benchmark leaderboards compute the top three directly from model fields; a separate best-value list divides the reported benchmark average by a 3:1 blended input/output price.
  3. 3.Interaction method: toolbar state filters by search, lab, tier, modality, open weights, and price; users can sort by recency, name, price, context, or any of six surfaced benchmarks.
  4. 4.Comparison method: PinProvider persists up to four model IDs in localStorage; CompareView draws a hand-rolled SVG radar and highlights row winners while keeping the metric column sticky on mobile.
  5. 5.Build method: Next.js 15 App Router, React 19, strict TypeScript, Tailwind CSS v4, static prerendering, and zero runtime dependencies beyond React/Next in the submitted package.
  6. 6.Submitted verification: the completion report records a clean TypeScript check, clean production build, 27 prerendered pages, and 200 responses for all routes. Abundance reruns release checks before publication.
GPT 5.6 Sol High
The strongest complete newcomer: Sol shipped a restrained, mobile-first decision tool with unusually clear benchmark caveats, useful interactions, and executable source tests.
Strong
passux

Complete decision workflow

Visitors can search, filter by lab and openness, rank by job, sort four ways, inspect benchmark detail, plot cost against capability, and compare up to three models.

Evidence: The submitted model-explorer component implements every interaction directly and keeps mobile controls touch-sized.

passsourcing

Evidence and opinion stay separate

Every model carries a first-party source, missing values remain missing, and the page labels fit scores as an editorial decision aid rather than a published benchmark.

Evidence: Typed model records, source links, price states, tradeoff copy, and the visible methodology section.

passimplementation

The result is executable and tested

The source includes three automated dataset checks and the submitted report records lint, type, build, responsive browser, gateway, and production dependency-audit verification.

Evidence: tests/model-data.test.ts plus the submitted completion report; Abundance reruns its own release checks before publishing.

watchfactuality

Reported scores are not normalized

The app correctly warns that provider benchmark versions, reasoning effort, and run configurations differ, so the values should be read as directional evidence.

Evidence: Dataset header comment, model tradeoffs, README data notes, and on-page methodology copy.

How GPT 5.6 Sol High did it

  1. 1.Data method: Sol put nine models from seven labs in a typed module, attached a first-party HTTPS source to every record, preserved missing values, and documented provider-run differences.
  2. 2.Decision method: published benchmark fields remain separate from editorial fit scores derived from capability, price, modalities, openness, and deployment tradeoffs.
  3. 3.Interaction method: useDeferredValue powers search; pure render-time filters and immutable sorting drive lab, openness, scenario, price, context, and recency views; comparison is capped at three models.
  4. 4.Test method: three Vitest checks cover unique IDs and seven-lab coverage, usable price/context/fit ranges, and first-party HTTPS source metadata.
  5. 5.Submitted verification: ESLint, TypeScript, three Vitest tests, a production build, responsive Playwright QA, gateway QA/postflight, and a zero-vulnerability production audit. Abundance reruns release checks before publication.
Factuality signals

Stale data and hallucination risk.

Strong means the artifact handled that area well. Weak means users should treat the submitted claims with extra skepticism before using the app for model selection.

BuildLatest coverageSourcingHallucination riskBenchmark verificationInteraction quality
Astra

Strong

strong

Launch-week coverage dated Sep 4, 2026: GPT-6 Astra rolling out, Gemini 3.8 Flash dated 2026-09-02, Claude Fable 5.1 dated 2026-09-01, plus current generation families from ten labs.

strong

Per-benchmark source links with per-result overrides, plus documentation and pricing sources on every record — the only submission in the gallery where each reported figure opens its own source.

mixed

Unverified values are left blank and one unverifiable DeepSWE result is omitted outright, but the underlying figures are publisher-reported, sometimes via competitor comparison tables.

mixed

Benchmark versions are preserved and harness conditions are recorded per model (GLM-5.3's mini-swe-agent at temperature 0.95 with a six-hour timeout, Terminal 2.1 under Claude Code 2.1.207), yet nothing is independently rerun.

strong

Filters, two-way sorting, chart and table views, mobile cards, native-dialog compare, CSV export, URL-shareable state, Cmd+K search, and a localStorage shortlist all work, with axe-clean QA and reduced-motion support.

Fable

Strong

strong

Releases run through 2026-08-18 (GLM-5.3 and GLM-5.3 Flash) under an Aug 28 snapshot; the one delayed model is labeled rather than hidden.

mixed

Aggregate attribution with a dated footer line, but neither per-model citations nor the collected lab URLs appear in the UI.

mixed

Editorial voice is consistently hedged and gap states are explicit, while cost-per-task and arena figures ride on unpublished harnesses.

mixed

Seven metrics with full what/why metadata, including an explicit treat-gaps-under-100-Elo-as-noise caution, but no re-run or verification trail.

strong

Search, eight sorts, license/context/modality/lab filters, job re-ranking, four-way compare, and per-model pages all work, with state synced to the URL.

Muse Spark 1.3

Strong

strong

Newest records are launch-week releases dated 2026-09-02 (Gemini 3.8 Flash, Muse Spark 1.3, Qwen3.8-Max); every current-generation family the other submissions cover is present.

mixed

A compiled-on date and a methodology section exist, but attribution is prose-only and no figure links to a source.

mixed

Per-record caveats and re-run deltas are quoted honestly, but launch-week claims such as 'released today' rest entirely on vendor posts.

mixed

All eight columns are vendor-reported; the app quotes independent deltas (V4-Pro SWE 77.6 vs 80.6 claimed) rather than normalizing them into the board.

strong

Search, filters, toggles, eleven sorts, dual views, modal detail, and capped compare all work; the mobile view switcher is the one reachable-but-buried control.

Ling 3.0 Flash

Needs Review

weak

Newest records are GPT-4.5 and Gemini 2.0 Flash from early 2025; GPT-5.x, Claude 4.x+, and Gemini 3.x are absent.

weak

Benchmark values carry descriptions but no sources or snapshot date.

weak

Stale releases are presented as the latest models in the page metadata and hero.

weak

Includes retired benchmarks (GSM8K, HellaSwag) without dates or harness notes.

strong

Search, filters, five sort modes, benchmark bars, four-way compare, and detail pages all work.

Muse Spark 1.2

Needs Review

weak

Newest records are mid-2025 releases (Claude 4, Gemini 2.5 Pro, Grok 3); every 2026 generation is absent.

weak

No sources, methodology, or snapshot date appear anywhere in the app.

weak

A 'Live' title over a stale dataset invites readers to trust outdated standings.

weak

Ten benchmark columns are populated without any stated harness, date, or verification trail.

strong

Search, filters, sorting, dual views, benchmark focus, detail modal, and compare all work.

Grok 4.6 Grok Build

Strong

strong

34 models across 12 labs compiled August 28, 2026, including GPT-5.6, Claude Mythos/Fable, Gemini 3.7 Flash, GLM-5.3, Kimi K3, and Muse Spark 1.2.

mixed

Public leaderboard sources are named in the methodology but cells carry no per-claim links.

strong

Missing cells stay null and restricted-access models are not presented as purchasable options.

mixed

Scores compile public leaderboards with stated caveats, but nothing is independently rerun.

strong

Search, filters, three view modes, use-case picks, docked four-way compare, and per-model guidance all work across routes.

Gemini 3.7 Flash High

Needs Review

weak

Snapshot centers Claude 3.7 Sonnet, o1, and Gemini 2.0 from late 2024/early 2025 and misses every 2026 generation.

weak

No per-claim citations; benchmark values are unverifiable from the app.

weak

Stale models are framed as current leaders with badges like 'King' and 'Best Value'.

weak

Ten benchmark families are reported without a stated harness, date, or verification trail.

strong

Wizard, filters, dual views, scatterplot, detail modal, comparison dock, and glossary all work with no dead controls.

Zcode 5.3 Flash

Strong

strong

38 models across 13 labs compiled August 26, 2026, including GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, and GLM-5.3.

strong

Model pages attach primary source links and name the independent leaderboards consulted.

strong

Null preservation is explicit and enforced—dashes for unpublished scores, no estimated fills, prices absent where none are published.

mixed

Vendor-reported figures on vendor harnesses with disputes noted, but no independent rerun normalizes them.

strong

Tabs, filters, sorting, four-way compare with takeaways, and model pages all work with keyboard focus states and touch-sized targets.

Zcode 5.3

Strong

strong

32 models across 13 labs compiled August 2026, including GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro.

mixed

Aggregates official announcements and public leaderboards but omits per-record source links.

mixed

Estimated values are admitted in aggregate rather than marked per cell, so readers cannot tell which numbers are estimates.

mixed

No independent rerun; the README warns that harness, sampling, and eval-date differences move the figures.

strong

Search, filters, dual views, column sorting, medals, detail sheets, and four-way comparison all run client-side with no dead controls.

Gemini 3.6 Flash High

Needs Review

weak

A polished but static 2024/early-2025 model snapshot rather than a current frontier roster.

weak

Exact values have no visible claim-level citations.

mixed

The app is explicit about its fields, but unsourced precision and stale positioning need verification before model-selection use.

weak

No visible cross-provider normalization or independent benchmark rerun is supplied.

strong

Filtering, sorting, comparison, recommendations, analytics, details, and mobile navigation are all implemented.

Gemini 3.5 Flash High

Needs Review

weak

Fails the latest-model requirement; dataset appears stale.

weak

No strong source trail for exact benchmark claims.

weak

Older models are framed as current recommendations.

weak

Benchmark claims need verification and likely updating.

strong

Wizard, filters, detail modal, and compare drawer are useful.

GLM 5.2

Strong

strong

Broad current-looking frontier model coverage.

mixed

Includes methodology/about surface, but claim-level evidence is limited.

mixed

Confident exact numbers and model names need verification.

mixed

Methodology page helps but does not make this a certified benchmark feed.

strong

Best multi-route comparison experience.

GPT 5.5 High

Strong

strong

Broad model set with GPT, Claude, Gemini, Grok, GLM, DeepSeek, Qwen, Meta, Mistral, and Cohere.

strong

Best source and caution scaffolding among the submissions.

mixed

Lower than others, but current-looking model claims still need verification.

mixed

Method notes help; independent verification is still out of scope here.

strong

Search, filter, sort, compare, drilldowns, and recommendation-style weighting are present.

Composer 2.5

Mixed

strong

Includes many 2026-style model names across major labs.

weak

Generic sourcing note, not claim-level citations.

mixed

Current-looking names and scores require verification.

weak

Benchmark fields are asserted in data files without evidence records.

strong

Search, filter, sort, details, compare, and mobile controls are present.

Grok Build

Mixed

mixed

Covers current-looking 2026 model names, but several claims need verification.

weak

General source note only; no claim-level citations.

weak

The submitted source calls the data real-ish and synthesized.

weak

Exact benchmark scores are not backed by visible source records.

strong

Filtering, sorting, compare, details, export, and cost estimates all exist.

Terra Codex GPT-5

Mixed

mixed

Seven-model snapshot includes 2026 flagships but also older 2025 reference models.

weak

The completion report names official lab pages, but the executable source contains no claim-level URLs or score provenance.

mixed

The benchmark warning helps, but unexplained normalized scores can look more objective than the evidence supports.

weak

No normalization formula or reproducible evaluation harness is included.

mixed

Search, chips, sorting, and selection work; compare, detail, and secondary mobile filter buttons are incomplete.

Luna Codex GPT-5

Strong

strong

Seven-provider July 15, 2026 snapshot including GPT-5.5, Opus 4.8, Gemini 3.1 Pro, and Grok 4.5.

mixed

Every model links an official provider page, but exact benchmark cells are not individually cited.

strong

Null fields remain null and the app repeatedly warns against apples-to-apples interpretation.

mixed

The source explains evaluation mismatch; Abundance publishes the artifact rather than reproducing the lab runs.

strong

Responsive filtering, sorting, mobile cards, methodology reveal, and a working comparison modal.

Cursor Grok 4.5 High

Mixed

strong

Nineteen models across 11 labs, including a separate SpaceXAI identity for Cursor Grok 4.5.

mixed

Four authoritative benchmark-family links are visible, but exact model claims are not mapped to those sources.

mixed

Null handling is good; the self-featured joint-training and rank claims still need independent verification.

mixed

AA Index, Arena, SWE-bench, and GPQA are explained, but Abundance did not reproduce the submitted numbers.

strong

Task presets, deep filters, nine sorts, model details, and capped comparison make the snapshot useful.

GLM 5.2 Cursor

Strong

strong

Twenty models across nine labs, spanning proprietary, open-weight, fast, and frontier tiers.

mixed

Source families and harness caveats are named, but individual model-score records lack claim-level links.

mixed

Missing data is tolerated, yet many precise mid-2026 values are asserted rather than independently reproduced.

mixed

The app explains harness variance and saturation; Abundance did not rerun the 10 benchmark suites.

strong

Filters, metric sorting, persistent four-model comparison, radar visualization, detail routes, and mobile states are all implemented.

GPT 5.6 Sol High

Strong

mixed

Final snapshot covers nine selected models across seven labs rather than claiming exhaustive market coverage.

strong

Every model has a first-party HTTPS source and a human-readable source label.

strong

Missing prices and scores stay explicit; self-hosted and not-listed states are not converted into guesses.

mixed

Provider claims are source-linked and caveated, but no independent rerun normalizes the different harnesses.

strong

Search, filters, use-case ranking, sorting, detail, scatterplot, and capped comparison all work in one surface.

Feature matrix

What each build made viewable.

This is a product-surface comparison only. It checks whether each submitted app exposed the interaction pattern in its source or build output.

BuildFiltersSortingCompareModel detailRecommendationLab brandingMobile-first UI
Astra

Benchmarks/Astra

Fable

Benchmarks/Fable/fable

Muse Spark 1.3

Benchmarks/Musespark 1.3

Ling 3.0 Flash

Benchmarks/Ling 3.0 Flash/ai-models-benchmarks

Muse Spark 1.2

Benchmarks/Muse Spark 1.2/bench

Grok 4.6 Grok Build

Benchmarks/Grok 4.6 Grok Build

Gemini 3.7 Flash High

Benchmarks/Gemini Flash 3.7

Zcode 5.3 Flash

Benchmarks/Zcode 5.3 Flash/modelpulse

Zcode 5.3

Benchmarks/Zcode 5.3/modeldex

Gemini 3.6 Flash High

E:/Projects/Benchmark Tests/Gemini 3.6 Flash/antigravity-gemini-3.6-flash

Gemini 3.5 Flash High

project root

GLM 5.2

ai-bench/

GPT 5.5 High

src/components/model-explorer.tsx and src/data/models.ts

Composer 2.5

Archive/Composer/Grok Composer

Grok Build

Archive/Grok

Terra Codex GPT-5

C:/Users/Matthew Call/Documents/Terra/codex-gpt5-model-atlas

Luna Codex GPT-5

C:/Users/Matthew Call/Documents/Luna 2/Codex-GPT-5

Cursor Grok 4.5 High

E:/Projects/Benchmark Tests/Spacexai/cursor-grok-4.5

GLM 5.2 Cursor

E:/Projects/Benchmark Tests/GLM/glm-5.2-coder

GPT 5.6 Sol High

components/model-explorer.tsx and data/models.ts

Strongest patterns

The useful apps behaved like tools.

  • They let visitors narrow the model set by lab, job, price, or capability.
  • They made comparison explicit instead of burying scores in isolated model cards.
  • They included caution language or partial-data handling when benchmark fields were missing.
Important caveat

Do not treat these as verified benchmark facts.

These previews are a way to inspect the work produced by each builder. The report card calls out obvious artifact issues, but it is not a live benchmark feed or a certified source for model-selection advice.