v0.5.0 (2026-07-13)
v0.5.0 is a broad release that pulls in three directions at once.
First, the model catalog stops being a file you edit and becomes a service. Where v0.4.1 still shipped a committed model dataset regenerated by a CLI build step, the catalog is now owned by the backend, served from its store, curated live in the app, and baked into the public snapshot the gallery consumes. Adding or changing a model no longer touches the repo or a release.
Second, the product learns to operate at scale. Neither gallery loads the whole run history anymore: the console and the static site page, filter, and sort runs against lightweight summary cards and fetch a full record only when you open it. Launching runs becomes a matrix you fan out across harness/model combinations, with automatic retry on infrastructure failure and per-harness parallelism limits. Reviewing becomes a planned, tracked activity, where you declare the coverage you want, work an unreviewed queue, and read a live dashboard of what’s still missing against reference implementations, the authored correct build of each case a run can be compared to.
Third, a case can ask for more. A new full-stack test type has the model build a game and produce the game’s own 2D assets during the run, and asset generation gains a Blender-authored family of props, mechanisms, and skinned characters emitted as standard glTF.
This remains pre-1.0 software, built first for the people working on the project.
Features
Section titled “Features”The full-stack test type
Section titled “The full-stack test type”A new full-stack test type joins end-to-end and asset-generation. It is the same long-horizon, autonomous “build one playable static-site game from scratch” task, judged by a human who plays it and scored identically, with one addition: the model must also produce the game’s own 2D assets during the run (sprites, sprite sheets, particle effects, audio) rather than being handed them. An end-to-end case that needs real art pre-seeds it, so the test stays about code; a full-stack case removes that seam, and one model both draws the art and writes the code that ships it.
Unlike an asset-generation run, there is no separately-scored asset and no
action-log replay. The produced files are ordinary build inputs, judged only as
part of the running game. A case declares type = "full-stack" and nothing else
to opt in; the type alone selects the run image. A standing quality directive is
auto-prepended to every full-stack prompt (no placeholder rectangles or silence;
the build must be self-contained), so authors don’t restate it. See
manifests and
evaluation.
Runs execute in a new test-cabinet-full-stack-2d image, with the 2D
asset-generation binaries (draw, draw-sheet, particle-2d, sfx-synth,
sfx-sample, and music) baked onto PATH alongside the pinned combat-core
sample pack and gm-lite instrument bank. Those tools are present only while
the run is live, absent at the validation rebuild and the published-source
rebuild, so a build must commit its produced files and its npm run build must
not shell out to them.
The image is built on a new shared base-wasm layer (a pinned Rust → wasm32
toolchain with wasm-bindgen, wasm-pack, and wasm-opt), which end-to-end
runs now resolve too, so any e2e or full-stack build may author its core
simulation in Rust and ship a committed .wasm. The Junction case requires
exactly that: its city simulation must be written in Rust, compiled to
WebAssembly once during the run, committed as a .wasm build input, and
consumed by a Node-only build that never invokes cargo.
Six cases ship: Coil (easy) and five medium sim games, the Hollowdeep exemplar, Holdfast, Junction, Midway, and Valence. The console and static site gain a Full-stack catalog category to browse them.
The model catalog is backend-managed
Section titled “The model catalog is backend-managed”The model catalog now lives in the backend store, in the model, model_alias,
and model_price tables, instead of a models/ directory of per-model
TOML/Markdown files and a bundled packages/ui/src/app/data/models.json. The
backend serves it at GET /models and bakes it into the public R2 snapshot as
models.json (referenced by the snapshot index.json’s modelsKey), so the
static site still renders model metadata and prices without a backend
round-trip. There is no committed dataset and no build step in the middle.
Every run’s model appears automatically
Section titled “Every run’s model appears automatically”Any model with at least one recorded run shows up in the Models section, curated
or not. An uncurated derived model appears under its canonical model id,
resolved from the run record: the openrouter/ routing prefix is stripped for
the harnesses that require it (OpenCode and Kilo Code), and a trailing
OpenRouter variant tag such as :free is stripped for OpenRouter-accessed
harnesses (every harness except Codex, Claude Code, and Antigravity). That
normalization keeps one underlying model from splitting into phantom duplicate
entries.
Models are curated in the app
Section titled “Models are curated in the app”Curated model configuration is now edited in the web console and desktop app, in
the Models section (a write, so it requires sign-in). A model carries a
Test-Cabinet display name, one or more aliases (the run-record model ids it
covers, so a single entry spans the different ids harnesses report the same
model under), a provider, a provider logo supplied as an
svgl.app https:// URL and fetched and sanitized
server-side, a markdown description, and an OpenRouter slug. There are two ways
to add one: a blank Add model form, or Add this model seeded from an existing
run of an unknown model. Adding always goes through the form and an explicit
Save, the display name is required and never auto-generated, and it takes effect
immediately with no recompile or release.
Model slugs are paired with a harness family
Section titled “Model slugs are paired with a harness family”Each of a model’s aliases is now tagged with the harness family it is usable
with (Claude Code, Codex, Antigravity, or Others (OpenRouter)) because a slug
only resolves for the harnesses that speak its namespace. A Claude Code id like
claude-opus-4-8 means nothing to Codex; an OpenRouter id like
anthropic/claude-opus-4.8 only works through the OpenRouter-routed harnesses.
The Add model form pairs every model id with a family; New Run and Coverage then
filter the model dropdown to just the slugs the selected harness can launch, and
re-map the selection when you switch harnesses. Existing catalog data is
migrated in place on startup: native ids are reassigned to their family from run
evidence and slug shape, and every OpenRouter provider/model id stays in the
Others family. The separate OpenRouter slug field remains, used only for
comparable pricing.
Prices are recorded as a per-model history
Section titled “Prices are recorded as a per-model history”Comparable cost is still computed from OpenRouter’s per-token prices exactly as
before, but the price is now fetched by the backend rather than the removed
tcab catalog build. The backend records a model’s current OpenRouter price
when a run completes, capturing the rate in effect at that moment so promotional
pricing such as a launch-week discount is preserved for the runs that ran under
it, and again on a 24-hour periodic refresh, appending a new observation only
when the price actually changed. A :free-tagged OpenRouter run is priced at
the model’s base rate, never $0. A model’s detail page shows the retained
history as a graph and a table, one row per newly-observed price.
Neither gallery loads the whole cabinet
Section titled “Neither gallery loads the whole cabinet”Both hosts previously loaded the full set of runs: the console drained every
page of GET /runs into one array, and the static site inlined every full run
record into its JS bundle, then filtered, sorted, and paginated client-side.
That does not scale as the run history grows. Now a lightweight RunSummary
card (id, subject including the test type, timestamps, metrics, state, aggregate
rating, review count, and links) backs every run list, the home page, the
leaderboards, and the metrics views. A run’s full
record and its reviews load lazily, one run at
a time, only when its detail page opens. The backend serves the cards from a new
GET /runs?fields=summary projection; the
static site ships the snapshot’s summary
index as its in-memory
dataset and emits each full record as a lazily-fetched runs/<id>.json asset
instead of bundling it.
The console pages, filters, and sorts runs on the server
Section titled “The console pages, filters, and sorts runs on the server”The console’s run and model list pages are now server-paged. GET /runs?fields=summary gained a numbered-offset mode, offset + limit
returning { runs, total }, alongside the existing before cursor (still used
by the public-snapshot drain). It accepts testCase/model/harness filters,
a q free-text search over the lifted subject columns, and sort (date,
runtime, tokens, cost, rating, test type, test case, harness, model, or variant)
with dir.
To keep those DB-native, a versioned migration with an idempotent startup
backfill lifts the test type, run time, total tokens, comparable cost, rating,
and review count onto the run row, kept current as runs are recorded, reviewed,
and published. The list pages drive a numbered pager off total, debounce
search, re-query and reset to page 0 on a column-header sort, and pin
produced/in-progress runs to page 0. The static site runs the identical
filter/sort/paging over its in-memory summary index client-side.
The q search matches the raw recorded subject ids, not a model’s resolved
display name, and the case/variant pages scope to the bounded case set (the
summary filter has no per-variant filter).
Runs launch as a matrix, and retry themselves
Section titled “Runs launch as a matrix, and retry themselves”The New Run form now launches a matrix. A shared test (case, version, variant,
orchestrator, max runtime) sits above a list of harness × model combination
rows, and a run-count multiplier (default 1, up to 20) launches that many copies
of each, so a single submit fans out combinations × runCount runs. The fan-out
is client-side and sequential, each launch isolated in its own try/catch, so a
batch renders a per-run launch summary with links and partial failures stay
visible instead of aborting the rest.
Every launch also carries a retry count (default 1, clamped to 10). When a run
ends in an infrastructure or catastrophic failure, meaning Test Cabinet’s own
plumbing broke or the harness exited clean but the build won’t load, the backend
automatically re-enqueues a fresh job with the same request and an incremented
attempt, up to the retry budget. A timed-out or completed run (the model’s own
outcome) and a user cancel are never retried. The retry chain is bounded by a
strictly-monotonic per-job attempt column and a guard that fires the decision
only the first time a job crosses terminal, and it’s available on the CLI as
tcab run --retry-count N.
Picking the model is now a combobox: type a brand-new id or choose from “Known”
(the curated catalog) and “Previously used” (derived from prior runs). For the
provider-routed harnesses (OpenCode and Kilo Code), a provider dropdown applies
the openrouter/ prefix at launch so you enter the bare model id.
Per-harness parallelism, and two new run states
Section titled “Per-harness parallelism, and two new run states”A harness can now be given a maximum parallelism, edited in the console’s
Harnesses settings and stored in a new harness_config table (an unset value
means unlimited). The backend enforces it at claim time: it hands a waiting job
to a dispatcher only when its harness is under its cap, still oldest-first
across harnesses and still subject to the dispatcher’s global in-flight limit.
To make that legible, the run lifecycle gains two states.
Pending marks a run that is enqueued but
deliberately held back because its harness is at capacity, distinct from Queued
(simply next in line). Starting marks the window after a driver pod is up but
before the harness session begins, while it connects to the container runtime
and materializes the served definition, distinct from Running (the session
actually executing). Both new states count as in-flight for the active-run list,
coverage, and cancellation; starting is reaped if the backend restarts,
pending (like queued) is not.
Reviewing is a planned, tracked activity
Section titled “Reviewing is a planned, tracked activity”Reviewers can now declare the coverage they want and track progress toward it. A
per-account coverage plan (a ReviewPlan) lists harness × model combinations,
version-pinned cases, and a single runs-per-cell target; GET /review-plan/coverage expands it into a matrix of case × combination cells,
each reporting desired, completed, in_flight, and remaining counts
(global, so any run for a cell counts whoever launched it) plus a stale flag
when the pinned version is no longer the newest ingested.
The Coverage dashboard renders that matrix grouped by case, each cell and case
header showing a segmented progress bar (completed in green, in-flight in amber)
and a Trigger one / Trigger all missing button that launches exactly the
deficient runs through a connected worker. A separate Unreviewed worklist tab
(GET /runs?state=unreviewed) queues completed runs with zero reviews, newest
first, so nothing produced sits unlooked-at. A run’s aggregate points, earned
checklist weight over available weight averaged across its reviews, surface as a
selectable Points column on the run log.
The review editor itself is tightened. The pass/fail verdict is now a radio-like
control (arrow-key navigable, click-to-clear); a checklist item’s rail mark is
derived from its verdicts (a check when fully addressed and passing, a cross
when any fail); a Mark unplayable button sets every verdict to fail and every
rating to the worst tier in one step for a run that won’t load; and a review
item may carry name-only sub-items that each score as an independent point,
splitting the parent item’s weight evenly (weight × passed ÷ total) and keyed
by a composite <item>.<sub> id. Run-log rows also gain a right-click menu:
open in a new tab, open the test case or model, copy a shareable link, and
(where allowed) delete the run.
Reference implementations
Section titled “Reference implementations”Each case variant can now ship a reference implementation, an authored, in-repo,
buildable static game that is the correct implementation of the variant, “the
answer key rather than a model’s attempt.” It’s shown on a Reference tab on the
case page (in the console, desktop app, and public site alike) so a reviewer can
compare a model’s run against a known-good build. A reference is never seeded
into a run and never enters the prompt, because handing a model the answer would
defeat the case, and it is deployed out-of-band by a person rather than as part
of the run lifecycle. A variant opts in with reference_implementation = "reference-impl/<variant>", built with the case’s own [build] commands.
Publication follows a pull model. tcab publish-reference --env prod|staging
builds each variant, scrubs it with the run publisher’s secret-redaction pass,
deploys it to Cloudflare Pages, and records the served URL in a committed
test-cases/reference-builds.lock.json (keyed env → slug → version → variant). The backend ingests that lockfile during catalog reconciliation,
taking the entries for its own environment and syncing them into a
case_reference_build table, so scripts/reingest-cluster.sh refreshes
references the same way it refreshes catalog edits, tracking the environment
branch (staging → staging, prod → master). The earlier authenticated PUT …/reference-builds/… endpoint is removed in favor of this.
As of this release, the non-experimental cases’ references are published to staging. Relatedly, Cascade and Fathom now derive their reference screenshots by capturing their playable reference builds with Playwright, rather than from hand-authored HTML mockups. See results.
Blender props, mechanisms, and characters
Section titled “Blender props, mechanisms, and characters”Asset generation gains a Blender-authored family, sitting beside the in-house
voxel/meshed/skinned pipelines. Three new asset_kinds share one channel, a
headless Blender driven by a seeded build.py (bpy) script through the
sanctioned tcab-blend tool, and emit standard native glTF 2.0 rather than a
Test-Cabinet rig.json:
blender-propis static geometry, no armature (emitsmodel.glb).blender-mechanismis parented parts animated as glTF node-hierarchy clips, no skin (rigid machinery: deploy/fire/stow).blender-characteris a skinned mesh with an armature, weights, and pose-bone Actions (emits a skinned, animatedcharacter.glb).
Where the existing mc/sn/dc skinned path composites a signed-distance field
under a constrained op-log tool and animates from an inline rig.json, the
Blender kinds drive the industry-standard character pipeline and the emitted
glTF is authoritative, with rig and clips baked in and build.py itself the
recorded trace, re-executed for provenance. Prop and mechanism kinds can expose
runtime-drivable caller DOFs, parameters a consuming game sets each frame to aim
the asset (a turret’s yaw and pitch, a soldier’s aim), declared as
[[model.joint]] limits and carried in the glTF node’s extras where the
runtime reads them as a live slider. Three example cases ship: the
siege-rifleman character, the siege-rifle prop, and the siege-turret mechanism.
The kind runs in a new test-cabinet-blender image, the only run image not
built on the shared base, built on ubuntu:26.04 with Blender 5.0.x from apt
so it has both amd64 and arm64 builds. The console gives Blender assets their
own catalog tab and a native glTF viewer (clip picker plus live caller-DOF
sliders, serving all three kinds). See Blender
binaries.
Experimental cases are gated behind an env var
Section titled “Experimental cases are gated behind an env var”A test-case manifest gains an optional experimental flag (all types; absent
means false) to mark a case that’s still being iterated on versus ready to
publish runs for. The backend only offers experimental versions to the UI when
the deployment sets the truthy TCAB_BACKEND_ALLOW_EXPERIMENTAL; otherwise
they’re hidden from the catalog and refuse to resolve, so they’re never run or
published. The local overlay enables the flag; production and the bundled Tauri
app overlay leave it unset. (Ingest still walks the unfiltered case list, so
its reconciliation keeps pruning every stored version.)
The on-disk catalog is grouped by type and difficulty
Section titled “The on-disk catalog is grouped by type and difficulty”The test-case catalog moves from a flat test-cases/<slug>/<version>/ layout to
test-cases/<type>/<difficulty>/<slug>/<version>/, with all 104 case folders
regrouped under end-to-end/, full-stack/, asset-generation/,
adversarial/, and performance/ by easy/medium/hard. The two new levels
are organizational only: a case’s identity, type, and difficulty still come from
its test-case.toml, never its path. Core discovers cases by walking the two
grouping levels; reingest.sh and the various path references were re-nested
accordingly.
New and revised test cases
Section titled “New and revised test cases”Alongside the six full-stack cases and the three Blender cases:
- End-to-end cases got substantial revisions. Carom (v1.1.0) drops the Frenzy mode and folds the Gyre and Multi rules into the base game; Fathom (v1.1.0) models the sonar as a travelling wavefront, adds named predators and a redesigned AI, StarCraft-style fog, and a Kindle sensing variant; Meltdown is heavily reworked around a richer heat system with a Play → mode-select flow (three Containment difficulties plus special modes), rebalanced so mazing and heat are both required, and backed by a headless balance-sim harness.
- New asset-generation cases: the Caldera mechanism/prop set (
caldera-boiler,caldera-mortar,caldera-lance,caldera-runner, and more), the Sunfront muzzle-flash and structure cases, and four faction music themes (ironbound-theme,meridian-theme,geode-theme, and a retargetedthunderhead-themereimagined as a faction-neutral main-menu overture). Music and asset briefs now set mood and role only, with no prescribed instrument, key, or tempo, so they test creativity rather than instruction-following.
Console polish
Section titled “Console polish”- The Metrics tab labels each model’s bars with its catalog display name and gains a version scope control (current minor, current major, a specific version, or all).
- The Inputs tab distinguishes a seeded starter Script (a new presentation-only
kindon the[[spec]]table, used by the Blender case’sbuild.py) from a spec, and surfaces a case’s declared runtime packages. - List pagination and filters live in the URL, so a filtered/paged view is linkable and survives a refresh.
- The test-case leaderboard expands its columns, and a run’s start time shows on the Metadata tab.
The gallery deploy hook reports its outcome
Section titled “The gallery deploy hook reports its outcome”Publishing a run triggers a Cloudflare Pages deploy hook to rebuild the static
gallery, but a hook that returned without queuing a rebuild failed silently. The
backend now surfaces the deploy-hook outcome (deployHookFired), so a
no-redeploy is visible rather than presenting as a gallery that never updates.
Coverage counts the launched model id
Section titled “Coverage counts the launched model id”Coverage cells count runs against the launched (provider-prefixed) model id, so provider-routed cells, an OpenRouter model through OpenCode or Kilo Code, no longer read zero completed runs against a target they’ve actually met.
The leaderboard ranks published runs by their enriched score
Section titled “The leaderboard ranks published runs by their enriched score”A regression on the summary-backed leaderboard ranked published runs before their score was enriched from the case catalog; runs are now ranked from the enriched summary score.
Deployment
Section titled “Deployment”Passwordless Postgres auth with Microsoft Entra
Section titled “Passwordless Postgres auth with Microsoft Entra”A new crates/db-auth crate lets the backend and auth service authenticate to
Azure Database for PostgreSQL without a stored password. The pod authenticates
as a user-assigned managed identity via Azure Workload Identity, and the
Postgres “password” is a short-lived Entra access token, refreshed by a
background task that swaps in a fresh connection pool before the token expires
(in-flight queries drain on the old pool). Each service opts in with a
backward-compatible flag, TCAB_BACKEND_DB_AZURE_AD / TCAB_AUTH_DB_AZURE_AD,
defaulting off, so the existing static-password path is unchanged.
A new kustomize component, components/postgres-azure-ad, wires the
workload-identity ServiceAccounts (backend → backend DB, auth → auth DB, least
privilege) and sets the flags. It is authored and shipped but not yet enabled in
either overlay: the component and its companion patch stay commented pending the
cutover, which rolls the db-auth images first, then creates the in-DB Entra
principals and switches the Key Vault DB-URL secrets to the passwordless shape.
The full runbook is in Kubernetes deployment.
azure-staging is a full mirror of azure-prod
Section titled “azure-staging is a full mirror of azure-prod”The staging overlay, previously a thin variant, now mirrors production: it gains
the Key Vault CSI + workload-identity secret population, the VPN-only internal
ingress at *.staging.tcab.testcabinet.ai, subscription-auth for Codex and
Claude Code, the catalog-ingest sidecar tracking the staging branch tip
(superseding the base CronJob), Grafana anon-admin lockdown, and
public-URL/web-config repointing. It is a faithful prod rehearsal differing only
in namespace, environment, hostnames, and image tags. Single-replica
coordinator, auth, and Key-Vault-sync pods also gain a safe-to-evict: "false"
annotation so the cluster autoscaler can’t drain the node under them.
Image pins
Section titled “Image pins”azure-staging rolls to the v0.5.0 candidate build (9bb4858e), exercising
this release end to end before production. azure-prod stays pinned to v0.4.1
(b3c2b322) and rolls separately once staging has proven the build, with
service and run-container images pinned independently as always. See rolling
prod service images.
Removed: tcab catalog, models/, and add-model.mjs
Section titled “Removed: tcab catalog, models/, and add-model.mjs”With the catalog owned by the backend, the on-disk pieces are gone: the tcab catalog CLI subcommand, the scripts/add-model.mjs scaffolder, the models/
directory of per-model <slug>.toml/.md files, and the bundled
packages/ui/src/app/data/models.json. Curating a model is now an in-app edit;
see Adding or Updating a Model.