v0.6.0 (2026-07-25)
v0.6.0 is a release about evidence.
Every prior version treated a run’s verdict as something a person produced from scratch: play the build, work the checklist, decide. That scales badly, and it is unreproducible, because two reviewers reading the same rule can disagree about whether a build met it and neither can show their work. This release changes the shape of that. A case can now demand that a build expose a debug API, and the platform drives that API to decide the checklist itself, arranging scenarios, injecting real input, sampling the pixels the game actually painted, and recording every assertion it made. What arrives at a reviewer is a pre-filled form rather than a blank one, each verdict backed by a list of checks and a video of the model’s build beside the same scenario running on the reference implementation.
Second, the catalog learns to age. A version with runs recorded against it is now mechanically immutable, because editing one silently invalidates every score already computed from it. Immutable is only livable if there is a way to correct a version without replacing it, so a version can carry errata, and an erratum can retire a review point from scoring when the point itself turns out to be wrong. Review edits, likewise, stop being silent overwrites and become a public revision history.
Third, a case can now decline to specify anything at all. The new game jam type hands a model a theme and nothing else, with no spec, no reference, and no checklist to conform to, and grades the result on a 💩→💎 scale instead of pass/fail.
This remains pre-1.0 software, built first for the people working on the project.
Features
Section titled “Features”Instrumentation: a case can require a debug API
Section titled “Instrumentation: a case can require a debug API”A case declares an [instrumentation]
table naming the window property its build must install a debug API on:
[instrumentation]handle = "__carom"The handle is reporter-side and never seeded. The case’s seeded spec documents
the same API independently, as an ordinary game debug affordance. It must never
name The Test Cabinet, a test, or validation (no __tcab, __test,
__validate), because a build that knows it is being graded is a build that can
game the grader.
Three core operations are identical across the catalog, so the driver’s
lifecycle is uniform: reset(options?) (return to a known state; options.seed
seeds all randomness), step(seconds) (advance the simulation by exactly that
much sim time through the build’s fixed timestep, without waiting on real time),
and snapshot() (a pure read returning a JSON-serializable object of the full
observable state).
Beyond those, a case enumerates its own control operations by name and
signature, such as Carom’s startMatch(mode), serve(), setScore(p1, p2),
and setPaddle(side, state). It also enumerates its input operations
keyDown(code) / keyUp(code) / press(code), which take a standard
KeyboardEvent.code and route through the build’s real key handling, so a check
exercises the actual bindings rather than bypassing them.
The governing rule is the precondition guardrail: a control op may establish
preconditions and fast-forward setup, routed through the real systems, and it
may never fabricate the outcome a check observes. Set the score to 10–8 and then
drive a real point; never expose a declareWinner(). Alongside the API a case
mandates a deterministic core (fixed timestep, render-free, seedable RNG) and a
read-only debug overlay, off by default on a documented toggle key.
The manual clock
Section titled “The manual clock”The contract that makes all of this exact is the clock. A build normally
advances on a fixed timestep fed by the animation loop from the wall clock, so
it plays in real time for a person. reset() and step() both switch the build
to manual stepping: the wall clock stops feeding the simulation, and from that
point step() is the only thing that moves it. Successive steps advance by
exactly the time asked for, with no stray wall-clock frames slipping in between
calls. setAutoStep(true) hands the clock back.
That split is what lets one script be both exact and watchable. A check asserts
under the manual clock, where the result is identical no matter what else the
machine is doing; then it calls setAutoStep(true) and waits in real time so
the render loop actually animates for the video capture. The verdict comes from
the deterministic half, the proof clip from the real-time half.
Validation scripts decide a checklist point
Section titled “Validation scripts decide a checklist point”A graded checklist unit, meaning an item scored as a whole or each sub-item of
one that is broken down, opts in with a validation table naming a driver
script and the media it produces:
validation = { script = "validation/ball-spin/stationary.mjs", outputs = [ { id = "straight", name = "Straight return, no curve", kind = "video" },] }Scripts live under <version>/validation/, are reporter-side and never seeded,
and the whole directory is materialized into a run’s served definition so
sibling imports like _helpers.mjs resolve. Each is an ES module
default-exporting async function drive(api, ttc).
The api is a thin wrapper over the declared handle: reset, step,
snapshot, a generic call(method, ...args) into any control op, wait(ms)
for real-time animation, screenshot(id) for a declared image output, and the
anti-cheat read pixel(u, v), which samples the largest canvas’s backing store
at normalized coordinates. A check that reads pixel sees what the build
actually drew, so a build cannot pass a color requirement by reporting a palette
it never paints.
The ttc kit is the assertion half, deliberately outside the model’s api and
never seeded. It follows googletest semantics: expect* records a failure and
continues so one run collects every problem, assert* stops the script and
unwinds into an ordinary failed verdict rather than a conformance error.
Matchers (Eq, Ne, Gt, Ge, Lt, Le, Close, Ok) each record an
Assertion { label, pass, expected?, actual? }, passing ones too, because the
kept assertions are the machine-readable proof of the verdict. A unit passes iff
every assertion passed.
Thirteen cases ship 876 validation scripts between them.
Reference baselines, and side-by-side review
Section titled “Reference baselines, and side-by-side review”An automated verdict answers “did it pass”. It does not answer “what should this have looked like”, so every scripted check is driven twice, against two different builds.
The model’s build is driven per run, writing its media into the collected tree
under .vendor/validation/. The reference implementation is driven separately by
tcab capture-baselines, once per case version, writing into
<version>/validation-baseline/<variant>/, committed to the repo. A run never
re-drives the reference; it just points at the committed baseline. Both sides
use the same flat <verdict>__<output>.<ext> naming, so only the directory
distinguishes them. There are 1,219 committed baseline files today.
Live, the actual media is served at GET /runs/{id}/validation/{file} and the
baseline at GET /test-cases/{slug}/versions/{version}/validation-baseline/{variant}/{file}.
Both are published into the public snapshot and, as with proof clips, transcoded
from .webm to H.264 .mp4 at snapshot time so they play on iOS and Safari.
The side-by-side is composed at render time rather than at capture time, and nothing in the pipeline stitches the two halves into a combined file. Each stays an ordinary standalone clip or still under its own root, and the reviewer pairs them: each output renders as Reference beside This run, inline beneath the exact verdict it backs. For video the pair shares one Play/Pause control, which restarts both clips from the top and runs them together so the two advance frame-for-frame, both looping, both muted so playing them together isn’t a cacophony.
The auto-fail gate
Section titled “The auto-fail gate”A drive records whether it ran, true only against a conformant build: the
handle was installed within a bounded 10 s wait, every call returned, the return
value was well-formed, and every declared output was produced. If any gating
script did not run, the run is classified Catastrophic rather than
Completed, because the harness exited cleanly and the model claimed completion
but the output cannot be evaluated. Such a run is publishable, carrying its
broken source, and is never put in front of a reviewer, since there is no
checklist to score.
Two escape hatches keep the gate honest. It degrades rather than fires on infrastructure trouble: a host with no browser, or a serve failure, produces no results at all, and an empty result set never trips the gate. And an erratum can un-gate a single point (below), so a check that turns out to be buggy stops failing runs outright.
tcab capture-baselines
Section titled “tcab capture-baselines”Baseline capture previously lived inside tcab publish-reference, a deployment
command needing --env, a Cloudflare project, and credentials, which coupled a
purely local authoring step to a deploy. The new tcab capture-baselines is that capture and nothing else:
it needs only a browser and the case’s own toolchain. It resolves the case,
builds each targeted variant’s reference implementation with the case’s
[build] commands, drives every scripted item against it, and writes the
declared outputs. Conversely publish-reference gained --skip-baselines,
since the capture dominates its runtime. Run it whenever you add or change a
validation script, or change the reference implementation it is driven against.
Review checklists gain a categories grammar
Section titled “Review checklists gain a categories grammar”A case opts into a new checklist grammar with [review] format = 2, replacing
[[review_item]] with [[review.categories]] and
[[review.categories.items]]. The two grammars are mutually exclusive per case,
and the opt-in is purely additive, leaving existing manifests untouched.
A category carries only id and title; it is a grouping, not a verdict. The
item beneath it is the scored leaf, recorded under the composite verdict id
<category>.<item>, so item ids need only be unique within their category. The
scoring change is the point of the exercise: a category’s weight used to be
split evenly across its sub-items, making every verdict a fraction. Now each
item carries its own weight and is credited by it, and the category’s weight is
the sum, so with the default weight = 1 each point is worth exactly one whole
point and items within a category can be weighted independently, which the even
split could not express.
A variant may extend a common category by reusing its id, folding its items in and summing the weight, rather than forming a second same-id group.
Asset-generation cases are judged on one overall rating
Section titled “Asset-generation cases are judged on one overall rating”An asset-generation case asks for a single artifact, a sprite, a walk cycle, a
material, and whether it reads as the thing the brief describes, whether the
walk carries weight, and whether the material sits right under light are
judgments about the asset as a whole. Scoring them point by point invents a
precision the judgment does not have, so the whole review is now one rating.
Every asset-generation case drops its [[review_item]]s and declares a single
overall scoring domain in place of whatever domains it carried; being the only
domain, the reviewer’s one rating becomes the run’s rating. These runs carry a
rating and a writeup but no point score, which puts the weight squarely on the
brief, both what the model is asked to satisfy and the only thing the rating is
given against.
floe-bear is the one published asset-generation case, so its frozen v1.0.0
keeps the checklist its runs were judged against; its unreleased revision
becomes v1.1.0 (a scoring change, not the prompt rewording it started as) and
carries the new single-rating model. The asset-generation manifest and
evaluation pages, every asset-kind authoring guide and quickstart, and the
variant-creation guides, whose “add review items” step is now gone, follow.
Errata: correcting a version without replacing it
Section titled “Errata: correcting a version without replacing it”A run is grouped in the metrics by its exact (slug, version) key, so any
scoring-affecting fix forces a version bump, and the bump moves every existing
run onto a version nothing else references, quietly dropping it from that
version’s graphs. An erratum acknowledges the problem while the version, and its
runs, stay put.
Errata are not part of test-case.toml. A version folder may carry an optional
errata.toml beside its manifest, auto-discovered with no manifest key
declaring it, so it can be added to an already-reviewed version without touching
the reviewed definition. Like the changelog it is site-facing and never seeded,
and the mechanism is shared by every test type. Each [[erratum]] carries an
id, title, Markdown body, and optionally a date, a severity
(info/minor/major), an affects_scoring flag, a resolved_in version
(which need not exist yet, since the fix may merely be planned), a variant
scope, and a review link naming a verdict id.
Because errata live in the tree the backend ingests from a git checkout,
publishing one needs no tcab release: commit, push, re-ingest. They surface as
an Errata tab on the case detail page and as a “Known errata for this version”
callout on a run’s Verdict view, resolved by the run’s version and variant so a
reviewer weighs known issues before scoring. See Publish
errata.
Retiring a review point from scoring
Section titled “Retiring a review point from scoring”An erratum with a review link may set exclude_from_score, which does two
things. It stamps the point unscored, so score_checklist skips it in both
earned and total, leaving the point visible, still checked and still driven,
but no longer counting. And it un-gates the point’s automated drive: a
validation script on an excluded point can no longer auto-fail the run and rate
it broken.
That pairing is the whole design. When a review point turns out to be ambiguous, or its validation script turns out to be wrong, the fix would otherwise be a version bump that evicts every recorded run from the version’s metrics. Now it is an erratum.
Frozen versions
Section titled “Frozen versions”The backend’s runs reference a case by slug and version but do not snapshot the prompt and specs they were produced from, so editing a version directory in place silently invalidates every run already scored against it. The runs stay in the metrics and on the leaderboard, produced from inputs that no longer exist, with nothing in the data saying so. Valence v1.0.0 was edited exactly that way this cycle, and had to be recovered by splitting the changes out into v2.0.0.
The rule (“to revise a case, add a new version”) already held. What failed was
remembering, at the moment of editing, that a version had been used, so it is
now enforced mechanically. A version directory carries a .frozen marker
recording a digest, a timestamp, and a reason; the digest is a SHA-256 over the
directory’s git ls-files -s listing, which covers every tracked path, blob
hash, and file mode, so any edit, addition, deletion, rename, or mode change
moves it.
Two gates read the git index rather than diffing against a base branch, which is
why the check needs no merge base, no fetch depth, and no toolchain, and holds
on every branch and every history shape. A pre-commit hook catches the mistake
as it is made, and a CI check backstops --no-verify and machines without hooks
installed.
Unfreezing is deleting the marker in its own reviewable commit, deliberately not a flag or an override, because it means accepting that recorded runs no longer match their inputs, and that should not be a quiet decision. Every superseded version of the seven non-experimental multi-version cases is now frozen. See Frozen versions.
Review edits become a public revision history
Section titled “Review edits become a public revision history”Editing a submitted review worked, but as a silent in-place overwrite. It is now
auditable. A changed re-submit requires an edit note (422 without one), records
a row in a new review_revision table, and stamps a new edited_at, so
reviewed_at now always means “first submitted”. An identical re-submit is a
no-op.
The diff is computed rather than written: a structured prior→new ReviewDiff of
rating, verdict, and writeup changes, where a rating change carries from = None for a newly rated domain and to = None for a dropped one, and a verdict
change carries a note_changed flag so a note-only edit registers even when the
status held. Both the read shape and the snapshot shape carry editedAt and
revisions, so the history is public, rendered as a ReviewHistory panel on
the single-review page, with an “edited” marker on review headers. Revisions are
stored newest-last, so replaying their diffs from the original walks the review
forward.
The edit gate also relaxed: a reviewer may revise their own review on any run the console can submit to, not only locally-produced runs.
The game jam test type
Section titled “The game jam test type”A game jam is a full-stack case with its specification removed. It hands the model a theme and asks it to invent and build a complete, playable, enjoyable game of any genre, producing its own 2D assets during the run exactly as a full-stack case does. It is the most open-ended type in the cabinet: it measures design, scoping, and taste rather than adherence to a spec.
A jam is not a test case. It lives in a sibling top-level
game-jams/ directory laid out
game-jams/<slug>/<version>/, with no type/difficulty grouping, and parses
through its own game-jam.toml schema. It has no difficulty, because a jam is
inherently unclassified once the model decides what to build, and no variants,
because a jam is one theme and a differently themed jam is a different jam. It
declares no specs, references, or domains; any test-case-only key is a parse
error.
Jams run in their own test-cabinet-game-jam image (overridable with
TCAB_CONTAINER_IMAGE_GAME_JAM), built from the full-stack 2D image so it
carries the six asset-generation binaries, the baked audio packs, and the
Rust→wasm toolchain, plus date so a model can read the clock and pace itself
against its 8-hour budget. Discovery folds them into the same catalog, and they
surface under a new Other console section (Game Jams and Tournaments) rather
than on the Test Cases page.
Jams are graded, not passed or failed
Section titled “Jams are graded, not passed or failed”There is no spec to conform to, so there is no pass/fail checklist. A jam is
graded on a five-tier scale, 💩 Broken (0), 🙁 Not great (1), 😐 Neutral (3), 😀
Great (5), 💎 Incredible (10), across seven default categories: Playability, Fun,
Theme, Presentation, Audio, Polish, and Creativity. A category is worth weight × 10 and earns its tier’s points times its weight; a run’s score is the total,
and the leaderboard ranks by average across reviews, the same points-based
ranking every other type uses.
The reviewer also gives the whole game one overall grade on the same scale,
supplied directly and never derived from the categories. It becomes the run’s
rating badge, standing in for the per-domain rating a jam does not carry, and
with multiple reviews the displayed value is the worst any reviewer gave,
mirroring how a domain-scored run takes the worst across its domains. It rides
the checklist under the reserved id overall, so it needs no separate storage,
and being undeclared it is excluded from the point score.
Eight jams
Section titled “Eight jams”Every jam ships at v1.0.0 with an 8-hour budget:
- Dead Man’s Switch 🚂 is a game of any genre built around both trains and tension, neither reduced to garnish. (Renamed from “Trains & Tension”.)
- Well, Well, Well 🌀 makes gravity the whole game, in a star system that won’t hold still: worlds with real mass, in motion, pulling on each other and the player, so no path travels straight.
- Grace Period 🛡️ is a strategic bullet hell: slow, readable, telegraphed, with defense as a toolkit of distinct answers rather than one catch-all button.
- Band of Bots 🤖 is a single-player battle royale where the player leads an AI squad. It is a capability probe on three axes at once: a large field of independently acting AI running smoothly, allies that read as competent, and a low-friction human↔AI coordination UI.
- Outside the Box 🧩 is a puzzle game of any kind, with at least eight hand-designed levels and a real difficulty ramp. It probes whether a model can predict how hard its own puzzles are.
- Dealer’s Choice 🃏 is an original card game with custom cards. The central requirement is a negative one: it must not play like Magic, Hearthstone, Slay the Spire, or Balatro. Originality must be mechanical.
- Comfort Zone 🏡 is a cozy game: a warm, low-pressure world a player doesn’t want to leave. It probes whether a model can hit a feeling rather than a mechanic.
- Plot Twist 🔀 is a choice-driven narrative game whose choices genuinely branch, the line between a game and a book, with depth preferred over length.
The standing jam preamble was slimmed to carry no build or tooling detail (that moved into each jam’s own prompt) and now tells entries they are judged competitively against other models’ entries from the same theme, on presentation, polish, theme, audio, and creativity, the one place jam framing is deliberately comparative.
It also dropped its “a clear way to win or lose” mandate, which belongs to individual jams rather than the type, since a cozy or open-ended build may legitimately have neither. Every jam’s verify step was reworded to the inclusive “play it through to a natural end (a win, a loss, a cleared goal, or a satisfying session)”, the explicit expectation was re-added to the two jams that had been relying on the preamble for it, and the default Playability review category was reworded to match, so a reviewer grading a jam that deliberately has no win state is no longer reading a requirement the prompt stopped making.
Repeated jam runs must build something distinct
Section titled “Repeated jam runs must build something distinct”A jam can be run against the same model repeatedly, which invites near-copies.
Every jam prompt now requires a player-facing README.md (premise, goal,
controls, core loop, with no implementation detail) and that README is captured
into the run record (gameJamReadme, truncated at 16 KB on a char boundary)
whether or not the run is ever published.
On a later run of the same jam with the same harness and model, the driver
fetches those prior READMEs from GET /game-jams/{slug}/prior-readmes and seeds
them into a previous-entries/ folder, alongside a prompt section requiring an
entry that is genuinely distinct: a different core idea, genre, or central
mechanic, not a reskin or a sequel. The folder is reference material rather than
part of the submission, so it is excluded via .git/info/exclude, a local,
uncommitted ignore chosen so the model’s own .gitignore stays the model’s file
to own. A first run sees no prior entries and renders exactly as before, and a
lookup failure degrades to no entries rather than failing the run.
Accounts: pictures, a Reviews tab, and activity charts
Section titled “Accounts: pictures, a Reviews tab, and activity charts”Reviews are attributable work, so reviewers get a face. A profile picture is
stored on the auth service (PUT/DELETE /auth/profile/picture, open GET /auth/users/{id}/picture), center-cropped and downscaled client-side before
upload, and Account gains a pictureUpdatedAt that doubles as a has-picture
flag and a cache-bust version. The live console reads avatars straight from the
auth service, so it works in a local cluster with no R2 credentials; the public
site gets them through the existing snapshot bake, which emits a content-stable
pfp/<id> object each review references by pictureKey. A shared Avatar
primitive renders picture-or-initials in the top bar and beside reviewer names,
and the standalone “Reviewing as…” notice moved into the Verdict action row.
The account section also gained a Reviews tab, the account’s submitted reviews
as a proper run-log-style table with show/hide and drag-to-resize columns
(headers deliberately don’t sort, since the endpoint is a fixed newest-first
server page). It also gained a rebuilt Profile tab: a full-width identity card
above three ring charts breaking down recent review activity by test case,
model, and rating given, backed by a new GET /account/review-stats over the
100 most recent reviews.
Coverage becomes plans and groups
Section titled “Coverage becomes plans and groups”Reviewer coverage was one plan per account. It is now reusable groups plus
multiple named plans, in a tabbed account section (Profile | Reviews | Coverage
| Groups). A group is a named, reusable set of combinations or cases; a plan
references groups as pointers, so editing a group reshapes every plan using it,
and may also pin one-off members, which the backend unions and de-duplicates.
Each plan carries its own runs_per_cell. Cases in a plan’s dashboard are
collapsible, showing only the overall progress bar until expanded. Existing
plans are carried across the upgrade by an idempotent startup backfill that
copies each legacy plan into one named “My coverage plan”.
A “passable” rating tier
Section titled “A “passable” rating tier”A fifth run-quality tier lands between great and scuffed, making the scale
flawless > great > passable > scuffed > broken. It covers a run implemented to
spec and playable, but with rough edges beyond a great run’s minor issues,
noticeable though not enough to deviate from the spec or impair playability. It
is defined once in the core and regenerated into the contract, with the UI
mirrors (badge, chart colors, metrics tally) updated in lockstep. It is fully
additive: existing great and scuffed runs keep their tokens, so there is no
data migration.
The Reviewing tab
Section titled “The Reviewing tab”A read-only Reviewing tab on the case detail page shows what a run of that case
would be graded against: the scoring domains, the rating scale, and the weighted
checklist, with no verdicts, since it is tied to no run. The run Verdict tab’s
breakdown was extracted into a shared ReviewChecklist component with a
“definition” mode so both surfaces render identically.
Harness telemetry joins the run’s trace
Section titled “Harness telemetry joins the run’s trace”Test Cabinet’s own processes were instrumented, but the harness, the third-party
CLI doing the actual work inside the run container, was a hole in every trace. A
run now resolves the ambient OTLP configuration once and translates it per
harness into container environment variables and config files, gated entirely on
the existing OTEL_EXPORTER_OTLP_ENDPOINT master switch. A deployment that
already exports telemetry exports its harnesses’ telemetry with nothing extra to
turn on.
Claude Code, OpenCode, and Goose emit traces, metrics, and logs; Codex and Kilo
Code emit traces and logs. Claude Code and OpenCode join the run’s own trace
through a propagated traceparent; the rest are correlatable by resource
attribute (tcab.harness, tcab.test_case, tcab.variant, tcab.model,
percent-encoded since model ids routinely contain / and :). Cline, Pi, and
Antigravity are documented as unsupported with the reason each cannot be wired
up.
Two details are worth calling out. Codex’s metrics exporter is set explicitly,
because its default ships run metrics to a vendor, and log_user_prompt is
forced off. Export intervals are shortened to 1 s, because a run is short-lived
and default batching loses the session tail. Each harness now has its own
Telemetry documentation page.
Batch enqueue, and a coverage matrix that loads
Section titled “Batch enqueue, and a coverage matrix that loads”Triggering a coverage matrix’s missing runs launched them one at a time, so a
“trigger all missing” of ~1,600 runs was ~1,600 serial round-trips with the UI
frozen throughout. A new POST /jobs/batch carries many launches in one
request, enqueued in one chunked bulk insert. Each run is validated and minted
independently, so one malformed entry is reported as its own error without
aborting the rest, and the response returns one result per run aligned by index.
The coverage handler itself looped over every case×combination cell awaiting two
sequential COUNT(*) queries each, so a large plan fanned out into 1,000+
serial round-trips, taking 5–10 s to load and re-running in full after every
trigger click. It is now two grouped queries, each a single COUNT(*) … GROUP BY scoped to the plan’s slugs, looked up in memory.
Harness Error is its own run state
Section titled “Harness Error is its own run state”A model that drives the harness to exit non-zero is a real, reportable model
outcome rather than a Test Cabinet fault, but it was being recorded as an
infrastructure failure, which is ours and never publishable. harness_error
splits it out.
It is publishable, but as a statistic only: a new publishes_artifacts()
predicate covers just Completed, Catastrophic, and TimedOut, so a harness
error releases no source repo and no playable build. It is deliberately
retryable, because a subscription auth-token refresh surfaces here and can
self-heal on a bounded retry, while a model that genuinely crashes the harness
burns its retries and then settles as a recorded harness error. Classification
happens at run time and is persisted, so existing records keep their
infrastructure state; only new runs are affected.
A model’s Stats tab gained a ReliabilityRing breaking its published runs into
completed / harness errors / timeouts, the two publishable failure tiers
alongside clean completion.
The v2 case wave
Section titled “The v2 case wave”Thirteen cases were rebuilt onto the instrumented format: Carom, Cascade,
Fathom, Floe, Shatter, Spectra, Wireworm, and Meltdown (end-to-end), and Coil,
Valence, Arc Foundry, Deepcore, and Locomotivation (full-stack). A v2 case
carries self-contained specs (each file states its own rules rather than
cross-referencing), a mandatory specs/instrumentation.md, a format = 2
checklist whose items each assert exactly one observable behavior, a
validation/ tree, a per-variant reference-impl/, and committed baseline
media.
The upgrade landed in two shapes. Six cases (Carom, Cascade, Fathom, Floe,
Shatter, and Valence) had runs recorded against their v1, so they got a
genuinely new v2.0.0 directory and their v1 was preserved verbatim and frozen.
The other seven had no graded runs to protect and were upgraded in place at
v1.0.0.
Genuinely subjective points such as art, HUD fit, and audio deliberately carry no validation script; UI-state points carry one only to reach and capture the screen for human judgment. A new spec-and-prompt editorial guide covers writing this kind of case.
Three new full-stack cases
Section titled “Three new full-stack cases”- Arc Foundry (medium) is an electro-industrial GemTD reskin. Rocks placed at a scrap press roll a random salvaged component at a random quality tier; you keep exactly one per level while the rest harden into inert maze blockers, climb a five-rung quality ladder, and fold recipes into a dozen combination towers while the Load paths an ordered-waypoint maze.
- Deepcore (medium) is a Motherload reskin. Drill a lone prospector down through banded rock, haul ore up against a jetpack-fuel-and-weight budget, scan for two buried exotics, and race an unstable Core Sample up on a 90-second timer to fabricate a five-part escape rocket.
- Locomotivation (medium) is a ¾-overhead rail-yard hauling dash: a carry-weight speed curve, recharging sprint, three telegraphed train kinds, color-matched freight, a shift clock, and three lives. It borrows Frogger’s lethal-on-any-touch rule but is intentionally original rather than a reskin.
Valence, rebuilt on the Bloons TD 6 model
Section titled “Valence, rebuilt on the Bloons TD 6 model”Valence’s wave composition was a weighted random draw from an unlock pool on a flat metronome, so every post-unlock round was the same soup. Rather than tune from scratch, v2.0.0 inherits a proven balance wholesale.
The economy pays for damage, not kills. Each shell stripped pays 1; overkill past the last shell pays nothing; a bond pool pays nothing while draining and its whole value on the breaking hit. Per-type bounties are gone, so a round’s total income equals its total shells, Bloons’ RBE, which means starting energy, the round-clear bonus, and all seven tower costs transfer directly. There is no per-round scaling: a Dimer in round 3 is identical to one in round 38, and difficulty comes entirely from the round table, whose rounds 1–40 reproduce Bloons’ RBE curve round for round.
The roster gains the Lattice (a thin bond pool over sixteen full atoms, so it opens fast and floods the strippers), and inert becomes a modifier any type can carry rather than a property of three fixed types. Macromass becomes a fission chain rather than a big health bar, shedding daughter isotopes across six decay steps; round 40 is a single Macromass and the campaign’s only boss. The sim now draws no random numbers at all, so scenarios replay exactly.
The debug API’s spawnUnit gained an inert flag routed through the same
construction the wave system uses, so a scenario can pose a shielded Dimer or
Lattice rather than only observe one mid-round.
Meltdown: single-type waves
Section titled “Meltdown: single-type waves”Every wave now fields one intruder type instead of a mix, so each wave presses one specific answer. Motes want sustained volume, Sprints want slowing or a long kill-box, Swarms want splash, Hulks want concentrated heat, and Drifts want anti-air, and milestone waves become pure Core boss waves. The Hundred is exempt, being one continuous onslaught rather than a wave sequence.
Pure waves removed the concentrated overload that mixed waves created, so heat stopped deciding matches: the heat-ignoring twin could trip 80+ times in a wave and still leak nothing. Restoring that required raising the HP slope 0.2 → 0.62 and reshaping the count curves to a low base with a steep slope. The rebalanced state has no-maze flank, no-maze battery, and heat-ignored all losing, while maze-plus-heat wins 20/20.
Foray: large seeds and respawning jelly
Section titled “Foray: large seeds and respawning jelly”The adversarial case remains experimental for this release, but two rule changes reshape it:
- Large seeds. Each half holds two objects worth, and weighing, three
ordinary caches, which drift a tile at a time toward the border, so 30% of a
half’s value walks around on its own. The agent contract gained
large_seedsandcarrying_large; a hauling agent carries nothing incarryingbut three units of load. - Respawning royal jelly, and immunity-settled tagging. Jelly nodes now respawn, and tagging is decided entirely by immunity: neither immune, the soldier tags the raider; exactly one immune, the immune one tags the other; both immune, nothing happens. Previously a soldier in its own half could never be tagged; now an enemy raider running active jelly can kill it.
A new easy asset-generation case, foray-large-seed, produces the sprite the
adversarial case now ships in its sheet, the first time one case’s produced art
becomes another case’s committed input.
Reference views come from the reference implementations
Section titled “Reference views come from the reference implementations”The four non-experimental cases still serving HTML-mockup reference views
migrated each view from a rendered path to a committed media PNG, and the
mockup sources (reference/*.html plus their large hand-maintained theme.css
files) were deleted, a net −5,241 lines. The reference implementation is now the
single source of truth with no second artifact to keep in sync.
Every frame was driven through the case’s own debug API against a real build, never fabricated: Arc Foundry’s board is 26 towers played to wave 9 of 50; Deepcore’s game over resolves through the actual hull-destroyed death path; Carom’s is a real match point out the right goal. Relatedly, proof folders and proof-capture scripts were dropped from reference implementations altogether, since proof captures are evidence models submit and a reference implementation is verified directly.
Console polish
Section titled “Console polish”- The Metrics tab gained a Ratings chart, and chart tooltips became readable (a dark box rather than light-on-white), trigger across the whole bar, and highlight the hovered column.
- Runs can be multi-selected for batch open, kill, or delete, with the whole gutter cell as the click target.
- Detail titles gained a back chevron, account page titles moved above the tabs bar, and the test-case and model detail titles were calmed down.
- Run and review tables show a model’s catalog display name rather than its raw id.
- The new-run form defaults to end-to-end, splits the test-case picker into type and case dropdowns, and requires an explicit model choice.
- The public site hides never-run models from its Models page.
- The Proof tab is hidden when a run has no proof.
- The full-panel loading spinner is framed as an arcade-cabinet screen: a lit, moulded bezel around a vignetted glass panel with the squadron flying inside, CRT scanlines over the lot, and an amber-phosphor “Loading…” caption burned in beneath.
- The public About pages were rewritten.
The public gallery could freeze on an empty definition store
Section titled “The public gallery could freeze on an empty definition store”On the managed-Postgres deployment shape the ingested definition store lives on an ephemeral volume that a pod reschedule empties and an ingest sidecar re-populates, while runs stay durable in Postgres. If the public snapshot was regenerated while that store was momentarily empty, it emitted every published run but no per-case metadata, so the gallery listed a run whose test case 404’d and could not be opened, its card’s name fallen back to the raw slug.
Ingest only queued a snapshot refresh on reference-build or sheet changes, never on a definition re-ingest, so once the store healed nothing re-published a corrected snapshot and the gallery stayed frozen on the empty-store build. An ingest that actually (re)ingests a version now queues a coalesced refresh; a no-op scan queues nothing, so the periodic ingest does not rebuild the gallery each cycle, and a forced re-ingest always refreshes.
Template specs leaked raw Handlebars into the console and gallery
Section titled “Template specs leaked raw Handlebars into the console and gallery”Variant-branching .hbs specs, Carom v2.0.0’s playfield.md.hbs and the others
that moved per-variant differences from separate .md files into shared
templates, leaked raw {{#if (eq variant.slug …)}} into the console’s Inputs
tab and the static gallery. Only the prompt was ever rendered on the display
paths; seeded spec bodies were served and inlined as raw template text, which
stayed hidden until spec bodies started branching.
Specs now get the same per-variant render treatment the prompt already had: a
new GET /test-cases/{slug}/versions/{version}/specs/{variant} returns the
full, seed-ordered, per-variant-rendered set, and each variant’s snapshot
seededInputs is its complete rendered set. The old commonSeededInputs is
retired, since a template spec has no single shared form.
Successful runs failed by their own artifact collection
Section titled “Successful runs failed by their own artifact collection”The Kubernetes collector streamed the entire /work tree out of the run pod
with tar -c, node_modules included, unpacked it on the host, and only then
dropped node_modules, which it had always intended to discard. For Vite builds
that tree is large and full of platform-specific native binaries and
package-manager symlinks, and the host-side unpack choked partway through,
failing the whole run as an infrastructure error even though the run had
succeeded and its produced tree was intact. Three prod runs hit it, each on a
different file. The dependency directories are now excluded at pack time,
sharing one exclusion list with the copy step, so they never enter the archive
at all.
A run killed at its deadline is a timeout, not a harness error
Section titled “A run killed at its deadline is a timeout, not a harness error”When a run hits its maximum runtime the container tears the harness down, which surfaces as a non-zero runner exit, and that was being classified as a harness invocation failure. A run that did hours of useful work before the cap was reported as a harness error. A non-zero exit at or past the deadline now returns a timeout, mirroring the in-process runtime cap, so it publishes with its code and build and the model’s work is kept.
Harness error detail was being swallowed
Section titled “Harness error detail was being swallowed”OpenCode and other structured-error harnesses carry the real text under
error.data.message alongside a machine-readable name, but only
error.message, message, and error were checked, so every such failure
collapsed to the generic “harness reported an error”. The nested shapes are now
dug out, and an error object with no string anywhere is serialized rather than
discarded.
The traceparent never reached the harness
Section titled “The traceparent never reached the harness”It was being set on the host-side docker/podman client, which does not
forward its environment across the daemon, and the Kubernetes exec API carries
no environment at all. It is now set on the container at start. (The
observability documentation had claimed this already worked; that claim was
corrected.) Relatedly, the LGTM NetworkPolicy admitted the services and driver
Jobs but not the sandbox pods, which is where the harness actually runs, so on
an enforcing CNI harness export would have been dropped while driver spans
arrived, presenting as an uninstrumented harness.
Foray playback ran an extra tick
Section titled “Foray playback ran an extra tick”The engine ended execution on a forfeit action but still pushed it to the log,
so playback ran one unexecuted tick and could display a score that didn’t match
the committed result. The forfeit action is no longer logged, leaving the log a
strict record of executed ticks. A tournament summary’s replay_key also now
carries the match id, so replays are viewable after a tournament run.
A particle effect could be authored dense enough to stall the reviewer
Section titled “A particle effect could be authored dense enough to stall the reviewer”A particle system is simulated live by everything that plays it, so the count of
particles alive at once is a cost the reviewer’s browser pays every frame, and
nothing in the authoring flags said so. --rate 20000 --lifetime 1600 reads
like two ordinary numbers and means thirty-two thousand live particles; the
binary’s own preview couldn’t show the difference, because it draws at most
8,000 billboards a frame however many the system holds. Effects came out of runs
looking right and made the run page stutter.
particle-2d and particle-3d now hold a system to 10,000 live particles, and
enforce it where it can still be acted on. Every operation is projected forward
to the peak count the system would settle at (rate x lifetime for a rate
emitter, the count for a burst re-fired each cycle on a looping timeline, and
sub-emitter children projected from the traffic their parent hands them,
generation by generation) and an operation that would push the system over is
rejected rather than recorded, reporting the projection, the emitters spending
it, and the flags to turn down.
The ceiling costs an effect nothing visible: ten thousand particles is already
denser than any preview a model can see, and density comes from particle size,
opacity, and color, which are free. The simulator and the browser runtime
enforce the same ceiling as a backstop, so a system.json recorded before the
budget existed is bounded when it is played too.
Curl-noise turbulence made a particle effect unplayable in the browser
Section titled “Curl-noise turbulence made a particle effect unplayable in the browser”A particle count is not the only thing a live simulation spends. Turbulence cost
about 70 µs per particle per frame in the browser runtime, so a system with
turbulence on stalled the run page at a couple of frames a second while its
preview GIF played perfectly, because the Rust simulator does the identical work
in native 64-bit arithmetic and never noticed. Curl noise is the curl of a
hash-based potential, and every lattice corner it reads is a SplitMix64 hash
carried in BigInt to stay exact: roughly 400 wrapping 64-bit multiplies to
move one particle one frame.
Two things fixed it without changing a single value. The hash lattice depends only on integer coordinates and never changes over time, so the runtime now memoizes it in a fixed-size open-addressed table shared by every particle for the whole play, with each hit verified against its stored key. And the curl of a 3-vector potential reads only two components of each partial derivative, where the straightforward formulation computed all three, so taking just the six that are used drops 18 value-noise samples per particle to 12. Both are arithmetically identical to the naive form, so a seeded browser play still matches the binary’s own simulation. A turbulent system at the 10,000-particle ceiling went from ~670 ms a frame to ~10 ms.
Smaller fixes
Section titled “Smaller fixes”- The reviewer checklist aligns across sub-itemed and whole items, and its category accordion animates.
- Automated validation is hidden on the Verdict tab until a review exists, so it doesn’t bury the form, and the gate now tests the signed-in reviewer’s own review, where it had been revealed to everyone by any one reviewer’s review.
- Errata score exclusions now surface in the review UIs, so a point retired from scoring reads as excluded where it is reviewed.
- Adversarial match replays render only on the Results tab; a merge had briefly restored a duplicate Proof-tab view of them.
- A pre-publish build link resolves on a cold deep-link to Play.
- Post-auth visitors land on the home page.
- The coverage plan and group pages read as detail pages.
- The browser-driver assertion kit is baked into the service images and
allowlisted in
.dockerignore.
Development
Section titled “Development”Pre-commit hooks
Section titled “Pre-commit hooks”Fast gates now run on every commit, with the slow ones left to CI. Setup is one
idempotent command, run automatically on devcontainer create. The set covers
file hygiene (JSON/YAML/TOML/XML validity, merge conflicts, large files, case
and shebang checks), shellcheck, cargo fmt --check, clippy (mirroring CI’s
invocation exactly), cspell, and markdownlint, the last two scoped to authored
test-case and game-jam prose, which is how over-length spec lines had been
slipping in.
Large-file exemptions were added for validation-baseline/ and
reference/screenshots/, both of which are deliberately tracked media that
exceed the 500 KB default.
cargo-nextest, and an exactly-pinned toolchain
Section titled “cargo-nextest, and an exactly-pinned toolchain”The Rust suite moved to cargo-nextest (doctests stay on cargo test --doc),
with a pinned version installed at devcontainer build time, a cross-platform CI
installer for the Linux/Windows/macOS binary smoke legs, and a config setting
one retry, no fail-fast, and a per-test hard timeout.
Separately, nothing that builds this repo should float. rust-toolchain.toml
tracked stable, and when stable moved 1.96 → 1.97 the clippy::question_mark
lint widened and a -D warnings gate failed on untouched code. The same drift
was latent in the container builders and CI’s Node. Rust is now pinned to 1.97.1
everywhere it is named, including the builder stage of all 22 container images,
where base-wasm had been on 1.96.0 and silently disagreeing with CI, and Node to
the exact patch each config already resolved to. Deliberately left floating:
GitHub Action major tags, and the internal :latest base-image defaults that
the build script always overrides with an explicit tag.