Terminology
Artifact service
Section titled “Artifact service”The artifact service serves the produced run trees off a persistent volume: a run’s playable build and its proof and asset media. The driver uploads each run’s tree to it, and a console reads it from there to play and review the run. It is a data-plane peer of the backend, so artifact bytes never transit the control plane.
Attachment pivot
Section titled “Attachment pivot”A part’s attachment pivot is the point, in its parent’s local voxel
coordinates, at which it hangs off its parent in a rig. Posing the
parent moves the child about this point, so a turret’s pivot is where it sits
on the chassis. For the root part the pivot is its origin in world space.
Auth service
Section titled “Auth service”The auth service is the standalone private service that holds The Test Cabinet’s user accounts. It handles open self-registration and password login, and mints the opaque bearer tokens the backend verifies on mutating run requests. It keeps its own database, so credential storage stays out of the backend.
Backend
Section titled “Backend”The backend is The Test Cabinet’s central private service. It is the canonical source of test case definitions for runners and the system of record for run results through storage, review, and publication. It verifies the auth service’s bearer tokens rather than storing credentials itself.
Catalog
Section titled “Catalog”The catalog is The Test Cabinet’s full set of test cases.
Climber
Section titled “Climber”A climber is one combination enrolled on a ladder, the thing that actually does the climbing. Each climber’s progress is tracked separately, so a model added to a standing ladder starts at rung one while the others carry on from wherever they had reached. A climber is climbing, awaiting review, walled, held (stopped by hand), or topped out. Climber and combination name the same thing: the first is the role it plays on a ladder, the second is what it is.
Combination
Section titled “Combination”A combination is the thing a run is executed by, as opposed to the test case it is executed on. It takes one of two shapes: a harness and model pair, plus a provider for a provider-routed harness; or a gg configuration and a model for each launch slot that configuration asks for. It is the unit a coverage plan crosses with its cases to form a cell, the unit a ladder enrolls as a climber, and the unit a reusable coverage group holds.
Coverage
Section titled “Coverage”Coverage carries three meanings in The Test Cabinet, and they are not the same thing:
- The measurement: how much of a declared matrix actually has runs. A coverage plan declares test cases pinned to a version, variant and engine, crossed with combinations and a target run count per cell, and its coverage is how many of those cells have met their target. This is the older and narrower sense. See Coverage plans.
- The feature area: the reviewer scheduling surface as a whole, which is plans, ladders, the reusable groups both draw their members from, the account-wide review buffer, and the pause and halt controls. This is the sense in which the console has a Coverage section and the backend a coverage API, and it takes in ladders, which are not plans and aim at no matrix at all.
- Code coverage: how much of a produced implementation’s own
src/the tests the model wrote reached when they ran. It is measured by istanbul while the case’s[toolchain]test command runs, and recorded on the run record’stoolchain.test.coverageblock.
So “a ladder is part of coverage” and “a ladder has no coverage target” are both true, in the first two senses. When it matters, say coverage plan for the first, the coverage surface for the second, and code coverage for the third.
Code coverage shares only the word with the other two. It is a property of one run’s produced tree, it measures the model’s own code with the tests the model wrote, and the test case’s validators contribute nothing to it because their own suite has coverage disabled. The static analyzer’s walk diagnostics on a run’s Code tab are headed Analysis notes rather than Coverage, since they describe the analysis and measure nothing about the code.
Dispatcher
Section titled “Dispatcher”The dispatcher is a thin, stateless
controller that drains the backend’s run queue. It claims each
queued run and creates one Kubernetes Job running a driver to
execute it. The backend’s job table is the source of truth, so concurrency
scales with the cluster.
Domain
Section titled “Domain”A scoring domain is a facet of a test case the reviewer rates independently, such as a game’s single-player and versus modes. A case declares common domains that every variant is rated on, and a variant may add its own, so the effective set for a run is the common domains plus its variant’s. Each domain carries a functional rating, and the run’s functional rating is the worst across that effective set. A review item names the domains its failure lowers on a validator-rated version, and on a legacy version may roll up to a domain or stay general when it applies to every mode.
Driver
Section titled “Driver”The driver is the per-run executor: a one-shot
process, created by the dispatcher as a Kubernetes Job, that
runs exactly one test case. It resolves the definition from the
backend, drives the run through the
core in an untrusted sandbox pod, streams live
progress back to the backend, uploads the produced tree to the artifact
service, and exits.
Engine
Section titled “Engine”In the context of The Test Cabinet, “engine” refers to two elements:
- The runtime a produced game is built on, selected per run and recorded on it
- The WebAssembly module a model submits for a performance test case
The first is provided to a run and the second is the deliverable of one. A performance case’s engine is the artifact under test, so it is never selected and a performance case supports no engine in the first sense.
Failure cap
Section titled “Failure cap”A failure cap is a review item’s declared failure_cap: the highest functional
rating the item’s domains may reach while the item’s validator
fails, one of broken, scuffed, passable, or great. A domain’s functional
rating is the lowest cap among its failing items, so the caps let the validators
decide the rating without a reviewer. Every item of a
validator-rated case version declares one.
Harness
Section titled “Harness”In the context of The Test Cabinet, “harness” refers to two elements:
- The Test Cabinet itself
- Agentic harnesses used to drive models
The Test Cabinet runs other harnesses. It hits no LLM API and implements no agentic loop; that responsibility lies entirely with the agentic harnesses it drives.
A joint is one named degree of freedom on a rig part: a
rotation in radians about an axis through a pivot, or a translation in voxel
units along an axis, bounded by a min/max/rest range. A caller-driven
joint takes its value from a consuming game at runtime, for example
turret_yaw. An auto joint is driven only by the model’s animation tracks and
holds at its rest value otherwise. Joints are model-invented: a case declares
the animations its rig must carry, and the model devises whatever joints carry
them.
Ladder
Section titled “Ladder”A ladder is an ordered series of test cases that climbers ascend one rung at a time, stopping at the first rung they cannot clear, so the rung a model stops at is the result. Where a coverage plan asks whether a cell has been run yet and treats its cells as an unordered set, a ladder asks how far a model gets and treats its steps as a sequence in which each is harder than the last. Whether a climber advances is decided by the ladder’s gate, a single rule parameterised by a rating floor and a threshold. See Ladders.
Leaderboard
Section titled “Leaderboard”Each test case has a per-variant leaderboard over its scored runs. Each row is one harness and model pairing, ranked by average score, then by rating, then by recency. gg runs are excluded, because a gg run’s agents may span several models.
Models are the large language models that determine the actions an agentic harness takes.
Orchestrator
Section titled “Orchestrator”An orchestrator decides how a run’s harness sessions are conducted:
how many sessions to drive, what each is told, and when the work is done. The
harness still owns each individual session. An orchestrator is selected per run,
is harness-agnostic, and defaults to one-shot, a single session. See
Orchestrators.
A part is one named voxel component of a rig, for example a
tank’s chassis, turret, or barrel. Parts form a parent/child hierarchy,
each attached to its parent at an attachment pivot, and
each is sculpted independently with voxel-anim --part <name>. Posing a parent
moves its children with it. Parts are model-invented: the model creates each
part at run time with define-part.
Publishing
Section titled “Publishing”Publishing releases a reviewed run and makes it public. It releases
the run’s source to a public GitHub repo and its playable build to Cloudflare
Pages, and adds the run to the public snapshot and gallery. It is
the second of two steps, review then publish, and the release runs
asynchronously in a per-publish tcab-publisher Job. The CLI’s tcab publish
performs both steps at once for a solo operator.
A completed validator-rated run is publishable from the moment it completes, on its functional rating and score. A completed legacy run needs at least one review before it can be published, with two waivers. A run in a publishable failure state (catastrophic, timed out, harness error) has no checklist to complete. An auto-validated comparison run carries automated verdicts in place of a review.
A produced run is stored privately on the backend as soon as it finishes, and its build is playable for review off the artifact service. Publishing is what first releases it publicly. See Results.
Rating
Section titled “Rating”A run is rated on two channels, each a five-tier scale. The functional channel is rated per domain; the aesthetic channel is rated once for the whole run.
The functional rating says how faithfully a domain implements the spec:
flawless, great, passable, scuffed, or broken, and a run’s is the
worst across its domains. On a validator-rated run the
validators decide each domain’s rating from its failing items’ failure
caps; a review that overrides verdicts recomputes
those ratings for itself, and a run with reviews takes the worst effective
rating across them. On a legacy run each review carries one per domain and the
run’s is the worst across every review.
The aesthetic rating says how the build looks, sounds, and feels to play:
legendary, amazing, good, okay, or slop. amazing is the normal
maximum and legendary is reserved for an exceptionally beautiful build. Only
a validator-rated run has one; each review carries a single run-wide tier and
the run’s is the worst across every review. See
Evaluation.
Reporters
Section titled “Reporters”Reporters are The Test Cabinet components capable of reporting run results. Only GUI reporters allow users to interact with test case implementations. The web console is both a reporter and a launcher of runs.
Review
Section titled “Review”A review is a person’s assessment of a run after playing its build, providing the feedback automation cannot. Reviews are subjective, since games do not map cleanly to a rigid grading scale.
A review carries a prose writeup, a rating, and the identity of the account that wrote it. On a validator-rated run the rating is a single run-wide aesthetic tier, and the review may also override individual reviewer-checklist verdicts the validators decided. On a legacy run the rating is the functional one per domain, the review also carries a verdict on each reviewer-checklist item the case declares, and those verdicts and item weights produce the review’s numeric score, averaged across the run’s reviews. A run may carry one review per account, typically from people other than the operator who produced it.
Review buffer
Section titled “Review buffer”The review buffer is how many runs a coverage plan or ladder may leave waiting on you before it stops enqueueing: everything in flight, plus everything finished that you have not reviewed. Its size is a property of the reviewer, an account-wide setting overridable per plan or ladder, rather than of any one plan, because it describes how much work you want to come back to. It exists so the first few reviews can still steer a plan, where firing an entire matrix at once spends the whole budget before anyone has looked at a single run. Refilling it is called a top-up. The setting is either a bound on outstanding runs or no limit, which makes a top-up enqueue every missing run at once.
Reviewer checklist
Section titled “Reviewer checklist”A test case may declare a reviewer checklist: a list of major, observable requirements that every reviewer verifies by playing the build. Each item carries a point weight. An item may break into name-only sub-items, each judged on its own, with the item’s weight split evenly across them. A game-jam case grades its items on a five-level scale instead of pass/fail.
On a validator-rated run the validators decide every verdict, and a review may override any point’s verdict with a binary pass or fail and an optional note. Points a review leaves untouched keep the validators’ verdicts, so a review’s effective checklist is the validators’ verdicts overlaid with its overrides. Overriding is the exception, for a validator whose precondition could not be met or a build that does the right thing despite broken instrumentation. On a legacy run the web console presents the checklist as a guided review with a completeness gate: every item and sub-item needs a verdict before a review can be saved or the run published. The checklist is reporter-side and is never seeded, so it stays out of the model’s input.
A rig is the posable structure of a voxel-animation model: its named
parts in a hierarchy, the named joints a consuming game
drives, and the model-authored animations those joints carry. The rig is
model-invented: a case’s [model] table declares the required animations by
name, and the model devises whatever parts and joints carry them. The produced
rig.json carries everything the model built, and the
voxel-runtime poses it for both the
review viewer and real games.
Run records
Section titled “Run records”A run record is produced each time a test case runs to completion. It records all information from the run, such as its run time, version information, and token and cost data.
A rung is one step of a ladder: exactly one test case, pinned to an exact version and variant, with an optional override of how many runs it takes to judge. The rungs’ order is the climb. Each rung carries a stable opaque id rather than being identified by its position, because rungs get reordered and re-pinned and every recorded verdict references that id. A positional identifier would silently reattribute a climber’s history to a different case.
Runners
Section titled “Runners”A runner is the component that actually executes a test case. There is exactly one: the per-run driver a dispatcher creates for each run, built on the core. The CLI and the web console enqueue a run at the backend and watch it.
A score is earned points over the points available. Each reviewer-checklist item is worth a weight, a pass earns that weight, and a fail earns none, so the total is the sum of every declared item’s weight. An item with sub-items earns the fraction of its weight whose sub-items passed, so an earned score can be fractional. A validator-rated run scores from its validators’ verdicts the moment it completes; each of its reviews scores from its effective checklist, and a run with reviews scores the average of them. A legacy run scores per review and a run carrying several reviews scores the average. The run’s score is shown alongside its functional rating, and is what the per-case leaderboard ranks on.
A comparison scores a run from its automated validators alone, restricting both the earned points and the available points to the machine-checkable ones.
Snapshot
Section titled “Snapshot”A snapshot is the public export the backend produces from its published results. The static public site is built from this snapshot, so the gallery keeps no live dependency on the backend.
Test case
Section titled “Test case”Test cases provide the scenarios used for testing. Each test case represents an isolated task that a harness and model must perform.
Topped out
Section titled “Topped out”A climber has topped out when it has cleared every rung of its ladder: there is nothing left to climb, and the ladder has no further question to ask of that combination. It is the only one of the five climber states that is nobody’s move, the opposite end of the ladder from a wall, and distinct from held, which is a stop the reviewer chose.
User account
Section titled “User account”A user account is a real, registered identity in the auth service, created by open self-registration with a username, password, and display name. Logging in mints a bearer token that authenticates the mutating run actions, review and publishing, so every review a run carries is attributed to the account that wrote it. Accounts are an identity layer on top of the private network. Reading the gallery or the backend needs no account.
Validation
Section titled “Validation”Automated validation checks everything a machine can check: that an implementation builds and loads, how closely a view matches its reference image, and, through the instrumentation a case requires the build to expose, whether the spelled-out mechanics work when the build is driven into the states that exercise them. A build that fails the mandated debug-API contract fails automatically. Each validator’s pass or fail is visible per checklist point on a run’s Verdict tab, to every visitor including the public gallery. A game’s feel and quality are left to a human review.
Validator-rated
Section titled “Validator-rated”A case version is validator-rated when it is on the engine-supported manifest
format, the [workspaces] / engines / [[engine]] spelling of its starter
project, and is not a game jam. A run is validator-rated when its case version
is. Its validators decide its functional rating and score
at completion, so it is publishable with no review, and those figures stand
while the run has none. Its reviewers supply a single run-wide aesthetic rating
and may override individual checklist verdicts, and a run with reviews averages
their effective scores and takes their worst effective functional rating. Every
other version is a legacy version, whose reviewers supply the functional rating
and checklist verdicts.
Variant
Section titled “Variant”Test cases may define multiple variants, which identify modifications to make to the specifications provided as input for the test. These variants may change game mechanics, add or remove content, and may noticeably affect the difficulty of a test case.
A voxel is a single opaque-#rrggbb cell in a 3D grid, the 3D counterpart of a
pixel. The two 3D asset-generation kinds
sculpt into a fixed voxel volume, which always starts empty: voxel-model
produces a static model, and voxel-animation produces a rig.
A voxel run’s authoritative output is the data its voxel binary emits: the
meshed geometry as a per-part .glb, and a rendered preview. The validator
parses and validates that emitted data rather than regenerating it, and the
frontend renders an interactive 3D model with three.js.
A wall is the rung a climber failed and therefore stopped at, so “walled at rung four” is a ladder’s headline result for one model. It is a verdict the gate computed from your reviews of that rung’s runs, so it is an opinion rather than a fact about the model: a reviewer can promote a climber past a wall by hand, and the automatic verdict is kept underneath rather than overwritten, so clearing the override restores exactly what the gate said. A failed or canceled job is never a wall, because infrastructure failures are retried and only completed runs are evidence.
Web console
Section titled “Web console”The web console is The Test Cabinet’s
runner/reporter GUI, delivered as a static web app that runs in a plain browser.
It mounts the routed application from the UI
library. It enqueues runs at the backend,
which a dispatcher drains into per-run driver Jobs.
It is an operator tool served on the private network, not a public site.