Overview
The Test Cabinet’s CLI is the tcab binary, an enqueue-and-watch client of the
backend. It exposes the backend’s run-queue
control plane on the command line so test case runs can be scripted and
benchmark sweeps run in batch.
Every run executes remotely. tcab run enqueues a run on the backend’s job
queue, a dispatcher claims it, and a per-run
driver pod executes it and streams progress back
through the backend. The CLI therefore needs no container runtime of its own,
only a reachable backend (TCAB_BACKEND_URL) and a logged-in account. See
Execution.
Commands
Section titled “Commands”Running
Section titled “Running”-
runenqueues a run for a test case version, variant, harness, model, orchestrator and engine, streams the run’s live event stream until it finishes, then reads the produced run record back.--max-runtimeoverrides the case’s per-run cap,--auth-modeselects the harness authentication mode, and--retry-countsets how many times the backend retries the run after a terminal infrastructure error or a catastrophic build, with0disabling retries.--out-diralso writes the fetched record as<record-id>.jsonthere. RequiresTCAB_BACKEND_URLand a logged-in account.--harnessselects one of the third-party CLI harnesses. A gg run carries a capability set and is launched from a console.--engineselects the runtime the game is built on, defaulting tonone, andseedandvalidatetake the same flag. -
seedruns only the seeding step for a chosen variant and leaves the result on disk, so the exact inputs a harness would receive can be inspected without launching a container. -
promptrenders and prints the prompt a run would hand the harness for a given variant, without seeding or launching anything. -
validateruns validation over a produced implementation and reports whether the tree satisfied everything the case declares. It leaves the directory as it found it, apart from the install and build it runs there and the media it synthesizes under.vendor/, so it is safe to point at a case’s committed reference implementation. See Exit codes for what makes it fail. -
harnesseslists the supported agent harnesses and whether each one’s resolved authentication mode has the credentials it needs. -
orchestratorslists the built-in orchestrators and what each one does. -
engineslists the built-in engines and what each one provides. -
test-case-groupslists the test-case groups and each one’s member cases.
These listings accept --json.
Accounts
Section titled “Accounts”registercreates an account on the auth service (--username,--display-name,--passwordorTCAB_PASSWORD), then logs in and stores the resulting token.loginlogs in to an existing account and stores the bearer token.logoutdiscards the stored token.
Reviewing and publishing
Section titled “Reviewing and publishing”review <run-id> [--writeup writeup.md]submits a review for a produced run, attributed to the logged-in account, from a writeup the reviewer authored locally. It defaults towriteup.mdin the working directory. For a validator-rated run the writeup carries a single run-wideaestheticrating and may carryreview.<id>lines overriding individual validator verdicts; for a legacy run it carriesrating.<domain>ratings with checklist verdicts. A run may carry several reviews, one per account.publish <run-id>...does self-review and publish in one step: it submits the operator’s own review from a<run-id>.mdwriteup in the working directory, enqueues the publish, and streams the release’s progress until it finishes. Publishing a legacy run requires at least one review, which the self-review satisfies. A validator-rated run with no writeup is published without a self-review, and the command says so. Whether a writeup-less run is validator-rated is decided by asking the backend for the run’s case version, so a run that lacks a writeup while the backend cannot be reached is refused with that reason. Every run’s writeup is gated before anything is submitted, so a batch is never left half-published.--dry-runprints the plan instead. Where different people review, usereviewand have an operator publish.
Reference implementations and baselines
Section titled “Reference implementations and baselines”publish-reference --env <prod|staging> <slug> [<version>] [--variant <slug>] [--engine <slug>] [--all-variants] deploys a case’s reference
implementations and
records where they landed. --env is required, so a publish always names the
environment it targets. It selects the Cloudflare Pages project, prod’s
test-cabinet-references or staging’s test-cabinet-references-staging.
--dry-run prints the plan without building, deploying or writing anything.
The unit of work is the variant-on-an-engine pair, resolved from each targeted
variant’s reference_implementation key. For
each pair the command runs the case’s [build] install then build in that
pair’s reference directory, scrubs the output with the same secret-redaction
pass the publisher uses, deploys
the static build to that Pages project under a per-variant, per-engine branch
alias, reads the served URL back out of wrangler’s output, and writes it into
the committed test-cases/reference-builds.lock.json under the --env key. The
URL is parsed rather than constructed because Cloudflare truncates long
subdomains.
The command contacts no backend, so it requires only wrangler. The private
backends ingest the lockfile from their own checkout on the next
scripts/reingest-cluster.sh, which upserts the case_reference_build table the
version response and public snapshot read. The command also refreshes each
variant’s committed baseline validation media from the build it deploys.
--skip-baselines deploys without re-capturing when that media is current.
An asset-generation case declares no
[build] table and produces no site, so the same command takes a different path
for it: it seeds a scratch workspace from the manifest, runs the variant’s
reference-impl/<variant>/draw.sh with the case’s drawing binary on PATH, and
uploads the frames and action logs to the public snapshot bucket under
media/references/<slug>/<version>/<variant>/frames/. That path needs the
target environment’s TCAB_R2_* credentials rather than wrangler. It writes
no lockfile because the keys are constructible, so the backend discovers what
exists by listing that prefix at ingest.
capture-baselines <slug> [<version>] [--variant <slug>] [--all-variants] [--engine <slug>] [--dry-run] regenerates a case version’s committed baseline
validation media. A variant has one
reference implementation per engine, so the unit it
works in is the variant/engine pair. For each targeted pair it runs the case’s
[build] install then build in that reference-implementation directory,
produces every scripted review
item’s declared outputs
from it, and writes them under the version’s
validation-baseline/<engine>/<variant>/ in the cold-storage submodule,
together with the shared image
store the recordings among
them draw from. That directory is regenerated wholesale, so a renamed or removed
output never lingers. The media is the expected-behavior half of the comparison
a reviewer makes and is a fixed property of the case version.
The cold-storage root defaults to cold-storage/ beside the catalog’s
test-cases/, and TCAB_COLD_STORAGE_DIR overrides it (see where baselines
live). A capture into the
default root refuses to run until the submodule is checked out.
How the outputs are produced follows the case, exactly as it does per run: a case shipping a validator project for the engine has its baseline recorded by running those suites against the reference implementation, and a case shipping none has its reference build served and driven in a browser. Both halves of that comparison therefore come from the same scenario driven the same way.
A reference implementation is the case’s own answer, so every unit is expected
to run clean against it. A unit that does not fails its target. The sweep still
runs every remaining target first, so one pass reports every fault.
publish-reference shares that rule through the same capture, and a target
whose baseline it could not produce is never deployed.
capture-baselines deploys nothing and writes no lockfile, so it takes no
--env and requires only the case’s toolchain, plus a browser for a case
decided by browser scripts.
Analysis
Section titled “Analysis”analyze runs the static code analyzer over a
directory and prints what it found:
analyze <dir> [--seed-commit <sha>] [--tree-basis <pre-validation|post-validation>] [--top <n>] [--json]It is the same analysis a run records about its produced tree, pointed at any tree on disk, so it needs no run, container, backend or credentials, and it executes nothing in the tree it reads.
--seed-commit makes the authored set exact when the directory is a seeded run
workspace. Left off, every file in the tree is treated as authored, which is the
right answer for an ordinary source tree. --json prints the full analysis
document for piping onward.
Exit codes
Section titled “Exit codes”tcab exits 0 only when the thing it was asked to do succeeded, so a script
or a CI step reads the status rather than the log. Anything that stopped a
command from doing its job, such as an unresolvable case, an unreachable
backend, a missing browser or a rejected login, exits non-zero with the reason
on standard error.
Two commands additionally carry a verdict: they ran to completion and the answer they arrived at is itself a pass or a fail. Both report every fault they found.
validate exits non-zero when the tree it was pointed at failed the case. Each
of these is a fault the tree earned:
- The implementation did not load.
- A required install or build step failed or was never reached, for a case whose validation runs them.
- A declared check could not be reached. The similarity a reached check records is a signal rather than a threshold, so it never decides the exit code.
- A declared proof-of-implementation artifact is missing.
- A gating validator decided a verdict against the build.
- A gating validator did not run against the build, whether it failed the
debug-API contract or was recorded
inconclusive. Scoring skips an
inconclusive unit, but
validateasks whether the tree satisfied everything the case declares, and a unit that decided nothing has not. The final line separates the two: contract failures by verdict id, and inconclusive units grouped by kind and reason, with the ids listed when there are few and one shared reason quoted once when every unit failed for it. A point an erratum excludes from scoring costs nothing. - An adversarial submission forfeited its match, which is a failure to present a playable controller. A loss or a draw is a result rather than a fault.
capture-baselines exits non-zero when any targeted reference build failed to
build or left a unit that did not run clean. The sweep finishes every target
first, so one pass reports every fault. publish-reference decides a baseline
capture by the same rule and skips deploying the target whose media it could not
produce.
Authentication
Section titled “Authentication”The CLI deals with several independent kinds of credential and keeps them apart.
- Harness API keys are supplied to the run’s container as secrets so the agent harness can reach its model provider. See Authentication.
- Reading definitions and runs from the backend is handled at the network layer. The CLI must be on the backend’s private network and presents no token to read.
- Account credentials authenticate the mutating backend calls (launching a run,
reviewing, publishing) and the launch gate.
tcab loginortcab registersigns in to the auth service (TCAB_AUTH_URL) and stores the resulting bearer token at~/.config/tcab/credentials.json, overridable withTCAB_CONFIG_DIR. The CLI sends it on every launch, review and publish so the account is recorded.TCAB_TOKENoverrides the stored token for non-interactive use, and a password may be supplied with--passwordorTCAB_PASSWORD. - Release credentials, the repository-host and Cloudflare tokens used to
release a run’s code and playable build,
live with the backend’s publisher Job rather than with
tcab.