Skip to content

Building

This page is the authoritative source for the repository layout and for the build, format, lint, and test commands of both workspaces. The Test Cabinet is an independent, open-source benchmark and depends only on public crates.io and npm packages.

The repository is rendered from the k8s standard workspace template, which supplies the dev container, the gates and their runner, the pipeline’s gate jobs and its staging publish and deploy. .copier-answers.yml records the template’s source, its version and this project’s answers; see The workspace template.

Setting up a machine that executes test cases (container runtime, run-container image, credentials) is covered by First Time Setup. Running the services on your own machine is covered by Running.

The repository lives on Azure Repos and is mirrored to GitHub, and a clone from either host works the same way. Each submodule is a separate repository, such as cold-storage.

A plain clone downloads the superproject alone and leaves each submodule directory empty. It builds and passes every gate, because nothing in the build, the tests, or CI reads a submodule’s contents. Fetch a submodule into it later when a task needs one:

Terminal window
git clone <superproject-url>
git submodule update --init --depth 1 cold-storage

A recursive clone also downloads every submodule at the commit the superproject pins, which for cold-storage is about 2 GB of media. --shallow-submodules fetches only each pinned commit rather than the submodule’s history:

Terminal window
git clone --recurse-submodules --shallow-submodules <superproject-url>

.gitmodules names every submodule by a URL relative to the superproject, such as ../cold-storage. Git resolves it against the remote the superproject was cloned from. On Azure that is the sibling repository in the genyume/the-test-cabinet project, and on GitHub it is TheClockwyrks/<name>. A submodule repository therefore has the same name on both hosts, while the superproject’s name may differ.

Each submodule repository carries its own Azure pipeline, .azure-pipelines/mirror.yml, whose one job force-pushes master to the GitHub repository of the same name. That job is the only writer of the mirror.

A submodule commit may be pinned once it is on that submodule’s master. Push the submodule commit to master first, then commit the moved pointer here. scripts/ci/submodule-pins.sh checks this, and the pipeline’s submodule_pins job runs it on every push: it fails when a pinned commit is not an ancestor of the submodule’s master on the host the superproject was cloned from. A superproject commit that reaches the GitHub mirror therefore names only submodule commits the mirror already holds. The check fetches commits without trees or blobs, so it finishes in seconds, and scripts/ci/submodule-pins.test.sh is its offline table test.

No gate reads a submodule, so the template’s gate jobs check out none: the pipeline_repositories answer is empty, which also keeps the baseline media out of every devcontainer start. The one job that reads cold-storage is submodule_pins. It checks this repository out without submodules and adds an inline checkout of cold-storage, which is what puts that repository in the scope of the job’s access token, then runs scripts/ci/submodule-pins.sh with the token in SYSTEM_ACCESSTOKEN.

The repository is both a Cargo workspace (Rust) and an npm workspace (TypeScript). Around them:

  • .copier-answers.yml: the workspace template’s source, the version the repository was last rendered against, and every answer.
  • .azure/: the pipeline templates. .azure/publish-image.yml is the template’s; .azure/project/*.yml is the project’s own pipeline work, which the template rendered once and never renders over; .azure/tcab/ holds the project’s job and step templates, shared by azure-pipelines.yml and azure-pipelines-release.yml.
  • ci/: the gates, one file per gate under ci/gates/, the ci uv project that runs them, and under ci/images/ the CI images the gate jobs run in.
  • .devcontainer/: the dev container, entirely the template’s.
  • .claude/: the harness settings, the template’s hooks and skills, and the project skills.
  • scripts/: the template’s hook installer and devcontainer check, the project’s scripts, and under scripts/ci/ the pipeline’s scripts, the template’s and the project’s side by side.
  • tasks/: the issue board.
  • crates/core: test-cabinet-core (lib test_cabinet_core). The headless core that owns all orchestration: resolving a test case version, seeding a run’s repository, executing the run in a container, invoking the agent harness, collecting metrics, running validation, and writing the run record.
  • crates/cli: test-cabinet-cli (binary tcab). The command line interface over the core.
  • crates/gg: test-cabinet-gg (binary gg). The in-container coding harness this repository ships itself.
  • crates/dispatcher: test-cabinet-dispatcher (binary tcab-dispatcher). The dispatcher, which claims queued runs from the backend and creates a per-run Kubernetes Job.
  • crates/driver: test-cabinet-driver (binary tcab-driver). The driver, the per-run executor that runs one run server-side.
  • crates/artifacts: test-cabinet-artifacts (binary tcab-artifacts). The artifact service, which retains and serves the trees a run produces.
  • crates/arena: test-cabinet-arena (binary tcab-arena). The arena service, which runs adversarial matches and tournaments off the backend.
  • crates/backend: test-cabinet-backend (binary tcab-backend). The backend, the private definition/run store and API. Its system of record is a SeaORM database: embedded SQLite by default, or PostgreSQL when TCAB_BACKEND_DATABASE_URL points at one.
  • crates/auth-service: test-cabinet-auth-service (binary tcab-auth-service). The auth service, which holds accounts and mints the bearer tokens the backend verifies.
  • crates/telemetry: test-cabinet-telemetry. The shared OpenTelemetry wiring every long-lived binary initializes at startup.

The remaining members are the asset-generation tools (draw, voxel, mc, sn, dc, paint, particle-*, sfx-*, music and the libraries they share), the adversarial and performance engines (foray-*, lattice-*), the store’s entities/migration crates, and the ten gg-sandbox-artifacts arm crates. Shared dependency versions are declared once in the root Cargo.toml under [workspace.dependencies] and inherited with { workspace = true }.

  • packages/run-record: @clockwyrks/run-record. Shared TypeScript types and JSON Schema for the run record, the central data contract.
  • packages/asset-contract: @clockwyrks/asset-contract. The rig and F-curve shapes a produced model is described by, generated in the same pass from the same Rust types. Its own package because it is the only slice of the contract that may be seeded: the voxel and particle runtimes depend on it and are vendored into a model’s workspace, so whatever they depend on travels with them. run-record re-exports these types, so a console keeps importing them from there. The seeded-contract gate (ci/gates/seeded-contract.py, which runs scripts/ci/seeded-contract-check.sh) keeps the evaluation half out of a run.
  • packages/run-stats: @clockwyrks/run-stats. The framework-free rules for scoring a reviewed run, each mirroring a counterpart in crates/core/src/review.rs, plus the set-level rollup that keeps a figure frozen at one moment comparable with the same figure recomputed later. It has no runtime dependencies and imports only types from run-record, so it runs in a bundle, a build script, or a worker alike. packages/ui’s ratings module re-exports the scoring half alongside its display metadata.
  • packages/browser-driver: @clockwyrks/browser-driver. The Playwright driver script the validator shells out to, used to render reference mockups and to drive and screenshot a produced implementation.
  • packages/case-harness: @clockwyrks/case-harness. The shared harness every engineless (none) validator project is built on: the browser lifecycle, the driven-frame loop, the injected draw-command recorder and audio probe, and the readings a check makes over them. It ships TypeScript source rather than a built dist/, because the validator stages it beside a case’s own validator files, where nothing compiles. Its own suite lives in test/ as *.spec.ts files, held outside the files the package publishes, so the harness’s tests stay with the harness while a case’s validation/ tree carries only the suites its review items name.
  • packages/ui: @clockwyrks/ui. The shared UI library hosting the routed gallery application and the presentational primitives; the site and web console are thin hosts over it.
  • packages/voxel-runtime and packages/particle-runtime. The runtimes that pose and render a produced voxel rig and simulate a produced particle system.
  • packages/simple-2d: @clockwyrks/simple-2d. The Simple 2D engine, providing the frame loop, input actions, audio, assets, and diagnostics a produced game builds on, plus the host interface a driver binds to. It is staged into the host package store and vendored into the run repository at seed time, and its package version is the engine version recorded on the run.
  • packages/structured-2d: @clockwyrks/structured-2d. The Structured 2D engine, providing a gameplay framework of worlds, levels, game modes, actors, and controllers, with engine-owned rendering and collision, around a 2D game written in TypeScript. It is staged and vendored the same way as packages/simple-2d, and its package version is the engine version recorded on the run.
  • packages/gg-sandbox: the TypeScript and JavaScript arm of gg’s responses-as-code sandbox.
  • apps/site: @clockwyrks/site. The static gallery site that displays published run records.
  • apps/web: @clockwyrks/web. The browser web console that enqueues runs at the backend.
  • apps/docs: @the-test-cabinet/docs. This Astro Starlight documentation site. Its package.json is the template’s, which pins the site’s dependencies; its astro.config.mjs and pages are the project’s.

cold-storage/ at the root is a git submodule holding the captured baseline validation media, the bulk of the repository’s bytes. Its tree mirrors this one’s, so a test-case version’s baselines live at cold-storage/test-cases/<type>/<difficulty>/<slug>/<version>/validation-baseline/<engine>/<variant>/ (see where baselines live). Fetch it only to capture or review baselines, or to ingest a catalog that serves them.

A reference implementation under a case version’s references/ is its own npm project rather than a member of the workspace, so it is installed from its own directory. An engine-backed one resolves its engine from packages/simple-2d or packages/structured-2d through a relative file: dependency that npm installs as a symlink, so the reference builds and tests against the engine’s current source.

That symlink points at the package directory, and what the reference imports is the package’s dist/, so the root workspace must be installed and built first:

Terminal window
npm ci && npm run build:packages # at the repository root
npm ci # in the reference's own directory

A reference installed over an unbuilt engine installs cleanly and fails at npm run typecheck instead, reporting TS2307: Cannot find module '@clockwyrks/<slug>' ahead of the property errors that the failed resolution produces.

The template renders most of what surrounds the code: .devcontainer/, ci/, azure-pipelines.yml, azure-pipelines-ci-images.yml, .pre-commit-config.yaml, the lint and format configurations, Makefile, rust-toolchain.toml, .config/nextest.toml, .claude/settings.json with its hooks and four skills, and a few scripts. Those files are taken exactly as rendered and are never edited here: a change one of them needs is made in the template under a new version and brought in with copier update. The answer seed_files: project makes the files whose content is this project’s own (this site’s pages and configuration, README.md, CLAUDE.md, the manifests and lockfiles, the web app’s sources, crates/backend and the deployment seeds) seeds: written once, never rendered over.

What the project needs beyond the template has these homes:

  • an answer in .copier-answers.yml: the extra gates, git and lint ignores, the packages the Rust CI image carries, the notes the drain skill carries;
  • a project gate, ci/gates/<id>.py, registered through extra_gates;
  • the project’s pipeline work in .azure/project/*.yml;
  • scripts/ci/post-deploy.sh, which the template’s deploy runs after the backend’s rollout;
  • any path the template does not render.

One command runs every check a commit runs, over the whole tree:

Terminal window
make gate

Every check is a gate: one file under ci/gates/ whose stem is the gate’s id, run by the gate runner of the ci project. That id is what a commit hook, a pipeline step and an issue all call the check by.

Terminal window
uv run --quiet --project ci gate list # every gate and what it checks
uv run --quiet --project ci gate run <id>... # run some of them
uv run --quiet --project ci gate run --all # what make gate runs

The template’s gates:

IdWhat it checks
no-nul-bytesNo NUL byte in a file .gitattributes does not declare binary
devcontainer-declarationThe checkout is mounted at the folder the container works in
shell-testsEvery *.test.sh under scripts/ and scripts/ci/
rust-fmtcargo fmt --all -- --check
rust-clippycargo clippy --locked --workspace --all-targets -- -D warnings
rust-doccargo doc --locked --workspace --no-deps, with private items (see below)
rust-testcargo nextest run --locked --workspace, which skips gg’s unit tests; no hook
rust-doctestcargo test --locked --workspace --doc; no hook
k8s-manifestsThe template’s render of every overlay; see Kubernetes
web-lintnpm run lint: ESLint over the TypeScript
web-typecheckThe web console’s tsc -b
web-testThe web console’s Vitest run, under jsdom
web-browser-testThe web console’s Vitest run, in real browser engines
web-buildThe web console’s vite build
markdownlintThe Markdown style, from .markdownlint-cli2.yaml, at 90 columns
cspellThe prose’s spelling, from cspell.json and .cspell/project-words.txt
formatPrettier, with .prettierignore as its scope (plus the k8s manifests and the case-harness fixture page) and the frozen versions left out
docs-typecheckThis site’s astro check
docs-buildThis site’s build
python-lintruff over ci/
ci-testspytest over ci/

The project’s gates, answered in extra_gates:

IdWhat it checks
frozen-pathsNo change to a frozen test-case version
seeded-contractNo evaluation vocabulary in the packages seeded into a model’s workspace
spec-vocabularyNo evaluation vocabulary in a non-frozen version’s seeded specs
spec-prosemarkdownlint and cspell over the test-case, game-jam and group prose
audio-packsEvery version’s [audio] packs resolves against the pack registry
build-contextEvery project Dockerfile’s COPY sources, and the containers/ Rust pins
k8s-deploy-setsThe staging and prod deploy sets, pinned to a commit
ci-image-pins.azure/project/jobs.yml pins the Rust CI image ci/images/tags.yml names
scripts-testnode --test over scripts/lib
file-endingsEvery tracked file outside a frozen version ends in one newline; no hook
workspace-testEvery npm workspace’s Vitest run but the console’s and this site’s; no hook
validators-typechecktsc over every test case’s validator projects; no hook
site-buildThe gallery’s build; no hook
contract-driftThe generated data contract is current; no hook

.pre-commit-config.yaml runs each gate on every commit as one hook, triggered by the files it judges, except the ones marked “no hook”, which run in make gate and the pipeline only. scripts/setup-hooks.sh installs the hook, and the dev container runs it when the container is created. The same file carries the checks the hook framework brings from its pinned upstream repositories, the file checks of pre-commit-hooks and shellcheck; they are hooks, not gates, so make gate does not run them and pre-commit run --all-files does.

Two commit-time stopgaps stand until the template takes project excludes and hookless gates:

  • The upstream check-added-large-files hook cannot exclude test-case media, so a commit adding a reference screenshot, showcase video or clip over 500 kB sets SKIP=check-added-large-files.
  • rust-clippy, rust-doc, web-typecheck, web-test, web-browser-test, web-build, docs-typecheck, docs-build and ci-tests run as commit hooks. A commit whose change has already been gated may name them in SKIP=.

A reference implementation’s scripts are model output, so each slug directory that holds one carries a .shellcheckrc turning shellcheck off beneath it.

Terminal window
cargo build --workspace

The toolchain is pinned to an exact patch release in rust-toolchain.toml, and every checkout builds with that release. Format, lint, and test with:

Terminal window
cargo fmt --all
cargo clippy --workspace --all-targets -- -D warnings
cargo doc --workspace --no-deps --document-private-items
cargo nextest run --workspace
cargo test --workspace --doc

Tests run under cargo-nextest, configured by .config/nextest.toml, which the dev container and the Rust CI image both carry. nextest does not execute doctests, so cargo test --doc runs those separately. cargo doc is what checks intra-doc links, and --document-private-items is required for it to reach crates whose public surface is small. The project’s .cargo/config.toml sets it as build.rustdocflags, so the rust-doc gate documents private items too. Cargo reads that file for every run under the repository, the doctests and gg’s Rust signature reflection included (whose catalogue it leaves unchanged), and an empty RUSTDOCFLAGS in the environment overrides it.

The gates run the same commands, and are the ones to run locally when reproducing a CI failure:

Terminal window
uv run --quiet --project ci gate run rust-fmt rust-clippy rust-doc
uv run --quiet --project ci gate run rust-test rust-doctest contract-drift
scripts/ci/rust-build.sh # cargo build, every crate and target
scripts/ci/gg-test-build.sh # cargo build of gg's unit tests
scripts/ci/gg-test.sh [k/N] # gg's unit tests, whole or one hash partition

gg’s unit tests are the one part of the suite no gate runs; see Testing gg.

Terminal window
uv run --quiet --project ci gate run audio-packs

The audio-packs gate runs at commit and in CI. It requires an [audio] packs declaration on every full-stack and game-jam version that is not frozen, resolves every declared ref against containers/sample-packs/ for name, version, kind, and published clips, and prints the defaults each version’s pack order resolves to. It reads the committed manifests only, so it needs no credentials.

Terminal window
uv run --quiet --project ci gate run spec-vocabulary

The spec-vocabulary gate also runs at commit and in CI. It reads every non-frozen version’s prompt.hbs and specs/**, plus the shared prompt preambles in crates/core/src/prompt.rs, and fails on any word that would tell a model it is being evaluated or that this project exists; the list is in Keeping evaluation out of the seeded set. Hits in frozen versions are reported, not failed. It is dependency-free and finishes in under a second.

rust-toolchain.toml is the template’s, and pins the compiler to an exact patch release, together with the RUST_VERSION build argument in .devcontainer/docker-compose.yml and the CI images. A new compiler therefore arrives with a template update. The images under containers/ build with the same compiler: each FROM rust:<version>-bookworm and each ARG RUST_VERSION=<version> default there must equal the toolchain file’s channel, and the build-context gate, whose hook also fires on rust-toolchain.toml, fails naming the file, the line and both versions until they do.

Two things follow a bump without anything to do by hand:

  • gg’s Rust arm compiles a model’s program against a set of .rlib files, a compiler-private format rustc refuses from any other release (E0514). The set is generated by crates/gg-sandbox-artifacts/rust into its OUT_DIR, and rust-toolchain.toml is in that crate’s rerun set, so the next build re-cuts it with the new compiler.
  • The toolchain file lists no targets. The wasm target gg’s Rust arm compiles to is installed by scripts/ci/install-rust-wasm.sh instead, because a target named there is fetched on the first cargo invocation of every checkout, including the containers/ builder images that cross-compile nothing.

gg drives a model in one of eleven program languages, and building test-cabinet-gg produces two things for each of them.

A signature catalogue is what a model is told an arm offers: every module, signature, argument, type, and type member, reflected out of that language’s own SDK by that language’s own documentation tool. crates/gg/build.rs generates all eleven into the build’s OUT_DIR, and the arm modules include_str! them from there. A build of gg therefore describes the SDK sources in its own checkout. See Program languages.

An arm’s build artifacts are what a model’s program is compiled and evaluated against: a guest component, an SDK jar, a compiler, a library set. Each arm has a crate under crates/gg-sandbox-artifacts/ whose build script runs that arm’s packages/gg-sandbox*/build.sh into its own OUT_DIR, and crates/gg embeds the result from there. Ten crates serve the eleven arms, because packages/gg-sandbox serves TypeScript and JavaScript from one build. Separate crates give each arm its own rerun set, so editing the Java SDK re-cuts the Java jar alone, and cargo runs the independent arms concurrently.

Both are generated rather than committed. crates/gg/src/sandbox/checkers/ holds five files that are gg’s own source: three hand-written Java compiler drivers and the two JVM arms’ pin declarations. A .gitignore allowlist names exactly those five and ignores everything else in that directory.

Anything that compiles crates/gg needs those toolchains present. That is cargo build --workspace, cargo clippy --workspace, cargo doc --workspace, scripts/build-gg-static.sh, and scripts/gg-reference.sh. Nothing else in the workspace depends on test-cabinet-gg, so a package-scoped build such as -p test-cabinet-cli and the release binaries need none of it.

Install them once, before the first build:

Terminal window
scripts/ci/install-gg-toolchains.sh # the eleven arms, ~1.9 GB, idempotent
scripts/ci/install-gg-build-toolchains.sh # the C# guest's build tools, idempotent
npm ci # the pinned tsc two arms reflect with

The two lists stay separate. The first is every toolchain that a gg run and gg’s reflectors execute; it is installed in the dev container by scripts/devcontainer-setup.sh, in the pipeline by scripts/ci/gg-ci-toolchains.sh, and in the run images. The second is a .NET SDK and an unpruned wasi-sdk that exactly one arm’s artifact build needs, since C#‘s guest is Mono’s IL interpreter, relinked. It installs under its own prefix ~/.local/share/tcab/gg-build/ so the eleven-arm list keeps its meaning and the run images keep their size. With the second prefix populated, packages/gg-sandbox-csharp/build.sh resolves both from it instead of fetching its own copies.

The dev container’s image is the template’s, and carries none of them. After the container is created, the root postinstall (run by the template’s npm ci) starts scripts/devcontainer-setup.sh, which provisions in the background: the apt packages gg’s installers and the asset tools need (ruby, libicu-dev, ffmpeg, cmake, iproute2, lsof, procps), both installers above, and wrangler. That is about 3.4 GB and up to an hour on a first create. It logs to ~/.cache/tcab-devcontainer-setup.log and writes ~/.cache/tcab-devcontainer-setup.done once everything installed; the marker holds a digest of the provisioning inputs, so a moved pin provisions again. Building crates/gg, and the rust-clippy and rust-doc hooks, need it done.

A container stop, a rebuild or a failed install leaves no marker, and nothing restarts the provisioning on the next start. With no marker and no provisioner running, run it again:

Terminal window
bash scripts/devcontainer-setup.sh

In CI, scripts/ci/gg-ci-toolchains.sh is the one provisioner for every job that builds or tests Rust. It installs the pinned Node and both lists into a directory the job caches, with the default install locations under $HOME linked into it, so a warm cache downloads only the wasm32 standard library.

A missing toolchain fails the build with the arm named and the fix stated. A failure naming node_modules is fixed by npm ci.

An interrupted install is safe to re-run, and it is cheap. Between them these scripts fetch about 3 GB — the Swift toolchain alone is 1.05 GB — and every one of those fetches goes through gg_fetch (scripts/ci/fetch.sh), which retries a dropped connection and resumes rather than restarting. The partial archive is kept under ~/.cache/tcab/downloads (the same path the service-image build mounts a cache over), so re-running the installer after a failure picks up where it stopped instead of paying for the whole transfer again; each installer deletes its own archive once the tree it feeds has verified itself. Nothing under an install prefix is touched until every byte is on disk, so a fetch that fails leaves a toolchain a machine already had rather than half of one.

To read an arm’s build artifacts, or a catalogue, run the same scripts the build runs:

Terminal window
scripts/gg-artifacts.sh # -> target/gg-artifacts/, every arm
scripts/gg-signatures.sh # -> target/gg-signatures/<language>.signatures.json

Both accept a destination argument and default to a gitignored directory under target/. Both read scripts/gg-arms.sh, the only enumeration of gg’s arms in the repository, and build.rs calls gg-signatures.sh rather than repeating it. A catalogue is also where reflector bugs surface, such as a dropped @return paragraph or a truncated parameter description.

nextest runs each test in its own process, and most of gg’s tests need a compiled guest component. The test build keeps compiled components on disk in component-cache under the crate’s OUT_DIR, so a guest is compiled a handful of times per build rather than once per test. An entry is keyed on the component’s bytes and wasmtime’s compatibility hash, is stored only once its bytes have been compiled twice, and is evicted least recently used past 4 GiB. cargo clean removes it. Production builds have no such cache.

Test builds compile components on one thread, since nextest already runs one test per core. The root manifest optimizes the Cranelift and wasmtime crates in the dev profile, which is where that compile time goes.

Most of gg’s tests compile a program in one of its language arms, so the suite is the bulk of the workspace’s test time, more than the template’s one Rust gate job can hold. crates/gg/Cargo.toml therefore sets test = false on gg’s lib and bin and doctest = false on the lib (gg has no compiled doctest), so the rust-test gate and make gate skip the suite, while rust-clippy still lints its #[cfg(test)] code. scripts/ci/gg-test.sh runs it with --lib, which overrides test = false, whole with no argument or as one k/N hash partition of it; the pipeline’s gg_tests_<k>_of_4 jobs run one partition each. Run it after any change under crates/gg/:

Terminal window
scripts/ci/gg-test.sh # the whole suite
scripts/ci/gg-test.sh 2/4 # one hash partition of it

The default build dynamically links against glibc and the generic FHS dynamic loader (/lib64/ld-linux-x86-64.so.2), which suits mainstream Linux such as Ubuntu. Distributions that ship no such loader, notably NixOS, need a fully static binary built against the musl target instead. This is opt-in and leaves the default build unchanged.

Prerequisites:

Terminal window
rustup target add x86_64-unknown-linux-musl
# plus a musl C toolchain on PATH (`musl-gcc`), for the little C that the `ring`
# TLS backend and the bundled SQLite compile. On Debian/Ubuntu: `musl-tools`.

The aliases are pinned to x86_64-unknown-linux-musl, so an aarch64 host cross-compiles and needs an x86_64 musl cross toolchain on top of that. The dev container installs the musl target for its own architecture, which is what scripts/build-gg-static.sh builds against.

Then build with the aliases defined in .cargo/config.toml, each of which writes to target/x86_64-unknown-linux-musl/release/:

AliasBinary
cargo build-portabletcab
cargo build-portable-backendtcab-backend
cargo build-portable-dispatchertcab-dispatcher
cargo build-portable-drivertcab-driver
cargo build-portable-artifactstcab-artifacts
cargo build-portable-gggg

The backend links statically too: its SeaORM SQLite driver compiles SQLite from vendored C source with the same musl toolchain, and its PostgreSQL driver is pure Rust over rustls. Under its cli runtime the driver still shells out to a host container runtime, so that host needs Podman or Docker on PATH and the harness API keys for the harnesses it runs; see the driver configuration.

gg executes inside the run container, and a run’s image varies by test type: most are glibc Debian bookworm, and the blender image is Ubuntu. The static musl build is what lets one binary run across every run-container image.

Terminal window
# Release/CI build (x86_64), matching the other `build-portable-*` aliases:
cargo build-portable-gg
# Arch-adaptive build for the HOST architecture. Use this in the dev container
# (aarch64 on Apple Silicon) for a local k3d cluster. It picks the musl target,
# wires up `musl-gcc`, and prints the built binary's path:
scripts/build-gg-static.sh # -> …/<host-arch>-unknown-linux-musl/release/gg
scripts/build-gg-static.sh /out/gg # …and also copies it to /out/gg

The driver image bakes gg in from the same script, installing it at /usr/local/lib/tcab/gg with TCAB_GG_BINARY pointed at it, so make -C deployments/local local-up produces a driver image carrying gg and a Kubernetes run installs gg locally. Set TCAB_GG_INSTALL=release to pull a published release instead. See gg installation.

Install all workspace dependencies from the repository root, then build every workspace that defines a build:

Terminal window
npm install
npm run build

The other root scripts delegate to each workspace that defines them: npm run dev, npm run test, and npm run typecheck. npm run lint is ESLint over the whole workspace from the root eslint.config.js (the web-lint gate).

npm run test runs vitest in each workspace. Iterate on the gallery’s own suite with npm run test -w @clockwyrks/ui. On a clean checkout, build the workspace runtime packages first with npm run build:packages, since the tests import them from a built dist/.

packages/ui/src/app/pages/routeSmoke.test.tsx mounts the routed app at every path in routePatterns and asserts that each one renders, meaning the page error boundary caught nothing and the page put content on screen. It walks that table against three hosts: a console with a stocked catalog, a console holding nothing, and the read-only static gallery.

This suite covers whether a page loads at all; the per-page suites cover whether it behaves correctly. Adding a route to routePatterns adds it to the walk.

The page error boundary in packages/ui/src/app/components/PageErrorBoundary.tsx is the runtime counterpart. It wraps the routed body, so a page that throws is contained to itself: the chrome and section nav stay, the panel names the failure, and navigating to another page clears it.

Type-check every test case’s validators with:

Terminal window
npm run typecheck:validators

It runs tsc --noEmit over each test-cases/**/<version>/validation/<engine>/ project. Nothing else compiles them: a run stages a project into the produced tree and vitest transpiles it without checking types, so this is the only gate on a validator that fails to compile. It is the validators-typecheck gate, a step of the web job.

A project resolves the build’s ../src/* against the case’s reference implementation and the shared harness’s ./case-harness/* against packages/case-harness/src, through the rootDirs its config declares. See Writing Case Specifications and Prompts.

Lint the authored prose with:

Terminal window
npm run lint:specs

It runs the spec-prose gate: markdownlint-cli2 over the Markdown in test-cases/**, game-jams/** and test-case-groups/**, then cspell, from .cspell/specs.json, over the Markdown, HTML, CSS, TOML, and Handlebars there that a case or jam ships to a model. Add a legitimate domain term to .cspell/project-words.txt when cspell flags it. This is the linter the test-case authoring and variant guides refer to under “Validate your work”. The rest of the repository’s Markdown, this site included, is the markdownlint and cspell gates’ (npm run lint:docs).

Check formatting over the whole checkout with:

Terminal window
npm run format:check # check, the format gate
npm run format # rewrite

It runs prettier --check over every file prettier has a parser for: the packages, the front ends, and every test case’s validators, seeded workspaces and reference implementations. A formatting warning anywhere fails the check. Two things are left out: what .prettierignore names (build trees, vendored copies, the Handlebars templates prettier does not parse, and prose, which markdownlint and cspell own), and the frozen test-case versions, which scripts/format-check.mjs derives from their .frozen markers on each run. .prettierignore comes from the project template, and two of its rules are broader than this project: the Kubernetes manifests under deployments/k8s/ and the case-harness fixture page under packages/case-harness/test/build/ are still checked, in a pass that honours .gitignore alone. It is the format gate.

The run-record (and arena, job-API, backend) data contract has a single source of truth: the Rust types that derive ts_rs::TS and schemars::JsonSchema behind their contract feature, in crates/core and crates/backend. The TypeScript bindings under packages/run-record/src/ and packages/asset-contract/src/ — one generator, two packages, because only the latter may be seeded into a run — and the JSON Schemas under apps/docs/public/schema/ are generated from those types by crates/contract-codegen. After changing any contract type, regenerate and commit:

Terminal window
npm run gen:contract

It runs the generator, formats the output with Prettier, mirrors gg’s built-in system-prompt templates into the TypeScript contract, and compiles the regenerated package. The contract-drift gate regenerates and fails on any diff, so the Rust, TypeScript, and JSON Schema representations stay in step.

It compiles test-cabinet-core and test-cabinet-backend only, so it needs none of gg’s program-language toolchains.

The console’s gg Reference section is served by the backend from files gg projects: the whole tool vocabulary, and every responses-as-code function and type on each of the eleven arms, carrying the bytes a model’s own documentation lookup renders. The backend must not link test-cabinet-gg, so the projection crosses that boundary as files:

Terminal window
scripts/gg-reference.sh # -> target/gg-reference/

That writes index.json plus one <language>.json per program language. The backend finds them there by default through TCAB_GG_REFERENCE, which resolves under TCAB_BACKEND_CHECKOUT when unset. The backend image bakes the same files at /opt/gg-reference and the driver image bakes the binary that projected them, so a deployed console and the harness its runs execute come from one build of one checkout. A backend with no documents serves everything else normally and answers 503 on the two reference endpoints.

The pipeline builds that binary once, in the gg_<arch> jobs, and publishes it as a per-architecture artifact together with the documents it projects. The backend and driver image builds consume that artifact in place of the Dockerfile stage that would link a second copy, which is what makes the run images’ self-check, the driver’s baked harness and the console’s reference one binary per commit. The Dockerfile stage remains what an offline make -C deployments/local images builds.

The script builds gg, so it needs the toolchains described under gg and its eleven toolchains. Running the projection itself needs none of them, because every arm’s catalogue is compiled into the binary.

Every language arm compiles inside the run container, against the toolchain tree the -gg image variants carry, so the question a build cannot answer is whether that tree works where it is copied to. gg selfcheck answers it: it drives every arm’s bootstrap program through the real preparation, guest and views in whatever environment the binary is running in, and exits non-zero when any arm fails.

Terminal window
make -C deployments/local run-images-gg-selfcheck # build gg, build the five, check them

That target builds the static binary with scripts/build-gg-static.sh and hands it to containers/build.sh --gg-selfcheck, which drives the check inside each image between building it and pushing it — installing the binary the way a run does, with docker cp to /tmp/gg and an exec as the unprivileged run user. Naming the images runs the same command by hand:

Terminal window
containers/build.sh --gg-selfcheck target/gg-selfcheck/gg \
sprite-gg base-wasm-gg voxel-gg full-stack-3d-gg blender-gg

Five images answer for all twenty-seven variants, because each copies one builder image’s /opt/gg to one absolute path and so carries the same tree byte for byte, and what differs is the environment it runs in: the image the lineage is rooted at plus every package a run image installs on the way down. build.sh re-derives that grouping from the Dockerfiles on every gated build. CI passes the same --gg-selfcheck flag post-merge, before the images are published.

The tree being identical is a fact about the files, not about how the registry stores them. Each variant carries its own layer for those bytes, so a host that pulls two variants pulls the tree twice, and fifty-five images with twenty-seven copies of a 2.2 GB tree do not fit on a build agent. So the CI build sets RECLAIM=1 alongside PUSH=1, and containers/build.sh then removes each image from the local store once nothing later is FROM it and prunes the builder cache as it advances. RECLAIM is deliberately a separate switch: the builder prune takes every cache mount on the daemon, including ones belonging to other Dockerfiles, so a publish from a development machine does not set it. A local build sets neither and keeps everything, because there the images are what is being produced.

cargo run -p test-cabinet-gg -- selfcheck asks the same question of this machine’s own toolchains, which is the form to run while working on an arm. See the self-check.

Azure Pipelines is the project’s only CI. azure-pipelines.yml is the template’s: it gates a commit, and on staging publishes the backend image and deploys the staging overlay. The project’s own jobs and stages come from the files under .azure/project/, which it includes. Every job delegates to a script, so a failure reproduces locally by running the same script; scripts/ci/README.md lists them. The YAML names the image a job runs inside, its caches, the credentials a step runs under, and which registry round trips are retried, and nothing else.

Pushes to master, staging and nightly trigger it, batched: pushes that arrive while a run of the same branch is in flight get one run, for the newest commit, so an intermediate commit is not gated, mirrored or deployed on its own. Pull requests into master and staging run it through build validation policies on those branches, because Azure Repos ignores a pr: block. Tags do not trigger it; a v* tag runs the release pipeline instead (see The release pipeline).

StageRuns onWhat it does
gatesevery runThe template’s rust and web jobs, and the project’s jobs below; on master and staging also the static gg builds and every image; the GitHub mirror on master, staging and nightly
publishstagingThe template’s: builds deployments/images/backend.Dockerfile and pushes the-test-cabinet-backend:<commit>
deploystagingThe template’s: rolls the staging overlay onto the commit; see Kubernetes
prodmasterThe project’s: the same backend publish, then scripts/ci/deploy-environment.sh prod <commit>
gg_releasemasterThe project’s: uploads the static gg binaries; see Releasing gg
docsmaster, stagingThe project’s: deploys this site

Every stage after gates depends on it, so a red job in the gates stage (a gate, a project job, or a template job past its 60 minutes) stops every image job, the mirror, the publish and deploy, prod, gg_release and docs.

The template’s two gate jobs run every gate, one step per gate, in the CI image of their track: rust runs the Rust gates and contract-drift, and web runs the rest, every project gate but contract-drift included. Each step runs uv run --quiet --project ci gate run <id>, the command a commit hook and make gate run it with, so a failing check is one red step named for it. Each upstream hook runs as pre-commit run <id> --all-files in a step of its own in the web job. Each job is capped at 60 minutes.

The project adds steps to both through .azure/project/setup-steps.yml and .azure/project/steps.yml:

  • In rust: scripts/ci/free-disk-linux.sh, which reclaims what the template’s own reclaim leaves; the gg toolchains, restored from the gg-toolchains cache and completed by scripts/ci/gg-ci-toolchains.sh, since the Rust CI image carries no Node and no gg toolchains; and, last, scripts/ci/cargo-target-prune.sh, which removes the executables cargo linked before the template saves its target/ cache, so the cache holds the compiled libraries and fits beside the tree on the agent’s disk.
  • In web: npm run build:packages, and a step that sets SKIP=end-of-file-fixer and logs a warning saying so on every run. The upstream end-of-file-fixer hook takes no project excludes, and 109 files inside frozen versions, which may never change, fail it. The hookless file-endings gate runs the same pinned hook over every tracked file outside the frozen versions instead, and fails if the project’s pipeline files set any other SKIP. This is the only step the project skips.

.azure/project/jobs.yml adds these jobs to the gates stage. Their bodies live in .azure/tcab/, shared with the release pipeline. The amd64 jobs that compile Rust run as container jobs in the template’s Rust CI image, with the template rust job’s paths, variables and cargo cache, so they share its caches and seed them. A jobs template cannot read the pipeline’s image tag variable, so jobs.yml names that image as a literal parameter default: scripts/ci/tcab-image-pin.sh writes it from ci/images/tags.yml, and the ci-image-pins gate fails when the two differ.

JobRuns onWhat it does
gg_tests_<k>_of_4every rungg’s unit tests, one hash partition per job: gg-test-build.sh, then gg-test.sh k/4
rust_buildevery runrust-build.sh, the link of every target. When no seed exists for this Cargo.lock, it also runs the template gates’ clippy, doc and test builds and saves target/ under a key the template rust job restores
binary_linuxevery runThe release build and tests of tcab, its doctests, and the binary smoke
binary_windowsevery runThe same on windows-2022
submodule_pinsevery runsubmodule-pins.sh; see Submodules in CI
gg_amd64, gg_arm64master, stagingThe static gg binary per architecture (gg-dist.sh)
audiostore_<arch>, runimages_<arch>, services_<arch> and their manifestsmaster, stagingEvery image, built natively per architecture and fused by manifest.sh; the eight service images of an architecture on one builder, so their Rust compile happens once
mirrormaster, staging, nightlyForce-pushes the branch to GitHub, after every check job

rust_build exists because the template’s rust job is one job capped at 60 minutes, and a pipeline cache is saved only after every step of a job succeeded: a cold rust job that times out never warms its own cache. The first run after a template update that moves the compiler, or after the cache expires, may still time out once; queue a second run.

The image jobs sit in the gates stage on purpose: the template’s publish stage retags tcab-backend:<commit>, which they build, and the staging deploy pins every image they push. amd64 builds on Microsoft-hosted ubuntu-24.04 agents and arm64 on the organisation’s arm64 pool pool-dev-linux-arm64-wus3-4c-eph-01. Each architecture pushes <image>:<sha>-<arch> to testcabinet.azurecr.io, and manifest.sh fuses the two into the multi-arch <image>:<sha> a deployment pins. The audio store is built first, because the driver image bakes test-cabinet-audio-store:<sha>, and the run images are pushed only after gg selfcheck passes inside each -gg environment. The eight service images of an architecture are built by one job on one builder (service-images.sh), because seven of them share one cargo build of the service binaries; a separate job per image compiled it seven times over. The 27 -gg run images share one /opt/gg layer, so the registry stores it once and each push after the first mounts it.

The mirror job force-pushes the branch with its tags to github.com/TheClockwyrks/TheTestCabinet, with the deploy key held in the secure file github-mirror-key. It is the only thing that pushes there, so the mirror holds gated commits only.

azure-pipelines-release.yml is the project’s second pipeline, triggered by v* tags and nothing else. Its first job, gated, runs scripts/ci/require-gated-commit.sh: it looks for a run of the main pipeline at the tagged commit, on any branch, and waits until that run’s check jobs (rust, web, the gg test partitions, rust_build, both binary jobs and submodule_pins) have passed. The images, publish, deploy and docs results never gate a release. When no run exists at that commit, it queues one on the tag ref, where every branch-conditional job and every later stage skips, so that run is the gates stage alone. Then it builds and publishes the tcab binaries and gg, and mirrors the tag. See Releasing.

The template’s gate jobs, and the project’s Rust container jobs, run inside the template’s CI images:

ImageJobs
testcabinet.azurecr.io/ubuntu-the-test-cabinet-rust-cicdrust, gg_tests_<k>_of_4, rust_build, binary_linux, gg_amd64
testcabinet.azurecr.io/ubuntu-the-test-cabinet-web-cicdweb

ci/images/README.md describes each. The Rust image carries the pinned toolchain, nextest, the musl target, and the apt packages the rust_ci_packages answer names (cmake, ffmpeg, libicu-dev, python3, ruby, zip); gg’s toolchains and Node come from the gg-toolchains cache. An image’s tag is a digest of the files it is built from, and ci/images/tags.yml holds it. After changing an image input (a template update does), run scripts/ci/ci-image.sh tag, commit the file, and let the image pipeline, azure-pipelines-ci-images.yml, push the image before the gates run. scripts/ci/tcab-image-pin.sh then writes the new tag into jobs.yml.

A job pulls its CI image from the registry before its first step runs, so a registry answer that times out would fail the job before it has checked anything. Every project container job therefore sets VSTSAGENT_DOCKER_ACTION_RETRIES, which makes the agent retry each Docker login, pull and start, three attempts ten seconds apart. The template’s rust and web jobs take no job variable from the project, so the main pipeline’s definition in Azure DevOps (the-test-cabinet) sets the same variable, as a pipeline variable, for them. Every registry login and manifest fuse in the image jobs carries a step retry of its own. Only those registry round trips are retried; a build runs once, so a failed one keeps its reason in the log.

No job holds a stored credential. The pipelines authenticate through workload-identity-federated service connections and a few secret pipeline variables:

NameKindUsed for
the-test-cabinet-acrDocker Registry connectionAcrPush on testcabinet.azurecr.io: pulls the CI images, pushes them from the image pipeline, and pushes every image
tcab-deployAzure Resource ManagerThe publish stage’s registry sign-in (AcrPush) and the cluster deploys: the custom “Test Cabinet AKS Command Invoke” role on each cluster, “Azure Kubernetes Service RBAC Admin” on its application namespace
tcab-gg-publishAzure Resource ManagerStorage Blob Data Contributor on testcabinetartifacts
github-mirror-keySecure fileThe GitHub mirror’s write deploy key
CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_IDSecret pipeline variablesThe docs deploy
CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_AUDIO_R2_BUCKET, CLOUDFLARE_AUDIO_R2_PRESIGN_ACCESS_KEY_ID, CLOUDFLARE_AUDIO_R2_PRESIGN_SECRET_ACCESS_KEYSecret pipeline variablesStaging the audio store out of the audio object store (read-only)

The secret variables must be set on the pipeline before master or staging builds can complete their images and deploy stages. The release pipeline needs the-test-cabinet-acr, tcab-gg-publish and github-mirror-key authorized for it, and its build identity allowed to view and queue runs of the main pipeline.

Each pipeline is created once, from the repository’s default branch, and then follows the YAML on whichever branch triggers it. The gates pipeline is the-test-cabinet and the image pipeline tcab-ci-images; the release pipeline is created beside them:

Terminal window
az pipelines create --org https://dev.azure.com/genyume -p the-test-cabinet \
--name the-test-cabinet-release --repository the-test-cabinet \
--repository-type tfsgit --branch master \
--yml-path azure-pipelines-release.yml --skip-first-run

The release pipeline names the main pipeline by its name in its resources.pipelines block (the-test-cabinet); a main pipeline registered under another name needs that source changed to match.