Skip to content

Architecture

The Test Cabinet is a headless core with a set of components layered on top of it. The core owns run orchestration: resolving a test case version, seeding a run’s repository, executing the run in a container, invoking the agent harness, collecting metrics, running validation, writing the run record, and publishing. Every other component links against that library and exposes it under the interface it provides, whether that is a CLI, an HTTP API, or a browser console. Keeping orchestration in the core is what makes batch runs, unattended sweeps, and remote execution possible.

ComponentWhat it is
CoreThe Rust library that implements the run lifecycle and owns the data contracts.
CLIThe tcab binary. Exposes the core so runs can be scripted and swept in batch.
DispatcherA controller that claims queued runs from the backend and creates one driver Job per run.
DriverThe per-run executor. It runs exactly one test case in a Job, streams its progress to the backend, and exits.
ArtifactsA data-plane service that serves produced run trees (playable builds, proof and asset media) off a persistent volume.
ArenaA data-plane service that runs the adversarial test type’s matches and tournaments off the backend.
Web consoleThe browser console: the primary interactive way to launch runs, watch them live, review them, and publish.
BackendA private Rust server that distributes test case definitions, owns the run queue, and stores run results.
Auth serviceA standalone Rust server for user accounts: self-registration, password login, and the bearer tokens the backend verifies.
SiteThe public static gallery at testcabinet.ai where published runs are browsed and played.
UI libraryShared frontend code (@clockwyrks/ui): the routed gallery application both GUIs mount, the primitives they render, and the backend client interfaces.
Voxel runtimePoses and renders a produced voxel rig.
Particle runtimeSimulates and renders a produced particle system.
DocsThis documentation site.

Two roles recur across the components.

  • A runner executes a test case. The driver is the only one. It resolves the requested test case version from the backend, drives the run through the core, creates an untrusted sandbox pod through the Kubernetes API, and reports the result back to the backend.
  • A reporter displays run results: the web console and the public site. Reporters read stored results and let a person interact with the produced implementations.

The web console is a launcher and reporter in one. It enqueues runs, watches them live, reviews them, and shows results in one place. The web console is the primary way The Test Cabinet is used.

Both GUIs mount the same routed gallery application from the UI library. The console is that application with the launch surface enabled, and the public site is the same application with it off.

A run launched from the CLI or the web console executes on the cluster rather than on the launcher’s machine. The launcher enqueues the run at the backend, which owns the run queue. A dispatcher claims the queued run and creates one Kubernetes Job running a driver.

The driver executes the run, creating an untrusted sandbox pod through the Kubernetes API. It streams live progress back to the backend, which relays it to the launcher. It uploads the produced tree to the artifact service and reports the produced record, which the backend stores privately.

Each run is one schedulable Job, so concurrency scales with the cluster. A launcher needs a reachable backend and an account, and no container runtime of its own. Local development runs the same manifests on a k3d cluster, so a run is a Job everywhere. See Kubernetes: run plane.

The backend records run results and serves as the canonical copy of the test case definitions runners need. It has no public write surface and sits on a private network, so reaching it is the first line of access control.

User accounts, held in the standalone auth service, identify who acts, so that every review is attributed to a person. The backend verifies the auth service’s bearer tokens on the mutating run endpoints, review and publish. Reads stay open.

The public site is a fully static, backend-less deployment. Publishing exports a public snapshot of the published runs that the site builds from, so the gallery has no live dependency on the private backend.

At a high level, launching a run must:

  • Select a test case version, an agent harness, and a model, resolving the version from the backend.
  • Seed a fresh git repository with the selected variant’s data.
  • Start a container and invoke the agent harness against the seeded repository.
  • Surface the harness’s activity as a live stream of harness events while the run is in progress.
  • Record metrics as the run proceeds and collect the produced repository when it finishes.
  • Run validation over the produced implementation.
  • Write a run record and report it to the backend, which stores it privately.

A stored run reaches the gallery through two explicit steps. Review collects assessments of the produced run from people other than the operator, and publish releases the produced code and flips the reviewed run public. See Results.

Progress that happens inside the run container reaches a watching viewer in real time over a dedicated channel. See Live Streaming.

The word harness is used two ways throughout these docs.

  • The testing harness is The Test Cabinet’s own application that runs benchmarks.
  • An agent harness is a coding tool, for example Claude Code or Codex, that drives a model through a test case. See Agent Harnesses.