Skip to content

v0.7.0 (2026-09-25)

The v0.7.0 release follows v0.6.3 and its centre is gg, a coding harness authored inside the repository and executed inside the run container, which lets the cabinet vary one part of an agent at a time: how it answers a turn (tool calls or a whole program), which of eleven languages it writes that program in, how it compacts its context, what it remembers, and how it delegates. Every gg run streams a typed telemetry record that the console reads natively, so a run can be read turn by turn, request by request, without a second tool.

The second centre is the test cases themselves. A run now names an engine beside its case, variant, harness and model, and eighteen playable cases ship a version on which every checklist point is decided by a validator, so a run’s functional rating is known the moment it completes and a reviewer supplies only the aesthetic judgement. The captured baseline media those validators compare against left the repository for a cold-storage submodule, and every case declares the audio packs its run container is staged with.

Underneath both, every model carries its developer’s list price and every gg run is pinned to that developer’s own provider, so a run’s cost and its session duration become properties of the model rather than of the route a request happened to take.

The cabinet also moved house. Azure Pipelines is the one pipeline, images are pushed to the Test Cabinet container registry by commit sha, gg’s release binaries live on Azure Blob Storage, and the GitHub repository is a mirror. The Tauri desktop app is gone, and the web console is the only console.

This remains pre-1.0 software.

The sections below are what an operator has to do or know before the first run on v0.7.0. Backend and console must deploy together, because the release adds telemetry event kinds, enum values (canceled, limit_exceeded, the aesthetic ratings) and RunEvent fields the console reads. The schema changes are listed at the end of this page.

A run of a model can be enqueued only when the model’s catalog entry carries a dated list price: uncached input, cached input and output rates per Mtok, and the date they were taken. This applies to every harness, including Claude Code, Codex and Antigravity, and to every enqueue path (the run form, a gg launch, a coverage top-up and a retry). The seeded catalog carries no list price, so a freshly upgraded deployment refuses every launch until an operator enters one on each model’s edit page. The Fill from OpenRouter control seeds the three rates from the model’s official endpoint for confirmation. See adding or updating a model.

A gg launch additionally needs a resolvable context window and a non-empty provider candidate list for every model it binds; a model whose developer has no endpoint on OpenRouter is refused at enqueue, and providerPin on the model entry names the developer’s provider where its name differs from the model id’s author segment.

With TCAB_CONTAINER_REGISTRY unset a run resolves its image from testcabinet.azurecr.io, the registry the pipeline pushes every run image to. The compiled default tag remains latest, which the pipeline never pushes. A pipeline deploy sets TCAB_CONTAINER_TAG to the commit sha automatically; any run outside that, such as a host-side tcab run or a hand-applied overlay, must set TCAB_CONTAINER_TAG to a sha that master or staging has published. The generic staging and prod overlays carry the placeholder REPLACE_SHA.

Every run image must be rebuilt and pushed at the v0.7.0 commit before the first run, for three reasons. The base image now creates /opt/audio and writes the .tcab-audio-contract marker, and a full-stack, game-jam or audio run reads that marker back at container start and fails, naming the image and the pin, against an image that predates staged audio. Every run image other than base publishes a <name>-gg variant carrying gg’s language toolchains, which a gg run resolves unconditionally. And the new test-cabinet-full-stack-3d image must exist for Gantry, the first case with asset_dimension = "3d", or every run of it fails on voxel: not found.

containers/build.sh builds a -gg variant through its parent, runs gg selfcheck inside each variant before pushing, and rebuilds a parent whose Dockerfile is newer than the image, cascading to its children. The per-image override for a variant is TCAB_CONTAINER_IMAGE_<NAME>_GG.

The audio store and the Cloudflare secrets

Section titled “The audio store and the Cloudflare secrets”

Audio is no longer baked into run images. The driver image copies the published packs to /opt/tcab-audio from the data-only test-cabinet-audio-store image named by its AUDIO_STORE_IMAGE build arg, which defaults to scratch, so a driver built by hand must pass a published store or one built from the checkout (make -C deployments/local audio-store). TCAB_AUDIO_STORE points a driver at a store elsewhere. For a local tcab run or tcab validate on a case that declares packs, run az acr login --name testcabinet and scripts/fetch-audio-store.sh once, then export the TCAB_AUDIO_STORE it prints; --stage builds the store straight out of the object store instead.

The pipeline’s audio-store job reads the secret variables CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_AUDIO_R2_BUCKET, CLOUDFLARE_AUDIO_R2_PRESIGN_ACCESS_KEY_ID and CLOUDFLARE_AUDIO_R2_PRESIGN_SECRET_ACCESS_KEY, and the docs job reads CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID. No run-image or service-image build needs an audio credential any more. The TCAB_SAMPLE_PACK, TCAB_INSTRUMENT_BANK, TCAB_SAMPLE_PACK_DIR and TCAB_INSTRUMENT_BANK_DIR variables are gone from the audio binaries and the images; a palette is chosen only by the case’s [audio] packs and, host-side, TCAB_AUDIO_DIR.

The driver image bakes gg as a static musl binary at /usr/local/lib/tcab/gg and core copies it into each sandbox, so a cluster run installs gg with no network egress. Otherwise core downloads v<version>/gg-<target> from TCAB_GG_RELEASE_URL, which defaults to the gg-releases container at https://testcabinetartifacts.blob.core.windows.net/gg-releases. TCAB_GG_RELEASE_REPO is no longer read. TCAB_GG_INSTALL (local or release), TCAB_GG_BINARY, TCAB_GG_RELEASE_VERSION and TCAB_GG_RELEASE_TARGET keep their meanings, and the dispatcher forwards all five into every driver Job. A prerelease tag such as v0.7.0-rc1 equals no crate version, so a deployment on one sets TCAB_GG_RELEASE_VERSION by hand.

The backend serves GET /gg/reference from TCAB_GG_REFERENCE; the backend image bakes the documents at /opt/gg-reference. A backend deployed from tarballs needs v<version>/gg-reference.tar.gz unpacked and the variable set, or the console’s gg Reference section answers 503 while everything else works. OPENROUTER_API_KEY must reach core for injection into the run container, since gg reads it from the environment and nothing else.

TCAB_ARTIFACTS_URL is the in-cluster artifact service address every backend-originated call uses, distinct from the browser-facing TCAB_ARTIFACTS_PUBLIC_URL; without it a run delete cannot prune its tree. TCAB_ARTIFACT_SWEEP_INTERVAL_HOURS (default 6, 0 disables) and TCAB_ARTIFACT_SWEEP_GRACE_HOURS (default 24) tune the sweep that reclaims orphaned trees, and the artifact service’s GET /runs and DELETE /runs/{id}/artifacts require TCAB_BACKEND_SERVICE_TOKEN. Model probes bill to TCAB_OPENROUTER_API_KEY on the backend, which is distinct from the runners’ OPENROUTER_API_KEY; without it a probe trigger fails with openrouter_key_missing. TCAB_VITEST_TIMEOUT_SECS overrides the 45-minute cap on a case’s whole validator suite run.

TCAB_K8S_RUN_MEMORY_LIMIT is removed from the base manifest and the local overlay and left unset (commented out) in the env examples, and TCAB_DISPATCHER_DRIVER_MEMORY_LIMIT no longer defaults; a dispatcher test refuses any manifest under deployments/k8s that sets either. Size TCAB_K8S_RUN_MEMORY_REQUEST and TCAB_DISPATCHER_DRIVER_MEMORY_REQUEST (default 2Gi) to the heaviest case’s real peak, because a run pod that outgrows its request on a genuinely full node is still evicted. TCAB_DISPATCHER_DRIVER_CPU_LIMIT defaults to 2, which bounds how wide the post-run toolchain fans out and so the driver’s memory peak; over-limit CPU is throttled, never killed. See the run plane.

The catalog is re-ingested, and some stored records become unreadable

Section titled “The catalog is re-ingested, and some stored records become unreadable”

The definition store format is 2 and stored definitions carry an engine support set, so the backend holds itself unready and GET /test-cases answers 503 until a whole-catalog ingest rewrites the store. The overlays’ startup sidecar posts a forced ingest on every backend start, so a pipeline deploy does this itself; a hand-run deployment runs scripts/reingest.sh. Case showcases and reference builds keyed by engine appear only after that ingest.

The run-record format is 3. Stored records that lack the gg openingTurn field or the per-test toolchain report read as unreadable after upgrade, appear only in the console’s Unreadable worklist at /runs/unreadable, and can be deleted there; a re-push with a readable record restores one. Saved gg configurations and recorded gg runs from before stable agent ids do not parse. Scores and ratings of existing published runs can also change: a run is now scored against its own version’s and engine’s checklist rather than the case’s latest, a validator-rated run that did not complete is unrated rather than Flawless, and the toolchain typecheck gate applies at read time.

Validation baselines live in the cold-storage submodule

Section titled “Validation baselines live in the cold-storage submodule”

The captured baseline media is no longer in the tree. Fetch it with git submodule update --init --depth 1 cold-storage to capture or review baselines; tcab capture-baselines and tcab publish-reference refuse to run into an uninitialized submodule and name that command, and TCAB_COLD_STORAGE_DIR redirects the root. The deployed ingest sidecar fetches the submodule itself on every start, and a backend ingesting from a checkout without it serves versions with no baseline media. A submodule commit must be on cold-storage’s master before its pin is committed here, or the submodulepins gate fails.

Reference builds for the new case versions

Section titled “Reference builds for the new case versions”

test-cases/reference-builds.lock.json is keyed <env>.<slug>.<version>.<variant>.<engine>, so a reference build is recorded per engine. A case version’s reference build is recorded with tcab publish-reference --env <prod|staging> <slug> <version> --all-variants (after npm ci && npm run build:packages at the root, with wrangler, CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID), the lockfile is committed to the branch the environment tracks, and scripts/reingest-cluster.sh --env <env> is run, before the Reference tab and the catalog carry it. A reference-capable case has a recorded reference build by the time the release that makes it non-experimental goes live. See publishing a reference implementation.

Before the first pipeline deploy to an environment: grant the arm64 agent pool pool-dev-linux-arm64-wus3-4c-eph-01 to the project, because the AKS nodes are arm64 and Azure has no hosted arm64 agents; create the service connections tcab-acr (Docker Registry, AcrPush on testcabinet.azurecr.io), tcab-deploy and tcab-gg-publish (Azure Resource Manager, workload identity federation); upload the secure file github-mirror-key; and set the secret variables named above. tcab-gg-publish needs Storage Blob Data Contributor on testcabinetartifacts.

Create the custom role from deployments/azure/aks-command-invoke.role.json and assign tcab-deploy both “Test Cabinet AKS Command Invoke” on each cluster and “Azure Kubernetes Service RBAC Admin” scoped to <cluster>/namespaces/tcab-<env>. Both clusters’ kubelet identities need AcrPull on the registry. Apply the cluster-scoped bootstrap by hand once per environment (kubectl apply -k deployments/k8s/cluster/azure-prod or azure-staging) after installing ingress-nginx and cert-manager through Helm and creating the Key Vault secrets, and re-apply it whenever that folder changes.

Configure build validation policies on master and staging in Azure Repos so pull requests run the gates, because Azure ignores the pipeline’s pr: block. See the Kubernetes overview.

There is no desktop console and no macOS build of tcab. The web console at apps/web is the only console, and a release’s tcab binaries are the tcab-linux and tcab-windows artifacts of the tag’s pipeline run; every other platform builds from source. TCAB_DESKTOP_IMAGE_TAG is no longer read anywhere.

sample_pack and instrument_bank are rejected; the packs a case is staged with are [audio] packs = ["name@version", ...], required on every non-frozen full-stack version and every game jam, and a manifest with no [audio] table takes the pinned four-pack default. On a per-engine version every graded point must carry validation, domains and a failure_cap, which a legacy single-workspace version rejects by name. A gg configuration must carry openingTurn on every responses-as-code agent and a registered language, and a configuration that still names healing, ownership, planning or a flat capability set is refused at launch. Frozen versions are unaffected by all of it.

Devcontainer users should rebuild the image, since gg’s toolchains, Playwright’s Chromium and several apt tools are baked in, and re-run .devcontainer/setup-host.sh --force, because the host-runtime socket mount moved out of the shared compose file and a copied docker-compose.local.yml silently loses it. Re-run scripts/setup-hooks.sh: the hooks gained the audio-pack, spec-vocabulary, seeded-contract and build-context gates and lost clippy and rustdoc. The GitHub repository is a force-pushed mirror, so nothing should be committed or opened as a pull request there.

gg: a harness authored inside the repository

Section titled “gg: a harness authored inside the repository”

gg (crate test-cabinet-gg, binary gg, version 0.7.0) is the first harness authored inside the repository rather than integrated from a third party. Unlike the CLI harnesses, gg is the executor: core invokes gg --config <invocation> inside the run container, and the binary holds the model client, the agent turn loop, tool dispatch and the telemetry emitter. Its one secret, OPENROUTER_API_KEY, comes from the environment core injects. The orchestrator dimension does not apply, because gg continues within one session through compaction, and the metadata tab of a gg run shows its capability set in place of an orchestrator.

A run is configured as a named, account-scoped configuration: a capability set of one or more per-agent profiles, each with its own capabilities, implementations, params, model binding, prompt-cache lifetime, reasoning setting and loop-detection knobs. It launches from the ordinary New run page by picking gg in the Orchestrator selector, for every test type. Twenty-two capability ids are registered: shell, read-file, write-file, edit-file, list-dir, search, context-window-override, autoload-specs, agent-persistence, skills, memories, tasks, compaction, agent-managed-context, project-management, subagents, fsm, exec, fork, responses-as-code, program-library and docview-close.

gg checks the whole set before the first turn and refuses the launch on any value it cannot honour exactly as written, naming every defect at once and substituting nothing: an unknown capability id, an unknown param or field, a required value left absent, a mock/… model outside tests. It exits 0 on a natural end, 3 when any agent breached a ceiling (recorded limit_exceeded, never retried) and 1 on a launch failure, a refused credential, a provider-failed root call or a defect of its own (harness_error, retryable). See the gg overview and configurations.

An agent whose type is RaC (capability responses-as-code) answers each turn by writing a whole program over gg’s typed SDK, submitted as the program string of a forced submit_program tool call. That is the only tool such a request offers, pinned through tool_choice where the provider accepts it and auto where it refuses with a 400, downgraded once per model per run. gg runs the string exactly as sent in a wasmtime component, with no fence stripping, no repair pass and no prologue or import written around it. The language’s own compiler or parser accepts or refuses it; a rejected program is a Compiler error turn carrying the compiler’s diagnostic and nothing else, and an uncaught throw or sandbox stop is a Runtime error. A reply with several submit_program calls runs all of them in order and counts at most one error.

Views are the only channel back to the model. gg.views.openFile, openText and openDocsView and gg.docs.search push one message per open view into the next prompt; logged lines go to the operator and return values are discarded. The system prompt names only module briefs, so an agent discovers functions by searching, and a per-agent opening turn seeds the fresh window with a module listing, chosen documentation views and optionally a workspace tree. Every SDK function is compiled and callable whatever the run granted, and a withheld call fails at run time as unavailable, naming the call to make instead. Sessions end only through role-shaped calls: gg.session.finish, or a reviewer’s approve and requestChanges.

Four params are required per RaC profile: language, timeoutSecs (the guest execution ceiling, enforced by wasmtime epoch interruption and excluding time parked in host calls), maxMemoryBytes (the linear-memory cap) and docViewTypes (independent return, parameters and errors flags; gg refuses an absent or partial set, and the console seeds a fresh capability with return and errors on and parameters off). Guests get the full WASI p2 surface with / preopened, stderr is captured with 4 KiB kept at each end, GG_SANDBOX_DEADLINE_MS is exported, and the ECMAScript guest shadows setTimeout, fetch and their kin with named throwers. Composed calls stream as ordinary ToolCall/ToolResult pairs with ids program:{ordinal}:{tool}. See responses as code, programs and views.

The program-library capability keeps the source of every program gg ran under a cuid2 id minted when the submission is acknowledged, and the acknowledgement’s body is that bare id. programs.get(id) takes the id as a required argument, a miss is not-found naming the ids held and the turn each ran in, programs.rerun(source) lets a program patch and hand over a predecessor, at most four programs per submission (the model’s own plus three hand-overs), a fork clones the library and an exec or FSM edge moves it. A per-agent idLength param (default 4, range 2 to 32) sets the id length. See the program library.

Eleven program languages, each with a hand-written SDK

Section titled “Eleven program languages, each with a hand-written SDK”

The language param is a required per-agent configuration axis with eleven registered values. typescript is type-checked by tsc, erased and evaluated as an ES module by the quickjs-ng ECMAScript guest; javascript uses the same guest with no compiler; python runs on a guest carrying CPython 3.14; ruby is compiled to JavaScript on the host by an embedded Opal and evaluated by a guest carrying Opal’s runtime; purescript goes through purs and a bundle to the ECMAScript guest; java and kotlin are compiled by javac or kotlinc and then TeaVM inside a pooled warm JVM into one wasm component per turn; rust is rustc to a component per turn; swift is swiftc; cpp is clang++ with every #include written by the program; and csharp is Roslyn to IL, run by a guest holding Mono’s interpreter.

All arms bind one language-neutral WIT world (crates/gg/wit/gg-sandbox.wit, one interface per family) so they differ only in spelling. The surface is 52 operations in thirteen modules (files, shell, board, tasks, memories, docs, views, context, delegation, skills, programs, session and core), each bound by a gg capability and the agent’s allowlist, by the agent’s ending role, by its machine state, or always. Every SDK is static: every function is compiled, linked and callable in every program, a call the agent was not granted is refused at the membrane with a typed ApiError of code unavailable whose message names the alternative in the program’s own language, no gg name is reachable without an import line the model wrote, and every call is synchronous and typed rather than JSON. A failure reaches the model in the language’s own words at the model’s own line.

Signature catalogues and guest components are generated by the build, reflected from each SDK’s own documentation comments by the language’s own documentation generator, and never committed. The run’s root language lands on summary.programLanguage and each instance’s on its agent_surface event, so a cross-language study slices on one scalar. See languages, static SDKs and the ECMAScript guest.

gg’s prompts name no catalogued function. The docs family (gg.docs.search and the calls that open and close documentation views) is published on all eleven arms in each arm’s own spelling, search indexes parameter names and descriptions so a query for an argument finds its function, and search reaches three kinds, module, function and type, with a module leading a tie. A module’s view is its path, the line a program writes to bring it into scope, its prose and a line per function this agent binds, and every function and type view carries a line stating where the symbol is defined and the exact import line that reaches it. Search and views are filtered by the same predicate the membrane refuses calls with, so what a model can find and what gg will service are one set.

Every arm’s files module gains tree (and a tree tool under list-dir for tool-calling agents): a depth-bounded rendering of the directory beneath a root, depth 2 by default and at most 10, honouring the same ignore files search does, bounded at 1,000 entries and 16 KiB. A RaC agent’s opening turn can open a tree at a configured depth, so a program reads the layout instead of guessing at a path. See the API surface and the filesystem.

Compilation is measured and isolated, and a compiler crash is gg’s fault

Section titled “Compilation is measured and isolated, and a compiler crash is gg’s fault”

Each arm declares a checker (tsc, opal, purs, javac, kotlinc, rustc, swiftc, clang++, csc, or none for JavaScript and Python), interpolated into the system prompt. Compile time is recorded apart from the sandbox clock as compileMs on every code_execution event and summary.compileMs run-wide, charging a turn for its program, every hand-over, and the code skills, memories and on-use scripts it was first to load. A compiler that rejects a program yields a transpile_syntax, transpile_compile or transpile_unsupported error turn with the diagnostic verbatim, bounded by deletion only: eight diagnostics on most arms closing with a dropped count, and four errors with three notes each on C++. A compiler that could not finish (a crash, its own timeout, a binary missing from the image) tells the model nothing, counts against no ceiling and ends the run under internal_error.

Every agent holds one private compile workspace for its lifetime (a working directory, an output directory, a redirected HOME, XDG_* and TMPDIR). A response is compiled once, a loaded code module is compiled once per session under its binding key and linked from recorded output on later turns, and a source-level gate forbids Command::new and temp_dir outside the seam. A warm-up compile of each language’s guest component fires once per run concurrently with the first model request, and long-lived compiler JVMs are pooled. Compilers gg carries itself (tsc, Opal) run through Node (TCAB_GG_NODE); the rest come from the run image’s toolchain tree. See compilation.

Every run image gains a gg variant, gated by gg selfcheck

Section titled “Every run image gains a gg variant, gated by gg selfcheck”

A gg run resolves the <name>-gg variant of whatever image the test case would otherwise use: the parent image plus one COPY of the toolchain tree built by containers/gg-toolchains/Dockerfile at /opt/gg. containers/image-names.sh holds 55 names, 28 parent images and 27 variants, derived by suffix rather than listed, so a gg run of any test type resolves an image it can compile in while no other harness’s image carries the toolchains. Each toolchain carries every shared library the run images do not supply (the C# tree vendors the ICU trio, the wasi-sdk tree vendors libtinfo), because the variants form four distinct library environments.

gg selfcheck [--language <id>] drives every arm’s real bootstrap turn in whatever environment the binary runs in and exits non-zero on any broken arm. containers/build.sh runs it inside each built variant before pushing, scripts/ci/run-images.sh requires each representative’s gg selfcheck ok: line by name, and make -C deployments/local run-images-gg-selfcheck does the same locally. core::harness::resolve_run_image now takes the harness slug and RUN_IMAGE_OVERRIDE_ENVS is a LazyLock<Vec<String>>, which breaks in-repo callers only. See selfcheck.

The gg model client chooses the provider for every request

Section titled “The gg model client chooses the provider for every request”

gg talks to every live model through OpenRouter but chooses the provider itself. The backend builds an ordered candidate list per bound model at enqueue (native quantization, priced at or below the developer endpoint, cache-read priced, supporting tools, tool_choice and reasoning as needed, not banned; the developer first, then by fault rate and price) and every request names exactly one candidate through provider.only, provider.quantizations and allow_fallbacks: false. The run moves to the next candidate when the retry schedule is spent, at once on a 404 (a provider excluded by account privacy settings), or after providerCacheMissLimit (default 2) unexpected cache misses; each move is a provider_switch event and each stall or miss a provider_fault. A reply served by any other provider ends the run as model_provider_mismatch.

gg mints one cuid2 routing key per run and sends it as both session_id and prompt_cache_key on every request of every agent. For the Anthropic family it stamps up to four cache_control breakpoints (an anchor, two grid-snapped rolling points and the tail) and serializes every message as a one-element content array so a moving marker never changes wire shape; other providers get no markers. Each agent picks a prompt-cache lifetime (five minutes by default, or one hour) and an optional reasoning setting (effort from xhigh to none, or maxTokens), sent as OpenRouter’s unified object and never applied to a handoff summarizer’s model. Every request streams with usage included.

Failed calls (429, 5xx, transport, stall) retry on a configurable schedule: maxModelRetries (default 10) against modelRetryMaxDelaySecs (default 60), starting at one second and doubling, with Retry-After honoured. modelCallTimeoutSecs (default 900) bounds one attempt and yields a model_timeout error turn, and modelStreamIdleSecs (default 60) cancels a stream that delivers no delta. A refused pinned tool_choice is re-sent on auto and remembered run-wide, and requests carry gg’s own name and site as OpenRouter attribution. See execution limits.

Loop detection abandons a reply that has started repeating itself

Section titled “Loop detection abandons a reply that has started repeating itself”

A per-agent loopDetection setting watches delta.content as it streams and drops a reply when minOffenders distinct words each exceed repeatThreshold occurrences in a rolling windowWords window for minSaturatedRun consecutive words, or when it passes maxResponseChars; the console seeds 256, 32, 2, 3000 and 250,000. The sustained-run term is what keeps a 1,500-entry tilemap literal from tripping it. A trip is handled like a 5xx inside the client’s retry loop: the stream is closed unread, a fresh detector judges the retry, and nothing of the discarded reply enters the context or the session record. A model that loops on every attempt costs an error turn (model_response_loop), not the run.

Discarded attempts ride the turn_outcome event as loopAborts, loopAbortWords and loopAbortChars and roll up on the summary. Because a cancelled stream never delivers usage, gg keeps each abandoned reply’s generation id and, once at session end, looks it up on OpenRouter’s GET /api/v1/generation ledger (concurrently, a 404 retried for up to 30 seconds) and adds the price to the slot’s slotCosts and the run’s total cost. loopAbortUnpriced counts the ones it could not price, so a deployment whose egress blocks that endpoint records them rather than failing the run. See loop detection.

Execution limits, one error definition and cancellation

Section titled “Execution limits, one error definition and cancellation”

capabilitySet.limits carries five ceilings that are armed only when written: maxTurns per agent (exhausted), maxRuntimeSecs and maxCost run-wide (timed_out and limit_exceeded), and maxConsecutiveErrors and maxErrorRate with errorRateWindow per agent (limit_exceeded). Two keys are required, maxParallel (the run’s global agent pool) and replayMaxBytes (the capture journal bound), and five carry defaults: modelCallTimeoutSecs 900, modelStreamIdleSecs 60, providerCacheMissLimit 2, maxModelRetries 10 and modelRetryMaxDelaySecs 60. A fresh configuration seeds five consecutive errors and a 0.2 error rate over 50 turns, and the turn, runtime and cost ceilings start unarmed. Nothing about a ceiling or the detector is ever told to the model.

A turn is an error only when its declared work could not be carried out: a failed, timed-out, length-capped, unparseable or looping model call, a program that did not compile or threw uncaught, a sandbox stop, or a tool-calling turn that requested nothing. A failure reported back into a program that carried on is not. The judgement is made once per turn at the seam where an outcome is recorded against an agent’s ceilings, emitted as turn_outcome with a base error kind and one of twenty-one specific errorType leaves, and rolled up on the summary as errors. Failures of gg’s own machinery disqualify the run under internal_error with exit 1, releasing every suspended wait, and a refused credential is auth_error.

The host cancels a run by creating the cancelFile path named in the invocation. Every agent checks it at its turn boundary, the epilogue (rollups, session record, summary) runs in full, and the status is canceled. See execution limits and turn outcomes.

Every gg run tags each window contribution with one of fifteen sources (System, User prompt, Assistant, Tool output, Compiler errors, Runtime errors, File views, Agent views, Documentation, Doc search, Skills, Memories, Task list, Board and History), estimated with one BPE tokenizer, and streams the breakdown every turn. The denominator is the bound model’s window from the backend model catalog, fetched live from OpenRouter for an unseen model and pushed in as modelWindows; there is no default and no guess, and backend, core and gg each refuse a run without one. The context-window-override capability’s windowLimit is a ceiling, measured against the smaller of it and the model’s window, which lets one configuration exercise compaction across models of different sizes.

The prompt is append-only to preserve provider prefix caching. Rebuilt state blocks (memories, tasks) keep their position when unchanged and are superseded rather than removed when changed, the system prompt and the context-usage signal are slots, each read_file appends a snapshot while re-opening a view selector supersedes it, and openDocsView of an already-open key does nothing. read_file on a PNG, JPEG, GIF or WebP (detected by magic number) returns the picture; a model declared without image in modelModalities is never sent one, an unknown model is tried and denied run-wide on refusal, and an inline image is charged by its dimensions.

The message log fingerprints every message and streams its body once as a context_message event, with each turn a prompt event holding ordered pointers into that pool plus the reply, the serving provider, finish reason, usage and latency; images are logged as descriptors, never bytes. The console renders a per-agent Requests file from it, and the Agents panel’s context spend folds it turn by turn: each turn’s input tokens and their cost are split across the request’s messages in proportion to estimated size and summed per message, band, view and tool, so what a file cost to keep in the window over the whole run is ranked rather than what it cost to bring in once. See context visibility and context spend.

The per-agent compaction capability summarizes a filled thread and restarts it, carrying pinned state verbatim: used skills, tasks, memories, locked autoloaded specs, re-derived documentation views, and files a compact call named. summaryHeadroom (0.0 to 0.9, required) is held back from the window for the summarization round trip and sets the trigger at 1 - summaryHeadroom. The implementation names one of five strategies: the in-loop self-summarization, self-compaction (the agent calls compact with a summary and up to twelve files to re-read) and memory-compaction (working state must be written as memories, so a writable memories capability is required), or the between-turn handoff-summarization and handoff-compaction on an optional model param that may defer to one of the agent’s slots.

A boundary that leaves the window still at the trigger is retried maxRetries times and then ends the agent under the failure status compaction_failed. A failed condensation restarts from a fixed note flagged as a fallback. Each boundary records the strategy, the trigger fullness, tokens before and after, per-source composition and the summary, and fires the pre- and post-compact hooks. The companion agent-managed-context capability renders a per-turn context-usage signal and grants evict_file_view, archive_thread, search_archive and, under RaC, gg.views.close; docview-close separately grants gg.docs.close and closeAll. See compaction and agent-managed context.

The shell capability runs sh -c in the workspace, or the agent’s worktree, in its own process group under a per-call timeout that defaults to 3600 seconds, clamped to 24 hours and to the run’s remaining wall clock; a non-zero exit is a result, not a failed call. Its implementation is inline (the whole output, capped at 16 KiB) or offload: both streams of every command are written to /tmp/gg-shell/cmd-*.stdout and .stderr, only the last maxLines and maxChars return inline, and a truncation note names the files with each one’s line count, p50, p95 and p99 line lengths and five longest lines so a model can aim a windowed read. One shell telemetry event per command (origin tool, program or hook, directory, exit code and 16 Ki-character tails) backs a Shell file per agent in the console.

The filesystem tools are one capability each. read-file is unlimited or default-cap with a required lineCap, pages with offset and limit under a [showing lines a-b of N; continue with offset: c] footer, backstops at 256 KiB, and reads images. write-file, edit-file (exactly one occurrence) and list-dir (list_dir plus tree) complete the set, and the default-on search is a Rust-syntax regular expression over the workspace honouring .gitignore, .ignore and .git/info/exclude, 50 hits by default and 200 at most, with lines clipped at 200 characters. Paths resolve as any process in the container would, relative against gg’s working directory and absolute as themselves; the container, not the workspace, is the boundary, so the shell’s own offloaded output is readable. See the shell, the filesystem and shell commands.

The skills capability loads a per-agent catalogue from a dir param (.gg/skills by default, stood up empty by gg) and lists every skill by name and description in the system prompt; read_skill brings one in, and a used prose skill is retained across compaction. Under RaC a skill or a memory may also carry a code module, compiled once when first used and importable as lib:<name> by every later program (import * as csvTools from "lib:csvTools" in TypeScript), plus an on-use script that runs once after the turn’s program and opens documentation views of the module’s exports. Loaded code costs no tokens, compaction never touches it, and a module that fails reaches the model as a module error naming the binding key rather than as the program’s own failure.

gg ships twelve generated built-in skills, one per function family (gg-filesystem, gg-shell, gg-project, gg-tasks, gg-memory, gg-skills, gg-context, gg-delegation, gg-docs, gg-views, gg-programs and gg-session), selected per agent by a withholding builtIns object where {} offers all and an authored skill of the same name wins. The memories capability stores model-curated notes inside gg under one of three strategies (scratchpad, markdown, keyword-search) with six required limits (maxCount, maxLenPerMemory, maxTotalLen, maxLenIndex, maxLenDescription and maxResults, 0 disabling one), refuses rather than truncates a breaching write, supports isolated, inherited and read-only scoping and linked instances between agents, and records every revision. See skills, memories and modules.

Every profile carries a roster of the profiles it may use, each entry scoped subagent, implementer or reviewer. subagents grants spawn_subagent, wait_for_subagents and send_message bounded by maxDepth; all agents share one runtime scheduler under the run’s maxParallel, where a suspended agent releases its slot and resuming beats starting. agent-persistence makes a profile a single long-lived worker whose next instance re-opens the previous one’s file, text and documentation views. exec replaces the running instance with another profile, transferring its modules (history, memories, tasks, board, skills, archive) live, and fork runs a private copy. See subagents, agent persistence and fork and exec.

fsm profiles are model-less state tables authored in the console, each state running another profile and each transition naming the modules the next state inherits. project-management gives the run one board of epics (three- to six-letter prefixes) and issues (AUTH-1, ISSUE-2) with scoped briefs, a blocked-by DAG, auto-dispatch to implementers named AUTH-1.0i in isolated git worktrees merged by a merge agent, reviewers AUTH-1.0i.0r whose approve and requestChanges verdicts gate acceptance, and wait_for_issue; the board is reachable through tools only and costs no context. tasks is a per-agent blocked-by DAG (simple or issues mode, maxTasks) pinned in the prompt and carried across compaction. See FSMs, project management and tasks.

Hooks run operator commands or scripts at ten lifecycle points (two session events on the set, eight per agent). A hook can block a write, a shell command or a stop, or put text in front of the model, and three ship built in: trace, refuse-empty-write and guard-destructive-shell. See hooks.

Prompts as templates, autoloaded specifications and the Reference

Section titled “Prompts as templates, autoloaded specifications and the Reference”

Everything gg says to a model is a Handlebars template under crates/gg/templates/ embedded in the binary: one system prompt per execution mode assembled from per-capability partials and a per-language segment, the pinned task and memory blocks, dispatch briefs, and in-loop messages. Each profile may add custom instructions or override the whole template. The autoload-specs capability seeds an agent’s opening context with every file the test case provided, specs in seeded order and then reference images as pictures when images is on and the model can see them, synthesized as well-formed reads and optionally locked across compaction. crates/gg/prompts/ ships operator-pasted system prompts for baseline, layered and studio dev-team configurations plus common agents. See prompts and autoload specifications.

gg reference [--out DIR] projects the exact tool definitions and every arm’s documentation views, which the backend serves at GET /gg/reference from TCAB_GG_REFERENCE, so the console’s Reference section (gg → Reference, with Tools and API tabs) cannot drift from what a run sends: all 37 tools with their live schemas and, per arm, every module, function and type with its documentation view body. gg probe-fixtures writes the per-language turn-1 fixtures the backend’s model probes replay. See the reference.

An always-on session record, salvaged from a hung run

Section titled “An always-on session record, salvaged from a hung run”

Every gg run appends a JSON line per event to .gg/replay.ndjson inside the container (model I/O, tool results, the prompt frame, shell and git subprocesses, cancel probes and clock reads, per agent), bounded by the required replayMaxBytes limit. The host folds it into the run tree’s replay.json.gz after the run and mirrors it to the backend at GET /runs/{id}/replay. Capture is not configurable. Before tearing down a run that hung or outran its cap, the engine copies the journal alone out of the container and assembles it with truncation reason session_killed, so the one outcome that collects nothing can still be explained. See the session record and session records in analysis.

gg’s telemetry stream, and the console built from it

Section titled “gg’s telemetry stream, and the console built from it”

gg emits a richer stream than the shared event contract, and the backend and console read it natively, live and after the run. It opens with session_started, announcing the resolved capability set and routing key, and carries per agent agent_spawned, agent_modules, agent_surface (execution mode, language, bound module stores and the tools or API functions offered, each with the operation it serves), usage deltas per model call, turn_outcome, turn_timing, context_message and prompt, context_breakdown, code_execution, shell, tool_call/tool_result and api_call/api_result, provider_fault and provider_switch, response_rejected, limit_exceeded, slot_usage, compaction records, and one session_summary immediately before session_ended. Every usage delta is attributed to the model and profile that spent it.

A live run is watched at /runs/gg/:jobId/live and a finished one read at /runs/:runId/gg, both rendering the same panels over one reduction of the stream. The Dashboard carries status (with the orchestrator’s setup stage named before gg’s first event), the token and cost tally with caching and reasoning rings, turns, tokens per second, an errors row, four clock tiles for runtime, limit, active and waiting, and per-profile and per-model spend bars. The Agents panel sums each configured profile’s instances; the Instances explorer is a per-agent filesystem (overview, prompt, tools or apis, context, requests, metrics, shell, compaction, a programs file for RaC agents and a modules folder) that renders successions, forks and machine states as one lineage; the Modules view reads a run by the stores it holds; and the Project tab is the board with issues rendered as Markdown. A dropped live stream resolves the run’s outcome or re-joins instead of surfacing a raw TypeError.

turn_timing partitions each turn’s wall clock into promptMs, requestMs and responseMs, which the console renders as a stacked bar per turn heading an agent’s Metrics file, windowed to about fifty bars, with four per-request line graphs beneath it: tokens per second, cost per request, cache-read share and reasoning share. code_execution reports per program ok, toolCalls, apiCalls, durationMs, the error, the capped logs, undocumentedCalls (calls written without an earlier documentation view) and the two compile figures. A RaC agent emits an api_call/api_result pair per call its programs make, naming the operation and no arguments, so a function the model never used reads a real 0x in the console. See the telemetry overview, the console, turn timing, code execution and the agent surface.

Two cost figures, and a bounded usage split

Section titled “Two cost figures, and a bounded usage split”

A gg run’s session summary carries its spend as two figures, per (profile, model) slot on slotCosts and run-wide as cost and workCost. The total cost sums every request that reported a price; the work cost sums only the turns whose reply produced a program gg ran or a tool call gg dispatched, decided by the same rules dispatch applies. Every usage delta names the figure its turn fed in a figure field, so summing the deltas marked work reproduces the work cost and summing all of them reproduces the total, and the maxCost ceiling reads the total. A length-capped reply is rejected whole, recorded on a response_rejected event and summed in rejectedResponses, inside the total and outside the work cost, and the compaction summarizer’s own calls stay outside both.

Each usage event carries the four normalized token classes, the call’s cost when the provider reported one, the serving provider, and wire, the provider’s usage object verbatim, so a disagreement between the provider’s numbers and the recorded split can be read off the row. gg bounds the provider’s output/reasoning split by the reply it received: output is at least gg’s own estimate of the reply’s text and tool calls, reasoning is the remainder of completion_tokens, and a row whose split gg had to rebuild carries reconciled: true. The summary’s providerStats slices calls, tokens, cost, turns, working turns, errors, stalls and cache misses per (provider, model), and its maxResponseChars and maxResponseOutputTokens fold over every turn except a length-capped one. See usage.

A run’s session is measured apart from its setup

Section titled “A run’s session is measured apart from its setup”

runTimeSeconds began before reference rendering and host seeding, so the figure described the host more than the model. Every run record now also carries metrics.sessionSeconds (the harness session alone, the figure that describes a model), setupSeconds (through container start, probe, harness install and the case’s init, with the cluster queueing wait subtracted), teardownSeconds (the remainder, so the three sum exactly to runTimeSeconds) and validationSeconds (outside the run time, null on a canceled run). Each stage is null on a record written before the stages were measured, distinct from 0.

Because every run of a model is now pinned to the model’s own provider, the session duration is a property of the model, so the test case Metrics tab charts it with the same bar and scatter control as tokens and cost, the Leaderboard tab tabulates its mean, the model overview folds it, the comparison detail page summarizes it per arm with median, spread and mean, the gg Discover dashboard reads metric.sessionSeconds, and the public site’s case and model pages show the same figures. The run’s Metrics page lists the four stages, and a run with no session figure contributes nothing. See metrics.

Comparable cost is priced at the model’s list price

Section titled “Comparable cost is priced at the model’s list price”

Every model’s catalog entry carries an operator-entered list price: uncached input, cached input and output rates per Mtok from the developer’s own pricing page, and the date taken (listPriceInputPerMtok, listPriceCachedInputPerMtok, listPriceOutputPerMtok and listPriceAsOf on POST and PUT /models, saved together or refused with 422). The backend resolves the list price at enqueue and stamps it on the launch, so a run is scored at the price its model carried when it was queued, and the comparable cost is derived from the token classes and that price with reasoning at the output rate. A class carrying tokens with an unknown rate makes the cost unknown, and a run enqueued before the catalog carried list prices records an unknown comparable cost. In v0.6.3 the comparable cost was computed from the per-token prices OpenRouter listed at run completion.

The actual cost is the harness’s own figure where it reports one (Claude Code’s total_cost_usd) and equals the comparable cost otherwise, and every harness metrics page was rewritten onto this rule. Beside the list price the backend keeps a billed-rate history observed from the model’s official OpenRouter endpoint, recorded on run completion, on a 24-hour refresh, and missing-only when a model is saved, first enqueued or at backend startup; the model’s Stats tab shows list price and billed rate side by side with their difference, and the billed rate never rewrites what a run is scored at. GET /models/openrouter?slug= seeds the three rates from the official endpoint for the operator to confirm. See metrics and the backend API.

The catalog is the single store of model facts

Section titled “The catalog is the single store of model facts”

The new-model form leads with the OpenRouter slug and a Fill from OpenRouter control that fills the display name, provider, description and the list-price seed figures from the model’s endpoints listing. Each billed-rate observation now also records the model’s context window, release date and input modalities. A gg run’s context window is resolved from the catalog at enqueue for every bound model and pushed onto the launch, fetched live from the per-model endpoint when the catalog has not observed the model, and a launch whose window cannot be resolved is refused per index in a batch. Input modalities travel the same way so a text-only model is never sent a reference image, and the model page shows them under Specs.

On top of the developer pin the backend builds each bound model’s provider candidate list at enqueue from the endpoints listing, filtered by the catalog entry’s policy (a hand-set native quantization, a price ceiling used when no developer endpoint is listed, banned providers, and providers accepted despite unknown quantization) and ordered by the fault rate recorded across every stored gg run. A model whose list comes out empty refuses the enqueue naming the filter that emptied it. GET /models/{slug}/candidates shows the list the next enqueue would build, the model edit page exposes the policy fields, the candidate table and the cache-miss limit, and provider names compare case- and punctuation-insensitively. See adding or updating a model.

A model probe is a responses-as-code readiness check run from the model’s Probes tab (/models/:modelId/probes). The backend replays gg’s RaC turn-1 request against the model through OpenRouter with the submit_program tool offered and forced, over two scenarios (baseline and missing-docview) across several prompts, per program-language arm or across every arm, and reduces the results to a ready or not-ready verdict at 80% of scored calls per group. Probes are append-only dated history in model_probe and model_probe_item, never reach the public snapshot, and run inside the backend process, so a restart fails a running probe. The endpoints are POST and GET /models/{slug}/probes, GET /model-probes/{id} and GET /models/{slug}/probe-providers.

The Models section gains a Providers tab (/models/providers) fed by two open reads. GET /stats/providers reports per provider and model the calls, tokens, cost, length-capped rejections, stalls, unexpected cache misses and turn error breakdowns from stored gg session summaries, kept apart from probe evidence, and GET /stats/model-accuracy reports per-model RaC turn validity and tool-calling dispatch outcomes, with approximate older folds marked. See probe a model.

gg launches from the New Run form under a saved configuration

Section titled “gg launches from the New Run form under a saved configuration”

gg is registered as a first-class harness subject (harness_slug: gg, routed through OpenRouter, outside the third-party catalogue). A gg run is launched from a saved configuration: gg_config rows per account with REST CRUD under /gg/configs, edited from the account section’s gg Configs tab at /account/gg, while POST /gg/runs enqueues a run whose capability set is lifted onto job.gg_config_json and whose configuration name and id are lifted onto the run row so the run log’s MODEL / CONFIG column, its search and its model sort agree. Agent profiles are authored once in a library (gg_agent, /gg/agents, the gg Agents tab), and a configuration imports one as a live reference with per-field overrides recorded beside its resolved set. Each profile carries a stable id and its name is display text.

The New Run form’s gg model picker is scoped to the OpenRouter family, a launched gg run is tracked in the runs list immediately, and the configuration editor exposes the run limits, per-agent reasoning setting and prompt-cache lifetime, an Agent type selector (Tools, RaC or FSM), an Opening Turn tab that chooses from what the APIs tab enabled, an FSM state editor, module ownership and memory scope, and hooks. A responses-as-code agent’s profile must carry an openingTurn naming the modules the seeded first program lists and the functions whose documentation it opens; a module is listable only when the agent holds at least one of its functions, an unknown or role-bound entry refuses the launch, and an empty opening turn seeds no program. See configurations and agents.

Coverage plans and ladders schedule gg configurations

Section titled “Coverage plans and ladders schedule gg configurations”

A plan’s, ladder’s or kind = "combo" group’s member is now one of two shapes: a harness combination as before, or a gg configuration the account has saved plus a model for each launch slot it declares. A gg cell is identified by the configuration’s id and the models the bound set runs on, lifted onto run.gg_config_id, job.gg_config_id, run.gg_models, job.gg_preset and job.gg_models, with a one-time startup attribution of older runs recorded in the new backfill_state table. Every enqueue path refuses a configuration the launching account does not own, and the runs listing gains a ggConfigId filter. A top-up lowers a gg member through the same code POST /gg/runs uses, so a scheduled and a hand-launched run cannot drift, and a member whose configuration was deleted, leaves a slot unbound, binds a model with no window or list price, or imports a saved agent edited since is reported unlaunchable and skipped. gg gains a harness parallelism lane.

A pinned case now names an engine after its variant (ladder_rung.engine and job.engine_slug; a pin with no engine means the engineless run, so stored plans keep their counts). The review buffer target is a tagged shape, {kind: bounded, runs} (clamped to 500) or {kind: unbounded}, on the account setting and the per-plan or ladder override, so running through everything is a native instruction. The plan page is three tabs under one layout route (/account/coverage/:planId, /reviews and /tests): a Dashboard of tiles and rings, the review queue as a worklist, and the matrix with a one-run and a whole-shortfall press per cell. The Pause/Resume button and auto-top-up checkbox become one Auto top-up switch (ladders get an Enabled switch), and the plan and ladder editors pick a case through the shared Test grid. See coverage plans and ladders.

A comparison holds a case coordinate constant and runs two or more arms, each a configuration of its own (a third-party harness plus its model, or a gg configuration plus a model per slot), so “gg against Pi” is statable. Core adds deterministic descriptive statistics (median, mean, quartiles, a bootstrap confidence interval on the median seeded from the arm’s sorted run ids, a Wilson pass-rate interval, a two-arm median ratio), an automated-only checklist score, per-arm aggregation and confound detection, plus two record captures every harness parser contributes: per-turn usage events (Pi, Kilo, OpenCode) and tool-call counts including consumed todo tools (RunRecord.tool_calls).

The backend stores the configuration in a comparison table with per-arm statistics computed on read (GET, POST, PUT and DELETE /comparisons and /comparisons/{id}), validates the sample size, arm set and engine, and POST /comparisons/{id}/publish marks it published, enqueues a publish job per publishable arm run and emits comparisons.json plus comparisons/<id>.json into the snapshot. The console has create, edit and detail pages (/comparisons/new, /comparisons/:id, /comparisons/:id/edit) with box-plot distributions labelled with n and a Trigger missing runs action that tops each arm up, and a Comparisons tab in the Runs section rendered read-only on the static site from the snapshot. The per-case Metrics and Leaderboard tabs split series by (harness, model) instead of merging harnesses, gain an Average points chart and a shared sort control, and exclude gg runs, which span several models. See comparisons.

TCQ: the gg query language, Discover, saved queries and dashboards

Section titled “TCQ: the gg query language, Discover, saved queries and dashboards”

The gg result-aggregation widget builder and POST /gg/aggregate are replaced by TCQ, a query language over a document model. Every gg run is one flat map of dotted typed fields (summary.*, cap.<id> totals over the catalog, sparse tool.<name>, metric.*, code.*, agent.<id>.*, model) and a query is a filter tree optionally piped to | stats <aggs> by <keys>. The backend serves it from an in-memory index of one document per gg run, reconciled per id against run.updated_at (a mutation timestamp every run mutator stamps) and invalidated when an ingest re-resolves checklist weights: POST /gg/query, POST /gg/query/batch (one read per dashboard) and GET /gg/fields, each result capped at 1000 rows and flagged truncated.

The console’s Discover page (/gg/query) is a highlighted editor with a field sidebar carrying document counts, worked examples, a document view that opens each run, and charts chosen from the query’s shape. Saved queries and dashboards (gg_saved_query and gg_dashboard; /gg/saved, /gg/dashboards/:id) store query source text per account; the /gg/aggregate widget builder and its URLs are gone, and Discover is the only address the analysis section answers to. The evaluator is mirrored in TypeScript and held to the crate’s own conformance fixture, which is what lets the public site run the same Discover page in the browser over gg-runs.json, a redacted export of every gg run’s document that is not gated on publication and fails closed on a case it cannot classify. See the query language.

Static code analysis of the tree a run produced

Section titled “Static code analysis of the tree a run produced”

A new crates/code-analysis crate reads the produced tree without executing anything: a .gitignore-honouring walk with a floor for the host-owned .tcab/ namespace, an authored set derived from the run’s seed commit (now recorded as RunRecord.seed_commit), complexity scoring through oxc and syn behind a derived-stack parser guard, a module graph with cycle detection, a clone detector and a two-tier rollup. It runs at a new post-run stage seam (after collection, before validation, outside the runtime cap, and on a canceled run too), writes the bounded summary onto RunRecord.code_analysis and the unbounded document to {run}/code-analysis.json.gz, mirrored into the backend store and served by GET and POST /runs/{id}/code-analysis. run.code_analyzer_version records the analyzer generation (currently 1) so a mixed corpus stays sliceable, and NULL means never analyzed.

The run’s Code tab (/runs/:runId/code) shows the provenance strip, a figure table driven by the CODE_METRICS catalog, a treemap explorer with per-file coverage, outlier rankings and cycles, plus the executed test and coverage bands; the analyzer’s own diagnostics are headed Analysis notes and its static test counts Test authorship so neither reads as executed coverage. tcab analyze <dir> [--seed-commit] [--tree-basis] [--top] [--json] runs the same pass over any directory. The snapshot lifts code lines, size Gini and mean cognitive complexity onto every run’s summary card and publishes the document under media/runs/<id>/code-analysis/v<gen>.json; the run log’s CODE column reads “not measured” rather than zero and is off by default. See code analysis.

Engines: a run selects the runtime its game is built on

Section titled “Engines: a run selects the runtime its game is built on”

A run now carries an engine beside its test case, variant, harness and model. The catalogue is closed and embedded in core from engines/<slug>/engine.toml: none (the default; the build supplies its own frame loop, input, audio, assets and diagnostics), simple-2d, structured-2d, simple-3d and structured-3d, each a @clockwyrks/<slug> npm package under packages/<slug>/ staged into the host package store and vendored into .vendor/engine/ at seed time, with its dependency written into the seeded workspace’s package.json and its documentation seeded at engine/ in the run root.

A case declares support per version with the engines = [...] root key and one [[engine]] table per runtime engine carrying a required inclusive min_version and an optional exclusive max_version. The engine’s version is read from the staged package at seed time and recorded on the run beside the slug, and a run whose engine is unsupported or whose version falls outside the range is refused before any container work. tcab run, seed and validate take --engine (default none), tcab engines lists the catalogue, every enqueue endpoint carries engine on its launch body, and the console’s New Run form offers the engines the resolved version supports. A game jam may declare engines as an optional key defaulting to ["none"].

The engine travels the whole system. The run record lifts engine_slug onto the run row, the run header names the engine beside the variant, the run log gains an ENGINE column, and GET /runs filters on engine, versions and testCases. A case’s detail page is anchored to one coordinate (version, variant, engine) carried in the URL, with the run aggregations scoped relative to it and widening across engines listing each engine separately, and a version’s prompt and specs render for a named engine (GET /test-cases/{slug}/versions/{version}/specs/{variant}?engine=), so a run’s Inputs tab shows what that run was given. See engines in core and the engines overview.

@clockwyrks/simple-2d (1.0.0) owns the frame loop and the delta time it hands the game, the letterboxed DPR-aware canvas fit, named input actions over keyboard bindings and a closed touch-layout catalogue, a synthesized audio cue bus with looping cues and the browser’s first-interaction unlock, asset resolution under a fixed root, a diagnostics overlay and an opt-in draw-command recorder; the game writes its own simulation and drawing. The engine holds the game’s state by value: update receives DeepReadonly<S> and returns the next state, initialize returns [state, debug] so the debug surface is handed back off engine.debug, engine.apply(transition) poses a game between frames, and Engine.diagnostics() reads back the registered sources. Its clock is an object, so the engine runs without a document.

@clockwyrks/structured-2d (1.0.0) adds a gameplay framework the game is written inside (worlds built from levels, game modes, actors and components, pawns and controllers), engine-owned rendering through render components that draw in world space or in screen space, a one-option image-smoothing switch, collision detection and the same recorder, with the debug surface returned from the game instance’s initialize and poses acting on the live engine.world. Both 2D engines clip the game’s drawing to the logical field so nothing paints into the letterbox bars, name the pressed button on a bare pointerdown’s samples, and read an event with no pointerId as pointer 0. Each engine documents itself under /engines/<slug>/ in four sections (APIs, concepts, usage and validators), and that documentation is what a run is seeded with. See Simple 2D and Structured 2D.

Simple 3D and Structured 3D, rendered through three.js

Section titled “Simple 3D and Structured 3D, rendered through three.js”

@clockwyrks/simple-3d and @clockwyrks/structured-3d (both 1.0.0) are the 3D members of the two families. Simple 3D owns the renderer over the canvas, the scene object and camera it renders through and a 2D screen layer composited over the picture while the game populates the scene; Structured 3D adds the same worlds, levels, game modes, actors and controllers framework with engine-owned rendering and collision, plus model loading and animation. three is a peer dependency the build declares itself so the engine, the build and @clockwyrks/voxel-runtime/three share one instance, at Three 0.185 across the repository.

A 3D recording is a video of the frames the engine drew, encoded with WebCodecs as VP9 in WebM and timestamped in simulated time with a keyframe at least every 60 frames, so a player steps it frame by frame; the console’s replay player draws 3D draw-command documents through a three.js drawer and refuses an unknown space or format by name. A Structured 3D component’s opacity multiplies the alpha the game last wrote onto a game-owned material, and each placement of a loaded model gets materials of its own so two components built from one model fade independently. Gantry v1.0.0 is the case that runs on them. See Simple 3D, Structured 3D and 3D recording.

Validators decide the functional rating; reviewers rate aesthetics

Section titled “Validators decide the functional rating; reviewers rate aesthetics”

A run now carries two rating channels. The functional rating (flawless, great, passable, scuffed, broken) says whether the build works, is rated per domain, and is the worst across the run’s effective domains; the aesthetic rating (legendary, amazing, good, okay, slop) is one tier over the whole run for how it looks, sounds and feels. A case version on the per-engine manifest spelling, other than a game jam, is validator-rated: every graded point carries a validation script together with domains and a failure_cap (broken, scuffed, passable or great, never flawless), and resolution names the point that lacks one. A failing point lowers each of its domains to at most its cap, a domain’s rating is the lowest cap among its failing points, an undecided point lowers nothing, and the toolchain typecheck gate applies on top.

The functional rating and score are derived from the run record at push time and shown the moment the run completes, so a validator-rated run is publishable with zero reviews, which tcab publish says when no writeup is present. A review of such a run carries a single aesthetic tier, a writeup, and optional overrides of individual validator verdicts; its effective checklist is the validators’ verdicts overlaid with its overrides, the run’s score is the average effective score and its functional rating the worst effective rating across reviews, and the validators’ own figures stand while there are no reviews. A validator-rated run that did not complete is left unrated. Legacy versions on the single workspace spelling keep the reviewer’s rating as the functional rating, have no aesthetic rating, and reject failure_cap and domains by name.

Console and site render a functional badge and an aesthetic badge side by side, every capped point carries its cap as a rating badge, the Verdict tab (/runs/:runId/verdict) becomes one per-item browser with the reference-versus-run media and per-pane downloads (a replay is rendered to WebM), a per-domain rating strip appears on the public gallery too, and the About section gains a Ratings tab. @clockwyrks/run-stats gains AESTHETIC_RATINGS, FAILURE_CAPS, FAILURE_CAP_RATING, reviewItemsForEngine and the effective-checklist scoring. See end-to-end evaluation and results.

Validators are per-engine vitest projects that hold the build in process

Section titled “Validators are per-engine vitest projects that hold the build in process”

A case that supports an engine ships one validator project per engine under validation/<engine>/, a TypeScript vitest project separate from the build’s own config. A 2D engine’s suites build the engine over a @napi-rs/canvas surface in the test process, import the engine and the build’s own game module, install a scripted clock and step an exact number of frames, observing the engine’s values, its events and what it drew; an engineless project reaches the build through a browser via the shared harness. The runner stages validation/<engine>/ into the collected tree at validation/, plus the shared harness at validation/case-harness/, runs vitest naming the project’s config and, as file filters, exactly the suites the run’s variant’s checklist declares, so a suite no item of that variant names is never loaded and a variant left with nothing to run is refused. The staged project is removed afterwards and a directory that already stood at that name is held aside and put back, so tcab validate is safe on a committed reference.

A validation’s optional engines key names the engines a validator decides its point on, defaulting to every engine the case supports. A point whose validator does not cover the run’s engine is excluded from that run’s checklist entirely, is shown to no reviewer and carries no weight, and review_items_for_engine in core and reviewItemsForEngine in TypeScript give backend and console one answer; this is how the debug overlay’s toggle is graded under none alone. Every graded point of a v0.7.0 case is decided by a validator, and every .test.ts in a project is one a review item names. Failure assertions are stored as real expected and actual pairs with stack frames stripped. See validation and the suite.

A validator run stopped at its budget is inconclusive, not a failed build

Section titled “A validator run stopped at its budget is inconclusive, not a failed build”

Three outcomes are held apart from a failed verdict and leave a point undecided rather than synthesizing a failure: an unmet precondition (a validator skipping every check, including a browser or page the suite could not obtain), a check that never ran (no validator project for the engine, no vitest in the tree, a suite the project lacks, an unreadable report), and a run that exceeded its wall-clock budget. Each is recorded with its reason on the run record, an undecided point lowers no rating and carries no weight, and the console’s Verdict tab says which of the three reasons an undecided point carries.

The whole suite run is capped by TCAB_VITEST_TIMEOUT, 45 minutes by default and overridden in whole seconds by TCAB_VITEST_TIMEOUT_SECS; crossing it stops the suite’s whole process tree, is recorded as a fact about the host, and leaves the build’s score intact. A validator that could run but did not complete against a conformant build (a missing module, a call that threw, a malformed return, an undeclared output) still fails its point. See validation.

One shared validation harness: @clockwyrks/case-harness

Section titled “One shared validation harness: @clockwyrks/case-harness”

The browser-driven validator harness and canvas draw recorder that had been copied into every engineless case, and the per-engine harness machinery each engine project carried, are replaced by one workspace package, packages/case-harness, that every validator project on every engine is built on. It ships its TypeScript source with no build step, core stages it beside the case’s project out of the host package store with a checkout fallback, and it is deliberately absent from SHIPPABLE_PACKAGES so no case can vendor the tests into a run. The engineless half serves the built site, holds one Chromium and drives the build through the case’s debug handle; the engine half constructs an engine the case injects as values (createEngine, clock, driver, projection), so the package names no engine and a fifth needs no change.

It carries the readers cases share (coalesced text runs through drewText and drewTextAnywhere, pixels, colors, draw calls), the replay and still writers, an asset host (installAssetHost: fetch over the workspace, createImageBitmap, a structuredClone that carries decoded images by reference, canvas creation over real rasterizers with page-like font resolution) and installAudioContext with throwing and tolerant WAV decoders, a CDP touch driver a case opts into with hasTouch, a beforeLoad hook, and an opt-in audio gesture delivered before the opening reset. Its recorder binds to the surface a build draws into rather than the attached canvas it blits to, and images are pooled per run into the shared store. All 54 validator projects type-check under npm run typecheck:validators, which CI runs. See writing debug APIs and validators.

Engine recordings are the evidence a validator produces

Section titled “Engine recordings are the evidence a validator produces”

A validator captures each media output its review item declares in the form the manifest gives it: image (a PNG of the surface as it stands) or replay. For a replay it arms the engine’s recorder once its scenario is posed and disarms it once the behavior has happened, so the evidence is exactly what the build submitted over that stretch, with the diagnostics overlay outside the bracket. A 2D recording is a draw-command document in recording format 1, stored gzipped as <verdict>__<output>.json.gz and served as application/json with Content-Encoding: gzip: every frame carries the whole drawing state it inherited, values the context produced and the bitmaps drawn are held in tables the recording shares, and numbers are written to nine significant digits, so seeking to any frame costs the same and two recordings of one scenario compare frame for frame. A 3D recording is <verdict>__<output>.webm.

Bitmaps are written once per run into a shared image store (img.<id>.png, the id derived from the bytes) beside the recordings and travel every route declared media travels; RGBA pixel buffers stay inline. A recording is kept whatever the verdict, an output that was never written is recorded absent rather than failing anything, and a written recording holds at most 300 frames. The baseline half is the same suites run against the reference implementation, so the reviewer’s side-by-side compares the build’s and the reference’s frames from one scenario driven the same way, and the console’s replay player draws 2D and 3D recordings frame by frame from the shared resource tables. See recording.

The toolchain gate and the verified dependency install

Section titled “The toolchain gate and the verified dependency install”

A case declares a [toolchain] table (typecheck required; lint, format and test optional), and a post-run stage runs the commands over the collected tree after the container is gone, together with a smoke check that builds the site and opens it headlessly. Each command is recorded on the run record’s toolchain block with its exit code and a bounded output excerpt. A typecheck that ran and exited non-zero rates the run broken and scores it zero, applied where the rating is derived so validator and reviewer verdicts stay as written; the other three gate nothing, and lint runs with --max-warnings 0 so a warning fails. Versions that declare no [toolchain], every frozen one among them, are neither checked nor gated.

The test command’s results and coverage are read from two report files the case’s build vitest config writes (coverage/test-report.json from the json reporter and coverage/coverage-summary.json from istanbul with reportOnFailure: true) rather than from stdout, including the individual tests with name, file, status and duration, bounded and flagged when cut; the run’s Code tab renders both bands and joins per-file coverage onto the explorer. Every dependency install of a collected tree is verified against the lockfile the way npm’s tree builder decides the expected set, retried up to three attempts with a delay, and recorded with its output and attempt count; validation reuses the recorded install, and a run whose install still fails ends as an infrastructure failure rather than catastrophic. Seeding admits .prettierrc.json and .prettierignore through the dotfile allowlist so the format command checks the case’s configuration. See end-to-end manifests.

The instrumentation contract: reconcile(), unconditional operations, posed draws

Section titled “The instrumentation contract: reconcile(), unconditional operations, posed draws”

The debug-API contract every playable case mandates changed in three ways, applied to the twenty editable case versions and left untouched on the 27 frozen playable ones. reconcile() joins the core operations: it re-derives every reading the snapshot reports from the state it depends on without advancing the clock, firing anything or correcting anything, so a driver poses a world, reconciles it and reads the world it posed. An operation is unconditional: it never declines on the strength of the route a player would have taken (which screen is up, which panel is open, where an actor stands), a pose applies the value it is given without clamping, and a call whose argument names nothing real still throws.

Randomness is posed rather than seeded. Specs state each draw as behavior (the set, the probability, when it is drawn), name no generator, carry no seed on reset and no generator state in the snapshot, and the API carries an operation that sets the outcome each validated draw decides (setNextSaucerEdge, setNextPod, setNextCrit, setLanePhase and their siblings) plus a gate that stops automatic draws while a scenario is posed (setWaveSpawning, setBearEmergence, setFishCadence, setTimerRunning), or performs one draw alone as a reading (drawVent, rollPress, generateBoard). Each operation of an engine-format version sets one field, places or removes one entity, reads the state or moves the clock; the compound and patch operations (startMatch, serve, setBoard, startGame, placeTower) are gone, snapshot() reports every field an operation can set, every entity carries a stable id, rosters can be emptied one at a time, and per-entity sense, travel and fire splits let a check on what a creature senses hold its body still.

The keyboard operations (keyDown, keyUp, press) are gone because the keyboard belongs to the runtime, and most cases drop setMuted in favour of the real mute binding. The browser driver serves the engineless contract only: setAutoStep is a required entry, and a read-only preflight refuses a build whose handle lacks step or setAutoStep. See instrumentation.

Every engine-format version’s menus accept a pointer and touch alongside the keyboard: moving a pointer onto an item selects it, a press and release inside one item confirms it, a press that begins on one item and ends on another confirms nothing, and a touch selects where it lands and confirms where it lifts. The build owns its layout and reports each item’s hit region through a new menuItemRect(index) reading, which the validators drive real Chromium mouse and touch input at, and setMenuIndex and menuIndex join the surface and state. The same pass pinned the transitions the older specs left open: Esc and P both open and both resume the pause menu, every menu opens on its first entry and wraps at both ends, and returning to the title selects the entry that led away from it.

Pointer-first cases go further. Refract’s and Facet’s screens carry on-screen BACK, CLEAR or PAUSE controls so a touch-only player can reach every screen, and Facet’s board is played entirely by hand with the keyboard reduced to menu actions. Carom v3.0.0 also makes the paddles movable during the pre-serve countdown and draws the serve’s vertical sign at random.

Section titled “Showcases: a run’s landing page and a case’s carousel”

A showcase is a directory holding showcase.md (store-page markdown, up to 64 KiB), showcase.toml (up to 10 [[media]] tables, each a file and a caption name) and the media flat beside them (.png image, .json.gz replay or .webm video, each up to 25 MiB). The run showcase is written by the model at showcase/ in its repository, instructed through the case’s specs rather than a manifest key; record assembly captures it onto the run record as showcase with lenient degradation (truncation, dropped entries, a warning, never a failed run), the driver uploads the directory, the backend and artifact service serve it at GET /runs/{id}/showcase/{file}, the snapshot publishes it under media/runs/<id>/showcase/, and the run’s Play tab becomes its landing page rendering the carousel over the description. A case that requires a showcase carries one review item (showcase.exists, weight 3, capped at great) that checks its existence alone; the media is the reviewer’s to judge.

The case showcase is declared per variant with showcase = "showcase/<variant>" and captured from the reference implementation: the leading entry is twenty to forty seconds of real play driven through the real input path (a Carom match to two points, a whole Cascade deal, a Gantry lift, a Meltdown floor played out), with capture drivers committed under showcase/capture/ that audition several takes and keep the best. It hard-fails version resolution on any malformed entry, so tcab capture-baselines <slug> --dry-run doubles as the validity check, is served at GET /test-cases/{slug}/versions/{version}/showcase/{variant}/{file}, drives the catalog preview stage, and is published content-addressed under media/cases/... with .webm transcoded to .mp4. Every new version except Volute declares one. See showcases and authoring a case showcase.

A redesigned catalog and a home page of legendary runs

Section titled “A redesigned catalog and a home page of legendary runs”

The test-case catalog is a resizable master-detail split with a sticky preview stage and filmstrip, cards show each case’s latest version, and the detail page’s Inputs tab is a grouped file tree with an inline highlighted source viewer shared by runs and game jams. The header shows the anchored version’s own description and a link naming the latest version when an older one is viewed, and the snapshot publishes each variant’s starter-workspace files under files/cases/....

Test-case groups (test-case-groups/<slug>/test-case-group.toml, served by GET /test-case-groups and carried by the snapshot) define ordered sets of related cases, and the public home page renders one cross-case leaderboard per group, ranked by mean score fraction so cases with different point totals stay comparable. Three ship: tower-defense (meltdown, valence, arc-foundry), arcade-physics (pong, fathom, spectra) and sim-economy (deepcore, coil); every member slug must resolve in the catalog, which the manifest test and backend ingest both check. The home page also shows the five most recent Legendary runs as looping showcase replays, a totals band and a 52-week activity chart from the open GET /stats/cabinet read (runs, tokens, comparable cost, distinct cases and models, weekly buckets), and the Other section is reachable on the public gallery. See test-case groups.

Eleven existing cases ship a new major version on the engine format

Section titled “Eleven existing cases ship a new major version on the engine format”

Every non-experimental end-to-end and full-stack case gains a version authored to the v0.7.0 shape: Carom v3.0.0 (slug pong), Cascade v3.0.0, Fathom v3.0.0, Floe v3.0.0, Shatter v3.0.0, Spectra v2.0.0, Wireworm v2.0.0, Meltdown v2.0.0, Coil v2.0.0, Arc Foundry v2.0.0 and Deepcore v2.0.0. Each seeds a complete TypeScript project (Vite, tsc, ESLint, Prettier, Vitest and an index.html) instead of a bare page, names one starter directory per engine in a [workspaces] table, and runs under three engines: the engineless none run, Simple 2D and Structured 2D at >= 1.0.0. Under none the seeded project holds no src/ at all and the build writes the frame loop, canvas fit, input, audio, overlay and its window.__<case> surface; under an engine the case seeds src/constants.ts and src/main.ts and the build writes src/game.ts, whose initialize returns the debug surface beside the state.

The tick_hz key is gone from every manifest: rates are per second and integrated against the frame’s delta time (Floe keeps a 120 Hz step as a rule stated in its own spec), and each validator constructs its own clock. Every checklist was re-cut to one observable behavior per item, which is why they grew (Cascade 47 to 292 common items, Floe 83 to 256, Shatter 99 to 241 common plus 54 variant, Meltdown 107 to 378, Wireworm 77 to 225, Deepcore 109 to 400), and domains were split so one failure no longer sinks a whole score: Cascade rates rules, handling, cascade and presentation, Fathom navigation, sensing, predators, progression and presentation, Floe hunter, crossing, run and presentation, and Arc Foundry adds an audio domain for its twelve produced cues. The standalone .mjs browser drivers of the previous versions are gone, each variant ships a reference implementation per engine under references/<engine>/<variant>/, and the previous versions stay frozen and untouched apart from their .frozen digests. See the end-to-end overview and writing workspaces and references.

Refract v1.0.0 (end-to-end, easy, 3 h) is a light-tracing puzzle on an optical bench with a 24-board authored campaign and an endless Cascade mode whose generated boards are held to measured per-tier difficulty floors; every screen is pointer-first with named, fingertip-sized targets the debug surface reports. Facet v1.0.0 (full-stack, easy, 4 h) is an 8x8 gem matcher where every clear strains the neighboring stones, cracked stones clear with anything beside them for double, three cuts carry looping auras, a move is offered on the drag and only played on the release, and each level ends on a levelclear tally screen. Kessler v1.0.0 (full-stack, easy, 6 h) is an orbital breakout: a deflector on a circular track bats a ball outward through three rings of derelict satellites while caught salvage pods grant tools.

Volute v1.0.0 (full-stack, easy, 6 h) is a channel shooter in a geothermal pump hall, where cores ride a winding channel toward an intake and a central injector groups like charges to pull them out. Wick v1.0.0 (full-stack, easy, 8 h) is a survivors-like: the lamplighter only moves, sixteen weapons fire on their own, thirteen enemy types arrive on a ten-minute schedule ending with the Dark, and its checklist is 943 common items across 30 categories. Gantry v1.0.0 (full-stack, medium, 8 h) is the first 3D full-stack case: rig a tower crane from struts, cables and rails, write an instruction tape for its four axes and run it under a structural simulation across six sites; it declares asset_dimension = "3d", depends on @clockwyrks/voxel-runtime, and runs on Simple 3D and Structured 3D rather than the 2D engines. Orrery v1.0.0 (full-stack, medium, 8 h) is a machine-building puzzle of brass arms and celestial motes on a hex field (sigils, looping tapes, constellation deliveries graded on cost, time and footprint) and carries the largest checklist at 1054 items.

All seven are validator-rated on every engine they name and none is experimental. Every full-stack one except Wick declares packages = ["@clockwyrks/particle-runtime"], and Gantry declares the voxel runtime instead. See the full-stack overview.

Case-by-case rule changes in the new versions

Section titled “Case-by-case rule changes in the new versions”

Beyond the shared rework, each version pins rules its predecessor left open. Carom v3.0.0 splits hit-edge and no-tunnel into per-face points, adds navigation and hud categories, states the sub-stepped physics exactly (the serve angle is exactly SERVE_ANGLE), moves six points multi cannot share into the variant files, and adds setMuted. Cascade v3.0.0 states that the waste remembers each turn as a set (wasteSets), fixes every pile’s drop rectangle and the leading-card-centre drop rule, gives DRAG_THRESHOLD (5) a stated effect, and makes a cleared table put down the run in hand. Fathom v3.0.0 adds setMaze fixtures, splits sensing from travelling per creature, adds a released flag for the den schedule, and pins the sonar wavefront at fourteen corridor steps per second. Floe v3.0.0 moves to stage coordinates everywhere and fixes the five bays at columns (3,4), (11,12), (19,20), (27,28) and (35,36).

Shatter v3.0.0 clears a wave only by shooting it, grades the saucer over twenty real crossings, reads the fragment fan 412 units from the well, and grades that RESTART begins a game clear of saucers. Spectra v2.0.0 replaces the flat 4.0 s Flux cycle with fluxHold(stage) and three sibling stage formulas, sub-steps a frame by SUBSTEP_MAX, states mute at source strength, and requires before-and-after captures for difference claims. Wireworm v2.0.0 states the worm cadence as wormStepInterval(level), fixes CURSOR_SPEED at 430 and GLITCH_DART_INTERVAL at 0.32 s, rules that a drop passes through whatever is in the tile below, and pins when the spawner clocks are set. Meltdown v2.0.0 drops the 60 Hz tick, accepts dist/ alone as build output, requires unit tests at src/**/*.test.ts, and pins seven open rules including closed-form wave composition, the trip as a crossing of 100, and the Rime as an ordinary emitter with base damage 4.

Arc Foundry v2.0.0 rounds a unit’s maximum HP to a whole number, fixes Medium’s surcharge constants at c = 0.28 and r = 1.145, and replaces the right- and shift-press with a modify action bound to Shift that every engine can deliver. Deepcore v2.0.0 adds setWorldSize, pointer and touch across the shell, and scrolls the shaft through the engine’s own camera under Structured 2D. Coil v2.0.0 makes the best score a session figure and moves the three points the Maze mode disagrees on into both variant files.

Carom v2.2.0 grades spin on the ball’s flight

Section titled “Carom v2.2.0 grades spin on the ball’s flight”

Carom also gains v2.2.0, still on the single-workspace legacy format with tick_hz = 120 and 71 review items, as a correction pass over v2.1.0 learned from runs. The spec states that there are two ways out of a pause (the pause key toggles, and the RESUME entry confirmed with Enter or Space), that every menu opens on its first entry, and that the pre-serve countdown is live play in which paddles move and the pause key works; a new Pause point resume grades all four routes. Every Spin point measures the ball’s perpendicular offset after the contact rather than only reading the spin scalar, comparing magnitudes so mirrored builds both pass, and a parked AI at contact makes spin.moving-solo-ai inconclusive rather than failed. Audio points arm the AudioContext with the bound W key and drive cues in real time, advances-in-real-time is judged from injected keys instead of control operations, and the spec states who owns the clock at boot, after a control operation, and under injected input.

Lattice (performance, hard) loses its experimental = true line, so v0.7.0 is the first release in which Lattice runs are recordable and the case appears on the site as a real case. Its scored set was redesigned around dense main-bus factories: medium is a designer-authored 48x32 factory of 630 entities run over 300,000 ticks, large copies it onto a 72x40 grid and extends it to 1,027 entities building the whole machine tree over 360,000 ticks, and small keeps the simple lines layout as the fast correctness confirmation. Raw ore is emitted only at sources and smelted in a new 2x2 furnace entity that burns a new coal item, three machine recipes (transport-belt, inserter, assembler) form a dependency tree, belts move at tiered slow, fast and express speeds, and every splitter, curve and side-load does real work.

The engine rules were corrected in step: the splitter is a lane-preserving, item-agnostic balancer, an inserter’s empty return takes real time and a crafter-loading inserter picks the item the recipe still needs, belt movement is defined over a run rather than a tile, a pure curve continues its run with both lanes preserved, and a side-load lands each feeder lane at its contact point. The fuel ceiling is 40,000,000,000, under which the transport reference passes and a naive engine exhausts its limit, a returned checksum must match the returned state, and every oracle and checksum was regenerated. The case page gains a Reference tab that plays the scored factories, and the performance documentation was rewritten in step. See Lattice.

Two Lattice sprite-sheet cases, and the existing sheets grow

Section titled “Two Lattice sprite-sheet cases, and the existing sheets grow”

medium/lattice-furnace (twelve 64x64 frames of a 2x2 top-down smelter, an off idle in frames 0 to 3 and a smelting loop in frames 4 to 11) and medium/lattice-lane-splitter (eight 32x64 frames of a single-input machine that unzips one belt’s two lanes onto two outputs) are new experimental asset-generation cases, each with a draw.sh reference implementation, which exist to seed Lattice’s sprites. The existing Lattice sheets grew: lattice-belt goes from 16 to 48 frames (straight and curved forms at three tiers), lattice-assembler from 8 to 24, lattice-inserter from 12 to 36, and lattice-items from 7 to 17 icons. lattice-splitter gains a weight-4 fidelity item requiring the housing, output arrow and moving part on a registered mechanism layer over a continuous belt bed, and the lane splitter carries the same rule. The asset-generation catalog stands at 132 versions. See sprite cases.

No live case version seeds a canonical palette, a typeface requirement, HUD coordinates or reference screenshots any more, and none declares [[reference]] views, [[proof]] artifacts or [[check]] comparisons. The specs instead state a legibility table of what a player must read at a glance, and the presentation validators assert that a thing was drawn and that two things are told apart against a stated RGB distance, never a hex value; how good it looks is the run-wide aesthetic rating. The nine experimental legacy-format versions (Caldera, Siege, Sunfront, Thunderhead, Holdfast, Hollowdeep, Junction, Midway and Valence v2.0.0) also drop their [[proof]] and [[reference]] tables, specs/proof.md, mockup sources and capture instructions while keeping every review item, and a case declaring no mockup needs no headless browser at ingest or rendered reference at run start. The three features stay supported in core because the 27 frozen versions still declare them. See writing case specifications.

Audio packs are declared per case and staged into the run

Section titled “Audio packs are declared per case and staged into the run”

[audio] packs replaces the sample_pack and instrument_bank manifest keys: an ordered list of name@version refs, valid on audio asset-generation cases, full-stack cases and game jams, with ManifestAudio denying unknown keys so a stale line is a loud parse error. An sfx-sample case declares exactly one sample pack, a music case exactly one instrument bank, an sfx-synth case none, and order decides which pack an unqualified call plays. A full-stack version or jam with no [audio] table receives the pinned four-ref DEFAULT_AUDIO_PACKS ([email protected], [email protected], [email protected], [email protected]), a literal rather than a scan of the registry, so publishing a pack cannot widen a frozen case. The fourteen non-frozen full-stack versions each declare a set chosen against their own specs/assets.md, all eight game jams declare the full four-pack set, and scripts/ci/audio-packs-check.mjs requires the table on every unfrozen full-stack version and jam.

At container start core’s audio_stage resolves the declared refs against the host audio store (TCAB_AUDIO_STORE, default /opt/tcab-audio), verifies every clip against the store’s objects.lock.json by sha256 and byte length, and materializes packs.json, one packs/<name>@<version>/pack.toml per pack and the union of their clips under clips/ at /opt/audio inside the container, outside the seeded repository. The tree travels as a host directory on the new ContainerSpec.dirs, one copy per container rather than per-file bytes, and a pack ref’s halves are held to plain name characters. A run reads /opt/audio/.tcab-audio-contract back from its started container and fails at start, naming the image and pin, when the image predates staged delivery. See execution and full-stack manifests.

Audio packs are collections of registered clips in an object store

Section titled “Audio packs are collections of registered clips in an object store”

A pack is no longer an opaque tarball. containers/sample-packs/clips.toml is the clip registry, 70 distinct clips keyed by the sha256 of each clip’s original source bytes, holding its license, provenance and recorded pitch, and each of the four pack manifests is a collection of those ids plus the name, tags and description a model browses, so a clip shared between packs is stored once. Clip bytes live in a private Cloudflare R2 bucket as the original source and as the output normalized under each pack’s profile, and objects.lock.json records the 140 published objects. The publish tooling is split over scripts/lib/audio-store.mjs: curate-instrument-bank.mjs is the only script that contacts Freesound, build-sample-pack.mjs publishes a pack’s normalized objects, and stage-audio-store.mjs materializes every published pack into one verifiable tree; presign-sample-pack.mjs and packs.lock.json are gone.

That tree ships as the data-only test-cabinet-audio-store image, built ahead of the run images by scripts/ci/audio-store-image.sh and fused per architecture, and the driver service image copies it in through the AUDIO_STORE_IMAGE build arg to /opt/tcab-audio. The sfx-sample, music, full-stack-2d and game-jam run images bake no audio at all. scripts/fetch-audio-store.sh puts the store on a laptop for a local run, pulling the image from the registry or staging straight out of R2 with the read-scoped presign pair, defaulting to ~/.cache/tcab/audio-store, and the R2 endpoint may be given as CLOUDFLARE_AUDIO_R2_S3_URL instead of being derived from the account id. See publishing an audio sample pack.

The audio binaries load the staged palette and fail closed

Section titled “The audio binaries load the staged palette and fail closed”

sfx-sample and music read their palette from /opt/audio through a shared staged contract in crates/audio-core, so no run inherits a palette from its image and a tool config can only reach a pack its case declared; a config that names no pack takes the staged default of its kind, and --config on both tools takes a pack ref. A palette that is missing, unparseable, empty, of the wrong kind or not the pinned name@version fails the run instead of collapsing to an empty library, audio whose duration or sample rate disagrees with its manifest is rejected, and at render time a placed sample absent from the library, or a track instrument that is neither a synth waveform nor a bank entry, is an error naming what was missing rather than silence.

music gains list-instruments [--tag] and instrument-info --name, mirroring sfx-sample’s list-samples and sample-info, which every music brief already told the model to use. TCAB_AUDIO_DIR points a host-side tool run at a fetched audio store, read as a palette of every pack it holds. See the audio binaries.

A full-stack case names its asset dimension

Section titled “A full-stack case names its asset dimension”

asset_dimension is a full-stack-only root manifest key, "2d" (the default, so every existing case resolves unchanged) or "3d", and it selects the run image the way an asset-generation case’s asset_kind does. "2d" runs in test-cabinet-full-stack-2d with draw, draw-sheet, particle-2d, sfx-synth, sfx-sample and music; "3d" runs in the new test-cabinet-full-stack-3d, which adds voxel, voxel-anim and particle-3d plus the Mesa software Vulkan runtime their preview PNGs render through. The key travels the backend wire so dispatched runs resolve the same image a local tcab run does, the standing full-stack quality directive names the binaries of the image the run is in, and resolution rejects the key on every other test type. Gantry v1.0.0 is the first 3d case, and Orrery moved from end-to-end to full-stack so its build produces its own sky and brass. See full-stack manifests.

Reference implementations and baselines are per variant per engine

Section titled “Reference implementations and baselines are per variant per engine”

A variant declares one reference implementation per engine, by convention references/<engine>/<variant>/, or a bare path for a single-workspace case. tcab publish-reference --env <prod|staging> <slug> [<version>] [--variant] [--engine] [--all-variants] [--dry-run] [--skip-baselines] builds each variant and engine pair with the case’s [build] commands, re-captures its baselines, scrubs and deploys it to Cloudflare Pages under the alias <slug>-<version>-<variant>-<engine>, and records the URL in test-cases/reference-builds.lock.json at <env>.<slug>.<version>.<variant>.<engine>. The backend’s case_reference_build table is keyed by (slug, version, variant, engine), the case page’s Play tab launches the build recorded for its anchored engine, and the Reference tab switches between engines.

capture-baselines records an engine-backed case’s baseline by running that engine’s validator suites against the reference implementation in process, and a browser-driven case by serving and driving its build, so both panes of the reviewer’s comparison come from the same scenario. Engine-backed references depend on the repository’s own packages/<slug>/ by a relative file: path instead of a vendored copy, so npm ci && npm run build:packages at the root is a prerequisite of both commands. A new guide covers buildable, script and bundled references and the release gate. See publishing a reference implementation.

Writing Debug APIs and Validators fixes the debug-API design rules (atomic scalar operations, no patch operations, compound sequences in the validator harness, the declared state as the whole state, a build reporting layout a spec leaves loose, random draws posed as outcomes) and the validator rules: assert the specification never the reference, grade the build’s code not the engine’s, transcribe every asserted figure into the project’s own constants.ts with the build’s module imported at one site, one requirement per validator, pose an isolated world, drive simulated time only, sample a probability through a lone-draw operation inside a six-sigma band, read state before pixels with presence floors rather than thresholds, always reach a verdict, and finish inside budget (under 3 s per validator on an engine, 5 s engineless, 10 s hard cap, 15 minutes for the whole suite on a two-core host).

Writing Workspaces and Reference Implementations states what a seeded workspace and a reference hold: the four toolchain scripts, Prettier and ESLint configured in both, specs/ ignored by both, an inert <link rel="icon" href="data:,">, an engineless workspace of configuration only, and dependencies pinned to the newest usable version. Writing Case Specifications gained the never-help-the-model rule, the keep-evaluation-out rule enforced by scripts/ci/spec-vocabulary-check.sh, menus that take pointer and touch with every transition stated, one observable behavior per item, and that appearance is loose and reviewed while behavior is exact and validated. The end-to-end and full-stack guides and quickstarts were rewritten for the per-engine layout. See writing debug APIs and validators, writing workspaces and references and writing case specifications.

A killed run is canceled when it is a gg run, and destroyed otherwise

Section titled “A killed run is canceled when it is a gg run, and destroyed otherwise”

RunState::Canceled joins the run-record contract as a terminal state: never publishable, released nothing, never retried, and absent from every model statistic. Killing a gg run whose session has been launched is a cooperative wind-down: the driver raises a latch and keeps awaiting the run under a 20-minute grace, the engine writes the in-container sentinel that every agent checks at its turn boundary, the session finishes the turn in flight and runs its whole epilogue, and the run walks its ordinary post-session path (tree collection, metrics, session summary, artifact uploads), skipping only validation, so the record carries real tokens, cost and the tree as it stood. The driver posts a canceled status carrying the record, the only status the backend accepts on an already-canceled job, and the backend persists it without a completion notification or retry.

A killed run of any other harness, and a gg run killed before its session was launched, is destroyed: the sandbox is deleted, which is what stops it and frees the scheduling slot, nothing is recorded, and the job stays canceled with no record. state=publishable refuses every never-publishable state, so a reviewed canceled run is no longer listed for publish, the console keeps a canceled run visible and deletable, and a run event fires for a cancellation while no notification does. The CLI container runtime gains the job-id label so a canceled run is found and removed under both runtimes. See the driver and run records.

limit_exceeded joins the terminal states for a gg run stopped on one of its configured execution ceilings; it publishes like a harness_error (a statistic only) and is never retried. A 401 or 403 from the provider ends the session under auth_error and exits non-zero so core records a harness error rather than blaming a model that never ran, and a failed gg run’s status quotes the error lines gg logged instead of a bare exit code. A run’s compact GgSessionSummary (terminal status, agents, compactions, provider stats, per-slot costs) is emitted before session_ended and lifted onto RunSubject.gg_summary, which is what GET /stats/providers and the query language read. The run record and per-turn usage events carry the routing key, the provider pin and the reasoning setting, each optional so stored records keep reading. See results.

Deleting a run prunes its artifact tree, and a sweep reclaims orphaned trees

Section titled “Deleting a run prunes its artifact tree, and a sweep reclaims orphaned trees”

The backend’s artifact URL is split in two: TCAB_ARTIFACTS_PUBLIC_URL stays the browser-facing value advertised through GET /config, and the new TCAB_ARTIFACTS_URL is the in-cluster address every backend-originated call uses, so DELETE /runs/{id} now reaches the artifact service to prune the tree (best-effort; the delete succeeds regardless). A periodic sweep lists the service’s stored trees over a new service-token-gated GET /runs on the artifact service, keeps every tree whose id still has a run row, and deletes the rest once older than a grace window; a pass that reads no runs or fails its listing abandons itself so a fresh database beside a populated volume leaves the volume intact. The artifact service’s whole-run archive and the publisher’s tree.tar pull now stream instead of materializing the tree in memory, so a large tree can no longer exhaust the service. See the artifact service.

In-progress rows report their engine, start time and a ticking duration

Section titled “In-progress rows report their engine, start time and a ticking duration”

A pinned in-flight row used to dash its ENGINE, STARTED and DURATION cells. job.started_at is stamped once on the transition into starting, the moment the driver takes the start its record is measured from, so queued time is structurally excluded; RunEvent and GET /jobs/active carry engine, startedAt and ggPreset, and the run log ticks every moving duration off one shared clock held only while one is on screen. Every server-paged listing (the runs index, the worklists, case and model Runs tabs, gg Sessions, the home page window) re-queries on the runtime’s refresh token when a run finishes, is published, is killed or is deleted, so a finished run settles into the filtered page it was on instead of vanishing. Clear pending, Kill active and Stop all report their counts in a toast, another run finishing no longer resets the page you are on, and the testCase sort orders by the case’s display name. See the UI library.

Every run-detail tab is its own route, so leaving one re-fetched its assets behind a spinner; one bounded LRU cache seeded synchronously now serves replays, prepared frames, pixel buffers, meshes, particle systems, skinned meshes, workspace files, engine modules, case details, variants and run records. A read in flight shows a loading state, a failed read reports a failure, and only a read that settled with nothing reports absence, decided by the store’s status rather than error text, and a failed catalog refresh keeps the catalog on screen. Numeric fields hold what was typed and refuse out-of-range values with a reason, and the backend refuses the same bounds rather than correcting them.

A change that can cost work is staged and confirmed through the themed dialog, and every window.confirm site uses it. A page error boundary keeps the chrome when a page throws, a route smoke test mounts the app at every declared route against a stocked console, an empty one and the static gallery, and the front-end tests run in the pipeline’s webtest job. SubmitNotice renders a form’s outcome beside the button that raised it, nine pages hand their create action to the header’s titleActions slot, the “(starts off)” annotations are replaced by opening on the real default with a reset beside any moved field, and the loading mark states its aspect ratio before its SVG arrives. See the UI library.

tcab answers in its exit status, and gains analyze, engines and test-case-groups

Section titled “tcab answers in its exit status, and gains analyze, engines and test-case-groups”

tcab validate used to print a failing verdict and exit 0. It now exits non-zero when the tree did not load, a required install or build step failed or was never reached, a declared check was unreached, a declared proof is missing, a gating validator decided against the build or did not run at all (with inconclusive units grouped by kind and reason on the final line, so a broken host reads apart from a broken build), or an adversarial submission forfeited; a loss, a draw and a similarity figure decide nothing, and a point an erratum excludes costs nothing. capture-baselines fails a target any of whose units did not run clean, after sweeping every target, and publish-reference skips deploying a target whose baseline it could not produce. tcab validate also leaves the directory as it found it apart from the install, build and media under .vendor/.

run, seed and validate take --engine (default none); engines and test-case-groups list the built-in engines and the groups; analyze <dir> runs the static analyzer; publish-reference and capture-baselines work per variant and engine pair. The CLI’s log lines move to stderr so tcab analyze --json | jq parses, and tcab no longer links the gg harness library or an HTTP client. See the CLI.

gg conducts multi-step work through its own executor (compaction, memories, the project board, subagents), which supersedes looping a stateless third-party harness across sessions through the filesystem. The orchestrators/ralph/ data directory, its embedded built-in registration, its catalogue page and its entry in the run-launch picker are gone; one-shot is the only built-in orchestrator, so the picker’s program-building gate goes with it. A session loop remains expressible as an external orchestrator through --orchestrator-dir. See orchestrators.

The dispatcher no longer kills run pods for memory

Section titled “The dispatcher no longer kills run pods for memory”

Driver pods were dying with OOMKilled (exit 137) after hours of paid API calls, because the driver container carried a memory limit equal to its request and the sandbox pod a 4Gi limit, while the driver runs the case’s whole toolchain (install, tsc, linters, vitest, bundler, headless Chromium) over the collected tree in its own cgroup after the sandbox is gone.

Both memory limits are removed and the requests kept as the node reservation. The driver keeps a CPU limit of 2, which bounds Node’s worker fan-out and throttles rather than kills, with its memory request at 2Gi. The always-on services and the publisher Job keep memory request equal to limit, because a kill there is a restart, not lost spend. See the dispatcher.

The driver collects the run tree over its own verified channel

Section titled “The driver collects the run tree over its own verified channel”

Kubernetes can drop the tail of a large exec stdout under load while the exit status still reports success, which lost completed runs. Under the Kubernetes runtime the driver now binds a TCP listener on an ephemeral port, advertises its pod IP and a per-run token to a Node uploader it ships in its own binary and execs into the sandbox, and accepts the produced tree only when the stream’s terminator arrives with a byte count and SHA-256 digest matching what it received; a truncated or stalled stream, a mismatch, a failed tar or an uploader that never connects is retried with a fresh upload, and a verified archive that fails to unpack fails the collection outright with the innermost cause named.

Seeding the sandbox is judged by tar’s own verdict rather than the stdin write, so a broken pipe no longer masks tar’s real error (a run image without /opt/audio, for instance), a remote that dies without draining stdin no longer hangs the run at “starting the run container” because the wait is bounded on progress, and the kube client’s stdin and stdout pipes are sized up from 1 KiB. See the driver.

A run listing’s total matches its rows, and unreadable records get a tab

Section titled “A run listing’s total matches its rows, and unreadable records get a tab”

Db::assemble used to skip a run whose stored record_json no longer deserialized while the count still counted it, so the console offered empty pages and the run answered 404 while remaining deletable only through the database. Every run row now carries record_readable and the record_format generation it was decided under; every listing and its COUNT run one predicate, and a build whose RUN_RECORD_FORMAT (now 3) differs from a row’s stamp re-decides that row once at startup. GET /runs/unreadable lists the runs this build cannot read with the error their record produces, the console’s Unreadable worklist appears while any exist and is where they are deleted, and DELETE /runs/{id} acts on the row, so it deletes an unreadable published run too. The contract-shape pin now digests every schema the record references transitively, the gg capability set included, so a required field added to a gg agent profile moves the generation instead of silently marking rows unreadable one listing at a time.

A stored definition the running build cannot read makes the backend unready

Section titled “A stored definition the running build cannot read makes the backend unready”

After a stored-shape change the catalog listing used to skip each case whose latest manifest failed to parse and answer with the rest, so the console showed one test case while the backend reported itself ready. The definition store now records the format its contents were written in (STORE_FORMAT, currently 2); a store stamped with any other format holds the backend unready with that reason and makes GET /test-cases answer 503 naming the repair, and any ingest scan meeting such a store is promoted to a forced whole-catalog one, which is what the cluster overlays’ startup sidecar posts. scripts/reingest.sh asks /healthz first and ignores its baseline when the store is unservable, and local-rebuild ingests between the restart and the readiness wait. See the backend.

gg’s own defects end the run as internal_error

Section titled “gg’s own defects end the run as internal_error”

A guest artifact that would not instantiate, a wasmtime host fault, a panicking sandbox task, a compiler that crashed or was missing from the image, and gg failing to lower a source it had already accepted all used to end the session as model_error or feed back to the model as a compile error counted against its ceilings. Each now reports internal_error, and PrepareError::Lowering and the TranspileLowering error type are deleted from the published vocabulary.

A run-wide fault latch is read by every agent at its turn boundary, so a defect in any child, issue agent, reviewer or successor ends the whole run rather than letting the root finish and be scored, and a suspended wait_for_issue selects on the latch so a fault can no longer leave waiters parked until the idle watchdog files the run as hung with no record. An error turn taken under a raised latch is recorded as gg’s (fatal) rather than in the model’s program_fault bucket, and a wait on an issue blocked by a failed issue ends as a refusal naming the blocker.

A compile failure is the model’s only when its program caused it

Section titled “A compile failure is the model’s only when its program caused it”

On the C#, Java, Kotlin, Swift and C++ arms a loaded module’s rebuild diagnostic was returned in the model’s band, asking it to rewrite a program gg had accepted; it is re-attributed to a lowering fault of gg’s on every arm that rebuilds a module beside a program. Roslyn’s arrangement diagnostics are classified as toolchain failures, the response file’s arguments are quoted so a workspace or home path with a space compiles, and the C# guest carries the thrown exception’s code and operation up so an uncaught gg API failure is program_api_error and a withheld capability is program_unknown_name like every other reporting arm. A Compiler error body is the arm’s rendered diagnostic and nothing else: the supporting material, import-candidate lists and matching rule were removed on every arm, and the catalogue’s libraries section is the record of what an arm may import.

PureScript arm fixes, and a struck stack closes with a count

Section titled “PureScript arm fixes, and a struck stack closes with a count”

The PureScript arm’s hand-written module-header scanner was deleted in favour of the name purs emits. Its code mask panicked on a multi-byte character inside a block comment, which reached module preparation and ended the run; a PureScript runtime failure returned frames from inside gg’s SDK instead of the model’s own program line as every other arm does; and signature analysis for module documentation pages recognized only the ASCII -> and => spellings, not → and ⇒. All four are fixed. On every arm, a stack trace gg has struck frames from now closes with the plain line … and N more frames (external code) at both strike sites, which is visible in every Runtime error body a model receives.

A replay’s images travel beside it, and a recording binds to the drawn surface

Section titled “A replay’s images travel beside it, and a recording binds to the drawn surface”

The injected canvas recorder bound to the largest attached canvas, so a build that renders into an unattached fixed-resolution canvas and blits it produced recordings that were almost entirely base64 PNG and failed validator points over drawing the build did. It now follows a pass-through blit to the source canvas, images are pooled across a run in a shared store beside the recordings and deduplicated by content, and a per-run budget bounds what recordings hold. @napi-rs/canvas moves to the 1.x line because 0.1 leaked one copy of the source buffer on every drawImage(canvas). The player carries the canvas specification’s initial values and stays quiet about a refused assignment of a default while still reporting a non-default value it cannot apply, and a replay pane’s frame count sits on its label row.

Validators read what the spec fixes, off the wall clock, inside a budget

Section titled “Validators read what the spec fixes, off the wall clock, inside a budget”

A repository-wide audit of the new validator suites landed as one fix(test-cases) commit per case plus a cross-case sweep for the fault classes it found. Validators now transcribe every asserted figure from specs/ into their own project constants instead of importing the build’s src/constants.ts, with each case’s own lint rule refusing a validator that reads the build’s src. They read copy off coalesced text runs and grouped figures as one figure, and they drive simulated frames rather than waiting on real time, so a loaded host cannot fail a conformant build. Heavy drives were re-costed to what they read (Cascade’s trail-* points step at the runout rate instead of 240 Hz).

A still whose pose throws fails its point instead of shipping an un-posed picture, and an entity a check dereferences is hard-asserted first. Appearance thresholds became presence floors, engine-owned behavior is scoped to none, and withheld-asset shims were removed with the specs requiring every produced file to load. Each case passes its whole suite against every reference on every engine, and baselines were regenerated after the sweep.

Catalog discovery requires a test-case.toml or game-jam.toml before treating a subdirectory as a version, so a leftover reference-impl/ or dist/ on disk no longer makes the whole catalog fail to list. Whether a case may name an engine is decided by workspace against [workspaces], and the categories checklist grammar is opted into with [review] format = 2. The performance engine model is served as a run asset, the backend ingest test asserts every engine floor a manifest pins, case READMEs no longer list validation-baseline/ as part of the version folder, and seeded text that named the benchmark or the review UI was reworded (Foray’s baselines spec, fourteen particle briefs).

The local k3d stack keeps its images and reports what stalled

Section titled “The local k3d stack keeps its images and reports what stalled”

make -C deployments/local local-up on a cold state volume waited 600 s for a backend readiness that only its own local-ingest could produce; apply is split into apply-overlay and apply-wait with local-ingest between them, and local-ingest waits for a running, never Ready, backend pod before port-forwarding. The k3d node’s kubelet image garbage collection is disabled on cluster create and retrofitted onto an existing node through the new cluster-kubelet target, because under disk pressure it deleted the driver image and every locally imported run image and left runs in ImagePullBackOff; local-import re-imports every built image without a rebuild. One rollout-wait target over WAIT_WORKLOADS prints the pod table plus pods terminating past their grace and FailedKillPod events, so a node still reaping old pods is distinguished from a broken image, and scripts/free-local-forward.sh kills an orphaned local-forward. See running the services locally.

  • Failure details across core report the error rather than the handling, for example run failed: gg execution ceiling hit (5 consecutive turns failed).
  • The Simple 2D and Structured 2D engines clip the game’s drawing to the logical field, name the pressed button on a bare pointerdown’s samples, and read an event with no pointerId as pointer 0; the recorder records a tinted sprite as drawn.
  • The particle binaries’ over-budget rejection names the flags to lower and keeps the rate x lifetime sizing rule, and error messages across the audio binaries state the fact rather than narrate it.
  • A .json.gz recording is served as JSON framed in gzip, so the browser inflates it and nothing sniffs bytes on the wire.
  • The observability page’s cAdvisor guidance says run pods carry no memory limit, so container_spec_memory_limit_bytes is meaningless for them and the sizing target is the sandbox memory request.
  • The gg analysis page says the runtime cap bounds the harness session and each in-container setup step separately, each with the full budget.
  • binary-smoke.sh resolves the built binary through ${CARGO_TARGET_DIR:-target}, so the release gate also runs inside the devcontainer.

azure-pipelines.yml runs three stages. gates runs on every branch: rust, binary (Linux and Windows), web, webtest, specs, format, validators, frozen, audiopacks, specvocabulary, buildcontext, contract, manifests and submodulepins; on master, staging and v* tags it also builds the static gg binaries natively on amd64 and arm64, and a mirror job force-pushes the gated master, staging or nightly commit, or tag, to GitHub. images runs on master, staging and v* tags: on the branches it builds the audio store, every run-container image and every service image natively per architecture, pushes <image>:<sha>-<arch> to testcabinet.azurecr.io and fuses them into the multi-arch <image>:<sha> with manifest.sh, and on master and tags its gg_publish job uploads gg’s release objects. deploy runs scripts/ci/deploy.sh on the tcab-staging environment for the staging branch and tcab-prod for master, and deploys the docs site to test-cabinet-docs or test-cabinet-docs-staging.

deploy.sh writes a throwaway kustomization over deployments/k8s/overlays/azure-<env> setting every service image to the commit’s tag plus TCAB_DRIVER_IMAGE, TCAB_PUBLISHER_IMAGE, TCAB_CONTAINER_REGISTRY and TCAB_CONTAINER_TAG=<sha>, renders it, refuses anything outside the namespace, and applies and waits inside the private cluster through az aks command invoke; a rollout that does not become ready is described, logged, undone and fails the stage. amd64 jobs run on Microsoft-hosted ubuntu-24.04 agents and arm64 jobs on the organization pool, because the AKS nodes are arm64 and emulated Rust and wasm builds are far too slow. No job holds a stored credential: registry pushes use the tcab-acr connection, deploys tcab-deploy, the gg upload tcab-gg-publish, the mirror the secure file github-mirror-key, and the docs and audio-store jobs the secret variables. See building and the Kubernetes overview.

The GitHub workflows are retired and GitHub is a mirror

Section titled “The GitHub workflows are retired and GitHub is a mirror”

Every file under .github/workflows/ is deleted: ci.yml, build-containers.yml, build-service-images.yml, deploy-docs.yml, publish-reference.yml, release.yml, release-promote.yml and binary-macos.yml. .github/ now holds only a README stating that the repository lives on Azure Repos, that TheClockwyrks/TheTestCabinet is a mirror receiving master, staging, nightly and v* tags from the pipeline’s mirror job once a commit has passed its gates, that the push is forced so anything committed on GitHub directly is lost, and that the mirror runs no CI. The macOS validation job and the desktop release job leave with the workflows, and no manifest, script or doc page names ghcr.io any more. Publishing a reference implementation stays the manual tcab publish-reference flow.

Images move to the Test Cabinet container registry, pinned by commit sha

Section titled “Images move to the Test Cabinet container registry, pinned by commit sha”

Every service image and every run-container image is pushed to testcabinet.azurecr.io by the pipeline, tagged by the commit sha and never as :latest, and core’s compiled default registry is the same host. The azure-staging and azure-prod overlays carry no images: block and no driver-image patch at all, because the pipeline’s deploy supplies every image and TCAB_CONTAINER_TAG, so services, driver, publisher and the run images a driver stages audio into always come from one build. Both clusters’ kubelet identities hold AcrPull, so no pull secret exists in any overlay. The gg-ci toolchain image and its hydrate script, which existed only for the GitHub workflow, are deleted.

Releases are cut through the Azure tag route

Section titled “Releases are cut through the Azure tag route”

A release is a vX.Y.Z tag on master in Azure Repos. Its pipeline run executes the gates, and on the tag the binary job keeps the tcab it just release-built, tested and smoke-tested as the tcab-linux and tcab-windows pipeline artifacts, so the released binary is exactly the one the gate checked. The same run uploads gg’s release objects and the mirror job pushes the tag to GitHub. A tag builds no service or site of its own: services ship as the images master deployed and the docs from the pipeline’s docs job. The gg version gate (scripts/ci/gg-version-gate.sh) fails a tag run whose gg --version differs from the tag with its v stripped, naming crates/gg and crates/core as the crates to bump. The documented sequence is rel/vX.Y.Z into nightly, nightly into staging as vX.Y.Z-rcN, staging into master, then the tag. See releasing and cut a release.

gg has a real version and its release binaries live on Azure Blob Storage

Section titled “gg has a real version and its release binaries live on Azure Blob Storage”

crates/gg and crates/core carry version 0.7.0, pinned to each other by a test, so gg --version no longer reports 0.0.0 and a run record’s subject.harnessVersion distinguishes harness builds. gg’s release host is the gg-releases container of the storage account testcabinetartifacts, anonymous-read with no listing, laid out as v<version>/gg-x86_64-unknown-linux-musl, v<version>/gg-aarch64-unknown-linux-musl and v<version>/gg-reference.tar.gz. scripts/build-gg-static.sh (arch-adaptive; cargo build-portable-gg is its fixed x86_64 form) builds gg as a fully static musl executable so one binary runs across the glibc Debian and Ubuntu run images; the pipeline’s gg_amd64 and gg_arm64 gate jobs build both with scripts/ci/gg-dist.sh, and gg_publish runs scripts/ci/publish-gg.sh on every master build and every v* tag, re-running the version gate and reading each upload back anonymously. The install runs in the run’s setup stage and is recorded as setup, not session time.

Building test-cabinet-gg, and therefore the driver image, scripts/gg-reference.sh and the static build, needs all eleven language toolchains present: scripts/ci/install-gg-toolchains.sh (about 1.9 GB, idempotent) and scripts/ci/install-gg-build-toolchains.sh for the C# guest, because every arm’s guest component, catalogue and SDK library set are generated by the build. A developer’s TCAB_GG_DOTNET_HOME tree that predates the vendored ICU libraries is reinstalled by the installer on its next reconcile. See releasing.

Cluster-scoped objects are a hand-applied bootstrap

Section titled “Cluster-scoped objects are a hand-applied bootstrap”

The Namespace, the observability stack’s tcab-lgtm-node-metrics ClusterRole and ClusterRoleBinding, and the cert-manager ClusterIssuer letsencrypt-internal moved out of the base and overlays into deployments/k8s/cluster/ with per-environment bootstraps deployments/k8s/cluster/azure-staging and azure-prod, applied once by a cluster administrator and re-applied whenever anything under that folder changes. The azure-* overlays therefore render namespaced objects only, while the generic staging, prod and local overlays still include the namespace and observability bootstraps so one apply creates them. The pipeline’s tcab-deploy identity holds exactly two roles per cluster, the custom command-invoke role and RBAC Admin scoped to its namespace, deploy.sh refuses a render holding anything outside the namespace, and scripts/ci/k8s-manifests.sh gates every commit on the same bytes, so a compromised pipeline can at worst rewrite its own namespace. See the Kubernetes overview.

The ingest sidecar force-ingests the branch tip once per deploy

Section titled “The ingest sidecar force-ingests the branch tip once per deploy”

The azure-staging and azure-prod overlays’ ingest sidecar runs one forced POST /ingest of its branch tip (staging or master, cloned from the GitHub mirror into the shared /state/checkout) when the backend pod starts, then idles; every pipeline deploy changes the backend image tag and so restarts the pod, so a deploy also publishes the catalog that shipped with it. Before ingesting, the sidecar runs git submodule sync and git submodule update --init --depth 1 to fetch the cold-storage baselines shallowly, and falls back to ingesting without baseline media if that fails. Catalog changes between deploys are published with scripts/reingest-cluster.sh --env <env>, and the sidecar waits on /healthz rather than /readyz to avoid deadlocking against itself. See the control plane.

crates/desktop, apps/desktop, the deployments/k8s/overlays/app overlay it stood its own k3d cluster up from, the desktop build and sidecar scripts, the desktop CI and release jobs, the Cargo workspace member with its dev profile and Tauri dependency pins, the apps/desktop npm workspace entry and every TCAB_DESKTOP_IMAGE_TAG reference are deleted. The docs site loses the components/tauri section, the Arena component page takes its place in the Components navigation, and every page that described a desktop console is rewritten around the web console. In the UI package the harnessAuth capability, the built-in local worker handle and the solo review-and-publish path go, and core drops select_mode, SubscriptionFileStatus and subscription_files; the canExecute flag stays because the gallery mounts the shared application with it false. See the web console.

Validation baselines move into the cold-storage submodule

Section titled “Validation baselines move into the cold-storage submodule”

The captured baseline validation media (1.85 GiB, about 31k files, the bulk of the repository) left test-cases/ for a cold-storage submodule at the repository root whose tree mirrors the test-case tree, so a version’s baselines live at cold-storage/test-cases/<type>/<difficulty>/<slug>/<version>/validation-baseline/<engine>/<variant>/. A plain clone leaves the directory empty and still builds and passes every gate, and --recurse-submodules --shallow-submodules clones about 2 GB of media up front. One resolver in core, ColdStorage::validation_baseline_dir, maps a version root to its baseline directory beneath the cold-storage root; tcab capture-baselines and tcab publish-reference write there and backend ingest copies the counterpart into the staged version when present, so store, API and snapshot are unchanged and an absent submodule yields versions with no baseline media.

Because baselines are outside the version folder, the frozen digest never covers them, so recapturing a frozen version’s baselines is allowed and changes no score. Every frozen marker was rewritten with scripts/freeze.sh in the migration commit, and freeze.sh now refuses only unstaged or untracked changes under the directory, accepts several directories at once and an optional --reason for the marker. The four frozen versions that depend on @clockwyrks/particle-runtime (Spectra v1.0.0, Arc Foundry v1.0.0, Deepcore v1.0.0 and Valence v1.0.0) also had their prompt and reference implementation edited for the npm-scope rename, with the digest re-baselined and the reason recorded on the marker. CI checkouts leave submodules off, and the pre-commit large-file exemption for validation-baseline/ is gone. See frozen versions and where baselines live.

Submodules are addressed by relative URL, mirrored and gated on their master

Section titled “Submodules are addressed by relative URL, mirrored and gated on their master”

.gitmodules names every submodule by a URL relative to the superproject (../cold-storage), which git resolves against the remote the superproject was cloned from: the sibling repository in the Azure project, or TheClockwyrks/<name> on GitHub, so a recursive clone works from either host. Each submodule repository carries its own .azure-pipelines/mirror.yml that force-pushes its master to a GitHub repository of the same name. A new gate, scripts/ci/submodule-pins.sh (job submodulepins, with an offline table test), fails a superproject commit whose pinned submodule commit is not an ancestor of that submodule’s master, so a commit that reaches the GitHub mirror only ever names submodule commits the mirror already holds. A pipeline job that reaches a submodule repository names it under its uses.repositories and the pipeline lists it under resources.repositories, because the project scopes the job token to referenced repositories; every checkout leaves submodules off, and the pin gate reads the submodule through the Azure Repos API. See building.

Image builds compile once, rebuild stale bases and ship an exact context

Section titled “Image builds compile once, rebuild stale bases and ship an exact context”

The seven per-service Dockerfiles under deployments/images/ are replaced by one multi-stage deployments/images/services.Dockerfile, of which the pipeline builds each service as a stage, and containers/ gains tools/ (the shared asset-tool builder), gg/, gg-toolchains/, audio-store/ and full-stack-3d/ Dockerfiles. Every asset tool is compiled once in the shared tools builder and copied into the 22 run images that bake an asset binary (twenty asset kinds plus the two full-stack images). Service images each get their own cargo target cache and the services and gg stages their own rustup caches, fixing a could not rename downloaded file race on every cold builder and the six-fold recompile of the workspace crates per make images.

The root .dockerignore allowlist re-excludes ignored shapes and lists each gg helper script rather than a glob, because a wildcard negation makes BuildKit walk the whole tree. scripts/ci/build-context.sh, now also a pre-commit hook, asserts that every Dockerfile COPY, every gg guest package, every include_str! tree, every toolchain-installer input and every package stage-tcab-packages.mjs bakes survives every allowlist that can apply, that no allowlist re-includes a wildcard family, and that the context holds no git-ignored file. See building.

New CI gates, and the validator-constants gate is dropped

Section titled “New CI gates, and the validator-constants gate is dropped”

The whole checkout is prettier-clean and gated: npm run lint:format (scripts/format-check.mjs, run by the format job) checks every git-tracked file prettier has a parser for, minus .prettierignore and the frozen versions it derives from .frozen markers, sharded across cores with a content-keyed cache, and npm run lint runs it after the workspace linters.

scripts/ci/spec-vocabulary-check.sh (specvocabulary, also pre-commit) fails a non-frozen version whose prompt.hbs, specs/ or the shared preambles in core name the project, bare tcab, benchmark, test case, evaluation, run record, a review surface, the manifest, or a link to the repository or gallery; frozen hits are reported, not failed. npm run typecheck:validators (validators) runs tsc --noEmit over all 54 case validator projects, scripts/ci/k8s-manifests.sh (manifests) renders every kustomization under overlays/ and cluster/ and checks the deploy set is namespaced and names every image at the registry and the commit, scripts/ci/seeded-contract-check.sh runs on pre-commit and in the pipeline, and the gg version gate runs on tags.

The validator-constants gate, its pre-commit hook, its Azure and GitHub jobs and its README rows are removed as the wrong instrument: an awk scan of TypeScript cannot decide whether a validator reads the build, and the rule it enforced stays in the authoring guide, with Carom and Refract enforcing the boundary through their own ESLint rules. The pre-commit hooks no longer run clippy or rustdoc on commit; scripts/setup-hooks.sh installs cargo fmt --check, the markdownlint and cspell prose gates, the frozen-version check, the audio-pack lint, and the spec-vocabulary, seeded-contract, NUL-byte and build-context gates. The large-file hook’s exemptions for crates/gg/src/sandbox/guests/ and checkers/ are removed so a multi-megabyte generated guest can never be committed again, and the reference-implementation exemptions name both the reference-impl/ and references/<engine>/ layouts plus test-cases/**/showcase/. See building.

The tasks/ issue board and the repository skills

Section titled “The tasks/ issue board and the repository skills”

tasks/ is the issue board, one Markdown file per issue sorted into an area folder (backend, backlog, gg-client, gg-languages, gg-sandbox, gg-sdk, gg-telemetry, gg-tools, metrics, test-cases, v0.7.0) with completed issues moved to a done/ folder beside them. CLAUDE.md states that nothing under tasks/ is authoritative and a landed issue’s durable conclusions belong in apps/docs/, that all commits use Conventional Commits with an imperative subject, and that multi-agent workflows are standing-authorized. The .claude tree gains four PreToolUse hooks (block-done-task-writes, which makes completed issues immutable, block-bare-sleep, block-cargo-test and block-self-matching-pgrep) and eight skills (agent-prompts, analyzing-run-costs, documentation, drain-issues-queue, driving-gg-directly, gg-sdk-documentation, repo-tasks and test-cases). The root README.md is reduced to a pointer at the documentation site.

The @clockwyrks scope, and the evaluation contract kept out of a workspace

Section titled “The @clockwyrks scope, and the evaluation contract kept out of a workspace”

The npm scope of every in-repo package is @clockwyrks (seventeen workspace packages at the release branch), so a seeded workspace’s package.json no longer names the project as @test-cabinet/<package> resolved from ./.tcab/..., which the spec guide forbids a model from ever seeing. The directory the harness vendors into a run repository is .vendor (.vendor/packages/ for runtimes a case declares in packages, .vendor/engine/ for the engine), and the seed commit’s git identity says Clockwyrks; serve_validation_file falls back to the old .tcab/validation/ path for already-collected trees. Anything importing a runtime must use @clockwyrks/voxel-runtime, @clockwyrks/particle-runtime or @clockwyrks/asset-contract.

The voxel and particle runtimes are vendored into a model’s workspace at seed time, and both depended on @clockwyrks/run-record, whose file: closure dragged the evaluation contract into a case’s workspace. The ten rig and F-curve types they need (ModelSpec, PartSpec, JointSpec, JointKindSpec, AxisSpec, DriveKindSpec, InterpSpec, KeyframeSpec, AnimationSpec, AnimationTrackSpec) now live in the new generated package packages/asset-contract, emitted by contract-codegen alongside run-record from the same Rust types; run-record re-exports them and is no longer staged into the host package store at all. The references that vendor a runtime now vendor the contract beside it, and scripts/ci/seeded-contract-check.sh greps the generated package under a strict evaluation-vocabulary ban and walks the real manifests asserting no seeded package reaches run-record. See run records.

Every workspace package except @clockwyrks/gg-sandbox, whose checker is the TypeScript gg’s TS arm judges a program against and stays pinned at 5.9.3, moves to TypeScript 6 and Vitest 5, the five Three consumers and the root overrides to three ~0.185.0, and Playwright, @types/node and jsdom to the newest the workspace accepts, with jest-dom held below 7 and jsdom below 30 for stated reasons; the engines stay at 1.0.0, and the voxel and particle runtime sources changed only by the rename, the contract import and the prettier pass. The root build:packages also builds the four engine packages and the root gains format, lint:format, test:scripts and typecheck:validators. Every reference lockfile was regenerated and proved with a from-scratch npm ci, and the seeded build vitest configs write coverage/test-report.json and coverage/coverage-summary.json.

The Cargo workspace gains crates/code-analysis, crates/gg, ten crates/gg-sandbox-artifacts/<arm> members and their shared build-support, loses crates/desktop, swaps uuid for cuid2 (every id the services mint is a CUID2), adds semver, flate2, wasmtime-wasi 45, wit-component and wit-parser 0.248, tiktoken-rs, oxc and sourcemap, enables ts-rs’s no-serde-warnings feature workspace-wide, and sets opt-level 3 on the wasmtime and cranelift graph in the dev profile. rust-toolchain.toml stays at 1.97.1, with a header explaining that a bump re-cuts gg’s Rust arm rlibs automatically. The pipeline provisions Node 22 while the devcontainer ships Node 24.16.0. See building.

The devcontainer supports four host runtimes and bakes gg’s toolchains

Section titled “The devcontainer supports four host runtimes and bakes gg’s toolchains”

The host variant table is Docker on Linux, rootless Podman on Linux including a DGX Spark, Docker on macOS covering Docker Desktop and OrbStack, and the new Podman on macOS row; .devcontainer/setup-host.sh [variant] [--force] [--print] detects the host and writes the uncommitted compose and env pair, reading the podman socket directory and user-namespace mode off a real machine. The host runtime socket mount moved out of the shared docker-compose.yml into the three overrides that have one, and on macOS with Podman it is a bind-backed named volume resolved inside the machine VM so k3d and the local stack still work.

gg’s eleven program-language toolchains are installed by the image build itself, with postCreateCommand left as an idempotent reconciler, because crates/gg/build.rs reflects every arm’s catalogue as a step of cargo build --workspace. Playwright’s Chromium and its system libraries are baked in so the front-end gate can launch a browser after a rebuild, every per-platform download resolves its architecture from uname -m, gh is pinned (2.97.0, GH_VERSION overrides), the macOS Docker row binds the runtime’s synthesized SSH agent socket, ffmpeg is declared because build-sample-pack.mjs silently produces silent packs without it, and lsof, procps and iproute2 are in the apt list. Pinned versions: Rust 1.97.1, Node 24.16.0, nextest 0.9.140, lazygit 0.61.1.

Testing gg: gates, the offline mock provider and driving gg by hand

Section titled “Testing gg: gates, the offline mock provider and driving gg by hand”

gg’s test suite runs under nextest and carries structural gates rather than only unit tests: an authorship gate that compares each arm’s prepared bytes with the model’s, an isolation gate driven three wide sequentially and concurrently per language against five deliberately broken preparations, a build-count gate proving modules compile once, a docs gate that compiles every fenced program on the language pages through its arm, a prompt gate reading what the templates say, an agreement gate between catalogues and the operations table, and per-arm failure-shape cells. Compiled components are memory-mapped from a test-only disk cache, compiler JVMs run on the C1 JIT sized for two processors, and the JVM and Swift substrate tests were split to stay inside the nextest bound.

Any mock/… model id selects the scripted offline client and TCAB_GG_FAKE_MODEL overrides the provider, so a configuration can be validated for free, and the driving-gg-directly skill documents launching the bare binary from a dev box with scripts/model-windows.sh printing modelWindows, modelModalities and modelProviders for an invocation.

Twenty-seven closed issues give each of the eleven arms program-level coverage of gg’s surface: per arm, a workspace test file drives every files and shell operation and a session-side file drives every context, delegation, programs, docs, views and session operation from a short program, one test for the successful call and one per runtime failure mode the arm’s own SDK documentation declares, plus a feedback test per arm and Java and Kotlin tests driving all fourteen functions of the test-cabinet:gg/math interface. Thirteen more give the tool crate call-level coverage: every name in ALL_TOOL_NAMES reached through ToolRegistry::dispatch is driven with a synthesized ToolCall through a real registry, each loop-intercepted tool is asserted to answer an ordinary dispatch with the gg-defect refusal its declaration promises, and the board, task, memory, filesystem, tree, search and compact tools each gain a test per argument diagnostic, unknown id and store refusal a model can reach. See invariants.

apps/lattice-designer (@clockwyrks/lattice-designer) is a new dev-only React and Vite workspace under apps/: a drag-and-drop designer that lays out a Lattice factory, runs it through the real lattice-core engine, opens and saves the case’s own scenarios in place, imports layouts, and exports a scenario.json. It carries grid presets and recipe explanations, simulates the lane splitter, and stages a board-size change through the themed dialog because it discards the draft. Its README states that it is not part of the shipped product and that a saved scored scenario still needs its oracle re-solved and its playback bundle regenerated.

test-cabinet-telemetry’s Config gains with_stderr_logging(), which writes the human-readable fmt layer to standard error instead of standard output, and the tcab CLI calls it so a subcommand whose stdout is piped (--json) is not interleaved with log lines; services keep stdout. The observability page adds the auth, artifact and arena services to the instrumented-process table and drops the desktop app, and OTLP export remains opt-in on OTEL_EXPORTER_OTLP_ENDPOINT with no new variables. See observability.

The documentation site is rewritten and reorganized

Section titled “The documentation site is rewritten and reorganized”

Every page under apps/docs/src/content/docs was rewritten to the documentation skill’s standard and checked against the code. The site gained an Engines section with one folder per engine split into APIs, concepts, usage, validators and examples and a none page; the asset-generation manifest reference was split from one page into an overview plus per-kind pages; deployment/kubernetes.md split into deployment/kubernetes/ (overview, control plane, run plane, postgres, internal ingress); and development/building.md gained Cloning, Submodule URLs, Mirrors and pins, Cold storage, the gg toolchains, portable builds and validator projects. The docs tree is linted with markdownlint for the first time, the observability page registers PromQL and TraceQL grammars so its query blocks highlight, and the docs deploy from the pipeline’s docs job to test-cabinet-docs on master and to the new test-cabinet-docs-staging project on staging.

Response healing, the replay capability and fidelity axis, tcab gg-run, tcab gg-replay and tcab gg-playback, the ownership param, the planning capability, the tdd and plan-first machines, speculative best-of-K execution and the per-slot provider override were deleted before v0.7.0 and are not part of what ships. Because every configuration container refuses unknown keys, a saved configuration still carrying healing, ownership or planning is refused as an unknown field.

Upgrading to v0.7.0 applies 27 migrations, taking the registered list from 24 to 51. All are additive or defaulted except three: m20260824_000035 drops two probe columns, m20260824_000036 drops and recreates both model-probe tables, so probe history from earlier development builds is discarded, and m20260820_000032_drop_gg_agent_model_slots drops a column.

  • m20260723_000018_add_job_gg_config: nullable job.gg_config_json holding a gg run’s capability set.
  • m20260724_000019_create_gg_config: the per-account gg_config table of saved gg configurations.
  • m20260724_000020_add_model_price_input_modalities: nullable model_price.input_modalities.
  • m20260727_000021_create_comparison: the per-account comparison table.
  • m20260731_000022_add_run_gg_preset: nullable run.gg_preset, the configuration name a gg run was launched from.
  • m20260801_000024_add_run_updated_at: run.updated_at, seeded from finished_at in SQL.
  • m20260801_000025_add_run_code_analyzer_version: nullable run.code_analyzer_version.
  • m20260801_000026_create_gg_saved_views: the gg_saved_query and gg_dashboard tables.
  • m20260818_000030_create_gg_agent: the per-account gg_agent library.
  • m20260818_000031_add_gg_config_agent_sources: nullable gg_config.agent_sources_json.
  • m20260820_000032_add_case_reference_build_engine: case_reference_build.engine, taken into the primary key with existing rows backfilled none.
  • m20260820_000032_drop_gg_agent_model_slots: drops gg_agent.model_slots_json.
  • m20260822_000033_add_run_engine_slug: nullable run.engine_slug, backfilled at boot in batches.
  • m20260823_000034_create_model_probe: the model_probe and model_probe_item tables.
  • m20260824_000035_probe_forced_submission: reshapes the probe tables for the forced submit_program protocol.
  • m20260824_000036_probe_docview_scenarios: drops and recreates both probe tables for the scenario design.
  • m20260824_000037_add_validator_ratings: run.validator_rated (false for existing rows), nullable run.aesthetic, and review.aesthetics ('[]' for existing rows), the reviewer’s per-domain aesthetic ratings.
  • m20260827_000038_add_review_aesthetic: nullable review.aesthetic; a legacy per-domain aesthetics JSON collapses to its worst tier on read.
  • m20260830_000039_add_run_record_readability: run.record_readable (true) and run.record_format (0, so every existing row is re-decided at the first boot).
  • m20260901_000040_add_gg_cell_identity: run.gg_models, job.gg_preset and job.gg_models, backfilled at startup.
  • m20260901_000041_add_gg_config_id: run.gg_config_id and job.gg_config_id.
  • m20260901_000042_create_backfill_state: the backfill_state table recording the one-time gg_config_id attribution pass.
  • m20260901_000043_add_engine_pins: ladder_rung.engine and job.engine_slug.
  • m20260906_000044_add_job_started_at: nullable job.started_at.
  • m20260922_000045_add_provider_pin: nullable model_price.provider_pin (the observed pin) and model.provider_pin (the hand-set override).
  • m20260923_000046_add_model_list_price: five nullable list-price columns on model (input, cached input and output, stored per token and edited per Mtok in the console, the date taken and the source), so every existing row reads as not yet priced.
  • m20260923_000047_add_model_provider_policy: nullable provider policy columns on model (native quantization, input and output price ceilings, banned providers and the unknown-quantization allow list).