v0.7.0 (2026-09-25)
The v0.7.0 release follows v0.6.3 and its centre is gg, a coding harness authored inside the repository and executed inside the run container, which lets the cabinet vary one part of an agent at a time: how it answers a turn (tool calls or a whole program), which of eleven languages it writes that program in, how it compacts its context, what it remembers, and how it delegates. Every gg run streams a typed telemetry record that the console reads natively, so a run can be read turn by turn, request by request, without a second tool.
The second centre is the test cases themselves. A run now names an engine beside its case, variant, harness and model, and eighteen playable cases ship a version on which every checklist point is decided by a validator, so a run’s functional rating is known the moment it completes and a reviewer supplies only the aesthetic judgement. The captured baseline media those validators compare against left the repository for a cold-storage submodule, and every case declares the audio packs its run container is staged with.
Underneath both, every model carries its developer’s list price and every gg run is pinned to that developer’s own provider, so a run’s cost and its session duration become properties of the model rather than of the route a request happened to take.
The cabinet also moved house. Azure Pipelines is the one pipeline, images are pushed to the Test Cabinet container registry by commit sha, gg’s release binaries live on Azure Blob Storage, and the GitHub repository is a mirror. The Tauri desktop app is gone, and the web console is the only console.
This remains pre-1.0 software.
Upgrading
Section titled “Upgrading”The sections below are what an operator has to do or know before the first run
on v0.7.0. Backend and console must deploy together, because the release adds
telemetry event kinds, enum values (canceled, limit_exceeded, the aesthetic
ratings) and RunEvent fields the console reads. The schema changes are listed
at the end of this page.
Every model needs a list price
Section titled “Every model needs a list price”A run of a model can be enqueued only when the model’s catalog entry carries a dated list price: uncached input, cached input and output rates per Mtok, and the date they were taken. This applies to every harness, including Claude Code, Codex and Antigravity, and to every enqueue path (the run form, a gg launch, a coverage top-up and a retry). The seeded catalog carries no list price, so a freshly upgraded deployment refuses every launch until an operator enters one on each model’s edit page. The Fill from OpenRouter control seeds the three rates from the model’s official endpoint for confirmation. See adding or updating a model.
A gg launch additionally needs a resolvable context window and a non-empty
provider candidate list for every model it binds; a model whose developer has
no endpoint on OpenRouter is refused at enqueue, and providerPin on the model
entry names the developer’s provider where its name differs from the model id’s
author segment.
The run-container registry and tag
Section titled “The run-container registry and tag”With TCAB_CONTAINER_REGISTRY unset a run resolves its image from
testcabinet.azurecr.io, the registry the pipeline pushes every run image to.
The compiled default tag remains latest, which the pipeline never pushes. A
pipeline deploy sets TCAB_CONTAINER_TAG to the commit sha automatically; any
run outside that, such as a host-side tcab run or a hand-applied overlay, must
set TCAB_CONTAINER_TAG to a sha that master or staging has published. The
generic staging and prod overlays carry the placeholder REPLACE_SHA.
Run images must be rebuilt
Section titled “Run images must be rebuilt”Every run image must be rebuilt and pushed at the v0.7.0 commit before the
first run, for three reasons. The base image now creates /opt/audio and
writes the .tcab-audio-contract marker, and a full-stack, game-jam or audio
run reads that marker back at container start and fails, naming the image and
the pin, against an image that predates staged audio. Every run image other
than base publishes a <name>-gg variant carrying gg’s language toolchains,
which a gg run resolves unconditionally. And the new
test-cabinet-full-stack-3d image must exist for Gantry, the first case with
asset_dimension = "3d", or every run of it fails on voxel: not found.
containers/build.sh builds a -gg variant through its parent, runs
gg selfcheck inside each variant before pushing, and rebuilds a parent whose
Dockerfile is newer than the image, cascading to its children. The per-image
override for a variant is TCAB_CONTAINER_IMAGE_<NAME>_GG.
The audio store and the Cloudflare secrets
Section titled “The audio store and the Cloudflare secrets”Audio is no longer baked into run images. The driver image copies the published
packs to /opt/tcab-audio from the data-only test-cabinet-audio-store image
named by its AUDIO_STORE_IMAGE build arg, which defaults to scratch, so a
driver built by hand must pass a published store or one built from the checkout
(make -C deployments/local audio-store). TCAB_AUDIO_STORE points a driver at
a store elsewhere. For a local tcab run or tcab validate on a case that
declares packs, run az acr login --name testcabinet and
scripts/fetch-audio-store.sh once, then export the TCAB_AUDIO_STORE it
prints; --stage builds the store straight out of the object store instead.
The pipeline’s audio-store job reads the secret variables
CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_AUDIO_R2_BUCKET,
CLOUDFLARE_AUDIO_R2_PRESIGN_ACCESS_KEY_ID and
CLOUDFLARE_AUDIO_R2_PRESIGN_SECRET_ACCESS_KEY, and the docs job reads
CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID. No run-image or
service-image build needs an audio credential any more. The TCAB_SAMPLE_PACK,
TCAB_INSTRUMENT_BANK, TCAB_SAMPLE_PACK_DIR and TCAB_INSTRUMENT_BANK_DIR
variables are gone from the audio binaries and the images; a palette is chosen
only by the case’s [audio] packs and, host-side, TCAB_AUDIO_DIR.
Where gg comes from
Section titled “Where gg comes from”The driver image bakes gg as a static musl binary at /usr/local/lib/tcab/gg
and core copies it into each sandbox, so a cluster run installs gg with no
network egress. Otherwise core downloads v<version>/gg-<target> from
TCAB_GG_RELEASE_URL, which defaults to the gg-releases container at
https://testcabinetartifacts.blob.core.windows.net/gg-releases.
TCAB_GG_RELEASE_REPO is no longer read. TCAB_GG_INSTALL (local or
release), TCAB_GG_BINARY, TCAB_GG_RELEASE_VERSION and
TCAB_GG_RELEASE_TARGET keep their meanings, and the dispatcher forwards all
five into every driver Job. A prerelease tag such as v0.7.0-rc1 equals no
crate version, so a deployment on one sets TCAB_GG_RELEASE_VERSION by hand.
The backend serves GET /gg/reference from TCAB_GG_REFERENCE; the backend
image bakes the documents at /opt/gg-reference. A backend deployed from
tarballs needs v<version>/gg-reference.tar.gz unpacked and the variable set,
or the console’s gg Reference section answers 503 while everything else works.
OPENROUTER_API_KEY must reach core for injection into the run container, since
gg reads it from the environment and nothing else.
New backend variables
Section titled “New backend variables”TCAB_ARTIFACTS_URL is the in-cluster artifact service address every
backend-originated call uses, distinct from the browser-facing
TCAB_ARTIFACTS_PUBLIC_URL; without it a run delete cannot prune its tree.
TCAB_ARTIFACT_SWEEP_INTERVAL_HOURS (default 6, 0 disables) and
TCAB_ARTIFACT_SWEEP_GRACE_HOURS (default 24) tune the sweep that reclaims
orphaned trees, and the artifact service’s GET /runs and
DELETE /runs/{id}/artifacts require TCAB_BACKEND_SERVICE_TOKEN. Model probes
bill to TCAB_OPENROUTER_API_KEY on the backend, which is distinct from the
runners’ OPENROUTER_API_KEY; without it a probe trigger fails with
openrouter_key_missing. TCAB_VITEST_TIMEOUT_SECS overrides the 45-minute
cap on a case’s whole validator suite run.
Memory limits leave the run plane
Section titled “Memory limits leave the run plane”TCAB_K8S_RUN_MEMORY_LIMIT is removed from the base manifest and the local
overlay and left unset (commented out) in the env examples, and
TCAB_DISPATCHER_DRIVER_MEMORY_LIMIT no longer defaults; a dispatcher test
refuses any manifest under deployments/k8s that sets either. Size
TCAB_K8S_RUN_MEMORY_REQUEST and TCAB_DISPATCHER_DRIVER_MEMORY_REQUEST
(default 2Gi) to the heaviest case’s real peak, because a run pod that outgrows
its request on a genuinely full node is still evicted.
TCAB_DISPATCHER_DRIVER_CPU_LIMIT defaults to 2, which bounds how wide the
post-run toolchain fans out and so the driver’s memory peak; over-limit CPU is
throttled, never killed. See the run plane.
The catalog is re-ingested, and some stored records become unreadable
Section titled “The catalog is re-ingested, and some stored records become unreadable”The definition store format is 2 and stored definitions carry an engine
support set, so the backend holds itself unready and GET /test-cases answers
503 until a whole-catalog ingest rewrites the store. The overlays’ startup
sidecar posts a forced ingest on every backend start, so a pipeline deploy does
this itself; a hand-run deployment runs scripts/reingest.sh. Case showcases
and reference builds keyed by engine appear only after that ingest.
The run-record format is 3. Stored records that lack the gg openingTurn field
or the per-test toolchain report read as unreadable after upgrade, appear only
in the console’s Unreadable worklist at /runs/unreadable, and can be deleted
there; a re-push with a readable record restores one. Saved gg configurations
and recorded gg runs from before stable agent ids do not parse. Scores and
ratings of existing published runs can also change: a run is now scored against
its own version’s and engine’s checklist rather than the case’s latest, a
validator-rated run that did not complete is unrated rather than Flawless, and
the toolchain typecheck gate applies at read time.
Validation baselines live in the cold-storage submodule
Section titled “Validation baselines live in the cold-storage submodule”The captured baseline media is no longer in the tree. Fetch it with
git submodule update --init --depth 1 cold-storage to capture or review
baselines; tcab capture-baselines and tcab publish-reference refuse to run
into an uninitialized submodule and name that command, and
TCAB_COLD_STORAGE_DIR redirects the root. The deployed ingest sidecar fetches
the submodule itself on every start, and a backend ingesting from a checkout
without it serves versions with no baseline media. A submodule commit must be
on cold-storage’s master before its pin is committed here, or the
submodulepins gate fails.
Reference builds for the new case versions
Section titled “Reference builds for the new case versions”test-cases/reference-builds.lock.json is keyed
<env>.<slug>.<version>.<variant>.<engine>, so a reference build is recorded
per engine. A case version’s reference build is recorded with
tcab publish-reference --env <prod|staging> <slug> <version> --all-variants
(after npm ci && npm run build:packages at the root, with wrangler,
CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID), the lockfile is committed
to the branch the environment tracks, and
scripts/reingest-cluster.sh --env <env> is run, before the Reference tab and
the catalog carry it. A reference-capable case has a recorded
reference build by the time the release that makes it non-experimental goes
live. See
publishing a reference implementation.
The pipeline’s prerequisites in Azure
Section titled “The pipeline’s prerequisites in Azure”Before the first pipeline deploy to an environment: grant the arm64 agent pool
pool-dev-linux-arm64-wus3-4c-eph-01 to the project, because the AKS nodes are
arm64 and Azure has no hosted arm64 agents; create the service connections
tcab-acr (Docker Registry, AcrPush on testcabinet.azurecr.io), tcab-deploy
and tcab-gg-publish (Azure Resource Manager, workload identity federation);
upload the secure file github-mirror-key; and set the secret variables named
above. tcab-gg-publish needs Storage Blob Data Contributor on
testcabinetartifacts.
Create the custom role from deployments/azure/aks-command-invoke.role.json and
assign tcab-deploy both “Test Cabinet AKS Command Invoke” on each cluster and
“Azure Kubernetes Service RBAC Admin” scoped to
<cluster>/namespaces/tcab-<env>. Both clusters’ kubelet identities need
AcrPull on the registry. Apply the cluster-scoped bootstrap by hand once per
environment (kubectl apply -k deployments/k8s/cluster/azure-prod or
azure-staging) after installing ingress-nginx and cert-manager through Helm
and creating the Key Vault secrets, and re-apply it whenever that folder
changes.
Configure build validation policies on master and staging in Azure Repos so
pull requests run the gates, because Azure ignores the pipeline’s pr: block.
See the Kubernetes overview.
The desktop app is removed
Section titled “The desktop app is removed”There is no desktop console and no macOS build of tcab. The web console at
apps/web is the only console, and a release’s tcab binaries are the
tcab-linux and tcab-windows artifacts of the tag’s pipeline run; every other
platform builds from source. TCAB_DESKTOP_IMAGE_TAG is no longer read
anywhere.
Manifest changes case authors must make
Section titled “Manifest changes case authors must make”sample_pack and instrument_bank are rejected; the packs a case is staged
with are [audio] packs = ["name@version", ...], required on every non-frozen
full-stack version and every game jam, and a manifest with no [audio] table
takes the pinned four-pack default. On a per-engine version every graded point
must carry validation, domains and a failure_cap, which a legacy
single-workspace version rejects by name. A gg configuration must carry
openingTurn on every responses-as-code agent and a registered language, and
a configuration that still names healing, ownership, planning or a flat
capability set is refused at launch. Frozen versions are unaffected by all of
it.
Devcontainer and hooks
Section titled “Devcontainer and hooks”Devcontainer users should rebuild the image, since gg’s toolchains, Playwright’s
Chromium and several apt tools are baked in, and re-run
.devcontainer/setup-host.sh --force, because the host-runtime socket mount
moved out of the shared compose file and a copied docker-compose.local.yml
silently loses it. Re-run scripts/setup-hooks.sh: the hooks gained the
audio-pack, spec-vocabulary, seeded-contract and build-context gates and lost
clippy and rustdoc. The GitHub repository is a force-pushed mirror, so nothing
should be committed or opened as a pull request there.
Features
Section titled “Features”gg: a harness authored inside the repository
Section titled “gg: a harness authored inside the repository”gg (crate test-cabinet-gg, binary gg, version 0.7.0) is the first harness
authored inside the repository rather than integrated from a third party. Unlike
the CLI harnesses, gg is the executor: core invokes gg --config <invocation>
inside the run container, and the binary holds the model client, the agent turn
loop, tool dispatch and the telemetry emitter. Its one secret,
OPENROUTER_API_KEY, comes from the environment core injects. The orchestrator
dimension does not apply, because gg continues within one session through
compaction, and the metadata tab of a gg run shows its capability set in place
of an orchestrator.
A run is configured as a named, account-scoped configuration: a capability set
of one or more per-agent profiles, each with its own capabilities,
implementations, params, model binding, prompt-cache lifetime, reasoning setting
and loop-detection knobs. It launches from the ordinary New run page by picking
gg in the Orchestrator selector, for every test type. Twenty-two capability ids
are registered: shell, read-file, write-file, edit-file, list-dir,
search, context-window-override, autoload-specs, agent-persistence,
skills, memories, tasks, compaction, agent-managed-context,
project-management, subagents, fsm, exec, fork, responses-as-code,
program-library and docview-close.
gg checks the whole set before the first turn and refuses the launch on any
value it cannot honour exactly as written, naming every defect at once and
substituting nothing: an unknown capability id, an unknown param or field, a
required value left absent, a mock/… model outside tests. It exits 0 on a
natural end, 3 when any agent breached a ceiling (recorded limit_exceeded,
never retried) and 1 on a launch failure, a refused credential, a
provider-failed root call or a defect of its own (harness_error, retryable).
See the gg overview and
configurations.
Responses as code
Section titled “Responses as code”An agent whose type is RaC (capability responses-as-code) answers each turn by
writing a whole program over gg’s typed SDK, submitted as the program string
of a forced submit_program tool call. That is the only tool such a request
offers, pinned through tool_choice where the provider accepts it and auto
where it refuses with a 400, downgraded once per model per run. gg runs the
string exactly as sent in a wasmtime component, with no fence stripping, no
repair pass and no prologue or import written around it. The language’s own
compiler or parser accepts or refuses it; a rejected program is a Compiler error
turn carrying the compiler’s diagnostic and nothing else, and an uncaught throw
or sandbox stop is a Runtime error. A reply with several submit_program calls
runs all of them in order and counts at most one error.
Views are the only channel back to the model. gg.views.openFile, openText
and openDocsView and gg.docs.search push one message per open view into the
next prompt; logged lines go to the operator and return values are discarded.
The system prompt names only module briefs, so an agent discovers functions by
searching, and a per-agent opening turn seeds the fresh window with a module
listing, chosen documentation views and optionally a workspace tree. Every SDK
function is compiled and callable whatever the run granted, and a withheld call
fails at run time as unavailable, naming the call to make instead. Sessions
end only through role-shaped calls: gg.session.finish, or a reviewer’s
approve and requestChanges.
Four params are required per RaC profile: language, timeoutSecs (the guest
execution ceiling, enforced by wasmtime epoch interruption and excluding time
parked in host calls), maxMemoryBytes (the linear-memory cap) and
docViewTypes (independent return, parameters and errors flags; gg
refuses an absent or partial set, and the console seeds a fresh capability with
return and errors on and parameters off). Guests get the full WASI p2
surface with / preopened, stderr is captured with 4 KiB kept at each end,
GG_SANDBOX_DEADLINE_MS is exported, and the ECMAScript guest shadows
setTimeout, fetch and their kin with named throwers. Composed calls stream
as ordinary ToolCall/ToolResult pairs with ids program:{ordinal}:{tool}.
See responses as code,
programs and
views.
The program-library capability keeps the source of every program gg ran under
a cuid2 id minted when the submission is acknowledged, and the acknowledgement’s
body is that bare id. programs.get(id) takes the id as a required argument, a
miss is not-found naming the ids held and the turn each ran in,
programs.rerun(source) lets a program patch and hand over a predecessor, at
most four programs per submission (the model’s own plus three hand-overs), a
fork clones the library and an exec or FSM edge moves it. A per-agent idLength
param (default 4, range 2 to 32) sets the id length. See the program
library.
Eleven program languages, each with a hand-written SDK
Section titled “Eleven program languages, each with a hand-written SDK”The language param is a required per-agent configuration axis with eleven
registered values. typescript is type-checked by tsc, erased and evaluated
as an ES module by the quickjs-ng ECMAScript guest; javascript uses the same
guest with no compiler; python runs on a guest carrying CPython 3.14; ruby
is compiled to JavaScript on the host by an embedded Opal and evaluated by a
guest carrying Opal’s runtime; purescript goes through purs and a bundle to
the ECMAScript guest; java and kotlin are compiled by javac or kotlinc
and then TeaVM inside a pooled warm JVM into one wasm component per turn;
rust is rustc to a component per turn; swift is swiftc; cpp is
clang++ with every #include written by the program; and csharp is Roslyn
to IL, run by a guest holding Mono’s interpreter.
All arms bind one language-neutral WIT world (crates/gg/wit/gg-sandbox.wit,
one interface per family) so they differ only in spelling. The surface is 52
operations in thirteen modules (files, shell, board, tasks, memories,
docs, views, context, delegation, skills, programs, session and
core), each bound by a gg capability and the agent’s allowlist, by the
agent’s ending role, by its machine state, or always. Every SDK is static: every
function is compiled, linked and callable in every program, a call the agent
was not granted is refused at the membrane with a typed ApiError of code
unavailable whose message names the alternative in the program’s own
language, no gg name is reachable without an import line the model wrote, and
every call is synchronous and typed rather than JSON. A failure reaches the
model in the language’s own words at the model’s own line.
Signature catalogues and guest components are generated by the build, reflected
from each SDK’s own documentation comments by the language’s own documentation
generator, and never committed. The run’s root language lands on
summary.programLanguage and each instance’s on its agent_surface event, so
a cross-language study slices on one scalar. See
languages, static SDKs
and the ECMAScript guest.
A model finds the SDK by searching it
Section titled “A model finds the SDK by searching it”gg’s prompts name no catalogued function. The docs family (gg.docs.search
and the calls that open and close documentation views) is published on all
eleven arms in each arm’s own spelling, search indexes parameter names and
descriptions so a query for an argument finds its function, and search reaches
three kinds, module, function and type, with a module leading a tie. A
module’s view is its path, the line a program writes to bring it into scope, its
prose and a line per function this agent binds, and every function and type view
carries a line stating where the symbol is defined and the exact import line
that reaches it. Search and views are filtered by the same predicate the
membrane refuses calls with, so what a model can find and what gg will service
are one set.
Every arm’s files module gains tree (and a tree tool under list-dir for
tool-calling agents): a depth-bounded rendering of the directory beneath a root,
depth 2 by default and at most 10, honouring the same ignore files search does,
bounded at 1,000 entries and 16 KiB. A RaC agent’s opening turn can open a tree
at a configured depth, so a program reads the layout instead of guessing at a
path. See the API surface and the
filesystem.
Compilation is measured and isolated, and a compiler crash is gg’s fault
Section titled “Compilation is measured and isolated, and a compiler crash is gg’s fault”Each arm declares a checker (tsc, opal, purs, javac, kotlinc,
rustc, swiftc, clang++, csc, or none for JavaScript and Python),
interpolated into the system prompt. Compile time is recorded apart from the
sandbox clock as compileMs on every code_execution event and
summary.compileMs run-wide, charging a turn for its program, every hand-over,
and the code skills, memories and on-use scripts it was first to load. A
compiler that rejects a program yields a transpile_syntax,
transpile_compile or transpile_unsupported error turn with the diagnostic
verbatim, bounded by deletion only: eight diagnostics on most arms closing with
a dropped count, and four errors with three notes each on C++. A compiler that
could not finish (a crash, its own timeout, a binary missing from the image)
tells the model nothing, counts against no ceiling and ends the run under
internal_error.
Every agent holds one private compile workspace for its lifetime (a working
directory, an output directory, a redirected HOME, XDG_* and TMPDIR). A
response is compiled once, a loaded code module is compiled once per session
under its binding key and linked from recorded output on later turns, and a
source-level gate forbids Command::new and temp_dir outside the seam. A
warm-up compile of each language’s guest component fires once per run
concurrently with the first model request, and long-lived compiler JVMs are
pooled. Compilers gg carries itself (tsc, Opal) run through Node
(TCAB_GG_NODE); the rest come from the run image’s toolchain tree. See
compilation.
Every run image gains a gg variant, gated by gg selfcheck
Section titled “Every run image gains a gg variant, gated by gg selfcheck”A gg run resolves the <name>-gg variant of whatever image the test case would
otherwise use: the parent image plus one COPY of the toolchain tree built by
containers/gg-toolchains/Dockerfile at /opt/gg. containers/image-names.sh
holds 55 names, 28 parent images and 27 variants, derived by suffix rather than
listed, so a gg run of any test type resolves an image it can compile in while
no other harness’s image carries the toolchains. Each toolchain carries every
shared library the run images do not supply (the C# tree vendors the ICU trio,
the wasi-sdk tree vendors libtinfo), because the variants form four distinct
library environments.
gg selfcheck [--language <id>] drives every arm’s real bootstrap turn in
whatever environment the binary runs in and exits non-zero on any broken arm.
containers/build.sh runs it inside each built variant before pushing,
scripts/ci/run-images.sh requires each representative’s gg selfcheck ok:
line by name, and make -C deployments/local run-images-gg-selfcheck does the
same locally. core::harness::resolve_run_image now takes the harness slug and
RUN_IMAGE_OVERRIDE_ENVS is a LazyLock<Vec<String>>, which breaks in-repo
callers only. See selfcheck.
The gg model client chooses the provider for every request
Section titled “The gg model client chooses the provider for every request”gg talks to every live model through OpenRouter but chooses the provider
itself. The backend builds an ordered candidate list per bound model at enqueue
(native quantization, priced at or below the developer endpoint, cache-read
priced, supporting tools, tool_choice and reasoning as needed, not banned;
the developer first, then by fault rate and price) and every request names
exactly one candidate through provider.only, provider.quantizations and
allow_fallbacks: false. The run moves to the next candidate when the retry
schedule is spent, at once on a 404 (a provider excluded by account privacy
settings), or after providerCacheMissLimit (default 2) unexpected cache
misses; each move is a provider_switch event and each stall or miss a
provider_fault. A reply served by any other provider ends the run as
model_provider_mismatch.
gg mints one cuid2 routing key per run and sends it as both session_id and
prompt_cache_key on every request of every agent. For the Anthropic family it
stamps up to four cache_control breakpoints (an anchor, two grid-snapped
rolling points and the tail) and serializes every message as a one-element
content array so a moving marker never changes wire shape; other providers get
no markers. Each agent picks a prompt-cache lifetime (five minutes by default,
or one hour) and an optional reasoning setting (effort from xhigh to
none, or maxTokens), sent as OpenRouter’s unified object and never applied
to a handoff summarizer’s model. Every request streams with usage included.
Failed calls (429, 5xx, transport, stall) retry on a configurable schedule:
maxModelRetries (default 10) against modelRetryMaxDelaySecs (default 60),
starting at one second and doubling, with Retry-After honoured.
modelCallTimeoutSecs (default 900) bounds one attempt and yields a
model_timeout error turn, and modelStreamIdleSecs (default 60) cancels a
stream that delivers no delta. A refused pinned tool_choice is re-sent on
auto and remembered run-wide, and requests carry gg’s own name and site as
OpenRouter attribution. See execution limits.
Loop detection abandons a reply that has started repeating itself
Section titled “Loop detection abandons a reply that has started repeating itself”A per-agent loopDetection setting watches delta.content as it streams and
drops a reply when minOffenders distinct words each exceed repeatThreshold
occurrences in a rolling windowWords window for minSaturatedRun consecutive
words, or when it passes maxResponseChars; the console seeds 256, 32, 2, 3000
and 250,000. The sustained-run term is what keeps a 1,500-entry tilemap literal
from tripping it. A trip is handled like a 5xx inside the client’s retry loop:
the stream is closed unread, a fresh detector judges the retry, and nothing of
the discarded reply enters the context or the session record. A model that
loops on every attempt costs an error turn (model_response_loop), not the run.
Discarded attempts ride the turn_outcome event as loopAborts,
loopAbortWords and loopAbortChars and roll up on the summary. Because a
cancelled stream never delivers usage, gg keeps each abandoned reply’s
generation id and, once at session end, looks it up on OpenRouter’s
GET /api/v1/generation ledger (concurrently, a 404 retried for up to 30
seconds) and adds the price to the slot’s slotCosts and the run’s total cost.
loopAbortUnpriced counts the ones it could not price, so a deployment whose
egress blocks that endpoint records them rather than failing the run. See
loop detection.
Execution limits, one error definition and cancellation
Section titled “Execution limits, one error definition and cancellation”capabilitySet.limits carries five ceilings that are armed only when written:
maxTurns per agent (exhausted), maxRuntimeSecs and maxCost run-wide
(timed_out and limit_exceeded), and maxConsecutiveErrors and
maxErrorRate with errorRateWindow per agent (limit_exceeded). Two keys are
required, maxParallel (the run’s global agent pool) and replayMaxBytes (the
capture journal bound), and five carry defaults: modelCallTimeoutSecs 900,
modelStreamIdleSecs 60, providerCacheMissLimit 2, maxModelRetries 10 and
modelRetryMaxDelaySecs 60. A fresh configuration seeds five consecutive errors
and a 0.2 error rate over 50 turns, and the turn, runtime and cost ceilings
start unarmed. Nothing about a ceiling or the detector is ever told to the
model.
A turn is an error only when its declared work could not be carried out: a
failed, timed-out, length-capped, unparseable or looping model call, a program
that did not compile or threw uncaught, a sandbox stop, or a tool-calling turn
that requested nothing. A failure reported back into a program that carried on
is not. The judgement is made once per turn at the seam where an outcome is
recorded against an agent’s ceilings, emitted as turn_outcome with a base
error kind and one of twenty-one specific errorType leaves, and rolled up on
the summary as errors. Failures of gg’s own machinery disqualify the run under
internal_error with exit 1, releasing every suspended wait, and a refused
credential is auth_error.
The host cancels a run by creating the cancelFile path named in the
invocation. Every agent checks it at its turn boundary, the epilogue (rollups,
session record, summary) runs in full, and the status is canceled. See
execution limits and
turn outcomes.
Context accounting
Section titled “Context accounting”Every gg run tags each window contribution with one of fifteen sources (System,
User prompt, Assistant, Tool output, Compiler errors, Runtime errors, File
views, Agent views, Documentation, Doc search, Skills, Memories, Task list,
Board and History), estimated with one BPE tokenizer, and streams the breakdown
every turn. The denominator is the bound model’s window from the backend model
catalog, fetched live from OpenRouter for an unseen model and pushed in as
modelWindows; there is no default and no guess, and backend, core and gg each
refuse a run without one. The context-window-override capability’s
windowLimit is a ceiling, measured against the smaller of it and the model’s
window, which lets one configuration exercise compaction across models of
different sizes.
The prompt is append-only to preserve provider prefix caching. Rebuilt state
blocks (memories, tasks) keep their position when unchanged and are superseded
rather than removed when changed, the system prompt and the context-usage
signal are slots, each read_file appends a snapshot while re-opening a view
selector supersedes it, and openDocsView of an already-open key does nothing.
read_file on a PNG, JPEG, GIF or WebP (detected by magic number) returns the
picture; a model declared without image in modelModalities is never sent
one, an unknown model is tried and denied run-wide on refusal, and an inline
image is charged by its dimensions.
The message log fingerprints every message and streams its body once as a
context_message event, with each turn a prompt event holding ordered
pointers into that pool plus the reply, the serving provider, finish reason,
usage and latency; images are logged as descriptors, never bytes. The console
renders a per-agent Requests file from it, and the Agents panel’s context spend
folds it turn by turn: each turn’s input tokens and their cost are split across
the request’s messages in proportion to estimated size and summed per message,
band, view and tool, so what a file cost to keep in the window over the whole
run is ranked rather than what it cost to bring in once. See
context visibility and
context spend.
Compaction with five strategies
Section titled “Compaction with five strategies”The per-agent compaction capability summarizes a filled thread and restarts
it, carrying pinned state verbatim: used skills, tasks, memories, locked
autoloaded specs, re-derived documentation views, and files a compact call
named. summaryHeadroom (0.0 to 0.9, required) is held back from the window
for the summarization round trip and sets the trigger at 1 - summaryHeadroom.
The implementation names one of five strategies: the in-loop
self-summarization, self-compaction (the agent calls compact with a
summary and up to twelve files to re-read) and memory-compaction (working
state must be written as memories, so a writable memories capability is
required), or the between-turn handoff-summarization and handoff-compaction
on an optional model param that may defer to one of the agent’s slots.
A boundary that leaves the window still at the trigger is retried maxRetries
times and then ends the agent under the failure status compaction_failed. A
failed condensation restarts from a fixed note flagged as a fallback. Each
boundary records the strategy, the trigger fullness, tokens before and after,
per-source composition and the summary, and fires the pre- and post-compact
hooks. The companion agent-managed-context capability renders a per-turn
context-usage signal and grants evict_file_view, archive_thread,
search_archive and, under RaC, gg.views.close; docview-close separately
grants gg.docs.close and closeAll. See compaction and
agent-managed context.
Shell and filesystem tools
Section titled “Shell and filesystem tools”The shell capability runs sh -c in the workspace, or the agent’s worktree,
in its own process group under a per-call timeout that defaults to 3600 seconds,
clamped to 24 hours and to the run’s remaining wall clock; a non-zero exit is a
result, not a failed call. Its implementation is inline (the whole output,
capped at 16 KiB) or offload: both streams of every command are written to
/tmp/gg-shell/cmd-*.stdout and .stderr, only the last maxLines and
maxChars return inline, and a truncation note names the files with each one’s
line count, p50, p95 and p99 line lengths and five longest lines so a model can
aim a windowed read. One shell telemetry event per command (origin tool,
program or hook, directory, exit code and 16 Ki-character tails) backs a
Shell file per agent in the console.
The filesystem tools are one capability each. read-file is unlimited or
default-cap with a required lineCap, pages with offset and limit under a
[showing lines a-b of N; continue with offset: c] footer, backstops at 256
KiB, and reads images. write-file, edit-file (exactly one occurrence) and
list-dir (list_dir plus tree) complete the set, and the default-on
search is a Rust-syntax regular expression over the workspace honouring
.gitignore, .ignore and .git/info/exclude, 50 hits by default and 200 at
most, with lines clipped at 200 characters. Paths resolve as any process in the
container would, relative against gg’s working directory and absolute as
themselves; the container, not the workspace, is the boundary, so the shell’s
own offloaded output is readable. See the shell,
the filesystem and
shell commands.
Skills and memories, prose or code
Section titled “Skills and memories, prose or code”The skills capability loads a per-agent catalogue from a dir param
(.gg/skills by default, stood up empty by gg) and lists every skill by name
and description in the system prompt; read_skill brings one in, and a used
prose skill is retained across compaction. Under RaC a skill or a memory may
also carry a code module, compiled once when first used and importable as
lib:<name> by every later program (import * as csvTools from "lib:csvTools"
in TypeScript), plus an on-use script that runs once after the turn’s program
and opens documentation views of the module’s exports. Loaded code costs no
tokens, compaction never touches it, and a module that fails reaches the model
as a module error naming the binding key rather than as the program’s own
failure.
gg ships twelve generated built-in skills, one per function family
(gg-filesystem, gg-shell, gg-project, gg-tasks, gg-memory,
gg-skills, gg-context, gg-delegation, gg-docs, gg-views,
gg-programs and gg-session), selected per agent by a withholding
builtIns object where {} offers all and an authored skill of the same name
wins. The memories capability stores model-curated notes inside gg under one
of three strategies (scratchpad, markdown, keyword-search) with six
required limits (maxCount, maxLenPerMemory, maxTotalLen, maxLenIndex,
maxLenDescription and maxResults, 0 disabling one), refuses rather than
truncates a breaching write, supports isolated, inherited and read-only scoping
and linked instances between agents, and records every revision. See
skills, memories and modules.
Delegation and process
Section titled “Delegation and process”Every profile carries a roster of the profiles it may use, each entry scoped
subagent, implementer or reviewer. subagents grants spawn_subagent,
wait_for_subagents and send_message bounded by maxDepth; all agents share
one runtime scheduler under the run’s maxParallel, where a suspended agent
releases its slot and resuming beats starting. agent-persistence makes a
profile a single long-lived worker whose next instance re-opens the previous
one’s file, text and documentation views. exec replaces the running instance
with another profile, transferring its modules (history, memories, tasks,
board, skills, archive) live, and fork runs a private copy. See
subagents, agent persistence and
fork and exec.
fsm profiles are model-less state tables authored in the console, each state
running another profile and each transition naming the modules the next state
inherits. project-management gives the run one board of epics (three- to
six-letter prefixes) and issues (AUTH-1, ISSUE-2) with scoped briefs, a
blocked-by DAG, auto-dispatch to implementers named AUTH-1.0i in isolated git
worktrees merged by a merge agent, reviewers AUTH-1.0i.0r whose approve and
requestChanges verdicts gate acceptance, and wait_for_issue; the board is
reachable through tools only and costs no context. tasks is a per-agent
blocked-by DAG (simple or issues mode, maxTasks) pinned in the prompt and
carried across compaction. See FSMs,
project management and tasks.
Hooks run operator commands or scripts at ten lifecycle points (two session
events on the set, eight per agent). A hook can block a write, a shell command
or a stop, or put text in front of the model, and three ship built in: trace,
refuse-empty-write and guard-destructive-shell. See hooks.
Prompts as templates, autoloaded specifications and the Reference
Section titled “Prompts as templates, autoloaded specifications and the Reference”Everything gg says to a model is a Handlebars template under
crates/gg/templates/ embedded in the binary: one system prompt per execution
mode assembled from per-capability partials and a per-language segment, the
pinned task and memory blocks, dispatch briefs, and in-loop messages. Each
profile may add custom instructions or override the whole template. The
autoload-specs capability seeds an agent’s opening context with every file
the test case provided, specs in seeded order and then reference images as
pictures when images is on and the model can see them, synthesized as
well-formed reads and optionally locked across compaction. crates/gg/prompts/
ships operator-pasted system prompts for baseline, layered and studio dev-team
configurations plus common agents. See prompts and
autoload specifications.
gg reference [--out DIR] projects the exact tool definitions and every arm’s
documentation views, which the backend serves at GET /gg/reference from
TCAB_GG_REFERENCE, so the console’s Reference section (gg → Reference, with
Tools and API tabs) cannot drift from what a run sends: all 37 tools with their
live schemas and, per arm, every module, function and type with its
documentation view body. gg probe-fixtures writes the per-language turn-1
fixtures the backend’s model probes replay. See the reference.
An always-on session record, salvaged from a hung run
Section titled “An always-on session record, salvaged from a hung run”Every gg run appends a JSON line per event to .gg/replay.ndjson inside the
container (model I/O, tool results, the prompt frame, shell and git
subprocesses, cancel probes and clock reads, per agent), bounded by the required
replayMaxBytes limit. The host folds it into the run tree’s replay.json.gz
after the run and mirrors it to the backend at GET /runs/{id}/replay. Capture
is not configurable. Before tearing down a run that hung or outran its cap, the
engine copies the journal alone out of the container and assembles it with
truncation reason session_killed, so the one outcome that collects nothing can
still be explained. See the session record and
session records in analysis.
gg’s telemetry stream, and the console built from it
Section titled “gg’s telemetry stream, and the console built from it”gg emits a richer stream than the shared event contract, and the backend and
console read it natively, live and after the run. It opens with
session_started, announcing the resolved capability set and routing key, and
carries per agent agent_spawned, agent_modules, agent_surface (execution
mode, language, bound module stores and the tools or API functions offered,
each with the operation it serves), usage deltas per model call,
turn_outcome, turn_timing, context_message and prompt,
context_breakdown, code_execution, shell, tool_call/tool_result and
api_call/api_result, provider_fault and provider_switch,
response_rejected, limit_exceeded, slot_usage, compaction records, and
one session_summary immediately before session_ended. Every usage delta is
attributed to the model and profile that spent it.
A live run is watched at /runs/gg/:jobId/live and a finished one read at
/runs/:runId/gg, both rendering the same panels over one reduction of the
stream. The Dashboard carries status (with the orchestrator’s setup stage named
before gg’s first event), the token and cost tally with caching and reasoning
rings, turns, tokens per second, an errors row, four clock tiles for runtime,
limit, active and waiting, and per-profile and per-model spend bars. The Agents
panel sums each configured profile’s instances; the Instances explorer is a
per-agent filesystem (overview, prompt, tools or apis, context, requests,
metrics, shell, compaction, a programs file for RaC agents and a modules
folder) that renders successions, forks and machine states as one lineage; the
Modules view reads a run by the stores it holds; and the Project tab is the
board with issues rendered as Markdown. A dropped live stream resolves the run’s
outcome or re-joins instead of surfacing a raw TypeError.
turn_timing partitions each turn’s wall clock into promptMs, requestMs and
responseMs, which the console renders as a stacked bar per turn heading an
agent’s Metrics file, windowed to about fifty bars, with four per-request line
graphs beneath it: tokens per second, cost per request, cache-read share and
reasoning share. code_execution reports per program ok, toolCalls,
apiCalls, durationMs, the error, the capped logs, undocumentedCalls (calls
written without an earlier documentation view) and the two compile figures. A
RaC agent emits an api_call/api_result pair per call its programs make,
naming the operation and no arguments, so a function the model never used reads
a real 0x in the console. See the telemetry
overview, the console, turn
timing, code
execution and the agent
surface.
Two cost figures, and a bounded usage split
Section titled “Two cost figures, and a bounded usage split”A gg run’s session summary carries its spend as two figures, per
(profile, model) slot on slotCosts and run-wide as cost and workCost.
The total cost sums every request that reported a price; the work cost sums
only the turns whose reply produced a program gg ran or a tool call gg
dispatched, decided by the same rules dispatch applies. Every usage delta
names the figure its turn fed in a figure field, so summing the deltas marked
work reproduces the work cost and summing all of them reproduces the total,
and the maxCost ceiling reads the total. A length-capped reply is rejected
whole, recorded on a response_rejected event and summed in
rejectedResponses, inside the total and outside the work cost, and the
compaction summarizer’s own calls stay outside both.
Each usage event carries the four normalized token classes, the call’s cost
when the provider reported one, the serving provider, and wire, the
provider’s usage object verbatim, so a disagreement between the provider’s
numbers and the recorded split can be read off the row. gg bounds the
provider’s output/reasoning split by the reply it received: output is at least
gg’s own estimate of the reply’s text and tool calls, reasoning is the remainder
of completion_tokens, and a row whose split gg had to rebuild carries
reconciled: true. The summary’s providerStats slices calls, tokens, cost,
turns, working turns, errors, stalls and cache misses per (provider, model),
and its maxResponseChars and maxResponseOutputTokens fold over every turn
except a length-capped one. See usage.
A run’s session is measured apart from its setup
Section titled “A run’s session is measured apart from its setup”runTimeSeconds began before reference rendering and host seeding, so the
figure described the host more than the model. Every run record now also carries
metrics.sessionSeconds (the harness session alone, the figure that describes a
model), setupSeconds (through container start, probe, harness install and the
case’s init, with the cluster queueing wait subtracted), teardownSeconds
(the remainder, so the three sum exactly to runTimeSeconds) and
validationSeconds (outside the run time, null on a canceled run). Each stage
is null on a record written before the stages were measured, distinct from
0.
Because every run of a model is now pinned to the model’s own provider, the
session duration is a property of the model, so the test case Metrics tab
charts it with the same bar and scatter control as tokens and cost, the
Leaderboard tab tabulates its mean, the model overview folds it, the comparison
detail page summarizes it per arm with median, spread and mean, the gg Discover
dashboard reads metric.sessionSeconds, and the public site’s case and model
pages show the same figures. The run’s Metrics page lists the four stages, and a
run with no session figure contributes nothing. See
metrics.
Comparable cost is priced at the model’s list price
Section titled “Comparable cost is priced at the model’s list price”Every model’s catalog entry carries an operator-entered list price: uncached
input, cached input and output rates per Mtok from the developer’s own pricing
page, and the date taken (listPriceInputPerMtok,
listPriceCachedInputPerMtok, listPriceOutputPerMtok and listPriceAsOf on
POST and PUT /models, saved together or refused with 422). The backend
resolves the list price at enqueue and stamps it on the launch, so a run is
scored at the price its model carried when it was queued, and the comparable
cost is derived from the token classes and that price with reasoning at the
output rate. A class carrying tokens with an unknown rate makes the cost
unknown, and a run enqueued before the catalog carried list prices records an
unknown comparable cost. In v0.6.3 the comparable cost was computed from the
per-token prices OpenRouter listed at run completion.
The actual cost is the harness’s own figure where it reports one (Claude Code’s
total_cost_usd) and equals the comparable cost otherwise, and every harness
metrics page was rewritten onto this rule. Beside the list price the backend
keeps a billed-rate history observed from the model’s official OpenRouter
endpoint, recorded on run completion, on a 24-hour refresh, and missing-only
when a model is saved, first enqueued or at backend startup; the model’s Stats
tab shows list price and billed rate side by side with their difference, and
the billed rate never rewrites what a run is scored at.
GET /models/openrouter?slug= seeds the three rates from the official endpoint
for the operator to confirm. See metrics and
the backend API.
The catalog is the single store of model facts
Section titled “The catalog is the single store of model facts”The new-model form leads with the OpenRouter slug and a Fill from OpenRouter control that fills the display name, provider, description and the list-price seed figures from the model’s endpoints listing. Each billed-rate observation now also records the model’s context window, release date and input modalities. A gg run’s context window is resolved from the catalog at enqueue for every bound model and pushed onto the launch, fetched live from the per-model endpoint when the catalog has not observed the model, and a launch whose window cannot be resolved is refused per index in a batch. Input modalities travel the same way so a text-only model is never sent a reference image, and the model page shows them under Specs.
On top of the developer pin the backend builds each bound model’s provider
candidate list at enqueue from the endpoints listing, filtered by the catalog
entry’s policy (a hand-set native quantization, a price ceiling used when no
developer endpoint is listed, banned providers, and providers accepted despite
unknown quantization) and ordered by the fault rate recorded across every
stored gg run. A model whose list comes out empty refuses the enqueue naming the
filter that emptied it. GET /models/{slug}/candidates shows the list the next
enqueue would build, the model edit page exposes the policy fields, the
candidate table and the cache-miss limit, and provider names compare case- and
punctuation-insensitively. See
adding or updating a model.
Model probes and the Providers tab
Section titled “Model probes and the Providers tab”A model probe is a responses-as-code readiness check run from the model’s
Probes tab (/models/:modelId/probes). The backend replays gg’s RaC turn-1
request against the model through OpenRouter with the submit_program tool
offered and forced, over two scenarios (baseline and missing-docview) across
several prompts, per program-language arm or across every arm, and reduces the
results to a ready or not-ready verdict at 80% of scored calls per group.
Probes are append-only dated history in model_probe and model_probe_item,
never reach the public snapshot, and run inside the backend process, so a
restart fails a running probe. The endpoints are POST and
GET /models/{slug}/probes, GET /model-probes/{id} and
GET /models/{slug}/probe-providers.
The Models section gains a Providers tab (/models/providers) fed by two open
reads. GET /stats/providers reports per provider and model the calls, tokens,
cost, length-capped rejections, stalls, unexpected cache misses and turn error
breakdowns from stored gg session summaries, kept apart from probe evidence, and
GET /stats/model-accuracy reports per-model RaC turn validity and tool-calling
dispatch outcomes, with approximate older folds marked. See
probe a model.
gg launches from the New Run form under a saved configuration
Section titled “gg launches from the New Run form under a saved configuration”gg is registered as a first-class harness subject (harness_slug: gg, routed
through OpenRouter, outside the third-party catalogue). A gg run is launched
from a saved configuration: gg_config rows per account with REST CRUD under
/gg/configs, edited from the account section’s gg Configs tab at
/account/gg, while POST /gg/runs enqueues a run whose capability set is
lifted onto job.gg_config_json and whose configuration name and id are lifted
onto the run row so the run log’s MODEL / CONFIG column, its search and its
model sort agree. Agent profiles are authored once in a library (gg_agent,
/gg/agents, the gg Agents tab), and a configuration imports one as a live
reference with per-field overrides recorded beside its resolved set. Each
profile carries a stable id and its name is display text.
The New Run form’s gg model picker is scoped to the OpenRouter family, a
launched gg run is tracked in the runs list immediately, and the configuration
editor exposes the run limits, per-agent reasoning setting and prompt-cache
lifetime, an Agent type selector (Tools, RaC or FSM), an Opening Turn tab that
chooses from what the APIs tab enabled, an FSM state editor, module ownership
and memory scope, and hooks. A responses-as-code agent’s profile must carry an
openingTurn naming the modules the seeded first program lists and the
functions whose documentation it opens; a module is listable only when the
agent holds at least one of its functions, an unknown or role-bound entry
refuses the launch, and an empty opening turn seeds no program. See
configurations and agents.
Coverage plans and ladders schedule gg configurations
Section titled “Coverage plans and ladders schedule gg configurations”A plan’s, ladder’s or kind = "combo" group’s member is now one of two shapes:
a harness combination as before, or a gg configuration the account has saved
plus a model for each launch slot it declares. A gg cell is identified by the
configuration’s id and the models the bound set runs on, lifted onto
run.gg_config_id, job.gg_config_id, run.gg_models, job.gg_preset and
job.gg_models, with a one-time startup attribution of older runs recorded in
the new backfill_state table. Every enqueue path refuses a configuration the
launching account does not own, and the runs listing gains a ggConfigId
filter. A top-up lowers a gg member through the same code POST /gg/runs uses,
so a scheduled and a hand-launched run cannot drift, and a member whose
configuration was deleted, leaves a slot unbound, binds a model with no window
or list price, or imports a saved agent edited since is reported unlaunchable
and skipped. gg gains a harness parallelism lane.
A pinned case now names an engine after its variant (ladder_rung.engine and
job.engine_slug; a pin with no engine means the engineless run, so stored
plans keep their counts). The review buffer target is a tagged shape,
{kind: bounded, runs} (clamped to 500) or {kind: unbounded}, on the account
setting and the per-plan or ladder override, so running through everything is
a native instruction. The plan page is three tabs under one layout route
(/account/coverage/:planId, /reviews and /tests): a Dashboard of tiles
and rings, the review queue as a worklist, and the matrix with a one-run and a
whole-shortfall press per cell. The Pause/Resume button and auto-top-up
checkbox become one Auto top-up switch (ladders get an Enabled switch), and the
plan and ladder editors pick a case through the shared Test grid. See
coverage plans and
ladders.
Harness comparisons
Section titled “Harness comparisons”A comparison holds a case coordinate constant and runs two or more arms, each a
configuration of its own (a third-party harness plus its model, or a gg
configuration plus a model per slot), so “gg against Pi” is statable. Core adds
deterministic descriptive statistics (median, mean, quartiles, a bootstrap
confidence interval on the median seeded from the arm’s sorted run ids, a
Wilson pass-rate interval, a two-arm median ratio), an automated-only checklist
score, per-arm aggregation and confound detection, plus two record captures
every harness parser contributes: per-turn usage events (Pi, Kilo, OpenCode)
and tool-call counts including consumed todo tools (RunRecord.tool_calls).
The backend stores the configuration in a comparison table with per-arm
statistics computed on read (GET, POST, PUT and DELETE /comparisons and
/comparisons/{id}), validates the sample size, arm set and engine, and
POST /comparisons/{id}/publish marks it published, enqueues a publish job per
publishable arm run and emits comparisons.json plus comparisons/<id>.json
into the snapshot. The console has create, edit and detail pages
(/comparisons/new, /comparisons/:id, /comparisons/:id/edit) with box-plot
distributions labelled with n and a Trigger missing runs action that tops each
arm up, and a Comparisons tab in the Runs section rendered read-only on the
static site from the snapshot. The per-case Metrics and Leaderboard tabs split
series by (harness, model) instead of merging harnesses, gain an Average points
chart and a shared sort control, and exclude gg runs, which span several
models. See comparisons.
TCQ: the gg query language, Discover, saved queries and dashboards
Section titled “TCQ: the gg query language, Discover, saved queries and dashboards”The gg result-aggregation widget builder and POST /gg/aggregate are replaced
by TCQ, a query language over a document model. Every gg run is one flat map of
dotted typed fields (summary.*, cap.<id> totals over the catalog, sparse
tool.<name>, metric.*, code.*, agent.<id>.*, model) and a query is a
filter tree optionally piped to | stats <aggs> by <keys>. The backend serves
it from an in-memory index of one document per gg run, reconciled per id against
run.updated_at (a mutation timestamp every run mutator stamps) and invalidated
when an ingest re-resolves checklist weights: POST /gg/query, POST /gg/query/batch (one read per dashboard) and GET /gg/fields, each result
capped at 1000 rows and flagged truncated.
The console’s Discover page (/gg/query) is a highlighted editor with a field
sidebar carrying document counts, worked examples, a document view that opens
each run, and charts chosen from the query’s shape. Saved queries and dashboards
(gg_saved_query and gg_dashboard; /gg/saved, /gg/dashboards/:id) store
query source text per account; the /gg/aggregate widget builder and its URLs
are gone, and Discover is the only address the analysis section answers to. The
evaluator is mirrored in TypeScript and held to the crate’s own conformance
fixture, which is what lets the public site run the same Discover page in the
browser over gg-runs.json, a redacted export of every gg run’s document that
is not gated on publication and fails closed on a case it cannot classify. See
the query language.
Static code analysis of the tree a run produced
Section titled “Static code analysis of the tree a run produced”A new crates/code-analysis crate reads the produced tree without executing
anything: a .gitignore-honouring walk with a floor for the host-owned .tcab/
namespace, an authored set derived from the run’s seed commit (now recorded as
RunRecord.seed_commit), complexity scoring through oxc and syn behind a
derived-stack parser guard, a module graph with cycle detection, a clone
detector and a two-tier rollup. It runs at a new post-run stage seam (after
collection, before validation, outside the runtime cap, and on a canceled run
too), writes the bounded summary onto RunRecord.code_analysis and the
unbounded document to {run}/code-analysis.json.gz, mirrored into the backend
store and served by GET and POST /runs/{id}/code-analysis.
run.code_analyzer_version records the analyzer generation (currently 1) so a
mixed corpus stays sliceable, and NULL means never analyzed.
The run’s Code tab (/runs/:runId/code) shows the provenance strip, a figure
table driven by the CODE_METRICS catalog, a treemap explorer with per-file
coverage, outlier rankings and cycles, plus the executed test and coverage
bands; the analyzer’s own diagnostics are headed Analysis notes and its static
test counts Test authorship so neither reads as executed coverage.
tcab analyze <dir> [--seed-commit] [--tree-basis] [--top] [--json] runs the
same pass over any directory. The snapshot lifts code lines, size Gini and mean
cognitive complexity onto every run’s summary card and publishes the document
under media/runs/<id>/code-analysis/v<gen>.json; the run log’s CODE column
reads “not measured” rather than zero and is off by default. See
code analysis.
Engines: a run selects the runtime its game is built on
Section titled “Engines: a run selects the runtime its game is built on”A run now carries an engine beside its test case, variant, harness and model.
The catalogue is closed and embedded in core from engines/<slug>/engine.toml:
none (the default; the build supplies its own frame loop, input, audio, assets
and diagnostics), simple-2d, structured-2d, simple-3d and structured-3d,
each a @clockwyrks/<slug> npm package under packages/<slug>/ staged into the
host package store and vendored into .vendor/engine/ at seed time, with its
dependency written into the seeded workspace’s package.json and its
documentation seeded at engine/ in the run root.
A case declares support per version with the engines = [...] root key and one
[[engine]] table per runtime engine carrying a required inclusive
min_version and an optional exclusive max_version. The engine’s version is
read from the staged package at seed time and recorded on the run beside the
slug, and a run whose engine is unsupported or whose version falls outside the
range is refused before any container work. tcab run, seed and validate
take --engine (default none), tcab engines lists the catalogue, every
enqueue endpoint carries engine on its launch body, and the console’s New Run
form offers the engines the resolved version supports. A game jam may declare
engines as an optional key defaulting to ["none"].
The engine travels the whole system. The run record lifts engine_slug onto
the run row, the run header names the engine beside the variant, the run log
gains an ENGINE column, and GET /runs filters on engine, versions and
testCases. A case’s detail page is anchored to one coordinate (version,
variant, engine) carried in the URL, with the run aggregations scoped relative
to it and widening across engines listing each engine separately, and a
version’s prompt and specs render for a named engine
(GET /test-cases/{slug}/versions/{version}/specs/{variant}?engine=), so a
run’s Inputs tab shows what that run was given. See
engines in core and
the engines overview.
Simple 2D and Structured 2D
Section titled “Simple 2D and Structured 2D”@clockwyrks/simple-2d (1.0.0) owns the frame loop and the delta time it hands
the game, the letterboxed DPR-aware canvas fit, named input actions over
keyboard bindings and a closed touch-layout catalogue, a synthesized audio cue
bus with looping cues and the browser’s first-interaction unlock, asset
resolution under a fixed root, a diagnostics overlay and an opt-in draw-command
recorder; the game writes its own simulation and drawing. The engine holds the
game’s state by value: update receives DeepReadonly<S> and returns the next
state, initialize returns [state, debug] so the debug surface is handed back
off engine.debug, engine.apply(transition) poses a game between frames, and
Engine.diagnostics() reads back the registered sources. Its clock is an
object, so the engine runs without a document.
@clockwyrks/structured-2d (1.0.0) adds a gameplay framework the game is
written inside (worlds built from levels, game modes, actors and components,
pawns and controllers), engine-owned rendering through render components that
draw in world space or in screen space, a one-option image-smoothing switch,
collision detection and the same recorder, with the debug surface returned from
the game instance’s initialize and poses acting on the live engine.world.
Both 2D engines clip the game’s drawing to the logical field so nothing paints
into the letterbox bars, name the pressed button on a bare pointerdown’s
samples, and read an event with no pointerId as pointer 0. Each engine
documents itself under /engines/<slug>/ in four sections (APIs, concepts,
usage and validators), and that documentation is what a run is seeded with.
See Simple 2D and
Structured 2D.
Simple 3D and Structured 3D, rendered through three.js
Section titled “Simple 3D and Structured 3D, rendered through three.js”@clockwyrks/simple-3d and @clockwyrks/structured-3d (both 1.0.0) are the 3D
members of the two families. Simple 3D owns the renderer over the canvas, the
scene object and camera it renders through and a 2D screen layer composited over
the picture while the game populates the scene; Structured 3D adds the same
worlds, levels, game modes, actors and controllers framework with engine-owned
rendering and collision, plus model loading and animation. three is a peer
dependency the build declares itself so the engine, the build and
@clockwyrks/voxel-runtime/three share one instance, at Three 0.185 across the
repository.
A 3D recording is a video of the frames the engine drew, encoded with WebCodecs as VP9 in WebM and timestamped in simulated time with a keyframe at least every 60 frames, so a player steps it frame by frame; the console’s replay player draws 3D draw-command documents through a three.js drawer and refuses an unknown space or format by name. A Structured 3D component’s opacity multiplies the alpha the game last wrote onto a game-owned material, and each placement of a loaded model gets materials of its own so two components built from one model fade independently. Gantry v1.0.0 is the case that runs on them. See Simple 3D, Structured 3D and 3D recording.
Validators decide the functional rating; reviewers rate aesthetics
Section titled “Validators decide the functional rating; reviewers rate aesthetics”A run now carries two rating channels. The functional rating (flawless, great,
passable, scuffed, broken) says whether the build works, is rated per domain,
and is the worst across the run’s effective domains; the aesthetic rating
(legendary, amazing, good, okay, slop) is one tier over the whole run for how it
looks, sounds and feels. A case version on the per-engine manifest spelling,
other than a game jam, is validator-rated: every graded point carries a
validation script together with domains and a failure_cap (broken,
scuffed, passable or great, never flawless), and resolution names the point that
lacks one. A failing point lowers each of its domains to at most its cap, a
domain’s rating is the lowest cap among its failing points, an undecided point
lowers nothing, and the toolchain typecheck gate applies on top.
The functional rating and score are derived from the run record at push time
and shown the moment the run completes, so a validator-rated run is publishable
with zero reviews, which tcab publish says when no writeup is present. A
review of such a run carries a single aesthetic tier, a writeup, and optional
overrides of individual validator verdicts; its effective checklist is the
validators’ verdicts overlaid with its overrides, the run’s score is the
average effective score and its functional rating the worst effective rating
across reviews, and the validators’ own figures stand while there are no
reviews. A validator-rated run that did not complete is left unrated. Legacy
versions on the single workspace spelling keep the reviewer’s rating as the
functional rating, have no aesthetic rating, and reject failure_cap and
domains by name.
Console and site render a functional badge and an aesthetic badge side by
side, every capped point carries its cap as a rating badge, the Verdict tab
(/runs/:runId/verdict) becomes one per-item browser with the
reference-versus-run media and per-pane downloads (a replay is rendered to
WebM), a per-domain rating strip appears on the public gallery too, and the
About section gains a Ratings tab. @clockwyrks/run-stats gains
AESTHETIC_RATINGS, FAILURE_CAPS, FAILURE_CAP_RATING,
reviewItemsForEngine and the effective-checklist scoring. See
end-to-end evaluation and
results.
Validators are per-engine vitest projects that hold the build in process
Section titled “Validators are per-engine vitest projects that hold the build in process”A case that supports an engine ships one validator project per engine under
validation/<engine>/, a TypeScript vitest project separate from the build’s
own config. A 2D engine’s suites build the engine over a @napi-rs/canvas
surface in the test process, import the engine and the build’s own game module,
install a scripted clock and step an exact number of frames, observing the
engine’s values, its events and what it drew; an engineless project reaches the
build through a browser via the shared harness. The runner stages
validation/<engine>/ into the collected tree at validation/, plus the
shared harness at validation/case-harness/, runs vitest naming the project’s
config and, as file filters, exactly the suites the run’s variant’s checklist
declares, so a suite no item of that variant names is never loaded and a
variant left with nothing to run is refused. The staged project is removed
afterwards and a directory that already stood at that name is held aside and
put back, so tcab validate is safe on a committed reference.
A validation’s optional engines key names the engines a validator decides its
point on, defaulting to every engine the case supports. A point whose validator
does not cover the run’s engine is excluded from that run’s checklist entirely,
is shown to no reviewer and carries no weight, and
review_items_for_engine in core and reviewItemsForEngine in TypeScript give
backend and console one answer; this is how the debug overlay’s toggle is
graded under none alone. Every graded point of a v0.7.0 case is decided by a
validator, and every .test.ts in a project is one a review item names.
Failure assertions are stored as real expected and actual pairs with stack
frames stripped. See validation and
the suite.
A validator run stopped at its budget is inconclusive, not a failed build
Section titled “A validator run stopped at its budget is inconclusive, not a failed build”Three outcomes are held apart from a failed verdict and leave a point undecided rather than synthesizing a failure: an unmet precondition (a validator skipping every check, including a browser or page the suite could not obtain), a check that never ran (no validator project for the engine, no vitest in the tree, a suite the project lacks, an unreadable report), and a run that exceeded its wall-clock budget. Each is recorded with its reason on the run record, an undecided point lowers no rating and carries no weight, and the console’s Verdict tab says which of the three reasons an undecided point carries.
The whole suite run is capped by TCAB_VITEST_TIMEOUT, 45 minutes by default
and overridden in whole seconds by TCAB_VITEST_TIMEOUT_SECS; crossing it stops
the suite’s whole process tree, is recorded as a fact about the host, and
leaves the build’s score intact. A validator that could run but did not
complete against a conformant build (a missing module, a call that threw, a
malformed return, an undeclared output) still fails its point. See
validation.
One shared validation harness: @clockwyrks/case-harness
Section titled “One shared validation harness: @clockwyrks/case-harness”The browser-driven validator harness and canvas draw recorder that had been
copied into every engineless case, and the per-engine harness machinery each
engine project carried, are replaced by one workspace package,
packages/case-harness, that every validator project on every engine is built
on. It ships its TypeScript source with no build step, core stages it beside the
case’s project out of the host package store with a checkout fallback, and it is
deliberately absent from SHIPPABLE_PACKAGES so no case can vendor the tests
into a run. The engineless half serves the built site, holds one Chromium and
drives the build through the case’s debug handle; the engine half constructs an
engine the case injects as values (createEngine, clock, driver, projection),
so the package names no engine and a fifth needs no change.
It carries the readers cases share (coalesced text runs through drewText and
drewTextAnywhere, pixels, colors, draw calls), the replay and still writers,
an asset host (installAssetHost: fetch over the workspace,
createImageBitmap, a structuredClone that carries decoded images by reference,
canvas creation over real rasterizers with page-like font resolution) and
installAudioContext with throwing and tolerant WAV decoders, a CDP touch
driver a case opts into with hasTouch, a beforeLoad hook, and an opt-in
audio gesture delivered before the opening reset. Its recorder binds to the
surface a build draws into rather than the attached canvas it blits to, and
images are pooled per run into the shared store. All 54 validator projects
type-check under npm run typecheck:validators, which CI runs. See writing
debug APIs and
validators.
Engine recordings are the evidence a validator produces
Section titled “Engine recordings are the evidence a validator produces”A validator captures each media output its review item declares in the form the
manifest gives it: image (a PNG of the surface as it stands) or replay. For
a replay it arms the engine’s recorder once its scenario is posed and disarms it
once the behavior has happened, so the evidence is exactly what the build
submitted over that stretch, with the diagnostics overlay outside the bracket. A
2D recording is a draw-command document in recording format 1, stored gzipped as
<verdict>__<output>.json.gz and served as application/json with
Content-Encoding: gzip: every frame carries the whole drawing state it
inherited, values the context produced and the bitmaps drawn are held in tables
the recording shares, and numbers are written to nine significant digits, so
seeking to any frame costs the same and two recordings of one scenario compare
frame for frame. A 3D recording is <verdict>__<output>.webm.
Bitmaps are written once per run into a shared image store (img.<id>.png,
the id derived from the bytes) beside the recordings and travel every route
declared media travels; RGBA pixel buffers stay inline. A recording is kept
whatever the verdict, an output that was never written is recorded absent
rather than failing anything, and a written recording holds at most 300 frames.
The baseline half is the same suites run against the reference implementation,
so the reviewer’s side-by-side compares the build’s and the reference’s frames
from one scenario driven the same way, and the console’s replay player draws
2D and 3D recordings frame by frame from the shared resource tables. See
recording.
The toolchain gate and the verified dependency install
Section titled “The toolchain gate and the verified dependency install”A case declares a [toolchain] table (typecheck required; lint, format
and test optional), and a post-run stage runs the commands over the collected
tree after the container is gone, together with a smoke check that builds the
site and opens it headlessly. Each command is recorded on the run record’s
toolchain block with its exit code and a bounded output excerpt. A
typecheck that ran and exited non-zero rates the run broken and scores it
zero, applied where the rating is derived so validator and reviewer verdicts
stay as written; the other three gate nothing, and lint runs with
--max-warnings 0 so a warning fails. Versions that declare no [toolchain],
every frozen one among them, are neither checked nor gated.
The test command’s results and coverage are read from two report files the
case’s build vitest config writes (coverage/test-report.json from the json
reporter and coverage/coverage-summary.json from istanbul with
reportOnFailure: true) rather than from stdout, including the individual tests
with name, file, status and duration, bounded and flagged when cut; the run’s
Code tab renders both bands and joins per-file coverage onto the explorer.
Every dependency install of a collected tree is verified against the lockfile
the way npm’s tree builder decides the expected set, retried up to three
attempts with a delay, and recorded with its output and attempt count;
validation reuses the recorded install, and a run whose install still fails
ends as an infrastructure failure rather than catastrophic. Seeding admits
.prettierrc.json and .prettierignore through the dotfile allowlist so the
format command checks the case’s configuration. See
end-to-end manifests.
The instrumentation contract: reconcile(), unconditional operations, posed draws
Section titled “The instrumentation contract: reconcile(), unconditional operations, posed draws”The debug-API contract every playable case mandates changed in three ways,
applied to the twenty editable case versions and left untouched on the 27
frozen playable ones. reconcile() joins the core operations: it re-derives every
reading the snapshot reports from the state it depends on without advancing the
clock, firing anything or correcting anything, so a driver poses a world,
reconciles it and reads the world it posed. An operation is unconditional: it
never declines on the strength of the route a player would have taken (which
screen is up, which panel is open, where an actor stands), a pose applies the
value it is given without clamping, and a call whose argument names nothing
real still throws.
Randomness is posed rather than seeded. Specs state each draw as behavior (the
set, the probability, when it is drawn), name no generator, carry no seed on
reset and no generator state in the snapshot, and the API carries an operation
that sets the outcome each validated draw decides (setNextSaucerEdge,
setNextPod, setNextCrit, setLanePhase and their siblings) plus a gate that
stops automatic draws while a scenario is posed (setWaveSpawning,
setBearEmergence, setFishCadence, setTimerRunning), or performs one draw
alone as a reading (drawVent, rollPress, generateBoard). Each operation of
an engine-format version sets one field, places or removes one entity, reads the
state or moves the clock; the compound and patch operations (startMatch,
serve, setBoard, startGame, placeTower) are gone, snapshot() reports
every field an operation can set, every entity carries a stable id, rosters
can be emptied one at a time, and per-entity sense, travel and fire splits let a
check on what a creature senses hold its body still.
The keyboard operations (keyDown, keyUp, press) are gone because the
keyboard belongs to the runtime, and most cases drop setMuted in favour of
the real mute binding. The browser driver serves the engineless contract only:
setAutoStep is a required entry, and a read-only preflight refuses a build
whose handle lacks step or setAutoStep. See
instrumentation.
Menus take a mouse, a pen and a finger
Section titled “Menus take a mouse, a pen and a finger”Every engine-format version’s menus accept a pointer and touch alongside the
keyboard: moving a pointer onto an item selects it, a press and release inside
one item confirms it, a press that begins on one item and ends on another
confirms nothing, and a touch selects where it lands and confirms where it
lifts. The build owns its layout and reports each item’s hit region through a
new menuItemRect(index) reading, which the validators drive real Chromium
mouse and touch input at, and setMenuIndex and menuIndex join the surface
and state. The same pass pinned the transitions the older specs left open:
Esc and P both open and both resume the pause menu, every menu opens on its
first entry and wraps at both ends, and returning to the title selects the entry
that led away from it.
Pointer-first cases go further. Refract’s and Facet’s screens carry on-screen
BACK, CLEAR or PAUSE controls so a touch-only player can reach every
screen, and Facet’s board is played entirely by hand with the keyboard reduced
to menu actions. Carom v3.0.0 also makes the paddles movable during the
pre-serve countdown and draws the serve’s vertical sign at random.
Showcases: a run’s landing page and a case’s carousel
Section titled “Showcases: a run’s landing page and a case’s carousel”A showcase is a directory holding showcase.md (store-page markdown, up to 64
KiB), showcase.toml (up to 10 [[media]] tables, each a file and a caption
name) and the media flat beside them (.png image, .json.gz replay or
.webm video, each up to 25 MiB). The run showcase is written by the model at
showcase/ in its repository, instructed through the case’s specs rather than
a manifest key; record assembly captures it onto the run record as showcase
with lenient degradation (truncation, dropped entries, a warning, never a
failed run), the driver uploads the directory, the backend and artifact service
serve it at GET /runs/{id}/showcase/{file}, the snapshot publishes it under
media/runs/<id>/showcase/, and the run’s Play tab becomes its landing page
rendering the carousel over the description. A case that requires a showcase
carries one review item (showcase.exists, weight 3, capped at great) that
checks its existence alone; the media is the reviewer’s to judge.
The case showcase is declared per variant with showcase = "showcase/<variant>"
and captured from the reference implementation: the leading entry is twenty to
forty seconds of real play driven through the real input path (a Carom match to
two points, a whole Cascade deal, a Gantry lift, a Meltdown floor played out),
with capture drivers committed under showcase/capture/ that audition several
takes and keep the best. It hard-fails version resolution on any malformed
entry, so tcab capture-baselines <slug> --dry-run doubles as the validity
check, is served at
GET /test-cases/{slug}/versions/{version}/showcase/{variant}/{file}, drives
the catalog preview stage, and is published content-addressed under
media/cases/... with .webm transcoded to .mp4. Every new version except
Volute declares one. See showcases and
authoring a case showcase.
A redesigned catalog and a home page of legendary runs
Section titled “A redesigned catalog and a home page of legendary runs”The test-case catalog is a resizable master-detail split with a sticky preview
stage and filmstrip, cards show each case’s latest version, and the detail
page’s Inputs tab is a grouped file tree with an inline highlighted source
viewer shared by runs and game jams. The header shows the anchored version’s
own description and a link naming the latest version when an older one is
viewed, and the snapshot publishes each variant’s starter-workspace files under
files/cases/....
Test-case groups (test-case-groups/<slug>/test-case-group.toml, served by
GET /test-case-groups and carried by the snapshot) define ordered sets of
related cases, and the public home page renders one cross-case leaderboard per
group, ranked by mean score fraction so cases with different point totals stay
comparable. Three ship: tower-defense (meltdown, valence, arc-foundry),
arcade-physics (pong, fathom, spectra) and sim-economy (deepcore,
coil); every member slug must resolve in the catalog, which the manifest test
and backend ingest both check. The home page also shows the five most recent
Legendary runs as looping showcase replays, a totals band and a 52-week
activity chart from the open GET /stats/cabinet read (runs, tokens,
comparable cost, distinct cases and models, weekly buckets), and the Other
section is reachable on the public gallery. See
test-case groups.
Eleven existing cases ship a new major version on the engine format
Section titled “Eleven existing cases ship a new major version on the engine format”Every non-experimental end-to-end and full-stack case gains a version authored
to the v0.7.0 shape: Carom v3.0.0 (slug pong), Cascade v3.0.0, Fathom v3.0.0,
Floe v3.0.0, Shatter v3.0.0, Spectra v2.0.0, Wireworm v2.0.0, Meltdown v2.0.0,
Coil v2.0.0, Arc Foundry v2.0.0 and Deepcore v2.0.0. Each seeds a complete
TypeScript project (Vite, tsc, ESLint, Prettier, Vitest and an index.html)
instead of a bare page, names one starter directory per engine in a
[workspaces] table, and runs under three engines: the engineless none run,
Simple 2D and Structured 2D at >= 1.0.0. Under none the seeded project
holds no src/ at all and the build writes the frame loop, canvas fit, input,
audio, overlay and its window.__<case> surface; under an engine the case
seeds src/constants.ts and src/main.ts and the build writes src/game.ts,
whose initialize returns the debug surface beside the state.
The tick_hz key is gone from every manifest: rates are per second and
integrated against the frame’s delta time (Floe keeps a 120 Hz step as a rule
stated in its own spec), and each validator constructs its own clock. Every
checklist was re-cut to one observable behavior per item, which is why they grew
(Cascade 47 to 292 common items, Floe 83 to 256, Shatter 99 to 241 common plus
54 variant, Meltdown 107 to 378, Wireworm 77 to 225, Deepcore 109 to 400), and
domains were split so one failure no longer sinks a whole score: Cascade rates
rules, handling, cascade and presentation, Fathom navigation,
sensing, predators, progression and presentation, Floe hunter,
crossing, run and presentation, and Arc Foundry adds an audio domain for
its twelve produced cues. The standalone .mjs browser drivers of the previous
versions are gone, each variant ships a reference implementation per engine
under references/<engine>/<variant>/, and the previous versions stay frozen
and untouched apart from their .frozen digests. See the end-to-end
overview and writing workspaces and
references.
Seven new test cases
Section titled “Seven new test cases”Refract v1.0.0 (end-to-end, easy, 3 h) is a light-tracing puzzle on an optical
bench with a 24-board authored campaign and an endless Cascade mode whose
generated boards are held to measured per-tier difficulty floors; every screen
is pointer-first with named, fingertip-sized targets the debug surface reports.
Facet v1.0.0 (full-stack, easy, 4 h) is an 8x8 gem matcher where every clear
strains the neighboring stones, cracked stones clear with anything beside them
for double, three cuts carry looping auras, a move is offered on the drag and
only played on the release, and each level ends on a levelclear tally screen.
Kessler v1.0.0 (full-stack, easy, 6 h) is an orbital breakout: a deflector on a
circular track bats a ball outward through three rings of derelict satellites
while caught salvage pods grant tools.
Volute v1.0.0 (full-stack, easy, 6 h) is a channel shooter in a geothermal pump
hall, where cores ride a winding channel toward an intake and a central injector
groups like charges to pull them out. Wick v1.0.0 (full-stack, easy, 8 h) is a
survivors-like: the lamplighter only moves, sixteen weapons fire on their own,
thirteen enemy types arrive on a ten-minute schedule ending with the Dark, and
its checklist is 943 common items across 30 categories. Gantry v1.0.0
(full-stack, medium, 8 h) is the first 3D full-stack case: rig a tower crane
from struts, cables and rails, write an instruction tape for its four axes and
run it under a structural simulation across six sites; it declares
asset_dimension = "3d", depends on @clockwyrks/voxel-runtime, and runs on
Simple 3D and Structured 3D rather than the 2D engines. Orrery v1.0.0
(full-stack, medium, 8 h) is a machine-building puzzle of brass arms and
celestial motes on a hex field (sigils, looping tapes, constellation deliveries
graded on cost, time and footprint) and carries the largest checklist at 1054
items.
All seven are validator-rated on every engine they name and none is
experimental. Every full-stack one except Wick declares
packages = ["@clockwyrks/particle-runtime"], and Gantry declares the voxel
runtime instead. See the full-stack overview.
Case-by-case rule changes in the new versions
Section titled “Case-by-case rule changes in the new versions”Beyond the shared rework, each version pins rules its predecessor left open.
Carom v3.0.0 splits hit-edge and no-tunnel into per-face points, adds
navigation and hud categories, states the sub-stepped physics exactly (the
serve angle is exactly SERVE_ANGLE), moves six points multi cannot share
into the variant files, and adds setMuted. Cascade v3.0.0 states that the
waste remembers each turn as a set (wasteSets), fixes every pile’s drop
rectangle and the leading-card-centre drop rule, gives DRAG_THRESHOLD (5) a
stated effect, and makes a cleared table put down the run in hand. Fathom
v3.0.0 adds setMaze fixtures, splits sensing from travelling per creature,
adds a released flag for the den schedule, and pins the sonar wavefront at
fourteen corridor steps per second. Floe v3.0.0 moves to stage coordinates
everywhere and fixes the five bays at columns (3,4), (11,12), (19,20), (27,28)
and (35,36).
Shatter v3.0.0 clears a wave only by shooting it, grades the saucer over twenty
real crossings, reads the fragment fan 412 units from the well, and grades that
RESTART begins a game clear of saucers. Spectra v2.0.0 replaces the flat 4.0 s
Flux cycle with fluxHold(stage) and three sibling stage formulas, sub-steps a
frame by SUBSTEP_MAX, states mute at source strength, and requires
before-and-after captures for difference claims. Wireworm v2.0.0 states the
worm cadence as wormStepInterval(level), fixes CURSOR_SPEED at 430 and
GLITCH_DART_INTERVAL at 0.32 s, rules that a drop passes through whatever is
in the tile below, and pins when the spawner clocks are set. Meltdown v2.0.0
drops the 60 Hz tick, accepts dist/ alone as build output, requires unit
tests at src/**/*.test.ts, and pins seven open rules including closed-form
wave composition, the trip as a crossing of 100, and the Rime as an ordinary
emitter with base damage 4.
Arc Foundry v2.0.0 rounds a unit’s maximum HP to a whole number, fixes Medium’s
surcharge constants at c = 0.28 and r = 1.145, and replaces the right- and
shift-press with a modify action bound to Shift that every engine can deliver.
Deepcore v2.0.0 adds setWorldSize, pointer and touch across the shell, and
scrolls the shaft through the engine’s own camera under Structured 2D. Coil
v2.0.0 makes the best score a session figure and moves the three points the
Maze mode disagrees on into both variant files.
Carom v2.2.0 grades spin on the ball’s flight
Section titled “Carom v2.2.0 grades spin on the ball’s flight”Carom also gains v2.2.0, still on the single-workspace legacy format with
tick_hz = 120 and 71 review items, as a correction pass over v2.1.0 learned
from runs. The spec states that there are two ways out of a pause (the pause key
toggles, and the RESUME entry confirmed with Enter or Space), that every menu
opens on its first entry, and that the pre-serve countdown is live play in which
paddles move and the pause key works; a new Pause point resume grades all four
routes. Every Spin point measures the ball’s perpendicular offset after the
contact rather than only reading the spin scalar, comparing magnitudes so
mirrored builds both pass, and a parked AI at contact makes
spin.moving-solo-ai inconclusive rather than failed. Audio points arm the
AudioContext with the bound W key and drive cues in real time,
advances-in-real-time is judged from injected keys instead of control
operations, and the spec states who owns the clock at boot, after a control
operation, and under injected input.
Lattice graduates from experimental
Section titled “Lattice graduates from experimental”Lattice (performance, hard) loses its experimental = true line, so v0.7.0 is
the first release in which Lattice runs are recordable and the case appears on
the site as a real case. Its scored set was redesigned around dense main-bus
factories: medium is a designer-authored 48x32 factory of 630 entities run
over 300,000 ticks, large copies it onto a 72x40 grid and extends it to 1,027
entities building the whole machine tree over 360,000 ticks, and small keeps
the simple lines layout as the fast correctness confirmation. Raw ore is
emitted only at sources and smelted in a new 2x2 furnace entity that burns a
new coal item, three machine recipes (transport-belt, inserter,
assembler) form a dependency tree, belts move at tiered slow, fast and
express speeds, and every splitter, curve and side-load does real work.
The engine rules were corrected in step: the splitter is a lane-preserving, item-agnostic balancer, an inserter’s empty return takes real time and a crafter-loading inserter picks the item the recipe still needs, belt movement is defined over a run rather than a tile, a pure curve continues its run with both lanes preserved, and a side-load lands each feeder lane at its contact point. The fuel ceiling is 40,000,000,000, under which the transport reference passes and a naive engine exhausts its limit, a returned checksum must match the returned state, and every oracle and checksum was regenerated. The case page gains a Reference tab that plays the scored factories, and the performance documentation was rewritten in step. See Lattice.
Two Lattice sprite-sheet cases, and the existing sheets grow
Section titled “Two Lattice sprite-sheet cases, and the existing sheets grow”medium/lattice-furnace (twelve 64x64 frames of a 2x2 top-down smelter, an
off idle in frames 0 to 3 and a smelting loop in frames 4 to 11) and
medium/lattice-lane-splitter (eight 32x64 frames of a single-input machine
that unzips one belt’s two lanes onto two outputs) are new experimental
asset-generation cases, each with a draw.sh reference implementation, which
exist to seed Lattice’s sprites. The existing Lattice sheets grew:
lattice-belt goes from 16 to 48 frames (straight and curved forms at three
tiers), lattice-assembler from 8 to 24, lattice-inserter from 12 to 36, and
lattice-items from 7 to 17 icons. lattice-splitter gains a weight-4 fidelity
item requiring the housing, output arrow and moving part on a registered
mechanism layer over a continuous belt bed, and the lane splitter carries the
same rule. The asset-generation catalog stands at 132 versions. See
sprite cases.
Appearance is the build’s
Section titled “Appearance is the build’s”No live case version seeds a canonical palette, a typeface requirement, HUD
coordinates or reference screenshots any more, and none declares
[[reference]] views, [[proof]] artifacts or [[check]] comparisons. The
specs instead state a legibility table of what a player must read at a glance,
and the presentation validators assert that a thing was drawn and that two
things are told apart against a stated RGB distance, never a hex value; how
good it looks is the run-wide aesthetic rating. The nine experimental
legacy-format versions (Caldera, Siege, Sunfront, Thunderhead, Holdfast,
Hollowdeep, Junction, Midway and Valence v2.0.0) also drop their [[proof]] and
[[reference]] tables, specs/proof.md, mockup sources and capture
instructions while keeping every review item, and a case declaring no mockup
needs no headless browser at ingest or rendered reference at run start. The
three features stay supported in core because the 27 frozen versions still
declare them. See
writing case specifications.
Audio packs are declared per case and staged into the run
Section titled “Audio packs are declared per case and staged into the run”[audio] packs replaces the sample_pack and instrument_bank manifest keys:
an ordered list of name@version refs, valid on audio asset-generation cases,
full-stack cases and game jams, with ManifestAudio denying unknown keys so a
stale line is a loud parse error. An sfx-sample case declares exactly one
sample pack, a music case exactly one instrument bank, an sfx-synth case
none, and order decides which pack an unqualified call plays. A full-stack
version or jam with no [audio] table receives the pinned four-ref
DEFAULT_AUDIO_PACKS ([email protected], [email protected], [email protected],
[email protected]), a literal rather than a scan of the registry, so publishing
a pack cannot widen a frozen case. The fourteen non-frozen full-stack versions
each declare a set chosen against their own specs/assets.md, all eight game
jams declare the full four-pack set, and scripts/ci/audio-packs-check.mjs
requires the table on every unfrozen full-stack version and jam.
At container start core’s audio_stage resolves the declared refs against the
host audio store (TCAB_AUDIO_STORE, default /opt/tcab-audio), verifies every
clip against the store’s objects.lock.json by sha256 and byte length, and
materializes packs.json, one packs/<name>@<version>/pack.toml per pack and
the union of their clips under clips/ at /opt/audio inside the container,
outside the seeded repository. The tree travels as a host directory on the new
ContainerSpec.dirs, one copy per container rather than per-file bytes, and a
pack ref’s halves are held to plain name characters. A run reads
/opt/audio/.tcab-audio-contract back from its started container and fails at
start, naming the image and pin, when the image predates staged delivery. See
execution and
full-stack manifests.
Audio packs are collections of registered clips in an object store
Section titled “Audio packs are collections of registered clips in an object store”A pack is no longer an opaque tarball. containers/sample-packs/clips.toml is
the clip registry, 70 distinct clips keyed by the sha256 of each clip’s original
source bytes, holding its license, provenance and recorded pitch, and each of
the four pack manifests is a collection of those ids plus the name, tags and
description a model browses, so a clip shared between packs is stored once. Clip
bytes live in a private Cloudflare R2 bucket as the original source and as the
output normalized under each pack’s profile, and objects.lock.json records the
140 published objects. The publish tooling is split over
scripts/lib/audio-store.mjs: curate-instrument-bank.mjs is the only script
that contacts Freesound, build-sample-pack.mjs publishes a pack’s normalized
objects, and stage-audio-store.mjs materializes every published pack into one
verifiable tree; presign-sample-pack.mjs and packs.lock.json are gone.
That tree ships as the data-only test-cabinet-audio-store image, built ahead
of the run images by scripts/ci/audio-store-image.sh and fused per
architecture, and the driver service image copies it in through the
AUDIO_STORE_IMAGE build arg to /opt/tcab-audio. The sfx-sample, music,
full-stack-2d and game-jam run images bake no audio at all.
scripts/fetch-audio-store.sh puts the store on a laptop for a local run,
pulling the image from the registry or staging straight out of R2 with the
read-scoped presign pair, defaulting to ~/.cache/tcab/audio-store, and the R2
endpoint may be given as CLOUDFLARE_AUDIO_R2_S3_URL instead of being derived
from the account id. See publishing an audio sample
pack.
The audio binaries load the staged palette and fail closed
Section titled “The audio binaries load the staged palette and fail closed”sfx-sample and music read their palette from /opt/audio through a shared
staged contract in crates/audio-core, so no run inherits a palette from its
image and a tool config can only reach a pack its case declared; a config that
names no pack takes the staged default of its kind, and --config on both tools
takes a pack ref. A palette that is missing, unparseable, empty, of the wrong
kind or not the pinned name@version fails the run instead of collapsing to an
empty library, audio whose duration or sample rate disagrees with its manifest
is rejected, and at render time a placed sample absent from the library, or a
track instrument that is neither a synth waveform nor a bank entry, is an error
naming what was missing rather than silence.
music gains list-instruments [--tag] and instrument-info --name, mirroring
sfx-sample’s list-samples and sample-info, which every music brief already
told the model to use. TCAB_AUDIO_DIR points a host-side tool run at a fetched
audio store, read as a palette of every pack it holds. See the audio
binaries.
A full-stack case names its asset dimension
Section titled “A full-stack case names its asset dimension”asset_dimension is a full-stack-only root manifest key, "2d" (the default,
so every existing case resolves unchanged) or "3d", and it selects the run
image the way an asset-generation case’s asset_kind does. "2d" runs in
test-cabinet-full-stack-2d with draw, draw-sheet, particle-2d,
sfx-synth, sfx-sample and music; "3d" runs in the new
test-cabinet-full-stack-3d, which adds voxel, voxel-anim and particle-3d
plus the Mesa software Vulkan runtime their preview PNGs render through. The
key travels the backend wire so dispatched runs resolve the same image a local
tcab run does, the standing full-stack quality directive names the binaries of
the image the run is in, and resolution rejects the key on every other test
type. Gantry v1.0.0 is the first 3d case, and Orrery moved from end-to-end to
full-stack so its build produces its own sky and brass. See
full-stack manifests.
Reference implementations and baselines are per variant per engine
Section titled “Reference implementations and baselines are per variant per engine”A variant declares one reference implementation per engine, by convention
references/<engine>/<variant>/, or a bare path for a single-workspace case.
tcab publish-reference --env <prod|staging> <slug> [<version>] [--variant] [--engine] [--all-variants] [--dry-run] [--skip-baselines] builds each
variant and engine pair with the case’s [build] commands, re-captures its
baselines, scrubs and deploys it to Cloudflare Pages under the alias
<slug>-<version>-<variant>-<engine>, and records the URL in
test-cases/reference-builds.lock.json at
<env>.<slug>.<version>.<variant>.<engine>. The backend’s
case_reference_build table is keyed by (slug, version, variant, engine), the
case page’s Play tab launches the build recorded for its anchored engine, and
the Reference tab switches between engines.
capture-baselines records an engine-backed case’s baseline by running that
engine’s validator suites against the reference implementation in process, and a
browser-driven case by serving and driving its build, so both panes of the
reviewer’s comparison come from the same scenario. Engine-backed references
depend on the repository’s own packages/<slug>/ by a relative file: path
instead of a vendored copy, so npm ci && npm run build:packages at the root is
a prerequisite of both commands. A new guide covers buildable, script and
bundled references and the release gate. See publishing a reference
implementation.
The authoring standard: two new guides
Section titled “The authoring standard: two new guides”Writing Debug APIs and Validators fixes the debug-API design rules (atomic
scalar operations, no patch operations, compound sequences in the validator
harness, the declared state as the whole state, a build reporting layout a spec
leaves loose, random draws posed as outcomes) and the validator rules: assert
the specification never the reference, grade the build’s code not the engine’s,
transcribe every asserted figure into the project’s own constants.ts with the
build’s module imported at one site, one requirement per validator, pose an
isolated world, drive simulated time only, sample a probability through a
lone-draw operation inside a six-sigma band, read state before pixels with
presence floors rather than thresholds, always reach a verdict, and finish
inside budget (under 3 s per validator on an engine, 5 s engineless, 10 s hard
cap, 15 minutes for the whole suite on a two-core host).
Writing Workspaces and Reference Implementations states what a seeded workspace
and a reference hold: the four toolchain scripts, Prettier and ESLint configured
in both, specs/ ignored by both, an inert <link rel="icon" href="data:,">,
an engineless workspace of configuration only, and dependencies pinned to the
newest usable version. Writing Case Specifications gained the
never-help-the-model rule, the keep-evaluation-out rule enforced by
scripts/ci/spec-vocabulary-check.sh, menus that take pointer and touch with
every transition stated, one observable behavior per item, and that appearance
is loose and reviewed while behavior is exact and validated. The end-to-end and
full-stack guides and quickstarts were rewritten for the per-engine layout. See
writing debug APIs and
validators, writing
workspaces and references
and writing case
specifications.
A killed run is canceled when it is a gg run, and destroyed otherwise
Section titled “A killed run is canceled when it is a gg run, and destroyed otherwise”RunState::Canceled joins the run-record contract as a terminal state: never
publishable, released nothing, never retried, and absent from every model
statistic. Killing a gg run whose session has been launched is a cooperative
wind-down: the driver raises a latch and keeps awaiting the run under a
20-minute grace, the engine writes the in-container sentinel that every agent
checks at its turn boundary, the session finishes the turn in flight and runs
its whole epilogue, and the run walks its ordinary post-session path (tree
collection, metrics, session summary, artifact uploads), skipping only
validation, so the record carries real tokens, cost and the tree as it stood.
The driver posts a canceled status carrying the record, the only status the
backend accepts on an already-canceled job, and the backend persists it without
a completion notification or retry.
A killed run of any other harness, and a gg run killed before its session was
launched, is destroyed: the sandbox is deleted, which is what stops it and
frees the scheduling slot, nothing is recorded, and the job stays canceled
with no record. state=publishable refuses every never-publishable state, so a
reviewed canceled run is no longer listed for publish, the console keeps a
canceled run visible and deletable, and a run event fires for a cancellation
while no notification does. The CLI container runtime gains the job-id label so
a canceled run is found and removed under both runtimes. See
the driver and
run records.
gg run outcomes are recorded honestly
Section titled “gg run outcomes are recorded honestly”limit_exceeded joins the terminal states for a gg run stopped on one of its
configured execution ceilings; it publishes like a harness_error (a statistic
only) and is never retried. A 401 or 403 from the provider ends the session
under auth_error and exits non-zero so core records a harness error rather
than blaming a model that never ran, and a failed gg run’s status quotes the
error lines gg logged instead of a bare exit code. A run’s compact
GgSessionSummary (terminal status, agents, compactions, provider stats,
per-slot costs) is emitted before session_ended and lifted onto
RunSubject.gg_summary, which is what GET /stats/providers and the query
language read. The run record and per-turn usage events carry the routing key,
the provider pin and the reasoning setting, each optional so stored records keep
reading. See results.
Deleting a run prunes its artifact tree, and a sweep reclaims orphaned trees
Section titled “Deleting a run prunes its artifact tree, and a sweep reclaims orphaned trees”The backend’s artifact URL is split in two: TCAB_ARTIFACTS_PUBLIC_URL stays
the browser-facing value advertised through GET /config, and the new
TCAB_ARTIFACTS_URL is the in-cluster address every backend-originated call
uses, so DELETE /runs/{id} now reaches the artifact service to prune the tree
(best-effort; the delete succeeds regardless). A periodic sweep lists the
service’s stored trees over a new service-token-gated GET /runs on the
artifact service, keeps every tree whose id still has a run row, and deletes the
rest once older than a grace window; a pass that reads no runs or fails its
listing abandons itself so a fresh database beside a populated volume leaves the
volume intact. The artifact service’s whole-run archive and the publisher’s
tree.tar pull now stream instead of materializing the tree in memory, so a
large tree can no longer exhaust the service. See the artifact
service.
In-progress rows report their engine, start time and a ticking duration
Section titled “In-progress rows report their engine, start time and a ticking duration”A pinned in-flight row used to dash its ENGINE, STARTED and DURATION cells.
job.started_at is stamped once on the transition into starting, the moment
the driver takes the start its record is measured from, so queued time is
structurally excluded; RunEvent and GET /jobs/active carry engine,
startedAt and ggPreset, and the run log ticks every moving duration off one
shared clock held only while one is on screen. Every server-paged listing (the
runs index, the worklists, case and model Runs tabs, gg Sessions, the home page
window) re-queries on the runtime’s refresh token when a run finishes, is
published, is killed or is deleted, so a finished run settles into the filtered
page it was on instead of vanishing. Clear pending, Kill active and Stop all
report their counts in a toast, another run finishing no longer resets the page
you are on, and the testCase sort orders by the case’s display name. See
the UI library.
Console foundations
Section titled “Console foundations”Every run-detail tab is its own route, so leaving one re-fetched its assets behind a spinner; one bounded LRU cache seeded synchronously now serves replays, prepared frames, pixel buffers, meshes, particle systems, skinned meshes, workspace files, engine modules, case details, variants and run records. A read in flight shows a loading state, a failed read reports a failure, and only a read that settled with nothing reports absence, decided by the store’s status rather than error text, and a failed catalog refresh keeps the catalog on screen. Numeric fields hold what was typed and refuse out-of-range values with a reason, and the backend refuses the same bounds rather than correcting them.
A change that can cost work is staged and confirmed through the themed dialog,
and every window.confirm site uses it. A page error boundary keeps the chrome
when a page throws, a route smoke test mounts the app at every declared route
against a stocked console, an empty one and the static gallery, and the
front-end tests run in the pipeline’s webtest job. SubmitNotice renders a
form’s outcome beside the button that raised it, nine pages hand their create
action to the header’s titleActions slot, the “(starts off)” annotations are
replaced by opening on the real default with a reset beside any moved field, and
the loading mark states its aspect ratio before its SVG arrives. See the UI
library.
tcab answers in its exit status, and gains analyze, engines and test-case-groups
Section titled “tcab answers in its exit status, and gains analyze, engines and test-case-groups”tcab validate used to print a failing verdict and exit 0. It now exits
non-zero when the tree did not load, a required install or build step failed or
was never reached, a declared check was unreached, a declared proof is missing,
a gating validator decided against the build or did not run at all (with
inconclusive units grouped by kind and reason on the final line, so a broken
host reads apart from a broken build), or an adversarial submission forfeited;
a loss, a draw and a similarity figure decide nothing, and a point an erratum
excludes costs nothing. capture-baselines fails a target any of whose units
did not run clean, after sweeping every target, and publish-reference skips
deploying a target whose baseline it could not produce. tcab validate also
leaves the directory as it found it apart from the install, build and media
under .vendor/.
run, seed and validate take --engine (default none); engines and
test-case-groups list the built-in engines and the groups; analyze <dir>
runs the static analyzer; publish-reference and capture-baselines work per
variant and engine pair. The CLI’s log lines move to stderr so
tcab analyze --json | jq parses, and tcab no longer links the gg harness
library or an HTTP client. See the CLI.
The Ralph Loop orchestrator is removed
Section titled “The Ralph Loop orchestrator is removed”gg conducts multi-step work through its own executor (compaction, memories,
the project board, subagents), which supersedes looping a stateless third-party
harness across sessions through the filesystem. The orchestrators/ralph/ data
directory, its embedded built-in registration, its catalogue page and its entry
in the run-launch picker are gone; one-shot is the only built-in orchestrator,
so the picker’s program-building gate goes with it. A session loop remains
expressible as an external orchestrator through --orchestrator-dir. See
orchestrators.
The dispatcher no longer kills run pods for memory
Section titled “The dispatcher no longer kills run pods for memory”Driver pods were dying with OOMKilled (exit 137) after hours of paid API
calls, because the driver container carried a memory limit equal to its request
and the sandbox pod a 4Gi limit, while the driver runs the case’s whole
toolchain (install, tsc, linters, vitest, bundler, headless Chromium) over the
collected tree in its own cgroup after the sandbox is gone.
Both memory limits are removed and the requests kept as the node reservation. The driver keeps a CPU limit of 2, which bounds Node’s worker fan-out and throttles rather than kills, with its memory request at 2Gi. The always-on services and the publisher Job keep memory request equal to limit, because a kill there is a restart, not lost spend. See the dispatcher.
The driver collects the run tree over its own verified channel
Section titled “The driver collects the run tree over its own verified channel”Kubernetes can drop the tail of a large exec stdout under load while the exit status still reports success, which lost completed runs. Under the Kubernetes runtime the driver now binds a TCP listener on an ephemeral port, advertises its pod IP and a per-run token to a Node uploader it ships in its own binary and execs into the sandbox, and accepts the produced tree only when the stream’s terminator arrives with a byte count and SHA-256 digest matching what it received; a truncated or stalled stream, a mismatch, a failed tar or an uploader that never connects is retried with a fresh upload, and a verified archive that fails to unpack fails the collection outright with the innermost cause named.
Seeding the sandbox is judged by tar’s own verdict rather than the stdin write,
so a broken pipe no longer masks tar’s real error (a run image without
/opt/audio, for instance), a remote that dies without draining stdin no
longer hangs the run at “starting the run container” because the wait is
bounded on progress, and the kube client’s stdin and stdout pipes are sized up
from 1 KiB. See the driver.
A run listing’s total matches its rows, and unreadable records get a tab
Section titled “A run listing’s total matches its rows, and unreadable records get a tab”Db::assemble used to skip a run whose stored record_json no longer
deserialized while the count still counted it, so the console offered empty
pages and the run answered 404 while remaining deletable only through the
database. Every run row now carries record_readable and the
record_format generation it was decided under; every listing and its COUNT
run one predicate, and a build whose RUN_RECORD_FORMAT (now 3) differs from a
row’s stamp re-decides that row once at startup. GET /runs/unreadable lists
the runs this build cannot read with the error their record produces, the
console’s Unreadable worklist appears while any exist and is where they are
deleted, and DELETE /runs/{id} acts on the row, so it deletes an unreadable
published run too. The contract-shape pin now digests every schema the record
references transitively, the gg capability set included, so a required field
added to a gg agent profile moves the generation instead of silently marking
rows unreadable one listing at a time.
A stored definition the running build cannot read makes the backend unready
Section titled “A stored definition the running build cannot read makes the backend unready”After a stored-shape change the catalog listing used to skip each case whose
latest manifest failed to parse and answer with the rest, so the console showed
one test case while the backend reported itself ready. The definition store now
records the format its contents were written in (STORE_FORMAT, currently 2);
a store stamped with any other format holds the backend unready with that
reason and makes GET /test-cases answer 503 naming the repair, and any ingest
scan meeting such a store is promoted to a forced whole-catalog one, which is
what the cluster overlays’ startup sidecar posts. scripts/reingest.sh asks
/healthz first and ignores its baseline when the store is unservable, and
local-rebuild ingests between the restart and the readiness wait. See
the backend.
gg’s own defects end the run as internal_error
Section titled “gg’s own defects end the run as internal_error”A guest artifact that would not instantiate, a wasmtime host fault, a panicking
sandbox task, a compiler that crashed or was missing from the image, and gg
failing to lower a source it had already accepted all used to end the session as
model_error or feed back to the model as a compile error counted against its
ceilings. Each now reports internal_error, and PrepareError::Lowering and
the TranspileLowering error type are deleted from the published vocabulary.
A run-wide fault latch is read by every agent at its turn boundary, so a defect
in any child, issue agent, reviewer or successor ends the whole run rather than
letting the root finish and be scored, and a suspended wait_for_issue selects
on the latch so a fault can no longer leave waiters parked until the idle
watchdog files the run as hung with no record. An error turn taken under a
raised latch is recorded as gg’s (fatal) rather than in the model’s
program_fault bucket, and a wait on an issue blocked by a failed issue ends as
a refusal naming the blocker.
A compile failure is the model’s only when its program caused it
Section titled “A compile failure is the model’s only when its program caused it”On the C#, Java, Kotlin, Swift and C++ arms a loaded module’s rebuild diagnostic
was returned in the model’s band, asking it to rewrite a program gg had
accepted; it is re-attributed to a lowering fault of gg’s on every arm that
rebuilds a module beside a program. Roslyn’s arrangement diagnostics are
classified as toolchain failures, the response file’s arguments are quoted so a
workspace or home path with a space compiles, and the C# guest carries the
thrown exception’s code and operation up so an uncaught gg API failure is
program_api_error and a withheld capability is program_unknown_name like
every other reporting arm. A Compiler error body is the arm’s rendered
diagnostic and nothing else: the supporting material, import-candidate lists
and matching rule were removed on every arm, and the catalogue’s libraries
section is the record of what an arm may import.
PureScript arm fixes, and a struck stack closes with a count
Section titled “PureScript arm fixes, and a struck stack closes with a count”The PureScript arm’s hand-written module-header scanner was deleted in favour of
the name purs emits. Its code mask panicked on a multi-byte character inside a
block comment, which reached module preparation and ended the run; a PureScript
runtime failure returned frames from inside gg’s SDK instead of the model’s own
program line as every other arm does; and signature analysis for module
documentation pages recognized only the ASCII -> and => spellings, not →
and ⇒. All four are fixed. On every arm, a stack trace gg has struck frames
from now closes with the plain line … and N more frames (external code) at
both strike sites, which is visible in every Runtime error body a model
receives.
A replay’s images travel beside it, and a recording binds to the drawn surface
Section titled “A replay’s images travel beside it, and a recording binds to the drawn surface”The injected canvas recorder bound to the largest attached canvas, so a build
that renders into an unattached fixed-resolution canvas and blits it produced
recordings that were almost entirely base64 PNG and failed validator points over
drawing the build did. It now follows a pass-through blit to the source canvas,
images are pooled across a run in a shared store beside the recordings and
deduplicated by content, and a per-run budget bounds what recordings hold.
@napi-rs/canvas moves to the 1.x line because 0.1 leaked one copy of the
source buffer on every drawImage(canvas). The player carries the canvas
specification’s initial values and stays quiet about a refused assignment of a
default while still reporting a non-default value it cannot apply, and a replay
pane’s frame count sits on its label row.
Validators read what the spec fixes, off the wall clock, inside a budget
Section titled “Validators read what the spec fixes, off the wall clock, inside a budget”A repository-wide audit of the new validator suites landed as one
fix(test-cases) commit per case plus a cross-case sweep for the fault classes
it found. Validators now transcribe every asserted figure from specs/ into
their own project constants instead of importing the build’s src/constants.ts,
with each case’s own lint rule refusing a validator that reads the build’s
src. They read copy off coalesced text runs and grouped figures as one figure,
and they drive simulated frames rather than waiting on real time, so a loaded
host cannot fail a conformant build. Heavy drives were re-costed to what they
read (Cascade’s trail-* points step at the runout rate instead of 240 Hz).
A still whose pose throws fails its point instead of shipping an un-posed
picture, and an entity a check dereferences is hard-asserted first. Appearance
thresholds became presence floors, engine-owned behavior is scoped to none,
and withheld-asset shims were removed with the specs requiring every produced
file to load. Each case passes its whole suite against every reference on every
engine, and baselines were regenerated after the sweep.
Catalog and manifest fixes
Section titled “Catalog and manifest fixes”Catalog discovery requires a test-case.toml or game-jam.toml before treating
a subdirectory as a version, so a leftover reference-impl/ or dist/ on disk
no longer makes the whole catalog fail to list. Whether a case may name an
engine is decided by workspace against [workspaces], and the categories
checklist grammar is opted into with [review] format = 2. The performance
engine model is served as a run asset, the backend ingest test asserts every
engine floor a manifest pins, case READMEs no longer list validation-baseline/
as part of the version folder, and seeded text that named the benchmark or the
review UI was reworded (Foray’s baselines spec, fourteen particle briefs).
The local k3d stack keeps its images and reports what stalled
Section titled “The local k3d stack keeps its images and reports what stalled”make -C deployments/local local-up on a cold state volume waited 600 s for a
backend readiness that only its own local-ingest could produce; apply is
split into apply-overlay and apply-wait with local-ingest between them,
and local-ingest waits for a running, never Ready, backend pod before
port-forwarding. The k3d node’s kubelet image garbage collection is disabled on
cluster create and retrofitted onto an existing node through the new
cluster-kubelet target, because under disk pressure it deleted the driver
image and every locally imported run image and left runs in ImagePullBackOff;
local-import re-imports every built image without a rebuild. One
rollout-wait target over WAIT_WORKLOADS prints the pod table plus pods
terminating past their grace and FailedKillPod events, so a node still reaping
old pods is distinguished from a broken image, and
scripts/free-local-forward.sh kills an orphaned local-forward. See running
the services locally.
Smaller fixes
Section titled “Smaller fixes”- Failure details across core report the error rather than the handling, for
example
run failed: gg execution ceiling hit (5 consecutive turns failed). - The Simple 2D and Structured 2D engines clip the game’s drawing to the
logical field, name the pressed button on a bare pointerdown’s samples, and
read an event with no
pointerIdas pointer 0; the recorder records a tinted sprite as drawn. - The particle binaries’ over-budget rejection names the flags to lower and
keeps the
rate x lifetimesizing rule, and error messages across the audio binaries state the fact rather than narrate it. - A
.json.gzrecording is served as JSON framed in gzip, so the browser inflates it and nothing sniffs bytes on the wire. - The observability page’s cAdvisor guidance says run pods carry no memory
limit, so
container_spec_memory_limit_bytesis meaningless for them and the sizing target is the sandbox memory request. - The gg analysis page says the runtime cap bounds the harness session and each in-container setup step separately, each with the full budget.
binary-smoke.shresolves the built binary through${CARGO_TARGET_DIR:-target}, so the release gate also runs inside the devcontainer.
Development
Section titled “Development”Azure Pipelines is the one pipeline
Section titled “Azure Pipelines is the one pipeline”azure-pipelines.yml runs three stages. gates runs on every branch: rust,
binary (Linux and Windows), web, webtest, specs, format, validators, frozen,
audiopacks, specvocabulary, buildcontext, contract, manifests and submodulepins;
on master, staging and v* tags it also builds the static gg binaries
natively on amd64 and arm64, and a mirror job force-pushes the gated master,
staging or nightly commit, or tag, to GitHub. images runs on master,
staging and v* tags: on the branches it builds the audio store, every
run-container image and every service image natively per architecture, pushes
<image>:<sha>-<arch> to testcabinet.azurecr.io and fuses them into the
multi-arch <image>:<sha> with manifest.sh, and on master and tags its
gg_publish job uploads gg’s release objects. deploy runs
scripts/ci/deploy.sh on the tcab-staging environment for the staging branch
and tcab-prod for master, and deploys the docs site to test-cabinet-docs or
test-cabinet-docs-staging.
deploy.sh writes a throwaway kustomization over
deployments/k8s/overlays/azure-<env> setting every service image to the
commit’s tag plus TCAB_DRIVER_IMAGE, TCAB_PUBLISHER_IMAGE,
TCAB_CONTAINER_REGISTRY and TCAB_CONTAINER_TAG=<sha>, renders it, refuses
anything outside the namespace, and applies and waits inside the private cluster
through az aks command invoke; a rollout that does not become ready is
described, logged, undone and fails the stage. amd64 jobs run on
Microsoft-hosted ubuntu-24.04 agents and arm64 jobs on the organization pool,
because the AKS nodes are arm64 and emulated Rust and wasm builds are far too
slow. No job holds a stored credential: registry pushes use the tcab-acr
connection, deploys tcab-deploy, the gg upload tcab-gg-publish, the mirror
the secure file github-mirror-key, and the docs and audio-store jobs the
secret variables. See building and the Kubernetes
overview.
The GitHub workflows are retired and GitHub is a mirror
Section titled “The GitHub workflows are retired and GitHub is a mirror”Every file under .github/workflows/ is deleted: ci.yml,
build-containers.yml, build-service-images.yml, deploy-docs.yml,
publish-reference.yml, release.yml, release-promote.yml and
binary-macos.yml. .github/ now holds only a README stating that the
repository lives on Azure Repos, that TheClockwyrks/TheTestCabinet is a mirror
receiving master, staging, nightly and v* tags from the pipeline’s
mirror job once a commit has passed its gates, that the push is forced so
anything committed on GitHub directly is lost, and that the mirror runs no CI.
The macOS validation job and the desktop release job leave with the workflows,
and no manifest, script or doc page names ghcr.io any more. Publishing a
reference implementation stays the manual tcab publish-reference flow.
Images move to the Test Cabinet container registry, pinned by commit sha
Section titled “Images move to the Test Cabinet container registry, pinned by commit sha”Every service image and every run-container image is pushed to
testcabinet.azurecr.io by the pipeline, tagged by the commit sha and never as
:latest, and core’s compiled default registry is the same host. The
azure-staging and azure-prod overlays carry no images: block and no
driver-image patch at all, because the pipeline’s deploy supplies every image
and TCAB_CONTAINER_TAG, so services, driver, publisher and the run images a
driver stages audio into always come from one build. Both clusters’ kubelet
identities hold AcrPull, so no pull secret exists in any overlay. The gg-ci
toolchain image and its hydrate script, which existed only for the GitHub
workflow, are deleted.
Releases are cut through the Azure tag route
Section titled “Releases are cut through the Azure tag route”A release is a vX.Y.Z tag on master in Azure Repos. Its pipeline run
executes the gates, and on the tag the binary job keeps the tcab it just
release-built, tested and smoke-tested as the tcab-linux and tcab-windows
pipeline artifacts, so the released binary is exactly the one the gate checked.
The same run uploads gg’s release objects and the mirror job pushes the tag to
GitHub. A tag builds no service or site of its own: services ship as the images
master deployed and the docs from the pipeline’s docs job. The gg version gate
(scripts/ci/gg-version-gate.sh) fails a tag run whose gg --version differs
from the tag with its v stripped, naming crates/gg and crates/core as the
crates to bump. The documented sequence is rel/vX.Y.Z into nightly,
nightly into staging as vX.Y.Z-rcN, staging into master, then the
tag. See releasing and
cut a release.
gg has a real version and its release binaries live on Azure Blob Storage
Section titled “gg has a real version and its release binaries live on Azure Blob Storage”crates/gg and crates/core carry version 0.7.0, pinned to each other by a
test, so gg --version no longer reports 0.0.0 and a run record’s
subject.harnessVersion distinguishes harness builds. gg’s release host is the
gg-releases container of the storage account testcabinetartifacts,
anonymous-read with no listing, laid out as
v<version>/gg-x86_64-unknown-linux-musl,
v<version>/gg-aarch64-unknown-linux-musl and v<version>/gg-reference.tar.gz.
scripts/build-gg-static.sh (arch-adaptive; cargo build-portable-gg is its
fixed x86_64 form) builds gg as a fully static musl executable so one binary
runs across the glibc Debian and Ubuntu run images; the pipeline’s gg_amd64
and gg_arm64 gate jobs build both with scripts/ci/gg-dist.sh, and
gg_publish runs scripts/ci/publish-gg.sh on every master build and every
v* tag, re-running the version gate and reading each upload back anonymously.
The install runs in the run’s setup stage and is recorded as setup, not session
time.
Building test-cabinet-gg, and therefore the driver image,
scripts/gg-reference.sh and the static build, needs all eleven language
toolchains present:
scripts/ci/install-gg-toolchains.sh (about 1.9 GB, idempotent) and
scripts/ci/install-gg-build-toolchains.sh for the C# guest, because every
arm’s guest component, catalogue and SDK library set are generated by the
build. A developer’s TCAB_GG_DOTNET_HOME tree that predates the vendored ICU
libraries is reinstalled by the installer on its next reconcile. See
releasing.
Cluster-scoped objects are a hand-applied bootstrap
Section titled “Cluster-scoped objects are a hand-applied bootstrap”The Namespace, the observability stack’s tcab-lgtm-node-metrics ClusterRole
and ClusterRoleBinding, and the cert-manager ClusterIssuer
letsencrypt-internal moved out of the base and overlays into
deployments/k8s/cluster/ with per-environment bootstraps
deployments/k8s/cluster/azure-staging and azure-prod, applied once by a
cluster administrator and re-applied whenever anything under that folder
changes. The azure-* overlays therefore render namespaced objects only, while
the generic staging, prod and local overlays still include the namespace and
observability bootstraps so one apply creates them. The pipeline’s tcab-deploy
identity holds exactly two roles per cluster, the custom command-invoke role
and RBAC Admin scoped to its namespace, deploy.sh refuses a render holding
anything outside the namespace, and scripts/ci/k8s-manifests.sh gates every
commit on the same bytes, so a compromised pipeline can at worst rewrite its
own namespace. See
the Kubernetes overview.
The ingest sidecar force-ingests the branch tip once per deploy
Section titled “The ingest sidecar force-ingests the branch tip once per deploy”The azure-staging and azure-prod overlays’ ingest sidecar runs one forced
POST /ingest of its branch tip (staging or master, cloned from the GitHub
mirror into the shared /state/checkout) when the backend pod starts, then
idles; every pipeline deploy changes the backend image tag and so restarts the
pod, so a deploy also publishes the catalog that shipped with it. Before
ingesting, the sidecar runs git submodule sync and
git submodule update --init --depth 1 to fetch the cold-storage baselines
shallowly, and falls back to ingesting without baseline media if that fails.
Catalog changes between deploys are published with
scripts/reingest-cluster.sh --env <env>, and the sidecar waits on /healthz
rather than /readyz to avoid deadlocking against itself. See
the control plane.
The Tauri desktop app is removed
Section titled “The Tauri desktop app is removed”crates/desktop, apps/desktop, the deployments/k8s/overlays/app overlay it
stood its own k3d cluster up from, the desktop build and sidecar scripts, the
desktop CI and release jobs, the Cargo workspace member with its dev profile and
Tauri dependency pins, the apps/desktop npm workspace entry and every
TCAB_DESKTOP_IMAGE_TAG reference are deleted. The docs site loses the
components/tauri section, the Arena component page takes its place in the
Components navigation, and every page that described a desktop console is
rewritten around the web console. In the UI package the harnessAuth
capability, the built-in local worker handle and the solo review-and-publish
path go, and core drops select_mode, SubscriptionFileStatus and
subscription_files; the canExecute flag stays because the gallery mounts the
shared application with it false. See the web
console.
Validation baselines move into the cold-storage submodule
Section titled “Validation baselines move into the cold-storage submodule”The captured baseline validation media (1.85 GiB, about 31k files, the bulk of
the repository) left test-cases/ for a cold-storage submodule at the
repository root whose tree mirrors the test-case tree, so a version’s baselines
live at
cold-storage/test-cases/<type>/<difficulty>/<slug>/<version>/validation-baseline/<engine>/<variant>/.
A plain clone leaves the directory empty and still builds and passes every gate,
and --recurse-submodules --shallow-submodules clones about 2 GB of media up
front. One resolver in core, ColdStorage::validation_baseline_dir, maps a
version root to its baseline directory beneath the cold-storage root;
tcab capture-baselines and tcab publish-reference write there and backend
ingest copies the counterpart into the staged version when present, so store,
API and snapshot are unchanged and an absent submodule yields versions with no
baseline media.
Because baselines are outside the version folder, the frozen digest never covers
them, so recapturing a frozen version’s baselines is allowed and changes no
score. Every frozen marker was rewritten with scripts/freeze.sh in the
migration commit, and freeze.sh now refuses only unstaged or untracked changes
under the directory, accepts several directories at once and an optional
--reason for the marker. The four frozen versions that depend on
@clockwyrks/particle-runtime (Spectra v1.0.0, Arc Foundry v1.0.0, Deepcore
v1.0.0 and Valence v1.0.0) also had their prompt and reference implementation
edited for the npm-scope rename, with the digest re-baselined and the reason
recorded on the marker. CI checkouts leave submodules off, and the pre-commit
large-file exemption for validation-baseline/ is gone. See frozen
versions and where baselines
live.
Submodules are addressed by relative URL, mirrored and gated on their master
Section titled “Submodules are addressed by relative URL, mirrored and gated on their master”.gitmodules names every submodule by a URL relative to the superproject
(../cold-storage), which git resolves against the remote the superproject was
cloned from: the sibling repository in the Azure project, or
TheClockwyrks/<name> on GitHub, so a recursive clone works from either host.
Each submodule repository carries its own .azure-pipelines/mirror.yml that
force-pushes its master to a GitHub repository of the same name. A new gate,
scripts/ci/submodule-pins.sh (job submodulepins, with an offline table
test), fails a superproject commit whose pinned submodule commit is not an
ancestor of that submodule’s master, so a commit that reaches the GitHub mirror
only ever names submodule commits the mirror already holds. A pipeline job that
reaches a submodule repository names it under its uses.repositories and the
pipeline lists it under resources.repositories, because the project scopes the
job token to referenced repositories; every checkout leaves submodules off, and
the pin gate reads the submodule through the Azure Repos API. See
building.
Image builds compile once, rebuild stale bases and ship an exact context
Section titled “Image builds compile once, rebuild stale bases and ship an exact context”The seven per-service Dockerfiles under deployments/images/ are replaced by
one multi-stage deployments/images/services.Dockerfile, of which the pipeline
builds each service as a stage, and containers/ gains tools/ (the shared
asset-tool builder), gg/, gg-toolchains/, audio-store/ and
full-stack-3d/ Dockerfiles. Every asset tool is compiled once in the shared
tools builder and copied into the 22 run images that bake an asset binary
(twenty asset kinds plus the two full-stack images). Service images each get
their own cargo target cache and the services and gg stages their own rustup
caches, fixing a could not rename downloaded file race on every cold builder
and the six-fold recompile of the workspace crates per make images.
The root .dockerignore allowlist re-excludes ignored shapes and lists each gg
helper script rather than a glob, because a wildcard negation makes BuildKit
walk the whole tree. scripts/ci/build-context.sh, now also a pre-commit hook,
asserts that every Dockerfile COPY, every gg guest package, every
include_str! tree, every toolchain-installer input and every package
stage-tcab-packages.mjs bakes survives every allowlist that can apply, that no
allowlist re-includes a wildcard family, and that the context holds no
git-ignored file. See building.
New CI gates, and the validator-constants gate is dropped
Section titled “New CI gates, and the validator-constants gate is dropped”The whole checkout is prettier-clean and gated: npm run lint:format
(scripts/format-check.mjs, run by the format job) checks every git-tracked
file prettier has a parser for, minus .prettierignore and the frozen versions
it derives from .frozen markers, sharded across cores with a content-keyed
cache, and npm run lint runs it after the workspace linters.
scripts/ci/spec-vocabulary-check.sh (specvocabulary, also pre-commit) fails
a non-frozen version whose prompt.hbs, specs/ or the shared preambles in
core name the project, bare tcab, benchmark, test case, evaluation, run
record, a review surface, the manifest, or a link to the repository or gallery;
frozen hits are reported, not failed. npm run typecheck:validators
(validators) runs tsc --noEmit over all 54 case validator projects,
scripts/ci/k8s-manifests.sh (manifests) renders every kustomization under
overlays/ and cluster/ and checks the deploy set is namespaced and names
every image at the registry and the commit,
scripts/ci/seeded-contract-check.sh runs on pre-commit and in the pipeline,
and the gg version gate runs on tags.
The validator-constants gate, its pre-commit hook, its Azure and GitHub jobs
and its README rows are removed as the wrong instrument: an awk scan of
TypeScript cannot decide whether a validator reads the build, and the rule it
enforced stays in the authoring guide, with Carom and Refract enforcing the
boundary through their own ESLint rules. The pre-commit hooks no longer run
clippy or rustdoc on commit; scripts/setup-hooks.sh installs
cargo fmt --check, the markdownlint and cspell prose gates, the frozen-version
check, the audio-pack lint, and the spec-vocabulary, seeded-contract, NUL-byte
and build-context gates. The large-file hook’s exemptions for
crates/gg/src/sandbox/guests/ and checkers/ are removed so a multi-megabyte
generated guest can never be committed again, and the reference-implementation
exemptions name both the reference-impl/ and references/<engine>/ layouts
plus test-cases/**/showcase/. See building.
The tasks/ issue board and the repository skills
Section titled “The tasks/ issue board and the repository skills”tasks/ is the issue board, one Markdown file per issue sorted into an area
folder (backend, backlog, gg-client, gg-languages, gg-sandbox, gg-sdk,
gg-telemetry, gg-tools, metrics, test-cases, v0.7.0) with completed issues moved
to a done/ folder beside them. CLAUDE.md states that nothing under tasks/
is authoritative and a landed issue’s durable conclusions belong in
apps/docs/, that all commits use Conventional Commits with an imperative
subject, and that multi-agent workflows are standing-authorized. The .claude
tree gains four PreToolUse hooks (block-done-task-writes, which makes
completed issues immutable, block-bare-sleep, block-cargo-test and
block-self-matching-pgrep) and eight skills (agent-prompts,
analyzing-run-costs, documentation, drain-issues-queue,
driving-gg-directly, gg-sdk-documentation, repo-tasks and test-cases).
The root README.md is reduced to a pointer at the documentation site.
The @clockwyrks scope, and the evaluation contract kept out of a workspace
Section titled “The @clockwyrks scope, and the evaluation contract kept out of a workspace”The npm scope of every in-repo package is @clockwyrks (seventeen workspace
packages at the release branch), so a seeded workspace’s package.json no
longer names the project as @test-cabinet/<package> resolved from
./.tcab/..., which the spec guide forbids a model from ever seeing. The
directory the harness vendors into a run repository is .vendor
(.vendor/packages/ for runtimes a case declares in packages,
.vendor/engine/ for the engine), and the seed commit’s git identity says
Clockwyrks; serve_validation_file falls back to the old .tcab/validation/
path for already-collected trees. Anything importing a runtime must use
@clockwyrks/voxel-runtime, @clockwyrks/particle-runtime or
@clockwyrks/asset-contract.
The voxel and particle runtimes are vendored into a model’s workspace at seed
time, and both depended on @clockwyrks/run-record, whose file: closure
dragged the evaluation contract into a case’s workspace. The ten rig and F-curve
types they need (ModelSpec, PartSpec, JointSpec, JointKindSpec,
AxisSpec, DriveKindSpec, InterpSpec, KeyframeSpec, AnimationSpec,
AnimationTrackSpec) now live in the new generated package
packages/asset-contract, emitted by contract-codegen alongside run-record
from the same Rust types; run-record re-exports them and is no longer staged
into the host package store at all. The references that vendor a runtime now
vendor the contract beside it, and scripts/ci/seeded-contract-check.sh greps
the generated package under a strict evaluation-vocabulary ban and walks the
real manifests asserting no seeded package reaches run-record. See run
records.
Dependency and package bumps
Section titled “Dependency and package bumps”Every workspace package except @clockwyrks/gg-sandbox, whose checker is the
TypeScript gg’s TS arm judges a program against and stays pinned at 5.9.3, moves
to TypeScript 6 and Vitest 5, the five Three consumers and the root overrides
to three ~0.185.0, and Playwright, @types/node and jsdom to the newest the
workspace accepts, with jest-dom held below 7 and jsdom below 30 for stated
reasons; the engines stay at 1.0.0, and the voxel and particle runtime sources
changed only by the rename, the contract import and the prettier pass. The root
build:packages also builds the four engine packages and the root gains
format, lint:format, test:scripts and typecheck:validators. Every
reference lockfile was regenerated and proved with a from-scratch npm ci, and
the seeded build vitest configs write coverage/test-report.json and
coverage/coverage-summary.json.
The Cargo workspace gains crates/code-analysis, crates/gg, ten
crates/gg-sandbox-artifacts/<arm> members and their shared build-support,
loses crates/desktop, swaps uuid for cuid2 (every id the services mint is
a CUID2), adds semver, flate2, wasmtime-wasi 45, wit-component and
wit-parser 0.248, tiktoken-rs, oxc and sourcemap, enables ts-rs’s
no-serde-warnings feature workspace-wide, and sets opt-level 3 on the wasmtime
and cranelift graph in the dev profile. rust-toolchain.toml stays at 1.97.1,
with a header explaining that a bump re-cuts gg’s Rust arm rlibs automatically.
The pipeline provisions Node 22 while the devcontainer ships Node 24.16.0. See
building.
The devcontainer supports four host runtimes and bakes gg’s toolchains
Section titled “The devcontainer supports four host runtimes and bakes gg’s toolchains”The host variant table is Docker on Linux, rootless Podman on Linux including a
DGX Spark, Docker on macOS covering Docker Desktop and OrbStack, and the new
Podman on macOS row; .devcontainer/setup-host.sh [variant] [--force] [--print]
detects the host and writes the uncommitted compose and env pair, reading the
podman socket directory and user-namespace mode off a real machine. The host
runtime socket mount moved out of the shared docker-compose.yml into the three
overrides that have one, and on macOS with Podman it is a bind-backed named
volume resolved inside the machine VM so k3d and the local stack still work.
gg’s eleven program-language toolchains are installed by the image build itself,
with postCreateCommand left as an idempotent reconciler, because
crates/gg/build.rs reflects every arm’s catalogue as a step of cargo build --workspace. Playwright’s Chromium and its system libraries are baked in so the
front-end gate can launch a browser after a rebuild, every per-platform download
resolves its architecture from uname -m, gh is pinned (2.97.0, GH_VERSION
overrides), the macOS Docker row binds the runtime’s synthesized SSH agent
socket, ffmpeg is declared because build-sample-pack.mjs silently produces
silent packs without it, and lsof, procps and iproute2 are in the apt list.
Pinned versions: Rust 1.97.1, Node 24.16.0, nextest 0.9.140, lazygit 0.61.1.
Testing gg: gates, the offline mock provider and driving gg by hand
Section titled “Testing gg: gates, the offline mock provider and driving gg by hand”gg’s test suite runs under nextest and carries structural gates rather than only unit tests: an authorship gate that compares each arm’s prepared bytes with the model’s, an isolation gate driven three wide sequentially and concurrently per language against five deliberately broken preparations, a build-count gate proving modules compile once, a docs gate that compiles every fenced program on the language pages through its arm, a prompt gate reading what the templates say, an agreement gate between catalogues and the operations table, and per-arm failure-shape cells. Compiled components are memory-mapped from a test-only disk cache, compiler JVMs run on the C1 JIT sized for two processors, and the JVM and Swift substrate tests were split to stay inside the nextest bound.
Any mock/… model id selects the scripted offline client and
TCAB_GG_FAKE_MODEL overrides the provider, so a configuration can be validated
for free, and the driving-gg-directly skill documents launching the bare
binary from a dev box with scripts/model-windows.sh printing modelWindows,
modelModalities and modelProviders for an invocation.
Twenty-seven closed issues give each of the eleven arms program-level coverage
of gg’s surface: per arm, a workspace test file drives every files and shell
operation and a session-side file drives every context, delegation,
programs, docs, views and session operation from a short program, one
test for the successful call and one per runtime failure mode the arm’s own SDK
documentation declares, plus a feedback test per arm and Java and Kotlin tests
driving all fourteen functions of the test-cabinet:gg/math interface. Thirteen
more give the tool crate call-level coverage: every name in ALL_TOOL_NAMES
reached through ToolRegistry::dispatch is driven with a synthesized ToolCall
through a real registry, each loop-intercepted tool is asserted to answer an
ordinary dispatch with the gg-defect refusal its declaration promises, and the
board, task, memory, filesystem, tree, search and compact tools each gain a test
per argument diagnostic, unknown id and store refusal a model can reach. See
invariants.
The Lattice factory designer
Section titled “The Lattice factory designer”apps/lattice-designer (@clockwyrks/lattice-designer) is a new dev-only React
and Vite workspace under apps/: a drag-and-drop designer that lays out a
Lattice factory, runs it through the real lattice-core engine, opens and saves
the case’s own scenarios in place, imports layouts, and exports a
scenario.json. It carries grid presets and recipe explanations, simulates the
lane splitter, and stages a board-size change through the themed dialog because
it discards the draft. Its README states that it is not part of the shipped
product and that a saved scored scenario still needs its oracle re-solved and
its playback bundle regenerated.
The telemetry crate can log to stderr
Section titled “The telemetry crate can log to stderr”test-cabinet-telemetry’s Config gains with_stderr_logging(), which writes
the human-readable fmt layer to standard error instead of standard output, and
the tcab CLI calls it so a subcommand whose stdout is piped (--json) is not
interleaved with log lines; services keep stdout. The observability page adds
the auth, artifact and arena services to the instrumented-process table and
drops the desktop app, and OTLP export remains opt-in on
OTEL_EXPORTER_OTLP_ENDPOINT with no new variables. See
observability.
The documentation site is rewritten and reorganized
Section titled “The documentation site is rewritten and reorganized”Every page under apps/docs/src/content/docs was rewritten to the documentation
skill’s standard and checked against the code. The site gained an Engines
section with one folder per engine split into APIs, concepts, usage, validators
and examples and a none page; the asset-generation manifest reference was
split from one page into an overview plus per-kind pages;
deployment/kubernetes.md split into deployment/kubernetes/ (overview,
control plane, run plane, postgres, internal ingress); and
development/building.md gained Cloning, Submodule URLs, Mirrors and pins, Cold
storage, the gg toolchains, portable builds and validator projects. The docs
tree is linted with markdownlint for the first time, the observability page
registers PromQL and TraceQL grammars so its query blocks highlight, and the
docs deploy from the pipeline’s docs job to test-cabinet-docs on master and to
the new test-cabinet-docs-staging project on staging.
Mechanisms removed before release
Section titled “Mechanisms removed before release”Response healing, the replay capability and fidelity axis, tcab gg-run,
tcab gg-replay and tcab gg-playback, the ownership param, the planning
capability, the tdd and plan-first machines, speculative best-of-K execution
and the per-slot provider override were deleted before v0.7.0 and are not part
of what ships. Because every configuration container refuses unknown keys, a
saved configuration still carrying healing, ownership or planning is
refused as an unknown field.
The schema gains 27 migrations
Section titled “The schema gains 27 migrations”Upgrading to v0.7.0 applies 27 migrations, taking the registered list from 24
to 51. All are additive or defaulted except three: m20260824_000035 drops two
probe columns, m20260824_000036 drops and recreates both model-probe tables,
so probe history from earlier development builds is discarded, and
m20260820_000032_drop_gg_agent_model_slots drops a column.
m20260723_000018_add_job_gg_config: nullablejob.gg_config_jsonholding a gg run’s capability set.m20260724_000019_create_gg_config: the per-accountgg_configtable of saved gg configurations.m20260724_000020_add_model_price_input_modalities: nullablemodel_price.input_modalities.m20260727_000021_create_comparison: the per-accountcomparisontable.m20260731_000022_add_run_gg_preset: nullablerun.gg_preset, the configuration name a gg run was launched from.m20260801_000024_add_run_updated_at:run.updated_at, seeded fromfinished_atin SQL.m20260801_000025_add_run_code_analyzer_version: nullablerun.code_analyzer_version.m20260801_000026_create_gg_saved_views: thegg_saved_queryandgg_dashboardtables.m20260818_000030_create_gg_agent: the per-accountgg_agentlibrary.m20260818_000031_add_gg_config_agent_sources: nullablegg_config.agent_sources_json.m20260820_000032_add_case_reference_build_engine:case_reference_build.engine, taken into the primary key with existing rows backfillednone.m20260820_000032_drop_gg_agent_model_slots: dropsgg_agent.model_slots_json.m20260822_000033_add_run_engine_slug: nullablerun.engine_slug, backfilled at boot in batches.m20260823_000034_create_model_probe: themodel_probeandmodel_probe_itemtables.m20260824_000035_probe_forced_submission: reshapes the probe tables for the forcedsubmit_programprotocol.m20260824_000036_probe_docview_scenarios: drops and recreates both probe tables for the scenario design.m20260824_000037_add_validator_ratings:run.validator_rated(false for existing rows), nullablerun.aesthetic, andreview.aesthetics('[]'for existing rows), the reviewer’s per-domain aesthetic ratings.m20260827_000038_add_review_aesthetic: nullablereview.aesthetic; a legacy per-domainaestheticsJSON collapses to its worst tier on read.m20260830_000039_add_run_record_readability:run.record_readable(true) andrun.record_format(0, so every existing row is re-decided at the first boot).m20260901_000040_add_gg_cell_identity:run.gg_models,job.gg_presetandjob.gg_models, backfilled at startup.m20260901_000041_add_gg_config_id:run.gg_config_idandjob.gg_config_id.m20260901_000042_create_backfill_state: thebackfill_statetable recording the one-timegg_config_idattribution pass.m20260901_000043_add_engine_pins:ladder_rung.engineandjob.engine_slug.m20260906_000044_add_job_started_at: nullablejob.started_at.m20260922_000045_add_provider_pin: nullablemodel_price.provider_pin(the observed pin) andmodel.provider_pin(the hand-set override).m20260923_000046_add_model_list_price: five nullable list-price columns onmodel(input, cached input and output, stored per token and edited per Mtok in the console, the date taken and the source), so every existing row reads as not yet priced.m20260923_000047_add_model_provider_policy: nullable provider policy columns onmodel(native quantization, input and output price ceilings, banned providers and the unknown-quantization allow list).