Orchestrators
Overview
Section titled “Overview”An orchestrator decides how a test case’s harness sessions are conducted. It owns the loop around the harness: how many sessions to run, what each one is told, and when the work is done. The harness layer still owns each individual session, meaning how the harness is invoked, how its usage is parsed, and how its activity is translated into events.
Orchestration is harness-agnostic: an orchestrator drives sessions the same way
whichever harness is selected. It is therefore a distinct run dimension,
selected per run alongside the test case, variant, harness, and model, and
recorded as orchestratorSlug in the run
record.
Orchestrator structure
Section titled “Orchestrator structure”An orchestrator carries no in-tree code. It is a directory containing:
- a manifest,
orchestrator.toml, with the orchestrator’sslug,name,description, therunnerentrypoint, and any[params]the runner reads; - a runner script, the entrypoint named by the manifest.
The built-in orchestrators use exactly the same machinery as a custom one, so a new strategy can be supplied entirely from outside this repository. See External orchestrators.
The built-in orchestrators live under orchestrators/<slug>/ in the repo,
embedded into the core library at build time so a backend-driven worker with no
checkout resolves them the same way the CLI does. They are catalogued under
Orchestrators. There is one:
one-shot, a single harness session driven to completion. This is the default.
The execution model
Section titled “The execution model”All setup is shared and happens once, exactly as for a single-session run: the
run-container image is pulled, the container is started, authentication is
applied, the harness CLI is installed and probed, and the test case’s init
command runs. Only then does the orchestrator take over, in place of the single
harness invocation.
The orchestrator’s runner script runs inside the run container, so it drives
sessions and inspects progress from where the work happens. Its commands run
natively against the seeded workspace at /work, and it reads status and marker
files on the workspace filesystem directly. The whole runner script is bounded
by the run’s maximum runtime, just as a single
session is.
The tcab-session wrapper
Section titled “The tcab-session wrapper”A runner script must be able to invoke a harness session without knowing any
harness-specific details. The run therefore writes a tcab-session wrapper onto
the run user’s PATH inside the container before running the orchestrator.
Invoking tcab-session "<prompt>" runs the selected harness’s CLI with the
adapter’s exact session arguments, substituting the prompt. The harness’s output
flows back through the runner script’s stream, where it is parsed for usage and
translated into events exactly as a single session is.
The wrapper emits a sentinel line around each session so the run can segment the
output into sessions and sum each session’s reported usage into the run’s
totals. A one-shot run has exactly one segment, so its metrics are identical
to a run with no orchestration layer at all.
Where the usage is read from
Section titled “Where the usage is read from”The wrapper writes each session’s stdout, sentinels and all, to a session log
inside the container at ~/.tcab/sessions.log as well as streaming it back. The
streamed copy produces live events. The run’s
usage and cost are read from the log once the
runner has exited, which keeps a session’s reported usage independent of the
exec stream surviving to its last line.
If the log cannot be read, the streamed copy stands in. Either way a session that arrives without its closing sentinel is reported as a warning on the run’s event stream, with its tokens recorded as unknown rather than as none.
Runner environment contract
Section titled “Runner environment contract”The runner script is handed everything it needs through its environment:
| Variable | Meaning |
|---|---|
TCAB_PROMPT | The rendered test-case prompt (the goal). An orchestrator wraps this with its own protocol before passing it to tcab-session. |
TCAB_WORKSPACE | The seeded workspace directory (/work). |
TCAB_DEADLINE | Epoch seconds after which the run’s maximum runtime is exhausted. A multi-session runner checks this to stop gracefully before the hard cap. |
TCAB_PARAM_<KEY> | Each [params] entry from the manifest, upper-cased (for example marker_file becomes TCAB_PARAM_MARKER_FILE). |
one-shot’s runner is a single tcab-session "$TCAB_PROMPT".
Budget and timeouts
Section titled “Budget and timeouts”The run’s maximum runtime bounds the whole orchestrator the same way it bounds a
single session: when it is exceeded, the run is stopped. That hard cap is a
backstop. A multi-session orchestrator is expected to manage the budget itself
through TCAB_DEADLINE, stopping after the current session rather than starting
one it cannot finish, and to exit successfully with partial work when the budget
runs out. Because the runner exits normally, the produced workspace is still
collected and validated, so running out of
budget yields a likely-incomplete result rather than a discarded one.
An orchestrator’s scratch files, such as a status or marker file, live under a dot-directory in the workspace so they are easy to keep out of the collected implementation.
External orchestrators
Section titled “External orchestrators”Because an orchestrator is a directory of data, one can be supplied at run time from a directory anywhere on disk. The directory has the same shape as a built-in, and its manifest’s own slug is authoritative for the run record. A custom orchestrator is resolved purely at run time.
Selecting an orchestrator
Section titled “Selecting an orchestrator”An orchestrator is selected per run, defaulting to one-shot. Every runner
selects one and the resolved slug is recorded on the run: the
CLI through --orchestrator, the
driver through built-in slugs only, since it has
no access to a submitter’s local directory, and the run-execution UI.
A non-default orchestrator is limited to the test types that build a program over a working session: end-to-end, full-stack, and game-jam. Those are where a multi-session implementation is needed, as the other types build a single artifact in one pass. The run rejects a non-default orchestrator for any other test type before any container is started.