Skip to content

Turn outcomes

gg judges every turn on the sentence the execution ceilings are enforced on: was the work the turn declared carried out as declared? One turn_outcome rides on the stream per turn, carrying that judgement, emitted from the one seam where gg records an outcome against an agent’s ceilings. That seam is what makes it impossible for this event and the ceilings to disagree about what an error is. A run emits exactly one per model call it made.

{
"type": "turn_outcome",
"outcome": "error",
"error": "transpile",
"errorType": "transpile_syntax",
"consecutiveErrors": 2,
"turns": 31,
}
FieldWhat it carries
outcomeprogressed (the turn did its declared work), finished (the turn ended the session, which is never an error), error, or fatal (gg’s own machinery broke, recorded so the turn accounting stays whole and deliberately not charged to the model’s error budget).
errorWhy, at the base level, on an error outcome, and absent on every other, so error != null and outcome == "error" are the same statement. One of model_api, transpile, program_fault, sandbox_limit, missing_completion.
errorTypeWhy, specifically: the leaf of the two-level taxonomy. Present on exactly the turns error is, and its base is the error beside it, since gg holds one value and derives both halves when it emits the event.
consecutiveErrorsThis agent’s failing streak after this turn.
turnsHow many turns this agent has recorded, including this one. It is this agent’s own running total rather than the run’s, and the same figure its turn ceiling is measured against.
loopAbortsHow many replies loop detection discarded before this turn produced one. Omitted when zero.
loopAbortWords, loopAbortCharsHow much generated output those discarded replies produced, measured by the detector as they streamed. Present on exactly the turns loopAborts is. There is no token count and no price on these figures, since an abandoned stream reports no usage: the price is read back at session end and lands in the run’s total cost.
responseChars, responseOutputTokensThe reply’s size in two units: characters, and completion tokens (output plus reasoning, the unit a provider’s output cap is measured in). The character count covers the model’s raw text together with the program string of every submit_program call the turn made, which under responses as code carries nearly all of the turn’s output. Omitted when zero.

consecutiveErrors is 0 on every non-error turn, including a finished or fatal one that followed failures. gg’s internal counter is cleared only by a turn that carried out its declared work, so an agent that failed twice and then finished still holds a count of 2, and publishing that would put a streak on a turn that did not fail. The invariant the stream guarantees is consecutiveErrors > 0 if and only if outcome is error.

The count is carried rather than re-derived because it is per agent while the stream is run-wide. Turns from concurrently running agents interleave arbitrarily, so a reader folding the stream could only reconstruct a run-wide streak, which is an artefact of scheduling rather than a fact about any agent.

The summary folds maxResponseChars and maxResponseOutputTokens as maxima over every turn except one recorded model_length_capped, whatever the turn’s outcome. That one turn’s reply was cut off at the provider’s output cap and is the reply an output ceiling exists to cut, so a ceiling chosen from it would be chosen from itself. Every other reply was generated whole and is data, so a long program that then failed to compile still sets the figure a later ceiling is chosen from.

Under responses as code this event sits beside code_execution and does not duplicate it. That one reports what a program did and exists only in that mode; this is the mode-agnostic judgement of the turn. A tool-calling run emits turn_outcome too, which is what lets one error rate be compared across both modes. A tool-calling turn that ends with no tool call is reported here as missing_completion, since ending a session is always an explicit call.

The five error values are the base kinds: whose layer failed. They are what an execution ceiling acts on, what a cross-run comparison groups by, and what every stored run keys on, so they are stable. transpile in particular keeps a name wider than its meaning, because the value is what persisted records carry.

Underneath each base kind sits an errorType, and that is where the failure is named. There are twenty-one types, one per distinction gg makes.

Base kindTypes under it
model_apimodel_auth (the credential was refused), model_rejected (another non-retryable 4xx), model_retry_exhausted (the provider never served the request), model_response_loop (it served it and loop detection discarded every answer), model_vision_unsupported, model_parse (a 2xx reply gg could not read), model_provider_mismatch (a provider other than the candidate in force served the call), model_timeout (the call ran into gg’s per-call ceiling), model_length_capped (the reply hit the provider’s output cap and was rejected whole)
transpiletranspile_syntax, transpile_compile (the language’s compiler read the whole program and rejected it), transpile_unsupported
program_faultprogram_api_error (an uncaught failed call: the model is fighting the API rather than mis-writing it), program_unknown_name (it reached for something this run does not offer it, either a name that is not in scope or a call the host refused as unavailable), program_throw
sandbox_limitsandbox_timeout, sandbox_out_of_memory, sandbox_trap
missing_completionmissing_completion_no_call, missing_completion_compaction (a prose reply where a compaction was pending, which is answered differently), missing_completion_no_program (a responses-as-code reply that made no submit_program call)

Every type’s id names its base, because a “top error types” ranking shows one row per type with no heading over it. Each also carries a human-readable label. The labels live in Rust beside the variants and are generated into the TypeScript contract, so a type gg gains arrives already labelled and gg’s own error log line names the failure in the same words the console does.

Folded from those same events onto the session summary, so numerator and denominator can never come from different mechanisms.

"errors": { "turns": 96, "errors": 4, "maxConsecutive": 2,
"modelApi": 1, "transpile": 2, "programFault": 1,
"sandboxLimit": 0, "missingCompletion": 0,
"loopAborts": 7, "loopAbortWords": 21455,
"loopAbortChars": 136150,
"byType": { "model_response_loop": 1, "transpile_syntax": 2,
"program_api_error": 1 },
"toolFailures": { "not-found": 12, "invalid-argument": 3 } }

turns is the denominator and counts every turn whatever its outcome, a fatal one included, so the accounting stays whole even though no ceiling observes it. errors is exactly the sum of the five per-kind counters. maxConsecutive is the maximum over agents of the per-turn consecutiveErrors above, which is the only honest way to summarise a per-agent counter on a run-wide record, and the peak of the same counter maxConsecutiveErrors is enforced on.

No percentage is stored. The error rate is errors / turns and the reader divides. A stored rate is a figure that can disagree with its own denominator after a rounding change, a partially recorded run, or a reader that averages two runs’ rates, and the one thing that must be trustworthy here is that the numbers add up.

byType is the same errors split by their specific type, keyed by the errorType wire id. Two things hold: it sums to errors, and regrouping it by each type’s base reproduces the six named counters exactly. It is a map keyed by a string rather than by the enum, so that a run recorded by a newer gg still reads back in an older backend or console: an unknown enum key would fail the whole summary, where an unknown string degrades to one unlabelled row in a ranking. It is omitted from the wire when empty, which is exactly a run with no errors.

toolFailures counts calls rather than turns: every dispatched tool call that failed, by class, whether or not the program that made it caught the failure. It is a rollup of dispatches, so a code-mode call that never reached a tool is not in it. Those are on the stream as api_result with their class, where a console folds them. The summary’s top-level toolCalls counts every dispatched call, failed or not. It is the denominator toolFailures is read against, so toolCalls minus the failures is the count of calls that succeeded. It is omitted from the wire when zero, so a reader must treat a summary with failures but no toolCalls as one whose total was not recorded rather than as a contradiction.

A rejected length-capped reply is an error turn here and is additionally recorded on its own terms: a response_rejected event carries the reply’s size, usage, cost and serving provider, and the summary’s rejectedResponses rollup sums the count, tokens and cost. The same spend is in the run’s total cost and outside its work cost, so a degenerate generation never makes the run’s work look expensive.

Two things are deliberately excluded from the error counts. A tool call that failed inside a program that carried on is counted in toolFailures instead: the program handled it, which is the point of the typed surface, and charging it would make the one capability that expects failures the one that cannot survive them. Per-agent attribution is excluded because these are run-wide totals, and the per-agent breakdown lives on the stream, where every turn_outcome rides on its own agent’s id.

The console folds the same figures off the live stream for its Dashboard, beside the turn count they share a denominator with, because “seven” and “seven of two hundred” are not the same claim: errored turns, the rate they are of, the longest streak, the replies loop detection discarded when there were any, and the most common error types with their counts, each row carrying its base kind as a badge. An errored turn renders no row of its own on the live event feed, since gg already logs why a turn failed in the recorded type’s vocabulary.

The whole summary is flattened into the query language’s document, so every field above is directly queryable, including the open breakdowns, whose keys become fields of their own.

has.summary:true | stats avg(summary.errors.maxConsecutive) as streak by model
has.summary:true | stats sum(summary.errors.byType.program_api_error) as fights by model

Every request of a run names the one candidate in force for its model: provider.only carries its provider, provider.quantizations its level, and fallbacks are refused. Each call’s usage event records the provider that served it. A response from any other provider ends the run as a harness failure: the turn is recorded as model_provider_mismatch, and its error names the candidate and the served provider.

Two events record what the run saw of its providers. Both are emitted on the stream of the agent whose request saw the fault.

{ "type": "provider_fault", "modelId": "z-ai/glm-5.3",
"provider": "Z.AI", "fault": "stall" }
{ "type": "provider_switch", "modelId": "z-ai/glm-5.3",
"from": "Z.AI", "to": "Baidu", "fault": "failed_call",
"detail": "the stream stalled: no delta from the model for 60s (provider: Z.AI)" }
  • provider_fault is one stall or one unexpected cache miss, fault being stall or cache_miss, against the provider the request was sent to.
  • provider_switch is one move to the model’s next candidate. fault is failed_call for a spent retry schedule, unavailable for a candidate OpenRouter refused with 404, or cache_miss for a provider that reached providerCacheMissLimit. detail is the last failure’s cause, or the miss count that reached the limit.

OpenRouter names the serving provider on each response, and provider-specific failures are only diagnosable from a record that says who served what. The summary therefore carries providerStats: one slice per (provider, model) pair observed, folded from the same stream as the rollups above. A run that stayed on its first candidate has one slice per model, and a run that moved has one per provider it spent against.

"providerStats": [
{ "provider": "DeepInfra", "modelId": "qwen/qwen3.8-2.4t-a95b",
"calls": 41, "tokens": { "uncachedInput": 63167, "output": 71310 },
"cost": { "comparable": 0.70, "actual": 0.70 },
"turns": 41, "working": 39, "errors": { "transpile_compile": 2 },
"stalls": 1, "cacheMisses": 0 }
]

Each slice records the calls that reported usage (calls, with their summed tokens and cost), the length-capped replies the provider served (rejected), the turns attributed to it (turns, the working turns among them, and an errors map keyed by the same errorType wire ids byType uses), and the provider_fault events against it (stalls and cacheMisses). Two invariants hold: the slices’ turns sum to errors.turns, and a slice’s turns minus working minus its error count is its fatal turns.

A turn is attributed to the provider named by its own call’s usage, prompt or response_rejected event. A call that produced no reply — a model timeout — names no provider, so its turn lands on the slice with no provider key, as does any call whose gateway named none. The modelId is the one the agent’s usage deltas named, so a turn before an agent’s first usage report carries none. The array is omitted from the wire when no call, turn or rejection was ever folded into it.

limit_exceeded is emitted once by each agent that stops on an execution ceiling, immediately before its loop returns, carrying which ceiling, what it was set to, what was observed, and after how many turns. It is a structured event rather than only a log line, because “which ceiling ends my runs, at what value?” is a question a study asks of thousands of runs, and prose cannot be grouped by.