Skip to content

Memories

A memory is a durable note the model writes and maintains itself: a slug, a one-line description, and a body. It takes the same shape as a skill, with a description shown up front and a body that survives compaction. The model curates the set as the run goes.

Because the model controls them, gg bounds them, so that self-curated notes cannot crowd out the working context. A mutation that would breach a limit is refused with a message telling the model to revise, evict or delete. Nothing is truncated.

Under every strategy the memories live inside gg. The memory calls are the only way to create, read, revise or remove one, and nothing is written to the workspace. Two of the strategies talk about an index file and memory files because that is the mental model a model already has. Since a model cannot rewrite its own index by writing over a file, gg’s accounting of what the window holds always agrees with what is stored.

A slug is up to 64 characters of letters, digits, -, _ and ..

The capability’s implementation selects the memory strategy. The strategies differ in what is always in the context window. An enabled memories capability names one, whatever its scope, and a name gg does not recognize refuses the launch, naming the three that exist. What a profile binding somebody else’s store is saying by naming one is under Inheritance and the strategy.

StrategyAlways in contextTools
scratchpadevery memory, body and allwrite_memory, update_memory, delete_memory
markdownthe index (one slug — description line per memory)create_memory, read_memory, edit_memory, delete_memory
keyword-searchnothingcreate_memory, read_memory, edit_memory, delete_memory, search_memories

All three work in both execution modes. Under responses-as-code the same calls are the gg.memories module (gg.memories.createMemory, gg.memories.searchMemories), and a call the active strategy does not offer is refused as unavailable.

Every memory’s body is pinned in the window and crosses a compaction boundary verbatim. The count and the per-body length bound it, and the aggregate ceiling maxTotalLen bounds the pinned block as a whole.

A pinned index over memory files, one - `slug` — description line each. The bodies stay out of the window until read_memory brings one in. create_memory adds the entry and delete_memory removes it, so the model never writes the index directly.

The index bounds the population: a create whose entry would push the index over its limit is refused. There is no separate count limit, because every memory must have a line in the index anyway.

Memory files with no index, so nothing about them is in the window until the model looks. search_memories takes an array of keywords and ranks the matches by how many distinct keywords a memory mentions, then by how often, over case-insensitive substring matching across each memory’s slug, description and body. Each hit carries a short excerpt. A description is optional under this strategy and is shown with search results.

The two file-shaped strategies revise a memory by search/replace, exactly as edit_file does: the model quotes the text it is changing, which must appear exactly once. Appending means quoting the last line and replacing it with itself plus what is being added. An edit that would leave the memory empty is refused, and the model is told to delete it instead.

One slider sits in the capability’s Features box, per agent.

FeatureDefaultWhat switching it changes
Revise memoriesonOff withholds update_memory, edit_memory and delete_memory, so memories are append-only.

Every limit is set through the capability’s params, and an enabled memories capability writes all six whichever strategy it selects. Writing the whole block on every arm is what lets one sweep hand every arm the same params. 0 is how a limit is turned off, so a run that wants no aggregate ceiling writes maxTotalLen: 0. A missing param, a param name the capability does not know, and a value gg cannot read as a whole count each refuse the launch.

ParamApplies toWhat it bounds
maxCountscratchpad, keyword-searchHow many memories the set holds at once.
maxLenPerMemoryall threeCharacters in one memory’s body.
maxTotalLenscratchpadCharacters across every pinned body together.
maxLenIndexmarkdownCharacters in the pinned index.
maxLenDescriptionall threeCharacters in one memory’s description.
maxResultskeyword-searchHits one search_memories call returns.

Lengths are in characters of a memory’s body. A description is bounded separately, by maxLenDescription under every strategy, because it is the one field every turn pays for: on a markdown run every description is a line of the pinned index, and a search hit is mostly description. An over-long description is refused rather than truncated, and is checked before the body limits, so a call that breaches both is told about the cheaper fix first. A memory’s code counts against none of these limits.

For example, a markdown run with a small index and no per-memory limit:

{
"id": "memories",
"enabled": true,
"implementation": "markdown",
"params": {
"scope": "isolated",
"maxCount": 64,
"maxLenPerMemory": 0,
"maxTotalLen": 0,
"maxLenIndex": 4096,
"maxLenDescription": 256,
"maxResults": 25
}
}

A memory can carry code beside its body, exactly as a skill can. Both halves are responses-as-code only: the native tool schemas carry no code fields, so a tool-calling run sees a memory’s description and body alone.

import { memories } from "gg";
memories.writeMemory({
name: "csv-tools",
description: "Parsing the vendor CSV exports, which quote inconsistently.",
body: "The third column is sometimes quoted and sometimes not; parseCsv handles both.",
code: "export function parseCsv(text: string) { /* … */ }",
onUse:
'views.openText("csv-notes", "Row 1 is a header on exports after March.");',
});

createMemory and updateMemory take the same shape. An update replaces both halves, so omitting them clears them, on the rule the description and the body already follow. A native-mode update leaves both alone, since its schema has no way to say anything about them.

  • code is a module every program the agent writes from then on may import, key being the slug in camel case (csv-tools → csvTools), deduplicated if something else holds it. The load opens a documentation view per function it declares, each naming the key and the line a program writes to reach it.
  • onUse is a script gg runs on every use, after the turn’s program has ended, so the views it opens arrive on the next turn. It runs on the agent’s own grants, it cannot end the session, and its source is never shown back to the model.

A memory’s code loads by the rule its strategy already sets for what is in context.

StrategyThe code loads on
scratchpadthe write, since every memory is in the window from the moment it exists
markdown, keyword-searchthe read, since a memory’s code follows its body into context

A loaded module survives a compaction: it costs no tokens and is never summarized, so a boundary leaves it loaded and nothing makes an agent re-read a memory to get back a helper it already has. A fork and a succession start with nothing loaded.

Each half is bounded at 32 768 characters, refused in the same limit-exceeded voice every other cap uses. That limit protects the transpiler, which parses untrusted source.

Both halves are compiled at the moment they load, and what happens next depends on which moment that is.

  • A write that loads (the scratchpad’s) is refused, with a located diagnostic. Storing a module that can never load would store something that only fails later.
  • A read that loads (the two file-shaped strategies’) returns the memory, with the diagnostic appended to the body. The body is what the model asked for.

A compiler that could not finish is gg’s own defect and ends the run under internal_error, on the terms in gg’s own defects. The compiler’s crash detail goes to the run’s operator. The call that met it is refused as an io-error rather than an invalid-argument, since the model’s argument was never the thing that failed, and the turn that refusal fails is recorded against gg rather than against the model.

The scope param says which instance an agent binds. Under isolated a memory instance belongs to one agent instance: a subagent starts with an empty notebook, and nothing it writes is seen by anyone else. The other three link agents.

scopeWhich instance the agent binds
isolatedA fresh one, per agent instance
sharedOne per agent profile: every instance of it in the run, including those running in parallel
inheritedIts spawner’s, read/write, when it was spawned as a subagent; its own otherwise
read-onlyAs inherited, but this agent may not write
{
"id": "memories",
"enabled": true,
"implementation": "markdown",
"params": {
"scope": "inherited",
"maxCount": 64,
"maxLenPerMemory": 8192,
"maxTotalLen": 0,
"maxLenIndex": 4096,
"maxLenDescription": 256,
"maxResults": 25
}
}

A value gg does not recognize refuses the launch, and the refusal names every such value in the configuration at once. Setting scope on a capability that is switched off does not: a disabled capability records the configuration the arm would have used, so the on and off arms of one comparison stay symmetric.

Two rules make the four coherent. read-only restricts an inherited handle and nothing else: an agent that ends up with an instance of its own under read-only may write it. And write access belongs to the holder rather than to the store, so a read-only agent’s own inherited subagent gets a read/write handle onto that same store, and inheritance chains through a subagent of a subagent to whatever the top of the chain created.

inherited and read-only are defined against how the agent was started, so the four scopes read differently at each of the places gg starts one.

Started asisolatedsharedinheritedread-only
The run’s rootits ownthe profile’s instanceits own (nothing spawned it)its own, and writable
An issue’s implementerits ownthe profile’s instanceits own (it is top-level, not a subagent)its own, and writable
A subagentits ownthe profile’s instanceits spawner’s, read/write; its own if the spawner keeps no memoriesits spawner’s, read-only
A reviewer or merge agentits ownthe profile’s instanceits own (dispatched directly, not through a spawner)its own, and writable
A forkan independent copythe same instanceits forker’s, read/writeits forker’s, read-only
A successor (exec or an FSM transition)transferredits own profile’s instance, re-boundtransferredtransferred, read-only

The last row is a transfer rather than a binding, so it obeys the transfer rules first: a successor whose profile turns memories off gets none, and one that organizes them under a different strategy gets a fresh instance and is told why. shared overrides the transfer, because shared means bound to the profile: a successor running a shared profile re-binds that profile’s instance instead of keeping the one it was handed. Otherwise one profile would curate two notebooks at once, which is the situation the scope exists to prevent.

Exactly the read calls: read_memory under the two file-shaped strategies, plus search_memories under keyword-search. The write calls are never contributed to the toolset, so the model is shown no schema for one, no program has one in scope, and the system prompt says the memories are another agent’s. A write that reaches gg anyway is refused with an explanation naming the calls the agent does have.

Under scratchpad a read-only holder gets no memory tools at all. That strategy has no read call, because its memories are the pinned block; it reads them by having them in its window.

A profile configured for memory-compaction that declares its memories read-only refuses the launch. It has no call that could satisfy a memory compaction, and a run that could never satisfy its own compaction gate would wedge against a full window. An instance whose live scope resolves read-only under that strategy is gg’s own defect and ends the run under internal_error.

A fork copies almost everything its agent holds, and memories follow the scope rather than the copy. An isolated notebook is copied and the two diverge. Every scope that links agents stays linked across the fork: the original and its copy hold one store, and each is told what the other writes.

A store is read by the calls its own strategy offers, so an inherited or read-only agent organizes its memories the way its spawner does.

Every enabled memories capability names its implementation, an inheriting one included. Most of the rows above hand such an agent a store somebody else organized — but four of them hand it nothing, and it keeps a notebook of its own instead: the run’s root, an issue’s implementer, a reviewer or merge agent, and a subagent whose spawner keeps no memories. An arm gg picked for those would be a memories study running on an organization nobody wrote, so the profile names the one it works in.

Naming one is a claim about the store the agent ends up with, so the two ends of an inheritance have to agree. A roster pairing whose two profiles name different strategies refuses the launch. Where only the live spawner settles the pairing, a disagreement is gg’s own defect and ends the run under internal_error.

A scope is what the configuration asked for, and several rows above are cases where gg resolves it to something else. Every memory store therefore carries an id, and the console reads scoping off the ids rather than off the declaration. The Modules tab lists each store once with every instance holding it. A profile’s row on the Agents tab says whether its instances share one store or take one each, and names the divergence where that is not what the scope asked for. Each instance’s modules/memories file marks the store it bound and who else is in it.

Under every scope but isolated, several agents can hold one store at once. When one of them adds, revises or removes a memory, every other holder is told in its next prompt:

Another agent sharing your memories has made changes since your last turn:
- added `deploy-runbook` — how the staging cluster is rolled
- updated `api-conventions` — error envelope + pagination rules
- deleted `scratch-notes`
Read one with `read_memory` if it bears on what you are doing.

Deletions are included, since acting on a memory that has since been deleted is the failure the notice exists to prevent. Each memory gets one line saying where it ended up, however many writes touched it. Under scratchpad, where there is no read call, the notice carries each memory’s body inline.

Three properties hold:

  • The pinned index does not change. It is rebuilt at a compaction boundary and nowhere else. The notice is appended at the tail of the window as an ordinary message, so the whole previous request stays a byte-identical prefix of the next one and the prompt cache is undisturbed.
  • Each notice is delivered exactly once per holder. Every holder keeps its own watermark, a holder is never told about its own writes, and a holder that joins late starts at the store’s head rather than being handed a backlog.
  • Notices are ephemeral, so a compaction boundary sweeps them and they are not re-issued. The boundary rebuilds the pinned block first, so everything a notice announced crosses in the block or, under keyword-search, stays findable by search.

Memories survive a compaction boundary under every strategy. What crosses in the window is what the strategy pins: every body under scratchpad, the index alone under markdown, nothing under keyword-search, where the memories are still stored and still findable on the other side.

The memory-compaction strategy, which asks the agent to record its working state instead of writing a summary, asks for the calls the run’s memory strategy offers.

The pinned block is rebuilt at a boundary and only there. Between boundaries the thread carries the news, as the call the model made and the confirmation that answered it. What the thread cannot carry is a compaction, which drops the very tool results the model read its memories out of, so gg rebuilds the block immediately before the window is rewritten.

A pinned index can therefore lag: a memory created since the last compaction is stored and readable but is not listed there yet. The system prompt says so, so the model does not read a missing line as a lost memory.

An agent’s first turn is the other point at which the block is built. Some holders bind a store somebody else filled: an inheriting subagent, the second instance of a shared profile, a successor handed the module by a transfer. Such a holder opens on memories it never made a call for, so gg builds the block once as the window is opened. An empty store renders nothing.

Two events are streamed, and they answer different questions.

memory_state is the set as it stands: every memory currently held with its description, character count and line count, the totals, the limits in force, and the run’s peaks. The peaks are the most memories, characters and lines ever held at once, and they are reported because the live figures alone are misleading: a model that curates well spends its budget, prunes, and finishes holding almost nothing. The event also carries how the emitting agent holds the store, its scope and whether it may write, which distinguishes two agents curating one shared store from two agents that happen to hold the same notes.

memory_revision is one entry of an append-only record of what the model did: every successful mutation, in order, carrying the memory’s text as of that revision. It is what a snapshot cannot show, such as a memory written and later deleted, or the earlier wording of one that was revised. Revision numbers are per slug and keep counting across a delete, so a name that is discarded and re-created reads as the history it is. A refused mutation records nothing.

A revision is reported once, on the stream of the agent that made it, however many agents hold the store it landed in. Every other holder re-emits its own memory_state instead: its panel changed, but the write was not its work.

The console folds the two together into the Memories panel: current-and-peak totals, a treemap of every memory ever held sized by its character count, and each memory, deleted ones included, expandable into its revision history.