Files
shiro-neko/docs/memory.md
T
Muhammad Zakir Ramadhan 84c60f2022 Fix the loop stalling after compaction, add an external registry
The compaction bug, which is the important one:

beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.

The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.

So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.

Registry, via `/registry [list|search|add|remove|installed]`:

Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.

Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.

Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
  red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.

538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
2026-09-03 03:07:56 +07:00

189 lines
6.8 KiB
Markdown

# Memory and state
Four kinds of state, each with a different lifetime.
| State | Lives in | Survives |
|---|---|---|
| transcript | the message array | until compaction or `/clear` |
| task list | the system prompt, rebuilt each step | pruning and `/compact` |
| project memory | `~/.shiro-neko/memory/<hash>.json` | across sessions, forever |
| session record | `~/.shiro-neko/sessions/<uuid>.json` | until you delete it |
The split exists because compaction is destructive. `pruneMessages` deletes tool results and
`/compact` deletes the whole transcript, so anything recorded only in messages is lost
exactly when a long task needs it most.
## Project memory
Durable notes about the codebase, injected at the start of every session.
### `remember`
```
kind fact | decision | gotcha | command
text one self-contained line
```
- **fact** — how something is. "The API is versioned under `/v2`."
- **decision** — what was chosen and why. "We use snake_case for DB columns; the ORM
expects it."
- **gotcha** — a trap. "The migration must run before the seed or the FK fails."
- **command** — an invocation that works. "Tests run with `bun test`, not `npm test`."
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
### `recall`
Every term must appear. A match increments that entry's hit count, which protects it from
compaction later — an entry the agent actually uses is worth keeping verbatim.
### `forget`
Removes by substring, for a note that turned out wrong.
### What the agent sees
```
What you learned about this project in earlier sessions. Trust it, but verify anything
that contradicts what you can see in the code now:
- (command) tests run with bun test, not npm test
- (gotcha) the migration must run before the seed
- (decision) snake_case for DB columns, the ORM expects it
```
Top 20 by hit count, then recency. The "verify anything that contradicts" line matters:
memory goes stale and a confidently wrong note is worse than none.
### Compacting
`/memory` has the model merge entries. Two rules make it safe:
- Entries with at least one recall are kept verbatim and never merged.
- A model returning nothing parseable leaves the store untouched.
Without the second rule a bad response wipes everything the agent has learned.
`/notes` lists the store with hit counts. `--no-memory` disables loading and writing.
## Task list
`todo_write` replaces the whole list each call. Four states:
```
tasks ##########.............. 1/4 1 blocked
[x] read the pagination code
[~] fix the boundary
[ ] add a test
[!] update the docs (no write access to the wiki)
```
`blocked` requires a note saying what is blocking it. The tool warns when more than one task
is `in_progress`, when nothing is `in_progress` while work remains, or when a `blocked` task
has no note.
The list is re-rendered into the system prompt on **every step**, not once per turn — a
`todo_write` on step one has to be visible to step two. It is saved with the session and
restored by `-c` or `/resume`.
`/todos` shows it. The panel above the input shows it live.
## Sessions
Every turn autosaves, debounced 400 ms so a long tool loop does not hit the disk each step.
```json
{
"id": "0193ab2c-…",
"createdAt": "…", "updatedAt": "…",
"cwd": "/home/you/project",
"provider": "openai", "model": "gpt-5",
"title": "why does the pagination test fail?",
"inputTokens": 48210, "outputTokens": 3105,
"costUsd": 0.0913,
"notebook": { "todos": [ … ] },
"messages": [ … ]
}
```
```bash
shiro -c # newest session for this directory
shiro -r 0193ab2c # by id or unique prefix
```
```
/sessions list the last 15
/resume <id>
/save write now instead of waiting for the debounce
```
A corrupt session file is skipped rather than crashing the list.
## Compaction
Two mechanisms.
**Automatic**, at roughly 120k estimated tokens: `pruneMessages` strips reasoning and older
tool calls from what goes on the wire. Local history is untouched, so the transcript on your
screen stays complete. The turn reports it:
```
context compacted: 192 messages pruned to 15 on the wire
```
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
and outcomes, what remains — and it replaces the transcript entirely.
### The pruning repair
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
before it did.
**A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a
reasoning item and keeps the message item from the same response. That message carries a
provider `itemId`, and the OpenAI responses provider serialises anything with one as
`{ type: 'item_reference', id }` — a pointer to an item stored on their side, which depends on
the reasoning item that is now gone:
```
400 Item 'msg_…' of type 'message' was provided without its required 'reasoning' item: 'rs_…'
```
The two carry different ids, so they cannot be matched by id. What links them is the
assistant message they arrived in — one message is one response. `detachOrphanedItems` strips
the `itemId` from those parts. Without one the same content is serialised **inline**, which
carries no dependency on anything stored, so the turn survives intact.
Dropping the parts instead was the first attempt, and it was wrong in a way that only showed
up over a long turn: on a reasoning model every tool call carries an itemId, so after the
first compaction the model could no longer see what it had already run. It re-ran the same
tools until the step limit ended the turn. The history is the model's memory; compaction may
shorten it but must not blank it.
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
*messages*, not pairs, so the cut can land between the assistant message holding a `tool-call`
and the `tool` message answering it:
```
400 No tool call found for function call output with call_id call_…
```
`dropOrphanedResults` drops any result whose call id no longer survives. The reverse is left
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
like, and dropping it would break `/resume`.
## Prompt history
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the
input walk it; down past the newest restores what you were typing. While an `@` token is open
those keys move in the file picker instead.
Stored at `~/.shiro-neko/history/<hash>.json`, where the hash is a SHA-256 prefix of the
project path.
## The prompt queue
Not persisted, and deliberately so. A prompt typed during a turn lives in memory until the turn
ends, then runs. `esc` clears it along with aborting the turn, and quitting discards it — a
queued thought that fires on next launch, against a workspace that has since changed, is worse
than a lost one.