Expand the documentation with measured figures and operational detail

Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 09:26:37 +07:00
parent 84c60f2022
commit 9b978fdbe1
10 changed files with 583 additions and 72 deletions
+49 -1
View File
@@ -13,6 +13,9 @@ The split exists because compaction is destructive. `pruneMessages` deletes tool
`/compact` deletes the whole transcript, so anything recorded only in messages is lost
exactly when a long task needs it most.
The practical rule: if it should survive this turn, `todo_write` it. If it should survive this
session, `remember` it. The transcript is for the conversation, not for storage.
## Project memory
Durable notes about the codebase, injected at the start of every session.
@@ -32,11 +35,21 @@ text one self-contained line
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
The kinds are not decoration: they are what the model reads back at boot, and they set how much
to trust a note. A `command` is verifiable in one run. A `decision` explains why the obvious
alternative was not taken, which is the thing a newcomer most often gets wrong.
"One self-contained line" is the part that matters most. A note reading "use the new approach"
is worthless next session — there is no conversation left to say which approach.
### `recall`
Every term must appear. A match increments that entry's hit count, which protects it from
compaction later — an entry the agent actually uses is worth keeping verbatim.
AND rather than OR, on purpose: "migration seed order" should find the one note about that,
not every note mentioning any of the three words. Returns the 15 most recent matches.
### `forget`
Removes by substring, for a note that turned out wrong.
@@ -61,7 +74,9 @@ memory goes stale and a confidently wrong note is worse than none.
- Entries with at least one recall are kept verbatim and never merged.
- A model returning nothing parseable leaves the store untouched.
Without the second rule a bad response wipes everything the agent has learned.
Without the second rule a bad response wipes everything the agent has learned. The store also
has to be past 60 entries with at least two unused ones before anything happens, so `/memory`
on a small store is a deliberate no-op rather than a rewrite.
`/notes` lists the store with hit counts. `--no-memory` disables loading and writing.
@@ -118,6 +133,14 @@ shiro -r 0193ab2c # by id or unique prefix
A corrupt session file is skipped rather than crashing the list.
Ids are UUIDv7, so they sort by creation time and a prefix is usually enough to identify one.
`-c` matches on `cwd`, so it picks up the newest session **for this directory** rather than the
newest overall — two projects side by side do not steal each other's `-c`.
Resuming restores the messages and the task list, but not the model or agent variant: those come
from the current config and flags. A session started with `--agent deep` resumes as `default`
unless you pass it again.
## Compaction
Two mechanisms.
@@ -130,9 +153,24 @@ screen stays complete. The turn reports it:
context compacted: 192 messages pruned to 15 on the wire
```
The status bar warns before that happens: context is shown as a percentage of the threshold,
amber from two thirds, red at 90.
What gets discarded, in order: reasoning items first, then tool calls and their results older
than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not
conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what
lets a turn continue rather than restart.
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
and outcomes, what remains — and it replaces the transcript entirely.
The difference is which history is destroyed. Automatic pruning touches the wire only, so
scrolling back still shows everything and `/save` records everything. `/compact` replaces the
real message array, so it is irreversible for that session.
Use `/compact` when a session has drifted across several unrelated tasks and the early part is
noise. Let automatic pruning handle a single long task, since it keeps the recent work intact.
### The pruning repair
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
@@ -171,6 +209,16 @@ and the `tool` message answering it:
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
like, and dropping it would break `/resume`.
### What compaction still does not do
It tells the model the history was pruned but not what was in it. A decision from forty messages
ago can be contradicted with confidence, because from the model's side that span never existed.
Summarising the discarded part is [next on the list](../TODO.md).
The threshold is also measured with `JSON.stringify(messages).length / 4`, which is an estimate.
It is fine for deciding when to prune and wrong enough that it should not be read as a token
count — the real numbers in `/cost` come from the provider.
## Prompt history
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the