Expand the documentation with measured figures and operational detail

Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 09:26:37 +07:00
parent 84c60f2022
commit 9b978fdbe1
10 changed files with 583 additions and 72 deletions
+50 -1
View File
@@ -62,12 +62,35 @@ from.
| `high` | `high` | `budget_tokens: 38400` |
| `max` | `xhigh` | maximum budget |
Verified against both wire formats rather than assumed.
Verified against both wire formats rather than assumed. The vocabulary is deliberately ours:
`off` through `max` means the same thing whichever provider is configured, and switching
providers mid-session does not change what `/think high` asks for.
Higher costs more and takes longer. `off` on a hard problem produces confident wrong
answers; `max` on a rename wastes a few cents and several seconds. The variants pick
sensible defaults, so reach for `/think` only when a specific turn needs something else.
Reasoning is also charged as output tokens, so `max` shows up in `/cost` even on a turn where
the model wrote two lines. And reasoning is the **first thing compaction discards** — see
[memory](memory.md#compaction) — so a long turn at `max` pays for thinking that will not be on
the wire by the end of it.
## Steps
`maxSteps` caps how many model calls one turn may make. A step is one request: a tool call and
its result, or the final text.
| Variant | Steps |
|---|---|
| `quick` | 12 |
| `default`, `plan`, `review` | 50 |
| `deep` | 80 |
The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops
mid-work with whatever it has, which is why `quick`'s 12 suits a rename and would strand a
refactor. If turns regularly hit the cap on the same kind of task, the task wants `deep`
rather than a higher number.
## Overriding
`--agent deep --think low` gives you `deep`'s tools, steps, and appendix with a low thinking
@@ -97,3 +120,29 @@ told which tools need approval and to verify with the project's tests.
A prompt that describes a withheld tool teaches the model to attempt calls that cannot
succeed, which is why the description is generated from the live tool set.
Three rules flip on what is available:
| Condition | `default` says | `plan` says |
|---|---|---|
| can edit | "these need approval; if denied, stop and ask" | "you have no tools that change anything" |
| can run commands | "verify with the project's build or tests" | "say what should be run rather than claiming it passed" |
| can ask | "ask rather than guess when two readings differ" | (same, unless headless) |
The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for
the full set — cheaper per turn as well as safer.
## Which to reach for
- **`default`** for anything you have not thought about. It is the right answer most of the time.
- **`quick`** for a rename, a typo, a one-line fix. Its value is not the model being cheaper but
the absence of deliberation latency on work that needs none.
- **`deep`** when the first attempt already failed, or the cause is unclear. Asking for more than
one hypothesis is the actual difference; the thinking budget is secondary.
- **`plan`** before a change you are not sure about. Read-only means the plan cannot quietly
become a half-applied edit.
- **`review`** on a diff or a module. In headless CI this is the one that needs no `--yolo`,
because it holds no tool that can modify anything — see [headless](headless.md).
Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and
nothing about the history.
+100 -19
View File
@@ -48,21 +48,48 @@ Written by `/provider`, editable by hand. Every field is optional.
`/provider` offers these. Each sets `baseURL` and the wire protocol for you.
| Preset | Protocol | Endpoint |
|---|---|---|
| Anthropic | `anthropic` | `api.anthropic.com/v1` |
| OpenAI | `openai` | `api.openai.com/v1` |
| OpenRouter | `openai` | `openrouter.ai/api/v1` |
| Groq | `openai` | `api.groq.com/openai/v1` |
| DeepSeek | `openai` | `api.deepseek.com/v1` |
| xAI | `openai` | `api.x.ai/v1` |
| Ollama | `openai` | `localhost:11434/v1` |
| LM Studio | `openai` | `localhost:1234/v1` |
| Custom OpenAI-compatible | `openai` | you supply it |
| Custom Anthropic-compatible | `anthropic` | you supply it |
| Preset | Protocol | Endpoint | Env var checked |
|---|---|---|---|
| Anthropic | `anthropic` | `api.anthropic.com/v1` | `ANTHROPIC_API_KEY` |
| OpenAI | `openai` | `api.openai.com/v1` | `OPENAI_API_KEY` |
| OpenRouter | `openai` | `openrouter.ai/api/v1` | `OPENROUTER_API_KEY` |
| Groq | `openai` | `api.groq.com/openai/v1` | `GROQ_API_KEY` |
| DeepSeek | `openai` | `api.deepseek.com/v1` | `DEEPSEEK_API_KEY` |
| xAI | `openai` | `api.x.ai/v1` | `XAI_API_KEY` |
| Ollama | `openai` | `localhost:11434/v1` | none, keyless |
| LM Studio | `openai` | `localhost:1234/v1` | none, keyless |
| Custom OpenAI-compatible | `openai` | you supply it | none |
| Custom Anthropic-compatible | `anthropic` | you supply it | none |
After the key is entered, `GET /v1/models` is called and the list becomes a picker. If the
endpoint does not implement it, you type the model id instead — the setup still completes.
`provider` is the **wire protocol**, not the vendor. Groq, DeepSeek, xAI, OpenRouter, Ollama,
and LM Studio all speak `openai`; only Anthropic speaks `anthropic`. Two things differ between
them: the auth header (`Authorization: Bearer` versus `x-api-key`), and how thinking levels map.
After the key is entered, `GET /v1/models` is called and the list becomes a picker. Both
protocols expose that endpoint with the same `data[].id` shape, so one code path handles both.
If the endpoint does not implement it, a preset with a known model list falls back to that;
otherwise you type the model id and setup still completes.
Anything the picker offers is a model the endpoint actually reports, which is more reliable than
a hard-coded list — that is why the fallback lists are short and only exist for Anthropic and
OpenAI.
## Cost estimates
`/cost` and the status bar price a turn from a table in `src/pricing.ts`, matched by longest
prefix on the model id, so `claude-sonnet-4-5-20250929` resolves via `claude-sonnet-4-5`. An
OpenRouter-style `anthropic/claude-sonnet-4-5` has its vendor prefix stripped first.
An unknown model is reported as unpriced rather than guessed:
```
4210 in / 88 out tokens (llama-3.3-70b is unpriced)
```
Two limits worth knowing. The rates are hand-entered and drift as vendors change them, so treat
the figure as an estimate, not a bill. And the token counts come from the provider's usage
report, while `~ctx` in the status bar is `JSON.stringify(messages).length / 4` — good enough to
decide when to compact, wrong enough that it should not be read as a token count.
## Environment variables
@@ -74,12 +101,21 @@ endpoint does not implement it, you type the model id instead — the setup stil
| `SHIRO_API_KEY` | overrides `apiKey` |
| `ANTHROPIC_API_KEY` | used when `provider` is `anthropic` and no key is set |
| `OPENAI_API_KEY` | used when `provider` is `openai` and no key is set |
| `SHIRO_HOME` | relocates config, sessions, memory, history, and user skills |
| `SHIRO_HOME` | relocates config, sessions, memory, history, user skills, and installs |
| `SHIRO_INSTALL_DIR` | where `install:local` and the installers put the binary |
| `SHIRO_REPO` | which GitHub repo the installers download from |
| `SHIRO_VERSION` | pins the version the installers fetch |
`SHIRO_HOME` is what the test suite uses to keep a run out of your real config.
`SHIRO_HOME` is what the test suite uses to keep a run out of your real config. It is also the
way to run two isolated setups side by side — a work profile and a personal one — since it moves
every piece of state at once:
```bash
SHIRO_HOME=~/work-shiro shiro
```
A key on the command line ends up in your shell history and in `ps`. `SHIRO_API_KEY` in front of
one command is better; `/provider` writing to `config.json` is better still.
## Flags
@@ -103,13 +139,31 @@ cat file | shiro -p prompt read from stdin
| `--no-mcp` | skip MCP servers |
| `--no-subagent` | omit the `task` tool |
| `--no-instructions` | ignore `AGENTS.md` and friends |
| `--no-skills` | ignore builtin and project skills |
| `--no-plugins` | disable all plugins, including the guard |
| `--no-skills` | ignore builtin, installed, and project skills |
| `--no-plugins` | disable all plugins, builtin and installed, including the guard |
| `--no-memory` | do not load or write project memory |
| `--yolo` | skip every approval prompt |
| `-v`, `--version` | version, bun version, platform, source or compiled |
| `-h`, `--help` | usage |
The `--no-*` flags exist for isolating a problem. All six together strip the agent to its
built-in tools and nothing else, which answers "is this the loop or something layered on it?"
in one run:
```bash
shiro --no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp
```
`--no-plugins` also disables the guard, so `rm -rf` becomes an ordinary approval prompt.
Reasonable while debugging, not something to leave on.
An unknown value fails at startup with the valid list rather than falling back silently:
```
$ shiro --agent turbo
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review
```
## Where things live
```
@@ -133,7 +187,8 @@ Project files:
```
Memory and history file names are SHA-256 prefixes of the absolute project path, because a
path is not a safe filename.
path is not a safe filename. Two consequences: moving a project loses its memory and history,
and two checkouts of the same repo at different paths keep separate ones.
## OpenAI reasoning models
@@ -142,5 +197,31 @@ Newer OpenAI models reject function tools on `/v1/chat/completions` and require
on the first switches to the second, sticks for the rest of the session, and prints one
notice. Retryable failures — 429 and 5xx — are left to the SDK's backoff instead.
Only those six codes qualify, because they mean "this endpoint cannot serve this request shape".
A 401 is a wrong key and switching endpoints would only produce a second 401 with a more
confusing message.
The switch is sticky on purpose: once an endpoint rejects the shape it rejects every later step
too, so re-probing would waste a round trip per step of every turn.
Third-party endpoints get a plain chat-completions model with no fallback probe, since they
do not implement `/v1/responses`.
The two endpoints also differ in how they carry assistant history, which is where compaction gets
interesting — see [memory](memory.md#the-pruning-repair).
## Config that changes behaviour subtly
Three fields do more than they look like they do.
**`thinking`** costs money and latency on every turn, not just hard ones. `off` on a hard problem
produces confident wrong answers; `max` on a rename wastes cents and seconds. The agent variants
already pick sensible levels — see [agents](agents.md).
**`toolSets`** removes tools from the model's view entirely. If the agent stops using a tool you
expected, check the startup header for which sets loaded: an unrecognised name is dropped
silently, so `"gti"` reads as "git is off". See [tools](tools.md#tool-sets).
**`registryUrl`** is the whole trust decision for installed skills and plugins. There are no
signatures, so pointing it at an index means trusting whoever controls that URL — including for
whatever they publish later. See [registry](registry.md).
+57
View File
@@ -74,6 +74,29 @@ told and can respond to it.
if shiro -p "does this build?" --yolo; then echo ok; else echo failed; fi
```
That distinction is deliberate and it has a consequence: **a successful run says nothing about
whether the answer was yes.** `0` means the turn completed, not that the build passed. To gate CI
on the content, read the output:
```bash
shiro -p "Does this build? Answer only YES or NO." --json --yolo \
| jq -r 'select(.type=="text") | .text' | grep -q YES
```
Anything that must fail the build has to be asserted on text or, better, on the exit code of a
real command the agent ran.
## Timeouts
There is no wall-clock limit on a headless run. Three things bound it:
- `maxSteps` per variant — 12 for `quick`, 50 by default, 80 for `deep`.
- The `timeout` the model passes to `bash`, 120 s by default and 600 s at most.
- Whatever your CI runner enforces, which is the only hard stop.
Interactively `ctrl-c` kills one command and keeps the turn. Headless has no terminal for that, so
a signal ends the run. In CI, prefer `--agent quick` and a runner timeout over hoping.
## Sessions
Headless runs save like interactive ones, so `-c` picks up where one left off:
@@ -83,6 +106,14 @@ shiro -p "start the refactor" --yolo
shiro -p "now update the tests" --yolo -c
```
Useful, and worth knowing the shape of: each `-p` run is **one turn**, and `-c` resumes the newest
session for that directory. Two concurrent runs in the same directory therefore fight over the
same session, and the second overwrites the first. Pass `-r <id>` to keep parallel runs separate,
or point them at different `SHIRO_HOME` directories.
Memory also accumulates. An unattended loop calling `remember` writes to the project store like
any other run, so `--no-memory` is worth considering for a job that runs on every push.
## What is withheld
The `ask` tool is not offered at all, rather than being offered and left to hang. The model
@@ -131,8 +162,34 @@ env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
```
Two more worth having in a workflow. Trim the tool schema to what the job needs, since a CI run
pays for it on every step:
```yaml
- run: echo '{ "toolSets": [] }' > ~/.shiro-neko/config.json
```
And keep an unattended job from inheriting an installed skill nobody reviewed:
```yaml
- run: shiro -p "..." --agent review --no-skills --no-plugins
```
`--no-skills` matters more in CI than locally: a skill installed from a registry is instructions
in the system prompt, and CI is exactly where nobody is watching what it says. See
[registry](registry.md).
## Cost control
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
spend ceiling yet — see [TODO.md](../TODO.md).
What a run actually costs is in the `done` event, so a wrapper can total it:
```bash
shiro -p "..." --json --yolo | jq -r 'select(.type=="done" and .inputTokens) | "\(.inputTokens) in, \(.outputTokens) out"'
```
The token fields are optional: an aborted turn emits `done` with neither, which is why the filter
checks for one rather than assuming it.
+46 -5
View File
@@ -28,11 +28,24 @@ them in `~/.shiro-neko/config.json` and they appear alongside the builtins.
```
**stdio** servers take `command`, and optionally `args`, `env`, `cwd`. The process is spawned
at startup and closed on exit.
at startup and closed on exit. `env` is merged over the inherited environment, so a server
inherits your `PATH` unless you replace it.
**Remote** servers take `url`, and optionally `type` (`http` or `sse`, default `http`) and
`headers`.
A token in `headers` sits in `config.json` in plain text, same as `apiKey`. For anything beyond
a local dev token, prefer a stdio server that reads its own credential from the environment.
## Startup cost
Servers connect **in parallel**, so the slowest one sets how long startup takes rather than the
sum of them. `npx -y some-server` re-resolves the package on each launch; installing it and
calling the binary directly is usually the difference between a noticeable wait and none.
`--no-mcp` skips them all, which is also the quickest way to tell whether a slow start is MCP
or something else.
## Naming
Tools arrive as `mcp__<server>__<tool>`. A server named `fs` exposing `read_file` becomes
@@ -79,14 +92,42 @@ calling one.
## Cost
Each tool adds roughly 550 characters of schema to every request. A server exposing twenty
tools costs about 2,750 tokens per turn, sent whether or not the model uses any of them.
Each tool adds its name, description, and JSON schema to every request. The built-ins average
548 bytes; MCP tools vary with how verbose the server's schema is. A server exposing twenty
tools costs roughly 2,750 tokens per turn, sent whether or not the model uses any of them.
Prefer servers with a focused tool set. If one exposes many tools you never use, it is worth
finding a narrower server or writing one.
MCP tools are **not** covered by `toolSets` — that budget only governs the built-ins. There is
no per-server switch either, so the choice is a server or no server, and `--no-mcp` for all of
them. If one exposes many tools you never use, a narrower server is worth finding or writing.
`/tools` shows the count both ways:
```
tools
26 offered this turn of 26 registered
```
A gap between the two numbers means a tool set or a read-only agent variant is withholding
something. MCP tools never appear in that gap.
## Writing a server
Any MCP-compliant server works. A minimal stdio one needs three methods: `initialize`,
`tools/list`, and `tools/call`. The test suite includes one at
`test/fixtures/mcp-stub.ts` — about 50 lines, and useful as a starting point.
The suite runs it as a **real subprocess** rather than mocking the transport, because the parts
that break in practice are the handshake and the framing, and a mock asserts neither.
## Debugging a server
A server that starts but returns nothing useful is the harder case. In order of speed:
1. `/tools` — did the tools arrive at all? A server with no tools is a `tools/list` problem.
2. `shiro -p "call mcp__x__y with ..." --json --yolo` — the exact `tool-call` input and
`tool-result` output, one JSON object per line.
3. Run the server by hand: `echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | your-server`.
If that is wrong, nothing above it can be right.
For an HTTP server, `curl -X POST $URL -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'`
answers the same question without shiro in the way.
+49 -1
View File
@@ -13,6 +13,9 @@ The split exists because compaction is destructive. `pruneMessages` deletes tool
`/compact` deletes the whole transcript, so anything recorded only in messages is lost
exactly when a long task needs it most.
The practical rule: if it should survive this turn, `todo_write` it. If it should survive this
session, `remember` it. The transcript is for the conversation, not for storage.
## Project memory
Durable notes about the codebase, injected at the start of every session.
@@ -32,11 +35,21 @@ text one self-contained line
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
The kinds are not decoration: they are what the model reads back at boot, and they set how much
to trust a note. A `command` is verifiable in one run. A `decision` explains why the obvious
alternative was not taken, which is the thing a newcomer most often gets wrong.
"One self-contained line" is the part that matters most. A note reading "use the new approach"
is worthless next session — there is no conversation left to say which approach.
### `recall`
Every term must appear. A match increments that entry's hit count, which protects it from
compaction later — an entry the agent actually uses is worth keeping verbatim.
AND rather than OR, on purpose: "migration seed order" should find the one note about that,
not every note mentioning any of the three words. Returns the 15 most recent matches.
### `forget`
Removes by substring, for a note that turned out wrong.
@@ -61,7 +74,9 @@ memory goes stale and a confidently wrong note is worse than none.
- Entries with at least one recall are kept verbatim and never merged.
- A model returning nothing parseable leaves the store untouched.
Without the second rule a bad response wipes everything the agent has learned.
Without the second rule a bad response wipes everything the agent has learned. The store also
has to be past 60 entries with at least two unused ones before anything happens, so `/memory`
on a small store is a deliberate no-op rather than a rewrite.
`/notes` lists the store with hit counts. `--no-memory` disables loading and writing.
@@ -118,6 +133,14 @@ shiro -r 0193ab2c # by id or unique prefix
A corrupt session file is skipped rather than crashing the list.
Ids are UUIDv7, so they sort by creation time and a prefix is usually enough to identify one.
`-c` matches on `cwd`, so it picks up the newest session **for this directory** rather than the
newest overall — two projects side by side do not steal each other's `-c`.
Resuming restores the messages and the task list, but not the model or agent variant: those come
from the current config and flags. A session started with `--agent deep` resumes as `default`
unless you pass it again.
## Compaction
Two mechanisms.
@@ -130,9 +153,24 @@ screen stays complete. The turn reports it:
context compacted: 192 messages pruned to 15 on the wire
```
The status bar warns before that happens: context is shown as a percentage of the threshold,
amber from two thirds, red at 90.
What gets discarded, in order: reasoning items first, then tool calls and their results older
than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not
conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what
lets a turn continue rather than restart.
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
and outcomes, what remains — and it replaces the transcript entirely.
The difference is which history is destroyed. Automatic pruning touches the wire only, so
scrolling back still shows everything and `/save` records everything. `/compact` replaces the
real message array, so it is irreversible for that session.
Use `/compact` when a session has drifted across several unrelated tasks and the early part is
noise. Let automatic pruning handle a single long task, since it keeps the recent work intact.
### The pruning repair
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
@@ -171,6 +209,16 @@ and the `tool` message answering it:
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
like, and dropping it would break `/resume`.
### What compaction still does not do
It tells the model the history was pruned but not what was in it. A decision from forty messages
ago can be contradicted with confidence, because from the model's side that span never existed.
Summarising the discarded part is [next on the list](../TODO.md).
The threshold is also measured with `JSON.stringify(messages).length / 4`, which is an estimate.
It is fine for deciding when to prune and wrong enough that it should not be read as a token
count — the real numbers in `/cost` come from the provider.
## Prompt history
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the
+9 -2
View File
@@ -70,10 +70,10 @@ holding `a` through a batch of edits will approve one of these without reading i
| `rm -rf`, `rm -f` | recursive or forced delete |
| `git reset --hard` | discards uncommitted work |
| `git clean -f` | deletes untracked files |
| `git push --force`, `-f` | rewrites remote history |
| `git push --force`, `--force-with-lease`, `-f` | rewrites remote history |
| `git branch -D` | deletes a branch without a merge check |
| `DROP TABLE`, `TRUNCATE` | destroys database data |
| `mkfs`, `dd of=/dev/…` | writes to a raw device |
| `mkfs`, `dd of=/dev/…`, `> /dev/sd…` | writes to a raw device |
| `chmod 777` | makes files world-writable |
| `shutdown`, `reboot`, `halt` | affects the whole machine |
| `:(){ :\|:& };:` | fork bomb |
@@ -88,6 +88,13 @@ The model is told to relay the command rather than work around it. `rm build/one
`git push origin feature`, and `git commit` all pass — the patterns target irreversibility,
not the commands themselves.
Two honest limits. The patterns match the command **string**, so `bash -c "$(echo cm0gLXJm | base64 -d)"`
is not caught, and neither is a script the agent wrote and then ran. And it only inspects `bash`:
a `write_file` overwriting something important is an approval question, not a guard question.
The guard is the last line before a command runs; `ctrl-c` is the one after. A pattern the guard
does not know about is still interruptible by hand — see [tools](tools.md#bash).
### `time` (default on)
Adds `current_time`, returning ISO 8601 plus the local string. Auto-approved; it reads
+50
View File
@@ -126,3 +126,53 @@ add skill:review` disambiguates, and an ambiguous name is refused rather than gu
A private index is just a URL you control. There is no account, no token, and no telemetry —
`/registry` makes exactly one GET for the index and one for the entry you install.
## Publishing
Two files and a static host. GitHub raw works, and so does anything that serves JSON over https.
```
your-registry/
index.json
skills/migration.md
plugins/no-secrets.json
```
Three rules the validator enforces, so worth getting right first:
- The name in `index.json` must match the name inside the file. A skill's frontmatter `name` and a
plugin manifest's `name` are both checked against the index entry.
- Names are `^[a-z0-9][a-z0-9-]*$`. No uppercase, no dots, no slashes.
- A plugin needs at least one deny rule. A manifest with an `appendix` and no rules is prompt
text, which is what a skill is for.
Test it locally before publishing. `registryUrl` accepts `http://localhost`, so:
```bash
cd your-registry && python -m http.server 8000
```
```json
{ "registryUrl": "http://localhost:8000/index.json" }
```
`/registry` then exercises the real fetch, the real validation, and the real install path against
your files. That is the whole loop, without pushing anything.
## Troubleshooting
**"the registry index is malformed: …"** — the message names the first failing field. The usual
causes are an uppercase name, a `url` that is not https, or a plugin entry with no `deny`.
**"X calls itself Y but the index calls it X"** — the file's own name disagrees with the index.
Fix one of the two; the check exists so an index cannot serve something else under a name you
trusted.
**"invalid pattern …"** — a `pathPattern` or `commandPattern` is not a valid regex. Remember it is
JSON, so a backslash needs doubling: `\\.env$`, not `\.env$`.
**Installed but nothing happens** — installs load at startup. Restart, then check `/skills` or
`/plugins` for the entry and its origin.
**In `/plugins` with an error beside it** — the manifest on disk no longer validates. It is skipped
rather than fatal, so the agent still starts; `/registry remove` and reinstall.
+29 -3
View File
@@ -3,9 +3,9 @@
A skill is a markdown file with instructions for one kind of task. Only its name and
description sit in the system prompt; the body is loaded on demand.
That split matters. Four bundled skills are 4,659 characters of body but 681 characters of
catalogue. Putting every body in the prompt would cost that on every request, for
instructions that are relevant to one turn in twenty.
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
would cost that on every turn, for instructions relevant to one turn in twenty.
## Format
@@ -28,6 +28,14 @@ Never deploy from a dirty working tree.
description is what the model matches against, so write it as a trigger — "use when asked
to X" — not as a summary.
The frontmatter reader handles those two fields and nothing else. A real YAML parser would be a
dependency for two strings, so lists, nesting, and multi-line values are not supported: keep both
on one line. Quotes around a value are stripped. A body over 20,000 characters is truncated.
A file that fails to parse is skipped silently rather than reported, which is worth knowing when
a skill you wrote does not appear in `/skills` — the usual cause is a missing `---` fence or a
description spilling onto a second line.
## Where they load from
Four sources, later overriding earlier by name:
@@ -79,6 +87,24 @@ before you start working, and follow it as if the user had written it:
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
follow it for this task. The call needs no approval — it reads nothing outside the binary.
"Before you start working" is the load-bearing phrase. A skill loaded after the work is done is
wasted tokens, and the failure mode in practice is a model that reads the catalogue, decides it
already knows, and never calls the tool. A description written as a trigger is what prevents that.
Loading one costs its body, once, in that turn's context. A 3,000-character skill is cheaper than
one wrong approach it prevents, and more expensive than the catalogue line that would have been
enough.
## Verifying a skill loaded
```bash
shiro -p "fix the failing pagination test" --json --yolo | grep skill
```
`--json` shows the `tool-call` for `skill` with the name it chose, or its absence. If the model
never calls it on a task the skill was written for, the description is the thing to change — not
the body.
## Writing a good one
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled
+180 -31
View File
@@ -27,19 +27,42 @@ y allow once | a always allow edit_file | n deny
`a` whitelists that tool for the rest of the session. `n` tells the model it was denied and
to ask what to do instead. `--yolo` skips all prompts.
The approval is enforced by the SDK, not by the tools. A denied call **provably never
executes**: the SDK never reaches the tool's `execute`, so a tool cannot forget to honour a
denial or opt out of the check. See [architecture](architecture.md#why-approval-goes-through-the-sdk).
MCP tools are gated as a group because they are third-party code with unknown side effects —
`mcp__fs__read_file` sounds harmless and might not be. See [MCP](mcp.md).
**The guard runs before all of this.** It is not an approval — it is a refusal, and `--yolo`
does not reach it. See [plugins](plugins.md).
## Tool sets
Each tool costs roughly 550 characters of JSON schema on every request, and selection
accuracy drops as the list grows. Sets let you switch off what a project does not need:
Each tool costs its name, its description, and its JSON schema on **every request**. Measured
across the fourteen built-ins:
| Set | Tools |
|---|---|
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` |
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` |
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` |
| Tool | Bytes | Tool | Bytes |
|---|---|---|---|
| `read_many_files` | 972 | `git_blame` | 499 |
| `multi_edit` | 934 | `git_log` | 484 |
| `edit_file` | 618 | `git_diff` | 473 |
| `grep` | 595 | `bash` | 466 |
| `list_dir` | 594 | `git_show` | 432 |
| `read_file` | 526 | `git_status` | 292 |
| `glob` | 499 | `write_file` | 289 |
7,673 bytes for all fourteen, averaging 548. Roughly 1,900 tokens per request before your
prompt or the conversation. Selection accuracy also falls as the list grows: a model choosing
between six tools picks better than one choosing between twenty.
Sets let you switch off what a project does not need:
| Set | Tools | Cost |
|---|---|---|
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B |
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` | ~2,500 B |
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B |
```json
{ "toolSets": ["edit-plus"] }
@@ -50,7 +73,32 @@ agent is not an agent. A disabled set reaches neither the wire nor the system pr
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
Session, plugin, and MCP tools are not part of this budget and are never gated here.
`/tools` shows which set each live tool came from.
An unrecognised set name is dropped silently. The header line at startup shows which sets
actually loaded, so a typo reads as "that set is off" rather than as an error — worth checking
if a tool you expected is missing.
`/tools` shows which set each live tool came from:
```
tools
20 offered this turn of 22 registered
- `bash` core
- `git_diff` git
- `list_dir` edit-plus
- `remember`
```
A tool with no set is a session, plugin, or MCP tool.
### Which sets to keep
Both extra sets earn their place in most projects, but not all:
- **No git in the repo?** `git` is 2,180 bytes the model can never use. Switch it off.
- **A model that handles many tools badly?** `{ "toolSets": [] }` trims to six, which is the
smallest set that still lets the agent work.
- **Reading a lot, editing rarely?** Keep `edit-plus` for `list_dir` and `read_many_files`
alone; they pay for themselves in round trips saved.
## File tools
@@ -107,7 +155,14 @@ replaceAll replace every occurrence instead of requiring exactly one
`oldString` must match byte-for-byte and appear exactly once unless `replaceAll` is set.
An ambiguous match is an error naming the count, which pushes the model to add surrounding
context rather than guessing which occurrence it meant.
context rather than guessing which occurrence it meant:
```
oldString appears 3 times in src/users.ts. Add surrounding context or set replaceAll.
```
That error is deliberately specific. `edit failed` would leave the model to retry blind; the
count tells it what to do next.
### `multi_edit`
@@ -121,7 +176,14 @@ the previous one, so edits may build on each other.
Atomic: every edit is validated and applied in memory first, so a failure on the third edit
leaves the file exactly as it was rather than half-changed. The same uniqueness rule as
`edit_file` applies per edit, and the error names which edit failed.
`edit_file` applies per edit, and the error names which edit failed:
```
edit 2: oldString not found in src/users.ts. No edits were applied.
```
The last sentence matters. Without it a model reading the error has to guess whether edit 1
landed, and its next move — retry the whole batch, or only what failed — depends on the answer.
### `list_dir`
@@ -131,9 +193,19 @@ depth levels to descend, 1-6, default 2
includeIgnored also show files git ignores
```
Tree view honouring `.gitignore`. Directories end with `/`, files show their size. Past the
depth limit the containing directory is still listed, so the shape of the tree stays visible
without its contents. Capped at 300 entries.
Tree view honouring `.gitignore`. Directories end with `/`, files show their size:
```
.
README.md 2B
src/
app.ts 2K
ui/
```
Past the depth limit the containing directory is still listed, so the shape of the tree stays
visible without its contents — `src/ui/` above appears at `depth: 2` even though its files do
not. Capped at 300 entries.
### `glob`
@@ -145,9 +217,12 @@ includeIgnored also return files git ignores
Walks the tree honouring `.gitignore` and `.shiroignore`, skipping `.git` and
`node_modules` unconditionally. Nested ignore files apply only within their own directory,
as git does. Returns posix paths relative to the workspace root. A symlinked directory is
classified as a directory and not descended into, since it can point anywhere including
back into the tree.
as git does. Returns posix paths relative to the workspace root.
A symlinked directory is classified as a directory and not descended into. Both halves matter:
`readdir` reports a junction as a non-directory, so without the extra `stat` a symlinked
directory leaked past `dir/` ignore rules and was yielded as a file with a nonsense size. Not
descending is separate — a link can point anywhere, including back into the tree.
### `grep`
@@ -162,6 +237,15 @@ Shells out to ripgrep when it is on PATH — roughly 15x faster on a real repo
back to a JavaScript walker otherwise. Output is `path:line: text` either way, so the model
sees one format regardless. Skips binaries. Caps at 200 hits.
Two details keep the two paths in agreement. ripgrep is passed `--no-require-git`, because it
otherwise ignores `.gitignore` outside a repository while the JavaScript fallback always honours
it. And an rg exit code above 1 means rg could not run the search at all, so the fallback takes
over; exit 1 is simply "no matches" and is reported as such.
Regex syntax differs between the two: ripgrep is Rust regex, the fallback is JavaScript. A
pattern using look-around works in the fallback and fails under rg. An invalid pattern is
reported as `Invalid regex: <reason>` rather than returning an empty result set.
### `bash`
```
@@ -194,12 +278,25 @@ running quits as usual.
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
array rather than a shell string, so an argument like `--author="; rm -rf /"` can only ever
be a literal argument — which is what makes auto-approval safe.
be a literal argument — which is what makes auto-approval safe. A test asserts exactly that:
`git_log` with the path `; touch pwned.txt` creates no file.
Output is described rather than raw porcelain: `git_status` names the branch and says
`staged modified` or `untracked` per file instead of leaving the model to decode two columns
of flags. Outside a repository they fail with `<cwd> is not a git repository.` rather than
passing git's own error text through.
Output is described rather than raw porcelain. `git_status` names the branch and says
`staged modified` or `untracked` per file instead of leaving the model to decode porcelain's two
leading columns:
```
On main, 2 changed:
src/app.ts (staged modified, modified)
new.ts (untracked)
```
That file has a staged change *and* a later unstaged one, which the raw `MM` prefix conveys only
to a reader who knows the format.
Outside a repository they fail with `<cwd> is not a git repository.` rather than passing git's
own error text through. Other git failures do pass through, on purpose: `git_show no-such-ref`
reports what git said, because git's own message is the most useful thing available.
```
git_status branch, staged, modified, untracked
@@ -209,6 +306,13 @@ git_show ref path? one commit: message, author, diff
git_blame path startLine? endLine? who last changed each line
```
`git_log` defaults to 15 commits and caps at 40. `git_blame` without a range blames the whole
file; with `startLine` and no `endLine` it covers 40 lines from there.
Everything here is also reachable through `bash`. The reason the set exists anyway is the
approval boundary: `bash git diff` stops for a decision on every call, while `git_diff` cannot
mutate anything and so never needs one.
## Agent tools
### `task`
@@ -223,8 +327,17 @@ Spawns a read-only subagent with `read_file`, `glob`, and `grep` only. It return
report, so the parent pays for findings rather than the whole search transcript. It sees
none of the parent conversation, so its prompt has to stand alone.
Two properties follow from that tool set rather than from policy: it can never trigger an
approval prompt, because it has no gated tools; and the parent's context holds the conclusion
instead of the search. A subagent reading forty files to answer one question costs the parent
the answer, not the forty files.
`explore` finds and reports. `review` critiques code in severity order. Progress streams to
the subagent panel.
the subagent panel. Capped at 20 steps, and it shares the parent's model — an `explore` run
pays reasoning rates for what is really a search, which is [on the list](../TODO.md) to fix.
Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a
search spanning many files and loses on anything you could answer in one call.
### `ask`
@@ -235,9 +348,11 @@ multiple allow more than one
```
Stops the turn and puts the question on screen. With options it is a picker; without, free
text. `esc` skips, which tells the model to decide and state its assumption.
text. `esc` skips, which returns "the user dismissed this; use your best judgement" — so a
dismissal is an instruction to decide, not a dead end.
Withheld entirely in headless mode — a question with no one to answer it would hang.
Withheld entirely in headless mode — a question with no one to answer it would hang. The system
prompt says so, and tells the model to decide and state its assumption instead.
### `todo_write`
@@ -249,6 +364,9 @@ Statuses: `pending`, `in_progress`, `done`, `blocked`. Send the whole list each
replaces the previous one. Warns when more than one task is `in_progress`, when nothing is
`in_progress` while work remains, or when a `blocked` task has no note.
The list lives in the system prompt, rebuilt every step, so it survives both pruning and
`/compact`. See [memory](memory.md).
### `remember`, `recall`, `forget`
Durable per-project notes. See [memory](memory.md).
@@ -264,15 +382,46 @@ Loads the body of a skill. See [skills](skills.md).
## Path safety
Every path a tool receives goes through a jail: resolved against the workspace root, then
checked that it did not escape. `../../etc/passwd` and absolute paths outside the root are
both refused before any filesystem call.
checked that it did not escape.
```ts
export function jail(p: string, root = process.cwd()): string {
const abs = isAbsolute(p) ? resolve(p) : resolve(root, p);
const rel = relative(resolve(root), abs);
if (rel.startsWith('..') || isAbsolute(rel)) throw new Error(`Path escapes workspace: ${p}`);
return abs;
}
```
`../../etc/passwd`, `a/../../secret`, and absolute paths outside the root are all refused
before any filesystem call. Resolving first and comparing after is what catches the middle
case: string-prefix checks on the raw input miss `a/../../secret` entirely.
The model's output is a trust boundary. It can emit any string, so the check happens on
every call rather than being assumed.
every call rather than being assumed. That includes `read_many_files`, where a bad path is
reported in its block like any other unreadable file.
`jail` guards the workspace, not the shell. `bash` runs whatever it is given, which is why
every call needs approval and why the guard plugin exists — see [plugins](plugins.md).
## Output caps
Any single tool result is truncated at 30,000 characters with a note saying how much was
cut. `grep` stops at 200 hits, `glob` at 200 paths, `list_dir` at 300 entries,
`read_many_files` at 20 files, `read_file` at 2000 lines by default. Without caps one `grep`
for `function` can end a session.
Any single tool result is truncated at 30,000 characters with a note saying how much was cut:
```
... [truncated 41,233 chars]
```
| Tool | Cap |
|---|---|
| any result | 30,000 characters |
| `grep` | 200 hits, each line cut at 300 chars |
| `glob` | 200 paths |
| `list_dir` | 300 entries |
| `read_many_files` | 20 files |
| `read_file` | 2,000 lines by default |
| `bash` | 120 s default timeout, 600 s max |
Without caps one `grep` for `function` can end a session. The caps are per call, so a model
that needs more can narrow and ask again — which is cheaper than one call that fills the
context and forces compaction.