Files
shiro-neko/docs/skills.md
T
Muhammad Zakir Ramadhan 9b978fdbe1 Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
2026-09-03 09:26:37 +07:00

152 lines
5.8 KiB
Markdown

# Skills
A skill is a markdown file with instructions for one kind of task. Only its name and
description sit in the system prompt; the body is loaded on demand.
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
would cost that on every turn, for instructions relevant to one turn in twenty.
## Format
```markdown
---
name: deploy
description: Ship a release. Use when asked to deploy, cut a release, or publish a build.
---
# Deploy
1. Confirm the tests pass. Do not deploy on a red suite.
2. Tag with the version from `src/version.ts`, not by hand.
3. Push the tag. CI builds and publishes.
Never deploy from a dirty working tree.
```
`name` and `description` are both required; a file missing either is skipped. The
description is what the model matches against, so write it as a trigger — "use when asked
to X" — not as a summary.
The frontmatter reader handles those two fields and nothing else. A real YAML parser would be a
dependency for two strings, so lists, nesting, and multi-line values are not supported: keep both
on one line. Quotes around a value are stripped. A body over 20,000 characters is truncated.
A file that fails to parse is skipped silently rather than reported, which is worth knowing when
a skill you wrote does not appear in `/skills` — the usual cause is a missing `---` fence or a
description spilling onto a second line.
## Where they load from
Four sources, later overriding earlier by name:
1. **builtin** — compiled into the binary
2. **registry** — `~/.shiro-neko/registry/skills/*.md`, installed with `/registry add`
3. **user** — `~/.shiro-neko/skills/*.md`
4. **project** — `.shiro/skills/*.md`
A project skill named `debug` replaces the bundled one entirely. `/skills` shows what
loaded and where each came from — which matters most for `registry`, since that body came
from someone else and is now in your system prompt. See [registry](registry.md).
`--no-skills` skips all of them, builtin included.
## The bundled skills
**`debug`** — reproduce first, form three hypotheses, disprove them cheapest-first, fix the
cause not the symptom, write a test that failed before. After two failed attempts: re-read
the error literally and check whether the code you think is running is the code that is
running.
**`review`** — severity order: incorrect behaviour, missing validation at trust boundaries,
security, resource handling, then clarity. Say plainly when something is fine. Do not invent
findings to look thorough.
**`refactor`** — establish a safety net first, move in small steps with tests green between
each, do not fix bugs while refactoring, do not add abstraction for a single caller.
**`test`** — read two existing test files first and match them, assert on behaviour not
implementation, never weaken an assertion to make a test pass, a flaky test is a shared-state
problem and not something to retry around.
They are string constants in `src/skills-builtin.ts` rather than files, because
`bun build --compile` only embeds modules reachable through imports. A directory of `.md`
files would be missing from the shipped binary.
## How the agent uses one
The catalogue appears in the system prompt:
```
Skills available through the skill tool. Load one when its description matches the task,
before you start working, and follow it as if the user had written it:
- debug: Track down a bug whose cause is not obvious. Use when a test fails for unclear...
- refactor: Restructure code without changing behaviour. Use when asked to refactor...
```
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
follow it for this task. The call needs no approval — it reads nothing outside the binary.
"Before you start working" is the load-bearing phrase. A skill loaded after the work is done is
wasted tokens, and the failure mode in practice is a model that reads the catalogue, decides it
already knows, and never calls the tool. A description written as a trigger is what prevents that.
Loading one costs its body, once, in that turn's context. A 3,000-character skill is cheaper than
one wrong approach it prevents, and more expensive than the catalogue line that would have been
enough.
## Verifying a skill loaded
```bash
shiro -p "fix the failing pagination test" --json --yolo | grep skill
```
`--json` shows the `tool-call` for `skill` with the name it chose, or its absence. If the model
never calls it on a task the skill was written for, the description is the thing to change — not
the body.
## Writing a good one
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled
ones are generic on purpose; yours should not be.
Useful:
```markdown
---
name: migration
description: Write or run a database migration. Use when the schema changes.
---
Migrations live in `db/migrations/` and are timestamped, never renumbered.
Run `bun run db:migrate` locally first. Staging runs them automatically on deploy;
production needs `bun run db:migrate --env=prod` by hand, after the deploy is green.
Never edit a migration that has run anywhere. Write a new one.
```
Not useful:
```markdown
---
name: quality
description: Write good code.
---
Follow best practices. Write clean, maintainable code with good naming.
```
The second costs tokens and changes nothing.
## Skill or AGENTS.md?
`AGENTS.md` is always in the prompt. A skill is loaded when its description matches.
Put standing facts in `AGENTS.md`: build commands, layout, conventions that apply to every
change. Put task-specific procedure in a skill: how to deploy, how to add a migration, how
this project debugs its worker queue.
If it applies to every turn, it belongs in `AGENTS.md`. If it applies to one kind of turn,
make it a skill.