The compaction bug, which is the important one:
beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.
The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.
So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.
Registry, via `/registry [list|search|add|remove|installed]`:
Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.
Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.
Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.
538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
4.4 KiB
Skills
A skill is a markdown file with instructions for one kind of task. Only its name and description sit in the system prompt; the body is loaded on demand.
That split matters. Four bundled skills are 4,659 characters of body but 681 characters of catalogue. Putting every body in the prompt would cost that on every request, for instructions that are relevant to one turn in twenty.
Format
---
name: deploy
description: Ship a release. Use when asked to deploy, cut a release, or publish a build.
---
# Deploy
1. Confirm the tests pass. Do not deploy on a red suite.
2. Tag with the version from `src/version.ts`, not by hand.
3. Push the tag. CI builds and publishes.
Never deploy from a dirty working tree.
name and description are both required; a file missing either is skipped. The
description is what the model matches against, so write it as a trigger — "use when asked
to X" — not as a summary.
Where they load from
Four sources, later overriding earlier by name:
- builtin — compiled into the binary
- registry —
~/.shiro-neko/registry/skills/*.md, installed with/registry add - user —
~/.shiro-neko/skills/*.md - project —
.shiro/skills/*.md
A project skill named debug replaces the bundled one entirely. /skills shows what
loaded and where each came from — which matters most for registry, since that body came
from someone else and is now in your system prompt. See registry.
--no-skills skips all of them, builtin included.
The bundled skills
debug — reproduce first, form three hypotheses, disprove them cheapest-first, fix the
cause not the symptom, write a test that failed before. After two failed attempts: re-read
the error literally and check whether the code you think is running is the code that is
running.
review — severity order: incorrect behaviour, missing validation at trust boundaries,
security, resource handling, then clarity. Say plainly when something is fine. Do not invent
findings to look thorough.
refactor — establish a safety net first, move in small steps with tests green between
each, do not fix bugs while refactoring, do not add abstraction for a single caller.
test — read two existing test files first and match them, assert on behaviour not
implementation, never weaken an assertion to make a test pass, a flaky test is a shared-state
problem and not something to retry around.
They are string constants in src/skills-builtin.ts rather than files, because
bun build --compile only embeds modules reachable through imports. A directory of .md
files would be missing from the shipped binary.
How the agent uses one
The catalogue appears in the system prompt:
Skills available through the skill tool. Load one when its description matches the task,
before you start working, and follow it as if the user had written it:
- debug: Track down a bug whose cause is not obvious. Use when a test fails for unclear...
- refactor: Restructure code without changing behaviour. Use when asked to refactor...
When the model calls skill({ name: "debug" }) it gets the full body back and is told to
follow it for this task. The call needs no approval — it reads nothing outside the binary.
Writing a good one
Skills work when they encode what a newcomer to your project would get wrong. The bundled ones are generic on purpose; yours should not be.
Useful:
---
name: migration
description: Write or run a database migration. Use when the schema changes.
---
Migrations live in `db/migrations/` and are timestamped, never renumbered.
Run `bun run db:migrate` locally first. Staging runs them automatically on deploy;
production needs `bun run db:migrate --env=prod` by hand, after the deploy is green.
Never edit a migration that has run anywhere. Write a new one.
Not useful:
---
name: quality
description: Write good code.
---
Follow best practices. Write clean, maintainable code with good naming.
The second costs tokens and changes nothing.
Skill or AGENTS.md?
AGENTS.md is always in the prompt. A skill is loaded when its description matches.
Put standing facts in AGENTS.md: build commands, layout, conventions that apply to every
change. Put task-specific procedure in a skill: how to deploy, how to add a migration, how
this project debugs its worker queue.
If it applies to every turn, it belongs in AGENTS.md. If it applies to one kind of turn,
make it a skill.