The compaction bug, which is the important one:
beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.
The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.
So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.
Registry, via `/registry [list|search|add|remove|installed]`:
Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.
Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.
Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.
538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
181 lines
7.6 KiB
Markdown
181 lines
7.6 KiB
Markdown
# Development
|
|
|
|
## Setup
|
|
|
|
```bash
|
|
git clone https://github.com/zakirkun/shiro-neko
|
|
cd shiro-neko
|
|
bun install
|
|
```
|
|
|
|
Bun 1.3.14 or newer. Nothing else is required, though `rg` on PATH makes `grep` about 15x
|
|
faster and the fallback path is exercised without it.
|
|
|
|
## Commands
|
|
|
|
```bash
|
|
bun run shiro # run from source
|
|
bun run typecheck # tsc --noEmit
|
|
bun test # 538 tests
|
|
bun run build # single binary for this platform -> dist/shiro
|
|
bun run release # all five platforms -> dist/release + SHA256SUMS
|
|
bun run install:local # build, then copy onto PATH
|
|
```
|
|
|
|
`bun run install:local` copies the compiled binary. Do not use `bun link`: it writes a shim
|
|
that re-execs `bun`, which fails on any machine where bun was installed without `bun.exe` on
|
|
PATH — an npm install of bun, for instance. The compiled binary embeds its own runtime.
|
|
|
|
`SHIRO_INSTALL_DIR` overrides the target directory.
|
|
|
|
## Testing
|
|
|
|
No mocking framework. Everything is driven through a real boundary.
|
|
|
|
```ts
|
|
// The loop: a mock provider, asserting what crossed the wire.
|
|
const seen: LanguageModelV4CallOptions[] = [];
|
|
const model = new MockLanguageModelV4({
|
|
doStream: async (o) => { seen.push(o); return stream(text('ok')); },
|
|
});
|
|
const session = new Session({ model, askApproval: async () => 'deny', agent: variantByName('plan') });
|
|
for await (const _ of session.send('investigate')) void _;
|
|
|
|
const offered = (seen[0]?.tools ?? []).map((t) => t.name);
|
|
expect(offered).not.toContain('write_file');
|
|
```
|
|
|
|
```ts
|
|
// The UI: real keystrokes, asserting what is on screen.
|
|
const app = render(<App session={session} bridge={bridge} hooks={testHooks()} header="hdr" />);
|
|
app.stdin.write('/');
|
|
await wait(120);
|
|
expect(app.lastFrame()).toContain('/compact');
|
|
```
|
|
|
|
Provider wire formats are tested against a local `Bun.serve`. MCP is tested against a real
|
|
stdio subprocess. Tools are tested in a temp directory with `process.chdir`.
|
|
|
|
`SHIRO_HOME` points config, sessions, memory, and history at a temp directory, so a test run
|
|
never touches your real state.
|
|
|
|
### What to assert
|
|
|
|
Assert on what crossed a boundary: the request body, the rendered frame, the file on disk.
|
|
Not on internal calls.
|
|
|
|
That is not style. Several real bugs were caught this way and would have passed a
|
|
mock-verification test:
|
|
|
|
- `pruneMessages` leaving a message item without its reasoning item — visible only in the
|
|
request body
|
|
- `pruneMessages` leaving a tool result without its tool call — same, and it took a stub
|
|
endpoint that rejected the pairing to prove the fix
|
|
- Compaction blanking the model's memory of its own tool calls — invisible in any single
|
|
request, and visible only as "the loop ran to its step limit". Caught by asserting the loop
|
|
terminated because the model chose to, not that the messages had a particular shape
|
|
- `--json` serialising `Error` as `{}` — visible only in the printed output
|
|
- Automatic approval requests prompting the user — visible only in the event sequence
|
|
- `ctrl-c` killing `cmd /c` but not the command under it — visible only as elapsed time, since
|
|
the interrupt reported success while the command ran for another 19 seconds
|
|
|
|
## Adding a tool
|
|
|
|
1. Define it in `src/tools.ts` with a `zod` schema. Descriptions are read by the model, so
|
|
write them as guidance, not as documentation.
|
|
2. Add it to the `tools` object.
|
|
3. Add it to a set in `TOOL_SETS`. A tool in no set can never be gated off.
|
|
4. If it mutates anything, add it to `MUTATING_TOOLS` so it requires approval.
|
|
5. Add a line to `TOOL_DOCS` in `src/prompt.ts` saying *when* to reach for it.
|
|
6. If it is read-only, add it to `READ_ONLY` in `src/agents.ts` so `plan` and `review` can use
|
|
it.
|
|
7. Test the behaviour in a temp directory, including the failure path.
|
|
|
|
Steps 3 and 4 are two hand-maintained lists of tool names, which is a known weakness: a tool
|
|
added to one and forgotten in the other is a silently ungated write. Deriving both from the
|
|
tool definitions is on [TODO.md](../TODO.md).
|
|
|
|
Every tool costs roughly 550 characters of schema on every request. Fourteen built-in tools is
|
|
well past where selection accuracy starts to matter, which is why sets exist and why a new tool
|
|
needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
|
|
|
|
## Adding a slash command
|
|
|
|
`src/commands.ts` is the single source of truth. Add a `CommandSpec` to `COMMANDS`, a case to
|
|
`parseCommand`, and a case in `App.tsx`. The menu, `/help`, and the parser all read from that
|
|
one array, and a test asserts every entry parses and appears in help — they cannot drift.
|
|
|
|
## Code conventions
|
|
|
|
Sample a neighbouring file before inventing a pattern. Broadly:
|
|
|
|
- No comment that restates the code. Comments explain *why*, and usually only where something
|
|
non-obvious was forced by an external constraint.
|
|
- No `as any`, no `@ts-ignore`. `tsconfig.json` runs strict with
|
|
`noUncheckedIndexedAccess`.
|
|
- Validate at trust boundaries — model output, file contents, network responses. Not between
|
|
internal functions.
|
|
- Duplication over premature abstraction. No interface with one implementation.
|
|
- Errors carry what the reader needs to act. `oldString appears 3 times in src/x.ts` beats
|
|
`edit failed`.
|
|
|
|
## Releasing
|
|
|
|
The version lives in `src/version.ts`, compiled into the binary. `package.json` carries it too
|
|
for tooling, and `bun run release` refuses to build if the two disagree, or if a git tag
|
|
disagrees with either:
|
|
|
|
```
|
|
$ GITHUB_REF_NAME=v9.9.9 bun run release
|
|
tag v9.9.9 does not match src/version.ts (0.1.0-beta.3). Bump the version or retag.
|
|
```
|
|
|
|
A binary reporting the wrong version is worse than a failed release.
|
|
|
|
To cut one:
|
|
|
|
```bash
|
|
# bump src/version.ts and package.json to the same value
|
|
git commit -am "release 0.1.0-beta.4"
|
|
git tag v0.1.0-beta.4
|
|
git push --follow-tags
|
|
```
|
|
|
|
Cross-compilation is the part that only breaks in CI. `bun run release` on Windows
|
|
takes a different branch from the Ubuntu runner — `--windows-title` is accepted on a
|
|
Windows host and rejected everywhere else — so a green local release is not proof.
|
|
`buildArgs()` is unit-tested for both hosts because of exactly that.
|
|
|
|
`.github/workflows/release.yml` then runs typecheck and tests, cross-compiles all five
|
|
targets on one Ubuntu runner, asserts the built binary reports the expected version, and
|
|
publishes a GitHub release with the binaries and `SHA256SUMS`. A tag containing `-` is
|
|
published as a prerelease.
|
|
|
|
Bun cross-compiles from any host, which is why there is no build matrix. Verified: a working
|
|
`darwin-arm64` binary builds on Windows.
|
|
|
|
Publishing is gated on a `v*` tag, so a manual `workflow_dispatch` run produces artifacts
|
|
without releasing.
|
|
|
|
## CI
|
|
|
|
`.github/workflows/ci.yml` runs typecheck, tests, and a build on Ubuntu, macOS, and Windows
|
|
for every push and PR.
|
|
|
|
All three are necessary. The tools shell out to `rg`, `git`, and a platform shell, and path
|
|
handling differs — a Windows-only break is invisible on Linux until someone hits it.
|
|
|
|
## Debugging the agent itself
|
|
|
|
`--no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp` strips it to
|
|
the built-in tools alone, which isolates whether a problem is the loop or something layered on
|
|
it. `{ "toolSets": [] }` narrows it further, to the six core tools.
|
|
|
|
`--json` in headless mode shows the exact event sequence.
|
|
|
|
For provider issues, a local `Bun.serve` that logs the request body and returns a canned SSE
|
|
stream answers "what did we actually send" faster than any amount of reading. Several bugs in
|
|
this codebase were found that way. Making that stub *reject* the thing you think you fixed is
|
|
better still: the tool-pairing repair was confirmed by a stub that returned the real 400 for an
|
|
orphaned result, then stopped doing so.
|