release 1.0.0: cost control, 41 tools, 29 skills, custom commands, auto-load
This commit is contained in:
@@ -6,7 +6,7 @@ on:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
dry_run:
|
||||
description: Build the artifacts without publishing a release
|
||||
description: Build and verify the artifacts without publishing a release
|
||||
type: boolean
|
||||
default: true
|
||||
|
||||
@@ -53,7 +53,11 @@ jobs:
|
||||
|
||||
publish:
|
||||
needs: build
|
||||
if: startsWith(github.ref, 'refs/tags/v')
|
||||
# Tag pushes always publish. A manual dispatch publishes only when dry_run is
|
||||
# unchecked (false); the default true builds and verifies without a release.
|
||||
if: >-
|
||||
startsWith(github.ref, 'refs/tags/v') ||
|
||||
(github.event_name == 'workflow_dispatch' && inputs.dry_run == false)
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
# shiro-neko
|
||||
|
||||
Agentic coding CLI, built with Bun + TypeScript. The interactive UI is React rendered to the
|
||||
terminal with Ink; LLM access goes through the Vercel AI SDK (`ai`) with Anthropic, OpenAI,
|
||||
OpenAI-compatible, and MCP providers. Entry point and only executable is `src/cli.tsx` (bin `shiro`).
|
||||
|
||||
## Commands
|
||||
|
||||
- `bun install --frozen-lockfile` — install deps (CI uses this; lockfile is `bun.lock`)
|
||||
- `bun run shiro` — run the CLI from source (i.e. `bun run src/cli.tsx`)
|
||||
- `bun test` — full test suite (`bun:test`, no other runner)
|
||||
- `bun run typecheck` — `tsc --noEmit`; must pass before committing
|
||||
- `bun run build` — `bun build --compile` to `dist/shiro` (single native binary)
|
||||
- `bun run release` — cross-compile all five targets into `dist/release/`
|
||||
- No linter or formatter is configured; don't invent one.
|
||||
|
||||
CI (`.github/workflows/ci.yml`) runs install → typecheck → test → build on Ubuntu, macOS, and
|
||||
Windows, pinned to Bun 1.3.14. Everything must be cross-platform: the tools shell out to the
|
||||
platform shell, and paths in code and tests go through `node:path`, never hardcoded `/`.
|
||||
|
||||
## Layout
|
||||
|
||||
- `src/` — flat modules, one concern per file, lowercase names (`session.ts`, `prune.ts`).
|
||||
`src/ui/` holds the Ink components, PascalCase (`App.tsx`, `Panels.tsx`).
|
||||
- `test/` — one `<name>.test.ts` per `src/<name>.ts`; `*.test.tsx` for UI tests via
|
||||
`ink-testing-library`. `test/helpers.ts` has `testHooks()`, the standard App fixture.
|
||||
- `docs/` — user-facing docs, one per feature area.
|
||||
- `scripts/` — `release.ts`, `install.ts` (+ `.sh`/`.ps1` installers).
|
||||
- `src/version.ts` — hardcoded VERSION; the release workflow fails if it disagrees with the git tag.
|
||||
|
||||
## Conventions
|
||||
|
||||
- ES modules, `verbatimModuleSyntax` on: import types with `import type`. Strict TS with
|
||||
`noUncheckedIndexedAccess` — indexing gives `T | undefined`, so handle it (`arr[i]!` appears
|
||||
where provably safe).
|
||||
- Uses Bun APIs directly (`Bun.file`, `Bun.write`, `Bun.spawn`, `Bun.Glob`) — no fs-extra, no
|
||||
node shims. File tools read/write through `Bun.*`, not `fs`, where practical.
|
||||
- Path safety: every user/model-supplied path goes through `jail()` (in `src/ignore.ts`), which
|
||||
rejects escapes outside `process.cwd()`. Tools resolve paths against `process.cwd()`.
|
||||
- Error handling: tool `execute` functions return error text to the model or throw `Error` with a
|
||||
plain message — no error classes, no codes. Storage reads (`store.ts`, `memory.ts`) catch and
|
||||
degrade to empty rather than throw on corrupt JSON.
|
||||
- Tools are AI-SDK `tool()` objects with zod `inputSchema`. Any tool that mutates the workspace
|
||||
must be added to `MUTATING_TOOLS` in `src/tools.ts` — a test in `permission.test.ts` fails
|
||||
otherwise. The permission system (`src/permission.ts`) matches rules against the tool's subject
|
||||
(command for `bash`, path for file tools) with glob matching, and is pure/no-IO on purpose.
|
||||
- Comments explain *why*, often naming the failure being guarded against. Match that style.
|
||||
- Tests import from `bun:test`, build mock models with `MockLanguageModelV4` +
|
||||
`simulateReadableStream` from `ai/test`, and each test that touches the filesystem does
|
||||
`process.chdir()` into a fresh `mkdtemp` dir in `beforeEach` and restores in `afterEach`.
|
||||
|
||||
## Surprising / easy to break
|
||||
|
||||
- `bun test` runs the whole suite including `fallback-live.test.ts`, which spins up real local
|
||||
HTTP servers via `Bun.serve`, and `commit.test.ts`, which runs real `git` in temp repos — both
|
||||
need a working network stack and git on PATH.
|
||||
- Tool output is capped (`MAX_OUTPUT` in `src/tools.ts`); read_file returns NUL-sniffed binary
|
||||
files as an error. Don't remove these — they stop a model from burning its context.
|
||||
- Ink renders to the terminal, so library warnings are suppressed at the top of `cli.tsx`
|
||||
(`AI_SDK_LOG_WARNINGS = false`); anything written to stderr tears the UI.
|
||||
- Sessions, memory, and history live under `~/.shiro-neko/`, relocatable with `SHIRO_HOME`;
|
||||
tests depend on that env var to isolate state. Don't resolve the path eagerly at module load —
|
||||
`store.ts` resolves it per call for this reason.
|
||||
- Release tags must match `src/version.ts` exactly or `release.ts` stops the build.
|
||||
@@ -0,0 +1,201 @@
|
||||
# Audit
|
||||
|
||||
Checklist from a full audit of the codebase, run against `main` at `a22d8e1` ("release 0.1.0-beta.5").
|
||||
The greps cover every file under `src/`, `test/`, `docs/`, `.github/workflows/`, and `scripts/`.
|
||||
|
||||
Nothing here is a fix — it is a list. Items already tracked in `TODO.md` or `ROADMAP.md` say so;
|
||||
untracked items are marked **not yet tracked**.
|
||||
|
||||
---
|
||||
|
||||
## A. Clean findings (verified, no action needed)
|
||||
|
||||
- [x] **No TODO/FIXME/HACK markers in `src/`.** All 26 matches are false positives:
|
||||
placeholder attributes, the `TODO_MARK` export in `src/notebook.ts` (a literal string
|
||||
ingredient of the todos feature), and "later" in prose.
|
||||
- [x] **No `as any` / `@ts-ignore` / `@ts-expect-error` / `@ts-nocheck` / `: any` in `src/`.**
|
||||
Zero matches. The project's own `docs/development.md` rule is being kept.
|
||||
- [x] **No silently swallowed errors in `src/`.** Every `catch` was reviewed:
|
||||
- `src/registry.ts:116,155` — JSON parse failures become descriptive errors.
|
||||
- `src/registry.ts:120-123,159-162` — zod schema validation with named failure reasons.
|
||||
- `src/plugins.ts:59-63` — a throwing plugin hook fails **closed** (blocks the call).
|
||||
- `src/subagent.ts:233-237` — errors are reported and rethrown (never swallowed).
|
||||
- `src/tools.ts:80-82` — a batch read failure is reported in place, not thrown.
|
||||
- `src/headless.ts` serialization flattens `Error` before `JSON.stringify` (would emit `{}`).
|
||||
- `src/ui/App.tsx` — all 14 catch blocks surface the message in the UI.
|
||||
- `src/plugins-builtin.ts:213-215` — the only quiet `catch`, and it is deliberate,
|
||||
commented ("a missing binary is not worth interrupting the turn over").
|
||||
- [x] **No skipped tests.** No `.skip`, `xit`, or `xdescribe` in `test/` (one match was
|
||||
`process.exit(` containing "xit(").
|
||||
- [x] **Registry fetches are size-capped and schema-validated.** `src/registry.ts` caps
|
||||
content-length and body bytes (`fetchText`, lines 100-108), validates the index with
|
||||
`indexSchema` (line 120), and regex-validates every plugin manifest pattern (line 167).
|
||||
- [x] **Version/tag consistency is enforced twice.** `scripts/release.ts` refuses a build when
|
||||
the tag and `src/version.ts` disagree (line 129), and the release workflow asserts the
|
||||
built binary prints the expected version (`.github/workflows/release.yml:42-46`).
|
||||
- [x] **Install scripts match the build targets.** `test/ci.test.ts:85-101` iterates every
|
||||
`TARGETS` entry from `scripts/release.ts` and asserts the shell/PowerShell installers
|
||||
fetch exactly those asset names.
|
||||
- [x] **`.env`/`.pem` are refused on read;** the default permission table
|
||||
(`src/permission.ts:175-186`) matches the approvals banner in `README.md`. Unknown tools
|
||||
(MCP, plugins) default to `ask` rather than allow (line 235-236), so `mcp__*` needs no
|
||||
explicit rule.
|
||||
- [x] **`--yolo` cannot bypass the guard plugin.** Defaults fold `ask` into `allow` but never
|
||||
touch `deny` (`src/permission.ts:211-213`), and the guard refuses destructive commands
|
||||
in `beforeToolCall`, ahead of any approval.
|
||||
- [x] **Pinned toolchain.** Both workflows pin `bun-version: 1.3.14`, and `test/ci.test.ts:48-52`
|
||||
fails if the pin ever disagrees with the local `Bun.version`.
|
||||
- [x] **All three platforms in CI.** `ci.yml` runs the suite on ubuntu, macos, windows
|
||||
(required — the tools shell out to `rg`, git, and a platform shell).
|
||||
|
||||
---
|
||||
|
||||
## B. Bugs
|
||||
|
||||
- [ ] **`src/ui/App.tsx:613` — formatting glitch.** The `}` closing the `try` is jammed onto the
|
||||
same line as the preceding statement:
|
||||
`push({ kind: 'info', text: await hooks.summarizeMemory() }); } catch (e) {`
|
||||
Cosmetic only, but it is the kind of blemish left by an unformatted edit and reads as a
|
||||
slip. **not yet tracked**
|
||||
- [ ] **`.github/workflows/release.yml` — the `dry_run` input is dead.** `workflow_dispatch`
|
||||
declares `inputs.dry_run` (default `true`) but no step ever reads it. Nothing consults the
|
||||
value, so `dry_run=false` changes nothing, and because the `publish` job gates on
|
||||
`startsWith(github.ref, 'refs/tags/v')`, a manual run can never publish regardless of the
|
||||
input. Either wire the input into the `publish` `if`, or delete it and let the tag-only
|
||||
gate be the whole story. **not yet tracked**
|
||||
- [ ] **`README.md` says "Nineteen built-in tools" — it is now twenty.** `git_commit_message`
|
||||
(shipped in beta.5 via `src/commit.ts` + `cli.tsx:281`) is a built-in tool, and
|
||||
`TOOL_SETS.git` carries it (`src/tools-git.ts:189`). The count is one short;
|
||||
`docs/tools.md` already says "twenty" (line 63), so README is the stale one.
|
||||
**not yet tracked**
|
||||
|
||||
---
|
||||
|
||||
## C. Gaps / not implemented (official — tracked in TODO.md or ROADMAP.md)
|
||||
|
||||
### TODO.md "Now" — next up
|
||||
|
||||
- [ ] **Summarize the pruned span.** Compaction drops messages and tells the model nothing, so a
|
||||
decision from earlier in the session can be contradicted. (TODO.md `## Now`, first item)
|
||||
- [ ] **A spend ceiling.** `maxSpendUsd` in config, warn at 80%, refuse the next turn at 100%,
|
||||
headless exits non-zero naming the ceiling. Nothing stops a looping headless run today.
|
||||
(TODO.md `## Now`)
|
||||
- [ ] **A cheaper model for subagents.** `subagentModel` in config; an `explore` subagent is
|
||||
search, not reasoning, and today pays the parent's per-token rate. (TODO.md `## Now`)
|
||||
- [ ] **Hot-reload an installed entry.** `/registry add` writes the file and says restart; the
|
||||
skill catalogue and guard chain are assembled at boot. (TODO.md `## Now`)
|
||||
|
||||
### TODO.md "Next"
|
||||
|
||||
- [ ] **MCP without the schema tax.** Twenty MCP tools ≈ 2,750 tokens of schema per request;
|
||||
`toolSets` does not gate them. Plan: `mcp_list` / `mcp_inspect` / `mcp_call` meta-tools,
|
||||
prompt names servers not schemas. (TODO.md `## Next`; ROADMAP `## Next` + `## Later`)
|
||||
- [ ] **Custom commands from a file.** `.shiro/commands/*.md`, `$ARGUMENTS`, `$1`,
|
||||
`` !`cmd` `` shell substitution with the guard applied. (TODO.md `## Next`; ROADMAP `## Next`)
|
||||
- [ ] **Derive the tool-name lists.** `TOOL_SETS` and `MUTATING_TOOLS` are hand-maintained; a
|
||||
tool added to one and forgotten in the other is a silently ungated write. (TODO.md
|
||||
`## Next`; ROADMAP `## Next` "Derived tool metadata")
|
||||
- [ ] **Subagent parallelism.** Two independent searches run sequentially; the panel already
|
||||
renders several agents, the loop does not fan out. (TODO.md `## Next`; ROADMAP `## Later`)
|
||||
- [ ] **Undo a turn.** `/resume` restores a session but nothing walks one step back; `bash`
|
||||
effects cannot be snapshotted and the docs would say so. (TODO.md `## Next`; ROADMAP `## Next`)
|
||||
|
||||
### ROADMAP "Next" / "Later" — tracked, not yet scheduled in TODO.cpp-equivalent detail
|
||||
|
||||
- [ ] **Registry trust.** No signatures; `registryUrl` is the whole trust decision. Publisher
|
||||
keys + pinned digest per entry. (ROADMAP `## Next`; also TODO.md Known rough edges)
|
||||
- [ ] **Lossless-enough compaction** — same work as "Summarize the pruned span". (ROADMAP `## Next`)
|
||||
- [ ] **Session branching**, **structured diff review**, **plugin code from disk** (needs a
|
||||
sandbox story), **prompt caching** (stable prefix vs volatile suffix), **external hooks**
|
||||
(needs a trust story), **OS-level sandboxing** (Seatbelt/Landlock/Windows equivalent).
|
||||
(ROADMAP `## Later`)
|
||||
- [ ] **Deliberately declined** (do not "fix"): web UI, auto-commit, vector search, tool-call
|
||||
retries, client/server split, LSP integration — all recorded in ROADMAP `## Declined`.
|
||||
|
||||
---
|
||||
|
||||
## D. Documentation drift (untracked)
|
||||
|
||||
- [ ] **README tool count** — see Bug B.3. **not yet tracked**
|
||||
- [ ] **TODO.md "Done" is missing the rest of beta.5.** "Kept for one release, then deleted",
|
||||
but of the beta.5 batch (more tools incl. `git_branch`/`git_commit_message`/
|
||||
`move_file`/`delete_file`, the `protect` plugin, `security`/`perf`/`migrate` skills, the
|
||||
MCP panel wizard, the UI refinements, the farewell message) only "a dead provider item
|
||||
ends the turn" was checked off. Either the Done list gets the beta.5 items or it gets
|
||||
rotated, as the file's own rule says. **not yet tracked**
|
||||
- [ ] **`docs/registry.md` and `docs/headless.md`** are referenced by the README table and both
|
||||
exist — verified clean, no action.
|
||||
|
||||
---
|
||||
|
||||
## E. Test-coverage gaps
|
||||
|
||||
- [ ] **`src/cli.tsx` (562 lines) has no unit test.** Nothing in `test/` imports it. Its flag
|
||||
parsing (`-p`, `--json`, `--yolo`, `--resume`, provider setup, `/provider` wiring) is
|
||||
exercised only by hand or through `runHeadless` (`test/headless.test.ts`), which bypasses
|
||||
the argument surface. The largest module in `src/` outside the UI is the least tested one.
|
||||
**not yet tracked**
|
||||
- [ ] **`src/ui/Onboard.tsx` has no test.** The provider on-boarding wizard is never rendered in
|
||||
the suite. **not yet tracked**
|
||||
- [ ] **`src/ui/PromptInput.tsx` has no test** — the `@` completion input is only covered
|
||||
indirectly through `App`. (`src/complete.ts` itself is well tested.) **not yet tracked**
|
||||
- [ ] **`src/ui/panel-bodies.ts` and `src/ui/buses.ts` have no direct tests.**
|
||||
**not yet tracked**
|
||||
- [ ] **29 `as any` casts across 17 test files.** The identical mock `usage` object
|
||||
(`{ inputTokens: {...}, outputTokens: {...} } as any`) is copied verbatim in 7+ UI test
|
||||
files — a shared typed fixture in `test/helpers.ts` would remove the repetition and the
|
||||
casts in one move. Production `src/` remains clean; this is test-only debt. **not yet tracked**
|
||||
|
||||
---
|
||||
|
||||
## F. Maintenance debt
|
||||
|
||||
- [ ] **Pricing table is hand-entered with no source note or date** — `src/pricing.ts:8-22`.
|
||||
Rates drift; `estimateTokens` also divides JSON length by four (session.ts:90), which is
|
||||
fine as a compaction threshold but misleads in `/cost`. Both tracked in TODO.md
|
||||
`## Maintenance`.
|
||||
- [ ] **`listPaths` walks up to 5000 files once per session** — fine for a repo, wasteful in a
|
||||
monorepo, never notices a file created after the first `@`. Tracked in TODO.md.
|
||||
- [ ] **`MUTATING_TOOLS` (tools.ts:799-807) is only used by tests and docs.** The runtime gate
|
||||
is `DEFAULT_PERMISSIONS` + the unknown-tool `ask` default. Tracked in TODO.md — either
|
||||
delete it or make `DEFAULT_PERMISSIONS` derive from it.
|
||||
- [ ] **CI actions are about to leave Node 20.** The last release run annotated that actions on
|
||||
Node 20 are being forced onto Node 24. `actions/checkout@v4`, `setup-bun@v2`,
|
||||
`upload-artifact@v4`, `download-artifact@v4` still work, but the major-version bumps will
|
||||
become the silent fix; watch for the annotation to turn red. **not yet tracked**
|
||||
- [ ] **Oversized modules.** `src/ui/App.tsx` (863), `src/tools.ts` (717), `src/cli.tsx` (562),
|
||||
`src/session.ts` (500), `src/ui/Panels.tsx` (398). All of them grew past a comfortable
|
||||
review size during the beta.5 batch. Not a bug — a "who reads 863 lines" concern.
|
||||
**not yet tracked**
|
||||
|
||||
---
|
||||
|
||||
## G. Known rough edges (tracked in TODO.md, reproduced here for the record)
|
||||
|
||||
- [x] `/clear` wipes terminal scrollback (escape sequence takes earlier history with it).
|
||||
- [x] Memory has no conflict resolution; two contradictory notes both inject.
|
||||
- [x] Windows `cmd /c` vs `bash -lc` — the prompt names the platform, does not translate.
|
||||
- [x] An unknown name in `toolSets` is dropped silently — reads as "that set is off".
|
||||
- [x] Permission rules gate the call, not what it does — no OS sandbox around the shell.
|
||||
- [x] The reasoning panel is per-turn, not per-step.
|
||||
- [x] An interrupted command's effects are unknown, and the model is told so.
|
||||
- [x] `@` completion lists files, not directories.
|
||||
- [x] An installed skill is a stranger's words in the system prompt; nothing re-checks it later.
|
||||
- [x] A registry index is trusted for its contents, not its authorship.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Area | Items |
|
||||
|---|---|
|
||||
| Bugs | 3 (`App.tsx:613`, dead `dry_run` input, README tool count) |
|
||||
| Official gaps (Now/Next/Later) | 14 tracked in TODO.md/ROADMAP.md |
|
||||
| Documentation drift | 2 untracked |
|
||||
| Test-coverage gaps | 5 (of which `src/cli.tsx` is the significant one) |
|
||||
| Maintenance debt | 5 (3 tracked, 2 untracked) |
|
||||
| Known rough edges | 10 (all tracked) |
|
||||
| Verified clean | 9 areas, including zero `as any` and zero swallowed errors in `src/` |
|
||||
|
||||
The codebase is in good shape for a beta. The three bugs are each one-line fixes; the
|
||||
coverage gap on `cli.tsx` is the item that will actually bite.
|
||||
@@ -0,0 +1,89 @@
|
||||
# Changelog
|
||||
|
||||
All notable changes to this project are documented here. The format follows
|
||||
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and the project adheres to
|
||||
[Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
## [1.0.0]
|
||||
|
||||
The first stable release. Cost control, a larger tool and skill surface, custom slash
|
||||
commands, auto-loaded extensions, and a redesigned welcome interface, on top of the beta
|
||||
line's agent loop, approval model, and safety guarantees.
|
||||
|
||||
### Added
|
||||
|
||||
- **Spend ceiling** ("maxSpendUsd" in config). Checked before each turn: past the limit the
|
||||
model is never called, the turn is refused naming the ceiling, headless exits non-zero,
|
||||
and it warns once at 80% of the limit. Unpriced models cannot be measured, so the ceiling
|
||||
does not apply to them.
|
||||
- **Cheaper subagent model** ("subagentModel" in config). "explore" subagents — search, not
|
||||
reasoning — resolve against a configured cheaper model while "review" and "worker" keep the
|
||||
parent's. "/cost" reports subagent spend separately, priced against the subagent's model id.
|
||||
- **Twenty new built-in tools** (41 total) in a new "extra" tool set, across four families:
|
||||
- line edits: insert_lines, delete_lines, replace_lines, append_file, prepend_file, count_lines
|
||||
- filesystem: tree, file_info, find_files, recent_files, changed_files
|
||||
- git (read-only, argv-spawned): git_log_file, git_diff_commits, git_show_file, git_current_branch, git_changed_in_ref
|
||||
- code and environment: find_symbol, json_query, outline, read_symbol, env_info, count_tokens
|
||||
- **Twenty new bundled skills** (29 total), including plan, docs, api-design, ci-cd, db,
|
||||
docker, frontend, git-workflow, logging, optimize-sql, release, accessibility, data, i18n,
|
||||
deps, onboarding, ux-copy, readme, incident, perf-frontend.
|
||||
- **Ten new plugins**, all data. Safety refusals on by default — no-force-push, no-net-pipe,
|
||||
no-root, no-env-write — and opt-in workflow plugins — no-main-commit, no-git-config,
|
||||
confirm-delete, conventional-commit, tests-first, small-diffs.
|
||||
- **Custom slash commands** from Markdown files in ".shiro/commands/" and
|
||||
"~/.shiro-neko/commands/", with frontmatter "description"/"agent", "$ARGUMENTS" and
|
||||
positional "$1", and shell substitution passed through the guard. A custom command can
|
||||
never shadow a built-in.
|
||||
- **Auto-loaded external extensions** from "~/.shiro-neko/{skills,tools,plugins}" and
|
||||
".shiro/{skills,tools,plugins}". All data, never code: tools are bounded manifests (a shell
|
||||
template through the guard, an HTTPS fetch, or a workspace read), plugins are refusal
|
||||
manifests. Malformed files are reported and skipped, never fatal.
|
||||
|
||||
### Changed
|
||||
|
||||
- **Skills now live as Markdown files** in "src/skills-md/", embedded into the compiled binary
|
||||
by Bun text imports, replacing the previous TypeScript string constants. A format test
|
||||
enforces that each parses with valid frontmatter and a real body. The eleven pre-existing
|
||||
skills were also deepened.
|
||||
- **Welcome interface redesigned** into a structured dashboard: a session banner, a grouped
|
||||
environment panel with attention-worthy facts coloured out of the quiet layer, and a meta
|
||||
bar. The prompt input sits in a two-tone box with the agent and model row inside it and a
|
||||
split footer beneath.
|
||||
- **System prompt advanced** with a discrete failure-recovery loop, a delegation policy, and
|
||||
compaction awareness.
|
||||
|
||||
### Fixed
|
||||
|
||||
- **Release workflow "dry_run" input is now honoured.** A manual dispatch publishes only when
|
||||
it is unchecked; tag pushes always publish. Previously the input was declared but never read,
|
||||
so a manual run could never publish regardless of its value.
|
||||
- **Documentation drift** corrected across the tool count, the bundled-skill count, the plugin
|
||||
defaults, and the README quickstart.
|
||||
|
||||
## [0.1.0-beta.5]
|
||||
|
||||
- A dead provider item no longer ends the turn: a 404 naming a missing item rewrites the
|
||||
history inline and retries once.
|
||||
- More tools (git_commit_message, git_branch, move_file, delete_file), the "protect" plugin,
|
||||
the security/perf/migrate skills, the "/mcp add" wizard, and the farewell message.
|
||||
|
||||
## [0.1.0-beta.4]
|
||||
|
||||
- Compaction no longer stops the loop, and is bounded to keep the widest recent tool tail.
|
||||
- The external registry for skills and plugins.
|
||||
- Permission rules matched per command and path, replacing the per-tool list.
|
||||
- apply_patch, web_fetch, and writable worker subagents.
|
||||
|
||||
## [0.1.0-beta.3]
|
||||
|
||||
- Fourteen built-in tools with gateable tool sets.
|
||||
- Streaming reasoning display, the mid-turn prompt queue, multi_edit, list_dir, read-only git
|
||||
tools, batch reads, @file completion, and interruptible commands.
|
||||
|
||||
## [0.1.0-beta.1]
|
||||
|
||||
- The core agent loop with SDK-enforced tool approvals, endpoint fallback, and retries.
|
||||
- The initial tool set, Ink interface, agent variants, skills, plugins, per-project memory,
|
||||
session persistence, subagents, MCP, and five-platform builds.
|
||||
|
||||
[1.0.0]: https://github.com/zakirkun/shiro-neko/releases/tag/v1.0.0
|
||||
@@ -46,11 +46,11 @@ from the models that endpoint actually reports. Settings land in
|
||||
`~/.shiro-neko/config.json`. Run `/provider` any time to change them.
|
||||
|
||||
```
|
||||
shiro-neko 0.1.0-beta.5 openai/gpt-5 session 0193ab2c
|
||||
shiro-neko 1.0.0 openai/gpt-5 session 0193ab2c
|
||||
agent: default thinking: medium
|
||||
cwd: /home/you/project
|
||||
skills: commit, debug, migrate, perf, refactor, review, security, test, verify
|
||||
plugins: guard, secrets, protect, time
|
||||
skills: commit, debug, docs, migrate, perf, plan, refactor, review, security, test, verify
|
||||
plugins: guard, secrets, protect, time, no-force-push, no-net-pipe, no-root, no-env-write
|
||||
approvals: ask for write_file, edit_file, multi_edit, apply_patch, move_file, delete_file, bash, web_fetch, mcp__*
|
||||
/help for commands
|
||||
|
||||
@@ -113,7 +113,7 @@ record of what it already ran instead of repeating it.
|
||||
|
||||
**Runs headless.** `shiro -p "review this diff" --json` for scripts and CI.
|
||||
|
||||
**Keeps the tool list affordable.** Nineteen built-in tools, grouped into sets. Each costs
|
||||
**Keeps the tool list affordable.** Forty-one built-in tools, grouped into sets. Each costs
|
||||
about 550 characters of schema on every request, so `{ "toolSets": [] }` trims back to the six
|
||||
core ones and a disabled set reaches neither the wire nor the prompt.
|
||||
|
||||
@@ -131,6 +131,8 @@ the flags are.
|
||||
| [Skills](docs/skills.md) | the bundled skills, writing your own, why the catalogue is split |
|
||||
| [Plugins](docs/plugins.md) | the interface, the guard and its limits, builtin versus installed |
|
||||
| [Registry](docs/registry.md) | installing external skills and plugins, publishing your own |
|
||||
| [Custom commands](docs/custom-commands.md) | a Markdown file becomes a slash command, with arguments and shell substitution |
|
||||
| [Extensions](docs/extensions.md) | auto-loaded external skills, tools, and plugins — data, never code |
|
||||
| [Memory and state](docs/memory.md) | memory, task lists, sessions, compaction and its repair |
|
||||
| [MCP](docs/mcp.md) | connecting servers, namespacing, cost, debugging one |
|
||||
| [Headless mode](docs/headless.md) | `-p`, JSON events, exit codes, CI recipes |
|
||||
@@ -138,6 +140,7 @@ the flags are.
|
||||
| [Development](docs/development.md) | building, testing, adding a tool, releasing |
|
||||
| [Roadmap](ROADMAP.md) | what is next and what has been declined |
|
||||
| [TODO](TODO.md) | the current work list, with known rough edges |
|
||||
| [Changelog](CHANGELOG.md) | release history, newest first |
|
||||
|
||||
## Commands
|
||||
|
||||
@@ -156,15 +159,18 @@ workspace path. Up and down recall earlier prompts.
|
||||
|
||||
## Status
|
||||
|
||||
Working: the agent loop, tool approvals, subagents including the gated `worker` kind, skills,
|
||||
plugins, per-project memory, session persistence, MCP, markdown rendering, headless mode,
|
||||
five-platform builds, streaming reasoning display, the mid-turn prompt queue, gateable tool
|
||||
sets, read-only git tools, batch reads, `apply_patch`, `web_fetch`, `@file` completion,
|
||||
interruptible commands, and the external registry.
|
||||
Version 1.0 is stable. Working: the agent loop, per-call and per-command tool approvals with a
|
||||
guard that `--yolo` cannot bypass, subagents including the gated `worker` kind, a spend ceiling
|
||||
(`maxSpendUsd`) with a cheaper subagent model (`subagentModel`), 41 built-in tools across
|
||||
gateable sets, 29 bundled skills, built-in and data-only plugins, per-project memory, session
|
||||
persistence and resume, MCP servers, custom slash commands from markdown files, auto-loaded
|
||||
external skills/tools/plugins, markdown rendering, headless mode with JSON events for CI,
|
||||
five-platform builds, streaming reasoning, the mid-turn prompt queue, read-only git tools,
|
||||
batch reads, `apply_patch`, `web_fetch`, `@file` completion, interruptible commands, and the
|
||||
external registry.
|
||||
|
||||
Next up is in [TODO.md](TODO.md); the longer view and what has been declined are in
|
||||
[ROADMAP.md](ROADMAP.md). The short version of what is missing: a summary of what compaction
|
||||
discarded, a spend ceiling, and a cheaper model for subagent searches.
|
||||
[ROADMAP.md](ROADMAP.md); the release history is in [CHANGELOG.md](CHANGELOG.md).
|
||||
|
||||
## License
|
||||
|
||||
|
||||
+47
-12
@@ -189,6 +189,53 @@ each server's live tool count or its connection error, and `/mcp remove` takes o
|
||||
three write `config.json` directly; a new server connects on the next start, because
|
||||
connecting mid-turn would change the tool list under a running request.
|
||||
|
||||
### 1.0.0
|
||||
|
||||
The first stable release. The beta line's architecture held; this release rounds out cost
|
||||
control, extensibility, and the interface, and hardens the test suite to match.
|
||||
|
||||
**Cost control.** Two halves of one problem, both shipped. A **spend ceiling** (`maxSpendUsd`)
|
||||
checks before each turn: past the limit the model is never called, the turn is refused naming
|
||||
the ceiling, headless exits non-zero, and it warns once at 80%. A **cheaper subagent model**
|
||||
(`subagentModel`) runs `explore` — which is search, not reasoning — on a less expensive model
|
||||
while `review` and `worker` keep the parent's; `/cost` reports subagent spend as its own line,
|
||||
priced against the subagent's model id.
|
||||
|
||||
**Tools: 41 built-in.** Twenty new tools in four families, all path-jailed and ignore-aware,
|
||||
in a new `extra` tool set: precise line edits (`insert_lines`, `delete_lines`, `replace_lines`,
|
||||
`append_file`, `prepend_file`, `count_lines`), filesystem navigation (`tree`, `file_info`,
|
||||
`find_files`, `recent_files`, `changed_files`), read-only git extensions (`git_log_file`,
|
||||
`git_diff_commits`, `git_show_file`, `git_current_branch`, `git_changed_in_ref`), and code and
|
||||
environment reads (`find_symbol`, `json_query`, `outline`, `read_symbol`, `env_info`,
|
||||
`count_tokens`). Every git call still spawns the binary with a fixed argv, never a shell.
|
||||
|
||||
**Skills: 29 bundled, as Markdown.** The catalogue grew from nine to twenty-nine and every
|
||||
skill moved to a single source of truth: a Markdown file in `src/skills-md/`, frontmatter and
|
||||
body, embedded into the compiled binary by Bun text imports. A format test enforces that each
|
||||
one parses and carries a real body.
|
||||
|
||||
**Plugins: 10 more, all data.** Six narrow safety refusals (force push, pipe-to-shell, root
|
||||
elevation, env credential writes, main-branch commits, git config changes) and three advisory
|
||||
plugins (conventional commits, tests-first, small diffs) plus a delete guard for ambiguous
|
||||
paths. The safety refusals are on by default for the same reason the guard is; the opinionated
|
||||
ones are opt-in.
|
||||
|
||||
**Custom slash commands.** A Markdown file in `.shiro/commands/` or `~/.shiro-neko/commands/`
|
||||
becomes a slash command, with frontmatter `description`/`agent`, `$ARGUMENTS` and `$1`
|
||||
positionals, and `` !`cmd` `` substitution passed through the guard. A custom command can never
|
||||
shadow a built-in.
|
||||
|
||||
**Auto-loaded extensions.** External skills, tools, and plugins load from
|
||||
`~/.shiro-neko/<kind>/` and `.shiro/<kind>/` — all data, never code. An external tool is a
|
||||
bounded manifest (a shell template through the guard, an HTTPS fetch, or a workspace file
|
||||
read); an external plugin is a refusal manifest. A malformed file is reported and skipped,
|
||||
never fatal.
|
||||
|
||||
**Interface.** The welcome screen is a structured dashboard — a session banner, a grouped
|
||||
environment panel with attention-worthy facts lifted out of the quiet layer, and a meta bar —
|
||||
replacing a wall of dim text. The input sits in an OpenCode-style two-tone box with the
|
||||
agent·model row inside it and a split footer beneath.
|
||||
|
||||
---
|
||||
|
||||
## Next
|
||||
@@ -201,12 +248,6 @@ answer is three meta-tools — `mcp_list`, `mcp_inspect`, `mcp_call` — with th
|
||||
the servers, so a hundred servers cost almost nothing until one is called. Worth keeping direct
|
||||
registration as an option: for a two-tool server the indirection is the more expensive of the two.
|
||||
|
||||
### Custom commands from a file
|
||||
|
||||
A markdown file becoming a slash command, with `$ARGUMENTS`, `$1`, `` !`cmd` `` for shell output,
|
||||
and `@path` for a file. Every comparable CLI has this and none of it is hard; it is missing because
|
||||
nothing forced the issue.
|
||||
|
||||
### Undo a turn
|
||||
|
||||
opencode has `/undo` and `/redo`, Claude Code has `/rewind` over file checkpoints. There is
|
||||
@@ -220,12 +261,6 @@ Compaction keeps the model's memory of a turn now, but it still says nothing abo
|
||||
discarded, so the model can contradict its own earlier decision with confidence. A summary of the
|
||||
discarded span costs one cheap call and removes the whole class of problem.
|
||||
|
||||
### Cost control
|
||||
|
||||
Two halves of the same problem: an `explore` subagent pays the parent's reasoning rate for
|
||||
what is really a search, and nothing stops a headless run that loops. A cheaper subagent model
|
||||
and a per-session ceiling are both small changes on top of the pricing that already exists.
|
||||
|
||||
### Derived tool metadata
|
||||
|
||||
`TOOL_SETS` and `MUTATING_TOOLS` are hand-maintained lists of tool names. A tool added to one
|
||||
|
||||
@@ -19,24 +19,6 @@ confidence.
|
||||
- [ ] Budget it: a summary that grows with the session defeats the point
|
||||
- [ ] Test: a pruned decision is still recoverable from the summary
|
||||
|
||||
### A spend ceiling
|
||||
|
||||
A headless run that loops costs real money with nothing to stop it.
|
||||
|
||||
- [ ] `maxSpendUsd` in config, checked after every turn
|
||||
- [ ] Warn at 80%, refuse to start another turn at 100%
|
||||
- [ ] Headless exits non-zero with the ceiling named, rather than stopping silently
|
||||
- [ ] Test: a session past its ceiling refuses the next turn and says why
|
||||
|
||||
### A cheaper model for subagents
|
||||
|
||||
The subagent shares the parent's model. An `explore` run is search, not reasoning, and it
|
||||
currently pays the parent's per-token rate.
|
||||
|
||||
- [ ] `subagentModel` in config, defaulting to the parent
|
||||
- [ ] `/cost` separates parent from subagent spend
|
||||
- [ ] Test: the subagent's calls go to the configured model, the parent's do not
|
||||
|
||||
### Hot-reload an installed entry
|
||||
|
||||
`/registry add` writes the file and says to restart. The skill catalogue and the guard chain
|
||||
@@ -64,17 +46,6 @@ names only the servers. A hundred servers then cost almost nothing until one is
|
||||
- [ ] Keep per-tool registration as an option: a two-tool server is cheaper registered directly
|
||||
- [ ] Test: a configured server contributes no schema to the request until `mcp_call`
|
||||
|
||||
### Custom commands from a file
|
||||
|
||||
Every other CLI in this class has these and they are cheap: a markdown file becomes a slash
|
||||
command, with `$ARGUMENTS`, `$1`, `` !`cmd` `` for shell output, and `@path` for a file.
|
||||
|
||||
- [ ] `.shiro/commands/*.md` and `~/.shiro-neko/commands/*.md`, name from the filename
|
||||
- [ ] Frontmatter for `description` and `agent`
|
||||
- [ ] `$ARGUMENTS` and positional `$1`
|
||||
- [ ] `` !`cmd` `` substituted before the prompt is sent, with the guard applied to it
|
||||
- [ ] Test: a command with a shell substitution reaches the model with the output inlined
|
||||
|
||||
### Derive the tool-name lists
|
||||
|
||||
`TOOL_SETS` and `MUTATING_TOOLS` both list names by hand. A tool added to one and forgotten
|
||||
@@ -151,37 +122,25 @@ Not bugs exactly, but things that will bite someone.
|
||||
|
||||
## Done
|
||||
|
||||
Kept for one release, then deleted.
|
||||
Kept for one release, then deleted. The 1.0.0 release batch:
|
||||
|
||||
- [x] Reasoning streamed to a collapsed panel, `ctrl-r` to expand, dropped when the turn ends
|
||||
- [x] The tool in flight named on screen from `tool-input-start` until its result arrives
|
||||
- [x] Prompts typed during a turn queue and drain in order; `esc` clears the queue
|
||||
- [x] `toolSets` gating, so a disabled set reaches neither the wire nor the prompt
|
||||
- [x] `multi_edit`, atomic across several edits to one file
|
||||
- [x] `list_dir`, ignore-aware and depth-limited
|
||||
- [x] Read-only git tools: `git_status` `git_diff` `git_log` `git_show` `git_blame`
|
||||
- [x] Orphaned tool results dropped during pruning, fixing the 400 "No tool call found for
|
||||
function call output with call_id ..."
|
||||
- [x] `read_many_files`, concurrent, one labelled block per file, a bad path reported in place
|
||||
- [x] `@file` completion: picker fed by the ignore-aware walker, tab inserts a relative path
|
||||
- [x] `ctrl-c` kills the running command and keeps the turn. The kill takes the whole process
|
||||
tree: killing `cmd /c` alone left the real command holding both pipes open, so the
|
||||
interrupt appeared to do nothing for 19 seconds
|
||||
- [x] **Compaction no longer stops the loop.** Pruning used to drop any assistant part whose
|
||||
reasoning item it removed, which on a reasoning model is every tool call. The model lost
|
||||
its record of what it had run and re-ran it until the step limit. The repair strips the
|
||||
provider `itemId` instead of the part, so the same content is sent inline
|
||||
- [x] **A dead provider item no longer ends the turn.** An `item_reference` resolves only while
|
||||
the provider still stores that item, so a resumed session could fail on every attempt with
|
||||
404 "Item with id 'msg_...' not found". Compaction now sends the history inline, and a 404
|
||||
naming a missing item rewrites the history inline and retries once
|
||||
- [x] `/registry`: browse, search, install, and remove external skills and plugins. Skills are
|
||||
shown in full before install; plugins are a validated manifest of deny rules, never code
|
||||
- [x] Context shown as a percentage of the compaction threshold, amber at two thirds, red at 90
|
||||
- [x] **Permission rules per command and path**, replacing the per-tool list. `bash` was one
|
||||
yes/no for `git status` and `rm -rf`, so pressing `a` once removed the gate for both.
|
||||
Rules match the call's subject, `always` grants a pattern rather than the tool, `.env` and
|
||||
`.pem` are refused on read, and an identical call repeated three times in a turn asks even
|
||||
when allowed
|
||||
- [x] **`web_fetch`**, size-capped HTTP(S) to markdown in the opt-in `net` tool set, with
|
||||
redirect and private-address checks
|
||||
- [x] A spend ceiling (`maxSpendUsd`): checked before each turn, refused at 100% naming the
|
||||
ceiling, warns once at 80%, headless exits non-zero. Unpriced models are not enforced
|
||||
- [x] A cheaper subagent model (`subagentModel`): `explore` resolves against it, `review` and
|
||||
`worker` keep the parent's, `/cost` splits subagent spend by model id
|
||||
- [x] Twenty new built-in tools (41 total) in a new `extra` set: line edits, filesystem
|
||||
navigation, read-only git extensions, and code/environment reads
|
||||
- [x] Twenty new bundled skills (29 total) plus the eleven originals deepened; all moved to
|
||||
`src/skills-md/*.md` as the Markdown source of truth, embedded at build
|
||||
- [x] Ten new data-only plugins: safety refusals on by default (force push, pipe-to-shell,
|
||||
root, env credential writes) and opt-in workflow plugins (conventional commit,
|
||||
tests-first, small diffs, main-branch commits, git config, confirm-delete)
|
||||
- [x] Custom slash commands from Markdown files, with `$ARGUMENTS`/`$1` and guarded shell
|
||||
substitution; a custom command never shadows a built-in
|
||||
- [x] Auto-loaded external skills, tools, and plugins from `~/.shiro-neko/<kind>` and
|
||||
`.shiro/<kind>`, all data, never code; a bad file is reported and skipped
|
||||
- [x] The welcome interface redesigned into a structured dashboard with a session banner, a
|
||||
grouped environment panel, and a meta bar; the input in a two-tone box with a split footer
|
||||
- [x] The system prompt advanced: a failure-recovery loop, a delegation policy, compaction awareness
|
||||
- [x] The release workflow's dead `dry_run` input wired: manual dispatch publishes only when
|
||||
unchecked, tag pushes always publish
|
||||
|
||||
@@ -14,11 +14,13 @@
|
||||
"ink-select-input": "6.2.0",
|
||||
"ink-spinner": "^5.0.0",
|
||||
"ink-text-input": "6.0.0",
|
||||
"pngjs": "^7.0.0",
|
||||
"react": "19.2.8",
|
||||
"zod": "4.5.4",
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/bun": "^1.4.0",
|
||||
"@types/pngjs": "^6.0.5",
|
||||
"@types/react": "19.2.18",
|
||||
"ink-testing-library": "4.0.0",
|
||||
"react-devtools-core": "^7.0.1",
|
||||
@@ -51,6 +53,8 @@
|
||||
|
||||
"@types/node": ["@types/node@26.4.0", "", { "dependencies": { "undici-types": "~8.3.0" } }, "sha512-faiGnoIrLH/V8cibOMEAZ8pMw6oXqSukl29ra4mN8GdaB2ZewzeaLj+INpV5N+Z1eKWzY+IzaIZH2EIR6YZRNQ=="],
|
||||
|
||||
"@types/pngjs": ["@types/pngjs@6.0.5", "", { "dependencies": { "@types/node": "*" } }, "sha512-0k5eKfrA83JOZPppLtS2C7OUtyNAl2wKNxfyYl9Q5g9lPkgBl/9hNyAu6HuEH2J4XmIv2znEpkDd0SaZVxW6iQ=="],
|
||||
|
||||
"@types/react": ["@types/react@19.2.18", "", { "dependencies": { "csstype": "^3.2.2" } }, "sha512-AnzbBERsrLKtk2XSfTbYRLjQPdy116Sty4q+T+Bp3IC4l6jNBvreVPAHmpq9qhXQM7CXZPjLVmGMw9sy+hxQ3w=="],
|
||||
|
||||
"@vercel/oidc": ["@vercel/oidc@3.2.0", "", {}, "sha512-UycprH3T6n3jH0k44NHMa7pnFHGu/N05MjojYr+Mc6I7obkoLIJujSWwin1pCvdy/eOxrI/l3uDLQsmcrOb4ug=="],
|
||||
@@ -131,6 +135,8 @@
|
||||
|
||||
"pkce-challenge": ["pkce-challenge@5.0.1", "", {}, "sha512-wQ0b/W4Fr01qtpHlqSqspcj3EhBvimsdh0KlHhH8HRZnMsEa0ea2fTULOXOS9ccQr3om+GcGRk4e+isrZWV8qQ=="],
|
||||
|
||||
"pngjs": ["pngjs@7.0.0", "", {}, "sha512-LKWqWJRhstyYo9pGvgor/ivk2w94eSjE3RGVuzLGlr3NmD8bf7RcYGze1mNdEHRP6TRP6rMuDHk5t44hnTRyow=="],
|
||||
|
||||
"react": ["react@19.2.8", "", {}, "sha512-PWaYA1L/q9u2u7xYQi+Y3L3Yfnie7XyLeaJICV1MGD6LprsBxcAqGjYyr0eY3p+QdsA+x/Irkt4Qif8D63+Sbw=="],
|
||||
|
||||
"react-devtools-core": ["react-devtools-core@7.0.1", "", { "dependencies": { "shell-quote": "^1.6.1", "ws": "^7" } }, "sha512-C3yNvRHaizlpiASzy7b9vbnBGLrhvdhl1CbdU6EnZgxPNbai60szdLtl+VL76UNOt5bOoVTOz5rNWZxgGt+Gsw=="],
|
||||
|
||||
+24
-3
@@ -20,8 +20,10 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
"agent": "default",
|
||||
"thinking": "medium",
|
||||
"maxRetries": 3,
|
||||
"maxSpendUsd": 5,
|
||||
"subagentModel": "gpt-5-nano",
|
||||
"plugins": ["guard", "time"],
|
||||
"toolSets": ["edit-plus", "git"],
|
||||
"toolSets": ["edit-plus", "extra", "git"],
|
||||
"permission": {
|
||||
"bash": { "*": "ask", "git *": "allow" }
|
||||
},
|
||||
@@ -42,12 +44,31 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
| `agent` | default variant: `default`, `quick`, `deep`, `plan`, `review` |
|
||||
| `thinking` | default level: `off`, `low`, `medium`, `high`, `max` |
|
||||
| `maxRetries` | retries per model call for transient failures. Default 3 |
|
||||
| `plugins` | which builtin plugins to enable. Omit for `["guard", "secrets", "protect", "time"]` |
|
||||
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`, and `net`. Omit for the defaults; `net` is opt-in. See [tools](tools.md) |
|
||||
| `maxSpendUsd` | session spend ceiling: warn at 80%, refuse the next turn at 100%. Headless exits non-zero naming the ceiling. Only enforced on priced models |
|
||||
| `subagentModel` | model id for `explore` subagents, which search rather than reason. Omit to share the parent's model. `/cost` reports subagent spend separately |
|
||||
| `plugins` | which builtin plugins to enable. Omit for `["guard", "secrets", "protect", "time", "no-force-push", "no-net-pipe", "no-root", "no-env-write"]` |
|
||||
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `nav`, `extra`, `git`, and `net`. Omit for the defaults; `net` is opt-in. See [tools](tools.md) |
|
||||
| `permission` | which calls run, ask, or are refused, matched per command or path. See [permissions](permissions.md) |
|
||||
| `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) |
|
||||
| `mcpServers` | see [MCP](mcp.md) |
|
||||
|
||||
## Directories
|
||||
|
||||
Beyond the config file, these locations are read on every start:
|
||||
|
||||
| Path | Holds |
|
||||
|---|---|
|
||||
| `~/.shiro-neko/config.json` | the config above |
|
||||
| `~/.shiro-neko/skills/*.md` `.shiro/skills/*.md` | auto-loaded skills — [extensions](extensions.md) |
|
||||
| `~/.shiro-neko/tools/*.json` `.shiro/tools/*.json` | auto-loaded tool manifests |
|
||||
| `~/.shiro-neko/plugins/*.json` `.shiro/plugins/*.json` | auto-loaded plugin manifests |
|
||||
| `~/.shiro-neko/commands/*.md` `.shiro/commands/*.md` | custom slash commands — [custom commands](custom-commands.md) |
|
||||
| `~/.shiro-neko/registry/{skills,plugins}/` | entries installed with `/registry` |
|
||||
| `~/.shiro-neko/sessions/` | saved sessions, for `-c` / `-r` |
|
||||
|
||||
`SHIRO_HOME` overrides the home directory for all of these, which is also how the test suite
|
||||
isolates itself.
|
||||
|
||||
## Provider presets
|
||||
|
||||
`/provider` offers these. Each sets `baseURL` and the wire protocol for you.
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
# Custom slash commands
|
||||
|
||||
A Markdown file becomes a slash command. Write the prompt once, run it with `/name` any time.
|
||||
|
||||
Two directories are scanned, the project shadowing the user by name:
|
||||
|
||||
| Origin | Directory |
|
||||
|---|---|
|
||||
| user | `~/.shiro-neko/commands/*.md` |
|
||||
| project | `.shiro/commands/*.md` |
|
||||
|
||||
The filename is the command: `.shiro/commands/review-diff.md` becomes `/review-diff`. Names are
|
||||
letters, digits, dashes, and underscores; anything else is skipped. A custom command can never
|
||||
shadow a built-in — `/cost` always runs the built-in `/cost`.
|
||||
|
||||
## Format
|
||||
|
||||
```markdown
|
||||
---
|
||||
description: Review the staged diff for defects
|
||||
agent: review
|
||||
---
|
||||
|
||||
Review the staged changes. For each finding give file, line, what breaks, and the fix.
|
||||
```
|
||||
|
||||
Frontmatter is optional but useful:
|
||||
|
||||
- **`description`** — the one line shown in the `/` menu. Without it the first body line is used.
|
||||
- **`agent`** — run this command under a specific agent variant (`default`, `quick`, `deep`,
|
||||
`plan`, `review`). The variant is restored afterwards, so one command does not leak its agent
|
||||
into the rest of the session.
|
||||
|
||||
Everything after the frontmatter fence is the prompt. A file with an empty body is skipped, as
|
||||
is one that fails to parse.
|
||||
|
||||
## Arguments
|
||||
|
||||
The body is a template, expanded against whatever you type after the command:
|
||||
|
||||
- `$ARGUMENTS` — the whole argument string.
|
||||
- `$1`, `$2`, … — positional arguments. A missing positional expands to nothing.
|
||||
|
||||
```markdown
|
||||
Compare $1 against $2 and report the differences. Context: $ARGUMENTS
|
||||
```
|
||||
|
||||
`/compare src/a.ts src/b.ts` sends `Compare src/a.ts against src/b.ts … Context: src/a.ts src/b.ts`.
|
||||
|
||||
## Shell substitution
|
||||
|
||||
A `` !`command` `` inline runs the shell command and inlines its output before the prompt is
|
||||
sent:
|
||||
|
||||
```markdown
|
||||
Review this diff:
|
||||
|
||||
!`git diff --staged`
|
||||
```
|
||||
|
||||
Every substitution runs through the **guard** before executing, exactly as a direct `bash` call
|
||||
is — so a custom command cannot smuggle a destructive command past you. A substitution that
|
||||
exits non-zero, or one the guard refuses, fails the command with the reason named.
|
||||
|
||||
## When to write one
|
||||
|
||||
- A prompt you find yourself retyping: a review shape, a release checklist, a project-specific
|
||||
"how we test".
|
||||
- A prompt that should pin an agent: a read-only review command that always runs under `review`.
|
||||
- Project conventions the whole team should share: commit `.shiro/commands/` so everyone gets
|
||||
the same commands.
|
||||
|
||||
For behaviour that must survive across sessions rather than be invoked on demand, use
|
||||
[memory](memory.md). For instructions the agent loads by task rather than by name, use a
|
||||
[skill](skills.md). For extensions that add tools or refusal rules rather than prompts, see
|
||||
[extensions](extensions.md).
|
||||
@@ -0,0 +1,101 @@
|
||||
# Extensions: auto-loaded skills, tools, and plugins
|
||||
|
||||
External extensions load automatically from two directories on every start, the project
|
||||
shadowing the user by name:
|
||||
|
||||
| Origin | Directories |
|
||||
|---|---|
|
||||
| user | `~/.shiro-neko/skills` `~/.shiro-neko/tools` `~/.shiro-neko/plugins` |
|
||||
| project | `.shiro/skills` `.shiro/tools` `.shiro/plugins` |
|
||||
|
||||
Drop a file in and it is live on the next start. No registry, no install command, no restart
|
||||
of anything but the CLI itself.
|
||||
|
||||
**Everything here is data, never code.** That is the same rule the [registry](registry.md)
|
||||
enforces, and it is the whole security model. An external extension can add instructions, a
|
||||
bounded tool, or a refusal rule — it cannot run arbitrary code, so it cannot read every file
|
||||
the agent can read or lie about what it blocks. A malformed file is reported on the welcome
|
||||
dashboard and skipped, never fatal.
|
||||
|
||||
## Skills
|
||||
|
||||
A skill is a Markdown file with frontmatter, exactly like a bundled one:
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: deploy
|
||||
description: Ship a release. Use when asked to deploy or cut a release.
|
||||
---
|
||||
|
||||
# Deploy
|
||||
|
||||
1. Confirm the tests pass. Do not deploy on a red suite.
|
||||
2. Tag with the version from src/version.ts, not by hand.
|
||||
```
|
||||
|
||||
Skills merge by name with the precedence `builtin < registry < user < project`, so your own
|
||||
`debug.md` overrides the bundled `debug`. See [skills](skills.md) for the full format.
|
||||
|
||||
## Tools
|
||||
|
||||
A tool is a JSON manifest describing one bounded operation. Three kinds, each with a ceiling
|
||||
on what it can do:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "recent-changes",
|
||||
"description": "List the ten most recently changed files",
|
||||
"kind": "shell",
|
||||
"command": "git diff --name-only HEAD~10",
|
||||
"autoApprove": true
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Meaning |
|
||||
|---|---|
|
||||
| `name` | The tool name the model calls. Letters, digits, dashes, underscores. |
|
||||
| `description` | What the model reads to decide when to use it. |
|
||||
| `kind` | `shell`, `http`, or `read`. |
|
||||
| `command` | For `shell`: the template to run, with an optional `{arg}` placeholder. |
|
||||
| `url` | For `http`: the URL to fetch, with an optional `{arg}` placeholder. HTTPS only. |
|
||||
| `path` | For `read`: the workspace file to return, with an optional `{arg}` placeholder. |
|
||||
| `autoApprove` | `false` to require approval before running. Default `true`. |
|
||||
|
||||
The model passes a single optional `arg` string, substituted into `{arg}`.
|
||||
|
||||
**The limits are the point.** A `shell` tool runs a fixed template through the **guard** and
|
||||
the platform shell — the same chain a built-in `bash` call goes through, so an installed tool
|
||||
cannot do what the agent itself may not. An `http` tool fetches one HTTPS URL. A `read` tool
|
||||
returns one workspace file, jailed to the workspace. None of them executes code from the
|
||||
manifest.
|
||||
|
||||
## Plugins
|
||||
|
||||
A plugin is a refusal manifest — the same shape the registry installs — a name, an optional
|
||||
prompt appendix, and deny rules matched against tool input:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "no-prod-config",
|
||||
"description": "refuses to edit production config",
|
||||
"appendix": "Production config is changed by hand, never by the agent.",
|
||||
"deny": [
|
||||
{ "tools": ["write_file", "edit_file"], "pathPattern": "config/production", "reason": "production config is hand-edited" },
|
||||
{ "tools": ["bash"], "commandPattern": "kubectl\\s+apply", "reason": "deploys to the cluster" }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
A rule names the tools it covers and either a `pathPattern` (matched against the path a file
|
||||
tool carries) or a `commandPattern` (matched against a `bash` command), both as case-insensitive
|
||||
regexes, plus the `reason` handed to the model when it blocks. Patterns are validated on load;
|
||||
an invalid regex is a reported error, not a crash.
|
||||
|
||||
Refusal plugins compose with the built-in [plugins](plugins.md) — the first block wins.
|
||||
|
||||
## Relationship to the registry
|
||||
|
||||
The [registry](registry.md) fetches the same kinds of files over HTTPS with a confirmation
|
||||
step. Auto-load is for your own and your project's files, which need no confirmation because
|
||||
you wrote them. The two mechanisms share the loaders and the safety model; they differ only in
|
||||
where the file comes from.
|
||||
+44
-1
@@ -19,7 +19,7 @@ the agent can read. That is a sandbox problem, not a loader problem — see
|
||||
## Enabling
|
||||
|
||||
```json
|
||||
{ "plugins": ["guard", "secrets", "protect", "time"] }
|
||||
{ "plugins": ["guard", "secrets", "protect", "time", "no-force-push", "no-net-pipe", "no-root", "no-env-write"] }
|
||||
```
|
||||
|
||||
That is also the default when the field is absent, and it lists **builtin** plugins only.
|
||||
@@ -128,6 +128,49 @@ write normally.
|
||||
Adds `current_time`, returning ISO 8601 plus the local string. Auto-approved; it reads
|
||||
nothing. Useful because models are confidently wrong about the date.
|
||||
|
||||
### The narrow safety refusals (default on)
|
||||
|
||||
Six small plugins, each blocking one irreversible class of mistake. They are on by default for
|
||||
the same reason the guard is: a safety check you have to opt into is not one. Each is data — a
|
||||
name, a pattern list, and a refusal message — matched against the `bash` command string (or, for
|
||||
`confirm-delete`, the path a `delete_file` carries).
|
||||
|
||||
| Plugin | Refuses | Why |
|
||||
|---|---|---|
|
||||
| `no-force-push` | `git push --force`, `--force-with-lease`, `-f`, `+<ref>` | rewrites remote history |
|
||||
| `no-net-pipe` | `curl … \| sh`, `wget … \| node`, `iex (iwr …)` | executes a download unseen |
|
||||
| `no-root` | `sudo …`, elevated `runas` / `Start-Process -Verb RunAs` | nothing the agent does should need root |
|
||||
| `no-env-write` | `export …KEY/TOKEN/SECRET/PASSWORD=…` | writes a credential into the environment |
|
||||
| `no-main-commit` | `git commit`/`git merge` naming `main`/`master` | touches the default branch directly |
|
||||
| `no-git-config` | `git config --global`, identity/runner keys | changes how git identifies or runs |
|
||||
|
||||
Normal commands pass: `git push origin feature`, `npm test`, `export NODE_ENV=production`. The
|
||||
patterns target the irreversible act, not the command family.
|
||||
|
||||
### The advisory plugins (opt in)
|
||||
|
||||
Three plugins carry only a prompt appendix — no blocking hook — so they shape behaviour without
|
||||
ever refusing a call. Enable them in config when you want the nudge:
|
||||
|
||||
| Plugin | Advises |
|
||||
|---|---|
|
||||
| `conventional-commit` | commit subjects as `type(scope): summary`, e.g. `fix(auth): reject expired tokens` |
|
||||
| `tests-first` | for a bug, pin it with a failing test before fixing; watch it fail, then pass |
|
||||
| `small-diffs` | one change does one thing; split a diff that is really two |
|
||||
|
||||
```json
|
||||
{ "plugins": ["guard", "secrets", "protect", "time", "conventional-commit", "small-diffs"] }
|
||||
```
|
||||
|
||||
### `confirm-delete` (opt in)
|
||||
|
||||
Refuses `delete_file` calls whose path is broad or ambiguous — a wildcard, a trailing slash, or
|
||||
an empty path — so a delete is always one explicit file:
|
||||
|
||||
```
|
||||
refusing to delete "src/*" (ambiguous or broad). Delete one explicit file.
|
||||
```
|
||||
|
||||
### `bell` (opt in)
|
||||
|
||||
Writes `\u0007` to stderr when a turn ends. Off by default — a bell after every turn is
|
||||
|
||||
+15
-6
@@ -3,9 +3,9 @@
|
||||
A skill is a markdown file with instructions for one kind of task. Only its name and
|
||||
description sit in the system prompt; the body is loaded on demand.
|
||||
|
||||
That split matters. The nine bundled skills are roughly 14,000 characters of body against about
|
||||
1,500 characters of catalogue — paid on every request. Putting every body in the prompt
|
||||
would cost that on every turn, for instructions relevant to one turn in twenty.
|
||||
That split matters. The twenty-nine bundled skills are tens of thousands of characters of body
|
||||
against a small catalogue of names and descriptions — paid on every request. Putting every body
|
||||
in the prompt would cost that on every turn, for instructions relevant to one turn in twenty.
|
||||
|
||||
## Format
|
||||
|
||||
@@ -88,9 +88,18 @@ thing at a time, and stop at a target stated up front. Report the baseline along
|
||||
CI, Dockerfiles, and docs), apply one shape of change rather than improving as you pass, and
|
||||
never hand-merge a lockfile.
|
||||
|
||||
They are string constants in `src/skills-builtin.ts` rather than files, because
|
||||
`bun build --compile` only embeds modules reachable through imports. A directory of `.md`
|
||||
files would be missing from the shipped binary.
|
||||
**`plan`** — break a non-trivial task into an ordered, verifiable sequence before writing code:
|
||||
order by dependency rather than by file, one step one verifiable outcome, keep it small, and
|
||||
replan when the ground moves.
|
||||
|
||||
**`docs`** — write documentation grounded in the source: verify every claim against the code,
|
||||
answer the reader's actual question, show a working example before describing one, and match
|
||||
the house style.
|
||||
|
||||
They are Markdown files in `src/skills-md/`, one per skill, loaded by `src/skills-builtin.ts`
|
||||
as Bun raw-text imports. The `.md` file is the single source of truth — frontmatter and body
|
||||
in proper Markdown — and Bun inlines every text import into the compiled binary, so the folder
|
||||
ships with `bun build --compile` rather than being left behind on disk.
|
||||
|
||||
## How the agent uses one
|
||||
|
||||
|
||||
+56
-1
@@ -46,7 +46,7 @@ Three more things sit around the rules:
|
||||
## Tool sets
|
||||
|
||||
Each tool costs its name, its description, and its JSON schema on **every request**. The current
|
||||
registry has nineteen built-ins. `/tools` shows the live set; disabling an optional set removes
|
||||
registry has forty-one built-ins. `/tools` shows the live set; disabling an optional set removes
|
||||
its schemas from both the request and the system prompt.
|
||||
|
||||
| Tool | Bytes | Tool | Bytes |
|
||||
@@ -68,6 +68,8 @@ Sets let you switch off what a project does not need:
|
||||
|---|---|---|
|
||||
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B |
|
||||
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` `apply_patch` `move_file` `delete_file` | patch and file ops |
|
||||
| `nav` | `find_symbol` `json_query` | navigation and structured reads |
|
||||
| `extra` | 20 tools: line edits, fs inspect, git extensions, code/env reads | on by default |
|
||||
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` `git_branch` `git_commit_message` | ~2,180 B + message |
|
||||
| `net` | `web_fetch` | opt in |
|
||||
|
||||
@@ -107,6 +109,59 @@ Both extra sets earn their place in most projects, but not all:
|
||||
- **Reading a lot, editing rarely?** Keep `edit-plus` for `list_dir` and `read_many_files`
|
||||
alone; they pay for themselves in round trips saved.
|
||||
|
||||
## The `extra` set
|
||||
|
||||
Twenty tools across four families, on by default. Each follows the same rules as the core
|
||||
tools: writes are jailed to the workspace, reads honour `.gitignore`, and every git call spawns
|
||||
the binary with a fixed argument array, never a shell string.
|
||||
|
||||
### Line edits
|
||||
|
||||
Precise edits by line number, for changes that need no full-file rewrite and no exact-string
|
||||
match. All refuse a path outside the workspace.
|
||||
|
||||
| Tool | Does |
|
||||
|---|---|
|
||||
| `insert_lines` | Insert a block before a 1-based line, pushing the rest down. One past the end appends. |
|
||||
| `delete_lines` | Delete an inclusive line range. Refuses the whole file — that is `delete_file`'s job. |
|
||||
| `replace_lines` | Replace an inclusive line range with new text in one write. |
|
||||
| `append_file` | Add text to the end of a file. |
|
||||
| `prepend_file` | Add text to the top of a file, e.g. a header or import block. |
|
||||
| `count_lines` | Line count for one file, or per file across a glob. A size read before opening something large. |
|
||||
|
||||
### Filesystem
|
||||
|
||||
| Tool | Does |
|
||||
|---|---|
|
||||
| `tree` | Indented directory tree, ignore-aware, directories first. A broad shape faster to scan than `list_dir`. |
|
||||
| `file_info` | Size, line count, modified time, text-or-binary for one file. |
|
||||
| `find_files` | Files whose *name* contains a substring (not a glob), e.g. `auth`. |
|
||||
| `recent_files` | Files modified most recently, newest first. Find what a tool just touched. |
|
||||
| `changed_files` | The working-tree delta git reports (modified, staged, untracked). |
|
||||
|
||||
### Git extensions (read-only)
|
||||
|
||||
Spawned with a fixed argv, so they are auto-approved like the core git tools.
|
||||
|
||||
| Tool | Does |
|
||||
|---|---|
|
||||
| `git_log_file` | Commits that touched one file, newest first, with hash, date, subject. |
|
||||
| `git_diff_commits` | Diff between two refs, optionally limited to one path. |
|
||||
| `git_show_file` | A file's contents at a ref, e.g. `auth.ts` at `HEAD~3`. |
|
||||
| `git_current_branch` | The current branch with its upstream and ahead/behind count. |
|
||||
| `git_changed_in_ref` | Files changed between a ref and the working tree, names only. |
|
||||
|
||||
### Code and environment
|
||||
|
||||
| Tool | Does |
|
||||
|---|---|
|
||||
| `find_symbol` | Where a function, class, or type is *defined* across JS/TS, Python, Go, Rust. Matches declarations, not uses. |
|
||||
| `json_query` | One value from a JSON file by dotted path (`scripts.build`), instead of reading it whole. |
|
||||
| `outline` | Top-level declarations of a source file as a structural map. Read before opening a large file. |
|
||||
| `read_symbol` | The full body of one top-level definition by name. |
|
||||
| `env_info` | Platform, shell, and which runtimes and package managers are installed, before writing a command. |
|
||||
| `count_tokens` | Estimate the token cost of a file or string (~4 chars per token) before sending it to the model. |
|
||||
|
||||
## File tools
|
||||
|
||||
### `read_file`
|
||||
|
||||
+3
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "shiro-neko",
|
||||
"version": "0.1.0-beta.5",
|
||||
"version": "1.0.0",
|
||||
"type": "module",
|
||||
"private": true,
|
||||
"bin": {
|
||||
@@ -16,6 +16,7 @@
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/bun": "^1.4.0",
|
||||
"@types/pngjs": "^6.0.5",
|
||||
"@types/react": "19.2.18",
|
||||
"ink-testing-library": "4.0.0",
|
||||
"react-devtools-core": "^7.0.1"
|
||||
@@ -33,6 +34,7 @@
|
||||
"ink-select-input": "6.2.0",
|
||||
"ink-spinner": "^5.0.0",
|
||||
"ink-text-input": "6.0.0",
|
||||
"pngjs": "^7.0.0",
|
||||
"react": "19.2.8",
|
||||
"zod": "4.5.4"
|
||||
}
|
||||
|
||||
+186
@@ -0,0 +1,186 @@
|
||||
import { tool, type ToolSet } from 'ai';
|
||||
import { homedir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { z } from 'zod';
|
||||
import { jail } from './ignore';
|
||||
import { manifestToPlugin, parseManifest, type PluginManifest } from './registry';
|
||||
import type { Plugin } from './plugins';
|
||||
|
||||
/**
|
||||
* Auto-registration and auto-loading of external skills, tools, and plugins.
|
||||
*
|
||||
* Everything here is *data*, never code — the same rule the registry enforces.
|
||||
* An external tool is a bounded manifest (a shell template through the guard, an
|
||||
* HTTP fetch, or a file read), an external plugin a refusal manifest, an external
|
||||
* skill a markdown body. Loading arbitrary code from disk would let an entry read
|
||||
* every file the agent can read and lie about what it blocks, so it is not offered.
|
||||
*
|
||||
* Directories, later shadowing earlier by name:
|
||||
* ~/.shiro-neko/{tools,plugins,skills} (user)
|
||||
* .shiro/{tools,plugins,skills} (project)
|
||||
* Skills already load through skills.ts; this module adds tools and plugins and
|
||||
* the one place cli turns them all on.
|
||||
*/
|
||||
|
||||
const home = () => process.env['SHIRO_HOME'] ?? homedir();
|
||||
|
||||
export type LoadError = { name: string; message: string };
|
||||
|
||||
function dirs(kind: 'tools' | 'plugins' | 'skills', cwd: string): string[] {
|
||||
return [join(home(), '.shiro-neko', kind), join(cwd, '.shiro', kind)];
|
||||
}
|
||||
|
||||
async function scan(dir: string, ext: string): Promise<string[]> {
|
||||
const files: string[] = [];
|
||||
try {
|
||||
for await (const f of new Bun.Glob(`*.${ext}`).scan({ cwd: dir, onlyFiles: true })) files.push(f);
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
return files.sort();
|
||||
}
|
||||
|
||||
const MAX_PATTERN = 200;
|
||||
const nameSchema = z.string().min(1).max(40).regex(/^[a-z0-9][a-z0-9-_]*$/i);
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// External tools, as bounded manifests.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Three kinds of tool, each with a ceiling on what it can do. None runs arbitrary
|
||||
* code: `shell` interpolates a fixed template and runs it through the guard and
|
||||
* the platform shell, `http` fetches a fixed URL, `read` returns a fixed file's
|
||||
* contents (jailed to the workspace). The input is a single optional `arg` string
|
||||
* substituted into a `{arg}` placeholder, so a manifest cannot take structure it
|
||||
* was not declared for.
|
||||
*/
|
||||
const toolManifestSchema = z.object({
|
||||
name: nameSchema,
|
||||
description: z.string().min(1).max(300),
|
||||
kind: z.enum(['shell', 'http', 'read']),
|
||||
/** The template with an optional `{arg}` placeholder. */
|
||||
command: z.string().max(500).optional(),
|
||||
url: z.string().max(500).optional(),
|
||||
path: z.string().max(300).optional(),
|
||||
/** Set false to require approval before running. Default true (auto-approved). */
|
||||
autoApprove: z.boolean().optional(),
|
||||
});
|
||||
|
||||
export type ToolManifest = z.infer<typeof toolManifestSchema>;
|
||||
|
||||
export function parseToolManifest(source: string): ToolManifest {
|
||||
let raw: unknown;
|
||||
try {
|
||||
raw = JSON.parse(source);
|
||||
} catch {
|
||||
throw new Error('the tool manifest is not valid JSON');
|
||||
}
|
||||
const parsed = toolManifestSchema.safeParse(raw);
|
||||
if (!parsed.success) {
|
||||
throw new Error(`the tool manifest is malformed: ${parsed.error.issues[0]?.message ?? 'unknown reason'}`);
|
||||
}
|
||||
const m = parsed.data;
|
||||
if (m.kind === 'shell' && !m.command) throw new Error(`shell tool "${m.name}" needs a command template`);
|
||||
if (m.kind === 'http' && !m.url) throw new Error(`http tool "${m.name}" needs a url`);
|
||||
if (m.kind === 'read' && !m.path) throw new Error(`read tool "${m.name}" needs a path`);
|
||||
return m;
|
||||
}
|
||||
|
||||
const MAX_TOOL_OUTPUT = 30_000;
|
||||
const cap = (s: string) => (s.length <= MAX_TOOL_OUTPUT ? s : `${s.slice(0, MAX_TOOL_OUTPUT)}\n... [truncated]`);
|
||||
|
||||
/** The guard an external shell tool runs through, supplied by cli so it shares the real chain. */
|
||||
export type ShellGuard = (command: string) => Promise<string | undefined>;
|
||||
|
||||
/**
|
||||
* A manifest as a live tool. The guard is applied to every `shell` invocation, so
|
||||
* an external tool cannot smuggle a destructive command past the user any more
|
||||
* than a built-in bash call can.
|
||||
*/
|
||||
export function manifestToTool(manifest: ToolManifest, guard: ShellGuard) {
|
||||
const inputSchema = z.object({ arg: z.string().optional().describe('optional argument substituted into {arg}') });
|
||||
const substitute = (template: string, arg: string) => template.replaceAll('{arg}', arg);
|
||||
|
||||
return tool({
|
||||
description: `${manifest.description} (external ${manifest.kind} tool)`,
|
||||
inputSchema,
|
||||
execute: async ({ arg = '' }) => {
|
||||
if (manifest.kind === 'read') {
|
||||
const abs = jail(substitute(manifest.path!, arg));
|
||||
const file = Bun.file(abs);
|
||||
if (!(await file.exists())) throw new Error(`no such file: ${manifest.path}`);
|
||||
return cap(await file.text());
|
||||
}
|
||||
|
||||
if (manifest.kind === 'http') {
|
||||
const url = substitute(manifest.url!, arg);
|
||||
if (!/^https:\/\//i.test(url)) throw new Error(`http tools may only fetch https URLs, got: ${url}`);
|
||||
const res = await fetch(url, { redirect: 'follow', signal: AbortSignal.timeout(20_000) });
|
||||
if (!res.ok) throw new Error(`${url} returned ${res.status}`);
|
||||
return cap(await res.text());
|
||||
}
|
||||
|
||||
const command = substitute(manifest.command!, arg);
|
||||
const blocked = await guard(command);
|
||||
if (blocked) throw new Error(`refused: ${blocked}`);
|
||||
const shell = process.platform === 'win32' ? ['cmd', '/c', command] : ['bash', '-lc', command];
|
||||
const proc = Bun.spawn(shell, { stdout: 'pipe', stderr: 'pipe' });
|
||||
const [out, err, code] = await Promise.all([
|
||||
new Response(proc.stdout).text(),
|
||||
new Response(proc.stderr).text(),
|
||||
proc.exited,
|
||||
]);
|
||||
if (code !== 0) throw new Error(`exited ${code}: ${err.trim().slice(0, 300)}`);
|
||||
return cap(out.trim() || '(no output)');
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
export type ExternalTools = { tools: ToolSet; autoApprove: string[]; errors: LoadError[] };
|
||||
|
||||
/** Loads every external tool manifest, project shadowing user by name. Bad files are reported and skipped. */
|
||||
export async function loadExternalTools(cwd: string, guard: ShellGuard): Promise<ExternalTools> {
|
||||
const tools: ToolSet = {};
|
||||
const autoApprove: string[] = [];
|
||||
const errors: LoadError[] = [];
|
||||
|
||||
for (const dir of dirs('tools', cwd)) {
|
||||
for (const file of await scan(dir, 'json')) {
|
||||
const fallback = file.replace(/\.json$/i, '');
|
||||
try {
|
||||
const manifest = parseToolManifest(await Bun.file(join(dir, file)).text());
|
||||
tools[manifest.name] = manifestToTool(manifest, guard);
|
||||
if (manifest.autoApprove !== false) autoApprove.push(manifest.name);
|
||||
} catch (e) {
|
||||
errors.push({ name: fallback, message: e instanceof Error ? e.message : String(e) });
|
||||
}
|
||||
}
|
||||
}
|
||||
return { tools, autoApprove, errors };
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// External plugins, as refusal manifests (same shape the registry installs).
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export type ExternalPlugins = { plugins: Plugin[]; errors: LoadError[] };
|
||||
|
||||
/** Loads refusal-manifest plugins from disk, merging with any already installed via the registry. */
|
||||
export async function loadExternalPlugins(cwd: string): Promise<ExternalPlugins> {
|
||||
const byName = new Map<string, Plugin>();
|
||||
const errors: LoadError[] = [];
|
||||
|
||||
for (const dir of dirs('plugins', cwd)) {
|
||||
for (const file of await scan(dir, 'json')) {
|
||||
const fallback = file.replace(/\.json$/i, '');
|
||||
try {
|
||||
const manifest: PluginManifest = parseManifest(await Bun.file(join(dir, file)).text());
|
||||
byName.set(manifest.name, manifestToPlugin(manifest));
|
||||
} catch (e) {
|
||||
errors.push({ name: fallback, message: e instanceof Error ? e.message : String(e) });
|
||||
}
|
||||
}
|
||||
}
|
||||
return { plugins: [...byName.values()], errors };
|
||||
}
|
||||
+88
-29
@@ -3,6 +3,7 @@ import { render } from 'ink';
|
||||
import React from 'react';
|
||||
import type { LanguageModel, ModelMessage } from 'ai';
|
||||
import { resolveAgent, VARIANTS, isThinkingLevel, type AgentVariant } from './agents';
|
||||
import { loadExternalPlugins, loadExternalTools } from './autoload';
|
||||
import { configPath, loadConfig, missingKeyMessage, resolveModel, writeConfigFile, type Config } from './config';
|
||||
import type { FallbackEvent } from './fallback';
|
||||
import { farewell } from './farewell';
|
||||
@@ -18,12 +19,14 @@ import { createHost } from './plugins';
|
||||
import { fetchModels, presetById } from './providers';
|
||||
import * as registry from './registry';
|
||||
import { Session } from './session';
|
||||
import { loadCustomCommands } from './custom-commands';
|
||||
import { loadSkills } from './skills';
|
||||
import * as store from './store';
|
||||
import { createTaskTool, type SubagentApproval } from './subagent';
|
||||
import { VERSION, versionLine } from './version';
|
||||
import { createAskBridge } from './ui/Ask';
|
||||
import { App, createApprovalBridge, createNoticeBus, createSubagentBus, type AppHooks } from './ui/App';
|
||||
import { Header, type HeaderFact } from './ui/Header';
|
||||
import type { RegistryRow as AppRegistryRow } from './ui/Panels';
|
||||
|
||||
// SDK warnings go straight to stderr, which tears up the Ink render.
|
||||
@@ -154,6 +157,7 @@ if (resumeArg) {
|
||||
const mcp = has('--no-mcp') || !cfg.mcpServers ? undefined : await connectMcp(cfg.mcpServers);
|
||||
const instructions = has('--no-instructions') ? [] : await loadInstructions();
|
||||
const skills = has('--no-skills') ? [] : await loadSkills();
|
||||
const customCommands = await loadCustomCommands();
|
||||
const promptHistory = await store.loadHistory();
|
||||
|
||||
const installedPlugins = has('--no-plugins') ? { plugins: [], errors: [] } : await registry.loadInstalledPlugins();
|
||||
@@ -202,9 +206,29 @@ const enabledPlugins = has('--no-plugins') ? [] : (cfg.plugins ?? DEFAULT_ENABLE
|
||||
const pluginErrors = enabledPlugins
|
||||
.filter((name) => !BUILTIN_PLUGINS.some((p) => p.name === name))
|
||||
.map((name) => ({ plugin: name, message: 'no such plugin' }));
|
||||
|
||||
// External skills, tools, and plugins auto-load from ~/.shiro-neko/<kind> and
|
||||
// .shiro/<kind>. All are data, never code; a bad file is reported, not fatal.
|
||||
const externalPlugins = has('--no-plugins') ? { plugins: [], errors: [] } : await loadExternalPlugins(process.cwd());
|
||||
|
||||
const plugins = createHost(
|
||||
[...BUILTIN_PLUGINS.filter((p) => enabledPlugins.includes(p.name)), ...installedPlugins.plugins],
|
||||
[...pluginErrors, ...installedPlugins.errors],
|
||||
[
|
||||
...BUILTIN_PLUGINS.filter((p) => enabledPlugins.includes(p.name)),
|
||||
...installedPlugins.plugins,
|
||||
...externalPlugins.plugins,
|
||||
],
|
||||
[
|
||||
...pluginErrors,
|
||||
...installedPlugins.errors,
|
||||
...externalPlugins.errors.map((e) => ({ plugin: e.name, message: e.message })),
|
||||
],
|
||||
);
|
||||
|
||||
// External shell tools run through the same guard chain as a built-in bash call,
|
||||
// so an installed tool cannot do what the agent itself may not. Late-bound because
|
||||
// the host above is what runs the chain.
|
||||
const externalTools = await loadExternalTools(process.cwd(), async (command) =>
|
||||
plugins.guard({ toolName: 'bash', input: { command }, cwd: process.cwd() }),
|
||||
);
|
||||
|
||||
const memory = has('--no-memory') ? undefined : new Memory(process.cwd(), languageModel);
|
||||
@@ -261,8 +285,23 @@ const subagentGate: SubagentApproval = (req) => {
|
||||
return approveSubagent(req);
|
||||
};
|
||||
|
||||
// A subagent doing search rather than reasoning can run on a cheaper model.
|
||||
// It resolves against the same provider and key, so a configured `subagentModel`
|
||||
// never needs a second credential.
|
||||
const subagentModel =
|
||||
cfg.subagentModel && cfg.subagentModel !== cfg.model && cfg.apiKey
|
||||
? resolveModel({ ...cfg, model: cfg.subagentModel }, reportFallback)
|
||||
: (languageModel ?? unconfiguredModel);
|
||||
|
||||
// Late-bound like `approveSubagent`: the task tool is built into `extraTools`
|
||||
// before the Session that owns the spend ledger exists, so the usage callback is
|
||||
// wired after construction.
|
||||
let recordSubagent: (usage: { inputTokens: number; outputTokens: number }) => void = () => {};
|
||||
|
||||
const session = new Session({
|
||||
model: languageModel ?? unconfiguredModel,
|
||||
modelId: cfg.model,
|
||||
...(cfg.subagentModel ? { subagentModelId: cfg.subagentModel } : {}),
|
||||
askApproval: bridge.ask,
|
||||
yolo,
|
||||
instructions,
|
||||
@@ -276,8 +315,10 @@ const session = new Session({
|
||||
...(memory ? { memory } : {}),
|
||||
...(record.notebook ? { notebook: record.notebook } : {}),
|
||||
...(cfg.maxRetries !== undefined ? { maxRetries: cfg.maxRetries } : {}),
|
||||
...(cfg.maxSpendUsd !== undefined ? { maxSpendUsd: cfg.maxSpendUsd } : {}),
|
||||
extraTools: {
|
||||
...(mcp?.tools ?? {}),
|
||||
...externalTools.tools,
|
||||
git_commit_message: createCommitMessageTool({
|
||||
model: languageModel ?? unconfiguredModel,
|
||||
...(headless ? {} : { cwd: process.cwd() }),
|
||||
@@ -287,6 +328,9 @@ const session = new Session({
|
||||
: {
|
||||
task: createTaskTool({
|
||||
model: languageModel ?? unconfiguredModel,
|
||||
subagentModel,
|
||||
subagentModelId: cfg.subagentModel,
|
||||
onUsage: (u) => recordSubagent(u),
|
||||
...(headless ? {} : { report: subagents.emit }),
|
||||
// A worker's writes go through the parent's rules and the parent's
|
||||
// prompt. Headless has nobody to answer, so `worker` is withheld there
|
||||
@@ -295,7 +339,7 @@ const session = new Session({
|
||||
}),
|
||||
}),
|
||||
},
|
||||
autoApprove: ['task', 'git_commit_message'],
|
||||
autoApprove: ['task', 'git_commit_message', ...externalTools.autoApprove],
|
||||
messages: [...record.messages],
|
||||
onChange: (messages) => {
|
||||
// Debounced so a long tool loop does not hit the disk on every step.
|
||||
@@ -305,6 +349,7 @@ const session = new Session({
|
||||
});
|
||||
|
||||
approveSubagent = session.approveForSubagent();
|
||||
recordSubagent = (u) => session.recordSubagentUsage(u);
|
||||
|
||||
async function shutdown(code: number): Promise<never> {
|
||||
clearTimeout(saveTimer);
|
||||
@@ -344,6 +389,7 @@ const hooks: AppHooks = {
|
||||
for await (const rel of walk({ limit: 5000 })) found.push(rel);
|
||||
return found;
|
||||
},
|
||||
customCommands: () => customCommands,
|
||||
registry: {
|
||||
list: async () => {
|
||||
const entries = await registry.fetchIndex(cfg.registryUrl);
|
||||
@@ -548,35 +594,46 @@ const hooks: AppHooks = {
|
||||
},
|
||||
};
|
||||
|
||||
const header = [
|
||||
needsProvider
|
||||
? `shiro-neko ${VERSION} no provider configured`
|
||||
: `shiro-neko ${VERSION} ${cfg.provider}/${record.model} session ${record.id.slice(0, 8)}`,
|
||||
`agent: ${agentVariant.name} thinking: ${agentVariant.thinking}`,
|
||||
`cwd: ${process.cwd()}`,
|
||||
restored ? `resumed ${record.messages.length} messages` : undefined,
|
||||
// The welcome dashboard's environment facts, in scan order. Anything that should
|
||||
// stop the user — a failed plugin, `--yolo`, a missing key — is given a tone so it
|
||||
// lifts out of the quiet metadata rather than blending into it.
|
||||
const facts: HeaderFact[] = [
|
||||
{ label: 'agent', value: `${agentVariant.name} thinking ${agentVariant.thinking}` },
|
||||
restored ? { label: 'resumed', value: `${record.messages.length} messages` } : undefined,
|
||||
instructions.length > 0
|
||||
? `instructions: ${instructions.map((i) => i.path.split(/[\\/]/).at(-1)).join(', ')}`
|
||||
: 'no AGENTS.md found - /init writes one',
|
||||
skills.length > 0 ? `skills: ${skills.map((s) => s.name).join(', ')}` : undefined,
|
||||
plugins.plugins.length > 0 ? `plugins: ${plugins.plugins.map((p) => p.name).join(', ')}` : undefined,
|
||||
...plugins.errors.map((e) => `plugin ${e.plugin}: ${e.message}`),
|
||||
memory && memory.all().length > 0 ? `memory: ${memory.all().length} notes about this project` : undefined,
|
||||
mcp && Object.keys(mcp.tools).length > 0 ? `mcp: ${Object.keys(mcp.tools).length} tools` : undefined,
|
||||
!mcp && cfg.mcpServers && Object.keys(cfg.mcpServers).length > 0
|
||||
? `mcp: ${Object.keys(cfg.mcpServers).length} configured, not connected (--no-mcp)`
|
||||
? { label: 'instructions', value: instructions.map((i) => i.path.split(/[\\/]/).at(-1)!).join(', ') }
|
||||
: { label: 'instructions', value: 'none - /init writes an AGENTS.md', tone: 'info' },
|
||||
skills.length > 0 ? { label: 'skills', value: skills.map((s) => s.name).join(', ') } : undefined,
|
||||
plugins.plugins.length > 0
|
||||
? { label: 'plugins', value: plugins.plugins.map((p) => p.name).join(', ') }
|
||||
: undefined,
|
||||
...(mcp?.errors ?? []).map((e) => `mcp ${e.server} failed: ${e.message}`),
|
||||
...plugins.errors.map((e) => ({ label: 'plugin error', value: `${e.plugin}: ${e.message}`, tone: 'err' as const })),
|
||||
memory && memory.all().length > 0
|
||||
? { label: 'memory', value: `${memory.all().length} notes about this project` }
|
||||
: undefined,
|
||||
mcp && Object.keys(mcp.tools).length > 0 ? { label: 'mcp', value: `${Object.keys(mcp.tools).length} tools` } : undefined,
|
||||
!mcp && cfg.mcpServers && Object.keys(cfg.mcpServers).length > 0
|
||||
? { label: 'mcp', value: `${Object.keys(cfg.mcpServers).length} configured, not connected (--no-mcp)`, tone: 'warn' as const }
|
||||
: undefined,
|
||||
...(mcp?.errors ?? []).map((e) => ({ label: 'mcp error', value: `${e.server}: ${e.message}`, tone: 'err' as const })),
|
||||
yolo
|
||||
? 'approvals: OFF (--yolo), but deny rules and the guard still apply'
|
||||
? { label: 'approvals', value: 'OFF (--yolo) - deny rules and the guard still apply', tone: 'warn' as const }
|
||||
: cfg.permission
|
||||
? `approvals: rules for ${Object.keys(cfg.permission).join(', ')}, defaults elsewhere`
|
||||
: 'approvals: ask for write_file, edit_file, multi_edit, apply_patch, move_file, delete_file, bash, web_fetch, mcp__*',
|
||||
cfg.toolSets ? `tool sets: core, ${cfg.toolSets.join(', ')}` : undefined,
|
||||
'/help for commands',
|
||||
]
|
||||
.filter(Boolean)
|
||||
.join('\n');
|
||||
? { label: 'approvals', value: `rules for ${Object.keys(cfg.permission).join(', ')}, defaults elsewhere` }
|
||||
: { label: 'approvals', value: 'ask for writes, bash, web_fetch, mcp' },
|
||||
cfg.toolSets ? { label: 'tool sets', value: `core, ${cfg.toolSets.join(', ')}` } : undefined,
|
||||
].filter((f): f is HeaderFact => f !== undefined);
|
||||
|
||||
const headerNode = (
|
||||
<Header
|
||||
version={VERSION}
|
||||
{...(needsProvider ? {} : { provider: cfg.provider, model: record.model })}
|
||||
sessionId={record.id.slice(0, 8)}
|
||||
cwd={process.cwd()}
|
||||
title={restored ? record.title : undefined}
|
||||
facts={facts}
|
||||
/>
|
||||
);
|
||||
|
||||
// ctrl-c has to reach the App: with a command running it kills that command and
|
||||
// keeps the turn. Ink's own handler would exit the process before we saw the key.
|
||||
@@ -584,7 +641,9 @@ const app = render(
|
||||
<App
|
||||
session={session}
|
||||
bridge={bridge}
|
||||
header={header}
|
||||
header=""
|
||||
headerNode={headerNode}
|
||||
version={VERSION}
|
||||
hooks={hooks}
|
||||
notices={notices}
|
||||
askBridge={askBridge}
|
||||
|
||||
+18
-6
@@ -1,3 +1,5 @@
|
||||
import type { CustomCommand } from './custom-commands';
|
||||
|
||||
export type CommandAction =
|
||||
| { type: 'none' }
|
||||
| { type: 'prompt'; text: string }
|
||||
@@ -24,6 +26,8 @@ export type CommandAction =
|
||||
| { type: 'info'; text: string }
|
||||
| { type: 'model'; model: string }
|
||||
| { type: 'resume'; id: string }
|
||||
/** A custom command from a markdown file, expanded against its arguments. */
|
||||
| { type: 'custom'; command: CustomCommand; args: string[] }
|
||||
| { type: 'unknown'; name: string };
|
||||
|
||||
export type CommandSpec = {
|
||||
@@ -81,11 +85,12 @@ export const HELP = [
|
||||
* An exact name sorts first so pressing enter on `/model` cannot run `/models`.
|
||||
* Aliases stay hidden to keep the list short.
|
||||
*/
|
||||
export function matchCommands(input: string): CommandSpec[] {
|
||||
export function matchCommands(input: string, custom: readonly CustomCommand[] = []): CommandSpec[] {
|
||||
if (!input.startsWith('/')) return [];
|
||||
const typed = input.slice(1).toLowerCase();
|
||||
if (typed.includes(' ')) return [];
|
||||
const hits = COMMANDS.filter((c) => c.name.startsWith(typed));
|
||||
const customSpecs: CommandSpec[] = custom.map((c) => ({ name: c.name, summary: c.description }));
|
||||
const hits = [...COMMANDS, ...customSpecs].filter((c) => c.name.startsWith(typed));
|
||||
const exact = hits.findIndex((c) => c.name === typed);
|
||||
return exact > 0 ? [hits[exact]!, ...hits.filter((_, i) => i !== exact)] : hits;
|
||||
}
|
||||
@@ -156,8 +161,13 @@ function parseMcp(arg: string): CommandAction {
|
||||
}
|
||||
}
|
||||
|
||||
/** Pure parser: no IO, so the TUI and headless mode share one definition. */
|
||||
export function parseCommand(raw: string): CommandAction {
|
||||
/**
|
||||
* Pure parser: no IO, so the TUI and headless mode share one definition.
|
||||
*
|
||||
* Custom commands are consulted only after every built-in name misses, so a
|
||||
* markdown file can add a command but never shadow one that ships with the binary.
|
||||
*/
|
||||
export function parseCommand(raw: string, custom: readonly CustomCommand[] = []): CommandAction {
|
||||
const input = raw.trim();
|
||||
if (!input) return { type: 'none' };
|
||||
if (!input.startsWith('/')) return { type: 'prompt', text: input };
|
||||
@@ -215,7 +225,9 @@ export function parseCommand(raw: string): CommandAction {
|
||||
return arg ? { type: 'model', model: arg } : { type: 'models' };
|
||||
case 'resume':
|
||||
return arg ? { type: 'resume', id: arg } : { type: 'info', text: 'usage: /resume <session-id>' };
|
||||
default:
|
||||
return { type: 'unknown', name };
|
||||
default: {
|
||||
const cmd = custom.find((c) => c.name === name);
|
||||
return cmd ? { type: 'custom', command: cmd, args: arg ? arg.split(/\s+/) : [] } : { type: 'unknown', name };
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -20,6 +20,10 @@ export type Config = {
|
||||
presetId?: string;
|
||||
/** Retries per model call for transient failures. SDK default is 2. */
|
||||
maxRetries?: number;
|
||||
/** USD ceiling for a session's spend: warn at 80%, refuse the next turn at 100%. */
|
||||
maxSpendUsd?: number;
|
||||
/** Model id for subagents; omit to share the parent's. */
|
||||
subagentModel?: string;
|
||||
/** Default agent variant name. */
|
||||
agent?: string;
|
||||
/** Default thinking level. */
|
||||
@@ -90,6 +94,8 @@ export async function loadConfig(): Promise<Config> {
|
||||
apiKey: process.env['SHIRO_API_KEY'] ?? file.apiKey ?? process.env[ENV_KEY[provider]],
|
||||
...(file.presetId ? { presetId: file.presetId } : {}),
|
||||
...(file.maxRetries !== undefined ? { maxRetries: file.maxRetries } : {}),
|
||||
...(typeof file.maxSpendUsd === 'number' && file.maxSpendUsd > 0 ? { maxSpendUsd: file.maxSpendUsd } : {}),
|
||||
...(file.subagentModel ? { subagentModel: file.subagentModel } : {}),
|
||||
...(file.agent ? { agent: file.agent } : {}),
|
||||
...(file.thinking ? { thinking: file.thinking } : {}),
|
||||
...(Array.isArray(file.plugins) ? { plugins: file.plugins } : {}),
|
||||
|
||||
@@ -0,0 +1,121 @@
|
||||
import { homedir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { guardPlugin } from './plugins-builtin';
|
||||
|
||||
/**
|
||||
* Custom slash commands read from markdown files.
|
||||
*
|
||||
* `.shiro/commands/<name>.md` in the project and `~/.shiro-neko/commands/<name>.md`
|
||||
* for the user. The filename is the command; the body becomes the prompt. A project
|
||||
* command shadows a user command of the same name, so a repo can specialise a
|
||||
* personal default.
|
||||
*/
|
||||
export type CustomCommand = {
|
||||
name: string;
|
||||
/** One-line summary for the `/` menu, from frontmatter or the first body line. */
|
||||
description: string;
|
||||
/** Agent to run it under, when frontmatter sets one. */
|
||||
agent?: string;
|
||||
/** The prompt template, before substitution. */
|
||||
body: string;
|
||||
origin: 'project' | 'user';
|
||||
path: string;
|
||||
};
|
||||
|
||||
const MAX_BODY = 20_000;
|
||||
|
||||
/** Reads frontmatter `description` and `agent`; everything after the `---` fence is the prompt. */
|
||||
function parse(name: string, source: string, origin: CustomCommand['origin'], path: string): CustomCommand | undefined {
|
||||
const match = /^---\r?\n([\s\S]*?)\r?\n---\r?\n?([\s\S]*)$/.exec(source.trimStart());
|
||||
const meta: Record<string, string> = {};
|
||||
let body = source;
|
||||
if (match) {
|
||||
for (const line of match[1]!.split(/\r?\n/)) {
|
||||
const kv = /^([A-Za-z_-]+)\s*:\s*(.*)$/.exec(line.trim());
|
||||
if (kv) meta[kv[1]!.toLowerCase()] = kv[2]!.replace(/^["']|["']$/g, '').trim();
|
||||
}
|
||||
body = match[2]!;
|
||||
}
|
||||
const trimmed = body.trim().slice(0, MAX_BODY);
|
||||
if (!trimmed) return undefined;
|
||||
const description = meta['description'] ?? trimmed.split('\n').find((l) => l.trim().length > 0)?.trim().slice(0, 60) ?? name;
|
||||
return {
|
||||
name,
|
||||
description,
|
||||
...(meta['agent'] ? { agent: meta['agent'] } : {}),
|
||||
body: trimmed,
|
||||
origin,
|
||||
path,
|
||||
};
|
||||
}
|
||||
|
||||
function commandDirs(cwd: string): { dir: string; origin: CustomCommand['origin'] }[] {
|
||||
const home = join(process.env['SHIRO_HOME'] ?? homedir(), '.shiro-neko');
|
||||
return [
|
||||
{ dir: join(home, 'commands'), origin: 'user' },
|
||||
{ dir: join(cwd, '.shiro', 'commands'), origin: 'project' },
|
||||
];
|
||||
}
|
||||
|
||||
/** Loads every custom command, project shadowing user by name. A file that fails to parse is skipped. */
|
||||
export async function loadCustomCommands(cwd = process.cwd()): Promise<CustomCommand[]> {
|
||||
const byName = new Map<string, CustomCommand>();
|
||||
for (const { dir, origin } of commandDirs(cwd)) {
|
||||
let files: string[] = [];
|
||||
try {
|
||||
for await (const f of new Bun.Glob('*.md').scan({ cwd: dir, onlyFiles: true })) files.push(f);
|
||||
} catch {
|
||||
continue;
|
||||
}
|
||||
for (const file of files.sort()) {
|
||||
const name = file.replace(/\.md$/i, '');
|
||||
if (!/^[a-z0-9][a-z0-9-_]*$/i.test(name)) continue;
|
||||
const path = join(dir, file);
|
||||
try {
|
||||
const cmd = parse(name, await Bun.file(path).text(), origin, path);
|
||||
if (cmd) byName.set(cmd.name, cmd);
|
||||
} catch {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
}
|
||||
return [...byName.values()].sort((a, b) => a.name.localeCompare(b.name));
|
||||
}
|
||||
|
||||
/** Runs a `` !`cmd` `` substitution through the guard before executing it. */
|
||||
async function runSubstitution(command: string): Promise<string> {
|
||||
const blocked = await guardPlugin.beforeToolCall!({ toolName: 'bash', input: { command }, cwd: process.cwd() });
|
||||
if (blocked) throw new Error(`shell substitution refused: ${blocked}`);
|
||||
|
||||
// The same shell bash uses, so a substitution and a bash call agree on syntax.
|
||||
const shell = process.platform === 'win32' ? ['cmd', '/c', command] : ['bash', '-lc', command];
|
||||
const proc = Bun.spawn(shell, { stdout: 'pipe', stderr: 'pipe' });
|
||||
const [out, err, code] = await Promise.all([
|
||||
new Response(proc.stdout).text(),
|
||||
new Response(proc.stderr).text(),
|
||||
proc.exited,
|
||||
]);
|
||||
if (code !== 0) throw new Error(`shell substitution \`!${command}\` exited ${code}: ${err.trim().slice(0, 200)}`);
|
||||
return out.trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Expands a command's body against the arguments it was typed with.
|
||||
*
|
||||
* `$ARGUMENTS` is the whole argument string, `$1`, `$2`, … the positionals, and
|
||||
* `` !`cmd` `` runs a shell command and inlines its output — each such command
|
||||
* passed through the guard first, so a custom command cannot smuggle a destructive
|
||||
* call past the user the way a plain bash call cannot.
|
||||
*/
|
||||
export async function expandCommand(cmd: CustomCommand, args: string[]): Promise<string> {
|
||||
let out = cmd.body;
|
||||
out = out.replaceAll('$ARGUMENTS', args.join(' '));
|
||||
out = out.replace(/\$(\d+)/g, (_, i) => args[Number(i) - 1] ?? '');
|
||||
|
||||
const substitutions = [...out.matchAll(/!`([^`]+)`/g)];
|
||||
for (const m of substitutions) {
|
||||
const value = await runSubstitution(m[1]!);
|
||||
out = out.replace(m[0], value);
|
||||
}
|
||||
return out.trim();
|
||||
}
|
||||
+133
-3
@@ -218,6 +218,122 @@ export const formatPlugin: Plugin = {
|
||||
},
|
||||
};
|
||||
|
||||
/**
|
||||
* Bash command patterns a guard refuses, shared by several small plugins.
|
||||
*
|
||||
* Each plugin owns one concern so it can be toggled alone; they are data (a name,
|
||||
* a pattern list, an appendix), never code beyond the matcher they all share.
|
||||
*/
|
||||
const bashRefusal = (patterns: { re: RegExp; why: string }[]) => {
|
||||
return ({ toolName, input }: Parameters<NonNullable<Plugin['beforeToolCall']>>[0]) => {
|
||||
if (toolName !== 'bash') return undefined;
|
||||
const command = String((input as { command?: unknown } | null)?.command ?? '');
|
||||
if (!command) return undefined;
|
||||
for (const { re, why } of patterns) {
|
||||
if (re.test(command)) return `refusing "${command.slice(0, 120)}" (${why}). Run it yourself if it is really needed.`;
|
||||
}
|
||||
return undefined;
|
||||
};
|
||||
};
|
||||
|
||||
export const noForcePushPlugin: Plugin = {
|
||||
name: 'no-force-push',
|
||||
description: 'refuses any push that rewrites remote history',
|
||||
appendix: 'The no-force-push plugin refuses force pushes. Ask the user to run one by hand if it is truly intended.',
|
||||
beforeToolCall: bashRefusal([
|
||||
{ re: /\bgit\s+push\b[^|]*(--force\b|--force-with-lease\b|\s-f\b)/, why: 'rewrites remote history' },
|
||||
{ re: /\bgit\s+push\b[^|]*\s+\+/, why: 'a force push via refspec' },
|
||||
]),
|
||||
};
|
||||
|
||||
export const noMainCommitPlugin: Plugin = {
|
||||
name: 'no-main-commit',
|
||||
description: 'refuses to commit directly to main or master',
|
||||
appendix: 'The no-main-commit plugin refuses to commit to main/master. Create a branch and commit there instead.',
|
||||
beforeToolCall: bashRefusal([
|
||||
{ re: /\bgit\s+(commit|merge)\b[^|]*\b(main|master)\b/, why: 'touches the default branch directly' },
|
||||
{ re: /\bgit\s+checkout\s+(main|master)\b[^|]*&&[^|]*\bcommit\b/, why: 'commits on the default branch' },
|
||||
]),
|
||||
};
|
||||
|
||||
export const noRootPlugin: Plugin = {
|
||||
name: 'no-root',
|
||||
description: 'refuses commands run with sudo or as an elevated shell',
|
||||
appendix: 'The no-root plugin refuses sudo and elevation. Nothing the agent does should need it; ask the user to run it themselves.',
|
||||
beforeToolCall: bashRefusal([
|
||||
{ re: /(^|\s)sudo\b/, why: 'elevated privileges' },
|
||||
{ re: /\brunas\b|\bStart-Process\b[^|]*-Verb\s+RunAs/i, why: 'an elevated process' },
|
||||
]),
|
||||
};
|
||||
|
||||
export const noNetPipePlugin: Plugin = {
|
||||
name: 'no-net-pipe',
|
||||
description: 'refuses to execute anything downloaded straight into a shell',
|
||||
appendix: 'The no-net-pipe plugin refuses piping a download into an interpreter. Download, review the file, then run it.',
|
||||
beforeToolCall: bashRefusal([
|
||||
{ re: /\b(curl|wget)\b[^|]*\|\s*(ba|z|k)?sh\b|\b(curl|wget)\b[^|]*\|\s*(node|python|ruby|perl|bun)\b/i, why: 'executes a download unseen' },
|
||||
{ re: /\biex\b|\bInvoke-Expression\b[^|]*\b(iwr|Invoke-WebRequest|curl)\b/i, why: 'executes a download unseen' },
|
||||
]),
|
||||
};
|
||||
|
||||
export const noGitConfigPlugin: Plugin = {
|
||||
name: 'no-git-config',
|
||||
description: 'refuses to change git configuration or global state',
|
||||
appendix: 'The no-git-config plugin refuses to edit git config. Tell the user the exact config change to make themselves.',
|
||||
beforeToolCall: bashRefusal([
|
||||
{ re: /\bgit\s+config\b[^|]*(--global|--system)/, why: 'changes global git configuration' },
|
||||
{ re: /\bgit\s+config\b[^|]*(user\.(name|email)|core\.(sshCommand|editor|pager))\s+\S/, why: 'changes how git identifies or runs' },
|
||||
]),
|
||||
};
|
||||
|
||||
export const noEnvWritePlugin: Plugin = {
|
||||
name: 'no-env-write',
|
||||
description: 'refuses to print or export secrets into the shell environment',
|
||||
appendix: 'The no-env-write plugin refuses to export or echo credentials into the environment. The user sets their own secrets.',
|
||||
beforeToolCall: bashRefusal([
|
||||
{ re: /\b(export|setx?)\s+[A-Z_]*(KEY|TOKEN|SECRET|PASSWORD|PASSWD)\s*=/i, why: 'writes a credential into the environment' },
|
||||
{ re: /\becho\b[^|]*\b(api[_-]?key|secret|token|password)\b[^|]*>>?\s*\S/i, why: 'writes a credential to a file' },
|
||||
]),
|
||||
};
|
||||
|
||||
export const conventionalCommitPlugin: Plugin = {
|
||||
name: 'conventional-commit',
|
||||
description: 'nudges commit messages toward the conventional format',
|
||||
appendix:
|
||||
'The conventional-commit plugin is advisory: write commit subjects as type(scope): summary, e.g. ' +
|
||||
'`fix(auth): reject expired tokens`. Types: feat, fix, docs, style, refactor, perf, test, build, ci, chore, revert.',
|
||||
};
|
||||
|
||||
export const testsFirstPlugin: Plugin = {
|
||||
name: 'tests-first',
|
||||
description: 'reminds the agent to pin behaviour with a failing test before fixing',
|
||||
appendix:
|
||||
'The tests-first plugin is advisory: for a bug, write or find the test that reproduces it before changing code. ' +
|
||||
'Watch it fail, then fix, then watch it pass. A fix without a failing-then-passing test is unverified.',
|
||||
};
|
||||
|
||||
export const smallDiffsPlugin: Plugin = {
|
||||
name: 'small-diffs',
|
||||
description: 'reminds the agent to keep a change focused on one thing',
|
||||
appendix:
|
||||
'The small-diffs plugin is advisory: one change does one thing. Do not tidy, rename, or reformat outside the ' +
|
||||
'task. A diff that is hard to review is usually two diffs wearing one coat — split it.',
|
||||
};
|
||||
|
||||
export const confirmDeletePlugin: Plugin = {
|
||||
name: 'confirm-delete',
|
||||
description: 'refuses delete calls that name broad or ambiguous paths',
|
||||
appendix: 'The confirm-delete plugin refuses deletes that name a directory or a wildcard. Delete one explicit file at a time.',
|
||||
beforeToolCall: ({ toolName, input }) => {
|
||||
if (toolName !== 'delete_file') return undefined;
|
||||
const path = String((input as { path?: unknown } | null)?.path ?? '');
|
||||
if (/[*?[\]]/.test(path) || path.endsWith('/') || path === '.' || path === '') {
|
||||
return `refusing to delete "${path}" (ambiguous or broad). Delete one explicit file.`;
|
||||
}
|
||||
return undefined;
|
||||
},
|
||||
};
|
||||
|
||||
export const BUILTIN_PLUGINS: Plugin[] = [
|
||||
guardPlugin,
|
||||
secretsPlugin,
|
||||
@@ -225,15 +341,29 @@ export const BUILTIN_PLUGINS: Plugin[] = [
|
||||
bellPlugin,
|
||||
timePlugin,
|
||||
formatPlugin,
|
||||
noForcePushPlugin,
|
||||
noMainCommitPlugin,
|
||||
noRootPlugin,
|
||||
noNetPipePlugin,
|
||||
noGitConfigPlugin,
|
||||
noEnvWritePlugin,
|
||||
conventionalCommitPlugin,
|
||||
testsFirstPlugin,
|
||||
smallDiffsPlugin,
|
||||
confirmDeletePlugin,
|
||||
];
|
||||
|
||||
/**
|
||||
* Enabled unless the config turns them off.
|
||||
*
|
||||
* `guard`, `secrets`, and `protect` are refusals, so they are on: a user who has to
|
||||
* opt into a safety check does not have it. `bell` and `format` both act on their
|
||||
* own — one makes noise, the other writes files — so they are opt-in.
|
||||
* opt into a safety check does not have it. The four narrow safety refusals
|
||||
* (`no-force-push`, `no-net-pipe`, `no-root`, `no-env-write`) are on for the same
|
||||
* reason — each blocks a single irreversible class of mistake. `bell` and `format`
|
||||
* act on their own, and the advisory/opinionated plugins (`no-main-commit`,
|
||||
* `conventional-commit`, `tests-first`, `small-diffs`, `confirm-delete`,
|
||||
* `no-git-config`) encode a workflow preference, so all of those are opt-in.
|
||||
*/
|
||||
export const DEFAULT_ENABLED = ['guard', 'secrets', 'protect', 'time'];
|
||||
export const DEFAULT_ENABLED = ['guard', 'secrets', 'protect', 'time', 'no-force-push', 'no-net-pipe', 'no-root', 'no-env-write'];
|
||||
|
||||
export { DESTRUCTIVE, SECRET_PATHS, PROTECTED_PATHS };
|
||||
|
||||
+49
-4
@@ -43,6 +43,28 @@ const TOOL_DOCS: ToolDoc[] = [
|
||||
name: 'grep',
|
||||
line: 'search contents. Prefer it over reading many files; scope with include to keep results small.',
|
||||
},
|
||||
{ name: 'find_symbol', line: 'jump to where a function, class, or type is defined. Use it before grep when you want a declaration, not every use.' },
|
||||
{ name: 'json_query', line: 'read one value from a JSON file by dotted path, e.g. scripts.build, instead of reading it whole.' },
|
||||
{ name: 'insert_lines', line: 'insert a block at a line number, pushing the rest down. Cheaper than a rewrite for adding to the middle of a file.' },
|
||||
{ name: 'delete_lines', line: 'delete a line range. Refuses the whole file; use delete_file for that.' },
|
||||
{ name: 'replace_lines', line: 'replace a line range with new text in one write.' },
|
||||
{ name: 'append_file', line: 'add to the end of a file without a full rewrite.' },
|
||||
{ name: 'prepend_file', line: 'add to the top of a file, e.g. a header or an import block.' },
|
||||
{ name: 'count_lines', line: 'line counts for one file or a glob. A size read before opening something large.' },
|
||||
{ name: 'tree', line: 'indented directory tree, ignore-aware. Scan a broad shape faster than list_dir.' },
|
||||
{ name: 'file_info', line: 'size, line count, modified time, text or binary, for one file.' },
|
||||
{ name: 'find_files', line: 'find files whose name contains a substring, e.g. "auth". Not a glob.' },
|
||||
{ name: 'recent_files', line: 'files modified most recently. Find what a tool just touched.' },
|
||||
{ name: 'changed_files', line: 'the working-tree delta git reports, at a glance.' },
|
||||
{ name: 'git_log_file', line: 'commits that touched one file, newest first.' },
|
||||
{ name: 'git_diff_commits', line: 'diff between two refs, optionally one path.' },
|
||||
{ name: 'git_show_file', line: 'a file\'s contents at a ref, e.g. auth.ts at HEAD~3.' },
|
||||
{ name: 'git_current_branch', line: 'current branch with upstream and ahead/behind.' },
|
||||
{ name: 'git_changed_in_ref', line: 'files changed between a ref and the working tree, names only.' },
|
||||
{ name: 'outline', line: 'top-level declarations of a source file. Read it before opening a large file.' },
|
||||
{ name: 'read_symbol', line: 'the full body of one definition by name.' },
|
||||
{ name: 'env_info', line: 'platform, shell, and which runtimes are installed, before writing a command.' },
|
||||
{ name: 'count_tokens', line: 'estimate the token cost of a file or string before sending it.' },
|
||||
{
|
||||
name: 'edit_file',
|
||||
line: 'oldString must match byte-for-byte including indentation, and be unique. Include surrounding lines to disambiguate. Prefer several small edits over one large rewrite.',
|
||||
@@ -143,6 +165,7 @@ export function systemPrompt(parts: PromptParts): string {
|
||||
|
||||
const toolNames = availableTools ?? TOOL_DOCS.map((d) => d.name);
|
||||
const canRun = toolNames.includes('bash');
|
||||
const canDelegate = toolNames.includes('task');
|
||||
const approvalTools = toolNames.filter((name) =>
|
||||
['write_file', 'edit_file', 'multi_edit', 'apply_patch', 'move_file', 'delete_file', 'bash', 'web_fetch'].includes(
|
||||
name,
|
||||
@@ -150,19 +173,35 @@ export function systemPrompt(parts: PromptParts): string {
|
||||
);
|
||||
|
||||
const workflow = [
|
||||
'- Read before you write. Ground every claim about the code in something you actually opened.',
|
||||
'- Make the smallest change that solves the task. A bugfix diff contains only the bug.',
|
||||
'- Read before you write. Ground every claim about the code in something you actually opened. Never describe code you have not read.',
|
||||
'- Make the smallest change that solves the task. A bugfix diff contains only the bug; a feature diff contains only the feature.',
|
||||
'- Match the existing style, libraries, and conventions. Sample a neighbouring file before inventing a pattern.',
|
||||
approvalTools.length > 0
|
||||
? `- ${approvalTools.join(', ')} need the user to approve each call. If one is denied, stop and ask what to do instead of working around it.`
|
||||
: '- You have no tools that change anything this turn. Investigate and report; do not describe edits as if you had made them.',
|
||||
canRun
|
||||
? "- After changing code, verify it: run the project's build or tests. \"Should work\" is not verification."
|
||||
? "- After changing code, verify it: run the project's build or tests. \"Should work\" is not verification; output you saw is."
|
||||
: '- You cannot run commands this turn, so say what should be run to verify rather than claiming it passes.',
|
||||
'- When something fails twice, stop and re-read the error literally. Check that the code you think is running is the code that is running.',
|
||||
].join('\n');
|
||||
|
||||
// The failure loop is its own block so a stuck model has a procedure, not a vague
|
||||
// instruction to "try harder". Written as discrete steps because a model in a loop
|
||||
// needs an exit, not encouragement.
|
||||
const recovery = [
|
||||
'- Fail once: read the error literally and fix the thing it names, not the thing you expected.',
|
||||
'- Fail twice on the same attempt: stop. Confirm the code running is the code you think — right file, fresh build, no stale cache or shadowed import.',
|
||||
'- Fail three times: change strategy, not parameters. Reproduce smaller, print the value at the failure point, or ask. Do not re-run the same call hoping for a different result.',
|
||||
].join('\n');
|
||||
|
||||
const delegation = canDelegate
|
||||
? `- Delegate with task for a search across many files or a self-contained change you need not watch. Its prompt must stand alone — it sees none of this conversation. Keep work you must supervise in your own turn.`
|
||||
: '';
|
||||
|
||||
const workflow2 = [
|
||||
canAsk
|
||||
? '- Ask rather than guess when two readings of the request lead to different work. Decide small things yourself and say what you assumed.'
|
||||
: '- No one can answer a question this run. Decide yourself and state the assumption plainly.',
|
||||
'- Long sessions compact as context fills. Record what stays true with remember; restate the goal on a long task.',
|
||||
].join('\n');
|
||||
|
||||
return `You are Shiro Neko, a coding agent working in the user's terminal.
|
||||
@@ -178,6 +217,12 @@ ${renderTools(toolNames)}
|
||||
How to work
|
||||
${workflow}
|
||||
|
||||
When something fails
|
||||
${recovery}
|
||||
${delegation ? `\nDelegating\n${delegation}\n` : ''}
|
||||
Working with the user
|
||||
${workflow2}
|
||||
|
||||
How to reply
|
||||
- Lead with the outcome. The user wants to know what happened, not what you are about to do.
|
||||
- No preamble, no restating the task, no summary of your own summary.
|
||||
|
||||
@@ -15,6 +15,7 @@ import type { Memory } from './memory';
|
||||
import { Notebook, type NotebookState } from './notebook';
|
||||
import { Permissions, type PermissionConfig } from './permission';
|
||||
import type { PluginHost } from './plugins';
|
||||
import { costOf, formatUsd } from './pricing';
|
||||
import { systemPrompt } from './prompt';
|
||||
import { detachProviderItems, pruneToFit } from './prune';
|
||||
import { createSkillTool, renderSkills, type Skill } from './skills';
|
||||
@@ -53,10 +54,16 @@ export type AgentEvent =
|
||||
|
||||
export type SessionOptions = {
|
||||
model: LanguageModel;
|
||||
/** Model id, for pricing the session's spend against the ceiling. */
|
||||
modelId?: string;
|
||||
/** Subagent model id, when it differs; its spend prices against this. */
|
||||
subagentModelId?: string;
|
||||
askApproval: (req: ApprovalRequest) => Promise<ApprovalDecision>;
|
||||
yolo?: boolean;
|
||||
cwd?: string;
|
||||
maxSteps?: number;
|
||||
/** USD ceiling: warn at 80%, refuse the next turn at 100%. */
|
||||
maxSpendUsd?: number;
|
||||
/** MCP and subagent tools merged on top of the built-ins. */
|
||||
extraTools?: ToolSet;
|
||||
/** Tool sets offered this session; omit for all of them. `core` is always on. */
|
||||
@@ -116,6 +123,9 @@ export class Session {
|
||||
readonly notebook: Notebook;
|
||||
inputTokens = 0;
|
||||
outputTokens = 0;
|
||||
/** Subagent token use, priced against the subagent's own model id in /cost. */
|
||||
subagentInputTokens = 0;
|
||||
subagentOutputTokens = 0;
|
||||
private model: LanguageModel;
|
||||
private variant: AgentVariant;
|
||||
private readonly permissions: Permissions;
|
||||
@@ -123,6 +133,8 @@ export class Session {
|
||||
private readonly seen = new Map<string, number>();
|
||||
/** One stale-item repair per turn, so a repeating 404 cannot loop the run. */
|
||||
private staleItemsRepaired = false;
|
||||
/** The 80% spend warning is shown once, not on every turn past the line. */
|
||||
private warnedSpend = false;
|
||||
private controller: AbortController | undefined;
|
||||
|
||||
constructor(private readonly opts: SessionOptions) {
|
||||
@@ -213,10 +225,19 @@ export class Session {
|
||||
this.messages.length = 0;
|
||||
this.inputTokens = 0;
|
||||
this.outputTokens = 0;
|
||||
this.subagentInputTokens = 0;
|
||||
this.subagentOutputTokens = 0;
|
||||
this.warnedSpend = false;
|
||||
this.notebook.clear();
|
||||
this.opts.onChange?.(this.messages);
|
||||
}
|
||||
|
||||
/** A subagent's finished run, folded into the session's spend and the /cost split. */
|
||||
recordSubagentUsage(usage: { inputTokens: number; outputTokens: number }): void {
|
||||
this.subagentInputTokens += usage.inputTokens;
|
||||
this.subagentOutputTokens += usage.outputTokens;
|
||||
}
|
||||
|
||||
replace(messages: ModelMessage[]): void {
|
||||
this.messages.length = 0;
|
||||
this.messages.push(...messages);
|
||||
@@ -236,6 +257,27 @@ export class Session {
|
||||
return this.opts.compactThreshold ?? DEFAULT_COMPACT_THRESHOLD;
|
||||
}
|
||||
|
||||
/**
|
||||
* The session's spend so far and the configured ceiling, for the UI's status
|
||||
* and the refuse-the-next-turn check. Unpriced models report no spend: a
|
||||
* ceiling cannot be enforced against a model we cannot price.
|
||||
*/
|
||||
spend(): { usd?: number; ceiling?: number; overWarn: boolean; overLimit: boolean } {
|
||||
const ceiling = this.opts.maxSpendUsd;
|
||||
const parent = costOf(this.opts.modelId ?? '', this.inputTokens, this.outputTokens);
|
||||
const sub =
|
||||
this.subagentInputTokens + this.subagentOutputTokens > 0
|
||||
? costOf(this.opts.subagentModelId ?? this.opts.modelId ?? '', this.subagentInputTokens, this.subagentOutputTokens)
|
||||
: 0;
|
||||
// Spend is only knowable when every part is priced; an unpriced piece means
|
||||
// the total is a lower bound, so the ceiling is not enforced against it.
|
||||
const usd = parent === undefined || sub === undefined ? undefined : parent + sub;
|
||||
if (ceiling === undefined || usd === undefined) {
|
||||
return { ...(usd !== undefined ? { usd } : {}), ...(ceiling !== undefined ? { ceiling } : {}), overWarn: false, overLimit: false };
|
||||
}
|
||||
return { usd, ceiling, overWarn: usd >= ceiling * 0.8, overLimit: usd >= ceiling };
|
||||
}
|
||||
|
||||
private systemFor(): string {
|
||||
return systemPrompt({
|
||||
cwd: this.opts.cwd ?? process.cwd(),
|
||||
@@ -337,6 +379,21 @@ export class Session {
|
||||
}
|
||||
|
||||
async *send(userText: string): AsyncGenerator<AgentEvent> {
|
||||
// The ceiling is checked before the model is: a turn started past the limit
|
||||
// would spend money the caller said not to. An unpriced model cannot be
|
||||
// measured, so it is never refused here — the ceiling simply cannot see it.
|
||||
const spend = this.spend();
|
||||
if (spend.overLimit) {
|
||||
yield {
|
||||
type: 'error',
|
||||
error: new Error(
|
||||
`spend ceiling reached: ${formatUsd(spend.usd ?? 0)} of ${formatUsd(spend.ceiling ?? 0)} used. Raise maxSpendUsd or start a new session.`,
|
||||
),
|
||||
};
|
||||
yield { type: 'done' };
|
||||
return;
|
||||
}
|
||||
|
||||
this.messages.push({ role: 'user', content: userText });
|
||||
this.opts.onChange?.(this.messages);
|
||||
this.controller = new AbortController();
|
||||
@@ -528,6 +585,16 @@ export class Session {
|
||||
const usage = await result.usage;
|
||||
this.inputTokens += usage.inputTokens ?? 0;
|
||||
this.outputTokens += usage.outputTokens ?? 0;
|
||||
// Warn as the ceiling comes into view, once, so a long session is not
|
||||
// surprised by a refusal it never saw coming.
|
||||
const spend = this.spend();
|
||||
if (spend.overWarn && !this.warnedSpend) {
|
||||
this.warnedSpend = true;
|
||||
yield {
|
||||
type: 'notice',
|
||||
text: `approaching spend ceiling: ${formatUsd(spend.usd ?? 0)} of ${formatUsd(spend.ceiling ?? 0)} used`,
|
||||
};
|
||||
}
|
||||
yield { type: 'done', inputTokens: usage.inputTokens, outputTokens: usage.outputTokens };
|
||||
return;
|
||||
}
|
||||
|
||||
+64
-414
@@ -1,420 +1,70 @@
|
||||
/**
|
||||
* Skills bundled with the binary.
|
||||
*
|
||||
* These are string constants rather than files on disk because `bun build --compile`
|
||||
* only embeds modules reachable through imports; a directory of .md files would be
|
||||
* missing from the shipped binary.
|
||||
* Each skill is a Markdown file in `src/skills-md/`, loaded here as a raw-text import.
|
||||
* The `.md` file is the single source of truth — frontmatter and body in proper
|
||||
* Markdown — so skills are edited and reviewed as Markdown, not as escaped strings
|
||||
* inside TypeScript. Bun inlines every text import into the compiled binary, so the
|
||||
* folder ships with `bun build --compile` exactly as the old string constants did.
|
||||
*/
|
||||
import accessibility from './skills-md/accessibility.md' with { type: 'text' };
|
||||
import apiDesign from './skills-md/api-design.md' with { type: 'text' };
|
||||
import ciCd from './skills-md/ci-cd.md' with { type: 'text' };
|
||||
import commit from './skills-md/commit.md' with { type: 'text' };
|
||||
import data from './skills-md/data.md' with { type: 'text' };
|
||||
import db from './skills-md/db.md' with { type: 'text' };
|
||||
import debug from './skills-md/debug.md' with { type: 'text' };
|
||||
import deps from './skills-md/deps.md' with { type: 'text' };
|
||||
import docker from './skills-md/docker.md' with { type: 'text' };
|
||||
import docs from './skills-md/docs.md' with { type: 'text' };
|
||||
import frontend from './skills-md/frontend.md' with { type: 'text' };
|
||||
import gitWorkflow from './skills-md/git-workflow.md' with { type: 'text' };
|
||||
import i18n from './skills-md/i18n.md' with { type: 'text' };
|
||||
import incident from './skills-md/incident.md' with { type: 'text' };
|
||||
import logging from './skills-md/logging.md' with { type: 'text' };
|
||||
import migrate from './skills-md/migrate.md' with { type: 'text' };
|
||||
import onboarding from './skills-md/onboarding.md' with { type: 'text' };
|
||||
import optimizeSql from './skills-md/optimize-sql.md' with { type: 'text' };
|
||||
import perf from './skills-md/perf.md' with { type: 'text' };
|
||||
import perfFrontend from './skills-md/perf-frontend.md' with { type: 'text' };
|
||||
import plan from './skills-md/plan.md' with { type: 'text' };
|
||||
import readme from './skills-md/readme.md' with { type: 'text' };
|
||||
import refactor from './skills-md/refactor.md' with { type: 'text' };
|
||||
import release from './skills-md/release.md' with { type: 'text' };
|
||||
import review from './skills-md/review.md' with { type: 'text' };
|
||||
import security from './skills-md/security.md' with { type: 'text' };
|
||||
import test from './skills-md/test.md' with { type: 'text' };
|
||||
import uxCopy from './skills-md/ux-copy.md' with { type: 'text' };
|
||||
import verify from './skills-md/verify.md' with { type: 'text' };
|
||||
|
||||
export const BUILTIN_SKILLS: { name: string; source: string }[] = [
|
||||
{
|
||||
name: 'debug',
|
||||
source: `---
|
||||
name: debug
|
||||
description: Track down a bug whose cause is not obvious. Use when a test fails for unclear reasons, behaviour differs between environments, or an earlier fix did not hold.
|
||||
---
|
||||
|
||||
# Debugging
|
||||
|
||||
Do not guess. A guess that happens to work leaves the real cause in place.
|
||||
|
||||
## Reproduce first
|
||||
|
||||
Find the smallest command that shows the failure and record it with \`remember\`. If you
|
||||
cannot reproduce it, say so and ask what the user did differently — do not proceed on a
|
||||
hypothesis you cannot test.
|
||||
|
||||
## Three hypotheses, then evidence
|
||||
|
||||
Write down at least three causes that would produce this exact symptom. Rank them by how
|
||||
cheap they are to disprove, then disprove them in that order. State which one you are
|
||||
testing before you test it.
|
||||
|
||||
Evidence means observed output: a log line, a failing assertion, a value printed at the
|
||||
point of failure. "It should be X" is not evidence.
|
||||
|
||||
## Bisect when the space is large
|
||||
|
||||
- Recent regression: check what changed last.
|
||||
- Unclear layer: assert the value at each boundary until one is wrong.
|
||||
- Intermittent: run it in a loop and capture the failing case, do not reason about it abstractly.
|
||||
|
||||
## Fix the cause
|
||||
|
||||
Once you know the cause, fix that and nothing else. Do not tidy surrounding code in the
|
||||
same change — a bugfix diff should contain only the bug.
|
||||
|
||||
Write a test that fails before the fix and passes after. If you cannot express the bug as
|
||||
a test, say why.
|
||||
|
||||
## After two failed attempts
|
||||
|
||||
Stop. Re-read the error text literally, character by character. Check your assumption
|
||||
about which code is actually running: the wrong file, a stale build, a shadowed import,
|
||||
or a cached dependency accounts for most "impossible" bugs.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'review',
|
||||
source: `---
|
||||
name: review
|
||||
description: Review a diff or a file for defects. Use when asked to review, critique, or check code before it ships.
|
||||
---
|
||||
|
||||
# Code review
|
||||
|
||||
Severity order. Do not lead with style.
|
||||
|
||||
1. **Incorrect behaviour** — wrong result, wrong edge case, wrong state after failure.
|
||||
2. **Missing validation at trust boundaries** — user input, network responses, file contents,
|
||||
anything crossing a process line. Internal calls need no defensive checks.
|
||||
3. **Security** — injection, path traversal, secrets in logs or errors, missing authz.
|
||||
4. **Resource handling** — unclosed handles, unbounded growth, unawaited promises.
|
||||
5. **Clarity** — only when it will cause a future defect.
|
||||
|
||||
## For each finding
|
||||
|
||||
State file and line, what breaks, and the change. Show the fix as code when it is short.
|
||||
|
||||
Skip anything a formatter would fix. Skip preference. If a choice is defensible, leave it.
|
||||
|
||||
## Say when it is fine
|
||||
|
||||
A review that invents problems to look thorough is worse than a short one. If the change
|
||||
is correct, say so and stop.
|
||||
|
||||
## Verify, do not assume
|
||||
|
||||
Read the surrounding code before calling something a bug. A "missing" null check often
|
||||
exists one level up. Run the tests if that is what settles it.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'refactor',
|
||||
source: `---
|
||||
name: refactor
|
||||
description: Restructure code without changing behaviour. Use when asked to refactor, clean up, extract, or reorganise.
|
||||
---
|
||||
|
||||
# Refactoring
|
||||
|
||||
Behaviour must not change. That is the whole constraint.
|
||||
|
||||
## Establish the safety net first
|
||||
|
||||
Run the existing tests and record that they pass. If the code has no tests, write one that
|
||||
pins current behaviour — including the ugly parts — before touching anything. Refactoring
|
||||
untested code is rewriting it.
|
||||
|
||||
## Then move in small steps
|
||||
|
||||
One transformation at a time, tests green between each. Rename, then extract, then move —
|
||||
not all three in one edit. A large refactor that fails leaves you unable to tell which step
|
||||
broke it.
|
||||
|
||||
## What not to do
|
||||
|
||||
- Do not fix bugs while refactoring. Note them, finish, fix separately.
|
||||
- Do not add abstraction for a single caller. Duplication beats a premature interface.
|
||||
- Do not widen the scope. The request was this code, not its neighbours.
|
||||
- Do not change public API unless asked; if it must change, say so first.
|
||||
|
||||
## Done means
|
||||
|
||||
Tests pass, behaviour is identical, and the diff is smaller than the reader feared.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'test',
|
||||
source: `---
|
||||
name: test
|
||||
description: Write or repair tests. Use when adding coverage, fixing a flaky test, or asked how something should be tested.
|
||||
---
|
||||
|
||||
# Testing
|
||||
|
||||
A test earns its place by failing when the code is wrong.
|
||||
|
||||
## Match the project
|
||||
|
||||
Read two existing test files first. Use their runner, their assertion style, their file
|
||||
layout, their naming. A test that looks foreign is a test nobody maintains.
|
||||
|
||||
## Test behaviour, not implementation
|
||||
|
||||
Assert on what a caller observes. A test that reaches into private state breaks on every
|
||||
refactor and catches nothing.
|
||||
|
||||
Cover: the normal case, the boundaries, and the failure. Failure cases catch more real
|
||||
defects than happy paths.
|
||||
|
||||
## Never do this
|
||||
|
||||
- Do not assert what the code currently returns without knowing it is correct — that pins
|
||||
the bug.
|
||||
- Do not weaken an assertion to make a test pass. If it fails, either the code or the
|
||||
expectation is wrong; find out which.
|
||||
- Do not delete a failing test. It is telling you something.
|
||||
|
||||
## Flaky tests
|
||||
|
||||
A test that passes alone and fails in a suite is a shared-state problem: a global, a
|
||||
temp directory, a port, an unawaited promise, or ordering. Find which, do not add a retry.
|
||||
|
||||
## Verify
|
||||
|
||||
Run the test and watch it fail before the fix, pass after. A test you never saw fail is
|
||||
not known to work.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'verify',
|
||||
source: `---
|
||||
name: verify
|
||||
description: Confirm a change actually works by using it, not by reading it. Use before reporting a task complete, or when asked whether something works.
|
||||
---
|
||||
|
||||
# Verification
|
||||
|
||||
A green test suite says the tests pass. It does not say the feature works.
|
||||
|
||||
## Run the artifact, not the source
|
||||
|
||||
Build it and use it the way a user would:
|
||||
|
||||
- **CLI** — build the binary and run it. Happy path, bad input, \`--help\`. Read the output.
|
||||
- **HTTP service** — start it and \`curl\` the endpoint. Check the status and the body.
|
||||
- **Library** — write a throwaway script that imports and calls the new code end to end.
|
||||
- **Script or job** — run it against real input and inspect what it produced.
|
||||
|
||||
Delete the throwaway afterwards.
|
||||
|
||||
## What counts as evidence
|
||||
|
||||
Command output you actually saw. Paste the relevant lines, not a summary of them.
|
||||
|
||||
These are not evidence:
|
||||
|
||||
- "The tests pass" for a change tests do not cover.
|
||||
- "The types check" for anything about runtime behaviour.
|
||||
- "It should work now" for anything at all.
|
||||
|
||||
## Check the failure path too
|
||||
|
||||
Feed it the input you expect to be rejected and confirm it is rejected, with a message
|
||||
that says why. A feature that works only on correct input is half-built.
|
||||
|
||||
## Report what you did not verify
|
||||
|
||||
Say plainly what you could not run and why: a missing credential, a service you cannot
|
||||
start, a platform you are not on. An honest gap is useful; a claim that hides one is not.
|
||||
|
||||
## When verification fails
|
||||
|
||||
The defect is yours to fix in this turn. Do not report the task complete with a note that
|
||||
it did not work.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'commit',
|
||||
source: `---
|
||||
name: commit
|
||||
description: Stage and commit work. Use when asked to commit, or to split existing changes into commits.
|
||||
---
|
||||
|
||||
# Committing
|
||||
|
||||
Never commit unless the user asked. If it is unclear whether they did, ask.
|
||||
|
||||
## Look before you stage
|
||||
|
||||
\`git_status\` and \`git_diff\` first. You are looking for two things:
|
||||
|
||||
1. Changes that are not yours. Another agent or the user may share this worktree, and
|
||||
\`git add .\` takes their half-finished work with yours.
|
||||
2. Files that should never be committed: \`.env\`, credentials, keys, large build output,
|
||||
anything a \`.gitignore\` rule was supposed to catch and did not. Flag these to the user
|
||||
rather than committing them.
|
||||
|
||||
Stage the specific paths you changed. \`git add .\` is how unrelated work ends up in a
|
||||
commit that then has to be reverted whole.
|
||||
|
||||
## One commit, one reason
|
||||
|
||||
If the diff does two unrelated things, make two commits. A commit that both fixes a bug and
|
||||
renames a module cannot be reverted, cherry-picked, or bisected usefully.
|
||||
|
||||
## The message
|
||||
|
||||
Match the repository's existing style — read \`git_log\` before writing one. Failing that:
|
||||
|
||||
- A subject line under 70 characters, imperative, saying what changed.
|
||||
- A body explaining *why*, when the reason is not obvious from the diff. Wrap at 72.
|
||||
- No "as requested", no restating the diff line by line, no emoji unless the repo uses them.
|
||||
|
||||
## Do not
|
||||
|
||||
- Do not \`--amend\` a commit that has been pushed. Write a new one.
|
||||
- Do not \`--no-verify\`. If a hook rejects the commit, the hook found something.
|
||||
- Do not \`git push\` unless asked, and never force-push without being asked explicitly.
|
||||
- Do not commit and then immediately fix it up with a second commit. Get it right, or say
|
||||
what is wrong.
|
||||
|
||||
## After committing
|
||||
|
||||
Report the short hash and the subject. If a hook rewrote files, say so and confirm the
|
||||
final state is what was intended.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'security',
|
||||
source: `---
|
||||
name: security
|
||||
description: Review code for security defects, or write code that handles untrusted input. Use when touching authentication, user input, file paths, shell commands, SQL, or anything reachable from the network.
|
||||
---
|
||||
|
||||
# Security
|
||||
|
||||
Find the trust boundary first. Everything crossing it is hostile until parsed.
|
||||
|
||||
## The boundaries in most codebases
|
||||
|
||||
- Request bodies, query strings, headers, cookies.
|
||||
- File contents and filenames, including paths a user supplied.
|
||||
- Environment variables in a multi-tenant deployment.
|
||||
- Anything a model or a third-party API returned.
|
||||
|
||||
Inside a boundary, values are already validated and re-checking them is noise. At the
|
||||
boundary, nothing is optional.
|
||||
|
||||
## What to look for, in order
|
||||
|
||||
1. **Injection.** String-built SQL, shell commands assembled from input, \`eval\`, template
|
||||
rendering with user data as the template rather than the data. The fix is parameters and
|
||||
argument arrays, never escaping.
|
||||
2. **Missing authorisation.** An endpoint that checks *who* you are but not *what* you may
|
||||
touch. Look for an id taken from the request and used without an ownership check.
|
||||
3. **Path traversal.** \`../\` in anything joined onto a filesystem root. Resolve, then verify
|
||||
the result is still inside the root — a prefix check on the raw input misses
|
||||
\`a/../../secret\`.
|
||||
4. **Secrets in the wrong place.** Keys in source, in logs, in error messages, in a commit.
|
||||
A secret that reached a log is a secret to rotate.
|
||||
5. **Server-side request forgery.** A URL from input, fetched. Block private and loopback
|
||||
addresses by *resolved* address, and re-check every redirect hop.
|
||||
6. **Weak crypto and hand-rolled auth.** Homemade token formats, \`Math.random\` for anything
|
||||
security-bearing, comparisons on secrets that are not constant time.
|
||||
|
||||
## What not to do
|
||||
|
||||
Do not report a finding you cannot trace to a concrete input. "This could be unsafe" without
|
||||
a path from an attacker-controlled value to the sink is noise that buries the real one.
|
||||
|
||||
Do not fix a symptom at one caller when the sink is shared. Grep every caller and fix the
|
||||
seam once.
|
||||
|
||||
## Reporting
|
||||
|
||||
File, line, the path from input to sink, and the fix. Say plainly when a thing that looks
|
||||
dangerous is actually fine, and why — a reviewer's confidence is worth as much as a finding.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'perf',
|
||||
source: `---
|
||||
name: perf
|
||||
description: Make something faster, or find out why it is slow. Use when a command, request, test suite, or build takes longer than it should.
|
||||
---
|
||||
|
||||
# Performance
|
||||
|
||||
Measure first. A change made without a number before it is a guess with extra steps.
|
||||
|
||||
## Get a number
|
||||
|
||||
Time the actual operation, not a proxy for it. \`time\`, the framework's own timing output,
|
||||
or a loop around the slow call with a timestamp either side. Record the baseline with
|
||||
\`remember\` so the comparison survives compaction.
|
||||
|
||||
If you cannot measure it, say so and stop. Optimising an unmeasured path is how a codebase
|
||||
accumulates complexity that buys nothing.
|
||||
|
||||
## Find where the time goes
|
||||
|
||||
- **Wall-clock dominated by one call?** Look there and nowhere else.
|
||||
- **Spread evenly?** Suspect the loop around it: an O(n²) walk, a query per row, a file read
|
||||
per iteration.
|
||||
- **Idle time?** It is waiting: a sequential chain of independent awaits, an unpooled
|
||||
connection, a lock.
|
||||
|
||||
The usual culprits, in the order they actually appear: N+1 queries, work repeated inside a
|
||||
loop that could be hoisted, a missing index, sequential awaits that could run together,
|
||||
reading a whole file to use one line, and re-parsing something that could be parsed once.
|
||||
|
||||
## Change one thing
|
||||
|
||||
One change, then re-measure. Two changes together and you do not know which one paid — and
|
||||
one of them may have cost.
|
||||
|
||||
## Stop when it is fast enough
|
||||
|
||||
State the target before you start: "the test suite under a minute", "the endpoint under
|
||||
200ms". Past the target, further work is complexity with no user on the other end of it.
|
||||
|
||||
## Report
|
||||
|
||||
Baseline, change, new number, and what you did not do. A 40% win with one line changed is a
|
||||
better report than a 45% win that restructured a module.
|
||||
`,
|
||||
},
|
||||
{
|
||||
name: 'migrate',
|
||||
source: `---
|
||||
name: migrate
|
||||
description: Upgrade a dependency, framework, or language version across a codebase. Use when a major version bump, a deprecation, or a breaking API change has to be applied.
|
||||
---
|
||||
|
||||
# Migration
|
||||
|
||||
The failure mode is a half-applied migration: it compiles, most tests pass, and one code
|
||||
path still uses the old API.
|
||||
|
||||
## Read the changelog before the code
|
||||
|
||||
Find what actually broke. A major version usually has a migration guide; read it and list
|
||||
the changes that apply to this codebase specifically. Below 1.0, treat a minor bump as
|
||||
breaking — semver promises nothing there.
|
||||
|
||||
## Find every call site before changing one
|
||||
|
||||
Grep for the old API across the whole repository, including tests, scripts, config, CI
|
||||
workflows, Dockerfiles, and documentation. A version literal pinned in a workflow while the
|
||||
manifest says something else is a split-brain deploy.
|
||||
|
||||
Write the list down with \`todo_write\`. The list is the migration; the edits are mechanical.
|
||||
|
||||
## Change in one shape
|
||||
|
||||
Apply the same transformation everywhere rather than improving each site as you pass
|
||||
through it. A migration mixed with refactoring cannot be reviewed, and cannot be reverted
|
||||
if the upgrade turns out to be wrong.
|
||||
|
||||
\`apply_patch\` is the tool for this: one atomic patch across the files that must land
|
||||
together.
|
||||
|
||||
## Verify at the boundary that broke
|
||||
|
||||
Type checks catch signature changes and miss behaviour changes. Run the tests, then actually
|
||||
use the thing that was upgraded: start the server, run the CLI, execute the query. A green
|
||||
suite over an untested upgrade path proves the suite did not cover it.
|
||||
|
||||
## Never hand-merge a lockfile
|
||||
|
||||
On a conflict, take either side whole and regenerate with the package manager. The resolver
|
||||
owns that file.
|
||||
|
||||
## Report
|
||||
|
||||
The version before and after, every file class touched, what you verified by running, and
|
||||
anything the changelog said applies that you deliberately did not do.
|
||||
`,
|
||||
},
|
||||
{ name: 'accessibility', source: accessibility },
|
||||
{ name: 'api-design', source: apiDesign },
|
||||
{ name: 'ci-cd', source: ciCd },
|
||||
{ name: 'commit', source: commit },
|
||||
{ name: 'data', source: data },
|
||||
{ name: 'db', source: db },
|
||||
{ name: 'debug', source: debug },
|
||||
{ name: 'deps', source: deps },
|
||||
{ name: 'docker', source: docker },
|
||||
{ name: 'docs', source: docs },
|
||||
{ name: 'frontend', source: frontend },
|
||||
{ name: 'git-workflow', source: gitWorkflow },
|
||||
{ name: 'i18n', source: i18n },
|
||||
{ name: 'incident', source: incident },
|
||||
{ name: 'logging', source: logging },
|
||||
{ name: 'migrate', source: migrate },
|
||||
{ name: 'onboarding', source: onboarding },
|
||||
{ name: 'optimize-sql', source: optimizeSql },
|
||||
{ name: 'perf', source: perf },
|
||||
{ name: 'perf-frontend', source: perfFrontend },
|
||||
{ name: 'plan', source: plan },
|
||||
{ name: 'readme', source: readme },
|
||||
{ name: 'refactor', source: refactor },
|
||||
{ name: 'release', source: release },
|
||||
{ name: 'review', source: review },
|
||||
{ name: 'security', source: security },
|
||||
{ name: 'test', source: test },
|
||||
{ name: 'ux-copy', source: uxCopy },
|
||||
{ name: 'verify', source: verify },
|
||||
];
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: accessibility
|
||||
description: Make a UI accessible. Use when adding a feature that must work with a keyboard or screen reader, fixing contrast or focus issues, or reviewing for WCAG.
|
||||
---
|
||||
|
||||
# Accessibility
|
||||
|
||||
Accessibility is usability for everyone, including people using a keyboard, a screen reader, a
|
||||
magnifier, or a noisy display. Build it in, not on.
|
||||
|
||||
## Semantic HTML does the heavy lifting
|
||||
|
||||
A `<button>`, `<a>`, `<input>`, `<nav>`, `<main>` carries behaviour and meaning for free
|
||||
that a `<div>` with a click handler does not. Reach for the native element first; add ARIA only
|
||||
when no native element fits. The first rule of ARIA is do not use ARIA if a native element exists.
|
||||
|
||||
## Keyboard is the baseline
|
||||
|
||||
- Every interactive element is reachable and operable with Tab and Enter/Space alone.
|
||||
- A visible focus indicator on everything — never `outline: none` without a replacement.
|
||||
- Logical tab order following the visual order, and focus managed into and out of modals,
|
||||
menus, and dialogs (trapped while open, returned to the trigger on close).
|
||||
|
||||
## Screen readers hear structure
|
||||
|
||||
- Headings in order (`h1` once, then down a level at a time) so the page has a navigable outline.
|
||||
- Every `<input>` has a `<label>`; every icon-only button has an accessible name; every image
|
||||
has alt text that conveys its point (or empty alt when it is purely decorative).
|
||||
- Dynamic changes announce themselves: a toast, an error, a loaded region uses a live region so
|
||||
it is heard, not just seen.
|
||||
|
||||
## Contrast and meaning
|
||||
|
||||
Text meets 4.5:1 against its background (3:1 for large text). Colour is never the only carrier
|
||||
of meaning — pair it with an icon, a label, or a pattern. A red-only "error" is invisible to a
|
||||
colour-blind user.
|
||||
|
||||
## Test it the way it is used
|
||||
|
||||
Tab through the whole flow. Turn on a screen reader and listen. Zoom to 200% and 400%. Run an
|
||||
automated checker for the mechanical half — then do the manual half it cannot cover, because
|
||||
most accessibility failures are not machine-detectable.
|
||||
@@ -0,0 +1,45 @@
|
||||
---
|
||||
name: api-design
|
||||
description: Design or revise an HTTP or library API. Use when adding an endpoint, shaping request/response bodies, naming resources, or reviewing an API for consistency.
|
||||
---
|
||||
|
||||
# API design
|
||||
|
||||
An API is a contract. Every choice is a promise you cannot take back without a major version.
|
||||
|
||||
## Resource before action
|
||||
|
||||
Name things, not verbs. `POST /users` to create, not `POST /createUser`. The URL is the
|
||||
noun; the method is the verb. When you reach for a verb in the path, that is a sign the
|
||||
resource is missing — `POST /users/:id/deactivations` reads better than `/deactivateUser`
|
||||
when the operation has state of its own.
|
||||
|
||||
## Shape the body for the reader
|
||||
|
||||
- Field names are `snake_case` or `camelCase`, picked once for the whole API. A body that
|
||||
mixes both is a body nobody documented.
|
||||
- Return the object, not a wrapper, unless the wrapper carries something: `{ "user": {...} }`
|
||||
only when there is also pagination, a cursor, or an error envelope.
|
||||
- Errors have a stable shape: a machine-readable `code`, a human `message`, and the field
|
||||
that failed. A client should never have to parse the message.
|
||||
|
||||
## Status codes mean something
|
||||
|
||||
- `201` for a created resource, with the resource in the body.
|
||||
- `204` for success with nothing to return.
|
||||
- `400` for a body that failed validation, `401` unauthenticated, `403` authenticated but
|
||||
not allowed, `404` not found or not allowed to know, `409` a conflict with current state,
|
||||
`422` well-formed but semantically wrong.
|
||||
- Never `200` with an error in the body. A client checking only the status will treat it as
|
||||
success.
|
||||
|
||||
## Idempotency and safety
|
||||
|
||||
GET, PUT, DELETE must be safe to retry: same request, same state. POST is not. If a client can
|
||||
double-submit, provide an idempotency key or a natural unique constraint, and say which.
|
||||
|
||||
## Version when you must, not before
|
||||
|
||||
Add fields freely; removing or renaming is a break. If you are not yet committed, say so with
|
||||
a `beta` or `v0` marker rather than locking a shape you have not used. Document the contract
|
||||
you guarantee, not the implementation that happens to produce it.
|
||||
@@ -0,0 +1,39 @@
|
||||
---
|
||||
name: ci-cd
|
||||
description: Write or repair CI/CD pipelines and workflow files. Use when a build fails in CI but not locally, when adding a workflow, or when caching, matrix, or deploy steps need design.
|
||||
---
|
||||
|
||||
# CI/CD
|
||||
|
||||
CI is a second machine that does not have your setup. "Works on my machine" means the pipeline
|
||||
is missing something your machine has.
|
||||
|
||||
## Reproduce the environment, not the symptom
|
||||
|
||||
When CI fails and local passes, the difference is the environment: the toolchain version, an
|
||||
uncommitted file, a cache, an env var, the OS. Diff those before touching the code. Read the
|
||||
failing log literally — the first error, not the last, which is usually a downstream echo.
|
||||
|
||||
## Pin everything that can move
|
||||
|
||||
- Toolchain versions (`node`, `bun`, `python`), action versions, base images. `latest`
|
||||
is a build that breaks on a day you did nothing.
|
||||
- Lockfiles go in the repo and the install respects them (`--frozen-lockfile`, `ci`). An
|
||||
install that re-resolves in CI is a different build from the one you tested.
|
||||
|
||||
## Cache the expensive, deterministic part
|
||||
|
||||
Dependencies are the cache; build output usually is not. Key the cache on the lockfile hash so
|
||||
a changed dependency invalidates it. A cache that is too broad serves stale artifacts; too
|
||||
narrow saves nothing.
|
||||
|
||||
## Fail fast, in the right order
|
||||
|
||||
Cheap checks first: lint and typecheck before the test matrix, tests before the deploy. A
|
||||
pipeline that deploys before it verifies publishes the bug it was built to catch.
|
||||
|
||||
## Secrets and deploys
|
||||
|
||||
Secrets live in the CI secret store, never in the file, and are masked in logs. A deploy step
|
||||
is gated: on a tag, on a protected branch, on a manual approval — never on every push. Assume
|
||||
every log line is public and write the pipeline accordingly.
|
||||
@@ -0,0 +1,54 @@
|
||||
---
|
||||
name: commit
|
||||
description: Stage and commit work. Use when asked to commit, or to split existing changes into commits.
|
||||
---
|
||||
|
||||
# Committing
|
||||
|
||||
Never commit unless the user asked. If it is unclear whether they did, ask. A commit is a
|
||||
durable statement about shared history, not a save-point.
|
||||
|
||||
## Look before you stage
|
||||
|
||||
`git_status` and `git_diff` first — read the whole diff you are about to commit. You are
|
||||
looking for three things:
|
||||
|
||||
1. **Changes that are not yours.** Another agent or the user may share this worktree, and
|
||||
`git add .` takes their half-finished work with yours.
|
||||
2. **Files that should never be committed:** `.env`, credentials, keys, large build
|
||||
output, anything a `.gitignore` rule was supposed to catch and did not. Flag these to
|
||||
the user rather than committing them — a committed secret is a secret to rotate.
|
||||
3. **Your own accidents:** debug prints, commented-out code, a stray `TODO`, a file you
|
||||
opened and saved by mistake. Revert them before staging, not in a follow-up commit.
|
||||
|
||||
Stage the specific paths you changed. `git add .` is how unrelated work ends up in a
|
||||
commit that then has to be reverted whole.
|
||||
|
||||
## One commit, one reason
|
||||
|
||||
If the diff does two unrelated things, make two commits. A commit that both fixes a bug
|
||||
and renames a module cannot be reverted, cherry-picked, or bisected usefully. Each commit
|
||||
should pass the tests on its own — a series of broken commits defeats `git bisect`.
|
||||
|
||||
## The message
|
||||
|
||||
Match the repository's existing style — read `git_log` before writing one. Failing that:
|
||||
|
||||
- A subject line under 70 characters, imperative mood, saying what changed: "Fix off-by-one
|
||||
in pagination", not "fixed a bug" or "changes".
|
||||
- A body explaining *why* when the reason is not obvious from the diff. Wrap at 72.
|
||||
- No "as requested", no restating the diff line by line, no emoji unless the repo uses them,
|
||||
no sign-off noise the repo does not already use.
|
||||
|
||||
## Do not
|
||||
|
||||
- Do not `--amend` a commit that has been pushed. Write a new one.
|
||||
- Do not `--no-verify`. If a hook rejects the commit, the hook found something — read it.
|
||||
- Do not `git push` unless asked, and never force-push without being asked explicitly.
|
||||
- Do not commit and then immediately fix it up with a second commit. Get it right, or say
|
||||
what is wrong.
|
||||
|
||||
## After committing
|
||||
|
||||
Report the short hash and the subject. If a hook rewrote files, say so, and confirm the
|
||||
final state — `git_status` again — is what was intended.
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
name: data
|
||||
description: Process, validate, or transform data. Use when parsing files, cleaning datasets, designing a data pipeline, or debugging a transform that produces wrong output.
|
||||
---
|
||||
|
||||
# Data
|
||||
|
||||
Bad data fails silently and far away from where it entered. Validate at the boundary, keep the
|
||||
raw, and make every transform checkable.
|
||||
|
||||
## Validate at the boundary
|
||||
|
||||
Parse and validate when data enters the system, not when it is used. A schema check at the edge
|
||||
turns a corrupt record into a clear rejection; skipping it turns the same record into a wrong
|
||||
answer three layers later. Reject loudly, with the record and the reason — never coerce and
|
||||
carry on.
|
||||
|
||||
## Keep the raw
|
||||
|
||||
Store the untransformed input alongside the derived. When a transform turns out to be wrong,
|
||||
the raw lets you recompute; without it, the information is gone. Derived data is rebuildable;
|
||||
source data is not.
|
||||
|
||||
## Transformations are pure and tested
|
||||
|
||||
A transform takes input and returns output with no hidden state, so it can be tested on a
|
||||
fixture and re-run safely. Test the edge cases that actually occur in data: the empty field,
|
||||
the wrong type, the unexpected null, the duplicate, the encoding that is not UTF-8.
|
||||
|
||||
## Duplicates, nulls, and ranges are the usual corruption
|
||||
|
||||
Check for: unexpected duplicates on a key, nulls where a value is required, values outside a
|
||||
sane range (a negative age, a date in the future), and referential breaks (an id pointing at
|
||||
nothing). These four catch most real-world data problems before they reach a report.
|
||||
|
||||
## Idempotent pipelines
|
||||
|
||||
A step that can be re-run without duplicating or corrupting its output is a step you can retry
|
||||
after a failure. Key on a stable id and upsert rather than blind-insert. A pipeline you cannot
|
||||
safely re-run is a pipeline you will one day have to fix by hand at 2am.
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
name: db
|
||||
description: Design schemas, write migrations, or fix query and data problems. Use when adding a table, writing a migration, debugging a slow query, or choosing keys and indexes.
|
||||
---
|
||||
|
||||
# Databases
|
||||
|
||||
The schema is the hardest thing to change in the whole system. Design it for the queries, not
|
||||
the object model.
|
||||
|
||||
## Keys and constraints are the real schema
|
||||
|
||||
- Every table has a primary key; prefer a surrogate `id` unless a natural key is truly stable.
|
||||
- Foreign keys and `NOT NULL` are not optional decoration — they are the constraints that stop
|
||||
bad data at the door instead of in application code six months later.
|
||||
- Unique constraints belong on the thing that must be unique (email, slug), enforced by the
|
||||
database, not by a check-then-insert that races.
|
||||
|
||||
## Migrations are one-way and additive where possible
|
||||
|
||||
- Never edit a migration that has run anywhere. Add a new one.
|
||||
- Destructive changes (drop column, rename, change type) are two migrations: add the new shape,
|
||||
deploy code that writes both, then remove the old in a later release. A single migration that
|
||||
renames a column breaks every old copy of the app still running.
|
||||
- Test a migration against real data volume. `ALTER` on ten rows is instant; on ten million it
|
||||
locks the table.
|
||||
|
||||
## Indexes follow the queries
|
||||
|
||||
Index the columns you filter and join on, in the order the query uses them. A composite index
|
||||
`(a, b)` serves `WHERE a` and `WHERE a, b` but not `WHERE b` alone. Read the query plan
|
||||
(`EXPLAIN`) before adding one — a guess is an index that costs writes and serves nothing.
|
||||
|
||||
## The N+1 is the default bug
|
||||
|
||||
A query per row in a loop is the most common database performance defect. Fetch the set with a
|
||||
join or a batched `WHERE id IN (...)`. If a page does one query per item, that is the fix
|
||||
before any caching.
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
name: debug
|
||||
description: Track down a bug whose cause is not obvious. Use when a test fails for unclear reasons, behaviour differs between environments, or an earlier fix did not hold.
|
||||
---
|
||||
|
||||
# Debugging
|
||||
|
||||
Do not guess. A guess that happens to work leaves the real cause in place, and it will
|
||||
fire again — usually in production, usually at a worse time.
|
||||
|
||||
## Reproduce first
|
||||
|
||||
Find the smallest command that shows the failure and record it with `remember`. If you
|
||||
cannot reproduce it, say so and ask what the user did differently — do not proceed on a
|
||||
hypothesis you cannot test.
|
||||
|
||||
Shrink the reproduction until it is minimal: one input, one call, one assertion. Every
|
||||
moving part you leave in is a place the bug can hide. A reproduction that takes thirty
|
||||
steps will not get run often enough to confirm the fix.
|
||||
|
||||
## Three hypotheses, then evidence
|
||||
|
||||
Write down at least three causes that would produce this exact symptom — not "the code is
|
||||
wrong" but specific mechanisms: "the offset is off by one when the page is empty", "the
|
||||
cache is read before the write lands". Rank them by how cheap they are to disprove, then
|
||||
disprove them in that order. State which one you are testing before you test it.
|
||||
|
||||
Evidence means observed output: a log line, a failing assertion, a value printed at the
|
||||
point of failure. "It should be X" is not evidence. When the evidence contradicts your
|
||||
favoured hypothesis, the hypothesis is wrong — do not explain the evidence away.
|
||||
|
||||
## Localise before you fix
|
||||
|
||||
Assert the value at each boundary until one is wrong. The bug lives between the last
|
||||
boundary where the value is right and the first where it is wrong. Fixing before you have
|
||||
that bracket means editing the wrong place and learning nothing.
|
||||
|
||||
## Bisect when the space is large
|
||||
|
||||
- Recent regression: `git bisect` or read what changed last. The bug arrived in a commit;
|
||||
find which one.
|
||||
- Unclear layer: assert the value at each boundary until one is wrong.
|
||||
- Intermittent: run it in a loop and capture the failing case with full logging. Do not
|
||||
reason about a race abstractly — make it happen on demand, then it is no longer
|
||||
intermittent.
|
||||
|
||||
## Fix the cause, not the symptom
|
||||
|
||||
Once you know the cause, fix that and nothing else. Do not tidy surrounding code in the
|
||||
same change — a bugfix diff should contain only the bug, so it can be reverted whole if it
|
||||
is wrong.
|
||||
|
||||
Write a test that fails before the fix and passes after. Watch it fail first; a test you
|
||||
never saw fail proves nothing. If you cannot express the bug as a test, say why — and say
|
||||
what you ran instead to confirm the fix.
|
||||
|
||||
## After two failed attempts
|
||||
|
||||
Stop. Re-read the error text literally, character by character — most "impossible" bugs
|
||||
are a misread message. Then check your assumption about which code is actually running:
|
||||
the wrong file, a stale build, a shadowed import, a cached dependency, or an env var that
|
||||
differs from your shell. Verify by printing something at the point you *think* executes;
|
||||
if it does not print, that is your answer.
|
||||
@@ -0,0 +1,39 @@
|
||||
---
|
||||
name: deps
|
||||
description: Manage dependencies: choosing, adding, updating, or removing them. Use when evaluating a library, resolving a version conflict, pruning unused deps, or hardening the supply chain.
|
||||
---
|
||||
|
||||
# Dependencies
|
||||
|
||||
Every dependency is code you did not write but now maintain. Add deliberately, prune regularly.
|
||||
|
||||
## Choose on maintenance, not features
|
||||
|
||||
Before adding: is it actively maintained (recent commits, responsive issues), widely used, and
|
||||
small enough to be worth it? A dependency that saves a day and is abandoned in a year costs a
|
||||
week. For something small and stable, a dozen lines in your own codebase often beats a package.
|
||||
|
||||
## Pin and lock
|
||||
|
||||
Exact versions in the manifest for anything that matters, a lockfile committed, and installs
|
||||
that respect it. A `^` range means your build tomorrow differs from your build today. The
|
||||
lockfile is the build's memory; do not delete it to "fix" a conflict — resolve the conflict.
|
||||
|
||||
## Update on a schedule, read the changelog
|
||||
|
||||
Routine small updates beat a yearly painful one. For a major bump: read the changelog and the
|
||||
migration guide, find every call site of the changed API, and apply one shape of change (see
|
||||
the migrate skill). Update one thing at a time so a regression has an obvious cause.
|
||||
|
||||
## Know your transitive tree
|
||||
|
||||
A direct dependency drags in dozens of transitive ones. Audit the tree for: known
|
||||
vulnerabilities (`audit`/SCA tooling), abandoned packages deep in it, and duplicate copies of
|
||||
the same library at different versions bloating the bundle. Remove what you no longer use — an
|
||||
unused dependency is attack surface and install time for nothing.
|
||||
|
||||
## Supply chain is a trust decision
|
||||
|
||||
A package runs its install scripts with your permissions. Prefer packages with provenance and a
|
||||
reproducible build, be wary of sudden ownership transfers, and pin so a hijacked publish does
|
||||
not reach you automatically. The lockfile is also your audit trail of exactly what shipped.
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
name: docker
|
||||
description: Write or fix Dockerfiles and container setups. Use when an image is too large, a build is slow, a container will not start, or layering and caching need design.
|
||||
---
|
||||
|
||||
# Docker
|
||||
|
||||
An image is a build artifact. Small, reproducible, and boring is the goal.
|
||||
|
||||
## Layer cache is the whole speed game
|
||||
|
||||
Order instructions from least to most frequently changed: base image, then dependency
|
||||
manifests, then `install`, then source copy, then build. Copying `.` before installing
|
||||
dependencies means every code change re-runs the install — the single most common Dockerfile
|
||||
mistake.
|
||||
|
||||
## Small images, on purpose
|
||||
|
||||
- Use multi-stage builds: build in a full toolchain stage, copy only the artifact into a slim
|
||||
runtime stage. The compiler does not ship to production.
|
||||
- Pick a slim or distroless base unless you need the tooling. Alpine is small but musl breaks
|
||||
some binaries; know why you chose it.
|
||||
- One `RUN` with `&&` for related steps, cleaning up in the same layer — a separate `RUN rm`
|
||||
does not shrink the image, the data is still in the earlier layer.
|
||||
|
||||
## The container is not a VM
|
||||
|
||||
- One process per container, as PID 1, so signals work. Use an init if the app spawns children.
|
||||
- Do not run as root. Add a user and `USER` it.
|
||||
- Read-only filesystem where possible; write to a mounted volume for anything that must persist.
|
||||
Nothing in the image is writable state.
|
||||
|
||||
## .dockerignore is as important as the Dockerfile
|
||||
|
||||
Exclude `.git`, `node_modules`, build output, and any secret file. A context that sends the
|
||||
whole repo is slow, and a secret copied into an image layer is a secret to rotate.
|
||||
|
||||
## Healthcheck and logs
|
||||
|
||||
The process logs to stdout/stderr, never to a file inside the container — the runtime collects
|
||||
it. Add a `HEALTHCHECK` that proves the service answers, not just that the process exists.
|
||||
@@ -0,0 +1,39 @@
|
||||
---
|
||||
name: docs
|
||||
description: Write or update documentation, READMEs, and guides. Use when asked to document a feature, write usage docs, or bring docs back in line with the code.
|
||||
---
|
||||
|
||||
# Documentation
|
||||
|
||||
Docs lie by omission. Write only what you have verified in the code.
|
||||
|
||||
## Ground every claim in the source
|
||||
|
||||
Before documenting a behaviour, read it. A flag, a default, an error message — open the
|
||||
code and quote what it actually does, not what the name suggests. The most damaging doc
|
||||
line is the confident one that was true two versions ago. If the code and the existing
|
||||
docs disagree, the code is right; say so and fix the doc.
|
||||
|
||||
## Answer the reader's actual question
|
||||
|
||||
A reader opens a doc with a task, not a desire for completeness. Lead with the thing they
|
||||
came to do, in the order they will do it:
|
||||
|
||||
- **A reference** lists what exists: every flag, every field, with its default and its type.
|
||||
- **A guide** walks one path to one outcome. Resist documenting every branch — link instead.
|
||||
- **A README** orients in sixty seconds: what it is, install, the first command that works.
|
||||
|
||||
## Show, then say
|
||||
|
||||
A working example beats a paragraph about one. Every command in the doc must be one you
|
||||
ran, with its real output. A snippet that was never executed is a bug waiting for a reader.
|
||||
|
||||
## Match the house style
|
||||
|
||||
Read the neighbouring docs first: their heading depth, their code-fence language tags,
|
||||
their tone. A doc that reads foreign is a doc nobody trusts enough to maintain.
|
||||
|
||||
## Keep it true over time
|
||||
|
||||
Document the stable contract, not the current implementation, unless the point is the
|
||||
implementation. The fewer specifics a doc pins down, the less it rots.
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
name: frontend
|
||||
description: Build or fix a web UI. Use when working on components, state, rendering performance, forms, or anything the user sees and interacts with in a browser.
|
||||
---
|
||||
|
||||
# Frontend
|
||||
|
||||
The user's experience is the metric. Fast, clear, and forgiving beats clever.
|
||||
|
||||
## State lives as low as it can
|
||||
|
||||
Lift state only as high as the components that share it. Global state for something two siblings
|
||||
need is re-render and complexity for everything. Server data is not client state — cache it with
|
||||
the data layer rather than duplicating it into a store you must keep in sync by hand.
|
||||
|
||||
## Rendering is the usual bottleneck
|
||||
|
||||
Before optimising, find what re-renders. A component that re-renders on every parent render
|
||||
because of an inline object or function prop is the common case. Memoize the expensive subtree,
|
||||
not everything — `useMemo` and `useCallback` have a cost too, and slapping them everywhere is
|
||||
its own slowdown.
|
||||
|
||||
## Forms respect the user
|
||||
|
||||
- Validate on blur or submit, not on every keystroke, and show the message at the field.
|
||||
- Never clear a form on an error. The user's input is the most expensive thing on the page.
|
||||
- Disable the submit while submitting, and say what is happening. A double-submitted form is a
|
||||
duplicate record.
|
||||
|
||||
## Accessibility is not a later pass
|
||||
|
||||
Semantic HTML first: a `<button>` that looks like a button beats a `<div>` with a click
|
||||
handler. Keyboard-reachable everything, visible focus, labels on inputs, alt text that conveys
|
||||
the point not the pixels. Colour is never the only carrier of meaning.
|
||||
|
||||
## Measure what the user feels
|
||||
|
||||
Load: get the first meaningful paint and the time-to-interactive down before micro-tuning.
|
||||
Bundle: split the route nobody opens, lazy-load the heavy component. A Lighthouse number is a
|
||||
proxy; the goal is that it never feels slow.
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
name: git-workflow
|
||||
description: Work with branches, rebases, merges, and history. Use when untangling a branch, preparing a PR, deciding rebase vs merge, or recovering from a git mistake.
|
||||
---
|
||||
|
||||
# Git workflow
|
||||
|
||||
History is a communication tool. Write it for the person who reads it in six months — usually you.
|
||||
|
||||
## One branch, one purpose
|
||||
|
||||
A branch that does two things produces a PR that can only be reviewed as all-or-nothing and
|
||||
reverted only whole. Keep it small and single-purpose; open the second thing as its own branch.
|
||||
|
||||
## Rebase to clean up, merge to preserve
|
||||
|
||||
- Rebase your own unpushed work freely: it makes a linear, readable history.
|
||||
- Never rebase a branch others have pulled — it rewrites commits they have, and the next pull
|
||||
becomes a mess. Merge shared branches instead.
|
||||
- Interactive rebase before opening the PR: squash the "fix typo" and "wip" commits into the
|
||||
change they belong to. The PR should read as a series of intentional steps, not a diary.
|
||||
|
||||
## Recover without panic
|
||||
|
||||
- `git reflog` finds almost anything you "lost": the branch you deleted, the commit you reset
|
||||
away. Nothing committed is truly gone for ~30 days.
|
||||
- A bad merge: `git merge --abort`. A bad rebase: `git rebase --abort`. Both stop cleanly
|
||||
rather than pushing forward into a worse state.
|
||||
- Committed to the wrong branch: `git reset --soft` to keep the work, switch, recommit.
|
||||
|
||||
## The commit message is the review's first page
|
||||
|
||||
Subject under 70 chars, imperative, says what changed. Body explains *why* when it is not
|
||||
obvious. A reviewer who cannot tell why a change exists from its message will ask, or worse,
|
||||
approve without understanding.
|
||||
|
||||
## Read the conflict, do not guess
|
||||
|
||||
On a conflict, open the file and understand both sides before resolving. Taking "ours" or
|
||||
"theirs" wholesale because it is faster is how a resolved conflict silently drops someone's
|
||||
work.
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
name: i18n
|
||||
description: Internationalise or localise a product. Use when extracting strings for translation, formatting dates and numbers for a locale, handling pluralisation, or fixing layout that breaks in another language.
|
||||
---
|
||||
|
||||
# Internationalisation
|
||||
|
||||
Hard-coded English is a bug in every other language. Externalise strings and never assume a
|
||||
grammar.
|
||||
|
||||
## Every user-facing string is a key
|
||||
|
||||
No string in the UI lives in code; it lives in a message catalogue under a key. Concatenating
|
||||
translated fragments is the classic bug: "You have " + n + " messages" cannot be reordered for
|
||||
a language whose grammar puts the number elsewhere. Use a format with named placeholders:
|
||||
`{count, plural, ...}`, translated as a whole.
|
||||
|
||||
## Pluralisation and gender are not English
|
||||
|
||||
Languages have one, two, several, or no plural forms, with rules that do not map to "1 vs other".
|
||||
Use the ICU plural machinery of your i18n library and let the translator fill in every form the
|
||||
locale needs. The same goes for gendered agreement.
|
||||
|
||||
## Format dates, numbers, and currencies by locale
|
||||
|
||||
Never `dd/mm/yyyy` by hand: `03/04/2025` is March 4th to one user and April 3rd to another.
|
||||
Use the platform's locale-aware formatter (`Intl.DateTimeFormat`, `Intl.NumberFormat`).
|
||||
Store and transmit ISO 8601 / UTC; format for display only.
|
||||
|
||||
## Layout breaks in translation
|
||||
|
||||
German runs ~30% longer than English; some scripts are right-to-left. Flexible layout, no fixed
|
||||
widths on translated text, and CSS logical properties (`margin-inline-start` not
|
||||
`margin-left`) so RTL mirrors correctly. Test with a pseudo-locale that lengthens and accents
|
||||
every string to find the overflows before a translator does.
|
||||
|
||||
## Sort and search correctly
|
||||
|
||||
String order is locale-dependent: `ä` sorts with `a` in German, after `z` in Swedish. Use
|
||||
locale-aware collation (`Intl.Collator` or the database's) rather than byte order, and normalise
|
||||
Unicode before comparing, because the same character has more than one byte representation.
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
name: incident
|
||||
description: Respond to a production incident. Use when something is down, degraded, or misbehaving in production and must be diagnosed and mitigated under time pressure.
|
||||
---
|
||||
|
||||
# Incident response
|
||||
|
||||
Mitigate first, diagnose second. Restore service, then find out why.
|
||||
|
||||
## Confirm and scope before touching anything
|
||||
|
||||
What is actually broken, for whom, since when? Check the signal, not the report: the dashboard,
|
||||
the error rate, the health endpoint. A wrong scope sends you chasing a symptom. State the impact
|
||||
plainly in one line before you start changing things.
|
||||
|
||||
## Recent change is the prime suspect
|
||||
|
||||
Most incidents follow a deploy, a config change, a flag flip, or a scaling event. What changed
|
||||
in the window before it broke? Check the deploy log and the diff. The fastest fix is usually to
|
||||
undo the last change, not to understand it.
|
||||
|
||||
## Mitigate, then understand
|
||||
|
||||
- Roll back the deploy, flip the flag off, fail over, scale up, restart the wedged process —
|
||||
whichever restores service fastest, even if you do not yet know the root cause.
|
||||
- A mitigation you can reverse beats a perfect diagnosis that takes an hour. Note what you did so
|
||||
it can be undone or made permanent later.
|
||||
- Do not deploy an unreviewed "fix" into the fire; it adds a second change to a system already
|
||||
misbehaving.
|
||||
|
||||
## Preserve evidence before it rotates away
|
||||
|
||||
Capture the logs, the error, the relevant metrics, a snapshot of the state — before a restart or
|
||||
a rollback destroys it. You will want it for the postmortem, and it may be the only copy.
|
||||
|
||||
## Communicate and follow up
|
||||
|
||||
Say what is broken, what you are doing, and when the next update is — to whoever is affected,
|
||||
in plain language, on a schedule. Afterwards: write the timeline, the root cause, and the
|
||||
follow-ups that stop it recurring. An incident with no follow-up is a loan against the next one.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: logging
|
||||
description: Add or improve logging and observability. Use when debugging in production, adding structured logs, choosing log levels, or making a system traceable.
|
||||
---
|
||||
|
||||
# Logging
|
||||
|
||||
Logs are how you debug a system you cannot attach a debugger to. Write them for the 3am
|
||||
incident, not the happy path.
|
||||
|
||||
## Structure over prose
|
||||
|
||||
Emit fields, not sentences: `{ user: id, action: "checkout", ms: 142, ok: false }`, not
|
||||
`"User checked out"`. Structured logs are searchable and aggregable; a sentence is neither.
|
||||
One event, one line, one level.
|
||||
|
||||
## Levels are a contract
|
||||
|
||||
- `error` — something is broken and someone should look. Not "a user gave bad input".
|
||||
- `warn` — unexpected but handled; worth a glance.
|
||||
- `info` — the meaningful state transitions: started, finished, the decision made. Sparse.
|
||||
- `debug` — everything you might want while diagnosing, off in production.
|
||||
|
||||
A log at the wrong level trains people to ignore the right one. If everything is `error`,
|
||||
nothing is.
|
||||
|
||||
## Log the decision points, not every line
|
||||
|
||||
At a boundary — a request in, a call out, a branch taken — log what was decided and the inputs
|
||||
that decided it, with a correlation id that follows the request across services. You should be
|
||||
able to trace one request end to end from the id alone.
|
||||
|
||||
## Never log a secret
|
||||
|
||||
No passwords, tokens, session ids, full card numbers, or personal data beyond what policy
|
||||
allows. Redact at the point of logging, not by hoping a downstream filter catches it. A secret
|
||||
in a log aggregator is a secret to rotate.
|
||||
|
||||
## Measure, do not just log
|
||||
|
||||
For anything with a latency or a rate, a metric answers "is it slow?" faster than a thousand
|
||||
log lines. Logs explain *why*; metrics tell you *that* something is wrong in the first place.
|
||||
@@ -0,0 +1,55 @@
|
||||
---
|
||||
name: migrate
|
||||
description: Upgrade a dependency, framework, or language version across a codebase. Use when a major version bump, a deprecation, or a breaking API change has to be applied.
|
||||
---
|
||||
|
||||
# Migration
|
||||
|
||||
The failure mode is a half-applied migration: it compiles, most tests pass, and one code
|
||||
path still uses the old API.
|
||||
|
||||
## Read the changelog before the code
|
||||
|
||||
Find what actually broke. A major version usually has a migration guide; read it and list
|
||||
the changes that apply to this codebase specifically. Below 1.0, treat a minor bump as
|
||||
breaking — semver promises nothing there.
|
||||
|
||||
## Find every call site before changing one
|
||||
|
||||
Grep for the old API across the whole repository, including tests, scripts, config, CI
|
||||
workflows, Dockerfiles, and documentation. A version literal pinned in a workflow while the
|
||||
manifest says something else is a split-brain deploy.
|
||||
|
||||
Write the list down with `todo_write`. The list is the migration; the edits are mechanical.
|
||||
|
||||
## Change in one shape
|
||||
|
||||
Apply the same transformation everywhere rather than improving each site as you pass
|
||||
through it. A migration mixed with refactoring cannot be reviewed, and cannot be reverted
|
||||
if the upgrade turns out to be wrong.
|
||||
|
||||
`apply_patch` is the tool for this: one atomic patch across the files that must land
|
||||
together.
|
||||
|
||||
## Verify at the boundary that broke
|
||||
|
||||
Type checks catch signature changes and miss behaviour changes — the two ways a migration
|
||||
actually breaks you. Run the tests, then actually *use* the thing that was upgraded: start
|
||||
the server, run the CLI, execute the query, hit the endpoint. A green suite over an
|
||||
untested upgrade path proves only that the suite did not cover it.
|
||||
|
||||
Pay special attention to silent behaviour changes: a default that flipped, a deprecated
|
||||
call that still runs but does something subtly different, an error type that changed shape.
|
||||
These compile, pass type checks, and still break production.
|
||||
|
||||
## Never hand-merge a lockfile
|
||||
|
||||
On a conflict, take either side whole and regenerate with the package manager. The resolver
|
||||
owns that file; a hand-merge is a split-brain dependency tree that installs differently on
|
||||
every machine.
|
||||
|
||||
## Report
|
||||
|
||||
The version before and after, every file class touched, what you verified by running, the
|
||||
behaviour changes you checked by hand, and anything the changelog said applies that you
|
||||
deliberately did not do — with the reason.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: onboarding
|
||||
description: Orient in an unfamiliar codebase. Use when dropped into a new project and asked to understand it, or when writing the docs that help a newcomer get productive.
|
||||
---
|
||||
|
||||
# Onboarding to a codebase
|
||||
|
||||
Understand the running system before the source. The goal is a correct mental model, not to have
|
||||
read every file.
|
||||
|
||||
## Get it running first
|
||||
|
||||
Build it, run it, run the tests. A project you can execute you can interrogate; one you have
|
||||
only read you can only guess at. The README and the `package.json`/`Makefile` scripts tell you
|
||||
the intended commands; if they do not work, that is your first finding.
|
||||
|
||||
## Trace one request end to end
|
||||
|
||||
Pick the central thing the system does and follow it: the entry point, the route or main, the
|
||||
handler, the data out and back. One full path teaches you the architecture faster than reading
|
||||
any single module. Note the layers you cross — that is the system's real structure.
|
||||
|
||||
## Read the structure, not the files
|
||||
|
||||
- The directory layout names the major components and their boundaries.
|
||||
- The dependency manifest names the frameworks and the big choices already made.
|
||||
- The tests show what the code is supposed to do, often better than the code does.
|
||||
- `git log` on a core file shows what changes often and why — the living parts versus the
|
||||
stable ones.
|
||||
|
||||
## Map the seams
|
||||
|
||||
Where does data enter and leave (HTTP, a queue, a file)? Where is state kept (a database,
|
||||
memory, a cache)? Where are the trust boundaries? Those are the places bugs and features both
|
||||
live. You do not need to know every file; you need to know where a change of a given kind would
|
||||
go.
|
||||
|
||||
## Ask the codebase, then a person
|
||||
|
||||
Grep and the outline/symbol tools answer most "where is X" faster than reading. When genuinely
|
||||
stuck on *why* something exists — that is a question for a person or the history, not more
|
||||
reading.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: optimize-sql
|
||||
description: Diagnose and fix a slow SQL query. Use when a query is slow, a page makes too many queries, or an execution plan needs reading.
|
||||
---
|
||||
|
||||
# SQL optimisation
|
||||
|
||||
Read the plan before changing anything. `EXPLAIN` (or `EXPLAIN ANALYZE`) tells you what the
|
||||
database actually does; guessing at it is how you add an index that helps nothing.
|
||||
|
||||
## Read the plan for the expensive node
|
||||
|
||||
Find the node with the highest cost: a sequential scan over a large table, a nested loop over
|
||||
many rows, a sort that spills to disk. Optimise that node. A plan with ten cheap nodes and one
|
||||
expensive one has exactly one thing to fix.
|
||||
|
||||
## The index that matches the query
|
||||
|
||||
- Index the columns in the `WHERE` and `JOIN` clauses, and for a sort, the `ORDER BY`.
|
||||
- A composite index `(a, b, c)` serves a leftmost prefix: `a`, `a,b`, `a,b,c` — not
|
||||
`b` alone. Order the columns by the equality filters first, then the range, then the sort.
|
||||
- A covering index includes every column the query reads, so the table is never touched. That
|
||||
is the fastest a read gets.
|
||||
|
||||
## Write the query so the index is usable
|
||||
|
||||
- `WHERE lower(email) = ...` cannot use a plain index on `email`; either store it lowered or
|
||||
use a functional index. A function on the column defeats the index.
|
||||
- Leading `LIKE '%x'` cannot use a B-tree index; `LIKE 'x%'` can.
|
||||
- `OR` across different columns often defeats an index; `UNION` of two indexed queries can be
|
||||
faster.
|
||||
|
||||
## Kill the N+1 first
|
||||
|
||||
Before any index: if the page runs one query per row, that is the fix. Batch with
|
||||
`WHERE id IN (...)` or a join. Ten queries become one beats ten individually-fast queries.
|
||||
|
||||
## Measure the change
|
||||
|
||||
`EXPLAIN ANALYZE` before and after, against realistic data volume. An index that helps a
|
||||
100-row table may not justify its write cost at 100 million. Report the timing you actually
|
||||
measured, not the improvement you expected.
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
name: perf-frontend
|
||||
description: Make a web page faster. Use when a page loads slowly, feels janky, fails Core Web Vitals, or ships too much JavaScript.
|
||||
---
|
||||
|
||||
# Frontend performance
|
||||
|
||||
Measure the user's experience first: a Lighthouse lab score and, better, real-user data. Optimise
|
||||
the metric that is actually failing, not the one easiest to move.
|
||||
|
||||
## The vitals and what drives them
|
||||
|
||||
- **LCP** (largest contentful paint) — almost always the hero image or a web font. Preload it,
|
||||
size it correctly, serve it in a modern format, and do not let render-blocking resources delay it.
|
||||
- **INP / responsiveness** — long tasks on the main thread. Break up work, defer non-urgent JS,
|
||||
and keep event handlers fast. A click that responds in 50ms feels instant; 300ms feels broken.
|
||||
- **CLS** (layout shift) — images and embeds without dimensions, late-injected banners, web fonts
|
||||
swapping. Reserve the space before the content arrives.
|
||||
|
||||
## Ship less JavaScript
|
||||
|
||||
The bundle is usually the problem. Route-level code splitting so a page loads only what it needs,
|
||||
lazy-load the heavy below-the-fold component, and audit the dependency tree for a large library
|
||||
imported for one function. Removing 100KB of JS beats most micro-optimisations.
|
||||
|
||||
## Network discipline
|
||||
|
||||
- Cache static assets with long, content-hashed lifetimes; the second visit should cost almost
|
||||
nothing.
|
||||
- Compress (brotli/gzip) and serve images at the size they are displayed, responsive `srcset`,
|
||||
not a 3000px original in a 300px slot.
|
||||
- Fetch in parallel, not in waterfalls: start independent requests together, and preload the
|
||||
critical few.
|
||||
|
||||
## Change one thing, measure it
|
||||
|
||||
Take a baseline (the metric, the page, the device class), make one change, re-measure on the same
|
||||
setup. Two changes at once and you do not know which paid. Report the before/after you actually
|
||||
measured, on a realistic device and connection — a developer's fast laptop and fiber hides what a
|
||||
mid-range phone on 4G feels.
|
||||
@@ -0,0 +1,49 @@
|
||||
---
|
||||
name: perf
|
||||
description: Make something faster, or find out why it is slow. Use when a command, request, test suite, or build takes longer than it should.
|
||||
---
|
||||
|
||||
# Performance
|
||||
|
||||
Measure first. A change made without a number before it is a guess with extra steps, and
|
||||
most "optimisations" made on a guess make the code worse and no faster.
|
||||
|
||||
## Get a number
|
||||
|
||||
Time the actual operation, not a proxy for it: `time`, the framework's own timing output,
|
||||
or a loop around the slow call with a timestamp either side. Use realistic input — a fast
|
||||
result on a tiny fixture tells you nothing about the production case. Record the baseline
|
||||
with `remember` so the comparison survives compaction, and run it enough times that a
|
||||
warm cache and jitter do not fool you.
|
||||
|
||||
If you cannot measure it, say so and stop. Optimising an unmeasured path is how a codebase
|
||||
accumulates complexity that buys nothing.
|
||||
|
||||
## Find where the time goes
|
||||
|
||||
- **Wall-clock dominated by one call?** Look there and nowhere else. The biggest node is
|
||||
the only one worth touching.
|
||||
- **Spread evenly?** Suspect the loop around it: an O(n²) walk, a query per row, a file
|
||||
read per iteration, an allocation per element.
|
||||
- **Idle time?** It is waiting, not computing: a sequential chain of independent awaits, an
|
||||
unpooled connection, a contended lock, a slow remote call.
|
||||
|
||||
The usual culprits, in the order they actually appear: N+1 queries, work repeated inside a
|
||||
loop that could be hoisted, a missing index, sequential awaits that could run together,
|
||||
reading a whole file to use one line, and re-parsing something that could be parsed once.
|
||||
|
||||
## Change one thing
|
||||
|
||||
One change, then re-measure on the same setup. Two changes together and you do not know
|
||||
which one paid — and one of them may have cost. If the number did not move, revert the
|
||||
change; an optimisation that does not measure is just complexity.
|
||||
|
||||
## Stop when it is fast enough
|
||||
|
||||
State the target before you start: "the test suite under a minute", "the endpoint under
|
||||
200ms". Past the target, further work is complexity with no user on the other end of it.
|
||||
|
||||
## Report
|
||||
|
||||
Baseline, the change, the new number, and what you deliberately did not do. A 40% win with
|
||||
one line changed is a better report than a 45% win that restructured a module.
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
name: plan
|
||||
description: Break a non-trivial task into an ordered, verifiable sequence before writing code. Use when a request is large, spans several files, or its steps depend on each other.
|
||||
---
|
||||
|
||||
# Planning
|
||||
|
||||
A plan that cannot be checked is a wish. Every step ends in something you can run.
|
||||
|
||||
## Understand before you sequence
|
||||
|
||||
Read enough to know the real shape of the work: the entry point, the data's path, the
|
||||
module that owns the behaviour. A plan made from filenames alone reorders itself the
|
||||
moment you open the first file. Grep the actual call sites; do not plan around a guess.
|
||||
|
||||
## Order by dependency, not by file
|
||||
|
||||
A step may depend on another's output: a type before its callers, a schema before its
|
||||
migration, a test helper before the tests that use it. Sequence so nothing references
|
||||
what does not exist yet. If two steps are independent, say so — the order between them
|
||||
is free and you may take the cheaper one first to derisk the rest.
|
||||
|
||||
## One step, one verifiable outcome
|
||||
|
||||
Each step names the command that proves it done: a test that passes, a build that
|
||||
compiles, a script that runs. "Wire it up" is not a step. Write the list with
|
||||
`todo_write`, then work it in order, marking done immediately — not in a batch at the end.
|
||||
|
||||
## Keep it small
|
||||
|
||||
The plan is a scaffold, not the building. If a step grows past "change these few files",
|
||||
split it. If the task turns out smaller than it looked, drop the remaining steps and say
|
||||
why rather than inventing work to fill them.
|
||||
|
||||
## Replan when the ground moves
|
||||
|
||||
New information that changes the order or the scope is a reason to rewrite the list, not
|
||||
to push through it. A stale plan followed faithfully is worse than no plan.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: readme
|
||||
description: Write or fix a project README. Use when creating a README, when a new user cannot get the project running from it, or when it has drifted from the code.
|
||||
---
|
||||
|
||||
# README
|
||||
|
||||
A README has sixty seconds to answer: what is this, do I want it, and how do I run it. Everything
|
||||
else is secondary to those three.
|
||||
|
||||
## The first screen answers three questions
|
||||
|
||||
1. **What it is** in one or two sentences, concrete about the problem it solves — not "a modern
|
||||
solution" but "a CLI that lints Terraform plans against your org's policies".
|
||||
2. **Install** — the one command that gets it.
|
||||
3. **The first thing that works** — the minimal command or snippet that produces visible output.
|
||||
If a new user cannot get a win in two minutes, most leave.
|
||||
|
||||
## Verify every command
|
||||
|
||||
Run each command in the README against a clean environment and paste its real output. The most
|
||||
common README defect is an install or quickstart that no longer works because the code moved and
|
||||
the doc did not. If you cannot run it, do not write it.
|
||||
|
||||
## Structure for scanning
|
||||
|
||||
After the quickstart, in the order a new user needs them: features as a short list of what it
|
||||
does (not how), the common tasks as copy-paste examples, configuration as a table of options with
|
||||
defaults, then links to deeper docs. Headings let a reader jump; a wall of prose gets skimmed
|
||||
past the thing they needed.
|
||||
|
||||
## Show, do not tell
|
||||
|
||||
A three-line example of real use beats a paragraph describing capability. Show the input and the
|
||||
output. A screenshot or asciinema of the actual tool running is worth a hundred adjectives —
|
||||
include one if the tool has any visual surface.
|
||||
|
||||
## Keep it true
|
||||
|
||||
Document the stable interface, not this week's implementation, or the README rots. Re-read it on
|
||||
every release: a README that contradicts the current version is worse than a short one, because
|
||||
it actively misleads.
|
||||
@@ -0,0 +1,44 @@
|
||||
---
|
||||
name: refactor
|
||||
description: Restructure code without changing behaviour. Use when asked to refactor, clean up, extract, or reorganise.
|
||||
---
|
||||
|
||||
# Refactoring
|
||||
|
||||
Behaviour must not change. That is the whole constraint — every other goal (clarity,
|
||||
structure, naming) is subordinate to it. The moment behaviour changes, you are no longer
|
||||
refactoring, you are editing, and the safety argument below stops holding.
|
||||
|
||||
## Establish the safety net first
|
||||
|
||||
Run the existing tests and record that they pass — with `remember`, so the baseline
|
||||
survives compaction. If the code has no tests, write one that pins current behaviour,
|
||||
*including the ugly parts*: the odd return value, the quirk callers depend on. You are not
|
||||
judging the behaviour, you are freezing it. Refactoring untested code is not refactoring;
|
||||
it is rewriting, and it belongs under the edit workflow with its own verification.
|
||||
|
||||
## Then move in small steps
|
||||
|
||||
One transformation at a time, tests green between each. Rename, then extract, then move —
|
||||
not all three in one edit. The mechanical refactorings are the safe ones: rename, extract
|
||||
function, inline, move. Compose them. A large refactor that fails leaves you unable to
|
||||
tell which of five steps broke it; a small one that fails tells you exactly which.
|
||||
|
||||
After each step, run the tests, not just the typechecker. Types catch signature drift;
|
||||
they do not catch a reordered conditional or a dropped early return.
|
||||
|
||||
## What not to do
|
||||
|
||||
- Do not fix bugs while refactoring. Note them, finish the refactor green, then fix in a
|
||||
separate change — otherwise a regression could be either the refactor or the fix.
|
||||
- Do not add abstraction for a single caller. Duplication beats a premature interface;
|
||||
the third caller is when the abstraction earns its name.
|
||||
- Do not widen the scope. The request was this code, not its neighbours. A refactor that
|
||||
"while we're here" touches five more files is five more files of unreviewable risk.
|
||||
- Do not change public API unless asked; if it must change, say so first and update every
|
||||
caller in the same change.
|
||||
|
||||
## Done means
|
||||
|
||||
Tests pass, behaviour is identical, and the diff is smaller than the reader feared. If the
|
||||
diff is larger than the code it moved, you abstracted too early — put it back.
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
name: release
|
||||
description: Cut a release: versioning, changelogs, tagging, publishing. Use when asked to release, bump a version, write release notes, or fix a broken publish.
|
||||
---
|
||||
|
||||
# Release
|
||||
|
||||
A release is a promise that a specific, identified state of the code works. Make it
|
||||
reproducible or do not make it.
|
||||
|
||||
## Version says what changed
|
||||
|
||||
Semver: breaking is a major, a feature is a minor, a fix is a patch. The number is a message to
|
||||
whoever upgrades, not a marketing choice. Below 1.0, say so plainly — semver promises nothing
|
||||
and the version should not pretend otherwise.
|
||||
|
||||
## The changelog is for the upgrader
|
||||
|
||||
- Group by what the reader must do: breaking changes and required actions first, then features,
|
||||
then fixes.
|
||||
- Write it as "you can now X" or "Y no longer Z", from the user's side, not the commit's. A
|
||||
changelog that is a git log is a changelog nobody reads.
|
||||
- Every breaking change names the migration: what to change to keep working.
|
||||
|
||||
## Verify before you tag
|
||||
|
||||
The release candidate builds clean from a fresh checkout, the tests pass, and the version string
|
||||
in the source matches the tag you are about to push. A version/tag mismatch published is the
|
||||
kind of thing that ships "0.4" labelled as "0.3" forever.
|
||||
|
||||
## Tag the commit, publish the artifact
|
||||
|
||||
Tag the exact commit that was verified, and build the artifact from that tag — not from a
|
||||
working tree that has since moved. The tag is immutable; never move it to a different commit.
|
||||
If a release is wrong, cut a new one with a new number; do not quietly re-tag.
|
||||
|
||||
## If it goes wrong
|
||||
|
||||
Have the rollback ready before you need it: the previous artifact still available, the deploy
|
||||
reversible. A bad release is fixed forward with a patch release, not by deleting the evidence.
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
name: review
|
||||
description: Review a diff or a file for defects. Use when asked to review, critique, or check code before it ships.
|
||||
---
|
||||
|
||||
# Code review
|
||||
|
||||
Severity order. Do not lead with style — a review that opens on naming while a real bug
|
||||
sits three lines down has failed at its one job.
|
||||
|
||||
1. **Incorrect behaviour** — wrong result, wrong edge case, wrong state after failure.
|
||||
2. **Missing validation at trust boundaries** — user input, network responses, file contents,
|
||||
anything crossing a process line. Internal calls need no defensive checks.
|
||||
3. **Security** — injection, path traversal, secrets in logs or errors, missing authz.
|
||||
4. **Resource handling** — unclosed handles, unbounded growth, unawaited promises.
|
||||
5. **Clarity** — only when it will cause a future defect.
|
||||
|
||||
## How to read the change
|
||||
|
||||
- Read the diff against its intent. Does it actually do what the title/commit says? A
|
||||
correct-looking diff that solves the wrong problem is the most expensive approval.
|
||||
- Read the *deleted* lines as carefully as the added ones. Behaviour is often lost in a
|
||||
removal, and diffs render deletions quietly.
|
||||
- Follow each new call one level into the callee. The assumption that breaks it is usually
|
||||
one level down, invisible in the diff itself.
|
||||
|
||||
## For each finding
|
||||
|
||||
State file and line, the concrete failure (what input makes it break, or why it always
|
||||
breaks), and the change. Show the fix as code when it is short. "This could be a problem"
|
||||
without a path to a real input is noise; either trace it or drop it.
|
||||
|
||||
Order findings by severity and lead with the worst. Skip anything a formatter would fix.
|
||||
Skip preference. If a choice is defensible, leave it — a review is not a place to impose
|
||||
your style on code that works.
|
||||
|
||||
## Say when it is fine
|
||||
|
||||
A review that invents problems to look thorough is worse than a short one. If the change
|
||||
is correct, say so plainly and stop. "Looks correct, and here is what I checked" is a
|
||||
complete and useful review.
|
||||
|
||||
## Verify, do not assume
|
||||
|
||||
Read the surrounding code before calling something a bug. A "missing" null check often
|
||||
exists one level up; a "redundant" guard often covers a caller you have not seen. Run the
|
||||
tests or write the failing input if that is what settles it. A finding you verified is
|
||||
worth ten you suspected.
|
||||
@@ -0,0 +1,53 @@
|
||||
---
|
||||
name: security
|
||||
description: Review code for security defects, or write code that handles untrusted input. Use when touching authentication, user input, file paths, shell commands, SQL, or anything reachable from the network.
|
||||
---
|
||||
|
||||
# Security
|
||||
|
||||
Find the trust boundary first. Everything crossing it is hostile until parsed.
|
||||
|
||||
## The boundaries in most codebases
|
||||
|
||||
- Request bodies, query strings, headers, cookies.
|
||||
- File contents and filenames, including paths a user supplied.
|
||||
- Environment variables in a multi-tenant deployment.
|
||||
- Anything a model or a third-party API returned.
|
||||
|
||||
Inside a boundary, values are already validated and re-checking them is noise. At the
|
||||
boundary, nothing is optional.
|
||||
|
||||
## What to look for, in order
|
||||
|
||||
1. **Injection.** String-built SQL, shell commands assembled from input, `eval`, template
|
||||
rendering with user data as the template rather than the data. The fix is parameters and
|
||||
argument arrays, never escaping.
|
||||
2. **Missing authorisation.** An endpoint that checks *who* you are but not *what* you may
|
||||
touch. Look for an id taken from the request and used without an ownership check.
|
||||
3. **Path traversal.** `../` in anything joined onto a filesystem root. Resolve, then verify
|
||||
the result is still inside the root — a prefix check on the raw input misses
|
||||
`a/../../secret`.
|
||||
4. **Secrets in the wrong place.** Keys in source, in logs, in error messages, in a commit.
|
||||
A secret that reached a log is a secret to rotate.
|
||||
5. **Server-side request forgery.** A URL from input, fetched. Block private and loopback
|
||||
addresses by *resolved* address, and re-check every redirect hop.
|
||||
6. **Weak crypto and hand-rolled auth.** Homemade token formats, `Math.random` for anything
|
||||
security-bearing, comparisons on secrets that are not constant time.
|
||||
|
||||
## Verify the path before reporting
|
||||
|
||||
Trace each candidate from an attacker-controlled value to the sink before you name it. A
|
||||
"this could be unsafe" without that path is noise that buries the real finding. If you
|
||||
cannot construct the malicious input that reaches the sink, either keep looking or say
|
||||
plainly that you could not confirm it.
|
||||
|
||||
Do not fix a symptom at one caller when the sink is shared. Grep every caller and fix the
|
||||
seam once — a sanitiser at one of five call sites is four open holes and one false sense
|
||||
of safety.
|
||||
|
||||
## Reporting
|
||||
|
||||
File, line, the path from input to sink, a concrete payload, and the fix. Rank by
|
||||
exploitability: a reachable injection beats a theoretical weakness in dead code. Say
|
||||
plainly when a thing that looks dangerous is actually fine, and why — a reviewer's
|
||||
confidence in the clean parts is worth as much as a finding.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user