diff --git a/README.md b/README.md index f2a5080..2f8fdf1 100644 --- a/README.md +++ b/README.md @@ -88,6 +88,10 @@ the agent puts a question on screen with options. message, so a search across forty files does not fill the main context. Its progress streams to a panel. +**Extensible from the prompt.** `/registry` browses external skills and plugins and installs +them with one confirmation. A skill is shown in full before its text joins your system prompt; +a plugin is a manifest of refusal rules, never code. + **Remembers between sessions.** Decisions, working commands, and traps go into per-project memory that is injected at the start of every future session. @@ -109,6 +113,7 @@ core ones and a disabled set reaches neither the wire nor the prompt. | [Agents and thinking](docs/agents.md) | variants, thinking levels, read-only modes | | [Skills](docs/skills.md) | the bundled skills and writing your own | | [Plugins](docs/plugins.md) | the plugin interface and the builtins | +| [Registry](docs/registry.md) | installing external skills and plugins | | [Memory and state](docs/memory.md) | memory, task lists, sessions, compaction | | [MCP](docs/mcp.md) | connecting Model Context Protocol servers | | [Headless mode](docs/headless.md) | `-p`, JSON events, exit codes, CI recipes | @@ -123,8 +128,9 @@ Type `/` and a menu appears, narrowing as you type. ``` /help /agent [name] /think [level] /provider /models /model -/skills /plugins /init /context /todos /notes /memory -/tools /compact /cost /sessions /resume /save /clear /exit +/skills /plugins /registry [search|add|remove] /init /context +/todos /notes /memory /tools /compact /cost +/sessions /resume /save /clear /exit ``` `esc` dismisses a panel, interrupts a running turn, and clears the queue. `ctrl-c` kills the @@ -136,11 +142,11 @@ workspace path. Up and down recall earlier prompts. Working: the agent loop, tool approvals, subagents, skills, plugins, per-project memory, session persistence, MCP, markdown rendering, headless mode, five-platform builds, streaming reasoning display, the mid-turn prompt queue, gateable tool sets, read-only git tools, batch -reads, `@file` completion, and interruptible commands. +reads, `@file` completion, interruptible commands, and the external registry. Next up is in [TODO.md](TODO.md); the longer view and what has been declined are in [ROADMAP.md](ROADMAP.md). The short version of what is missing: a summary of what compaction -discarded, `web_fetch`, and a cheaper model for subagent searches. +discarded, `web_fetch`, a spend ceiling, and a cheaper model for subagent searches. ## License diff --git a/ROADMAP.md b/ROADMAP.md index 5219ea0..3dcbdd9 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -95,15 +95,35 @@ fails with a message saying the command did not finish and its effects are unkno model takes its next step from there. The kill takes the whole process tree, because killing `cmd /c` alone leaves the real command holding both pipes open and the read never returns. +### 0.1.0-beta.4 (unreleased) + +**Compaction no longer stops the loop.** The beta.2 repair dropped any assistant part whose +reasoning item pruning had removed. On a reasoning model that is every tool call, so past the +threshold the model could no longer see what it had already run — and re-ran it until the step +limit ended the turn. The fix strips the provider `itemId` rather than the part: without one the +same content is serialised inline instead of as an `item_reference`, so the dependency on the +pruned reasoning item disappears while the history survives. Compaction may shorten the +history; it must not blank it. + +**External registry.** `/registry` browses, searches, installs, and removes skills and plugins +from an index over https. The two kinds are treated differently on purpose: a skill is prompt +text and is shown in full before it joins your system prompt, while a plugin is a validated +manifest of deny rules that the compiled guard evaluates. Loading code from a URL is declined +outright — a plugin that can block tool calls could otherwise lie about blocking them. + +**Interface.** Context shown as a percentage of the compaction threshold, amber from two +thirds and red at 90, so a turn about to lose history says so first. Aligned command menu and +registry tables, and `/skills` and `/plugins` name the origin of every entry. + --- ## Next ### Lossless-enough compaction -Compaction says the history was pruned but not what was in it, so the model can contradict -its own earlier decision with confidence. A summary of the discarded span costs one cheap -call and removes the whole class of problem. +Compaction keeps the model's memory of a turn now, but it still says nothing about the messages +it discarded, so the model can contradict its own earlier decision with confidence. A summary of +the discarded span costs one cheap call and removes the whole class of problem. ### Cost control @@ -117,6 +137,12 @@ and a per-session ceiling are both small changes on top of the pricing that alre and forgotten in the other is a silently ungated write. Marking each tool where it is defined, and checking the coverage in the suite, removes the failure mode rather than documenting it. +### Registry trust + +An index is trusted for its contents, not its authorship: `registryUrl` is the whole trust +decision, and there are no signatures. Publisher keys and a pinned digest per entry would make +"install this skill" a decision about a specific artifact rather than about a URL. + ### `web_fetch` URL to markdown, in a `net` set that is off by default — it is the one tool that leaves the @@ -139,8 +165,10 @@ losing the original. **Structured diff review.** Approve or reject individual hunks of an `edit_file` call rather than the whole thing. -**Plugin loading from disk.** Plugins are compiled in. Loading `.shiro/plugins/*.ts` needs a -sandbox story first — a plugin that can block tool calls can also lie about blocking them. +**Plugin code from disk.** Declarative manifests ship in beta.4, and that is the whole of it +for now. Loading `.shiro/plugins/*.ts` needs a sandbox story first — a plugin that can block +tool calls can also lie about blocking them, and one that can execute can read whatever the +agent can read. **Prompt caching.** Anthropic and OpenAI both support it. The system prompt is rebuilt every step for task-list freshness, which defeats a naive cache; splitting the stable prefix from diff --git a/TODO.md b/TODO.md index 4f7f20d..bce7fb3 100644 --- a/TODO.md +++ b/TODO.md @@ -10,23 +10,15 @@ Longer-term direction lives in [ROADMAP.md](ROADMAP.md). ### Summarize the pruned span -Compaction tells the model the history was pruned but not what was in it, so it can -confidently contradict a decision it made forty messages ago. +Compaction now keeps the model's memory of a turn, but it still tells the model nothing about +the messages it dropped, so a decision from forty messages ago can be contradicted with +confidence. - [ ] Summarize the discarded messages before dropping them - [ ] Inject the summary in place of the count - [ ] Budget it: a summary that grows with the session defeats the point - [ ] Test: a pruned decision is still recoverable from the summary -### A cheaper model for subagents - -The subagent shares the parent's model. An `explore` run is search, not reasoning, and it -currently pays the parent's per-token rate. - -- [ ] `subagentModel` in config, defaulting to the parent -- [ ] `/cost` separates parent from subagent spend -- [ ] Test: the subagent's calls go to the configured model, the parent's do not - ### A spend ceiling A headless run that loops costs real money with nothing to stop it. @@ -36,20 +28,28 @@ A headless run that loops costs real money with nothing to stop it. - [ ] Headless exits non-zero with the ceiling named, rather than stopping silently - [ ] Test: a session past its ceiling refuses the next turn and says why +### A cheaper model for subagents + +The subagent shares the parent's model. An `explore` run is search, not reasoning, and it +currently pays the parent's per-token rate. + +- [ ] `subagentModel` in config, defaulting to the parent +- [ ] `/cost` separates parent from subagent spend +- [ ] Test: the subagent's calls go to the configured model, the parent's do not + +### Hot-reload an installed entry + +`/registry add` writes the file and says to restart. The skill catalogue and the guard chain +are both assembled at boot, so a mid-session install does nothing until then. + +- [ ] Rebuild the skill list and plugin host after an install or removal +- [ ] Leave a turn in flight alone: its rules must not change underneath it +- [ ] Test: a skill installed mid-session is callable in the next turn without a restart + --- ## Next -### Summarize the pruned span - -Compaction tells the model the history was pruned but not what was in it, so it can -confidently contradict a decision it made forty messages ago. - -- [ ] Summarize the discarded messages before dropping them -- [ ] Inject the summary in place of the count -- [ ] Budget it: a summary that grows with the session defeats the point -- [ ] Test: a pruned decision is still recoverable from the summary - ### `web_fetch` - [ ] URL to markdown, size-capped @@ -106,6 +106,11 @@ Not bugs exactly, but things that will bite someone. how far a half-run migration got. - **`@` completion lists files, not directories.** `@src/` narrows correctly, but you cannot complete to `src/` itself, because the walker only yields files. +- **An installed skill is a stranger's words in your system prompt.** The install shows the + body first and `/skills` records the origin, but nothing re-checks it later: a registry that + changes a URL's contents affects the next install, not one already on disk. +- **A registry index is trusted for its contents, not its authorship.** There are no + signatures. `registryUrl` is the whole trust decision. --- @@ -127,3 +132,10 @@ Kept for one release, then deleted. - [x] `ctrl-c` kills the running command and keeps the turn. The kill takes the whole process tree: killing `cmd /c` alone left the real command holding both pipes open, so the interrupt appeared to do nothing for 19 seconds +- [x] **Compaction no longer stops the loop.** Pruning used to drop any assistant part whose + reasoning item it removed, which on a reasoning model is every tool call. The model lost + its record of what it had run and re-ran it until the step limit. The repair strips the + provider `itemId` instead of the part, so the same content is sent inline +- [x] `/registry`: browse, search, install, and remove external skills and plugins. Skills are + shown in full before install; plugins are a validated manifest of deny rules, never code +- [x] Context shown as a percentage of the compaction threshold, amber at two thirds, red at 90 diff --git a/docs/architecture.md b/docs/architecture.md index 6d47b67..73a9b66 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -176,16 +176,25 @@ Only `api.openai.com` gets the chain. Third-party endpoints do not implement `/v ## Compaction and its repair -Pruning breaks two different provider invariants, and `src/prune.ts` repairs both. +Pruning breaks two provider invariants, and `src/prune.ts` repairs both. -**A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a -reasoning item and keeps the message item from the same response. The responses API treats the -message as that reasoning item's dependent and returns 400. +**A message detached from its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a +reasoning item and keeps the message item from the same response. A part carrying a provider +`itemId` is not sent inline: the responses provider serialises it as +`{ type: 'item_reference', id }`, pointing at an item stored on their side, and that stored +item depends on the reasoning item that pruning just removed. The result is a 400. The two carry different ids, so they cannot be matched by id. What links them is the assistant message they arrived in: one message is one response, and its reasoning item covers every -other item in it. `dropOrphanedItems` drops the dependent parts of any turn whose reasoning was -removed — which costs nothing, since pruning was already discarding those turns. +other item in it. `detachOrphanedItems` strips the `itemId` from those parts, which is what +sends the same content inline instead — verified against the provider's own serialiser, where +a `text` part with an itemId goes out as `item_reference` and the identical part without one +goes out as `output_text`. + +Dropping the parts was the first attempt and it broke the loop. On a reasoning model every +tool call carries an itemId, so after the first compaction the model could not see what it had +already run, and re-ran the same tools until the step limit ended the turn. **Compaction may +shorten the history; it must not blank it.** **A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts *messages*, so the cut lands between an assistant `tool-call` and the `tool` message answering @@ -199,6 +208,17 @@ it. What reaches the wire is a `function_call_output` with no `function_call`: reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a suspended approval looks like, and dropping it would break resume. +## Registry + +`/registry` fetches an index of external skills and plugins over https. Skills are prompt text +and are shown in full before install; plugins are a JSON manifest of refusal rules, never code. +The guard evaluating those rules is compiled, identical for every installed plugin, so an entry +from a registry cannot execute anything. See [registry](registry.md) for the validation and +the reasoning. + +`src/registry.ts` has no UI and no side effects until `install()` is called, which is what lets +`stage()` show a body before it becomes part of every future prompt. + ## Module map | Module | Responsibility | @@ -208,6 +228,7 @@ suspended approval looks like, and dropping it would break resume. | `tools-git.ts` | read-only git tools, spawned with a fixed argv | | `ignore.ts` | gitignore-aware walker, path jail | | `complete.ts` | `@path` token extraction, ranking, insertion | +| `registry.ts` | external index, validation, install and removal | | `prompt.ts` | system prompt assembly from live state | | `agents.ts` | variants, thinking levels | | `skills.ts` | discovery, catalogue, `skill` tool | @@ -234,10 +255,15 @@ Every module is pure of the UI except `ui/`, and `ui/` never touches the SDK. Th ## Testing -482 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop; +538 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop; `ink-testing-library` drives the UI with real keystrokes; MCP is tested against a real stdio -server subprocess; provider wire formats are tested against a local HTTP server; the interrupt -path spawns a real subprocess and asserts it died early rather than ran out. +server subprocess; provider wire formats and the registry are tested against a local HTTP +server; the interrupt path spawns a real subprocess and asserts it died early rather than ran +out. The pattern throughout is to assert on what actually crossed a boundary — what went on the wire, what is on screen, what is on disk — rather than on internal calls. + +Compaction is asserted on **behaviour**, not shape: the loop must terminate because the model +chose to, and every call after the first must still carry the earlier exchange. A shape +assertion would have passed while the model was losing its memory. diff --git a/docs/configuration.md b/docs/configuration.md index 593702e..5499df8 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -22,6 +22,7 @@ Written by `/provider`, editable by hand. Every field is optional. "maxRetries": 3, "plugins": ["guard", "time"], "toolSets": ["edit-plus", "git"], + "registryUrl": "https://example.com/my-registry/index.json", "mcpServers": { "fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] } } @@ -38,8 +39,9 @@ Written by `/provider`, editable by hand. Every field is optional. | `agent` | default variant: `default`, `quick`, `deep`, `plan`, `review` | | `thinking` | default level: `off`, `low`, `medium`, `high`, `max` | | `maxRetries` | retries per model call for transient failures. Default 3 | -| `plugins` | which plugins to enable. Omit for `["guard", "time"]` | +| `plugins` | which builtin plugins to enable. Omit for `["guard", "time"]` | | `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`. Omit for all of them. See [tools](tools.md) | +| `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) | | `mcpServers` | see [MCP](mcp.md) | ## Provider presets @@ -117,6 +119,8 @@ cat file | shiro -p prompt read from stdin memory/.json durable per-project notes history/.json prompt history for up-arrow recall skills/*.md your own skills + registry/skills/*.md skills installed with /registry + registry/plugins/*.json plugin manifests installed with /registry ``` Project files: diff --git a/docs/development.md b/docs/development.md index 1abcb1c..d7d511c 100644 --- a/docs/development.md +++ b/docs/development.md @@ -16,7 +16,7 @@ faster and the fallback path is exercised without it. ```bash bun run shiro # run from source bun run typecheck # tsc --noEmit -bun test # 482 tests +bun test # 538 tests bun run build # single binary for this platform -> dist/shiro bun run release # all five platforms -> dist/release + SHA256SUMS bun run install:local # build, then copy onto PATH @@ -71,6 +71,9 @@ mock-verification test: request body - `pruneMessages` leaving a tool result without its tool call — same, and it took a stub endpoint that rejected the pairing to prove the fix +- Compaction blanking the model's memory of its own tool calls — invisible in any single + request, and visible only as "the loop ran to its step limit". Caught by asserting the loop + terminated because the model chose to, not that the messages had a particular shape - `--json` serialising `Error` as `{}` — visible only in the printed output - Automatic approval requests prompting the user — visible only in the event sequence - `ctrl-c` killing `cmd /c` but not the command under it — visible only as elapsed time, since diff --git a/docs/memory.md b/docs/memory.md index c04ea77..bb3dce2 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -139,17 +139,25 @@ Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both we before it did. **A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a -reasoning item and keeps the message item from the same response. The OpenAI responses API -treats the message as a dependent of that reasoning item and rejects the request: +reasoning item and keeps the message item from the same response. That message carries a +provider `itemId`, and the OpenAI responses provider serialises anything with one as +`{ type: 'item_reference', id }` — a pointer to an item stored on their side, which depends on +the reasoning item that is now gone: ``` 400 Item 'msg_…' of type 'message' was provided without its required 'reasoning' item: 'rs_…' ``` The two carry different ids, so they cannot be matched by id. What links them is the -assistant message they arrived in — one message is one response. `dropOrphanedItems` drops the -dependent parts of any turn whose reasoning was removed. That costs nothing, because pruning -was already discarding those turns. +assistant message they arrived in — one message is one response. `detachOrphanedItems` strips +the `itemId` from those parts. Without one the same content is serialised **inline**, which +carries no dependency on anything stored, so the turn survives intact. + +Dropping the parts instead was the first attempt, and it was wrong in a way that only showed +up over a long turn: on a reasoning model every tool call carries an itemId, so after the +first compaction the model could no longer see what it had already run. It re-ran the same +tools until the step limit ended the turn. The history is the model's memory; compaction may +shorten it but must not blank it. **A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts *messages*, not pairs, so the cut can land between the assistant message holding a `tool-call` diff --git a/docs/plugins.md b/docs/plugins.md index 4d3b7ba..4806d00 100644 --- a/docs/plugins.md +++ b/docs/plugins.md @@ -4,8 +4,17 @@ A plugin extends the agent in four ways: it can add tools, mark tools auto-appro a tool call before it runs, and append to the system prompt. It can also run something after each turn. -Plugins are compiled into the binary. Loading them from disk is deliberately not supported -yet — see [ROADMAP.md](../ROADMAP.md). +Two kinds exist, and only one can contain code: + +- **Builtin** plugins are compiled into the binary and may do anything in the interface below. +- **Installed** plugins come from a registry as a JSON manifest of refusal rules. They are + data: the guard evaluating them is compiled code, identical for every install. See + [registry](registry.md). + +Loading TypeScript from disk or a URL is deliberately not supported. A plugin that can block +tool calls can also lie about blocking them, and one that could execute could read every file +the agent can read. That is a sandbox problem, not a loader problem — see +[ROADMAP.md](../ROADMAP.md). ## Enabling @@ -13,9 +22,12 @@ yet — see [ROADMAP.md](../ROADMAP.md). { "plugins": ["guard", "time"] } ``` -That is also the default when the field is absent. `--no-plugins` disables all of them, -including the guard. `/plugins` lists what is active and reports any name that did not -resolve. +That is also the default when the field is absent, and it lists **builtin** plugins only. +Installed plugins are always active once present, because installing one was the decision to +enable it; remove it with `/registry remove `. + +`--no-plugins` disables everything, builtin and installed, including the guard. `/plugins` +lists what is active, marks installed entries, and reports any name that did not resolve. ## The interface @@ -92,7 +104,11 @@ intrusive, but it is genuinely useful when a turn takes minutes. ## Writing one -Plugins live in `src/plugins-builtin.ts` and are registered in `BUILTIN_PLUGINS`. +A refusal rule is usually better as an installed manifest: no rebuild, and nothing to review. +See [registry](registry.md) for the manifest shape. Reach for a builtin only when the plugin +needs to contribute a tool or run something after a turn. + +Builtin plugins live in `src/plugins-builtin.ts` and are registered in `BUILTIN_PLUGINS`. ```ts export const noSecretsPlugin: Plugin = { @@ -122,6 +138,6 @@ refusal it was never told about and tries to route around it. ## Ordering -Plugins run in the order they are enabled. The first `beforeToolCall` to block wins; -later hooks are not consulted. `afterTurn` runs every hook, and one throwing does not stop -the rest. +Builtin plugins run first, in the order they are enabled, then installed ones. The first +`beforeToolCall` to block wins; later hooks are not consulted. `afterTurn` runs every hook, and +one throwing does not stop the rest. diff --git a/docs/registry.md b/docs/registry.md new file mode 100644 index 0000000..ad28027 --- /dev/null +++ b/docs/registry.md @@ -0,0 +1,128 @@ +# Registry + +External skills and plugins, browsed and installed from the CLI. + +``` +/registry everything in the index +/registry search narrow by name or description +/registry installed what is already here +/registry add fetch, show, then install on confirmation +/registry remove delete an installed entry +``` + +``` +registry +3 of 3 available +S migration Write or run a database migration +S commit-style Write commits the way this team does +P no-secrets Refuses writes to credential files installed +S skill P plugin | /registry add +``` + +`esc` dismisses the panel. + +## The two kinds are not equally safe + +**A skill is instructions.** Installing one puts a stranger's words into the system prompt of +every future session in this project. That is prompt injection by invitation, so the install +shows the body first and `/skills` always says the origin: + +``` +install skill "migration"? +https://raw.githubusercontent.com/example/registry/main/skills/migration.md + + Migrations live in `db/migrations/` and are timestamped, never renumbered. + + Run `bun run db:migrate` locally first. Staging runs them on deploy. + +A skill is instructions the agent follows. This text joins your system prompt. +y install | n cancel +``` + +**A plugin is data.** Never code. A manifest declares refusal rules; the guard that evaluates +them is the same compiled code for every installed plugin: + +```json +{ + "name": "no-secrets", + "description": "Refuses writes to credential files", + "appendix": "The no-secrets plugin refuses writes to .env and credential files.", + "deny": [ + { + "tools": ["write_file", "edit_file", "multi_edit"], + "pathPattern": "(^|/)\\.env|credentials|\\.pem$", + "reason": "refusing to write a credential file; add secrets yourself" + } + ] +} +``` + +Loading TypeScript from a URL is not offered at any price. A plugin can block tool calls, so +one that could also execute could read every file the agent can read and lie about blocking +anything. See [plugins](plugins.md) for why this boundary exists. + +## What is validated before anything is written + +| Check | Why | +|---|---| +| `https` only, `localhost` for tests | `file:` would read a local path, `data:` would inline a payload | +| Name matches `^[a-z0-9][a-z0-9-]*$` | the name becomes a filename, so `../evil` must not parse | +| Index at most 256 KB, body at most 64 KB | a hostile index should not exhaust memory | +| Manifest against a strict schema | extra keys like `beforeToolCall` are dropped, not honoured | +| Every pattern compiles as a regex | a broken pattern would fail on the first tool call instead | +| Pattern at most 200 characters | it runs on every tool call; a pathological one is a denial of service | +| Body name matches the index name | an index entry cannot serve something else under a trusted name | +| At least one deny rule | a plugin with no rules is only prompt text, which is what a skill is for | + +A malformed installed plugin is reported by `/plugins` and skipped. One bad install does not +stop the agent from starting. + +## Where installs land + +``` +~/.shiro-neko/registry/ + skills/.md loaded as origin "registry" + plugins/.json loaded as a declarative plugin +``` + +Precedence for skills, low to high: **builtin → registry → user → project**. A skill you wrote +in `~/.shiro-neko/skills/` or `.shiro/skills/` always beats one fetched from a registry, so an +install can never silently shadow your own work. + +Installs take effect on the next start. A skill joins the system prompt and a plugin joins the +guard chain, and both are assembled once at boot; hot-swapping either mid-session would mean a +turn whose rules changed underneath it. + +## Pointing at your own index + +```json +{ "registryUrl": "https://example.com/my-registry/index.json" } +``` + +The index is one JSON document: + +```json +{ + "skills": [ + { + "name": "migration", + "description": "Write or run a database migration", + "url": "https://example.com/skills/migration.md", + "author": "you" + } + ], + "plugins": [ + { + "name": "no-secrets", + "description": "Refuses writes to credential files", + "url": "https://example.com/plugins/no-secrets.json" + } + ] +} +``` + +Both arrays are optional. A name may appear once as a skill and once as a plugin; `/registry +add skill:review` disambiguates, and an ambiguous name is refused rather than guessed. + +A private index is just a URL you control. There is no account, no token, and no telemetry — +`/registry` makes exactly one GET for the index and one for the entry you install. diff --git a/docs/skills.md b/docs/skills.md index b19df6f..7c6ddd8 100644 --- a/docs/skills.md +++ b/docs/skills.md @@ -30,14 +30,16 @@ to X" — not as a summary. ## Where they load from -Three sources, later overriding earlier by name: +Four sources, later overriding earlier by name: 1. **builtin** — compiled into the binary -2. **user** — `~/.shiro-neko/skills/*.md` -3. **project** — `.shiro/skills/*.md` +2. **registry** — `~/.shiro-neko/registry/skills/*.md`, installed with `/registry add` +3. **user** — `~/.shiro-neko/skills/*.md` +4. **project** — `.shiro/skills/*.md` A project skill named `debug` replaces the bundled one entirely. `/skills` shows what -loaded and where each came from. +loaded and where each came from — which matters most for `registry`, since that body came +from someone else and is now in your system prompt. See [registry](registry.md). `--no-skills` skips all of them, builtin included. diff --git a/src/cli.tsx b/src/cli.tsx index b29b01b..c877d55 100644 --- a/src/cli.tsx +++ b/src/cli.tsx @@ -14,6 +14,7 @@ import { costOf } from './pricing'; import { BUILTIN_PLUGINS, DEFAULT_ENABLED } from './plugins-builtin'; import { createHost } from './plugins'; import { fetchModels, presetById } from './providers'; +import * as registry from './registry'; import { Session } from './session'; import { loadSkills } from './skills'; import * as store from './store'; @@ -21,6 +22,7 @@ import { createTaskTool } from './subagent'; import { VERSION, versionLine } from './version'; import { createAskBridge } from './ui/Ask'; import { App, createApprovalBridge, createNoticeBus, createSubagentBus, type AppHooks } from './ui/App'; +import type { RegistryRow as AppRegistryRow } from './ui/Panels'; // SDK warnings go straight to stderr, which tears up the Ink render. (globalThis as { AI_SDK_LOG_WARNINGS?: boolean }).AI_SDK_LOG_WARNINGS = false; @@ -56,13 +58,14 @@ first run: start shiro with no key and it opens provider setup, or use /provider config: ${configPath()} { "provider": "openai", "model": "gpt-5", "apiKey": "...", "agent": "default", "thinking": "medium", "plugins": ["guard", "time"], - "toolSets": ["edit-plus", "git"], + "toolSets": ["edit-plus", "git"], "registryUrl": "https://...", "mcpServers": { "fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] } } } env: SHIRO_PROVIDER SHIRO_MODEL SHIRO_BASE_URL SHIRO_API_KEY ANTHROPIC_API_KEY OPENAI_API_KEY skills: builtin, plus ~/.shiro-neko/skills/*.md and .shiro/skills/*.md +registry: /registry to browse and install external skills and plugins sessions: ${store.sessionsDir()} in-session: /help for the command list`; @@ -150,6 +153,40 @@ const instructions = has('--no-instructions') ? [] : await loadInstructions(); const skills = has('--no-skills') ? [] : await loadSkills(); const promptHistory = await store.loadHistory(); +const installedPlugins = has('--no-plugins') ? { plugins: [], errors: [] } : await registry.loadInstalledPlugins(); + +/** Installed entries, as `kind:name`, so the registry list can mark what is already here. */ +async function installedNames(): Promise> { + const names = new Set(); + for (const s of skills) if (s.origin === 'registry') names.add(`skill:${s.name}`); + for (const p of installedPlugins.plugins) names.add(`plugin:${p.name}`); + return names; +} + +/** + * One entry by name, from the index. + * + * A name alone is ambiguous when a skill and a plugin share it, so `skill:name` + * disambiguates. Asking rather than guessing would be worse here: the two differ in + * what they can do, and picking one silently is the wrong default. + */ +async function findEntry(name: string): Promise { + const parsed = /^(skill|plugin):(.+)$/.exec(name.trim()); + const kind = parsed?.[1] as registry.RegistryKind | undefined; + const bare = (parsed?.[2] ?? name).trim().toLowerCase(); + + const entries = await registry.fetchIndex(cfg.registryUrl); + const hits = entries.filter((e) => e.name === bare && (kind === undefined || e.kind === kind)); + + if (hits.length === 0) throw new Error(`no registry entry named "${bare}". Try /registry search ${bare}`); + if (hits.length > 1) { + throw new Error( + `"${bare}" is both a skill and a plugin. Use /registry add skill:${bare} or /registry add plugin:${bare}`, + ); + } + return hits[0]!; +} + let agentVariant: AgentVariant; try { agentVariant = resolveAgent(flag('--agent') || cfg.agent, flag('--think') || cfg.thinking); @@ -163,8 +200,8 @@ const pluginErrors = enabledPlugins .filter((name) => !BUILTIN_PLUGINS.some((p) => p.name === name)) .map((name) => ({ plugin: name, message: 'no such plugin' })); const plugins = createHost( - BUILTIN_PLUGINS.filter((p) => enabledPlugins.includes(p.name)), - pluginErrors, + [...BUILTIN_PLUGINS.filter((p) => enabledPlugins.includes(p.name)), ...installedPlugins.plugins], + [...pluginErrors, ...installedPlugins.errors], ); const memory = has('--no-memory') ? undefined : new Memory(process.cwd(), languageModel); @@ -280,6 +317,55 @@ const hooks: AppHooks = { for await (const rel of walk({ limit: 5000 })) found.push(rel); return found; }, + registry: { + list: async () => { + const entries = await registry.fetchIndex(cfg.registryUrl); + const installed = await installedNames(); + return entries.map((e) => ({ + name: e.name, + kind: e.kind, + description: e.description, + ...(e.author ? { author: e.author } : {}), + installed: installed.has(`${e.kind}:${e.name}`), + })); + }, + installed: async () => { + const rows: AppRegistryRow[] = []; + for (const s of skills) { + if (s.origin === 'registry') rows.push({ name: s.name, kind: 'skill', description: s.description, installed: true }); + } + for (const p of installedPlugins.plugins) { + rows.push({ name: p.name, kind: 'plugin', description: p.description, installed: true }); + } + return rows.sort((a, b) => a.name.localeCompare(b.name)); + }, + stage: async (name) => { + const entry = await findEntry(name); + const { preview } = await registry.stage(entry); + return { + row: { name: entry.name, kind: entry.kind, description: entry.description }, + url: entry.url, + preview, + }; + }, + install: async (name) => { + const entry = await findEntry(name); + const { path } = await registry.install(entry); + // Loaded on the next start rather than hot-swapped: a skill joins the system + // prompt and a plugin joins the guard chain, and both are built once at boot. + return `installed ${entry.kind} ${entry.name} to ${path}\nrestart shiro to load it`; + }, + remove: async (name) => { + const parsed = /^(skill|plugin):(.+)$/.exec(name); + const kinds: registry.RegistryKind[] = parsed ? [parsed[1] as registry.RegistryKind] : ['skill', 'plugin']; + const bare = parsed ? parsed[2]! : name; + + for (const kind of kinds) { + if (await registry.uninstall(kind, bare)) return `removed ${kind} ${bare}\nrestart shiro to unload it`; + } + throw new Error(`nothing installed under the name "${bare}"`); + }, + }, initPrompt: INIT_PROMPT, history: promptHistory, recordPrompt: (text) => void store.appendHistory(text), @@ -302,13 +388,17 @@ const hooks: AppHooks = { }, listSkills: () => { if (skills.length === 0) return 'no skills loaded'; - return skills.map((s) => `${s.name.padEnd(10)} ${s.origin.padEnd(8)} ${s.description}`).join('\n'); + return [ + ...skills.map((s) => `${s.name.padEnd(12)} ${s.origin.padEnd(9)} ${s.description}`), + '', + '`/registry` to browse and install more', + ].join('\n'); }, listPlugins: () => { - const active = plugins.plugins.map((p) => `${p.name.padEnd(8)} ${p.description}`); - const failed = plugins.errors.map((e) => `${e.plugin.padEnd(8)} ${e.message}`); + const active = plugins.plugins.map((p) => `${p.name.padEnd(12)} ${p.description}`); + const failed = plugins.errors.map((e) => `${e.plugin.padEnd(12)} ${e.message}`); if (active.length === 0 && failed.length === 0) return 'no plugins active'; - return [...active, ...failed].join('\n'); + return [...active, ...failed, '', '`/registry` to browse and install more'].join('\n'); }, listMemory: async () => { if (!memory) return 'memory is disabled (--no-memory)'; diff --git a/src/commands.ts b/src/commands.ts index 6d4e86f..e9daaa8 100644 --- a/src/commands.ts +++ b/src/commands.ts @@ -16,6 +16,7 @@ export type CommandAction = | { type: 'notes' } | { type: 'skills' } | { type: 'plugins' } + | { type: 'registry'; action: 'list' | 'search' | 'add' | 'remove' | 'installed'; arg?: string } | { type: 'memory' } | { type: 'agent'; agent?: string } | { type: 'think'; level?: string } @@ -42,6 +43,7 @@ export const COMMANDS: CommandSpec[] = [ { name: 'model', arg: '', summary: 'switch model by name' }, { name: 'skills', summary: 'list loaded skills' }, { name: 'plugins', summary: 'list active plugins' }, + { name: 'registry', arg: '[search|add|remove] [name]', summary: 'browse and install external skills and plugins' }, { name: 'init', summary: 'have the agent write AGENTS.md for this project' }, { name: 'context', summary: 'show which instruction files are loaded' }, { name: 'todos', summary: "show the agent's task list" }, @@ -60,14 +62,14 @@ export const COMMANDS: CommandSpec[] = [ const usage = (c: CommandSpec) => `/${c.name}${c.arg ? ` ${c.arg}` : ''}`; export const HELP = [ - ...COMMANDS.map((c) => `${usage(c).padEnd(18)} ${c.summary}`), + ...COMMANDS.map((c) => `${usage(c).padEnd(22)} ${c.summary}`), '', - 'esc interrupt the running turn and clear the queue', - 'ctrl-c kill the running command, keeping the turn', - 'ctrl-r expand or collapse the reasoning panel', - 'tab complete the highlighted command or file path', - 'up / down recall earlier prompts, or move in an open menu', - '@ complete a workspace path', + 'esc interrupt the running turn and clear the queue', + 'ctrl-c kill the running command, keeping the turn', + 'ctrl-r expand or collapse the reasoning panel', + 'tab complete the highlighted command or file path', + 'up / down recall earlier prompts, or move in an open menu', + '@ complete a workspace path', '', 'typing during a turn queues the prompt; queued prompts run in order afterwards', ].join('\n'); @@ -89,6 +91,42 @@ export function matchCommands(input: string): CommandSpec[] { /** True while the input is a bare command name being typed, so the menu should show. */ export const isMenuOpen = (input: string) => input.startsWith('/') && !input.includes(' '); +/** + * `/registry [list|installed|search |add |remove ]`. + * + * A bare `/registry` lists everything. `add` and `remove` need a name, and saying + * so beats fetching the whole index to then complain. + */ +function parseRegistry(arg: string): CommandAction { + const [verb = '', ...rest] = arg.split(/\s+/).filter(Boolean); + const name = rest.join(' ').trim(); + + switch (verb) { + case '': + case 'list': + return { type: 'registry', action: 'list' }; + case 'installed': + return { type: 'registry', action: 'installed' }; + case 'search': + return name + ? { type: 'registry', action: 'search', arg: name } + : { type: 'info', text: 'usage: /registry search ' }; + case 'add': + case 'install': + return name + ? { type: 'registry', action: 'add', arg: name } + : { type: 'info', text: 'usage: /registry add ' }; + case 'remove': + case 'uninstall': + return name + ? { type: 'registry', action: 'remove', arg: name } + : { type: 'info', text: 'usage: /registry remove ' }; + default: + // A bare word is almost always a search, and guessing beats a usage line. + return { type: 'registry', action: 'search', arg: arg.trim() }; + } +} + /** Pure parser: no IO, so the TUI and headless mode share one definition. */ export function parseCommand(raw: string): CommandAction { const input = raw.trim(); @@ -134,6 +172,8 @@ export function parseCommand(raw: string): CommandAction { return { type: 'skills' }; case 'plugins': return { type: 'plugins' }; + case 'registry': + return parseRegistry(arg); case 'memory': return { type: 'memory' }; case 'agent': diff --git a/src/config.ts b/src/config.ts index 84d0654..59f60d9 100644 --- a/src/config.ts +++ b/src/config.ts @@ -27,6 +27,8 @@ export type Config = { plugins?: string[]; /** Optional tool sets to offer beyond `core`; omit for all of them. */ toolSets?: ToolSetName[]; + /** Index for `/registry`. Omit for the default one. */ + registryUrl?: string; mcpServers?: Record; }; @@ -89,6 +91,7 @@ export async function loadConfig(): Promise { ...(file.thinking ? { thinking: file.thinking } : {}), ...(Array.isArray(file.plugins) ? { plugins: file.plugins } : {}), ...(Array.isArray(file.toolSets) ? { toolSets: file.toolSets.filter(isToolSetName) } : {}), + ...(typeof file.registryUrl === 'string' ? { registryUrl: file.registryUrl } : {}), ...(file.mcpServers ? { mcpServers: file.mcpServers } : {}), }; } diff --git a/src/prune.ts b/src/prune.ts index 10479a6..7af79de 100644 --- a/src/prune.ts +++ b/src/prune.ts @@ -6,9 +6,6 @@ type Part = { providerOptions?: Record>; }; -/** Parts the OpenAI responses API refuses to accept without their reasoning item. */ -const DEPENDENT = new Set(['text', 'tool-call']); - function itemId(part: Part): string | undefined { for (const options of Object.values(part.providerOptions ?? {})) { const id = options['itemId']; @@ -20,23 +17,39 @@ function itemId(part: Part): string | undefined { const partsOf = (message: ModelMessage): Part[] => message.role === 'assistant' && Array.isArray(message.content) ? (message.content as Part[]) : []; +/** The part again, with every provider `itemId` removed. */ +function withoutItemId(part: Part): Part { + const providerOptions: Record> = {}; + for (const [provider, options] of Object.entries(part.providerOptions ?? {})) { + const { itemId: _dropped, ...rest } = options; + if (Object.keys(rest).length > 0) providerOptions[provider] = rest; + } + const next: Part = { ...part }; + if (Object.keys(providerOptions).length > 0) next.providerOptions = providerOptions; + else delete next.providerOptions; + return next; +} + /** - * Drops assistant parts left orphaned by reasoning removal. + * Detaches assistant parts from reasoning items that pruning removed. * - * The OpenAI responses API treats a `message` item as a dependent of the `reasoning` - * item from the same response: send the message without its reasoning and the request - * is rejected with 400 "was provided without its required 'reasoning' item". - * `pruneMessages({ reasoning: 'all' })` strips the reasoning and keeps the message, - * producing exactly that request. + * A part carrying a provider `itemId` is not sent inline. The OpenAI responses + * provider serialises it as `{ type: 'item_reference', id }`, pointing at an item + * stored on their side, and that stored item depends on the `reasoning` item from + * the same response. Send the reference without the reasoning and the request is + * rejected with 400 "was provided without its required 'reasoning' item". * - * The two carry different item ids, so they cannot be matched by id. What links them - * is the assistant message they arrived in: one message is one response, and its - * reasoning item covers every other item in it. + * The repair is to drop the `itemId`, not the part. Without it the same content is + * serialised inline — a plain assistant message, a plain `function_call` — which + * carries no dependency on anything stored. Verified against the provider's own + * serialiser: `text` with an itemId goes out as `item_reference`, and the identical + * part without one goes out as `output_text`. * - * Reasoning only disappears from turns pruning is already discarding, so dropping the - * orphaned text costs nothing pruning was not already spending. + * Dropping the part instead, which is what this used to do, cost the model its + * memory of the turn: after compaction it could no longer see the tool results it + * had just collected, so it called the same tools again until it hit the step limit. */ -export function dropOrphanedItems(before: ModelMessage[], after: ModelMessage[]): ModelMessage[] { +export function detachOrphanedItems(before: ModelMessage[], after: ModelMessage[]): ModelMessage[] { const survivingReasoning = new Set(); for (const message of after) { for (const part of partsOf(message)) { @@ -54,7 +67,6 @@ export function dropOrphanedItems(before: ModelMessage[], after: ModelMessage[]) if (reasoning.some((id) => id !== undefined && survivingReasoning.has(id))) continue; for (const part of parts) { - if (!DEPENDENT.has(part.type)) continue; const id = itemId(part); if (id) orphaned.add(id); } @@ -62,23 +74,20 @@ export function dropOrphanedItems(before: ModelMessage[], after: ModelMessage[]) if (orphaned.size === 0) return after; - const cleaned: ModelMessage[] = []; - for (const message of after) { + return after.map((message) => { const parts = partsOf(message); - if (parts.length === 0) { - cleaned.push(message); - continue; - } + if (parts.length === 0) return message; - const kept = parts.filter((part) => { + let changed = false; + const next = parts.map((part) => { const id = itemId(part); - return id === undefined || !orphaned.has(id); + if (id === undefined || !orphaned.has(id)) return part; + changed = true; + return withoutItemId(part); }); - if (kept.length > 0) cleaned.push({ ...message, content: kept } as ModelMessage); - } - - return cleaned; + return changed ? ({ ...message, content: next } as ModelMessage) : message; + }); } export type PruneOptions = Parameters[0]; @@ -93,11 +102,9 @@ const anyParts = (message: ModelMessage): Part[] => * * The OpenAI responses API rejects a `function_call_output` with no `function_call` * carrying the same call id: 400 "No tool call found for function call output with - * call_id ...". Two things strand a result that way, and both happen on a long turn: - * `pruneMessages({ toolCalls: 'before-last-3-messages' })` counts messages, so the - * cut can land between an assistant tool-call and the tool message answering it, and - * `dropOrphanedItems` removes a tool-call whose reasoning item did not survive while - * the result sits in a separate message it never looks at. + * call_id ...". `pruneMessages({ toolCalls: 'before-last-3-messages' })` counts + * messages, so the cut can land between an assistant tool-call and the tool message + * answering it. * * The reverse pairing is left alone on purpose: a call still awaiting its result is * exactly what a suspended approval looks like, and dropping it would break resume. @@ -132,5 +139,5 @@ export function dropOrphanedResults(messages: ModelMessage[]): ModelMessage[] { /** pruneMessages, then repair the provider-item dependencies it breaks. */ export function prunePreservingItems(options: PruneOptions): ModelMessage[] { const pruned = pruneMessages(options); - return dropOrphanedResults(dropOrphanedItems(options.messages, pruned)); + return dropOrphanedResults(detachOrphanedItems(options.messages, pruned)); } diff --git a/src/registry.ts b/src/registry.ts new file mode 100644 index 0000000..7759983 --- /dev/null +++ b/src/registry.ts @@ -0,0 +1,295 @@ +import { homedir } from 'node:os'; +import { join } from 'node:path'; +import { z } from 'zod'; +import type { Plugin } from './plugins'; +import { parseSkill, type Skill } from './skills'; + +/** + * External skills and plugins, fetched from an index over HTTPS. + * + * Two kinds, and they are not equally safe. + * + * A **skill** is prompt text. Installing one puts a stranger's words into the + * system prompt of every future session in this project, which is prompt injection + * by invitation. The command shows the body before writing it and the origin is + * recorded, so `/skills` always says where an instruction came from. + * + * A **plugin** is declarative: a name, a prompt appendix, and deny rules matched + * against tool input. Never code. Loading arbitrary TypeScript from a URL would let + * an entry read every file the agent can read and lie about blocking anything, so + * that is not offered at any price. + */ + +const MAX_INDEX_BYTES = 256 * 1024; +const MAX_BODY_BYTES = 64 * 1024; +/** A pattern from the internet runs on every tool call; a huge one is a denial of service. */ +const MAX_PATTERN = 200; + +/** + * https only, with localhost allowed so the suite can serve a real index. + * + * `file:` would read a local path and `data:` would inline a payload, neither of + * which is what "fetch this from a registry" means to the person typing it. + */ +export function isFetchable(url: string): boolean { + let parsed: URL; + try { + parsed = new URL(url); + } catch { + return false; + } + if (parsed.protocol === 'https:') return true; + return parsed.protocol === 'http:' && (parsed.hostname === 'localhost' || parsed.hostname === '127.0.0.1'); +} + +const entrySchema = z.object({ + name: z + .string() + .min(1) + .max(40) + // The name becomes a filename. Anything else is a path traversal. + .regex(/^[a-z0-9][a-z0-9-]*$/, 'lowercase letters, digits, and dashes only'), + description: z.string().min(1).max(300), + // z.url() accepts file: and data:, which would read a local path or inline a + // payload. Only https reaches the network the way the user expects. + url: z + .string() + .url() + .refine((u) => isFetchable(u), 'must be an https URL'), + author: z.string().max(80).optional(), +}); + +const indexSchema = z.object({ + skills: z.array(entrySchema).max(500).optional(), + plugins: z.array(entrySchema).max(500).optional(), +}); + +export type RegistryKind = 'skill' | 'plugin'; +export type RegistryEntry = z.infer & { kind: RegistryKind }; + +const denyRuleSchema = z + .object({ + tools: z.array(z.string().max(60)).min(1).max(30), + pathPattern: z.string().max(MAX_PATTERN).optional(), + commandPattern: z.string().max(MAX_PATTERN).optional(), + reason: z.string().min(1).max(300), + }) + .refine((r) => r.pathPattern !== undefined || r.commandPattern !== undefined, { + message: 'a deny rule needs pathPattern or commandPattern', + }); + +const manifestSchema = z.object({ + name: z.string().min(1).max(40).regex(/^[a-z0-9][a-z0-9-]*$/), + description: z.string().min(1).max(300), + appendix: z.string().max(2000).optional(), + deny: z.array(denyRuleSchema).min(1).max(50), +}); + +export type PluginManifest = z.infer; + +export const DEFAULT_INDEX_URL = + 'https://raw.githubusercontent.com/zakirkun/shiro-neko-registry/main/index.json'; + +const home = () => process.env['SHIRO_HOME'] ?? homedir(); + +export const skillsDir = () => join(home(), '.shiro-neko', 'registry', 'skills'); +export const pluginsDir = () => join(home(), '.shiro-neko', 'registry', 'plugins'); + +async function fetchText(url: string, limit: number): Promise { + if (!isFetchable(url)) throw new Error(`refusing a non-https URL: ${url}`); + const res = await fetch(url, { redirect: 'follow', signal: AbortSignal.timeout(20_000) }); + if (!res.ok) throw new Error(`${url} returned ${res.status} ${res.statusText}`); + + const declared = Number(res.headers.get('content-length') ?? 0); + if (declared > limit) throw new Error(`${url} is ${declared} bytes, over the ${limit} byte limit`); + + const text = await res.text(); + if (text.length > limit) throw new Error(`${url} is over the ${limit} byte limit`); + return text; +} + +/** Parses an index document. Exported so the shape can be tested without a network. */ +export function parseIndex(source: string): RegistryEntry[] { + let raw: unknown; + try { + raw = JSON.parse(source); + } catch { + throw new Error('the registry index is not valid JSON'); + } + + const parsed = indexSchema.safeParse(raw); + if (!parsed.success) { + throw new Error(`the registry index is malformed: ${parsed.error.issues[0]?.message ?? 'unknown reason'}`); + } + + const seen = new Set(); + const entries: RegistryEntry[] = []; + for (const kind of ['skill', 'plugin'] as const) { + for (const entry of parsed.data[`${kind}s`] ?? []) { + const key = `${kind}:${entry.name}`; + if (seen.has(key)) continue; + seen.add(key); + entries.push({ ...entry, kind }); + } + } + return entries; +} + +export async function fetchIndex(url = DEFAULT_INDEX_URL): Promise { + return parseIndex(await fetchText(url, MAX_INDEX_BYTES)); +} + +export const searchEntries = (entries: RegistryEntry[], query: string): RegistryEntry[] => { + const needle = query.trim().toLowerCase(); + if (!needle) return entries; + return entries.filter( + (e) => e.name.includes(needle) || e.description.toLowerCase().includes(needle), + ); +}; + +/** Validates a plugin manifest, including that every pattern is a usable regex. */ +export function parseManifest(source: string): PluginManifest { + let raw: unknown; + try { + raw = JSON.parse(source); + } catch { + throw new Error('the plugin manifest is not valid JSON'); + } + + const parsed = manifestSchema.safeParse(raw); + if (!parsed.success) { + throw new Error(`the plugin manifest is malformed: ${parsed.error.issues[0]?.message ?? 'unknown reason'}`); + } + + for (const rule of parsed.data.deny) { + for (const pattern of [rule.pathPattern, rule.commandPattern]) { + if (pattern === undefined) continue; + try { + new RegExp(pattern, 'i'); + } catch (e) { + throw new Error(`invalid pattern "${pattern}": ${(e as Error).message}`); + } + } + } + + return parsed.data; +} + +/** + * A manifest as a Plugin. + * + * Rules are data, so the guard is the same code for every installed plugin: match + * the tool name, then the path or command against a compiled regex. Nothing from + * the manifest is ever evaluated. + */ +export function manifestToPlugin(manifest: PluginManifest): Plugin { + const rules = manifest.deny.map((rule) => ({ + tools: new Set(rule.tools), + path: rule.pathPattern ? new RegExp(rule.pathPattern, 'i') : undefined, + command: rule.commandPattern ? new RegExp(rule.commandPattern, 'i') : undefined, + reason: rule.reason, + })); + + return { + name: manifest.name, + description: `${manifest.description} (installed)`, + ...(manifest.appendix ? { appendix: manifest.appendix } : {}), + beforeToolCall: ({ toolName, input }) => { + const o = (input ?? {}) as Record; + const path = typeof o['path'] === 'string' ? o['path'] : ''; + const command = typeof o['command'] === 'string' ? o['command'] : ''; + + for (const rule of rules) { + if (!rule.tools.has(toolName)) continue; + if (rule.path && path && rule.path.test(path)) return rule.reason; + if (rule.command && command && rule.command.test(command)) return rule.reason; + } + return undefined; + }, + }; +} + +export type Installed = { name: string; kind: RegistryKind; path: string }; + +/** + * Downloads an entry and returns what would be written, without writing it. + * + * Separated from the write so the caller can show the user a skill body before it + * becomes part of every future prompt. + */ +export async function stage(entry: RegistryEntry): Promise<{ path: string; content: string; preview: string }> { + const body = await fetchText(entry.url, MAX_BODY_BYTES); + + if (entry.kind === 'plugin') { + const manifest = parseManifest(body); + if (manifest.name !== entry.name) { + throw new Error(`the manifest calls itself "${manifest.name}" but the index calls it "${entry.name}"`); + } + const rules = manifest.deny + .map((r) => `- denies ${r.tools.join(', ')}: ${r.pathPattern ?? r.commandPattern}`) + .join('\n'); + return { + path: join(pluginsDir(), `${entry.name}.json`), + content: JSON.stringify(manifest, null, 2), + preview: `${manifest.description}\n\n${rules}`, + }; + } + + const skill = parseSkill(body, 'registry'); + if (!skill) throw new Error('that skill has no name/description frontmatter, so it cannot be loaded'); + if (skill.name !== entry.name) { + throw new Error(`the skill calls itself "${skill.name}" but the index calls it "${entry.name}"`); + } + + return { path: join(skillsDir(), `${entry.name}.md`), content: body, preview: skill.body }; +} + +export async function install(entry: RegistryEntry): Promise { + const { path, content } = await stage(entry); + await Bun.write(path, content); + return { name: entry.name, kind: entry.kind, path }; +} + +/** Removes an installed entry. Returns false when there was nothing to remove. */ +export async function uninstall(kind: RegistryKind, name: string): Promise { + if (!/^[a-z0-9][a-z0-9-]*$/.test(name)) throw new Error(`"${name}" is not a valid entry name`); + const path = kind === 'plugin' ? join(pluginsDir(), `${name}.json`) : join(skillsDir(), `${name}.md`); + const file = Bun.file(path); + if (!(await file.exists())) return false; + await file.delete(); + return true; +} + +/** + * Declarative plugins from disk. + * + * A malformed file is reported and skipped rather than fatal: one bad install + * should not stop the agent from starting. + */ +export async function loadInstalledPlugins(): Promise<{ plugins: Plugin[]; errors: { plugin: string; message: string }[] }> { + const plugins: Plugin[] = []; + const errors: { plugin: string; message: string }[] = []; + + let files: string[] = []; + try { + for await (const f of new Bun.Glob('*.json').scan({ cwd: pluginsDir(), onlyFiles: true })) files.push(f); + } catch { + return { plugins, errors }; + } + + for (const file of files.sort()) { + const name = file.replace(/\.json$/, ''); + try { + plugins.push(manifestToPlugin(parseManifest(await Bun.file(join(pluginsDir(), file)).text()))); + } catch (e) { + errors.push({ plugin: name, message: e instanceof Error ? e.message : String(e) }); + } + } + + return { plugins, errors }; +} + +/** Installed skills, for `/registry list` — the loader already merges them by name. */ +export function installedSkillNames(skills: Skill[]): string[] { + return skills.filter((s) => s.origin === 'registry').map((s) => s.name); +} diff --git a/src/session.ts b/src/session.ts index 6225598..4939631 100644 --- a/src/session.ts +++ b/src/session.ts @@ -83,6 +83,9 @@ export type SessionOptions = { const estimateTokens = (messages: ModelMessage[]) => Math.round(JSON.stringify(messages).length / 4); +/** Estimated tokens at which the wire history is pruned. */ +const DEFAULT_COMPACT_THRESHOLD = 120_000; + export class Session { readonly messages: ModelMessage[]; readonly tools: ToolSet; @@ -164,6 +167,11 @@ export class Session { return estimateTokens(this.messages); } + /** Where compaction kicks in, so the status bar can show how close it is. */ + compactThreshold(): number { + return this.opts.compactThreshold ?? DEFAULT_COMPACT_THRESHOLD; + } + private systemFor(): string { return systemPrompt({ cwd: this.opts.cwd ?? process.cwd(), @@ -234,7 +242,7 @@ export class Session { this.opts.onChange?.(this.messages); this.controller = new AbortController(); const signal = this.controller.signal; - const threshold = this.opts.compactThreshold ?? 120_000; + const threshold = this.compactThreshold(); const outputs: Extract[] = []; onBashOutput(({ toolCallId, chunk }) => { diff --git a/src/skills.ts b/src/skills.ts index bac7a68..3173142 100644 --- a/src/skills.ts +++ b/src/skills.ts @@ -4,7 +4,7 @@ import { join } from 'node:path'; import { z } from 'zod'; import { BUILTIN_SKILLS } from './skills-builtin'; -export type SkillOrigin = 'builtin' | 'user' | 'project'; +export type SkillOrigin = 'builtin' | 'registry' | 'user' | 'project'; export type Skill = { name: string; @@ -37,14 +37,23 @@ export function parseSkill(source: string, origin: SkillOrigin, path?: string): return { name, description, origin, ...(path ? { path } : {}), body: match[2]!.trim().slice(0, MAX_BODY) }; } -const skillDirs = (cwd: string) => [ - { dir: join(process.env['SHIRO_HOME'] ?? homedir(), '.shiro-neko', 'skills'), origin: 'user' as const }, - { dir: join(cwd, '.shiro', 'skills'), origin: 'project' as const }, -]; +/** + * Precedence, low to high. Installed skills sit below your own on purpose: a skill + * you wrote must never be shadowed by one fetched from a registry. + */ +const skillDirs = (cwd: string) => { + const home = join(process.env['SHIRO_HOME'] ?? homedir(), '.shiro-neko'); + return [ + { dir: join(home, 'registry', 'skills'), origin: 'registry' as const }, + { dir: join(home, 'skills'), origin: 'user' as const }, + { dir: join(cwd, '.shiro', 'skills'), origin: 'project' as const }, + ]; +}; /** - * Builtin, then user, then project. Later wins, so a project can override a - * bundled skill by using the same name. + * Builtin, then installed, then user, then project. Later wins, so a project can + * override a bundled or installed skill by using the same name — and a skill you + * wrote yourself always beats one fetched from a registry. */ export async function loadSkills(cwd = process.cwd()): Promise { const byName = new Map(); diff --git a/src/ui/App.tsx b/src/ui/App.tsx index cfa52d0..7d2f769 100644 --- a/src/ui/App.tsx +++ b/src/ui/App.tsx @@ -15,7 +15,7 @@ import { AskPanel, type AskBridge, type AskPending } from './Ask'; import { Diff } from './Diff'; import { Markdown } from './Markdown'; import { Onboard, type OnboardResult } from './Onboard'; -import { InfoPanel, OutputPanel, QueuePanel, StatusBar, SubagentPanel, ThinkingPanel, TodoPanel, ActiveTool, FileMenu, type SubagentView } from './Panels'; +import { InfoPanel, OutputPanel, QueuePanel, RegistryPanel, InstallPrompt, StatusBar, SubagentPanel, ThinkingPanel, TodoPanel, ActiveTool, FileMenu, type RegistryRow, type SubagentView } from './Panels'; import { PromptInput } from './PromptInput'; type Line = @@ -139,6 +139,15 @@ export type AppHooks = { instructionFiles: () => string[]; /** Ignore-aware workspace paths for `@` completion, loaded on first use. */ listPaths: () => Promise; + /** Registry index, installed set, and the install/remove actions. */ + registry: { + list: () => Promise; + installed: () => Promise; + /** Fetches and validates without writing, so the body can be shown first. */ + stage: (name: string) => Promise<{ row: RegistryRow; url: string; preview: string }>; + install: (name: string) => Promise; + remove: (name: string) => Promise; + }; /** Prompt to hand the model for /init. */ initPrompt: string; history: string[]; @@ -198,19 +207,42 @@ function Approval({ pending }: { pending: Pending }) { } function CommandMenu({ matches, index }: { matches: CommandSpec[]; index: number }) { + const width = Math.max(...matches.map((c) => `/${c.name}${c.arg ? ` ${c.arg}` : ''}`.length)) + 1; return ( {matches.map((c, i) => ( - - {i === index ? '> ' : ' '} - {`/${c.name}${c.arg ? ` ${c.arg}` : ''}`.padEnd(18)} {c.summary} - + + {i === index ? '> ' : ' '} + + {`/${c.name}${c.arg ? ` ${c.arg}` : ''}`.padEnd(width)} + + {c.summary} + ))} up/down move | tab complete | enter run | esc dismiss ); } +/** Keyboard wrapper around InstallPrompt, so the prompt itself stays presentational. */ +function InstallConfirm({ + staged, + onDone, +}: { + staged: { row: RegistryRow; url: string; preview: string }; + onDone: (yes: boolean) => void; +}) { + useInput((input, key) => { + const c = input.toLowerCase(); + if (c === 'y' || key.return) onDone(true); + else if (c === 'n' || key.escape) onDone(false); + }); + + return ( + + ); +} + export function App({ session, bridge, @@ -260,8 +292,12 @@ export function App({ const [notebook, setNotebook] = useState(session.notebook.state()); const [agents, setAgents] = useState([]); const [panel, setPanel] = useState<{ title: string; hint?: string; body: string } | undefined>(); + const [registry, setRegistry] = useState<{ title: string; hint?: string; rows: RegistryRow[] } | undefined>(); + const [installing, setInstalling] = useState< + { row: RegistryRow; url: string; preview: string } | undefined + >(); - const modal = pending !== undefined || asking !== undefined || onboarding; + const modal = pending !== undefined || asking !== undefined || onboarding || installing !== undefined; const anyPicker = modelPicker !== undefined || agentPicker || thinkPicker; const matches = matchCommands(draft); const menuOpen = matches.length > 0 && !menuDismissed && !busy && !modal && !anyPicker && !panel; @@ -371,6 +407,10 @@ export function App({ setPanel(undefined); return true; } + if (key.escape && registry) { + setRegistry(undefined); + return true; + } // The file picker gets first refusal: while an `@` token is open its keys // mean navigation, not history recall or command completion. @@ -425,7 +465,7 @@ export function App({ } return false; }, - [draft, fileMatches.length, fileOpen, highlighted, highlightedPath, matches.length, menuOpen, panel, token], + [draft, fileMatches.length, fileOpen, highlighted, highlightedPath, matches.length, menuOpen, panel, registry, token], ); const onDraftChange = useCallback((value: string, at: number) => { @@ -536,6 +576,7 @@ export function App({ setFileIndex(0); setFileDismissed(false); setPanel(undefined); + setRegistry(undefined); // Enter on an open menu runs the highlighted entry, so `/mo` + enter works. const chosen = menuOpen && highlighted ? `/${highlighted.name}` : raw; @@ -676,6 +717,50 @@ export function App({ push({ kind: 'user', text: chosen.trim() }); setPanel({ title: 'plugins', body: hooks.listPlugins() }); return; + case 'registry': { + push({ kind: 'user', text: chosen.trim() }); + setWorking(true); + try { + if (action.action === 'add') { + // Staged, not installed: nothing is written until the prompt is answered. + setInstalling(await hooks.registry.stage(action.arg!)); + return; + } + if (action.action === 'remove') { + push({ kind: 'info', text: await hooks.registry.remove(action.arg!) }); + return; + } + if (action.action === 'installed') { + const rows = await hooks.registry.installed(); + setRegistry({ + title: 'installed', + hint: rows.length === 0 ? 'nothing installed yet' : `${rows.length} from the registry`, + rows, + }); + return; + } + + const all = await hooks.registry.list(); + const rows = + action.action === 'search' && action.arg + ? all.filter( + (r) => + r.name.includes(action.arg!.toLowerCase()) || + r.description.toLowerCase().includes(action.arg!.toLowerCase()), + ) + : all; + setRegistry({ + title: action.arg ? `registry: ${action.arg}` : 'registry', + hint: `${rows.length} of ${all.length} available`, + rows, + }); + } catch (e) { + push({ kind: 'error', text: e instanceof Error ? e.message : String(e) }); + } finally { + setWorking(false); + } + return; + } case 'memory': { push({ kind: 'user', text: chosen.trim() }); setWorking(true); @@ -799,6 +884,29 @@ export function App({ )} + {registry && ( + + )} + + {installing && ( + { + const staged = installing; + setInstalling(undefined); + if (!yes) { + push({ kind: 'info', text: `install cancelled: ${staged.row.name}` }); + return; + } + try { + push({ kind: 'info', text: await hooks.registry.install(staged.row.name) }); + } catch (e) { + push({ kind: 'error', text: e instanceof Error ? e.message : String(e) }); + } + }} + /> + )} + {asking && } {pending && } @@ -937,6 +1045,7 @@ export function App({ agent={hooks.agentName()} thinking={hooks.thinkingLevel()} contextTokens={session.estimatedTokens()} + contextLimit={session.compactThreshold()} cost={(() => { const spend = costOf(hooks.config().model, session.inputTokens, session.outputTokens); return spend === undefined ? 'unpriced' : formatUsd(spend); diff --git a/src/ui/Panels.tsx b/src/ui/Panels.tsx index a4b864d..3410a40 100644 --- a/src/ui/Panels.tsx +++ b/src/ui/Panels.tsx @@ -216,6 +216,7 @@ export function StatusBar({ agent, thinking, contextTokens, + contextLimit, cost, toolCount, }: { @@ -223,14 +224,25 @@ export function StatusBar({ agent: string; thinking: string; contextTokens: number; + /** Threshold compaction fires at, so the bar means something. */ + contextLimit?: number; cost: string; toolCount: number; }) { + const pct = contextLimit ? Math.min(100, Math.round((contextTokens / contextLimit) * 100)) : undefined; + // Amber from two thirds, red once compaction is imminent: the point is to warn + // before a turn silently loses its history, not after. + const contextColor = pct === undefined ? undefined : pct >= 90 ? 'red' : pct >= 66 ? 'yellow' : undefined; + return ( {`${model} `} {agent} - {`/${thinking} ${toolCount} tools ~${contextTokens} ctx ${cost}`} + {`/${thinking} ${toolCount} tools `} + + {pct === undefined ? `~${contextTokens} ctx` : `${pct}% ctx`} + + {` ${cost}`} ); } @@ -258,3 +270,105 @@ export function InfoPanel({ title, hint, lines }: { title: string; hint?: string ); } + +export type RegistryRow = { + name: string; + kind: 'skill' | 'plugin'; + description: string; + author?: string; + installed?: boolean; +}; + +const KIND_COLOR: Record = { skill: 'green', plugin: 'magenta' }; + +/** + * The registry index as a table. + * + * `kind` is coloured rather than spelled out on every row: skill and plugin carry + * very different risk, and a reader scanning the list should see that at a glance. + */ +export function RegistryPanel({ + rows, + hint, + title = 'registry', +}: { + rows: readonly RegistryRow[]; + hint?: string; + title?: string; +}) { + if (rows.length === 0) { + return ; + } + + const width = Math.min(22, Math.max(...rows.map((r) => r.name.length)) + 1); + + return ( + + + {title} + + {hint && {hint}} + {rows.map((r) => ( + + {r.kind === 'skill' ? 'S' : 'P'} + {r.name.padEnd(width)} + {r.description.length > 58 ? `${r.description.slice(0, 58)}...` : r.description} + {r.installed && {' installed'}} + + ))} + {'S skill P plugin | /registry add '} + + ); +} + +/** + * Confirmation before an install writes anything. + * + * A skill body becomes part of the system prompt of every future session in this + * project, so it is shown in full first. The wording says that plainly rather than + * asking a generic "are you sure". + */ +export function InstallPrompt({ + name, + kind, + url, + preview, + lines = 14, +}: { + name: string; + kind: 'skill' | 'plugin'; + url: string; + preview: string; + lines?: number; +}) { + const body = preview.split('\n'); + const shown = body.slice(0, lines); + const hidden = body.length - shown.length; + + return ( + + + {`install ${kind} "${name}"?`} + + {url} + + {shown.map((l, i) => ( + + {` ${l}`} + + ))} + {hidden > 0 && {` ... ${hidden} more lines`}} + + + + {kind === 'skill' + ? 'A skill is instructions the agent follows. This text joins your system prompt.' + : 'A plugin adds refusal rules. It is data, not code: nothing here is executed.'} + + + y install | n cancel + + + + ); +} diff --git a/test/commands.test.ts b/test/commands.test.ts index 9cc2e47..83bedd1 100644 --- a/test/commands.test.ts +++ b/test/commands.test.ts @@ -112,3 +112,49 @@ test('aliases are hidden from the menu but still parse', () => { expect(matchCommands('/lo')).toEqual([]); expect(parseCommand('/login').type).toBe('provider'); }); + +test('/registry with no verb lists everything', () => { + expect(parseCommand('/registry')).toEqual({ type: 'registry', action: 'list' }); + expect(parseCommand('/registry list')).toEqual({ type: 'registry', action: 'list' }); +}); + +test('/registry installed asks for what is already here', () => { + expect(parseCommand('/registry installed')).toEqual({ type: 'registry', action: 'installed' }); +}); + +test('/registry search carries the query', () => { + expect(parseCommand('/registry search migration')).toEqual({ + type: 'registry', + action: 'search', + arg: 'migration', + }); +}); + +test('/registry add and remove carry the name, and their aliases work', () => { + expect(parseCommand('/registry add migration')).toEqual({ type: 'registry', action: 'add', arg: 'migration' }); + expect(parseCommand('/registry install migration')).toEqual({ type: 'registry', action: 'add', arg: 'migration' }); + expect(parseCommand('/registry remove migration')).toEqual({ type: 'registry', action: 'remove', arg: 'migration' }); + expect(parseCommand('/registry uninstall migration')).toEqual({ + type: 'registry', + action: 'remove', + arg: 'migration', + }); +}); + +test('a kind-qualified name survives parsing, since a name can be both', () => { + expect(parseCommand('/registry add plugin:review')).toEqual({ + type: 'registry', + action: 'add', + arg: 'plugin:review', + }); +}); + +test('/registry add with no name returns usage rather than fetching anything', () => { + expect(parseCommand('/registry add')).toEqual({ type: 'info', text: 'usage: /registry add ' }); + expect(parseCommand('/registry remove')).toEqual({ type: 'info', text: 'usage: /registry remove ' }); + expect(parseCommand('/registry search')).toEqual({ type: 'info', text: 'usage: /registry search ' }); +}); + +test('a bare word after /registry is treated as a search', () => { + expect(parseCommand('/registry migration')).toEqual({ type: 'registry', action: 'search', arg: 'migration' }); +}); diff --git a/test/compact.test.ts b/test/compact.test.ts index 9c51132..de4e1ab 100644 --- a/test/compact.test.ts +++ b/test/compact.test.ts @@ -2,6 +2,9 @@ import { expect, test } from 'bun:test'; import { MockLanguageModelV4, simulateReadableStream } from 'ai/test'; import type { LanguageModelV4CallOptions, LanguageModelV4StreamPart } from '@ai-sdk/provider'; import type { ModelMessage } from 'ai'; +import { mkdtempSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; import { Session, type AgentEvent } from '../src/session'; const usage = { @@ -173,3 +176,129 @@ test('onChange fires for every history mutation so autosave stays current', asyn expect(snapshots).toEqual([1, 2]); }); + +/** A reasoning model's turn: a reasoning item, then the tool call, both with item ids. */ +const reasoningToolStep = (n: number): LanguageModelV4StreamPart[] => [ + { type: 'reasoning-start', id: `r${n}`, providerMetadata: { openai: { itemId: `rs_${n}` } } } as never, + { type: 'reasoning-delta', id: `r${n}`, delta: 'deciding what to read next' } as never, + { type: 'reasoning-end', id: `r${n}`, providerMetadata: { openai: { itemId: `rs_${n}` } } } as never, + { type: 'tool-input-start', id: `c${n}`, toolName: 'read_file' }, + { type: 'tool-input-end', id: `c${n}` }, + { + type: 'tool-call', + toolCallId: `c${n}`, + toolName: 'read_file', + input: JSON.stringify({ path: 'big.txt' }), + providerMetadata: { openai: { itemId: `fc_${n}` } }, + } as never, + { type: 'finish', finishReason: { unified: 'tool-calls', raw: 'tool_use' }, usage }, +]; + +function inTempDir(fn: () => Promise): Promise { + const orig = process.cwd(); + const dir = mkdtempSync(join(tmpdir(), 'shiro-compact-')); + process.chdir(dir); + return fn().finally(() => { + process.chdir(orig); + rmSync(dir, { recursive: true, force: true }); + }); +} + +/** + * The bug this guards against: compaction used to drop any assistant part whose + * reasoning item it had pruned. On a reasoning model that is every tool call, so + * after the first compaction the model could no longer see what it had already + * run — and kept re-running it until the step limit stopped the turn. + */ +test('a compacted turn still shows the model the tool calls it already made', async () => + inTempDir(async () => { + await Bun.write(join(process.cwd(), 'big.txt'), 'lorem ipsum dolor sit amet\n'.repeat(1500)); + + const seen: LanguageModelV4CallOptions[] = []; + let call = 0; + const session = new Session({ + compactThreshold: 4000, + maxSteps: 8, + model: new MockLanguageModelV4({ + doStream: async (o) => { + seen.push(o); + const n = call++; + return n < 3 ? stream(reasoningToolStep(n)) : stream(text('read it three times')); + }, + }), + askApproval: async () => 'deny', + }); + + const events: string[] = []; + for await (const ev of session.send('read big.txt a few times')) events.push(ev.type); + + expect(events).toContain('compacted'); + expect(events).toContain('text'); + expect(events.at(-1)).toBe('done'); + // Four calls, not eight: the loop ended because the model chose to, not + // because maxSteps cut it off. + expect(call).toBe(4); + + const shapeOf = (o: LanguageModelV4CallOptions) => + o.prompt + .filter((m) => m.role !== 'system') + .flatMap((m) => (Array.isArray(m.content) ? (m.content as { type: string }[]).map((p) => p.type) : ['str'])); + + // Every call after the first has to carry the earlier exchange. + for (const o of seen.slice(1)) { + const shape = shapeOf(o); + expect(shape).toContain('tool-call'); + expect(shape).toContain('tool-result'); + } + }), 20_000); + +test('a compacted turn sends no assistant item reference whose reasoning was pruned', async () => + inTempDir(async () => { + await Bun.write(join(process.cwd(), 'big.txt'), 'lorem ipsum dolor sit amet\n'.repeat(1500)); + + const seen: LanguageModelV4CallOptions[] = []; + let call = 0; + const session = new Session({ + compactThreshold: 4000, + model: new MockLanguageModelV4({ + doStream: async (o) => { + seen.push(o); + const n = call++; + return n < 3 ? stream(reasoningToolStep(n)) : stream(text('done')); + }, + }), + askApproval: async () => 'deny', + }); + + for await (const _ of session.send('read it')) void _; + + // An itemId on an assistant part is serialised as `item_reference`, which the + // responses API resolves against a stored item that depends on its reasoning + // item. Send one without that reasoning and the request is a 400. An itemId on + // a tool result is harmless: it goes out as a plain function_call_output. + for (const [i, o] of seen.entries()) { + const assistant = o.prompt.filter((m) => m.role === 'assistant'); + const reasoningIds = new Set(); + for (const m of assistant) { + if (!Array.isArray(m.content)) continue; + for (const p of m.content as { type: string; providerOptions?: Record> }[]) { + if (p.type !== 'reasoning') continue; + for (const options of Object.values(p.providerOptions ?? {})) { + if (typeof options['itemId'] === 'string') reasoningIds.add(options['itemId']); + } + } + } + + for (const m of assistant) { + if (!Array.isArray(m.content)) continue; + for (const p of m.content as { type: string; providerOptions?: Record> }[]) { + if (p.type === 'reasoning') continue; + for (const options of Object.values(p.providerOptions ?? {})) { + const id = options['itemId']; + if (typeof id !== 'string') continue; + expect(reasoningIds.size, `call ${i}: ${p.type} references ${id} with no reasoning item`).toBeGreaterThan(0); + } + } + } + } + }), 20_000); diff --git a/test/helpers.ts b/test/helpers.ts index 1682803..492291e 100644 --- a/test/helpers.ts +++ b/test/helpers.ts @@ -21,6 +21,15 @@ export function testHooks(over: Partial = {}): AppHooks { saveSession: async () => 'saved', instructionFiles: () => [], listPaths: async () => [], + registry: { + list: async () => [], + installed: async () => [], + stage: async () => { + throw new Error('no registry in tests unless a test provides one'); + }, + install: async () => 'installed', + remove: async () => 'removed', + }, initPrompt: 'write AGENTS.md', history: [], recordPrompt: () => {}, diff --git a/test/prune.test.ts b/test/prune.test.ts index 0a2c6c2..4eece1c 100644 --- a/test/prune.test.ts +++ b/test/prune.test.ts @@ -1,10 +1,24 @@ import { expect, test } from 'bun:test'; import type { ModelMessage } from 'ai'; -import { dropOrphanedItems, dropOrphanedResults, prunePreservingItems } from '../src/prune'; +import { detachOrphanedItems, dropOrphanedResults, prunePreservingItems } from '../src/prune'; const kinds = (messages: ModelMessage[]) => messages.map((m) => (Array.isArray(m.content) ? `${m.role}:${m.content.map((p) => p.type).join('+')}` : m.role)); +/** Every provider itemId in a message tree, which is what the repair strips. */ +const itemIds = (messages: ModelMessage[]): string[] => { + const found: string[] = []; + for (const m of messages) { + if (!Array.isArray(m.content)) continue; + for (const p of m.content as { providerOptions?: Record> }[]) { + for (const options of Object.values(p.providerOptions ?? {})) { + if (typeof options['itemId'] === 'string') found.push(options['itemId']); + } + } + } + return found; +}; + /** An assistant turn as the OpenAI responses API returns it. */ const reasoningTurn = (rs: string, msg: string, text = 'answer'): ModelMessage => ({ role: 'assistant', @@ -28,24 +42,33 @@ const toolTurn = (rs: string, call: string): ModelMessage => ({ ], }); -test('a message left without its reasoning item is dropped', () => { +/** + * A part carrying an itemId is serialised as `{ type: 'item_reference', id }`, + * which depends on the stored reasoning item. Stripping the id sends the same + * content inline instead, so the turn survives without the dependency. + */ +test('a message left without its reasoning item keeps its text and loses its item id', () => { const before = [{ role: 'user' as const, content: 'q' }, reasoningTurn('rs_1', 'msg_1')]; - const after = [{ role: 'user' as const, content: 'q' }, { role: 'assistant' as const, content: [{ type: 'text' as const, text: 'answer', providerOptions: { openai: { itemId: 'msg_1' } } }] }]; + const after: ModelMessage[] = [ + { role: 'user', content: 'q' }, + { role: 'assistant', content: [{ type: 'text', text: 'answer', providerOptions: { openai: { itemId: 'msg_1' } } }] }, + ]; - const cleaned = dropOrphanedItems(before, after); - expect(JSON.stringify(cleaned)).not.toContain('msg_1'); - expect(kinds(cleaned)).toEqual(['user']); + const cleaned = detachOrphanedItems(before, after); + expect(itemIds(cleaned)).toEqual([]); + expect(kinds(cleaned)).toEqual(['user', 'assistant:text']); + expect(JSON.stringify(cleaned)).toContain('answer'); }); -test('a tool call left without its reasoning item is dropped too', () => { +test('a tool call left without its reasoning item survives, detached', () => { const before = [{ role: 'user' as const, content: 'q' }, toolTurn('rs_1', 'fc_1')]; - const after = [ - { role: 'user' as const, content: 'q' }, + const after: ModelMessage[] = [ + { role: 'user', content: 'q' }, { - role: 'assistant' as const, + role: 'assistant', content: [ { - type: 'tool-call' as const, + type: 'tool-call', toolCallId: 'tc1', toolName: 'grep', input: { pattern: 'x' }, @@ -55,12 +78,60 @@ test('a tool call left without its reasoning item is dropped too', () => { }, ]; - expect(JSON.stringify(dropOrphanedItems(before, after))).not.toContain('fc_1'); + const cleaned = detachOrphanedItems(before, after); + expect(itemIds(cleaned)).toEqual([]); + // The call itself has to stay, or its result is orphaned and the model loses + // any record of what it already ran. + expect(JSON.stringify(cleaned)).toContain('tc1'); + expect(kinds(cleaned)).toEqual(['user', 'assistant:tool-call']); +}); + +test('an empty providerOptions object is removed rather than left behind', () => { + const before = [toolTurn('rs_1', 'fc_1')]; + const after: ModelMessage[] = [ + { + role: 'assistant', + content: [ + { + type: 'tool-call', + toolCallId: 'tc1', + toolName: 'grep', + input: { pattern: 'x' }, + providerOptions: { openai: { itemId: 'fc_1' } }, + }, + ], + }, + ]; + + const part = (detachOrphanedItems(before, after)[0]!.content as Record[])[0]!; + expect('providerOptions' in part).toBe(false); +}); + +test('other provider options are kept when the item id is stripped', () => { + const before: ModelMessage[] = [ + { + role: 'assistant', + content: [ + { type: 'reasoning', text: 't', providerOptions: { openai: { itemId: 'rs_1' } } }, + { type: 'text', text: 'a', providerOptions: { openai: { itemId: 'msg_1', phase: 'final' } } }, + ], + }, + ]; + const after: ModelMessage[] = [ + { + role: 'assistant', + content: [{ type: 'text', text: 'a', providerOptions: { openai: { itemId: 'msg_1', phase: 'final' } } }], + }, + ]; + + const json = JSON.stringify(detachOrphanedItems(before, after)); + expect(json).not.toContain('msg_1'); + expect(json).toContain('final'); }); test('a turn whose reasoning survived is left alone', () => { const messages = [{ role: 'user' as const, content: 'q' }, reasoningTurn('rs_1', 'msg_1')]; - expect(dropOrphanedItems(messages, messages)).toEqual(messages); + expect(detachOrphanedItems(messages, messages)).toEqual(messages); }); test('nothing is touched when no reasoning was removed', () => { @@ -68,13 +139,13 @@ test('nothing is touched when no reasoning was removed', () => { { role: 'user', content: 'q' }, { role: 'assistant', content: 'plain answer' }, ]; - expect(dropOrphanedItems(before, before)).toEqual(before); + expect(detachOrphanedItems(before, before)).toEqual(before); }); -test('parts with no provider itemId are always kept', () => { +test('parts with no provider itemId are returned unchanged', () => { const before = [reasoningTurn('rs_1', 'msg_1')]; const after: ModelMessage[] = [{ role: 'assistant', content: [{ type: 'text', text: 'no item id here' }] }]; - expect(dropOrphanedItems(before, after)).toEqual(after); + expect(detachOrphanedItems(before, after)).toEqual(after); }); test('user and tool messages are never affected', () => { @@ -83,10 +154,10 @@ test('user and tool messages are never affected', () => { { role: 'user', content: 'q' }, { role: 'tool', content: [{ type: 'tool-result', toolCallId: 't1', toolName: 'grep', output: { type: 'text', value: 'hit' } }] }, ]; - expect(dropOrphanedItems(before, after)).toEqual(after); + expect(detachOrphanedItems(before, after)).toEqual(after); }); -test('one orphaned turn does not take a healthy one with it', () => { +test('one orphaned turn does not detach a healthy one with it', () => { const before = [ { role: 'user' as const, content: 'q1' }, reasoningTurn('rs_1', 'msg_1', 'old answer'), @@ -100,14 +171,15 @@ test('one orphaned turn does not take a healthy one with it', () => { reasoningTurn('rs_2', 'msg_2', 'new answer'), ]; - const cleaned = dropOrphanedItems(before, after); + const cleaned = detachOrphanedItems(before, after); const json = JSON.stringify(cleaned); expect(json).not.toContain('msg_1'); + expect(json).toContain('old answer'); expect(json).toContain('msg_2'); expect(json).toContain('rs_2'); }); -test('prunePreservingItems leaves no orphan behind on a real prune', () => { +test('prunePreservingItems leaves no item reference behind on a real prune', () => { const messages: ModelMessage[] = []; for (let i = 0; i < 6; i++) { messages.push({ role: 'user', content: `question ${i} ${'x'.repeat(3000)}` }); @@ -121,12 +193,11 @@ test('prunePreservingItems leaves no orphan behind on a real prune', () => { emptyMessages: 'remove', }); - // Every surviving text part must either have no item id or belong to a turn - // whose reasoning also survived. Since reasoning: 'all' removes them all, no - // itemId-bearing assistant part may remain. - const survivingIds = JSON.stringify(pruned); - for (let i = 0; i < 6; i++) expect(survivingIds).not.toContain(`msg_${i}`); + // reasoning: 'all' removes every reasoning item, so no surviving part may still + // reference one. The text itself stays: that is the model's memory of the turn. + expect(itemIds(pruned)).toEqual([]); expect(pruned.filter((m) => m.role === 'user')).toHaveLength(6); + expect(JSON.stringify(pruned)).toContain('answer 5'); }); test('prunePreservingItems is a no-op when nothing needs pruning', () => { @@ -150,7 +221,10 @@ test('a provider other than openai is handled the same way', () => { const after: ModelMessage[] = [ { role: 'assistant', content: [{ type: 'text', text: 'a', providerOptions: { someProvider: { itemId: 'm1' } } }] }, ]; - expect(dropOrphanedItems(before, after)).toEqual([]); + + const cleaned = detachOrphanedItems(before, after); + expect(itemIds(cleaned)).toEqual([]); + expect(JSON.stringify(cleaned)).toContain('"text":"a"'); }); /** The assistant tool-call plus the tool message answering it, as one exchange. */ diff --git a/test/registry-ui.test.tsx b/test/registry-ui.test.tsx new file mode 100644 index 0000000..01fc90f --- /dev/null +++ b/test/registry-ui.test.tsx @@ -0,0 +1,266 @@ +import { expect, test } from 'bun:test'; +import { render } from 'ink-testing-library'; +import React from 'react'; +import { MockLanguageModelV4, simulateReadableStream } from 'ai/test'; +import { Session } from '../src/session'; +import { App, createApprovalBridge, type AppHooks } from '../src/ui/App'; +import { InstallPrompt, RegistryPanel, type RegistryRow } from '../src/ui/Panels'; +import { testHooks } from './helpers'; + +const usage = { inputTokens: { total: 3, noCache: 3, cacheRead: 0, cacheWrite: 0 }, outputTokens: { total: 1 } } as any; + +const model = new MockLanguageModelV4({ + doStream: async () => + ({ + stream: simulateReadableStream({ + chunks: [ + { type: 'text-start', id: '0' }, + { type: 'text-delta', id: '0', delta: 'reply' }, + { type: 'text-end', id: '0' }, + { type: 'finish', finishReason: { unified: 'stop', raw: 'stop' }, usage }, + ], + chunkDelayInMs: null, + initialDelayInMs: null, + }), + }) as any, +}); + +const rows: RegistryRow[] = [ + { name: 'migration', kind: 'skill', description: 'Write a database migration' }, + { name: 'no-secrets', kind: 'plugin', description: 'Refuses credential writes', installed: true }, +]; + +const wait = (ms: number) => new Promise((r) => setTimeout(r, ms)); + +function mount(over: Partial = {}) { + const bridge = createApprovalBridge(); + const session = new Session({ model, askApproval: bridge.ask }); + const app = render(); + return { app, session }; +} + +async function run(app: ReturnType, command: string, settle = 400) { + for (const ch of command) { + app.stdin.write(ch); + await wait(30); + } + app.stdin.write('\r'); + await wait(settle); +} + +test('the registry panel marks skill and plugin differently and flags what is installed', () => { + const app = render(); + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('migration'); + expect(frame).toContain('no-secrets'); + expect(frame).toContain('installed'); + expect(frame).toContain('S skill P plugin'); + app.unmount(); +}); + +test('an empty registry panel says so rather than rendering an empty box', () => { + const app = render(); + expect(app.lastFrame()).toContain('nothing found'); + app.unmount(); +}); + +test('the install prompt shows the body and says what a skill actually is', () => { + const app = render( + , + ); + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('install skill "migration"?'); + expect(frame).toContain('https://example.com/m.md'); + expect(frame).toContain('Migrations live in db/migrations.'); + expect(frame).toContain('joins your system prompt'); + app.unmount(); +}); + +test('the install prompt says a plugin is data, not code', () => { + const app = render( + , + ); + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('nothing here is executed'); + expect(frame).toContain('denies write_file'); + app.unmount(); +}); + +test('a long body is truncated with a count rather than flooding the screen', () => { + const preview = Array.from({ length: 40 }, (_, i) => `line ${i}`).join('\n'); + const app = render(); + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('line 4'); + expect(frame).not.toContain('line 20'); + expect(frame).toContain('35 more lines'); + app.unmount(); +}); + +test('/registry lists the index', async () => { + const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } }); + await wait(150); + + await run(app, '/registry'); + + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('migration'); + expect(frame).toContain('2 of 2 available'); + + app.unmount(); +}, 20_000); + +test('/registry search narrows the list', async () => { + const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } }); + await wait(150); + + await run(app, '/registry search migra'); + + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('migration'); + expect(frame).not.toContain('no-secrets'); + expect(frame).toContain('1 of 2'); + + app.unmount(); +}, 20_000); + +test('a failing index surfaces the error instead of an empty panel', async () => { + const { app } = mount({ + registry: { + ...testHooks().registry, + list: async () => { + throw new Error('the registry index is not valid JSON'); + }, + }, + }); + await wait(150); + + await run(app, '/registry'); + expect(app.lastFrame()).toContain('not valid JSON'); + + app.unmount(); +}, 20_000); + +test('/registry add stages, shows the body, and installs only on y', async () => { + const installed: string[] = []; + const { app } = mount({ + registry: { + ...testHooks().registry, + stage: async (name) => ({ + row: { name, kind: 'skill', description: 'd' }, + url: 'https://example.com/m.md', + preview: 'Migrations live in db/migrations.', + }), + install: async (name) => { + installed.push(name); + return `installed skill ${name}`; + }, + }, + }); + await wait(150); + + await run(app, '/registry add migration'); + expect(app.lastFrame()).toContain('install skill "migration"?'); + expect(installed).toEqual([]); + + app.stdin.write('y'); + await wait(400); + + expect(installed).toEqual(['migration']); + expect(app.lastFrame()).toContain('installed skill migration'); + + app.unmount(); +}, 20_000); + +test('n cancels the install and nothing is written', async () => { + const installed: string[] = []; + const { app } = mount({ + registry: { + ...testHooks().registry, + stage: async (name) => ({ + row: { name, kind: 'skill', description: 'd' }, + url: 'https://example.com/m.md', + preview: 'body', + }), + install: async (name) => { + installed.push(name); + return 'installed'; + }, + }, + }); + await wait(150); + + await run(app, '/registry add migration'); + app.stdin.write('n'); + await wait(400); + + expect(installed).toEqual([]); + expect(app.lastFrame()).toContain('install cancelled: migration'); + + app.unmount(); +}, 20_000); + +test('a staging failure never reaches the confirmation prompt', async () => { + const { app } = mount({ + registry: { + ...testHooks().registry, + stage: async () => { + throw new Error('no registry entry named "nope"'); + }, + }, + }); + await wait(150); + + await run(app, '/registry add nope'); + + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('no registry entry named'); + expect(frame).not.toContain('install skill'); + + app.unmount(); +}, 20_000); + +test('/registry remove reports what it removed', async () => { + const { app } = mount({ + registry: { ...testHooks().registry, remove: async (name) => `removed skill ${name}` }, + }); + await wait(150); + + await run(app, '/registry remove migration'); + expect(app.lastFrame()).toContain('removed skill migration'); + + app.unmount(); +}, 20_000); + +test('/registry installed shows what is already here', async () => { + const { app } = mount({ + registry: { ...testHooks().registry, installed: async () => [rows[1]!] }, + }); + await wait(150); + + await run(app, '/registry installed'); + + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('no-secrets'); + expect(frame).toContain('1 from the registry'); + + app.unmount(); +}, 20_000); + +test('esc dismisses the registry panel', async () => { + const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } }); + await wait(150); + + await run(app, '/registry'); + expect(app.lastFrame()).toContain('migration'); + + app.stdin.write('\u001B'); + await wait(300); + expect(app.lastFrame()).not.toContain('S skill P plugin'); + + app.unmount(); +}, 20_000); diff --git a/test/registry.test.ts b/test/registry.test.ts new file mode 100644 index 0000000..30b9d15 --- /dev/null +++ b/test/registry.test.ts @@ -0,0 +1,314 @@ +import { afterEach, beforeEach, expect, test } from 'bun:test'; +import { mkdtempSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { + DEFAULT_INDEX_URL, + fetchIndex, + install, + loadInstalledPlugins, + manifestToPlugin, + parseIndex, + parseManifest, + pluginsDir, + searchEntries, + skillsDir, + stage, + uninstall, + type RegistryEntry, +} from '../src/registry'; + +let home: string; +let origHome: string | undefined; + +beforeEach(() => { + origHome = process.env['SHIRO_HOME']; + home = mkdtempSync(join(tmpdir(), 'shiro-reg-')); + process.env['SHIRO_HOME'] = home; +}); + +afterEach(() => { + if (origHome === undefined) delete process.env['SHIRO_HOME']; + else process.env['SHIRO_HOME'] = origHome; + rmSync(home, { recursive: true, force: true }); +}); + +const index = (over: Record = {}) => + JSON.stringify({ + skills: [{ name: 'migration', description: 'Write a database migration', url: 'https://example.com/m.md' }], + plugins: [{ name: 'no-secrets', description: 'Refuses credential writes', url: 'https://example.com/p.json' }], + ...over, + }); + +const manifest = (over: Record = {}) => + JSON.stringify({ + name: 'no-secrets', + description: 'Refuses credential writes', + deny: [{ tools: ['write_file', 'edit_file'], pathPattern: '\\.env$', reason: 'refusing to write a .env file' }], + ...over, + }); + +test('the default index is an https URL', () => { + expect(DEFAULT_INDEX_URL.startsWith('https://')).toBe(true); +}); + +test('an index parses into entries tagged with their kind', () => { + const entries = parseIndex(index()); + expect(entries).toHaveLength(2); + expect(entries.find((e) => e.name === 'migration')?.kind).toBe('skill'); + expect(entries.find((e) => e.name === 'no-secrets')?.kind).toBe('plugin'); +}); + +test('a malformed index is reported rather than half-loaded', () => { + expect(() => parseIndex('not json')).toThrow(/not valid JSON/); + expect(() => parseIndex(JSON.stringify({ skills: [{ name: 'x' }] }))).toThrow(/malformed/); +}); + +test('a name that could escape its directory is rejected', () => { + for (const name of ['../evil', 'a/b', 'UPPER', '.hidden', 'with space']) { + const bad = JSON.stringify({ skills: [{ name, description: 'd', url: 'https://e.com/x.md' }] }); + expect(() => parseIndex(bad), name).toThrow(/malformed/); + } +}); + +test('a non-http url is rejected by the schema', () => { + const bad = JSON.stringify({ skills: [{ name: 'x', description: 'd', url: 'file:///etc/passwd' }] }); + expect(() => parseIndex(bad)).toThrow(/malformed/); +}); + +test('a duplicate name within one kind keeps the first', () => { + const dup = JSON.stringify({ + skills: [ + { name: 'x', description: 'first', url: 'https://e.com/1.md' }, + { name: 'x', description: 'second', url: 'https://e.com/2.md' }, + ], + }); + const entries = parseIndex(dup); + expect(entries).toHaveLength(1); + expect(entries[0]?.description).toBe('first'); +}); + +test('the same name may exist as both a skill and a plugin', () => { + const both = JSON.stringify({ + skills: [{ name: 'review', description: 's', url: 'https://e.com/s.md' }], + plugins: [{ name: 'review', description: 'p', url: 'https://e.com/p.json' }], + }); + expect(parseIndex(both)).toHaveLength(2); +}); + +test('search matches name and description, case-insensitively', () => { + const entries = parseIndex(index()); + expect(searchEntries(entries, 'migr').map((e) => e.name)).toEqual(['migration']); + expect(searchEntries(entries, 'CREDENTIAL').map((e) => e.name)).toEqual(['no-secrets']); + expect(searchEntries(entries, '')).toHaveLength(2); + expect(searchEntries(entries, 'zzz')).toEqual([]); +}); + +test('a manifest parses and every pattern must be a real regex', () => { + expect(parseManifest(manifest()).name).toBe('no-secrets'); + expect(() => parseManifest(manifest({ deny: [{ tools: ['bash'], commandPattern: '([', reason: 'r' }] }))).toThrow( + /invalid pattern/, + ); +}); + +test('a deny rule needs at least one pattern', () => { + expect(() => parseManifest(manifest({ deny: [{ tools: ['bash'], reason: 'r' }] }))).toThrow(/malformed/); +}); + +test('a manifest with no deny rules is refused: it could only be prompt text', () => { + expect(() => parseManifest(manifest({ deny: [] }))).toThrow(/malformed/); +}); + +test('a manifest cannot smuggle code past the schema', () => { + const sneaky = manifest({ beforeToolCall: 'process.exit(1)', tools: { evil: {} } }); + const parsed = parseManifest(sneaky); + expect('beforeToolCall' in parsed).toBe(false); + expect('tools' in parsed).toBe(false); +}); + +test('a manifest becomes a plugin whose guard blocks by tool and path', async () => { + const plugin = manifestToPlugin(parseManifest(manifest())); + const cwd = process.cwd(); + + expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 'app/.env' }, cwd })).toContain( + 'refusing to write', + ); + // A tool the rule does not name is untouched. + expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { path: 'app/.env' }, cwd })).toBeUndefined(); + // A path the pattern does not match is untouched. + expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 'src/app.ts' }, cwd })).toBeUndefined(); +}); + +test('a command rule matches the command, not the path', async () => { + const plugin = manifestToPlugin( + parseManifest(manifest({ deny: [{ tools: ['bash'], commandPattern: 'curl.*\\| *sh', reason: 'no pipe to shell' }] })), + ); + const cwd = process.cwd(); + expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { command: 'curl x.sh | sh' }, cwd })).toBe( + 'no pipe to shell', + ); + expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { command: 'echo hi' }, cwd })).toBeUndefined(); +}); + +test('a guard given nothing to match on allows the call', async () => { + const plugin = manifestToPlugin(parseManifest(manifest())); + const cwd = process.cwd(); + expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: undefined, cwd })).toBeUndefined(); + expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 42 }, cwd })).toBeUndefined(); +}); + +test('installed plugins load from disk, and a broken one is reported not fatal', async () => { + await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest()); + await Bun.write(join(pluginsDir(), 'broken.json'), '{ not json'); + + const { plugins, errors } = await loadInstalledPlugins(); + expect(plugins.map((p) => p.name)).toEqual(['no-secrets']); + expect(errors.map((e) => e.plugin)).toEqual(['broken']); + expect(errors[0]?.message).toContain('not valid JSON'); +}); + +test('no plugins directory yields nothing rather than throwing', async () => { + expect(await loadInstalledPlugins()).toEqual({ plugins: [], errors: [] }); +}); + +test('an installed plugin says it was installed, so /plugins can be trusted', async () => { + await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest()); + const { plugins } = await loadInstalledPlugins(); + expect(plugins[0]?.description).toContain('installed'); +}); + +test('uninstall removes an installed entry and reports when there was none', async () => { + await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest()); + expect(await uninstall('plugin', 'no-secrets')).toBe(true); + expect(await Bun.file(join(pluginsDir(), 'no-secrets.json')).exists()).toBe(false); + expect(await uninstall('plugin', 'no-secrets')).toBe(false); +}); + +test('uninstall refuses a name that is not a plain entry name', async () => { + expect(uninstall('skill', '../../../etc/passwd')).rejects.toThrow(/not a valid entry name/); +}); + +/** A local server stands in for the registry: no network, real HTTP. */ +function serve(routes: Record) { + return Bun.serve({ + port: 0, + fetch(req) { + const path = new URL(req.url).pathname; + const hit = routes[path]; + if (!hit) return new Response('not found', { status: 404 }); + return new Response(hit.body, { status: hit.status ?? 200 }); + }, + }); +} + +test('fetchIndex reads an index over http', async () => { + const server = serve({ '/index.json': { body: index() } }); + try { + const entries = await fetchIndex(`${server.url}index.json`); + expect(entries.map((e) => e.name).sort()).toEqual(['migration', 'no-secrets']); + } finally { + server.stop(true); + } +}); + +test('an index that returns an error status is reported with the status', async () => { + const server = serve({ '/index.json': { body: 'nope', status: 503 } }); + try { + expect(fetchIndex(`${server.url}index.json`)).rejects.toThrow(/503/); + } finally { + server.stop(true); + } +}); + +const SKILL_BODY = `--- +name: migration +description: Write a database migration +--- + +Migrations live in db/migrations and are never renumbered.`; + +test('stage returns what would be written without writing it', async () => { + const server = serve({ '/m.md': { body: SKILL_BODY } }); + try { + const entry: RegistryEntry = { + kind: 'skill', + name: 'migration', + description: 'Write a database migration', + url: `${server.url}m.md`, + }; + + const staged = await stage(entry); + expect(staged.path).toBe(join(skillsDir(), 'migration.md')); + expect(staged.preview).toContain('never renumbered'); + expect(await Bun.file(staged.path).exists()).toBe(false); + } finally { + server.stop(true); + } +}); + +test('install writes the skill where the loader will find it', async () => { + const server = serve({ '/m.md': { body: SKILL_BODY } }); + try { + const installed = await install({ + kind: 'skill', + name: 'migration', + description: 'd', + url: `${server.url}m.md`, + }); + + expect(installed.path).toBe(join(skillsDir(), 'migration.md')); + expect(await Bun.file(installed.path).text()).toContain('name: migration'); + + const { loadSkills } = await import('../src/skills'); + const loaded = await loadSkills(home); + const skill = loaded.find((s) => s.name === 'migration'); + expect(skill?.origin).toBe('registry'); + } finally { + server.stop(true); + } +}); + +test('a body whose name disagrees with the index is refused', async () => { + const server = serve({ '/m.md': { body: SKILL_BODY } }); + try { + expect( + stage({ kind: 'skill', name: 'something-else', description: 'd', url: `${server.url}m.md` }), + ).rejects.toThrow(/calls itself "migration"/); + } finally { + server.stop(true); + } +}); + +test('a skill with no frontmatter is refused rather than installed as prose', async () => { + const server = serve({ '/m.md': { body: 'just some text, no frontmatter' } }); + try { + expect(stage({ kind: 'skill', name: 'migration', description: 'd', url: `${server.url}m.md` })).rejects.toThrow( + /frontmatter/, + ); + } finally { + server.stop(true); + } +}); + +test('a plugin manifest is validated before it can be staged', async () => { + const server = serve({ + '/good.json': { body: manifest() }, + '/bad.json': { body: manifest({ deny: [{ tools: ['bash'], commandPattern: '([', reason: 'r' }] }) }, + }); + try { + const staged = await stage({ + kind: 'plugin', + name: 'no-secrets', + description: 'd', + url: `${server.url}good.json`, + }); + expect(staged.path).toBe(join(pluginsDir(), 'no-secrets.json')); + expect(staged.preview).toContain('denies write_file'); + + expect( + stage({ kind: 'plugin', name: 'no-secrets', description: 'd', url: `${server.url}bad.json` }), + ).rejects.toThrow(/invalid pattern/); + } finally { + server.stop(true); + } +}); diff --git a/test/ui-panels.test.tsx b/test/ui-panels.test.tsx index c522990..f8b03eb 100644 --- a/test/ui-panels.test.tsx +++ b/test/ui-panels.test.tsx @@ -106,6 +106,40 @@ test('the status bar reports model, agent, thinking, context, and spend', () => app.unmount(); }); +test('a context limit turns the raw token count into a percentage', () => { + const app = render( + , + ); + const frame = app.lastFrame() ?? ''; + expect(frame).toContain('50% ctx'); + expect(frame).not.toContain('60000'); + app.unmount(); +}); + +test('the context percentage is capped at 100 rather than running over', () => { + const app = render( + , + ); + expect(app.lastFrame()).toContain('100% ctx'); + app.unmount(); +}); + test('the info panel renders a markdown body', () => { const app = render(); const frame = app.lastFrame() ?? '';