Add batch reads, @file completion, interruptible commands, tool sets
Tools, six built-in to fourteen: - read_many_files: up to 20 paths read concurrently, each with its own window. An unreadable path is reported in its own block instead of throwing. - multi_edit: several edits to one file, validated in memory first so a late failure cannot leave the file half-written. - list_dir: ignore-aware depth-limited tree. - git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather than a shell string, which is what makes them safe to auto-approve. toolSets gates them. core is always on; edit-plus and git are optional. A disabled set reaches neither the wire nor the system prompt, since a prompt naming an absent tool teaches calls that cannot succeed. Interface: - Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn ends: it is progress, not the answer. - The tool in flight is named from tool-input-start, before its arguments finish streaming, and cleared on its result. - Prompts typed mid-turn queue and drain in order. esc clears the queue as well as aborting. - @ opens a path picker fed by the ignore-aware walker. Prefix matches rank above substring matches, so @src/ means "under src/". The walk runs on the first @, not at startup. ctrl-c kills the command in flight and keeps the turn. The call throws rather than returning, so the model cannot read a killed command as one that ran and failed on its own terms. The kill takes the whole process tree: killing cmd /c alone left the real command holding both pipes open, so the read never returned and the interrupt did nothing for 19 seconds. Two pruning fixes: - A tool result whose tool call was pruned is now dropped with it. Pruning counts messages, so the cut landed between an assistant tool-call and the tool message answering it, producing 400 "No tool call found for function call output with call_id ...". The reverse pairing is left alone: a call awaiting its result is what a suspended approval looks like. - ignore.ts called statFs without importing it, so walk() crashed on the first symlink. 482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the two hand-maintained tool-name lists recorded as a known weakness.
This commit is contained in:
+17
-3
@@ -32,9 +32,23 @@ pure latency.
|
||||
**`deep`** asks for more than one hypothesis before acting, more reading before concluding,
|
||||
and findings recorded with `remember` so they survive compaction.
|
||||
|
||||
**`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, and `bash`
|
||||
are withheld from the model, not merely discouraged in prose — a model that cannot see a
|
||||
tool cannot call it. Their prompts also forbid describing edits as if they had been made.
|
||||
**`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, `multi_edit`,
|
||||
and `bash` are withheld from the model, not merely discouraged in prose — a model that cannot
|
||||
see a tool cannot call it. They keep everything that only reads, including `read_many_files`,
|
||||
`list_dir`, and the git tools. Their prompts also forbid describing edits as if they had been
|
||||
made.
|
||||
|
||||
## Variants and tool sets
|
||||
|
||||
Two separate things narrow the tool list, and they compose.
|
||||
|
||||
A variant withholds tools by *capability*: `plan` cannot write, whatever the config says.
|
||||
`toolSets` withholds them by *cost*: a project that never wants the git tools switches that set
|
||||
off for every variant. See [tools](tools.md#tool-sets).
|
||||
|
||||
Both go through one function, so a withheld tool is missing from the wire and from the system
|
||||
prompt together. `/tools` lists what is actually offered this turn, with the set each tool came
|
||||
from.
|
||||
|
||||
## Thinking levels
|
||||
|
||||
|
||||
+85
-12
@@ -62,7 +62,9 @@ That is not an optimisation. A `todo_write` on step one must be visible to step
|
||||
`prepareStep` is the only place per-step state can enter.
|
||||
|
||||
The prompt also describes only the tools actually offered this turn. A prompt that mentions a
|
||||
withheld tool teaches the model to attempt impossible calls.
|
||||
withheld tool teaches the model to attempt impossible calls. Two things narrow that set: a
|
||||
read-only agent variant, and `toolSets` in config. Both go through `activeTools()`, so a
|
||||
withheld tool is absent from the wire and from the prompt together.
|
||||
|
||||
## Rendering
|
||||
|
||||
@@ -74,6 +76,10 @@ Two things fix it:
|
||||
- Finished lines go into `<Static>`, rendered once and never redrawn.
|
||||
- Token deltas accumulate in a ref and flush on a 60 ms interval, not per token.
|
||||
|
||||
Answer text and reasoning text are separate refs on the same interval. Reasoning is shown
|
||||
collapsed as a token estimate, expandable with `ctrl-r`, and dropped when the turn ends: it is
|
||||
progress, not the answer, and keeping it would bury the reply it was leading up to.
|
||||
|
||||
Markdown is parsed on every flush. An unclosed fence renders as a code block that grows,
|
||||
which is what a reader expects while text is still arriving.
|
||||
|
||||
@@ -83,9 +89,59 @@ which is what a reader expects while text is still arriving.
|
||||
recall is impossible, and it only ever *shrinks* its internal cursor offset, so an externally
|
||||
set value leaves the cursor stranded mid-string.
|
||||
|
||||
`src/ui/PromptInput.tsx` owns the cursor. That also gives home, end, and ctrl-a/e/k/u/w for
|
||||
free. It hands up, down, tab, and escape to a parent callback first, so the command menu and
|
||||
open panels can claim them before the input treats them as editing keys.
|
||||
`src/ui/PromptInput.tsx` owns the cursor and reports it with every change, which is what makes
|
||||
`@path` completion possible at all. That also gives home, end, and ctrl-a/e/k/u/w for free. It
|
||||
hands up, down, tab, and escape to a parent callback first, so the file picker, the command
|
||||
menu, and open panels can claim them before the input treats them as editing keys.
|
||||
|
||||
The file picker claims those keys ahead of the command menu. While an `@` token is open, up and
|
||||
down mean "move in the list", not "recall an earlier prompt".
|
||||
|
||||
`src/complete.ts` holds the token extraction, ranking, and insertion as pure functions, so the
|
||||
rules are testable without a terminal. Two of them are decisions rather than mechanics:
|
||||
|
||||
- The `@` must start a word, or `user@host` opens a file picker.
|
||||
- Prefix matches rank above substring matches, because `@src/` means "under `src/`" and a
|
||||
substring hit on `vendor/src/` would bury what the user pointed at.
|
||||
|
||||
## The prompt queue
|
||||
|
||||
The input stays mounted while the model works. A prompt submitted mid-turn is pushed onto a
|
||||
queue and drained in order when the turn ends, going back through `submit` so a queued slash
|
||||
command behaves exactly as if it were typed at that moment.
|
||||
|
||||
The queue is a ref as well as state. The drain runs synchronously as the turn ends, between
|
||||
renders, and a closure over a stale array would silently lose a prompt. `busy` is mirrored into
|
||||
a ref for the same reason.
|
||||
|
||||
`esc` clears the queue as well as aborting. Interrupting and then watching two more prompts
|
||||
fire anyway is not what anyone means by interrupt.
|
||||
|
||||
## Interrupting one command
|
||||
|
||||
`esc` aborts the whole turn. That is the wrong tool for a runaway command, because it throws
|
||||
away the conversation to stop a `sleep`.
|
||||
|
||||
`ctrl-c` kills the command in flight and leaves the turn alive. `src/tools.ts` keeps the
|
||||
running processes by tool call id, and `interruptBash()` kills them and returns what it killed.
|
||||
The call then **throws** rather than returning:
|
||||
|
||||
```
|
||||
The user interrupted this command. It did not finish, so its effects are unknown.
|
||||
```
|
||||
|
||||
Throwing is the point. A returned `exit: 1` reads to the model as a command that ran and
|
||||
failed on its own terms, which is a different fact from a command that was stopped partway.
|
||||
The model gets a tool error, and the loop continues to the next step.
|
||||
|
||||
The kill has to take the whole process tree. `cmd /c` and `bash -lc` run the real command as a
|
||||
child, and killing the shell alone leaves that child holding both pipes open, so the read never
|
||||
returns — measured at 19 seconds for a `ping -n 20` that should have died instantly. On Windows
|
||||
that means `taskkill /T /F`. The kill is also awaited before the tool returns, because a
|
||||
surviving grandchild keeps the working directory locked.
|
||||
|
||||
Ink's own `exitOnCtrlC` is turned off in `cli.tsx` so the key reaches the app; with nothing
|
||||
running, the handler exits as usual.
|
||||
|
||||
## Subagents
|
||||
|
||||
@@ -120,22 +176,38 @@ Only `api.openai.com` gets the chain. Third-party endpoints do not implement `/v
|
||||
|
||||
## Compaction and its repair
|
||||
|
||||
`pruneMessages({ reasoning: 'all' })` strips a reasoning item and keeps the message item from
|
||||
the same response. The responses API treats the message as that reasoning item's dependent
|
||||
and returns 400.
|
||||
Pruning breaks two different provider invariants, and `src/prune.ts` repairs both.
|
||||
|
||||
**A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a
|
||||
reasoning item and keeps the message item from the same response. The responses API treats the
|
||||
message as that reasoning item's dependent and returns 400.
|
||||
|
||||
The two carry different ids, so they cannot be matched by id. What links them is the assistant
|
||||
message they arrived in: one message is one response, and its reasoning item covers every
|
||||
other item in it. `src/prune.ts` drops the dependent parts of any turn whose reasoning was
|
||||
other item in it. `dropOrphanedItems` drops the dependent parts of any turn whose reasoning was
|
||||
removed — which costs nothing, since pruning was already discarding those turns.
|
||||
|
||||
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
|
||||
*messages*, so the cut lands between an assistant `tool-call` and the `tool` message answering
|
||||
it. What reaches the wire is a `function_call_output` with no `function_call`:
|
||||
|
||||
```
|
||||
400 No tool call found for function call output with call_id call_…
|
||||
```
|
||||
|
||||
`dropOrphanedResults` collects the surviving call ids and drops any result that has none. The
|
||||
reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a
|
||||
suspended approval looks like, and dropping it would break resume.
|
||||
|
||||
## Module map
|
||||
|
||||
| Module | Responsibility |
|
||||
|---|---|
|
||||
| `session.ts` | the loop, approvals, compaction, event stream |
|
||||
| `tools.ts` | file and shell tools, ripgrep bridge, bash streaming |
|
||||
| `tools.ts` | file and shell tools, tool sets, ripgrep bridge, bash streaming and interrupt |
|
||||
| `tools-git.ts` | read-only git tools, spawned with a fixed argv |
|
||||
| `ignore.ts` | gitignore-aware walker, path jail |
|
||||
| `complete.ts` | `@path` token extraction, ranking, insertion |
|
||||
| `prompt.ts` | system prompt assembly from live state |
|
||||
| `agents.ts` | variants, thinking levels |
|
||||
| `skills.ts` | discovery, catalogue, `skill` tool |
|
||||
@@ -146,7 +218,7 @@ removed — which costs nothing, since pruning was already discarding those turn
|
||||
| `ask.ts` | the `ask` tool |
|
||||
| `mcp.ts` | MCP clients and namespacing |
|
||||
| `fallback.ts` | endpoint chain |
|
||||
| `prune.ts` | provider-item repair |
|
||||
| `prune.ts` | provider-item and tool-pairing repair |
|
||||
| `markdown.ts` | parser, no dependency |
|
||||
| `store.ts` | sessions, prompt history |
|
||||
| `config.ts` | resolution, model construction |
|
||||
@@ -162,9 +234,10 @@ Every module is pure of the UI except `ui/`, and `ui/` never touches the SDK. Th
|
||||
|
||||
## Testing
|
||||
|
||||
404 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop;
|
||||
482 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop;
|
||||
`ink-testing-library` drives the UI with real keystrokes; MCP is tested against a real stdio
|
||||
server subprocess; provider wire formats are tested against a local HTTP server.
|
||||
server subprocess; provider wire formats are tested against a local HTTP server; the interrupt
|
||||
path spawns a real subprocess and asserts it died early rather than ran out.
|
||||
|
||||
The pattern throughout is to assert on what actually crossed a boundary — what went on the
|
||||
wire, what is on screen, what is on disk — rather than on internal calls.
|
||||
|
||||
@@ -21,6 +21,7 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
"thinking": "medium",
|
||||
"maxRetries": 3,
|
||||
"plugins": ["guard", "time"],
|
||||
"toolSets": ["edit-plus", "git"],
|
||||
"mcpServers": {
|
||||
"fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] }
|
||||
}
|
||||
@@ -38,6 +39,7 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
| `thinking` | default level: `off`, `low`, `medium`, `high`, `max` |
|
||||
| `maxRetries` | retries per model call for transient failures. Default 3 |
|
||||
| `plugins` | which plugins to enable. Omit for `["guard", "time"]` |
|
||||
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`. Omit for all of them. See [tools](tools.md) |
|
||||
| `mcpServers` | see [MCP](mcp.md) |
|
||||
|
||||
## Provider presets
|
||||
|
||||
+23
-9
@@ -16,7 +16,7 @@ faster and the fallback path is exercised without it.
|
||||
```bash
|
||||
bun run shiro # run from source
|
||||
bun run typecheck # tsc --noEmit
|
||||
bun test # 404 tests
|
||||
bun test # 482 tests
|
||||
bun run build # single binary for this platform -> dist/shiro
|
||||
bun run release # all five platforms -> dist/release + SHA256SUMS
|
||||
bun run install:local # build, then copy onto PATH
|
||||
@@ -69,21 +69,32 @@ mock-verification test:
|
||||
|
||||
- `pruneMessages` leaving a message item without its reasoning item — visible only in the
|
||||
request body
|
||||
- `pruneMessages` leaving a tool result without its tool call — same, and it took a stub
|
||||
endpoint that rejected the pairing to prove the fix
|
||||
- `--json` serialising `Error` as `{}` — visible only in the printed output
|
||||
- Automatic approval requests prompting the user — visible only in the event sequence
|
||||
- `ctrl-c` killing `cmd /c` but not the command under it — visible only as elapsed time, since
|
||||
the interrupt reported success while the command ran for another 19 seconds
|
||||
|
||||
## Adding a tool
|
||||
|
||||
1. Define it in `src/tools.ts` with a `zod` schema. Descriptions are read by the model, so
|
||||
write them as guidance, not as documentation.
|
||||
2. Add it to the `tools` object.
|
||||
3. If it mutates anything, add it to `MUTATING_TOOLS` so it requires approval.
|
||||
4. Add a line to `TOOL_DOCS` in `src/prompt.ts` saying *when* to reach for it.
|
||||
5. Test the behaviour in a temp directory, including the failure path.
|
||||
3. Add it to a set in `TOOL_SETS`. A tool in no set can never be gated off.
|
||||
4. If it mutates anything, add it to `MUTATING_TOOLS` so it requires approval.
|
||||
5. Add a line to `TOOL_DOCS` in `src/prompt.ts` saying *when* to reach for it.
|
||||
6. If it is read-only, add it to `READ_ONLY` in `src/agents.ts` so `plan` and `review` can use
|
||||
it.
|
||||
7. Test the behaviour in a temp directory, including the failure path.
|
||||
|
||||
Every tool costs roughly 550 characters of schema on every request. Thirteen live tools is
|
||||
already where selection accuracy starts to matter, so a new tool needs to earn its place —
|
||||
see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
|
||||
Steps 3 and 4 are two hand-maintained lists of tool names, which is a known weakness: a tool
|
||||
added to one and forgotten in the other is a silently ungated write. Deriving both from the
|
||||
tool definitions is on [TODO.md](../TODO.md).
|
||||
|
||||
Every tool costs roughly 550 characters of schema on every request. Fourteen built-in tools is
|
||||
well past where selection accuracy starts to matter, which is why sets exist and why a new tool
|
||||
needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
|
||||
|
||||
## Adding a slash command
|
||||
|
||||
@@ -149,10 +160,13 @@ handling differs — a Windows-only break is invisible on Linux until someone hi
|
||||
## Debugging the agent itself
|
||||
|
||||
`--no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp` strips it to
|
||||
the seven core tools, which isolates whether a problem is the loop or something layered on it.
|
||||
the built-in tools alone, which isolates whether a problem is the loop or something layered on
|
||||
it. `{ "toolSets": [] }` narrows it further, to the six core tools.
|
||||
|
||||
`--json` in headless mode shows the exact event sequence.
|
||||
|
||||
For provider issues, a local `Bun.serve` that logs the request body and returns a canned SSE
|
||||
stream answers "what did we actually send" faster than any amount of reading. Several bugs in
|
||||
this codebase were found that way.
|
||||
this codebase were found that way. Making that stub *reject* the thing you think you fixed is
|
||||
better still: the tool-pairing repair was confirmed by a stub that returned the real 400 for an
|
||||
orphaned result, then stopped doing so.
|
||||
|
||||
+16
-6
@@ -16,13 +16,14 @@ There is no terminal to approve on, so every gated tool is denied unless `--yolo
|
||||
|
||||
```
|
||||
$ shiro -p "add a test for paginate()"
|
||||
shiro: headless denies write_file, edit_file, bash and mcp tools unless --yolo is passed
|
||||
shiro: headless denies write_file, edit_file, multi_edit, bash and mcp tools unless --yolo is passed
|
||||
[tool] write_file {"path":"test/paginate.test.ts",...}
|
||||
[denied] write_file (run with --yolo to allow tool use in headless mode)
|
||||
```
|
||||
|
||||
Read-only tools work either way, so `-p` without `--yolo` is a safe way to ask questions
|
||||
about a codebase from a script.
|
||||
about a codebase from a script. That includes `read_many_files`, `list_dir`, and the git tools,
|
||||
which is enough to review a diff or explain a module without any write access at all.
|
||||
|
||||
**`--yolo` does not disable plugin guards.** `rm -rf` is still refused.
|
||||
|
||||
@@ -47,14 +48,18 @@ src/prune.ts repairs provider-item dependencies after pruneMessages strips reaso
|
||||
|
||||
```bash
|
||||
$ shiro -p "count the tools" --json
|
||||
{"type":"tool-start","id":"c1","name":"grep"}
|
||||
{"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}}
|
||||
{"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."}
|
||||
{"type":"text","text":"There are 6 built-in file and shell tools."}
|
||||
{"type":"text","text":"There are 14 built-in tools."}
|
||||
{"type":"done","inputTokens":4210,"outputTokens":88}
|
||||
```
|
||||
|
||||
Event types: `text`, `reasoning`, `tool-call`, `tool-output`, `tool-result`, `tool-error`,
|
||||
`tool-denied`, `compacted`, `notice`, `error`, `done`.
|
||||
Event types: `text`, `reasoning`, `tool-start`, `tool-call`, `tool-output`, `tool-result`,
|
||||
`tool-error`, `tool-denied`, `compacted`, `notice`, `error`, `done`.
|
||||
|
||||
`tool-start` arrives before the arguments have finished streaming, so it carries the name but
|
||||
no input. Use `tool-call` when you need the arguments.
|
||||
|
||||
Errors are flattened to message strings, because `JSON.stringify` turns an `Error` into `{}`
|
||||
and a JSON stream that reports failures as empty objects is useless for the one case it
|
||||
@@ -85,6 +90,10 @@ is told to decide and state its assumption instead.
|
||||
|
||||
Subagent progress events are not emitted; the report still comes back.
|
||||
|
||||
There is no terminal, so `ctrl-c` cannot interrupt a single command the way it does
|
||||
interactively — a signal kills the run. Cap the risk with the `timeout` the model passes to
|
||||
`bash`, or with `--agent quick` to cap the step count.
|
||||
|
||||
## CI recipes
|
||||
|
||||
Review a pull request diff:
|
||||
@@ -125,4 +134,5 @@ env:
|
||||
## Cost control
|
||||
|
||||
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
|
||||
step count at 12. There is no spend ceiling yet — see [ROADMAP.md](../ROADMAP.md).
|
||||
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
|
||||
spend ceiling yet — see [TODO.md](../TODO.md).
|
||||
|
||||
+28
-5
@@ -135,23 +135,46 @@ and outcomes, what remains — and it replaces the transcript entirely.
|
||||
|
||||
### The pruning repair
|
||||
|
||||
`pruneMessages({ reasoning: 'all' })` strips a reasoning item and keeps the message item from
|
||||
the same response. The OpenAI responses API treats the message as a dependent of that
|
||||
reasoning item and rejects the request:
|
||||
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
|
||||
before it did.
|
||||
|
||||
**A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a
|
||||
reasoning item and keeps the message item from the same response. The OpenAI responses API
|
||||
treats the message as a dependent of that reasoning item and rejects the request:
|
||||
|
||||
```
|
||||
400 Item 'msg_…' of type 'message' was provided without its required 'reasoning' item: 'rs_…'
|
||||
```
|
||||
|
||||
The two carry different ids, so they cannot be matched by id. What links them is the
|
||||
assistant message they arrived in — one message is one response. `src/prune.ts` drops the
|
||||
assistant message they arrived in — one message is one response. `dropOrphanedItems` drops the
|
||||
dependent parts of any turn whose reasoning was removed. That costs nothing, because pruning
|
||||
was already discarding those turns.
|
||||
|
||||
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
|
||||
*messages*, not pairs, so the cut can land between the assistant message holding a `tool-call`
|
||||
and the `tool` message answering it:
|
||||
|
||||
```
|
||||
400 No tool call found for function call output with call_id call_…
|
||||
```
|
||||
|
||||
`dropOrphanedResults` drops any result whose call id no longer survives. The reverse is left
|
||||
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
|
||||
like, and dropping it would break `/resume`.
|
||||
|
||||
## Prompt history
|
||||
|
||||
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the
|
||||
input walk it; down past the newest restores what you were typing.
|
||||
input walk it; down past the newest restores what you were typing. While an `@` token is open
|
||||
those keys move in the file picker instead.
|
||||
|
||||
Stored at `~/.shiro-neko/history/<hash>.json`, where the hash is a SHA-256 prefix of the
|
||||
project path.
|
||||
|
||||
## The prompt queue
|
||||
|
||||
Not persisted, and deliberately so. A prompt typed during a turn lives in memory until the turn
|
||||
ends, then runs. `esc` clears it along with aborting the turn, and quitting discards it — a
|
||||
queued thought that fires on next launch, against a workspace that has since changed, is worse
|
||||
than a lost one.
|
||||
|
||||
+8
-1
@@ -42,6 +42,10 @@ Two decisions worth knowing about:
|
||||
**Blocks are checked before approval.** `--yolo` skips prompts; it does not skip guards. A
|
||||
plugin block is a refusal, not a permission question.
|
||||
|
||||
**A guard sees `bash` before the command runs, not while it runs.** The guard is the only thing
|
||||
that can refuse a command outright; once one is running, `ctrl-c` is what stops it. Both matter:
|
||||
a pattern the guard does not know about is still interruptible by hand.
|
||||
|
||||
## Builtins
|
||||
|
||||
### `guard` (default on)
|
||||
@@ -98,7 +102,7 @@ export const noSecretsPlugin: Plugin = {
|
||||
'The no-secrets plugin refuses writes to .env and credential files. Ask the user to ' +
|
||||
'add secrets themselves rather than working around it.',
|
||||
beforeToolCall: ({ toolName, input }) => {
|
||||
if (toolName !== 'write_file' && toolName !== 'edit_file') return undefined;
|
||||
if (toolName !== 'write_file' && toolName !== 'edit_file' && toolName !== 'multi_edit') return undefined;
|
||||
const path = String((input as { path?: unknown } | null)?.path ?? '');
|
||||
if (/(^|\/)\.env|credentials|\.pem$/.test(path)) {
|
||||
return `refusing to write ${path}; add secrets yourself`;
|
||||
@@ -110,6 +114,9 @@ export const noSecretsPlugin: Plugin = {
|
||||
|
||||
Then add it to `BUILTIN_PLUGINS` and, if it should be on by default, `DEFAULT_ENABLED`.
|
||||
|
||||
Note the three tool names. Every write tool has to be listed, and `multi_edit` is easy to miss
|
||||
— a guard that only checks `write_file` and `edit_file` is bypassed by a batch edit.
|
||||
|
||||
Write the `appendix` whenever the plugin can block something. Without it the model hits a
|
||||
refusal it was never told about and tries to route around it.
|
||||
|
||||
|
||||
+112
-6
@@ -4,14 +4,15 @@
|
||||
|
||||
Three categories.
|
||||
|
||||
**Free.** Read-only, no prompt: `read_file`, `glob`, `grep`, `task`.
|
||||
**Free.** Read-only, no prompt: `read_file`, `read_many_files`, `glob`, `grep`, `list_dir`,
|
||||
`task`, and the whole git set.
|
||||
|
||||
**Session tools.** Also free, because they touch the agent's own state rather than your
|
||||
files: `todo_write`, `remember`, `recall`, `forget`, `skill`, `ask`, and anything a plugin
|
||||
marks auto-approved.
|
||||
|
||||
**Gated.** Every call stops for a decision: `write_file`, `edit_file`, `bash`, and every
|
||||
`mcp__*` tool.
|
||||
**Gated.** Every call stops for a decision: `write_file`, `edit_file`, `multi_edit`, `bash`,
|
||||
and every `mcp__*` tool.
|
||||
|
||||
```
|
||||
edit_file wants to run
|
||||
@@ -29,6 +30,28 @@ to ask what to do instead. `--yolo` skips all prompts.
|
||||
**The guard runs before all of this.** It is not an approval — it is a refusal, and `--yolo`
|
||||
does not reach it. See [plugins](plugins.md).
|
||||
|
||||
## Tool sets
|
||||
|
||||
Each tool costs roughly 550 characters of JSON schema on every request, and selection
|
||||
accuracy drops as the list grows. Sets let you switch off what a project does not need:
|
||||
|
||||
| Set | Tools |
|
||||
|---|---|
|
||||
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` |
|
||||
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` |
|
||||
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` |
|
||||
|
||||
```json
|
||||
{ "toolSets": ["edit-plus"] }
|
||||
```
|
||||
|
||||
Omit `toolSets` for all of them. `core` is always on — without read, edit, and bash the
|
||||
agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since
|
||||
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
|
||||
Session, plugin, and MCP tools are not part of this budget and are never gated here.
|
||||
|
||||
`/tools` shows which set each live tool came from.
|
||||
|
||||
## File tools
|
||||
|
||||
### `read_file`
|
||||
@@ -43,6 +66,27 @@ Returns contents with 1-based line numbers. Refuses binaries: a NUL byte in the
|
||||
means the file is not text, and a model that reads a 90 MB executable has burned its whole
|
||||
context on nothing.
|
||||
|
||||
### `read_many_files`
|
||||
|
||||
```
|
||||
files [{ path, offset?, limit? }], at most 20
|
||||
```
|
||||
|
||||
One round trip for several files, each with its own window. Reads run concurrently and the
|
||||
blocks come back in the order given, labelled:
|
||||
|
||||
```
|
||||
===== src/app.ts =====
|
||||
1: export const port = 8080;
|
||||
|
||||
===== src/gone.ts =====
|
||||
[unreadable: No such file: src/gone.ts]
|
||||
```
|
||||
|
||||
A path that cannot be read is reported in its own block rather than throwing, so one wrong
|
||||
guess costs a line instead of the whole call. Numbering and binary refusal are the same code
|
||||
path as `read_file`, so a batch read cannot drift from a single one.
|
||||
|
||||
### `write_file`
|
||||
|
||||
```
|
||||
@@ -65,6 +109,32 @@ replaceAll replace every occurrence instead of requiring exactly one
|
||||
An ambiguous match is an error naming the count, which pushes the model to add surrounding
|
||||
context rather than guessing which occurrence it meant.
|
||||
|
||||
### `multi_edit`
|
||||
|
||||
```
|
||||
path file path
|
||||
edits [{ oldString, newString, replaceAll? }], in the order to apply them
|
||||
```
|
||||
|
||||
Several edits to one file in one call, one approval, one write. Each edit sees the result of
|
||||
the previous one, so edits may build on each other.
|
||||
|
||||
Atomic: every edit is validated and applied in memory first, so a failure on the third edit
|
||||
leaves the file exactly as it was rather than half-changed. The same uniqueness rule as
|
||||
`edit_file` applies per edit, and the error names which edit failed.
|
||||
|
||||
### `list_dir`
|
||||
|
||||
```
|
||||
path directory, relative to the workspace root, default the root
|
||||
depth levels to descend, 1-6, default 2
|
||||
includeIgnored also show files git ignores
|
||||
```
|
||||
|
||||
Tree view honouring `.gitignore`. Directories end with `/`, files show their size. Past the
|
||||
depth limit the containing directory is still listed, so the shape of the tree stays visible
|
||||
without its contents. Capped at 300 entries.
|
||||
|
||||
### `glob`
|
||||
|
||||
```
|
||||
@@ -75,7 +145,9 @@ includeIgnored also return files git ignores
|
||||
|
||||
Walks the tree honouring `.gitignore` and `.shiroignore`, skipping `.git` and
|
||||
`node_modules` unconditionally. Nested ignore files apply only within their own directory,
|
||||
as git does. Returns posix paths relative to the workspace root.
|
||||
as git does. Returns posix paths relative to the workspace root. A symlinked directory is
|
||||
classified as a directory and not descended into, since it can point anywhere including
|
||||
back into the tree.
|
||||
|
||||
### `grep`
|
||||
|
||||
@@ -104,6 +176,39 @@ command that fills one while you block on the other deadlocks.
|
||||
|
||||
Returns exit code, stdout, stderr, and a note if a signal killed it.
|
||||
|
||||
**`ctrl-c` interrupts the command, not the turn.** The shell and everything it started are
|
||||
killed — on Windows through `taskkill /T`, because killing `cmd` alone leaves the real command
|
||||
holding both pipes open and the read never ends. The call then fails rather than returning,
|
||||
so the model cannot mistake a killed command for one that ran and failed on its own:
|
||||
|
||||
```
|
||||
The user interrupted this command. It did not finish, so its effects are unknown.
|
||||
stdout:
|
||||
[whatever it printed first]
|
||||
```
|
||||
|
||||
The turn continues from there. `esc` still aborts everything, and `ctrl-c` with nothing
|
||||
running quits as usual.
|
||||
|
||||
## Git tools
|
||||
|
||||
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
|
||||
array rather than a shell string, so an argument like `--author="; rm -rf /"` can only ever
|
||||
be a literal argument — which is what makes auto-approval safe.
|
||||
|
||||
Output is described rather than raw porcelain: `git_status` names the branch and says
|
||||
`staged modified` or `untracked` per file instead of leaving the model to decode two columns
|
||||
of flags. Outside a repository they fail with `<cwd> is not a git repository.` rather than
|
||||
passing git's own error text through.
|
||||
|
||||
```
|
||||
git_status branch, staged, modified, untracked
|
||||
git_diff staged? path? unified diff of uncommitted changes
|
||||
git_log limit? path? hash, date, author, subject; newest first
|
||||
git_show ref path? one commit: message, author, diff
|
||||
git_blame path startLine? endLine? who last changed each line
|
||||
```
|
||||
|
||||
## Agent tools
|
||||
|
||||
### `task`
|
||||
@@ -168,5 +273,6 @@ every call rather than being assumed.
|
||||
## Output caps
|
||||
|
||||
Any single tool result is truncated at 30,000 characters with a note saying how much was
|
||||
cut. `grep` stops at 200 hits, `glob` at 200 paths, `read_file` at 2000 lines by default.
|
||||
Without caps one `grep` for `function` can end a session.
|
||||
cut. `grep` stops at 200 hits, `glob` at 200 paths, `list_dir` at 300 entries,
|
||||
`read_many_files` at 20 files, `read_file` at 2000 lines by default. Without caps one `grep`
|
||||
for `function` can end a session.
|
||||
|
||||
Reference in New Issue
Block a user