diff --git a/README.md b/README.md index aaf0ad9..e7a6661 100644 --- a/README.md +++ b/README.md @@ -49,9 +49,9 @@ from the models that endpoint actually reports. Settings land in shiro-neko 0.1.0-beta.3 openai/gpt-5 session 0193ab2c agent: default thinking: medium cwd: /home/you/project -skills: debug, refactor, review, test +skills: commit, debug, refactor, review, test, verify plugins: guard, time -approvals: ask for write_file, edit_file, multi_edit, bash, mcp__* +approvals: ask for write_file, edit_file, multi_edit, apply_patch, bash, web_fetch, mcp__* /help for commands > why does the pagination test fail? @@ -64,16 +64,19 @@ installed and honours `.gitignore`. `list_dir` gives an ignore-aware tree so it blindly to orient, and `read_many_files` pulls a batch in one round trip. `read_file` refuses binaries rather than filling the context with mojibake. -**Edits with your approval, gated per command.** `write_file`, `edit_file`, `multi_edit`, and -`bash` stop for a `y`/`a`/`n` decision, with a coloured diff for edits. Rules match the command or -path rather than the tool, so `git *` can run unprompted while everything else still asks — -answering `a` whitelists that pattern, not the whole tool. `.env` and `.pem` files are refused on -read outright. The `guard` plugin refuses irreversible commands ahead of any of it — `rm -rf`, -`git reset --hard`, force pushes, `DROP TABLE` — and `--yolo` cannot bypass it. +**Edits with your approval, gated per command.** `write_file`, `edit_file`, `multi_edit`, +`apply_patch`, and `bash` stop for a `y`/`a`/`n` decision, with a coloured diff for edits. +`apply_patch` lands one atomic patch across files — add, update, move, delete — and nothing is +written if any part of it fails. Rules match the command or path rather than the tool, so +`git *` can run unprompted while everything else still asks — answering `a` whitelists that +pattern, not the whole tool. `.env` and `.pem` files are refused on read outright. The `guard` +plugin refuses irreversible commands ahead of any of it — `rm -rf`, `git reset --hard`, force +pushes, `DROP TABLE` — and `--yolo` cannot bypass it. **Shows its work.** Reasoning streams to a collapsed panel you can expand with `ctrl-r`, the -tool in flight is named as it runs, and `bash` output streams live instead of arriving all at -once when the command exits. `ctrl-c` kills a runaway command without ending the turn. +tool in flight is named as it runs with the arguments that identify the call, and `bash` +output streams live instead of arriving all at once when the command exits. `ctrl-c` kills a +runaway command without ending the turn. **Takes prompts while it works.** Type during a turn and it queues; the queue drains in order when the turn ends. `esc` interrupts and clears it. `@` completes workspace paths. @@ -82,12 +85,18 @@ when the turn ends. `esc` interrupts and clears it. `@` completes workspace path `git_blame` are approval-free, because they spawn git with a fixed argument list and cannot mutate anything. +**Fetches docs when the codebase cannot answer.** `web_fetch` pulls a public page and returns +it as markdown — a changelog, an RFC, a migration guide — size-capped and stripped of anything +that is not text. It lives in the opt-in `net` tool set: the one tool that leaves the machine +is a decision rather than a default, and it asks before every call. + **Asks instead of guessing.** When a request has two readings that lead to different work, the agent puts a question on screen with options. -**Delegates searches.** `task` spawns a read-only subagent whose findings come back as one -message, so a search across forty files does not fill the main context. Its progress -streams to a panel. +**Delegates work.** `task` spawns a subagent with its own context window whose findings come +back as one message, so a search across forty files does not fill the main context. `explore` +and `review` are read-only; `worker` also edits and runs commands, and every one of its writes +stops at the same approval prompt as yours. Progress streams to a panel. **Extensible from the prompt.** `/registry` browses external skills and plugins and installs them with one confirmation. A skill is shown in full before its text joins your system prompt; @@ -97,11 +106,13 @@ a plugin is a manifest of refusal rules, never code. memory that is injected at the start of every future session. **Survives long tasks.** The task list and project memory live outside the message array, -so they survive both automatic pruning and `/compact`. +so they survive both automatic pruning and `/compact`. Pruning itself is bounded: it drops +reasoning first and keeps the widest recent tool tail that fits, so the model keeps its +record of what it already ran instead of repeating it. **Runs headless.** `shiro -p "review this diff" --json` for scripts and CI. -**Keeps the tool list affordable.** Fourteen built-in tools, grouped into sets. Each costs +**Keeps the tool list affordable.** Sixteen built-in tools, grouped into sets. Each costs about 550 characters of schema on every request, so `{ "toolSets": [] }` trims back to the six core ones and a disabled set reaches neither the wire nor the prompt. @@ -144,14 +155,15 @@ workspace path. Up and down recall earlier prompts. ## Status -Working: the agent loop, tool approvals, subagents, skills, plugins, per-project memory, -session persistence, MCP, markdown rendering, headless mode, five-platform builds, streaming -reasoning display, the mid-turn prompt queue, gateable tool sets, read-only git tools, batch -reads, `@file` completion, interruptible commands, and the external registry. +Working: the agent loop, tool approvals, subagents including the gated `worker` kind, skills, +plugins, per-project memory, session persistence, MCP, markdown rendering, headless mode, +five-platform builds, streaming reasoning display, the mid-turn prompt queue, gateable tool +sets, read-only git tools, batch reads, `apply_patch`, `web_fetch`, `@file` completion, +interruptible commands, and the external registry. Next up is in [TODO.md](TODO.md); the longer view and what has been declined are in [ROADMAP.md](ROADMAP.md). The short version of what is missing: a summary of what compaction -discarded, `web_fetch`, a spend ceiling, and a cheaper model for subagent searches. +discarded, a spend ceiling, and a cheaper model for subagent searches. ## License diff --git a/docs/agents.md b/docs/agents.md index 002ea39..27d2420 100644 --- a/docs/agents.md +++ b/docs/agents.md @@ -33,10 +33,10 @@ pure latency. and findings recorded with `remember` so they survive compaction. **`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, `multi_edit`, -and `bash` are withheld from the model, not merely discouraged in prose — a model that cannot -see a tool cannot call it. They keep everything that only reads, including `read_many_files`, -`list_dir`, and the git tools. Their prompts also forbid describing edits as if they had been -made. +`apply_patch`, and `bash` are withheld from the model, not merely discouraged in prose — a +model that cannot see a tool cannot call it. They keep everything that only reads, including +`read_many_files`, `list_dir`, the git tools, and `web_fetch` when the `net` set is enabled. +Their prompts also forbid describing edits as if they had been made. ## Variants and tool sets @@ -146,3 +146,36 @@ the full set — cheaper per turn as well as safer. Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and nothing about the history. + +## Delegating with `task` + +The `task` tool spans a separate axis from the variants: it runs a subagent with its own +context window, so the parent pays for one report rather than the whole search transcript. The +subagent sees none of the parent's conversation, so its prompt must stand alone. + +| Kind | Tools | Approval | For | +|---|---|---|---| +| `explore` (default) | read and search only | never prompts — structurally read-only | a search spanning many files | +| `review` | read and search only | never prompts | a critique of code or a diff | +| `worker` | everything, including writes | every write and command asks, through the parent's gate | a self-contained change whose steps you do not need to watch | + +Three properties of the `worker` kind are structural rather than policy: + +**The gate is the parent's.** A worker routes each gated call back through the same permission +rules, the same guard plugins, and the same approval prompt as a direct call — flagged `a +worker subagent wants to run ...` so you can tell who is asking. Answering `always` grants the +pattern for the session exactly as it does for you. A subagent that could approve its own +writes would be a way to launder a tool call past you, so there is no separate, weaker gate. + +**Denial stops the work.** The worker is told a denial is your decision: report it, do not work +around it. The tool descriptions say the same thing, so the rule survives compaction. + +**No `worker` without a channel.** In headless runs there is no one to answer a prompt, so the +`worker` kind is not offered at all — an unattended write is not something to fall into by +accident. The read-only kinds work everywhere. No subagent holds `web_fetch`; network access +stays with the main agent, where the approval prompt says what it is for. + +When not to delegate: a single grep, or anything you must supervise step by step — keep that in +your own turn, where every call is on screen. A worker wins when the intermediate steps are +noise: a mechanical rename across twenty files, a test scaffold written to match an existing +suite, a cleanup whose shape you already know. diff --git a/docs/architecture.md b/docs/architecture.md index 73a9b66..2d10809 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -145,12 +145,16 @@ running, the handler exits as usual. ## Subagents -`task` runs a nested `streamText` with only `read_file`, `glob`, and `grep`. It returns one -message. +`task` runs a nested `streamText` and returns one message. The subagent kinds hold different +tool sets: `explore` and `review` the read-only tools, `worker` those plus every write tool. -Two consequences follow from the tool set, not from policy: +The consequences follow from the tool set, not from policy: -- It can never need approval, because it has no gated tools. +- `explore` and `review` can never need approval, because they hold no gated tool. +- `worker` needs approval for exactly the calls a direct one would, so the parent owns the + gate: the subagent's `toolApproval` callback routes back through the parent's permission + rules, guard plugins, and prompt. A subagent with its own approval would be a way to launder + a tool call past the user. - The parent's context holds the findings, not the search transcript. Progress is reported through a callback, wired to a bus the panel subscribes to. Without the @@ -196,9 +200,9 @@ tool call carries an itemId, so after the first compaction the model could not s already run, and re-ran the same tools until the step limit ended the turn. **Compaction may shorten the history; it must not blank it.** -**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts -*messages*, so the cut lands between an assistant `tool-call` and the `tool` message answering -it. What reaches the wire is a `function_call_output` with no `function_call`: +**A tool result without its tool call.** Tool pruning counts messages, so a cut can land between +an assistant `tool-call` and the `tool` message answering it. What reaches the wire is a +`function_call_output` with no `function_call`: ``` 400 No tool call found for function call output with call_id call_… @@ -208,6 +212,10 @@ it. What reaches the wire is a `function_call_output` with no `function_call`: reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a suspended approval looks like, and dropping it would break resume. +The pruning ladder drops reasoning first and then keeps the widest recent tool tail that fits. +The SDK carries that returned message view into later steps, and the session reports compaction +once per turn rather than once per step. + ## Registry `/registry` fetches an index of external skills and plugins over https. Skills are prompt text @@ -226,6 +234,7 @@ the reasoning. | `session.ts` | the loop, approvals, compaction, event stream | | `tools.ts` | file and shell tools, tool sets, ripgrep bridge, bash streaming and interrupt | | `tools-git.ts` | read-only git tools, spawned with a fixed argv | +| `tools-net.ts` | `web_fetch`, private-address and redirect checks | | `ignore.ts` | gitignore-aware walker, path jail | | `complete.ts` | `@path` token extraction, ranking, insertion | | `registry.ts` | external index, validation, install and removal | @@ -255,11 +264,11 @@ Every module is pure of the UI except `ui/`, and `ui/` never touches the SDK. Th ## Testing -538 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop; -`ink-testing-library` drives the UI with real keystrokes; MCP is tested against a real stdio -server subprocess; provider wire formats and the registry are tested against a local HTTP -server; the interrupt path spawns a real subprocess and asserts it died early rather than ran -out. +538 tests became 647 as the suites grew; no mocking framework. `MockLanguageModelV4` from +`ai/test` drives the loop; `ink-testing-library` drives the UI with real keystrokes; MCP is +tested against a real stdio server subprocess; provider wire formats and the registry are +tested against a local HTTP server; the interrupt path spawns a real subprocess and asserts it +died early rather than ran out. The pattern throughout is to assert on what actually crossed a boundary — what went on the wire, what is on screen, what is on disk — rather than on internal calls. diff --git a/docs/configuration.md b/docs/configuration.md index f9c8b86..496bd89 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -43,7 +43,7 @@ Written by `/provider`, editable by hand. Every field is optional. | `thinking` | default level: `off`, `low`, `medium`, `high`, `max` | | `maxRetries` | retries per model call for transient failures. Default 3 | | `plugins` | which builtin plugins to enable. Omit for `["guard", "time"]` | -| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`. Omit for all of them. See [tools](tools.md) | +| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`, and `net`. Omit for the defaults; `net` is opt-in. See [tools](tools.md) | | `permission` | which calls run, ask, or are refused, matched per command or path. See [permissions](permissions.md) | | `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) | | `mcpServers` | see [MCP](mcp.md) | diff --git a/docs/development.md b/docs/development.md index d7d511c..6ef9174 100644 --- a/docs/development.md +++ b/docs/development.md @@ -16,7 +16,7 @@ faster and the fallback path is exercised without it. ```bash bun run shiro # run from source bun run typecheck # tsc --noEmit -bun test # 538 tests +bun test # 647 tests bun run build # single binary for this platform -> dist/shiro bun run release # all five platforms -> dist/release + SHA256SUMS bun run install:local # build, then copy onto PATH @@ -95,9 +95,10 @@ Steps 3 and 4 are two hand-maintained lists of tool names, which is a known weak added to one and forgotten in the other is a silently ungated write. Deriving both from the tool definitions is on [TODO.md](../TODO.md). -Every tool costs roughly 550 characters of schema on every request. Fourteen built-in tools is -well past where selection accuracy starts to matter, which is why sets exist and why a new tool +Every tool costs roughly 550 characters of schema on every request. Sixteen built-in tools is +past where selection accuracy starts to matter, which is why sets exist and why a new tool needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why. +One set, `net`, is opt-in rather than on: `web_fetch` is the one tool that leaves the machine. ## Adding a slash command diff --git a/docs/headless.md b/docs/headless.md index bd74af6..32e5123 100644 --- a/docs/headless.md +++ b/docs/headless.md @@ -16,7 +16,7 @@ There is no terminal to approve on, so every gated tool is denied unless `--yolo ``` $ shiro -p "add a test for paginate()" -shiro: headless denies write_file, edit_file, multi_edit, bash and mcp tools unless --yolo is passed +shiro: headless denies write_file, edit_file, multi_edit, apply_patch, bash, web_fetch and mcp tools unless --yolo is passed [tool] write_file {"path":"test/paginate.test.ts",...} [denied] write_file (run with --yolo to allow tool use in headless mode) ``` @@ -51,7 +51,7 @@ $ shiro -p "count the tools" --json {"type":"tool-start","id":"c1","name":"grep"} {"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}} {"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."} -{"type":"text","text":"There are 14 built-in tools."} +{"type":"text","text":"There are 16 built-in tools."} {"type":"done","inputTokens":4210,"outputTokens":88} ``` diff --git a/docs/memory.md b/docs/memory.md index 1f67c7a..24565f7 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -4,7 +4,7 @@ Four kinds of state, each with a different lifetime. | State | Lives in | Survives | |---|---|---| -| transcript | the message array | until compaction or `/clear` | +| transcript | the message array | until `/compact` or `/clear` | | task list | the system prompt, rebuilt each step | pruning and `/compact` | | project memory | `~/.shiro-neko/memory/.json` | across sessions, forever | | session record | `~/.shiro-neko/sessions/.json` | until you delete it | @@ -145,9 +145,10 @@ unless you pass it again. Two mechanisms. -**Automatic**, at roughly 120k estimated tokens: `pruneMessages` strips reasoning and older -tool calls from what goes on the wire. Local history is untouched, so the transcript on your -screen stays complete. The turn reports it: +**Automatic**, at roughly 120k estimated tokens: reasoning is stripped first, then older tool +content is removed in a bounded ladder until the request fits. The SDK keeps that pruned view +for later steps in the turn; local session history remains complete. One `compacted` event is +reported per turn: ``` context compacted: 192 messages pruned to 15 on the wire @@ -156,10 +157,8 @@ context compacted: 192 messages pruned to 15 on the wire The status bar warns before that happens: context is shown as a percentage of the threshold, amber from two thirds, red at 90. -What gets discarded, in order: reasoning items first, then tool calls and their results older -than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not -conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what -lets a turn continue rather than restart. +What gets discarded, in order: reasoning items first, then the oldest tool calls and results as +needed. Recent exchanges are kept by the ladder, which lets a turn continue rather than restart. **Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands and outcomes, what remains — and it replaces the transcript entirely. @@ -197,9 +196,9 @@ first compaction the model could no longer see what it had already run. It re-ra tools until the step limit ended the turn. The history is the model's memory; compaction may shorten it but must not blank it. -**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts -*messages*, not pairs, so the cut can land between the assistant message holding a `tool-call` -and the `tool` message answering it: +**A tool result without its tool call.** Tool pruning counts messages, not call/result pairs, so +the cut can land between the assistant message holding a `tool-call` and the `tool` message +answering it: ``` 400 No tool call found for function call output with call_id call_… diff --git a/docs/permissions.md b/docs/permissions.md index b7c3cc2..36dd3be 100644 --- a/docs/permissions.md +++ b/docs/permissions.md @@ -34,6 +34,8 @@ remain are the ones worth reading. |---|---| | `bash` | the command, e.g. `git status --porcelain` | | `read_file` `write_file` `edit_file` `multi_edit` `list_dir` | the path | +| `apply_patch` | every file marker path in the patch | +| `web_fetch` | the URL | | `read_many_files` | every path in the batch; one match is enough | | `glob` `grep` | the pattern | | `git_diff` `git_log` `git_blame` | the path, when given | @@ -96,7 +98,7 @@ With no `permission` config: | `glob` `grep` `list_dir` | `allow` | | the git tools | `allow` — they cannot mutate anything | | `task`, and every session tool | `allow` — they touch the agent's own state | -| `write_file` `edit_file` `multi_edit` `bash` | `ask` | +| `write_file` `edit_file` `multi_edit` `apply_patch` `bash` `web_fetch` | `ask` | | anything else, including every `mcp__*` tool | `ask` | Credentials are denied on read rather than gated, because there is no recovery. A model that @@ -228,7 +230,7 @@ unmatched and the tool on its default: withholding the tools, which is stronger; use rules when you want the tools present but inert. ```json -{ "permission": { "write_file": "deny", "edit_file": "deny", "multi_edit": "deny", "bash": "deny" } } +{ "permission": { "write_file": "deny", "edit_file": "deny", "multi_edit": "deny", "apply_patch": "deny", "bash": "deny" } } ``` **An unattended job that may commit but never push.** diff --git a/docs/plugins.md b/docs/plugins.md index 5390aba..8239ff7 100644 --- a/docs/plugins.md +++ b/docs/plugins.md @@ -125,7 +125,7 @@ export const noSecretsPlugin: Plugin = { 'The no-secrets plugin refuses writes to .env and credential files. Ask the user to ' + 'add secrets themselves rather than working around it.', beforeToolCall: ({ toolName, input }) => { - if (toolName !== 'write_file' && toolName !== 'edit_file' && toolName !== 'multi_edit') return undefined; + if (!['write_file', 'edit_file', 'multi_edit', 'apply_patch'].includes(toolName)) return undefined; const path = String((input as { path?: unknown } | null)?.path ?? ''); if (/(^|\/)\.env|credentials|\.pem$/.test(path)) { return `refusing to write ${path}; add secrets yourself`; @@ -137,7 +137,7 @@ export const noSecretsPlugin: Plugin = { Then add it to `BUILTIN_PLUGINS` and, if it should be on by default, `DEFAULT_ENABLED`. -Note the three tool names. Every write tool has to be listed, and `multi_edit` is easy to miss +Note the four tool names. Every write tool has to be listed, and `multi_edit` is easy to miss — a guard that only checks `write_file` and `edit_file` is bypassed by a batch edit. Write the `appendix` whenever the plugin can block something. Without it the model hits a diff --git a/docs/skills.md b/docs/skills.md index e90eee6..67ea508 100644 --- a/docs/skills.md +++ b/docs/skills.md @@ -3,8 +3,8 @@ A skill is a markdown file with instructions for one kind of task. Only its name and description sit in the system prompt; the body is loaded on demand. -That split matters. The four bundled skills are 5,284 characters of body against 681 characters -of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt +That split matters. The six bundled skills are 8,900 characters of body against roughly 1,000 +characters of catalogue — paid on every request. Putting every body in the prompt would cost that on every turn, for instructions relevant to one turn in twenty. ## Format @@ -69,6 +69,14 @@ each, do not fix bugs while refactoring, do not add abstraction for a single cal implementation, never weaken an assertion to make a test pass, a flaky test is a shared-state problem and not something to retry around. +**`verify`** — confirm a change works by running the artifact the way a user would, not by +reading the source. What counts as evidence, what to do with the failure path, and reporting +what was not verified. + +**`commit`** — stage and commit work: look at the diff before staging, one commit one reason, +match the repository's message style, and the refusals — no amending pushed commits, no +`--no-verify`, no push unless asked. + They are string constants in `src/skills-builtin.ts` rather than files, because `bun build --compile` only embeds modules reachable through imports. A directory of `.md` files would be missing from the shipped binary. diff --git a/docs/tools.md b/docs/tools.md index 0d60e9f..1791e8d 100644 --- a/docs/tools.md +++ b/docs/tools.md @@ -15,7 +15,8 @@ auto-approved. reaches the context is on the wire and in the session file, and there is no taking it back. `*.env.example` is allowed. -**Asked by default.** `write_file`, `edit_file`, `multi_edit`, `bash`, and every `mcp__*` tool. +**Asked by default.** `write_file`, `edit_file`, `multi_edit`, `apply_patch`, `bash`, `web_fetch`, +and every `mcp__*` tool. ``` bash wants to run @@ -44,8 +45,9 @@ Three more things sit around the rules: ## Tool sets -Each tool costs its name, its description, and its JSON schema on **every request**. Measured -across the fourteen built-ins: +Each tool costs its name, its description, and its JSON schema on **every request**. The current +registry has sixteen built-ins. `/tools` shows the live set; disabling an optional set removes +its schemas from both the request and the system prompt. | Tool | Bytes | Tool | Bytes | |---|---|---|---| @@ -57,24 +59,24 @@ across the fourteen built-ins: | `read_file` | 526 | `git_status` | 292 | | `glob` | 499 | `write_file` | 289 | -7,673 bytes for all fourteen, averaging 548. Roughly 1,900 tokens per request before your -prompt or the conversation. Selection accuracy also falls as the list grows: a model choosing -between six tools picks better than one choosing between twenty. +Selection accuracy also falls as the list grows: a model choosing between six tools picks better +than one choosing between twenty. Sets let you switch off what a project does not need: | Set | Tools | Cost | |---|---|---| | `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B | -| `edit-plus` | `multi_edit` `list_dir` `read_many_files` | ~2,500 B | +| `edit-plus` | `multi_edit` `list_dir` `read_many_files` `apply_patch` | patch included | | `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B | +| `net` | `web_fetch` | opt in | ```json { "toolSets": ["edit-plus"] } ``` -Omit `toolSets` for all of them. `core` is always on — without read, edit, and bash the -agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since +Omit `toolSets` for the default sets. Add `net` when the agent should fetch public pages. +`core` is always on — without read, edit, and bash the agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since a prompt that names an absent tool teaches the model to attempt calls that cannot succeed. Session, plugin, and MCP tools are not part of this budget and are never gated here. @@ -190,6 +192,17 @@ edit 2: oldString not found in src/users.ts. No edits were applied. The last sentence matters. Without it a model reading the error has to guess whether edit 1 landed, and its next move — retry the whole batch, or only what failed — depends on the answer. +### `apply_patch` + +``` +patch one envelope containing Add, Update, Move, and Delete file markers +``` + +All operations are validated before anything is written, so a failure leaves every file +unchanged. Use it when one change spans files that must land together; use `multi_edit` for +several edits to one file and `edit_file` for one edit. Paths stay inside the workspace and the +call asks for approval. + ### `list_dir` ``` @@ -279,6 +292,18 @@ stdout: The turn continues from there. `esc` still aborts everything, and `ctrl-c` with nothing running quits as usual. +## `web_fetch` + +``` +url absolute HTTP(S) URL +maxChars returned characters, default 30,000, max 30,000 +``` + +Fetches a public text page and converts HTML to markdown. HTTPS is required for public hosts; +private and loopback addresses are refused, redirects are checked one hop at a time, and the +body is capped. The result is untrusted page content, not an instruction, and the call asks for +approval. It belongs to the opt-in `net` set. + ## Git tools All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument @@ -325,21 +350,25 @@ mutate anything and so never needs one. ``` description short label shown to you prompt self-contained instructions -kind "explore" (default) or "review" +kind "explore" (default), "review", or "worker" ``` -Spawns a read-only subagent with `read_file`, `glob`, and `grep` only. It returns one -report, so the parent pays for findings rather than the whole search transcript. It sees -none of the parent conversation, so its prompt has to stand alone. +Spawns a subagent with its own context window. It returns one report, so the parent pays for +the findings rather than the whole search transcript, and it sees none of the parent +conversation, so its prompt has to stand alone. -Two properties follow from that tool set rather than from policy: it can never trigger an -approval prompt, because it has no gated tools; and the parent's context holds the conclusion -instead of the search. A subagent reading forty files to answer one question costs the parent -the answer, not the forty files. +`explore` finds and reports; `review` critiques code in severity order. Both are structurally +read-only — they hold no gated tool at all, so they cannot trigger an approval prompt whatever +the config says. The parent's context holds the conclusion instead of the search: a subagent +reading forty files to answer one question costs the parent the answer, not the forty files. -`explore` finds and reports. `review` critiques code in severity order. Progress streams to -the subagent panel. Capped at 20 steps, and it shares the parent's model — an `explore` run -pays reasoning rates for what is really a search, which is [on the list](../TODO.md) to fix. +`worker` holds the write tools as well, and every write and command routes through the +parent's approval gate — the same rules and the same prompt as a direct call. The full +delegation trade-offs are in [agents](agents.md#delegating-with-task). + +Progress streams to the subagent panel, with each call's outcome. Capped at 20 steps, and it +shares the parent's model — an `explore` run pays reasoning rates for what is really a search, +which is [on the list](../TODO.md) to fix. Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a search spanning many files and loses on anything you could answer in one call.