Add batch reads, @file completion, interruptible commands, tool sets

Tools, six built-in to fourteen:
- read_many_files: up to 20 paths read concurrently, each with its own window.
  An unreadable path is reported in its own block instead of throwing.
- multi_edit: several edits to one file, validated in memory first so a late
  failure cannot leave the file half-written.
- list_dir: ignore-aware depth-limited tree.
- git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather
  than a shell string, which is what makes them safe to auto-approve.

toolSets gates them. core is always on; edit-plus and git are optional. A
disabled set reaches neither the wire nor the system prompt, since a prompt
naming an absent tool teaches calls that cannot succeed.

Interface:
- Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn
  ends: it is progress, not the answer.
- The tool in flight is named from tool-input-start, before its arguments finish
  streaming, and cleared on its result.
- Prompts typed mid-turn queue and drain in order. esc clears the queue as well
  as aborting.
- @ opens a path picker fed by the ignore-aware walker. Prefix matches rank
  above substring matches, so @src/ means "under src/". The walk runs on the
  first @, not at startup.

ctrl-c kills the command in flight and keeps the turn. The call throws rather
than returning, so the model cannot read a killed command as one that ran and
failed on its own terms. The kill takes the whole process tree: killing cmd /c
alone left the real command holding both pipes open, so the read never returned
and the interrupt did nothing for 19 seconds.

Two pruning fixes:
- A tool result whose tool call was pruned is now dropped with it. Pruning
  counts messages, so the cut landed between an assistant tool-call and the tool
  message answering it, producing 400 "No tool call found for function call
  output with call_id ...". The reverse pairing is left alone: a call awaiting
  its result is what a suspended approval looks like.
- ignore.ts called statFs without importing it, so walk() crashed on the first
  symlink.

482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of
docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the
two hand-maintained tool-name lists recorded as a known weakness.
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 01:37:48 +07:00
parent a5ace7a23f
commit 2fa6ee247b
36 changed files with 2541 additions and 215 deletions
+56 -29
View File
@@ -49,48 +49,78 @@ reported rather than fatal.
**Distribution** — five-platform cross-compiled binaries, checksums, install scripts, CI on
three operating systems, tag-driven releases.
### 0.1.0-beta.2 (unreleased)
Fourteen built-in tools, up from six, with sets so the schema cost stays controllable.
**Visible process** — reasoning streams to a collapsed panel with an estimated token count,
`ctrl-r` expands it, and it leaves with the turn since it is progress rather than the answer.
The tool in flight is named from `tool-input-start`, before its arguments have finished
streaming, and cleared on its result.
**Message queue** — the input stays mounted while the model works. A prompt typed mid-turn
queues, the panel counts what is waiting, and the queue drains in order when the turn ends.
`esc` clears the queue as well as aborting. Queued slash commands replay as if typed.
**More tools** — `multi_edit` applies several edits to one file atomically, validating every
edit in memory first so a late failure cannot leave the file half-written. `list_dir` gives an
ignore-aware depth-limited tree. Five read-only git tools, spawned with a fixed argv rather
than a shell string, which is what makes them safe to auto-approve.
**`activeTools` gating** — `toolSets` in config: `core` always on, `edit-plus` and `git`
optional. A disabled set reaches neither the wire nor the system prompt. `/tools` names the
set each live tool came from.
**Pruning correctness** — a tool result whose tool call the pruner discarded is now dropped
with it. Message-counted pruning cut between an assistant tool-call and the tool message
answering it, and the OpenAI responses API rejects the result on its own with 400 "No tool
call found for function call output with call_id ...".
**Batch reads** — `read_many_files` takes up to twenty paths, each with its own window, and
runs them concurrently. An unreadable path is reported in its own block rather than throwing,
so one wrong guess costs a line instead of the call.
**`@file` completion** — `@` opens a picker fed by the ignore-aware walker, narrowing as you
type. Prefix matches rank above substring matches, so `@src/` means "under src/" rather than
"anything containing src/". Tab inserts a plain relative path. The walk happens on the first
`@` rather than at startup.
**Interruptible commands** — `ctrl-c` kills the command in flight and keeps the turn: the call
fails with a message saying the command did not finish and its effects are unknown, and the
model takes its next step from there. The kill takes the whole process tree, because killing
`cmd /c` alone leaves the real command holding both pipes open and the read never returns.
---
## Next
### Visible process
### Lossless-enough compaction
The agent's reasoning is discarded. `reasoning-delta` already arrives from the session; the
transcript drops it. A collapsed panel showing what the model is thinking, expandable with a
key, is the largest gap between this and a tool that feels responsive on a slow turn.
Compaction says the history was pruned but not what was in it, so the model can contradict
its own earlier decision with confidence. A summary of the discarded span costs one cheap
call and removes the whole class of problem.
Also missing: which file is being read or written as it happens. `tool-call` events carry
the path but the transcript only shows a one-line summary after the fact.
### Cost control
### Message queue
Two halves of the same problem: an `explore` subagent pays the parent's reasoning rate for
what is really a search, and nothing stops a headless run that loops. A cheaper subagent model
and a per-session ceiling are both small changes on top of the pricing that already exists.
Typing during a turn does nothing. It should queue and run when the turn ends. Interrupting
with `esc` then retyping loses the thought. Requires input to stay live while `busy`, which
means the prompt and the spinner have to coexist rather than swap.
### Derived tool metadata
### More tools
`TOOL_SETS` and `MUTATING_TOOLS` are hand-maintained lists of tool names. A tool added to one
and forgotten in the other is a silently ungated write. Marking each tool where it is defined,
and checking the coverage in the suite, removes the failure mode rather than documenting it.
Measured cost: 553 characters of schema per tool, sent every request. Thirteen live tools is
already at the point where selection accuracy starts to matter, so the next additions need
`activeTools` gating per set before the count grows.
### `web_fetch`
Ordered by value per line of code:
- `multi_edit` — several edits to one file atomically, killing read-edit-read-edit churn
- `list_dir` — a tree view, so the model stops globbing blindly to orient
- `read_many_files` — batch reads in one round trip
- git read-only set — `git_status`, `git_diff`, `git_log`, `git_show`, `git_blame`, all
approval-free because they cannot mutate
- `web_fetch` — URL to markdown
URL to markdown, in a `net` set that is off by default — it is the one tool that leaves the
machine.
Declined: wrappers around a single bash line with no added guarantee. `run_tests`,
`typecheck`, `lint`, `build` are five tools of pure schema tax when the real commands are
already in `AGENTS.md`.
### `@file` completion
Typing `@src/` should complete paths. The last real ergonomic gap in the input.
---
## Later
@@ -101,9 +131,6 @@ handles multiple agents; the loop does not fan out.
**Session branching.** Fork a session at a message to try a different approach without
losing the original.
**Cost budgets.** A per-session ceiling that warns, then stops. Pricing and accounting exist;
the limit does not.
**Structured diff review.** Approve or reject individual hunks of an `edit_file` call rather
than the whole thing.