Files
shiro-neko/TODO.md
T
Muhammad Zakir Ramadhan 2fa6ee247b Add batch reads, @file completion, interruptible commands, tool sets
Tools, six built-in to fourteen:
- read_many_files: up to 20 paths read concurrently, each with its own window.
  An unreadable path is reported in its own block instead of throwing.
- multi_edit: several edits to one file, validated in memory first so a late
  failure cannot leave the file half-written.
- list_dir: ignore-aware depth-limited tree.
- git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather
  than a shell string, which is what makes them safe to auto-approve.

toolSets gates them. core is always on; edit-plus and git are optional. A
disabled set reaches neither the wire nor the system prompt, since a prompt
naming an absent tool teaches calls that cannot succeed.

Interface:
- Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn
  ends: it is progress, not the answer.
- The tool in flight is named from tool-input-start, before its arguments finish
  streaming, and cleared on its result.
- Prompts typed mid-turn queue and drain in order. esc clears the queue as well
  as aborting.
- @ opens a path picker fed by the ignore-aware walker. Prefix matches rank
  above substring matches, so @src/ means "under src/". The walk runs on the
  first @, not at startup.

ctrl-c kills the command in flight and keeps the turn. The call throws rather
than returning, so the model cannot read a killed command as one that ran and
failed on its own terms. The kill takes the whole process tree: killing cmd /c
alone left the real command holding both pipes open, so the read never returned
and the interrupt did nothing for 19 seconds.

Two pruning fixes:
- A tool result whose tool call was pruned is now dropped with it. Pruning
  counts messages, so the cut landed between an assistant tool-call and the tool
  message answering it, producing 400 "No tool call found for function call
  output with call_id ...". The reverse pairing is left alone: a call awaiting
  its result is what a suspended approval looks like.
- ignore.ts called statFs without importing it, so walk() crashed on the first
  symlink.

482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of
docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the
two hand-maintained tool-name lists recorded as a known weakness.
2026-09-03 01:37:48 +07:00

5.4 KiB

TODO

Next up. One item, one outcome, verifiable when done.

Longer-term direction lives in ROADMAP.md.


Now

Summarize the pruned span

Compaction tells the model the history was pruned but not what was in it, so it can confidently contradict a decision it made forty messages ago.

  • Summarize the discarded messages before dropping them
  • Inject the summary in place of the count
  • Budget it: a summary that grows with the session defeats the point
  • Test: a pruned decision is still recoverable from the summary

A cheaper model for subagents

The subagent shares the parent's model. An explore run is search, not reasoning, and it currently pays the parent's per-token rate.

  • subagentModel in config, defaulting to the parent
  • /cost separates parent from subagent spend
  • Test: the subagent's calls go to the configured model, the parent's do not

A spend ceiling

A headless run that loops costs real money with nothing to stop it.

  • maxSpendUsd in config, checked after every turn
  • Warn at 80%, refuse to start another turn at 100%
  • Headless exits non-zero with the ceiling named, rather than stopping silently
  • Test: a session past its ceiling refuses the next turn and says why

Next

Summarize the pruned span

Compaction tells the model the history was pruned but not what was in it, so it can confidently contradict a decision it made forty messages ago.

  • Summarize the discarded messages before dropping them
  • Inject the summary in place of the count
  • Budget it: a summary that grows with the session defeats the point
  • Test: a pruned decision is still recoverable from the summary

web_fetch

  • URL to markdown, size-capped
  • Belongs to a net set, off by default — it is the one tool that leaves the machine
  • Test: a redirect is followed, an oversized body is truncated with a note

Derive the tool-name lists

TOOL_SETS and MUTATING_TOOLS both list names by hand. A tool added to one and forgotten in the other is a silently ungated write, which is the worst kind of bug this codebase can have.

  • Mark each tool as mutating where it is defined, not in a list beside it
  • TOOL_SETS covers every registered tool, checked rather than assumed
  • Test: a tool in no set, or a mutating tool outside MUTATING_TOOLS, fails the suite

Subagent parallelism

Two independent searches run sequentially. The panel already renders several agents; the loop does not fan out.

  • task accepts several investigations and runs them together
  • Test: two delegated searches overlap in time rather than queueing

Maintenance

  • Pricing table needs a source note and a date; rates drift and ours are hand-entered
  • estimateTokens divides JSON length by four. Good enough for a compaction threshold, wrong enough to mislead in /cost. Either label it an estimate everywhere or use a real tokenizer
  • listPaths walks up to 5000 files once per session. Fine for a repo, wasteful in a monorepo, and it never notices a file created after the first @

Known rough edges

Not bugs exactly, but things that will bite someone.

  • /clear wipes the terminal scrollback. <Static> output is already committed, so clearing React state alone leaves it on screen. The escape sequence works but takes the user's earlier terminal history with it.
  • Memory has no conflict resolution. Two contradictory notes both persist and both get injected. /memory may merge them, or may keep both.
  • Windows cmd /c differs from bash -lc. A command the model writes for one shell may fail on the other. The prompt states the platform; it does not translate.
  • An unknown name in toolSets is dropped silently. The header line shows which sets actually loaded, but a typo reads as "that set is off" rather than as a mistake.
  • The reasoning panel is per-turn, not per-step. Reasoning from an early step stays on screen through later ones until the turn ends.
  • An interrupted command's effects are unknown, and the model is told so. Nothing can know how far a half-run migration got.
  • @ completion lists files, not directories. @src/ narrows correctly, but you cannot complete to src/ itself, because the walker only yields files.

Done

Kept for one release, then deleted.

  • Reasoning streamed to a collapsed panel, ctrl-r to expand, dropped when the turn ends
  • The tool in flight named on screen from tool-input-start until its result arrives
  • Prompts typed during a turn queue and drain in order; esc clears the queue
  • toolSets gating, so a disabled set reaches neither the wire nor the prompt
  • multi_edit, atomic across several edits to one file
  • list_dir, ignore-aware and depth-limited
  • Read-only git tools: git_status git_diff git_log git_show git_blame
  • Orphaned tool results dropped during pruning, fixing the 400 "No tool call found for function call output with call_id ..."
  • read_many_files, concurrent, one labelled block per file, a bad path reported in place
  • @file completion: picker fed by the ignore-aware walker, tab inserts a relative path
  • ctrl-c kills the running command and keeps the turn. The kill takes the whole process tree: killing cmd /c alone left the real command holding both pipes open, so the interrupt appeared to do nothing for 19 seconds