Files
shiro-neko/TODO.md
T
Muhammad Zakir Ramadhan 7fd578e13b Gate tool calls per command and path, not per tool name
Approval was a list of tool names: `bash` needed it, `read_file` did not. That
fails in a specific way. `bash` covers `git status` and `rm -rf` equally, so a
user working through a batch presses `a` — always allow — on the first prompt and
every later command runs unasked, including the one they would have refused. The
gate was strongest when it mattered least and gone by the time it mattered.

Rules now match the *subject* of a call: the command for `bash`, the path for a
file tool, the pattern for a search.

    "permission": {
      "bash": { "*": "ask", "git *": "allow", "rm *": "deny" },
      "edit_file": { "*": "deny", "src/generated/*": "allow" }
    }

Plain last-match-wins, with no special case for deny. An earlier version made
deny win wherever it sat, on the theory that a refusal should be impossible to
undo by accident, and it made default-deny-with-exceptions unexpressible — which
is the shape a careful user actually writes, and the same shape as `*.env` denied
while `*.env.example` is allowed. Refusals that must never be configurable stay
in the guard plugin, which runs ahead of this and which --yolo cannot reach.

Three behaviours fall out of it:

- `always` grants the pattern the tool suggests, not the tool. Approving
  `git status` runs `git log` unprompted and still asks about `npm publish`.
- `.env`, `.env.*`, and `.pem` are denied on read by default. Not gated, refused:
  a secret that reaches the context is on the wire and in the session file, and
  there is no taking it back. `.env.example` stays allowed.
- A call repeated identically three times in one turn asks even when allowed. A
  model repeating itself is not making progress, and `bash: allow` is a statement
  about which commands are safe rather than permission to loop.

The prompt now says which rule matched and what `always` would grant:

    bash wants to run
    git status --porcelain
    y allow once | a always allow bash git * | n deny

A typo in a decision string is dropped at parse time, leaving the tool on its
default. Treating an unparseable value as `allow` would mean one misspelling
silently removing the gate.

Written after surveying Claude Code, Codex, opencode, and phi. Three of the four
had already moved to per-pattern rules; the credential deny and the repeat guard
come from opencode directly. ROADMAP records what was deliberately not taken and
why — OS sandboxing needs three platform implementations and is worse than
nothing if half-built, and opencode's own docs say LSP integration is often not a
net positive.

581 tests, up from 572. The engine is a pure function tested on its own, and the
loop is tested through a real Session: an allowed pattern never prompts, a denied
one never executes, and a `.env` read leaves no secret in the transcript.
2026-09-03 11:19:53 +07:00

9.3 KiB

TODO

Next up. One item, one outcome, verifiable when done.

Longer-term direction lives in ROADMAP.md.


Now

Summarize the pruned span

Compaction now keeps the model's memory of a turn, but it still tells the model nothing about the messages it dropped, so a decision from forty messages ago can be contradicted with confidence.

  • Summarize the discarded messages before dropping them
  • Inject the summary in place of the count
  • Budget it: a summary that grows with the session defeats the point
  • Test: a pruned decision is still recoverable from the summary

A spend ceiling

A headless run that loops costs real money with nothing to stop it.

  • maxSpendUsd in config, checked after every turn
  • Warn at 80%, refuse to start another turn at 100%
  • Headless exits non-zero with the ceiling named, rather than stopping silently
  • Test: a session past its ceiling refuses the next turn and says why

A cheaper model for subagents

The subagent shares the parent's model. An explore run is search, not reasoning, and it currently pays the parent's per-token rate.

  • subagentModel in config, defaulting to the parent
  • /cost separates parent from subagent spend
  • Test: the subagent's calls go to the configured model, the parent's do not

Hot-reload an installed entry

/registry add writes the file and says to restart. The skill catalogue and the guard chain are both assembled at boot, so a mid-session install does nothing until then.

  • Rebuild the skill list and plugin host after an install or removal
  • Leave a turn in flight alone: its rules must not change underneath it
  • Test: a skill installed mid-session is callable in the next turn without a restart

Next

MCP without the schema tax

Every MCP tool's schema goes into the prompt today, so twenty tools from one server cost roughly 2,750 tokens per request whether the model uses them or not. toolSets does not gate them.

phi solves this with three meta-tools — mcp_list, mcp_inspect, mcp_call — and a prompt that names only the servers. A hundred servers then cost almost nothing until one is called.

  • mcp_list / mcp_inspect / mcp_call replacing per-tool registration
  • The prompt lists server names, not schemas
  • Calls go through the same permission rules and guard as a built-in
  • Keep per-tool registration as an option: a two-tool server is cheaper registered directly
  • Test: a configured server contributes no schema to the request until mcp_call

Custom commands from a file

Every other CLI in this class has these and they are cheap: a markdown file becomes a slash command, with $ARGUMENTS, $1, !`cmd` for shell output, and @path for a file.

  • .shiro/commands/*.md and ~/.shiro-neko/commands/*.md, name from the filename
  • Frontmatter for description and agent
  • $ARGUMENTS and positional $1
  • !`cmd` substituted before the prompt is sent, with the guard applied to it
  • Test: a command with a shell substitution reaches the model with the output inlined

web_fetch

  • URL to markdown, size-capped
  • Belongs to a net set, off by default — it is the one tool that leaves the machine
  • Test: a redirect is followed, an oversized body is truncated with a note

Derive the tool-name lists

TOOL_SETS and MUTATING_TOOLS both list names by hand. A tool added to one and forgotten in the other is a silently ungated write, which is the worst kind of bug this codebase can have.

  • Mark each tool as mutating where it is defined, not in a list beside it
  • TOOL_SETS covers every registered tool, checked rather than assumed
  • Test: a tool in no set, or a mutating tool outside MUTATING_TOOLS, fails the suite

Subagent parallelism

Two independent searches run sequentially. The panel already renders several agents; the loop does not fan out.

  • task accepts several investigations and runs them together
  • Test: two delegated searches overlap in time rather than queueing

Undo a turn

Every comparable CLI has this: opencode /undo and /redo, Claude Code /rewind with checkpoints. There is /resume here, which restores a session, and nothing that walks one back.

  • Snapshot files before each prompt, capped at the 100 most recent
  • /undo restores files, conversation, or both; /redo reverses it
  • Say plainly what is not covered: a bash command's effects cannot be snapshotted
  • Test: an edit is reverted, and the model's own record of it goes with it

Maintenance

  • Pricing table needs a source note and a date; rates drift and ours are hand-entered
  • estimateTokens divides JSON length by four. Good enough for a compaction threshold, wrong enough to mislead in /cost. Either label it an estimate everywhere or use a real tokenizer
  • listPaths walks up to 5000 files once per session. Fine for a repo, wasteful in a monorepo, and it never notices a file created after the first @
  • MUTATING_TOOLS is now only used by tests and docs; the permission defaults are what actually gate a write. Either delete it or make the defaults derive from it

Known rough edges

Not bugs exactly, but things that will bite someone.

  • /clear wipes the terminal scrollback. <Static> output is already committed, so clearing React state alone leaves it on screen. The escape sequence works but takes the user's earlier terminal history with it.
  • Memory has no conflict resolution. Two contradictory notes both persist and both get injected. /memory may merge them, or may keep both.
  • Windows cmd /c differs from bash -lc. A command the model writes for one shell may fail on the other. The prompt states the platform; it does not translate.
  • An unknown name in toolSets is dropped silently. The header line shows which sets actually loaded, but a typo reads as "that set is off" rather than as a mistake.
  • Permission rules gate the call, not what it does. bash with git * allowed will run a git alias that shells out to anything, and there is no sandbox around the shell. Codex solves this with OS-level isolation — Seatbelt, Landlock, a Windows equivalent — which is three platform-specific implementations and not something to half-ship.
  • The reasoning panel is per-turn, not per-step. Reasoning from an early step stays on screen through later ones until the turn ends.
  • An interrupted command's effects are unknown, and the model is told so. Nothing can know how far a half-run migration got.
  • @ completion lists files, not directories. @src/ narrows correctly, but you cannot complete to src/ itself, because the walker only yields files.
  • An installed skill is a stranger's words in your system prompt. The install shows the body first and /skills records the origin, but nothing re-checks it later: a registry that changes a URL's contents affects the next install, not one already on disk.
  • A registry index is trusted for its contents, not its authorship. There are no signatures. registryUrl is the whole trust decision.

Done

Kept for one release, then deleted.

  • Reasoning streamed to a collapsed panel, ctrl-r to expand, dropped when the turn ends
  • The tool in flight named on screen from tool-input-start until its result arrives
  • Prompts typed during a turn queue and drain in order; esc clears the queue
  • toolSets gating, so a disabled set reaches neither the wire nor the prompt
  • multi_edit, atomic across several edits to one file
  • list_dir, ignore-aware and depth-limited
  • Read-only git tools: git_status git_diff git_log git_show git_blame
  • Orphaned tool results dropped during pruning, fixing the 400 "No tool call found for function call output with call_id ..."
  • read_many_files, concurrent, one labelled block per file, a bad path reported in place
  • @file completion: picker fed by the ignore-aware walker, tab inserts a relative path
  • ctrl-c kills the running command and keeps the turn. The kill takes the whole process tree: killing cmd /c alone left the real command holding both pipes open, so the interrupt appeared to do nothing for 19 seconds
  • Compaction no longer stops the loop. Pruning used to drop any assistant part whose reasoning item it removed, which on a reasoning model is every tool call. The model lost its record of what it had run and re-ran it until the step limit. The repair strips the provider itemId instead of the part, so the same content is sent inline
  • /registry: browse, search, install, and remove external skills and plugins. Skills are shown in full before install; plugins are a validated manifest of deny rules, never code
  • Context shown as a percentage of the compaction threshold, amber at two thirds, red at 90
  • Permission rules per command and path, replacing the per-tool list. bash was one yes/no for git status and rm -rf, so pressing a once removed the gate for both. Rules match the call's subject, always grants a pattern rather than the tool, .env and .pem are refused on read, and an identical call repeated three times in a turn asks even when allowed