features: /changes, /search, /fork, web_search, prompt memoization, per-turn spend cap, workspace refresh
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s

- /changes diffs the last turn's file snapshot (added/modified/deleted), reusing undo infra via SnapshotStack.peek()
- system prompt memoized behind version counters (notebook/memory/skills/plugins/tools/workspace); hit-rate in /cost, foundation for provider caching
- web_search: DuckDuckGo Lite, keyless, 5 results, SSRF-filtered, in the net set with ask permission
- /search <query>: full-text grep over saved sessions incl. tool-input JSON
- workspace file list re-walks at a turn boundary after writes
- maxSpendPerTurn: per-turn cap stops a runaway step with a notice
- /fork: branch the session at the last turn boundary, original untouched

821 tests pass, typecheck clean, build green
This commit is contained in:
asepharyana
2026-09-09 18:06:31 +07:00
parent 0bdf642672
commit e709737df1
24 changed files with 720 additions and 41 deletions
+36 -29
View File
@@ -9,6 +9,34 @@ Nothing here is a date. Items move to [TODO.md](TODO.md) when they are next up.
## Shipped
### Session-feature batch (post-1.0)
**`/changes`** — diff the last turn's file snapshot: added / modified / deleted, per
absolute path. The file side of `/undo` without undoing; bash effects are still out of
reach of either.
**System-prompt memoization** — version counters (notebook, memory, skills, plugins,
tools, workspace) gate a cached system prompt, so the string built on every step becomes
one build plus hits. The provider-side win it unlocks — splitting the stable prefix for
cache_control — is the remaining half of the old "Prompt caching" entry.
**`web_search`** — DuckDuckGo Lite, no API key, five results with title/URL/snippet,
re-checked through the same private-address filter as `web_fetch`, all inside the opt-in
`net` set.
**`/search <query>`** — full-text across saved sessions, matching transcript strings and
tool-input JSON. Deliberately no index: a session store fits in a grep.
**Workspace list refresh** — the boot-injected file list re-walks at a turn boundary when
that turn wrote files, so a path created mid-session shows up in the next prompt without a
restart.
**Per-turn spend cap** (`maxSpendPerTurn`) — a `deep` turn that runs away is stopped at a
step boundary by its own budget, complementing the session ceiling.
**`/fork`** — clone the session at the last turn boundary into a new saved session;
trying a different approach no longer costs the original.
### 0.1.0-beta.1
**Core loop** — `streamText` with tool approvals suspended and resumed through the SDK's
@@ -240,6 +268,14 @@ agent·model row inside it and a split footer beneath.
## Next
### Prompt caching
The system prompt is now memoized client-side, so the string is byte-identical across
steps when nothing volatile changed. The remaining half is provider-side: Anthropic
`cache_control` and OpenAI automatic prefix caching already reward that stable prefix, and
splitting the stable prefix from the volatile suffix (notebook/memory) would make the
cache unmissable even when a todo_write happens mid-turn.
### MCP without the schema tax
Every MCP tool's schema is in the prompt on every request, and `toolSets` does not gate them: a
@@ -248,25 +284,6 @@ answer is three meta-tools — `mcp_list`, `mcp_inspect`, `mcp_call` — with th
the servers, so a hundred servers cost almost nothing until one is called. Worth keeping direct
registration as an option: for a two-tool server the indirection is the more expensive of the two.
### Undo a turn
opencode has `/undo` and `/redo`, Claude Code has `/rewind` over file checkpoints. There is
`/resume` here, which restores a whole session, and nothing that steps one turn back. The honest
limit is the same for everyone: a `bash` command's effects cannot be snapshotted, so this covers
file-tool edits and says so.
### Lossless-enough compaction
Compaction keeps the model's memory of a turn now, but it still says nothing about the messages it
discarded, so the model can contradict its own earlier decision with confidence. A summary of the
discarded span costs one cheap call and removes the whole class of problem.
### Derived tool metadata
`TOOL_SETS` and `MUTATING_TOOLS` are hand-maintained lists of tool names. A tool added to one
and forgotten in the other is a silently ungated write. Marking each tool where it is defined,
and checking the coverage in the suite, removes the failure mode rather than documenting it.
### Registry trust
An index is trusted for its contents, not its authorship: `registryUrl` is the whole trust
@@ -283,12 +300,6 @@ when the real commands are already in `AGENTS.md`.
## Later
**Subagent parallelism.** Two independent searches run sequentially today. The panel already
handles multiple agents; the loop does not fan out.
**Session branching.** Fork a session at a message to try a different approach without
losing the original.
**Structured diff review.** Approve or reject individual hunks of an `edit_file` call rather
than the whole thing.
@@ -297,10 +308,6 @@ for now. Loading `.shiro/plugins/*.ts` needs a sandbox story first — a plugin
tool calls can also lie about blocking them, and one that can execute can read whatever the
agent can read.
**Prompt caching.** Anthropic and OpenAI both support it. The system prompt is rebuilt every
step for task-list freshness, which defeats a naive cache; splitting the stable prefix from
the volatile suffix would fix that.
**External hooks.** phi and both first-party CLIs let a script sit in the tool loop: a directory
with a manifest and an executable, one JSON object in on stdin, one out. phi's `pre_tool` can
rewrite the tool's input as well as allow or deny, which the compiled plugin interface here cannot