Enhance roadmap and documentation with new features and optimizations

- Add optimization pass details including subagent model override, system prompt caching, and memory recall improvements.
- Document per-path permission matching for apply_patch and read_many_files.
- Update configuration to include subagentModel for cost-effective task handling.
- Implement atomic config writes to prevent half-written JSON on crashes.
- Introduce empty-report subagent retry mechanism for improved reliability.
- Adjust memory entry limits and search functionality for better performance.
- Add tests for new features and ensure existing functionality remains intact.
This commit is contained in:
asepharyana
2026-09-04 22:37:38 +07:00
parent 1e2df2b751
commit e90bcac7c3
16 changed files with 406 additions and 74 deletions
+23
View File
@@ -179,3 +179,26 @@ When not to delegate: a single grep, or anything you must supervise step by step
your own turn, where every call is on screen. A worker wins when the intermediate steps are
noise: a mechanical rename across twenty files, a test scaffold written to match an existing
suite, a cleanup whose shape you already know.
### Subagent model
By default a subagent runs on the same model as the parent. For read-only explorations and
reviews that is usually overkill — an `explore` search spanning many files pays the parent's
reasoning rate for what is really a grep with better recall. Set `subagentModel` in config to a
cheaper, faster model and every `task` call uses it instead:
```json
{ "subagentModel": "claude-sonnet-4-5" }
```
Omit it (or set it to a model that fails to resolve) to fall back to the parent model. `worker`
subagents inherit the same override; a worker that needs the parent's reasoning can be written
to do the careful parts in the main turn and delegate only the mechanical shell.
### Empty reports
A subagent that returns nothing at all — no tool steps and no text — is almost always a
transient failure rather than a real "nothing found". The very first such blank response is
retried once with a nudge to report, so a swallowed provider error does not surface as an empty
result. A run that did real tool work but never wrote a final answer is not retried: repeating
it would just redo the work, so it is handed back as-is for the parent to decide.
+8
View File
@@ -61,6 +61,14 @@ That is not an optimisation. A `todo_write` on step one must be visible to step
`system:` on `streamText` is bound once for the whole run. Returning `instructions` from
`prepareStep` is the only place per-step state can enter.
Building it, though, is cheap to cache. `systemFor()` re-renders the same strings and re-serialises
the same prompt on every step when nothing changed, and on Anthropic that defeats prompt caching
(sending the identical prefix each request misses the cache hit). So `Session` keeps the built
prompt and reuses it whenever the underlying inputs are unchanged: same message count, same
notebook revision, same agent variant. A `todo_write`, a `/save`, or a model switch bumps one of
those and the next step rebuilds. The result is one system prompt string per actual state change,
and identical requests across steps for the unchanged ones.
The prompt also describes only the tools actually offered this turn. A prompt that mentions a
withheld tool teaches the model to attempt impossible calls. Two things narrow that set: a
read-only agent variant, and `toolSets` in config. Both go through `activeTools()`, so a
+2
View File
@@ -26,6 +26,7 @@ Written by `/provider`, editable by hand. Every field is optional.
"bash": { "*": "ask", "git *": "allow" }
},
"registryUrl": "https://example.com/my-registry/index.json",
"subagentModel": "claude-sonnet-4-5",
"mcpServers": {
"fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] }
}
@@ -46,6 +47,7 @@ Written by `/provider`, editable by hand. Every field is optional.
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`, and `net`. Omit for the defaults; `net` is opt-in. See [tools](tools.md) |
| `permission` | which calls run, ask, or are refused, matched per command or path. See [permissions](permissions.md) |
| `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) |
| `subagentModel` | model used for `task` subagents. Omit to reuse the parent model. A cheaper model here cuts subagent cost (and latency) sharply for read-only searches. See [agents](agents.md) |
| `mcpServers` | see [MCP](mcp.md) |
## Provider presets
+10 -3
View File
@@ -33,7 +33,7 @@ text one self-contained line
- **gotcha** — a trap. "The migration must run before the seed or the FK fails."
- **command** — an invocation that works. "Tests run with `bun test`, not `npm test`."
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
Duplicates are refused. Text is capped at 600 characters, the store at 300 entries.
The kinds are not decoration: they are what the model reads back at boot, and they set how much
to trust a note. A `command` is verifiable in one run. A `decision` explains why the obvious
@@ -44,8 +44,15 @@ is worthless next session — there is no conversation left to say which approac
### `recall`
Every term must appear. A match increments that entry's hit count, which protects it from
compaction later — an entry the agent actually uses is worth keeping verbatim.
Every term must appear, as a substring at first and then via character-trigram overlap as a
fallback, so inflection and word order do not hide a note: "migration" finds "migrate" and
"databases" finds "database" in either orientation. A match increments that entry's hit count,
which protects it from compaction later — an entry the agent actually uses is worth keeping
verbatim.
Stale notes eventually clear out on their own. On load, any entry older than 180 days that was
never recalled is dropped rather than carried forever; recalled entries are kept whatever their
age. An unparseable stored date never gets a note dropped over a parsing quirk.
AND rather than OR, on purpose: "migration seed order" should find the one note about that,
not every note mentioning any of the three words. Returns the 15 most recent matches.
+6
View File
@@ -44,6 +44,12 @@ remain are the ones worth reading.
| `skill` | the skill name |
| everything else | `*` only |
For `apply_patch` and `read_many_files` each path is its own subject, matched independently. A
deny like `src/generated/*` catches a patch that touches one generated file among four, and a
rule matching any one of a batch's paths decides the call — one bad path is enough. A path that
contains spaces stays a single subject rather than being split into two, so `*.ts` matches
`my file.ts` as one thing.
A tool with no subject — `git_status` takes no arguments — matches `*` and nothing narrower. That
is why a rule for it is a plain decision rather than a pattern table: