Tools, six built-in to fourteen: - read_many_files: up to 20 paths read concurrently, each with its own window. An unreadable path is reported in its own block instead of throwing. - multi_edit: several edits to one file, validated in memory first so a late failure cannot leave the file half-written. - list_dir: ignore-aware depth-limited tree. - git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather than a shell string, which is what makes them safe to auto-approve. toolSets gates them. core is always on; edit-plus and git are optional. A disabled set reaches neither the wire nor the system prompt, since a prompt naming an absent tool teaches calls that cannot succeed. Interface: - Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn ends: it is progress, not the answer. - The tool in flight is named from tool-input-start, before its arguments finish streaming, and cleared on its result. - Prompts typed mid-turn queue and drain in order. esc clears the queue as well as aborting. - @ opens a path picker fed by the ignore-aware walker. Prefix matches rank above substring matches, so @src/ means "under src/". The walk runs on the first @, not at startup. ctrl-c kills the command in flight and keeps the turn. The call throws rather than returning, so the model cannot read a killed command as one that ran and failed on its own terms. The kill takes the whole process tree: killing cmd /c alone left the real command holding both pipes open, so the read never returned and the interrupt did nothing for 19 seconds. Two pruning fixes: - A tool result whose tool call was pruned is now dropped with it. Pruning counts messages, so the cut landed between an assistant tool-call and the tool message answering it, producing 400 "No tool call found for function call output with call_id ...". The reverse pairing is left alone: a call awaiting its result is what a suspended approval looks like. - ignore.ts called statFs without importing it, so walk() crashed on the first symlink. 482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the two hand-maintained tool-name lists recorded as a known weakness.
100 lines
3.7 KiB
Markdown
100 lines
3.7 KiB
Markdown
# Agents and thinking
|
|
|
|
An agent variant sets three things: how much the model deliberates, which tools it is
|
|
offered, and a behaviour appendix in the system prompt.
|
|
|
|
```bash
|
|
shiro --agent deep # at launch
|
|
shiro --agent plan --think low # variant with an overridden level
|
|
```
|
|
|
|
```
|
|
/agent picker
|
|
/agent review direct
|
|
/think picker
|
|
/think max direct
|
|
```
|
|
|
|
## The variants
|
|
|
|
| Variant | Thinking | Tools | Steps | For |
|
|
|---|---|---|---|---|
|
|
| `default` | medium | all | 50 | ordinary work |
|
|
| `quick` | off | all | 12 | small, well-scoped edits |
|
|
| `deep` | max | all | 80 | hard problems, unclear causes |
|
|
| `plan` | high | read-only | 50 | investigate and propose |
|
|
| `review` | high | read-only | 50 | critique a change |
|
|
|
|
**`quick`** tells the model not to deliberate, not to write a task list, and not to explore
|
|
beyond what the change needs. Good for a rename or a one-line fix where thinking budget is
|
|
pure latency.
|
|
|
|
**`deep`** asks for more than one hypothesis before acting, more reading before concluding,
|
|
and findings recorded with `remember` so they survive compaction.
|
|
|
|
**`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, `multi_edit`,
|
|
and `bash` are withheld from the model, not merely discouraged in prose — a model that cannot
|
|
see a tool cannot call it. They keep everything that only reads, including `read_many_files`,
|
|
`list_dir`, and the git tools. Their prompts also forbid describing edits as if they had been
|
|
made.
|
|
|
|
## Variants and tool sets
|
|
|
|
Two separate things narrow the tool list, and they compose.
|
|
|
|
A variant withholds tools by *capability*: `plan` cannot write, whatever the config says.
|
|
`toolSets` withholds them by *cost*: a project that never wants the git tools switches that set
|
|
off for every variant. See [tools](tools.md#tool-sets).
|
|
|
|
Both go through one function, so a withheld tool is missing from the wire and from the system
|
|
prompt together. `/tools` lists what is actually offered this turn, with the set each tool came
|
|
from.
|
|
|
|
## Thinking levels
|
|
|
|
`off`, `low`, `medium`, `high`, `max`. They map to whatever the provider actually supports:
|
|
|
|
| Level | OpenAI `reasoning_effort` | Anthropic `thinking` |
|
|
|---|---|---|
|
|
| `off` | `none` | `{ type: "disabled" }` |
|
|
| `low` | `low` | `budget_tokens: 6400` |
|
|
| `medium` | `medium` | proportional budget |
|
|
| `high` | `high` | `budget_tokens: 38400` |
|
|
| `max` | `xhigh` | maximum budget |
|
|
|
|
Verified against both wire formats rather than assumed.
|
|
|
|
Higher costs more and takes longer. `off` on a hard problem produces confident wrong
|
|
answers; `max` on a rename wastes a few cents and several seconds. The variants pick
|
|
sensible defaults, so reach for `/think` only when a specific turn needs something else.
|
|
|
|
## Overriding
|
|
|
|
`--agent deep --think low` gives you `deep`'s tools, steps, and appendix with a low thinking
|
|
budget. The override clones the preset rather than mutating it, so a later `/agent deep` in
|
|
the same session still gets `max`.
|
|
|
|
## Defaults in config
|
|
|
|
```json
|
|
{ "agent": "deep", "thinking": "high" }
|
|
```
|
|
|
|
A flag beats the config file. An unknown name fails at startup with the valid list rather
|
|
than silently falling back:
|
|
|
|
```
|
|
$ shiro --agent turbo
|
|
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review
|
|
```
|
|
|
|
## What the variant changes in the prompt
|
|
|
|
The system prompt describes only the tools actually offered, and the workflow rules adapt.
|
|
Under `plan` the model is told it has no tools that change anything and that it cannot run
|
|
commands, so it should say what to run rather than claim it passed. Under `default` it is
|
|
told which tools need approval and to verify with the project's tests.
|
|
|
|
A prompt that describes a withheld tool teaches the model to attempt calls that cannot
|
|
succeed, which is why the description is generated from the live tool set.
|