Most of this replaces "roughly 550 characters per tool" with the actual per-tool measurements, and fills in the parts a reader hits after the happy path: what a specific error means, what a setting costs, what is not covered. Measured rather than estimated: - Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the full set, averaging 548. - Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the argument for loading bodies on demand. - Full system prompt 3,571 chars, core-only 2,045. New sections: - tools: which sets to keep and why, the jail function itself, an output-cap table, and the real error strings for edit_file and multi_edit. - configuration: env var per provider preset, cost-estimate limits, what each --no-* flag isolates, and three settings that do more than they look like. - agents: step caps per variant, which variant to reach for, and the fact that reasoning is charged as output and discarded first by compaction. - headless: exit code 0 means "the turn completed", not "the answer was yes" — with the jq pattern for gating on content. Timeouts, concurrent -c runs fighting over one session, CI recipes for --no-skills. - mcp: parallel connect, startup cost, a debugging ladder, and that toolSets does not gate MCP tools. - registry: publishing, local testing over http://localhost, and a troubleshooting section keyed on the actual validator messages. - memory: what compaction discards in what order, /compact versus automatic pruning, and that -c matches on cwd. - skills: the frontmatter reader's limits, and how to verify a skill loaded. Corrections found while cross-checking against the source: - The guard table was missing --force-with-lease and > /dev/sd… - The done event's token fields are optional, so the jq example filters on one rather than assuming it. Two honest limits now written down: the guard matches command strings, so a base64-decoded or script-wrapped command is not caught; and a registry index is trusted for its contents, not its authorship. Verified: all internal links and heading anchors resolve, every docs/ page is reachable from the README, 538 tests pass, typecheck clean.
6.2 KiB
Agents and thinking
An agent variant sets three things: how much the model deliberates, which tools it is offered, and a behaviour appendix in the system prompt.
shiro --agent deep # at launch
shiro --agent plan --think low # variant with an overridden level
/agent picker
/agent review direct
/think picker
/think max direct
The variants
| Variant | Thinking | Tools | Steps | For |
|---|---|---|---|---|
default |
medium | all | 50 | ordinary work |
quick |
off | all | 12 | small, well-scoped edits |
deep |
max | all | 80 | hard problems, unclear causes |
plan |
high | read-only | 50 | investigate and propose |
review |
high | read-only | 50 | critique a change |
quick tells the model not to deliberate, not to write a task list, and not to explore
beyond what the change needs. Good for a rename or a one-line fix where thinking budget is
pure latency.
deep asks for more than one hypothesis before acting, more reading before concluding,
and findings recorded with remember so they survive compaction.
plan and review are genuinely read-only. write_file, edit_file, multi_edit,
and bash are withheld from the model, not merely discouraged in prose — a model that cannot
see a tool cannot call it. They keep everything that only reads, including read_many_files,
list_dir, and the git tools. Their prompts also forbid describing edits as if they had been
made.
Variants and tool sets
Two separate things narrow the tool list, and they compose.
A variant withholds tools by capability: plan cannot write, whatever the config says.
toolSets withholds them by cost: a project that never wants the git tools switches that set
off for every variant. See tools.
Both go through one function, so a withheld tool is missing from the wire and from the system
prompt together. /tools lists what is actually offered this turn, with the set each tool came
from.
Thinking levels
off, low, medium, high, max. They map to whatever the provider actually supports:
| Level | OpenAI reasoning_effort |
Anthropic thinking |
|---|---|---|
off |
none |
{ type: "disabled" } |
low |
low |
budget_tokens: 6400 |
medium |
medium |
proportional budget |
high |
high |
budget_tokens: 38400 |
max |
xhigh |
maximum budget |
Verified against both wire formats rather than assumed. The vocabulary is deliberately ours:
off through max means the same thing whichever provider is configured, and switching
providers mid-session does not change what /think high asks for.
Higher costs more and takes longer. off on a hard problem produces confident wrong
answers; max on a rename wastes a few cents and several seconds. The variants pick
sensible defaults, so reach for /think only when a specific turn needs something else.
Reasoning is also charged as output tokens, so max shows up in /cost even on a turn where
the model wrote two lines. And reasoning is the first thing compaction discards — see
memory — so a long turn at max pays for thinking that will not be on
the wire by the end of it.
Steps
maxSteps caps how many model calls one turn may make. A step is one request: a tool call and
its result, or the final text.
| Variant | Steps |
|---|---|
quick |
12 |
default, plan, review |
50 |
deep |
80 |
The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops
mid-work with whatever it has, which is why quick's 12 suits a rename and would strand a
refactor. If turns regularly hit the cap on the same kind of task, the task wants deep
rather than a higher number.
Overriding
--agent deep --think low gives you deep's tools, steps, and appendix with a low thinking
budget. The override clones the preset rather than mutating it, so a later /agent deep in
the same session still gets max.
Defaults in config
{ "agent": "deep", "thinking": "high" }
A flag beats the config file. An unknown name fails at startup with the valid list rather than silently falling back:
$ shiro --agent turbo
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review
What the variant changes in the prompt
The system prompt describes only the tools actually offered, and the workflow rules adapt.
Under plan the model is told it has no tools that change anything and that it cannot run
commands, so it should say what to run rather than claim it passed. Under default it is
told which tools need approval and to verify with the project's tests.
A prompt that describes a withheld tool teaches the model to attempt calls that cannot succeed, which is why the description is generated from the live tool set.
Three rules flip on what is available:
| Condition | default says |
plan says |
|---|---|---|
| can edit | "these need approval; if denied, stop and ask" | "you have no tools that change anything" |
| can run commands | "verify with the project's build or tests" | "say what should be run rather than claiming it passed" |
| can ask | "ask rather than guess when two readings differ" | (same, unless headless) |
The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for the full set — cheaper per turn as well as safer.
Which to reach for
defaultfor anything you have not thought about. It is the right answer most of the time.quickfor a rename, a typo, a one-line fix. Its value is not the model being cheaper but the absence of deliberation latency on work that needs none.deepwhen the first attempt already failed, or the cause is unclear. Asking for more than one hypothesis is the actual difference; the thinking budget is secondary.planbefore a change you are not sure about. Read-only means the plan cannot quietly become a half-applied edit.reviewon a diff or a module. In headless CI this is the one that needs no--yolo, because it holds no tool that can modify anything — see headless.
Switching mid-session is fine and cheap: /agent changes the next turn's tools and prompt, and
nothing about the history.