Files
shiro-neko/docs/agents.md
T
Muhammad Zakir Ramadhan 9b978fdbe1 Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
2026-09-03 09:26:37 +07:00

6.2 KiB

Agents and thinking

An agent variant sets three things: how much the model deliberates, which tools it is offered, and a behaviour appendix in the system prompt.

shiro --agent deep              # at launch
shiro --agent plan --think low  # variant with an overridden level
/agent          picker
/agent review   direct
/think          picker
/think max      direct

The variants

Variant Thinking Tools Steps For
default medium all 50 ordinary work
quick off all 12 small, well-scoped edits
deep max all 80 hard problems, unclear causes
plan high read-only 50 investigate and propose
review high read-only 50 critique a change

quick tells the model not to deliberate, not to write a task list, and not to explore beyond what the change needs. Good for a rename or a one-line fix where thinking budget is pure latency.

deep asks for more than one hypothesis before acting, more reading before concluding, and findings recorded with remember so they survive compaction.

plan and review are genuinely read-only. write_file, edit_file, multi_edit, and bash are withheld from the model, not merely discouraged in prose — a model that cannot see a tool cannot call it. They keep everything that only reads, including read_many_files, list_dir, and the git tools. Their prompts also forbid describing edits as if they had been made.

Variants and tool sets

Two separate things narrow the tool list, and they compose.

A variant withholds tools by capability: plan cannot write, whatever the config says. toolSets withholds them by cost: a project that never wants the git tools switches that set off for every variant. See tools.

Both go through one function, so a withheld tool is missing from the wire and from the system prompt together. /tools lists what is actually offered this turn, with the set each tool came from.

Thinking levels

off, low, medium, high, max. They map to whatever the provider actually supports:

Level OpenAI reasoning_effort Anthropic thinking
off none { type: "disabled" }
low low budget_tokens: 6400
medium medium proportional budget
high high budget_tokens: 38400
max xhigh maximum budget

Verified against both wire formats rather than assumed. The vocabulary is deliberately ours: off through max means the same thing whichever provider is configured, and switching providers mid-session does not change what /think high asks for.

Higher costs more and takes longer. off on a hard problem produces confident wrong answers; max on a rename wastes a few cents and several seconds. The variants pick sensible defaults, so reach for /think only when a specific turn needs something else.

Reasoning is also charged as output tokens, so max shows up in /cost even on a turn where the model wrote two lines. And reasoning is the first thing compaction discards — see memory — so a long turn at max pays for thinking that will not be on the wire by the end of it.

Steps

maxSteps caps how many model calls one turn may make. A step is one request: a tool call and its result, or the final text.

Variant Steps
quick 12
default, plan, review 50
deep 80

The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops mid-work with whatever it has, which is why quick's 12 suits a rename and would strand a refactor. If turns regularly hit the cap on the same kind of task, the task wants deep rather than a higher number.

Overriding

--agent deep --think low gives you deep's tools, steps, and appendix with a low thinking budget. The override clones the preset rather than mutating it, so a later /agent deep in the same session still gets max.

Defaults in config

{ "agent": "deep", "thinking": "high" }

A flag beats the config file. An unknown name fails at startup with the valid list rather than silently falling back:

$ shiro --agent turbo
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review

What the variant changes in the prompt

The system prompt describes only the tools actually offered, and the workflow rules adapt. Under plan the model is told it has no tools that change anything and that it cannot run commands, so it should say what to run rather than claim it passed. Under default it is told which tools need approval and to verify with the project's tests.

A prompt that describes a withheld tool teaches the model to attempt calls that cannot succeed, which is why the description is generated from the live tool set.

Three rules flip on what is available:

Condition default says plan says
can edit "these need approval; if denied, stop and ask" "you have no tools that change anything"
can run commands "verify with the project's build or tests" "say what should be run rather than claiming it passed"
can ask "ask rather than guess when two readings differ" (same, unless headless)

The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for the full set — cheaper per turn as well as safer.

Which to reach for

  • default for anything you have not thought about. It is the right answer most of the time.
  • quick for a rename, a typo, a one-line fix. Its value is not the model being cheaper but the absence of deliberation latency on work that needs none.
  • deep when the first attempt already failed, or the cause is unclear. Asking for more than one hypothesis is the actual difference; the thinking budget is secondary.
  • plan before a change you are not sure about. Read-only means the plan cannot quietly become a half-applied edit.
  • review on a diff or a module. In headless CI this is the one that needs no --yolo, because it holds no tool that can modify anything — see headless.

Switching mid-session is fine and cheap: /agent changes the next turn's tools and prompt, and nothing about the history.