Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual per-tool measurements, and fills in the parts a reader hits after the happy path: what a specific error means, what a setting costs, what is not covered. Measured rather than estimated: - Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the full set, averaging 548. - Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the argument for loading bodies on demand. - Full system prompt 3,571 chars, core-only 2,045. New sections: - tools: which sets to keep and why, the jail function itself, an output-cap table, and the real error strings for edit_file and multi_edit. - configuration: env var per provider preset, cost-estimate limits, what each --no-* flag isolates, and three settings that do more than they look like. - agents: step caps per variant, which variant to reach for, and the fact that reasoning is charged as output and discarded first by compaction. - headless: exit code 0 means "the turn completed", not "the answer was yes" — with the jq pattern for gating on content. Timeouts, concurrent -c runs fighting over one session, CI recipes for --no-skills. - mcp: parallel connect, startup cost, a debugging ladder, and that toolSets does not gate MCP tools. - registry: publishing, local testing over http://localhost, and a troubleshooting section keyed on the actual validator messages. - memory: what compaction discards in what order, /compact versus automatic pruning, and that -c matches on cwd. - skills: the frontmatter reader's limits, and how to verify a skill loaded. Corrections found while cross-checking against the source: - The guard table was missing --force-with-lease and > /dev/sd… - The done event's token fields are optional, so the jq example filters on one rather than assuming it. Two honest limits now written down: the guard matches command strings, so a base64-decoded or script-wrapped command is not caught; and a registry index is trusted for its contents, not its authorship. Verified: all internal links and heading anchors resolve, every docs/ page is reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
+50
-1
@@ -62,12 +62,35 @@ from.
|
||||
| `high` | `high` | `budget_tokens: 38400` |
|
||||
| `max` | `xhigh` | maximum budget |
|
||||
|
||||
Verified against both wire formats rather than assumed.
|
||||
Verified against both wire formats rather than assumed. The vocabulary is deliberately ours:
|
||||
`off` through `max` means the same thing whichever provider is configured, and switching
|
||||
providers mid-session does not change what `/think high` asks for.
|
||||
|
||||
Higher costs more and takes longer. `off` on a hard problem produces confident wrong
|
||||
answers; `max` on a rename wastes a few cents and several seconds. The variants pick
|
||||
sensible defaults, so reach for `/think` only when a specific turn needs something else.
|
||||
|
||||
Reasoning is also charged as output tokens, so `max` shows up in `/cost` even on a turn where
|
||||
the model wrote two lines. And reasoning is the **first thing compaction discards** — see
|
||||
[memory](memory.md#compaction) — so a long turn at `max` pays for thinking that will not be on
|
||||
the wire by the end of it.
|
||||
|
||||
## Steps
|
||||
|
||||
`maxSteps` caps how many model calls one turn may make. A step is one request: a tool call and
|
||||
its result, or the final text.
|
||||
|
||||
| Variant | Steps |
|
||||
|---|---|
|
||||
| `quick` | 12 |
|
||||
| `default`, `plan`, `review` | 50 |
|
||||
| `deep` | 80 |
|
||||
|
||||
The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops
|
||||
mid-work with whatever it has, which is why `quick`'s 12 suits a rename and would strand a
|
||||
refactor. If turns regularly hit the cap on the same kind of task, the task wants `deep`
|
||||
rather than a higher number.
|
||||
|
||||
## Overriding
|
||||
|
||||
`--agent deep --think low` gives you `deep`'s tools, steps, and appendix with a low thinking
|
||||
@@ -97,3 +120,29 @@ told which tools need approval and to verify with the project's tests.
|
||||
|
||||
A prompt that describes a withheld tool teaches the model to attempt calls that cannot
|
||||
succeed, which is why the description is generated from the live tool set.
|
||||
|
||||
Three rules flip on what is available:
|
||||
|
||||
| Condition | `default` says | `plan` says |
|
||||
|---|---|---|
|
||||
| can edit | "these need approval; if denied, stop and ask" | "you have no tools that change anything" |
|
||||
| can run commands | "verify with the project's build or tests" | "say what should be run rather than claiming it passed" |
|
||||
| can ask | "ask rather than guess when two readings differ" | (same, unless headless) |
|
||||
|
||||
The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for
|
||||
the full set — cheaper per turn as well as safer.
|
||||
|
||||
## Which to reach for
|
||||
|
||||
- **`default`** for anything you have not thought about. It is the right answer most of the time.
|
||||
- **`quick`** for a rename, a typo, a one-line fix. Its value is not the model being cheaper but
|
||||
the absence of deliberation latency on work that needs none.
|
||||
- **`deep`** when the first attempt already failed, or the cause is unclear. Asking for more than
|
||||
one hypothesis is the actual difference; the thinking budget is secondary.
|
||||
- **`plan`** before a change you are not sure about. Read-only means the plan cannot quietly
|
||||
become a half-applied edit.
|
||||
- **`review`** on a diff or a module. In headless CI this is the one that needs no `--yolo`,
|
||||
because it holds no tool that can modify anything — see [headless](headless.md).
|
||||
|
||||
Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and
|
||||
nothing about the history.
|
||||
|
||||
Reference in New Issue
Block a user