Loop and ergonomics batch across Now/Next and Maintenance:
- /undo and /redo via pre-prompt file snapshots (snapshot.ts)
- task takes a tasks[] array and runs investigations concurrently (subagent.ts)
- lazy MCP tools: mcp_list/mcp_inspect/mcp_call meta-tools, eager opt-in (mcp.ts, config.ts)
- skill tool reads its list live so a mid-session install is callable next turn (skills.ts)
- tool-name lists (tool-kinds.ts) derived from a mutating() marker; gates previously ungated writes
- prune/session recovery path summarized, and step-back doom-loop primitive (step-back.ts)
- @file completion re-walks on a slow cooldown; estimateTokens and pricing labeled as estimates
Docs: README, CHANGELOG, docs/{mcp,architecture,development} updated to match.
CI/CD: bun install-store cache and concurrency gates on both workflows; release.yml now
composes file-based release notes via scripts/make-release-notes.ts and verifies every binary.
- /changes diffs the last turn's file snapshot (added/modified/deleted), reusing undo infra via SnapshotStack.peek()
- system prompt memoized behind version counters (notebook/memory/skills/plugins/tools/workspace); hit-rate in /cost, foundation for provider caching
- web_search: DuckDuckGo Lite, keyless, 5 results, SSRF-filtered, in the net set with ask permission
- /search <query>: full-text grep over saved sessions incl. tool-input JSON
- workspace file list re-walks at a turn boundary after writes
- maxSpendPerTurn: per-turn cap stops a runaway step with a notice
- /fork: branch the session at the last turn boundary, original untouched
821 tests pass, typecheck clean, build green
Every guide catches up with the sixteen-tool registry: apply_patch and web_fetch sections, the delegation guide with the worker kind, bounded compaction replacing the three-message window description, net in the tool-set tables, six bundled skills, and the module map gains tools-net.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.
Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.
New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
--no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
with the jq pattern for gating on content. Timeouts, concurrent -c runs
fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.
Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
rather than assuming it.
Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.
Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
Tools, six built-in to fourteen:
- read_many_files: up to 20 paths read concurrently, each with its own window.
An unreadable path is reported in its own block instead of throwing.
- multi_edit: several edits to one file, validated in memory first so a late
failure cannot leave the file half-written.
- list_dir: ignore-aware depth-limited tree.
- git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather
than a shell string, which is what makes them safe to auto-approve.
toolSets gates them. core is always on; edit-plus and git are optional. A
disabled set reaches neither the wire nor the system prompt, since a prompt
naming an absent tool teaches calls that cannot succeed.
Interface:
- Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn
ends: it is progress, not the answer.
- The tool in flight is named from tool-input-start, before its arguments finish
streaming, and cleared on its result.
- Prompts typed mid-turn queue and drain in order. esc clears the queue as well
as aborting.
- @ opens a path picker fed by the ignore-aware walker. Prefix matches rank
above substring matches, so @src/ means "under src/". The walk runs on the
first @, not at startup.
ctrl-c kills the command in flight and keeps the turn. The call throws rather
than returning, so the model cannot read a killed command as one that ran and
failed on its own terms. The kill takes the whole process tree: killing cmd /c
alone left the real command holding both pipes open, so the read never returned
and the interrupt did nothing for 19 seconds.
Two pruning fixes:
- A tool result whose tool call was pruned is now dropped with it. Pruning
counts messages, so the cut landed between an assistant tool-call and the tool
message answering it, producing 400 "No tool call found for function call
output with call_id ...". The reverse pairing is left alone: a call awaiting
its result is what a suspended approval looks like.
- ignore.ts called statFs without importing it, so walk() crashed on the first
symlink.
482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of
docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the
two hand-maintained tool-name lists recorded as a known weakness.
Agentic coding CLI on Bun, Ink, and the AI SDK.
Core: streamText loop with SDK-level tool approval so a denied call provably never executes; endpoint fallback for OpenAI reasoning models; retry with backoff.
Tools: read/write/edit/glob/grep/bash, path-jailed, gitignore-aware, ripgrep with a JS fallback, binary rejection, live bash streaming.
Agents: five variants crossing thinking level with tool restriction; plan and review withhold mutating tools from the model.
Extensibility: frontmatter skills with on-demand bodies, plugin host with blocking hooks, MCP stdio and HTTP, read-only subagents.
State: durable per-project memory, session task lists, session persistence, compaction that repairs provider-item dependencies.
Distribution: five-platform cross-compiled binaries with checksums, install scripts, CI on three operating systems.
404 tests, typecheck clean.