Loop and ergonomics batch across Now/Next and Maintenance:
- /undo and /redo via pre-prompt file snapshots (snapshot.ts)
- task takes a tasks[] array and runs investigations concurrently (subagent.ts)
- lazy MCP tools: mcp_list/mcp_inspect/mcp_call meta-tools, eager opt-in (mcp.ts, config.ts)
- skill tool reads its list live so a mid-session install is callable next turn (skills.ts)
- tool-name lists (tool-kinds.ts) derived from a mutating() marker; gates previously ungated writes
- prune/session recovery path summarized, and step-back doom-loop primitive (step-back.ts)
- @file completion re-walks on a slow cooldown; estimateTokens and pricing labeled as estimates
Docs: README, CHANGELOG, docs/{mcp,architecture,development} updated to match.
CI/CD: bun install-store cache and concurrency gates on both workflows; release.yml now
composes file-based release notes via scripts/make-release-notes.ts and verifies every binary.
The unreleased section gains bounded compaction, apply_patch, web_fetch, the worker subagent, and the tool-detail transcript; web_fetch leaves the roadmap's future list.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Approval was a list of tool names: `bash` needed it, `read_file` did not. That
fails in a specific way. `bash` covers `git status` and `rm -rf` equally, so a
user working through a batch presses `a` — always allow — on the first prompt and
every later command runs unasked, including the one they would have refused. The
gate was strongest when it mattered least and gone by the time it mattered.
Rules now match the *subject* of a call: the command for `bash`, the path for a
file tool, the pattern for a search.
"permission": {
"bash": { "*": "ask", "git *": "allow", "rm *": "deny" },
"edit_file": { "*": "deny", "src/generated/*": "allow" }
}
Plain last-match-wins, with no special case for deny. An earlier version made
deny win wherever it sat, on the theory that a refusal should be impossible to
undo by accident, and it made default-deny-with-exceptions unexpressible — which
is the shape a careful user actually writes, and the same shape as `*.env` denied
while `*.env.example` is allowed. Refusals that must never be configurable stay
in the guard plugin, which runs ahead of this and which --yolo cannot reach.
Three behaviours fall out of it:
- `always` grants the pattern the tool suggests, not the tool. Approving
`git status` runs `git log` unprompted and still asks about `npm publish`.
- `.env`, `.env.*`, and `.pem` are denied on read by default. Not gated, refused:
a secret that reaches the context is on the wire and in the session file, and
there is no taking it back. `.env.example` stays allowed.
- A call repeated identically three times in one turn asks even when allowed. A
model repeating itself is not making progress, and `bash: allow` is a statement
about which commands are safe rather than permission to loop.
The prompt now says which rule matched and what `always` would grant:
bash wants to run
git status --porcelain
y allow once | a always allow bash git * | n deny
A typo in a decision string is dropped at parse time, leaving the tool on its
default. Treating an unparseable value as `allow` would mean one misspelling
silently removing the gate.
Written after surveying Claude Code, Codex, opencode, and phi. Three of the four
had already moved to per-pattern rules; the credential deny and the repeat guard
come from opencode directly. ROADMAP records what was deliberately not taken and
why — OS sandboxing needs three platform implementations and is worse than
nothing if half-built, and opencode's own docs say LSP integration is often not a
net positive.
581 tests, up from 572. The engine is a pure function tested on its own, and the
loop is tested through a real Session: an allowed pattern never prompts, a denied
one never executes, and a `.env` read leaves no secret in the transcript.
The compaction bug, which is the important one:
beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.
The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.
So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.
Registry, via `/registry [list|search|add|remove|installed]`:
Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.
Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.
Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.
538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
v0.1.0-beta.2 was tagged but never published: the release workflow failed on the
windows-x64 build, so no release object and no artifacts exist under that tag.
The fix is in df5f9f8. Burning the version is cheaper than moving a tag that is
already on the remote.
ROADMAP records what happened, and development.md now says plainly that a green
local `bun run release` is not proof — the Windows host takes a different branch
from the Ubuntu runner CI releases from.
Bump src/version.ts and package.json together; release.ts refuses to build if
they disagree, or if the git tag disagrees with either.
Also updates the version shown in the README and MCP sample headers, and the
example in development.md so it points at the next release rather than this one.
Tools, six built-in to fourteen:
- read_many_files: up to 20 paths read concurrently, each with its own window.
An unreadable path is reported in its own block instead of throwing.
- multi_edit: several edits to one file, validated in memory first so a late
failure cannot leave the file half-written.
- list_dir: ignore-aware depth-limited tree.
- git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather
than a shell string, which is what makes them safe to auto-approve.
toolSets gates them. core is always on; edit-plus and git are optional. A
disabled set reaches neither the wire nor the system prompt, since a prompt
naming an absent tool teaches calls that cannot succeed.
Interface:
- Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn
ends: it is progress, not the answer.
- The tool in flight is named from tool-input-start, before its arguments finish
streaming, and cleared on its result.
- Prompts typed mid-turn queue and drain in order. esc clears the queue as well
as aborting.
- @ opens a path picker fed by the ignore-aware walker. Prefix matches rank
above substring matches, so @src/ means "under src/". The walk runs on the
first @, not at startup.
ctrl-c kills the command in flight and keeps the turn. The call throws rather
than returning, so the model cannot read a killed command as one that ran and
failed on its own terms. The kill takes the whole process tree: killing cmd /c
alone left the real command holding both pipes open, so the read never returned
and the interrupt did nothing for 19 seconds.
Two pruning fixes:
- A tool result whose tool call was pruned is now dropped with it. Pruning
counts messages, so the cut landed between an assistant tool-call and the tool
message answering it, producing 400 "No tool call found for function call
output with call_id ...". The reverse pairing is left alone: a call awaiting
its result is what a suspended approval looks like.
- ignore.ts called statFs without importing it, so walk() crashed on the first
symlink.
482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of
docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the
two hand-maintained tool-name lists recorded as a known weakness.
Agentic coding CLI on Bun, Ink, and the AI SDK.
Core: streamText loop with SDK-level tool approval so a denied call provably never executes; endpoint fallback for OpenAI reasoning models; retry with backoff.
Tools: read/write/edit/glob/grep/bash, path-jailed, gitignore-aware, ripgrep with a JS fallback, binary rejection, live bash streaming.
Agents: five variants crossing thinking level with tool restriction; plan and review withhold mutating tools from the model.
Extensibility: frontmatter skills with on-demand bodies, plugin host with blocking hooks, MCP stdio and HTTP, read-only subagents.
State: durable per-project memory, session task lists, session persistence, compaction that repairs provider-item dependencies.
Distribution: five-platform cross-compiled binaries with checksums, install scripts, CI on three operating systems.
404 tests, typecheck clean.