48 Commits
Author SHA1 Message Date
asepharyana 97834e1411 Merge remote-tracking branch 'refs/remotes/upstream/main'
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
# Conflicts:
#	ROADMAP.md
#	TODO.md
#	src/cli.tsx
#	src/config.ts
#	src/mcp.ts
#	src/permission.ts
#	src/prompt.ts
#	src/session.ts
#	src/snapshot.ts
#	src/subagent.ts
#	src/tools-extra.ts
#	src/tools.ts
#	src/ui/App.tsx
#	test/mcp.test.ts
#	test/prune.test.ts
#	test/session.test.ts
#	test/tools.test.ts
2026-09-21 20:43:35 +07:00
Muhammad Zakir Ramadhan 8eebc0f4d6 release 1.0.1: fix resume showing a blank transcript
A resumed session loaded its wire messages but never rebuilt the on-screen
history, so -r/-c and /resume rendered a blank transcript. Convert the stored
messages back into transcript lines and seed the panel with them at mount and
after /resume.
2026-09-17 19:26:42 +07:00
Muhammad Zakir Ramadhan f8c3cc2d8e Ship the batch: undo, parallel subagents, lazy MCP, hot-reloaded skills; docs and CI/CD
Loop and ergonomics batch across Now/Next and Maintenance:
- /undo and /redo via pre-prompt file snapshots (snapshot.ts)
- task takes a tasks[] array and runs investigations concurrently (subagent.ts)
- lazy MCP tools: mcp_list/mcp_inspect/mcp_call meta-tools, eager opt-in (mcp.ts, config.ts)
- skill tool reads its list live so a mid-session install is callable next turn (skills.ts)
- tool-name lists (tool-kinds.ts) derived from a mutating() marker; gates previously ungated writes
- prune/session recovery path summarized, and step-back doom-loop primitive (step-back.ts)
- @file completion re-walks on a slow cooldown; estimateTokens and pricing labeled as estimates

Docs: README, CHANGELOG, docs/{mcp,architecture,development} updated to match.
CI/CD: bun install-store cache and concurrency gates on both workflows; release.yml now
composes file-based release notes via scripts/make-release-notes.ts and verifies every binary.
2026-09-17 17:49:19 +07:00
asepharyana de89457cf7 feat: add diagnostics tool for live command output monitoring
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
- Implemented diagnostics functionality to start, stop, and monitor background check commands.
- Created a DiagnosticsPanel for real-time output display in the UI.
- Added support for auto-detecting default diagnostics commands based on project configuration.
- Introduced run_checks tool to execute project verification commands and report results.
- Enhanced tools-extra with functions to parse AGENTS.md and package.json for check commands.
- Added diff review functionality to visualize changes made in the last turn.
- Implemented tests for diagnostics and run_checks functionalities to ensure reliability.
2026-09-11 22:09:19 +07:00
asepharyana b0da909454 feat: prioritize batched reads for turn efficiency
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
Move read_many_files to core tool set so it is always offered. Strengthen
system-prompt guidance: read_file now redirects to read_many_files for
multiple files, read_many_files is framed as the primary reading tool with
an explicit batch range (2-20). Add a 'How to work' rule on read
efficiency, and a read-batching instruction in the deep agent variant.
861 tests pass; build clean.
2026-09-11 19:06:04 +07:00
asepharyana 7d77408a41 feat: implement auto-scaffolding for project workflow files in bare repos
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
2026-09-11 18:52:50 +07:00
asepharyana 44bb4bebf1 feat: enhance nudge logic to suppress notifications during todo_write 2026-09-11 18:52:50 +07:00
asepharyana 76043aeb38 Implement background command support for bash tool
- Added `background` option to `bash` tool to allow long-lived commands to run without blocking the turn.
- Introduced `bash_status` tool to check the status and output of background commands.
- Added `bash_stop` tool to explicitly stop running background commands.
- Updated command parsing to handle `/bash` commands for listing, stopping individual, and stopping all background commands.
- Enhanced session management to reap stale background commands on agent shutdown.
- Updated UI to reflect background command status and allow user interruption via ctrl-c.
- Added tests for background command functionality and command parsing.
- Documented new features and usage in background-bash.md.
2026-09-11 18:52:50 +07:00
asepharyana 7a1e415fea feat: implement provider-side prompt caching and workflow scaffolding
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
- Added support for caching the stable prefix of the system prompt for Anthropic models.
- Introduced a new option `cacheSystemPrefix` in SessionOptions to enable caching.
- Implemented a nudge escalation system that reminds users to update the TODO.md file, capped at three nudges.
- Created a new `scaffoldWorkflowFiles` function to generate TODO.md, ROADMAP.md, and docs/ directory when they are missing.
- Updated the App component to scaffold workflow files during initialization.
- Added hooks functionality to allow external scripts to modify tool input and manage approvals.
- Implemented tests for the new hooks functionality, ensuring proper approval and execution flow.
2026-09-09 20:11:23 +07:00
asepharyana 4aeb0d2455 workflow: polish — /context groups trackers, /workflow in README, memoize git root
- contextPanel splits 'instructions' from 'trackers' (TODO.md/ROADMAP.md), so
  /context no longer presents project trackers as standing orders, and its
  empty-state mentions both instruction files and trackers
- README: docs table gets Project workflow, command list gets the new commands
- session: gitRoot memoized (one sync fs walk per session, not per prompt
  build); workflowPolicy computed once per systemFor
- test: context panel grouping covered

10 workflow tests, 830+ suite green
2026-09-09 19:12:29 +07:00
asepharyana a3739eb628 workflow: project-driven agent workflow (docs-first, TODO/ROADMAP tracking, verify-before-done)
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
- system prompt gains a Project workflow block when the repo tracks its own
  progress (TODO.md/ROADMAP.md/docs): read the task list first and keep it
  current, spec-first for non-trivial work, complete unit tests, verify before
  declaring done
- TODO.md + ROADMAP.md loaded as 'Project tracker' instruction blocks, capped
  tighter than AGENTS.md so the agent sees the shape without drowning in it
- one soft nudge per session when a turn writes files without touching the
  task list; suppressed when todo_write was called; never a gate
- /workflow command + panel showing TODO/ROADMAP/docs presence, line/file
  counts, nudge state
- config: workflow.enabled (default true), workflow.docsDir (default docs)
- prompt-cache version key extended with wf: so toggling the workflow busts
  the cached system prompt
- docs: docs/workflow.md, configuration table row, ROADMAP + TODO entries
- tests: test/workflow.test.ts (9 tests)

830 tests pass, typecheck clean, build green, live headless demo verified
2026-09-09 19:00:22 +07:00
asepharyana e709737df1 features: /changes, /search, /fork, web_search, prompt memoization, per-turn spend cap, workspace refresh
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
- /changes diffs the last turn's file snapshot (added/modified/deleted), reusing undo infra via SnapshotStack.peek()
- system prompt memoized behind version counters (notebook/memory/skills/plugins/tools/workspace); hit-rate in /cost, foundation for provider caching
- web_search: DuckDuckGo Lite, keyless, 5 results, SSRF-filtered, in the net set with ask permission
- /search <query>: full-text grep over saved sessions incl. tool-input JSON
- workspace file list re-walks at a turn boundary after writes
- maxSpendPerTurn: per-turn cap stops a runaway step with a notice
- /fork: branch the session at the last turn boundary, original untouched

821 tests pass, typecheck clean, build green
2026-09-09 18:06:31 +07:00
asepharyana 0bdf642672 prompt: inject full ignore-aware file list at boot (gitignore-respected, capped 5000)
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
walk() at boot -> Session.workspaceFiles -> systemPrompt Environment
section + subagent system prompts. read_many_files still on-demand,
but the agent now sees the tree without a tool call.
2026-09-09 14:19:29 +07:00
asepharyana 4b4ddd0585 @: directories complete, not just files
walk({includeDirs}) yields src/ and src/ui/ with trailing slash, only
on the @ path. matchPaths ranks dirs before files for the same prefix
so @src/ surfaces the directory itself. Insertion keeps the slash.
2026-09-09 13:55:02 +07:00
asepharyana 5028ea6973 config: warn on unknown toolSets typo instead of dropping silently
unknownToolSetNames() + startup notice (stderr headless, NoticeBus
interactive) flags gti vs git typos at boot, not as that set off.
2026-09-09 13:28:24 +07:00
asepharyana dcede3c10a Maintenance: pricing date, token est. label, listPaths refresh, MUTATING derive 2026-09-09 12:43:03 +07:00
asepharyana 4f3a7eef7c TODO Next: undo a turn — /undo /redo + per-turn snapshots
File tools record before-write state via snapshot hook; Session
captures beforeFiles at turn start and afterFiles at turn end (cap
100 via SnapshotStack), and undo/redo restore files + messages
together. Bash not snapshotted (docs + notice).
2026-09-09 11:49:33 +07:00
asepharyana 7edaeafca0 TODO Next: subagent parallelism (task tasks batch)
task now accepts { description, prompt, kind } or
{ tasks: TaskSpec[], kind? } (up to 8) and fans out with
Promise.all. Each subagent keeps its own context window and
the progress panel receives start/step/result/end per id, so
independent searches overlap in wall time instead of queueing.

- src/subagent.ts: runOne extracted, createTaskTool uses union
  schema (single | batch), batch validates worker channel once
  then Promise.all, results joined as headings.
- docs/agents.md: delegation section notes tasks batch.
- 800 pass (added subagent-parallel.test.ts).
2026-09-09 10:51:01 +07:00
asepharyana 2587e03beb TODO Next: MCP without the schema tax (phi meta-tools)
- mcpExpose=phi|direct|auto on RawConfig + MCPConfig; phi default keeps
  direct as a measured option for small servers (2-tool cheaper direct)
- src/mcp.ts: buildMcpMetaTools + phi branch in connectMcp: lazy
  list/inspect/call behind 3 fixed schemas instead of N per-tool schemas;
  direct branch kept; bindMcpGuard wires permission+PluginHost guard for
  MCP calls through the same gate as built-ins
- prompt mcpServers names-only under phi; under direct schemas travel
  as before — 20-tool server ~2750 tok -> ~phi (names) until mcp_call
- session: isDirect mcpServerNames non-enumerable marker, activeTools
  hides mcp_call from pre-wired rules when inside mcp_call, /tools
  shows mcp_call and Enabled/Withheld reflects the 3 meta-names under phi
- permission: mcp_* as read (free), mcp_call mutating keyed by
  server.tool, same top-level suppression + ask-to-approve as other
  mutating tools
- cli: bindMcpGuard + /mcp list shows exposure + mcp_* listed;
  headless denies line mentions mcp_list
- commit.ts: git_commit_message kept in git set via toolSetOf/
  disabledToolNames overlay

Tests: 798 pass, 0 fail; tsc exit 0; mcp.test.ts covers phi stays
empty until mcp_call and direct still namespaced.
2026-09-09 10:33:18 +07:00
asepharyana 626450eb06 TODO Next: derive TOOL_SETS and MUTATING_TOOLS from tool definitions
Wrap every builtin tool with withMeta({set, mutating}) at its definition
site (src/tool-utils.ts). TOOL_SETS and MUTATING_TOOLS are now derived
via setsFrom/mutatingNames rather than hand-lists, so a new write cannot
be added ungated by forgetting a parallel list. Permission defaults now
cover the 5 extra mutating line-edit tools. Test suite covers coverage,
derived-equality, and mutating consistency (test/tool-derive.test.ts).
2026-09-09 00:29:23 +07:00
asepharyana c125cf36fa TODO Now: summarize pruned span + hot-reload installed entries
Summarize pruned span: droppedSpan + summarizeDiscarded (6k excerpt,
3-6 lines, one call per compaction) with retained note already landed;
flip TODO Now checkboxes to [x].

Hot-reload: Session.updateSkills/updatePlugins with pendingSkills/
pendingHost deferred during a turn, drainPendingHotReload at turn
boundary, live skill tool closure, cli.tsx rebuilds host from
builtin+registry+external on /registry add/remove. Guard now uses
live pluginHost. Test mid-session skill callability and plugin
enforcement without restart.

Spec: .hermes/plans/todo-now-next.md
2026-09-08 23:54:37 +07:00
asepharyana 4848db308a harness: auto-skill general + memory spesifik + lossless compaction
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
Memory balik spesifik per-project (keep paths, THIS repo), general jadi
skill auto-create ke ~/.shiro-neko/skills/auto-*.md. Compaction lossless
dengan head preserved + droppedSpan normalized + retained note. Learner
throttled tiap 3 turn / delta 8, pakai subagentModel (learnerModel) dan
ceiling guard, hash dedup, JSON fallback, TTL tiap 6 turn.
2026-09-08 22:52:31 +07:00
asepharyana 7db69f19da Enhance memory management and compaction features
- Implement global memory layer for cross-project patterns.
- Improve memory entry scoring with recency and tokenization.
- Adjust MAX_TEXT limit from 400 to 800 for better context retention.
- Add TTL pruning for stale memory entries.
- Capture and summarize dropped content during compaction.
- Update tests to reflect changes in memory behavior and compaction logic.
2026-09-08 21:27:47 +07:00
Muhammad Zakir Ramadhan ffa9a02c26 release 1.0.0: cost control, 41 tools, 29 skills, custom commands, auto-load 2026-09-07 19:59:58 +07:00
Muhammad Zakir RamadhanandSisyphus 5feec4b5c7 Print resume instructions on the way out
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:31:56 +07:00
Muhammad Zakir RamadhanandSisyphus e06fbc3050 Add the /mcp wizard for local and remote servers
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:31:38 +07:00
Muhammad Zakir RamadhanandSisyphus 3ef2d8a75a Number diff lines, render task lists, and add word navigation
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:31:00 +07:00
Muhammad Zakir RamadhanandSisyphus 940a1e7401 Draw Onboard with the shared frame
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:30:38 +07:00
Muhammad Zakir RamadhanandSisyphus a4edbcc71e Retry a dead provider item inline instead of ending the turn
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:30:00 +07:00
Muhammad Zakir RamadhanandSisyphus 4f6ed43fed Send a compacted history inline, with no provider item references
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:29:40 +07:00
Muhammad Zakir RamadhanandSisyphus a8a3ffcb5c Describe the new tools in the system prompt
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:29:20 +07:00
Muhammad Zakir RamadhanandSisyphus 9c4316e562 Bundle the security, perf, and migrate skills
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:29:01 +07:00
Muhammad Zakir RamadhanandSisyphus 26d62bc926 Add the protect plugin, refusing writes to tool-owned files
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:28:39 +07:00
Muhammad Zakir RamadhanandSisyphus e8b8b06a4b Add git_commit_message, written from the staged diff
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:28:00 +07:00
Muhammad Zakir RamadhanandSisyphus 6b67465723 Add git_branch and export the git runner
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:27:43 +07:00
Muhammad Zakir RamadhanandSisyphus 16d27e1611 Add move_file and delete_file, and flag a collapsed rewrite
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-05 17:26:49 +07:00
Muhammad Zakir RamadhanandSisyphus 55ebb40524 Show tool detail and outcomes in the transcript
Each tool line carries the arguments that identify the call - the paths a batch read is about to pull in, the files a patch touches - and its result line carries a one-line outcome. Worker approvals are labelled as subagent asks, and a subagent result attaches to its step in the panel.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:43:13 +07:00
Muhammad Zakir RamadhanandSisyphus 6df41b56d2 Bundle verify and commit skills
verify: confirm a change works by running the artifact the way a user would, and report what was not verified. commit: stage deliberately, one commit one reason, match the repository style, and the refusals around amending, hooks, and pushing.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:42:46 +07:00
Muhammad Zakir RamadhanandSisyphus 8ee4c5bd2c Document the new tools in the prompt and read-only variants
The tool guidance covers apply_patch and web_fetch, the approval line derives from the tools actually offered rather than a hardcoded edit-tool check, and plan and review gain approved web research when the net set is on.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:42:31 +07:00
Muhammad Zakir RamadhanandSisyphus 41546154fe Add the gated worker subagent kind
worker holds the write tools and routes every write and command through the parent's approval gate: same permission rules, same guard, same prompt, flagged as coming from a subagent. Without an approval channel, as in headless runs, the kind is not offered at all rather than silently downgraded to read-only.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:42:20 +07:00
Muhammad Zakir RamadhanandSisyphus 83f4399e64 Make compaction bounded and report it once per turn
The fixed three-message tool window collapsed long transcripts to a handful of messages: a 405-message run kept two of 202 tool calls, and the model re-ran what it could no longer see. Pruning now drops reasoning first and keeps the widest recent tool tail that fits a ladder, the SDK carries that view into later steps, and the turn emits one compaction event instead of one per step.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:42:02 +07:00
Muhammad Zakir RamadhanandSisyphus 9e03512cb0 Add apply_patch for atomic multi-file edits
One envelope carrying Add, Update, Move, and Delete file markers, validated in full before anything is written: a failure on the fourth file leaves the first three untouched. Permission rules match every path a patch touches, and the write guard and secret checks cover it alongside the other write tools.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:41:44 +07:00
Muhammad Zakir RamadhanandSisyphus 87219350f9 Add web_fetch in an opt-in net tool set
URL to markdown, size-capped, with redirect and private-address checks so a public URL cannot be walked into the cloud metadata endpoint. It belongs to a net set that is off unless asked for: it is the one tool that leaves the machine.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:41:26 +07:00
Muhammad Zakir Ramadhan 7fd578e13b Gate tool calls per command and path, not per tool name
Approval was a list of tool names: `bash` needed it, `read_file` did not. That
fails in a specific way. `bash` covers `git status` and `rm -rf` equally, so a
user working through a batch presses `a` — always allow — on the first prompt and
every later command runs unasked, including the one they would have refused. The
gate was strongest when it mattered least and gone by the time it mattered.

Rules now match the *subject* of a call: the command for `bash`, the path for a
file tool, the pattern for a search.

    "permission": {
      "bash": { "*": "ask", "git *": "allow", "rm *": "deny" },
      "edit_file": { "*": "deny", "src/generated/*": "allow" }
    }

Plain last-match-wins, with no special case for deny. An earlier version made
deny win wherever it sat, on the theory that a refusal should be impossible to
undo by accident, and it made default-deny-with-exceptions unexpressible — which
is the shape a careful user actually writes, and the same shape as `*.env` denied
while `*.env.example` is allowed. Refusals that must never be configurable stay
in the guard plugin, which runs ahead of this and which --yolo cannot reach.

Three behaviours fall out of it:

- `always` grants the pattern the tool suggests, not the tool. Approving
  `git status` runs `git log` unprompted and still asks about `npm publish`.
- `.env`, `.env.*`, and `.pem` are denied on read by default. Not gated, refused:
  a secret that reaches the context is on the wire and in the session file, and
  there is no taking it back. `.env.example` stays allowed.
- A call repeated identically three times in one turn asks even when allowed. A
  model repeating itself is not making progress, and `bash: allow` is a statement
  about which commands are safe rather than permission to loop.

The prompt now says which rule matched and what `always` would grant:

    bash wants to run
    git status --porcelain
    y allow once | a always allow bash git * | n deny

A typo in a decision string is dropped at parse time, leaving the tool on its
default. Treating an unparseable value as `allow` would mean one misspelling
silently removing the gate.

Written after surveying Claude Code, Codex, opencode, and phi. Three of the four
had already moved to per-pattern rules; the credential deny and the repeat guard
come from opencode directly. ROADMAP records what was deliberately not taken and
why — OS sandboxing needs three platform implementations and is worse than
nothing if half-built, and opencode's own docs say LSP integration is often not a
net positive.

581 tests, up from 572. The engine is a pure function tested on its own, and the
loop is tested through a real Session: an allowed pattern never prompts, a denied
one never executes, and a `.env` read leaves no secret in the transcript.
2026-09-03 11:19:53 +07:00
Muhammad Zakir Ramadhan 84c60f2022 Fix the loop stalling after compaction, add an external registry
The compaction bug, which is the important one:

beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.

The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.

So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.

Registry, via `/registry [list|search|add|remove|installed]`:

Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.

Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.

Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
  red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.

538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
2026-09-03 03:07:56 +07:00
Muhammad Zakir Ramadhan df5f9f83fd Fix the release build: skip windows metadata when cross-compiling
`bun build --compile --target=bun-windows-x64` rejects `--windows-title` unless
the host is Windows:

    error: Using --windows-title is only available when compiling on Windows
    build failed for windows-x64 (exit 1)

CI releases all five targets from one Ubuntu runner, so the flag failed the
release on the last of the five builds — after the other four had already been
produced. Locally on Windows it passed, which is why it shipped.

The metadata is now stamped only on a Windows host. The published binary goes
without it rather than the release failing.

Extracted `buildArgs()` so the host-dependent branch is testable: the previous
version could only be checked by running a five-platform build on two operating
systems.
2026-09-03 01:57:35 +07:00
Muhammad Zakir Ramadhan 2fa6ee247b Add batch reads, @file completion, interruptible commands, tool sets
Tools, six built-in to fourteen:
- read_many_files: up to 20 paths read concurrently, each with its own window.
  An unreadable path is reported in its own block instead of throwing.
- multi_edit: several edits to one file, validated in memory first so a late
  failure cannot leave the file half-written.
- list_dir: ignore-aware depth-limited tree.
- git_status/diff/log/show/blame: read-only, spawned with a fixed argv rather
  than a shell string, which is what makes them safe to auto-approve.

toolSets gates them. core is always on; edit-plus and git are optional. A
disabled set reaches neither the wire nor the system prompt, since a prompt
naming an absent tool teaches calls that cannot succeed.

Interface:
- Reasoning streams to a collapsed panel, ctrl-r expands, dropped when the turn
  ends: it is progress, not the answer.
- The tool in flight is named from tool-input-start, before its arguments finish
  streaming, and cleared on its result.
- Prompts typed mid-turn queue and drain in order. esc clears the queue as well
  as aborting.
- @ opens a path picker fed by the ignore-aware walker. Prefix matches rank
  above substring matches, so @src/ means "under src/". The walk runs on the
  first @, not at startup.

ctrl-c kills the command in flight and keeps the turn. The call throws rather
than returning, so the model cannot read a killed command as one that ran and
failed on its own terms. The kill takes the whole process tree: killing cmd /c
alone left the real command holding both pipes open, so the read never returned
and the interrupt did nothing for 19 seconds.

Two pruning fixes:
- A tool result whose tool call was pruned is now dropped with it. Pruning
  counts messages, so the cut landed between an assistant tool-call and the tool
  message answering it, producing 400 "No tool call found for function call
  output with call_id ...". The reverse pairing is left alone: a call awaiting
  its result is what a suspended approval looks like.
- ignore.ts called statFs without importing it, so walk() crashed on the first
  symlink.

482 tests, up from 404. Docs synced across README, ROADMAP, TODO, and all of
docs/: tool sets, the new tools, ctrl-c semantics, the tool-start event, and the
two hand-maintained tool-name lists recorded as a known weakness.
2026-09-03 01:37:48 +07:00
Muhammad Zakir Ramadhan 5b8503fcd9 Initial commit: shiro-neko 0.1.0-beta.1
Agentic coding CLI on Bun, Ink, and the AI SDK.

Core: streamText loop with SDK-level tool approval so a denied call provably never executes; endpoint fallback for OpenAI reasoning models; retry with backoff.

Tools: read/write/edit/glob/grep/bash, path-jailed, gitignore-aware, ripgrep with a JS fallback, binary rejection, live bash streaming.

Agents: five variants crossing thinking level with tool restriction; plan and review withhold mutating tools from the model.

Extensibility: frontmatter skills with on-demand bodies, plugin host with blocking hooks, MCP stdio and HTTP, read-only subagents.

State: durable per-project memory, session task lists, session persistence, compaction that repairs provider-item dependencies.

Distribution: five-platform cross-compiled binaries with checksums, install scripts, CI on three operating systems.

404 tests, typecheck clean.
2026-09-02 17:30:18 +07:00