A resumed session loaded its wire messages but never rebuilt the on-screen
history, so -r/-c and /resume rendered a blank transcript. Convert the stored
messages back into transcript lines and seed the panel with them at mount and
after /resume.
Loop and ergonomics batch across Now/Next and Maintenance:
- /undo and /redo via pre-prompt file snapshots (snapshot.ts)
- task takes a tasks[] array and runs investigations concurrently (subagent.ts)
- lazy MCP tools: mcp_list/mcp_inspect/mcp_call meta-tools, eager opt-in (mcp.ts, config.ts)
- skill tool reads its list live so a mid-session install is callable next turn (skills.ts)
- tool-name lists (tool-kinds.ts) derived from a mutating() marker; gates previously ungated writes
- prune/session recovery path summarized, and step-back doom-loop primitive (step-back.ts)
- @file completion re-walks on a slow cooldown; estimateTokens and pricing labeled as estimates
Docs: README, CHANGELOG, docs/{mcp,architecture,development} updated to match.
CI/CD: bun install-store cache and concurrency gates on both workflows; release.yml now
composes file-based release notes via scripts/make-release-notes.ts and verifies every binary.
- Implemented diagnostics functionality to start, stop, and monitor background check commands.
- Created a DiagnosticsPanel for real-time output display in the UI.
- Added support for auto-detecting default diagnostics commands based on project configuration.
- Introduced run_checks tool to execute project verification commands and report results.
- Enhanced tools-extra with functions to parse AGENTS.md and package.json for check commands.
- Added diff review functionality to visualize changes made in the last turn.
- Implemented tests for diagnostics and run_checks functionalities to ensure reliability.
Move read_many_files to core tool set so it is always offered. Strengthen
system-prompt guidance: read_file now redirects to read_many_files for
multiple files, read_many_files is framed as the primary reading tool with
an explicit batch range (2-20). Add a 'How to work' rule on read
efficiency, and a read-batching instruction in the deep agent variant.
861 tests pass; build clean.
- Added `background` option to `bash` tool to allow long-lived commands to run without blocking the turn.
- Introduced `bash_status` tool to check the status and output of background commands.
- Added `bash_stop` tool to explicitly stop running background commands.
- Updated command parsing to handle `/bash` commands for listing, stopping individual, and stopping all background commands.
- Enhanced session management to reap stale background commands on agent shutdown.
- Updated UI to reflect background command status and allow user interruption via ctrl-c.
- Added tests for background command functionality and command parsing.
- Documented new features and usage in background-bash.md.
- Added support for caching the stable prefix of the system prompt for Anthropic models.
- Introduced a new option `cacheSystemPrefix` in SessionOptions to enable caching.
- Implemented a nudge escalation system that reminds users to update the TODO.md file, capped at three nudges.
- Created a new `scaffoldWorkflowFiles` function to generate TODO.md, ROADMAP.md, and docs/ directory when they are missing.
- Updated the App component to scaffold workflow files during initialization.
- Added hooks functionality to allow external scripts to modify tool input and manage approvals.
- Implemented tests for the new hooks functionality, ensuring proper approval and execution flow.
- contextPanel splits 'instructions' from 'trackers' (TODO.md/ROADMAP.md), so
/context no longer presents project trackers as standing orders, and its
empty-state mentions both instruction files and trackers
- README: docs table gets Project workflow, command list gets the new commands
- session: gitRoot memoized (one sync fs walk per session, not per prompt
build); workflowPolicy computed once per systemFor
- test: context panel grouping covered
10 workflow tests, 830+ suite green
- system prompt gains a Project workflow block when the repo tracks its own
progress (TODO.md/ROADMAP.md/docs): read the task list first and keep it
current, spec-first for non-trivial work, complete unit tests, verify before
declaring done
- TODO.md + ROADMAP.md loaded as 'Project tracker' instruction blocks, capped
tighter than AGENTS.md so the agent sees the shape without drowning in it
- one soft nudge per session when a turn writes files without touching the
task list; suppressed when todo_write was called; never a gate
- /workflow command + panel showing TODO/ROADMAP/docs presence, line/file
counts, nudge state
- config: workflow.enabled (default true), workflow.docsDir (default docs)
- prompt-cache version key extended with wf: so toggling the workflow busts
the cached system prompt
- docs: docs/workflow.md, configuration table row, ROADMAP + TODO entries
- tests: test/workflow.test.ts (9 tests)
830 tests pass, typecheck clean, build green, live headless demo verified
- /changes diffs the last turn's file snapshot (added/modified/deleted), reusing undo infra via SnapshotStack.peek()
- system prompt memoized behind version counters (notebook/memory/skills/plugins/tools/workspace); hit-rate in /cost, foundation for provider caching
- web_search: DuckDuckGo Lite, keyless, 5 results, SSRF-filtered, in the net set with ask permission
- /search <query>: full-text grep over saved sessions incl. tool-input JSON
- workspace file list re-walks at a turn boundary after writes
- maxSpendPerTurn: per-turn cap stops a runaway step with a notice
- /fork: branch the session at the last turn boundary, original untouched
821 tests pass, typecheck clean, build green
walk() at boot -> Session.workspaceFiles -> systemPrompt Environment
section + subagent system prompts. read_many_files still on-demand,
but the agent now sees the tree without a tool call.
walk({includeDirs}) yields src/ and src/ui/ with trailing slash, only
on the @ path. matchPaths ranks dirs before files for the same prefix
so @src/ surfaces the directory itself. Insertion keeps the slash.
File tools record before-write state via snapshot hook; Session
captures beforeFiles at turn start and afterFiles at turn end (cap
100 via SnapshotStack), and undo/redo restore files + messages
together. Bash not snapshotted (docs + notice).
task now accepts { description, prompt, kind } or
{ tasks: TaskSpec[], kind? } (up to 8) and fans out with
Promise.all. Each subagent keeps its own context window and
the progress panel receives start/step/result/end per id, so
independent searches overlap in wall time instead of queueing.
- src/subagent.ts: runOne extracted, createTaskTool uses union
schema (single | batch), batch validates worker channel once
then Promise.all, results joined as headings.
- docs/agents.md: delegation section notes tasks batch.
- 800 pass (added subagent-parallel.test.ts).
- mcpExpose=phi|direct|auto on RawConfig + MCPConfig; phi default keeps
direct as a measured option for small servers (2-tool cheaper direct)
- src/mcp.ts: buildMcpMetaTools + phi branch in connectMcp: lazy
list/inspect/call behind 3 fixed schemas instead of N per-tool schemas;
direct branch kept; bindMcpGuard wires permission+PluginHost guard for
MCP calls through the same gate as built-ins
- prompt mcpServers names-only under phi; under direct schemas travel
as before — 20-tool server ~2750 tok -> ~phi (names) until mcp_call
- session: isDirect mcpServerNames non-enumerable marker, activeTools
hides mcp_call from pre-wired rules when inside mcp_call, /tools
shows mcp_call and Enabled/Withheld reflects the 3 meta-names under phi
- permission: mcp_* as read (free), mcp_call mutating keyed by
server.tool, same top-level suppression + ask-to-approve as other
mutating tools
- cli: bindMcpGuard + /mcp list shows exposure + mcp_* listed;
headless denies line mentions mcp_list
- commit.ts: git_commit_message kept in git set via toolSetOf/
disabledToolNames overlay
Tests: 798 pass, 0 fail; tsc exit 0; mcp.test.ts covers phi stays
empty until mcp_call and direct still namespaced.
Wrap every builtin tool with withMeta({set, mutating}) at its definition
site (src/tool-utils.ts). TOOL_SETS and MUTATING_TOOLS are now derived
via setsFrom/mutatingNames rather than hand-lists, so a new write cannot
be added ungated by forgetting a parallel list. Permission defaults now
cover the 5 extra mutating line-edit tools. Test suite covers coverage,
derived-equality, and mutating consistency (test/tool-derive.test.ts).
Summarize pruned span: droppedSpan + summarizeDiscarded (6k excerpt,
3-6 lines, one call per compaction) with retained note already landed;
flip TODO Now checkboxes to [x].
Hot-reload: Session.updateSkills/updatePlugins with pendingSkills/
pendingHost deferred during a turn, drainPendingHotReload at turn
boundary, live skill tool closure, cli.tsx rebuilds host from
builtin+registry+external on /registry add/remove. Guard now uses
live pluginHost. Test mid-session skill callability and plugin
enforcement without restart.
Spec: .hermes/plans/todo-now-next.md
- Implement global memory layer for cross-project patterns.
- Improve memory entry scoring with recency and tokenization.
- Adjust MAX_TEXT limit from 400 to 800 for better context retention.
- Add TTL pruning for stale memory entries.
- Capture and summarize dropped content during compaction.
- Update tests to reflect changes in memory behavior and compaction logic.
Keeps local claude-code preset (readClaudeCodeSettings from ~/.claude/settings.json) merged with upstream Pickers refactor.
Resolved Onboard.tsx: both readClaudeCodeSettings + Frame/Row imports.
Each tool line carries the arguments that identify the call - the paths a batch read is about to pull in, the files a patch touches - and its result line carries a one-line outcome. Worker approvals are labelled as subagent asks, and a subagent result attaches to its step in the panel.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
verify: confirm a change works by running the artifact the way a user would, and report what was not verified. commit: stage deliberately, one commit one reason, match the repository style, and the refusals around amending, hooks, and pushing.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The tool guidance covers apply_patch and web_fetch, the approval line derives from the tools actually offered rather than a hardcoded edit-tool check, and plan and review gain approved web research when the net set is on.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
worker holds the write tools and routes every write and command through the parent's approval gate: same permission rules, same guard, same prompt, flagged as coming from a subagent. Without an approval channel, as in headless runs, the kind is not offered at all rather than silently downgraded to read-only.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The fixed three-message tool window collapsed long transcripts to a handful of messages: a 405-message run kept two of 202 tool calls, and the model re-ran what it could no longer see. Pruning now drops reasoning first and keeps the widest recent tool tail that fits a ladder, the SDK carries that view into later steps, and the turn emits one compaction event instead of one per step.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
One envelope carrying Add, Update, Move, and Delete file markers, validated in full before anything is written: a failure on the fourth file leaves the first three untouched. Permission rules match every path a patch touches, and the write guard and secret checks cover it alongside the other write tools.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>