1402 lines
101 KiB
Markdown
1402 lines
101 KiB
Markdown
# Plan: zesdex — Rust TUI AI Coding + Security Agent (koma architectural clone)
|
||
|
||
**Source reference**: `/mnt/code/koma/` (Rust `src-agent` + Python `src-security` + Python `src-internet`)
|
||
**Complexity**: Large (multi-session, phased delivery)
|
||
|
||
## Summary
|
||
|
||
`zesdex` is a from-scratch Rust rebuild of koma's full architecture: a `ratatui` TUI
|
||
agent driven by an OpenRouter-compatible LLM backend, with a rich sandboxed tool set
|
||
(filesystem, shell, git, web, memory, sub-agents), a three-layer approval harness,
|
||
MCP client support, OAuth login, a detachable daemon/client IPC split for
|
||
multi-session use, and an opt-in security-tooling module wired to a Python sidecar
|
||
daemon (mirroring `koma_sec_daemon`) for **authorized** pentesting/CTF/research use
|
||
against systems the operator owns or is explicitly authorized to test.
|
||
|
||
Four explicit deviations from the literal 1:1 clone, all already confirmed with you:
|
||
|
||
1. **Approval harness stays, reshaped as an opt-in speed dial, not deleted.** Auto
|
||
mode is genuinely fast (WC + TAC run inline, no `y/n` prompts, mutations proceed
|
||
immediately) — that satisfies "unrestricted auto-approve at maximum speed." What
|
||
does **not** get removed: the deterministic **catastrophic-op guard**
|
||
(`check_destructive` in `git_operator`, the recursive-delete-outside-workspace
|
||
check, force-push-over-history, secret/credential-file exfil patterns in
|
||
`web_download`/`bash`). That guard is not a per-turn approval prompt — it never
|
||
blocks routine work — it is a hard-coded refusal on a short, named list of
|
||
irreversible operations, and it fires even in Yolo/Auto mode. Everything else
|
||
(file writes, edits, ordinary shell commands, git commits/branches, web fetches)
|
||
runs with zero prompts in Auto/Yolo mode, exactly as requested.
|
||
2. **Security toolkit ships as an explicit opt-in module** (`/security` panel,
|
||
disabled by default, requires the user to explicitly enable it per koma's own
|
||
`security_enabled`/`yolo_armed` gating pattern) with a startup banner and
|
||
`--security-install` flow that states the authorized-use scope, mirroring
|
||
koma's own "curated, opt-in suite... for authorized testing and research"
|
||
framing. `sec_*` tools remain harness-exempt (per koma's own design — these are
|
||
explicit user-invoked offensive actions, not silent mutations to gate), but the
|
||
catastrophic-op guard's network/destructive-target patterns still apply where
|
||
they'd protect the *operator's own machine* (e.g., a `sec_z3`/`sec_sage` script
|
||
that shells out to `rm -rf /`).
|
||
3. **koma's `task` subagent primitive (Phase 6) is replaced wholesale by a dynamic
|
||
workflow orchestration engine**, modeled directly on the `Workflow` tool this
|
||
very session runs on — not on anything in koma, which has no equivalent. See
|
||
"Dynamic Workflow Orchestration" in the digest below and the rewritten Phase 6.
|
||
Three decisions you confirmed:
|
||
- **Fan-out is model-decided, not forced.** Simple requests are handled inline by
|
||
the main agent with its own tools, exactly as today; the model only invokes the
|
||
workflow engine when it judges a task benefits from decomposition, parallelism,
|
||
or independent verification passes.
|
||
- **Full replacement, not a layer on top.** There is no separate simple
|
||
"delegate one subagent" tool anymore — a single-agent delegation is just a
|
||
one-line workflow script (`agent(...)` with no `parallel`/`pipeline` wrapper),
|
||
so the workflow engine is the *only* delegation surface, keeping one mental
|
||
model instead of two.
|
||
- **UI: a separate live progress-tree panel**, auto-shown the instant a workflow
|
||
starts (phase → agent → status), dismissible/re-summonable without killing the
|
||
run, while the Main Execution Log keeps behaving exactly as specified for
|
||
everything else.
|
||
4. **NEW — Reasoned mutations + self-learning quality loop**, added after your
|
||
latest request: every `write`/`edit` tool call must carry a required `reason`
|
||
argument (rejected deterministically, before touching the filesystem, if
|
||
missing/empty), so every mutation is auditable after the fact. Once an agentic
|
||
turn finishes and any files were touched, zesdex automatically — no model
|
||
decision involved — spawns a detached background subagent (reusing Phase 6's
|
||
subagent engine in koma's existing detached-nudge mode) that reviews the
|
||
touched files against their stated reasons and the surrounding code's own
|
||
conventions, and folds anything worth keeping back into per-project memory as
|
||
a new `lesson`-typed entry. Because the memory index is already injected into
|
||
every future turn's system prompt (koma's existing mechanism, unchanged), this
|
||
closes a real feedback loop: code quality compounds across turns instead of
|
||
the agent repeating the same mistake every session. See "Reasoned Mutations &
|
||
Self-Learning Quality Loop" in the digest below and the new Phase 7.
|
||
|
||
Following your explicit request for concrete self-learning feature
|
||
suggestions, four core additions and three advanced additions were confirmed
|
||
and are folded into the same digest section and Phase 7 (advanced ones note
|
||
their own phase where they land instead): **outcome-linked learning** (lessons
|
||
backed by real `build`/`test` results, not just LLM opinion), **lesson
|
||
lifecycle** (decay/staleness + contradiction detection, never silent auto-
|
||
delete), **human-in-the-loop calibration** (a keep/discard toast before a
|
||
lesson becomes permanent, auto-keep in Auto/Yolo but always visible/
|
||
reversible), and a **global cross-project lesson tier** (mirrors the Builtin→
|
||
Global→Session tiering Phase 6 already builds for agent definitions) as core
|
||
Phase 7 scope; **batch/offline pattern mining**, a **`/learning` dashboard**
|
||
(delivered in Phase 14 alongside the rest of the remaining UI surface), and
|
||
**escalation on repeated violations** as the advanced layer on top.
|
||
|
||
A follow-up round of suggestions added three more confirmed additions, all
|
||
still Phase 7/Phase 14 scope: **lesson graduation into a real lint check**
|
||
(a lesson that survives repeatedly at `confidence: verified` stops being
|
||
just injected text and becomes a deterministic pre-write check), a
|
||
**session-end retrospective with per-lesson provenance** (one holistic
|
||
summary per session, plus every lesson traceable back to the exact turn/
|
||
edit that produced it), and **user-authored lessons with consensus-gated
|
||
global promotion and token-budget visibility** (`/lesson` to author a
|
||
`confidence: human` entry directly, a style-guide import path, 2-of-2
|
||
independent-reviewer agreement required before any lesson reaches `global`
|
||
scope, and the reviewer loop's own token spend surfaced on the usage
|
||
dashboard). Per your explicit requirement, **every one of these — and
|
||
everything already in Phase 7 — is language-agnostic by design**: nothing
|
||
in the self-learning loop hard-codes Rust or any single ecosystem
|
||
(build/test detection, lint graduation, and lesson content all generalize
|
||
across whatever language a given project/workspace is written in — zesdex
|
||
itself being written in Rust is incidental to this feature, not a
|
||
constraint on it).
|
||
|
||
A third round added six more confirmed additions, still Phase 7/Phase 14
|
||
scope: **shadow-mode enforcement** (a lint-graduation candidate runs silent
|
||
for a trial window, logging would-have-fired counts before ever surfacing
|
||
to the model, so a blunt rule is caught before it annoys anyone), **cross-
|
||
subagent learning within one workflow run** (a finding from one `agent()`
|
||
thunk can reach sibling thunks still running in the *same* `workflow_run`
|
||
call, not just the next turn/session via memory), **adaptive review
|
||
frequency** (consecutive empty review passes throttle the reviewer down;
|
||
fresh violations un-throttle it back up — the token-spend counterpart to
|
||
escalation), **quality trend metrics on the usage dashboard** (build/test
|
||
pass-rate, escalation count, and per-lesson trigger count charted over
|
||
time, with lessons that fire often but never correlate with fewer
|
||
escalations flagged as `forget` candidates), **structured lesson bodies**
|
||
(a `before`/`after` snippet pulled straight from the triggering diff,
|
||
alongside the prose description, so a lesson is concretely actionable, not
|
||
just abstract), and **lesson-pack export/import** (an opt-in, portable
|
||
bundle of a project's lessons for moving to another machine or sharing
|
||
with a team, independent of a shared `~/.zesdex/`).
|
||
|
||
Per your request to focus on the plan itself before any code, Phase 7 —
|
||
which had grown to cover all seventeen self-learning additions across
|
||
three rounds of suggestions — is now **split into Phase 7 (MVP core:
|
||
`reason`, the ledger, the automatic reviewer) plus four independently
|
||
deferrable sub-phases 7B (outcome-linking/lifecycle/calibration/global
|
||
tier), 7C (escalation/adaptive frequency/batch mining), 7D (lint
|
||
graduation, shadow-mode gated), and 7E (retrospectives/provenance/
|
||
user-authored/consensus/export/trend metrics)**, restoring the
|
||
"independently buildable/testable" property every other phase in this
|
||
plan already has. Phase 7 alone is a complete, working answer to your
|
||
original request; 7B–7E are refinements you can build immediately after
|
||
or defer indefinitely. See the rewritten Phase 7 family in Tasks below.
|
||
|
||
No comments in generated Rust source, per your original instruction — self-documenting
|
||
naming instead. No stubs; every module in a phase is fully implemented before that
|
||
phase is considered done. Cargo.toml matches your requested metadata exactly.
|
||
|
||
## Reference Architecture Digest (from `/mnt/code/koma/`)
|
||
|
||
Captured here so later phases can cite it without re-reading the whole reference tree.
|
||
|
||
**Workspace**: Cargo workspace with a single `agent` binary crate (`src-agent`),
|
||
edition 2021, plus two out-of-process Python sidecars (`src-security`,
|
||
`src-internet`) invoked via subprocess/daemon, not compiled in.
|
||
|
||
**Core dependency set** (from `src-agent/Cargo.toml`, informs zesdex's `Cargo.toml`):
|
||
`ratatui 0.30`, `tokio` (`rt-multi-thread, macros, sync, time, net, io-util, signal`),
|
||
`reqwest` (`json, stream, blocking, native-tls-vendored`), `serde`/`serde_json`/
|
||
`serde_yaml_ng`, `anyhow`, `uuid` (`v4, v5`), `dirs`, `futures-util`,
|
||
`pulldown-cmark`, `syntect` (`default-fancy`), `rusqlite` (`bundled`), `ignore`,
|
||
`regex`, `globset`, `infer`, `base64`, `sha2`, `libc`, `rmcp` (MCP client:
|
||
`client, transport-child-process, transport-streamable-http-client-reqwest, macros`),
|
||
`include_dir`, plus web-scraping-adjacent crates (`dom_smoothie`, `fast_html2md`,
|
||
`scraper`, `url`, `percent-encoding`) for the lightweight (non-Python) web-fetch path.
|
||
|
||
**MVC data flow**: `KeyEvent → controller/input::handle_key → Action →
|
||
runtime/actions::apply_action → state mutation → view::draw`. Slash commands:
|
||
`controller/command::parse → Command → runtime/commands::apply_slash`. The
|
||
controller layer never mutates state; `apply_action`/`apply_slash` own all
|
||
mutation and async spawns; the view is read-only.
|
||
|
||
**Async/streaming bridge**: one `tokio::runtime::Runtime` created in `main.rs`; the
|
||
main loop is synchronous (`try_recv`/poll, no `.await` in the hot loop). One
|
||
`mpsc::unbounded_channel` per request stored in `active_rx`; harness verdicts ride
|
||
a **separate** `harness_rx` channel so they never interleave with stream tokens.
|
||
`StreamEvent` variants: `Token, Reasoning, Usage, ToolCalls, Done, Error, Compacted,
|
||
HarnessVerdict`.
|
||
|
||
**Agentic loop**: model streams `tool_calls`; on `Done`, `advance_turn` commits the
|
||
assistant message, then either ends the turn or runs `process_tools` → gate/execute
|
||
each call → `finish_tool_round` → re-invoke the model. Loop continues until no more
|
||
tool calls or `MAX_AGENT_STEPS` (40) is hit. No forced plan gate — tools run on the
|
||
first call unless the model explicitly requests Plan mode via `plan_enter`.
|
||
|
||
**`Tool` trait**: `name() -> &'static str`, `description() -> &'static str`,
|
||
`parameters() -> Value` (JSON Schema), `run(&self, ctx: &ToolCtx, args: &Value) ->
|
||
Result<String>`. `ToolCtx` carries `workspace`, `workspaces: Vec<PathBuf>`,
|
||
`dir_cache`, plus (per the fuller subagent report) memory dir, MCP manager, security
|
||
manager, bash-saving flag.
|
||
|
||
**Deterministic tool inventory** (`tool_is_risky`: `write | delete | edit | bash |
|
||
git_operator | web_download`): `read, grep, glob, write, edit, delete, bash,
|
||
bash_output, bash_kill, cd, dir_list, dir_cache_update, pong, remember, forget,
|
||
recall, task, task_output, task_kill, todowrite, web_fetch, web_search,
|
||
web_download, git_cred, git_operator, git_worktree, plan_enter, plan_ready,
|
||
seqthink` — plus dynamic `mcp__*` and `sec_*` tools. `DEFERRED_TOOLS` (blocking I/O
|
||
tools) run on a plain `std::thread`, never inline on the event loop or inside the
|
||
tokio runtime directly.
|
||
|
||
`write`/`edit` (koma reference, `tool/fs/write.rs` + `tool/fs/edit.rs`): `write`
|
||
takes `path` + `content`, always creates parent dirs, always writes/overwrites,
|
||
reindexes the dir cache. `edit` takes `path` + `old` + `new` + optional
|
||
`replace_all`, requires the file to exist, refuses (returns a string result, not an
|
||
error) if `old` is missing or non-unique unless `replace_all=true`. Neither carries
|
||
any notion of "why" in koma today — zesdex's `reason` requirement (see below) is a
|
||
net-new addition on top of this exact shape, not a koma pattern.
|
||
|
||
**Three-layer harness** (`app/harness.rs`): WC (workspace containment check,
|
||
deterministic, gates the whole turn) → PC (advisory prompt classifier, fire-and-
|
||
forget, fail-open) → TAC (per-risky-call classifier, intent-aware, mode-dependent
|
||
fail-open/fail-closed policy). `CLASSIFY_TIMEOUT = 12s`. Modes: `Auto, Normal,
|
||
Plan, Yolo` (`AgentMode::cycle`), Yolo requires explicit arming
|
||
(`yolo_armed`) and bypasses TAC/`y/n` entirely — only deterministic guards remain.
|
||
|
||
**`shell_filter`**: post-processes `bash`/`git_operator` stdout — smart per-command
|
||
filters (cargo, git status/log/diff) tried first, static regex spec table
|
||
(npm/pip/docker/make/wget-curl) as fallback, full raw output always teed to disk
|
||
when trimmed/truncated/non-zero-exit.
|
||
|
||
**`git_operator`/`git_worktree`**: direct argv exec (no shell), SSH key injection
|
||
via `git_cred`, `GIT_TERMINAL_PROMPT=0`, and a **named destructive-op guard**
|
||
(`check_destructive`) requiring explicit `confirm_destructive=true` for force-push,
|
||
`reset --hard`, `clean -f/-d/-x`, `branch -D`, force checkout/switch/restore,
|
||
`stash drop/clear`, `tag -d`, `update-ref -d`, `filter-branch`, `gc --prune`.
|
||
Worktrees live in a shadow dir outside the repo.
|
||
|
||
**Subagents**: up to 5 concurrent (`MAX_SUBAGENTS`), agent definitions are
|
||
frontmatter+markdown files merged Builtin→Global→Session tier, isolated
|
||
`Conversation` seed (own system prompt, no shared history with parent), own
|
||
restricted tool allow-list (always minus `task`/`task_output`/`task_kill`, and
|
||
zero MCP tools), engine runs the same commit/execute/classify loop non-
|
||
interactively with fail-closed risky-tool gating (no human to ask), completion
|
||
folds back via blocking-park / chat-append / **detached-nudge** depending on
|
||
invocation mode. The detached-nudge mode matters most for zesdex: it lets a
|
||
subagent run fully off the critical path and only surface a synthetic-user-turn
|
||
nudge (or, per Phase 7, a collapsed system note) once the session goes idle,
|
||
without blocking whatever the user does next.
|
||
|
||
**Background bash**: `run_in_background: true` bash calls spawn a plain OS thread
|
||
(not tokio — child process needs to run outside the runtime), answer the model
|
||
immediately with a job id, poll via `bash_output`, kill via `bash_kill`
|
||
(SIGTERM to the recorded pid), completion is toast + buffered synthetic-user-turn
|
||
nudge once the session goes idle.
|
||
|
||
**IPC/daemon**: **daemon-per-session**, not daemon-per-machine — each session UUID
|
||
gets its own headless process at `~/.<app>/run/<session_id>.sock`. Transport: Unix
|
||
domain socket, 4-byte-BE-length-prefixed JSON frames (`FrameReader` reassembles
|
||
split/coalesced reads, 64 MiB hard cap enforced at the prefix). `bind()` success is
|
||
the sole liveness oracle (never a PID file). Per-client connection task splits
|
||
into independent async read/write halves bridged to the **synchronous** daemon
|
||
loop via `std::sync::mpsc` (daemon loop stays sync so it shares turn-servicing
|
||
code with a hypothetical standalone/no-daemon mode). `ClientRequest`/`DaemonFrame`
|
||
message sets cover attach/detach/list/status/submit/shell/key/paste/approve/plan-
|
||
decision/new/quit/switch-foreground. Multiple clients can attach one daemon
|
||
(controller + viewers, promotion on controller disconnect). Self-exit after a
|
||
grace window once all sessions are tombstoned and no client is enrolled.
|
||
|
||
**Snapshot/diff**: daemon builds a pure `StateSnapshot` from live `AppState`;
|
||
per-client `diff()` decides `needs_full` (structural change → resend full
|
||
snapshot; correctness-first, "when in doubt, ask for a full snapshot") vs a
|
||
cheap `Vec<StateDelta>` for pure-append/status/scroll changes. Client applies
|
||
frames into a **shadow `AppState`** and reuses the *same* `view::draw` renderer,
|
||
so daemon and thin-client rendering can never diverge.
|
||
|
||
**Session locking**: one `session.lock` file per session holding the owner PID;
|
||
liveness via `kill(pid, 0)` (portable, not `/proc` parsing); stale locks
|
||
opportunistically cleared. Lock always written with the **daemon's** PID, never
|
||
the thin client's.
|
||
|
||
**MCP**: `rmcp`-based client, two transports (stdio child process, streamable
|
||
HTTP), tools namespaced `mcp__<sanitized_server>__<tool>`, `Local` vs `Proxy`
|
||
(shared global MCP daemon) backend behind one facade so multiple session-daemons
|
||
don't each spawn duplicate heavyweight MCP servers, sync↔async bridge for
|
||
`Tool::run` compatibility, 20s connect / 60s call timeouts, non-blocking
|
||
background connect/reconnect with a generation counter to discard stale in-flight
|
||
connects.
|
||
|
||
**OAuth**: PKCE for the primary provider (loopback `TcpListener` on a fixed local
|
||
port for the redirect, hand-parsed HTTP, CSRF `state` check), device-authorization
|
||
polling for a secondary provider, unverified local JWT payload decode for display
|
||
metadata only, single-flight per-provider refresh with a stale/near-expiry lead
|
||
window and an `unrecoverable` latch on `invalid_grant`.
|
||
|
||
**TUI layout**: header (mode indicator) → transcript (scrollable, markdown-
|
||
rendered assistant turns, plain-text user/streaming turns) → optional pending-
|
||
steer panel → model-name row → input box (dynamic height, capped at half the
|
||
frame) → status bar. Markdown rendering emits pre-wrapped `Vec<Vec<Span>>` so
|
||
line count matches on-screen rows exactly (needed for scroll math); `syntect`
|
||
lazy-singleton highlighting in fenced code blocks with hard-splitting on
|
||
over-width lines; tables/lists/blockquotes word-wrap via a shared char-level
|
||
wrapper. Theme: a `Palette` struct of ~12 semantic colors + a 5-stop heatmap
|
||
ramp, computed fresh per frame from config, several named palettes plus an
|
||
accent-color resolver.
|
||
|
||
**Mode state machine**: far more than "5 modes" — onboarding, provider-OAuth
|
||
onboarding, credentials form, session picker/hub, chat, loading splash, settings,
|
||
agents editor, MCP manager, help, effort picker, usage dashboard, message rewind,
|
||
quit confirm, security panel, background-bash panel, todo panel. Each mode is a
|
||
struct variant on one `Mode` enum, exactly one active at a time, stored per-session.
|
||
|
||
**Security sidecar** (`koma_sec_daemon` equivalent): long-lived Python child
|
||
process, newline-delimited JSON frames over stdin/stdout (not a socket) —
|
||
handshake with a shared token, `call`/`health`/`install` ops, single-threaded
|
||
serialized dispatch. Tool roster spans web (http, sqlmap, nuclei, ffuf, dalfox,
|
||
zap, xss-confirm), crypto (z3, sage, rsa, factordb, lattice/fpylll, hashcat,
|
||
hashid, generic decode), reverse-engineering (js deobfuscate/unminify, source-map
|
||
recovery, wasm decompile), and pwn (static triage, ROPgadget, pwntools remote
|
||
session, exploit-template scaffold). Sandbox = wall-clock timeout + optional
|
||
RLIMIT_AS/RLIMIT_CPU + optional (opt-in, currently-unused-upstream) bubblewrap
|
||
network unshare — no container/chroot/seccomp. Tiered installer (pip / GitHub-
|
||
release binary / gem / manual-detect-only) with a health-probe surfaced in the UI.
|
||
|
||
**Internet sidecar** (`scrapion_agent` equivalent): **stateless**, one-shot
|
||
subprocess per call (`python -m ... --json "<url-or-query>"`), Playwright-Firefox-
|
||
headless search+scrape with stealth tweaks, HTML→Markdown conversion, JSON out on
|
||
stdout. Rust side spawns it on a plain OS thread (not tokio) with a `recv_timeout`
|
||
guard and falls back to a raw HTTP fetch on any failure.
|
||
|
||
**Dynamic Workflow Orchestration** (new — not present in koma; modeled on this
|
||
session's own `Workflow` tool): a script-driven orchestration layer that lets the
|
||
main agent decompose a task into many subagents running in parallel, pipeline, or
|
||
phased stages, rather than koma's one-subagent-at-a-time `task` primitive.
|
||
|
||
- **Trigger**: exposed to the model as one new tool, `workflow_run(script, args)`.
|
||
There is no separate "spawn one subagent" tool — a single delegated task is
|
||
simply a one-`agent()`-call script, so `workflow_run` is the sole delegation
|
||
surface. The system prompt instructs the model to reach for it when a task
|
||
benefits from decomposition/parallelism/independent verification, and to skip
|
||
it (use its own `read`/`grep`/`edit`/`bash` directly) for anything simple —
|
||
this is model judgement, not a hard trigger rule, matching your "dynamic, model
|
||
decides" choice.
|
||
- **Script primitives** (a small embedded scripting surface, not a general
|
||
scripting language — deliberately narrow so it stays auditable and
|
||
serializable into the persisted transcript):
|
||
- `agent(prompt, opts?) -> AgentResult` — spawn one subagent (own isolated
|
||
conversation seed, own restricted tool allow-list, same non-interactive
|
||
engine loop as koma's subagent engine: commit/execute/classify with
|
||
fail-closed risky-tool gating since there's no human to ask). `opts` covers
|
||
`label`, `phase`, `model` override, `effort` override.
|
||
- `parallel(thunks) -> Vec<AgentResult>` — run N agent thunks concurrently,
|
||
barrier: all must finish before returning. Bounded by a concurrency cap
|
||
(mirrors koma's `MAX_SUBAGENTS`-style ceiling, default 5, configurable).
|
||
- `pipeline(items, stage1, stage2, ...) -> Vec<Any>` — run each item through
|
||
every stage independently with **no barrier between stages** (item A can be
|
||
on stage 3 while item B is on stage 1); the default recommended shape for
|
||
multi-stage fan-out because wall-clock is the slowest single chain, not the
|
||
sum of slowest-per-stage.
|
||
- `phase(title)` — label subsequent `agent()` calls under a named phase for
|
||
the progress-tree UI; purely cosmetic/organizational, no execution effect.
|
||
- `log(message)` — narrator line surfaced in the progress panel above the tree.
|
||
- **Execution model**: `workflow_run` itself is a `DEFERRED_TOOL` (per the
|
||
existing `DEFERRED_TOOLS` pattern) whose body runs on the tokio runtime (it
|
||
needs real concurrency for `parallel`/`pipeline`, unlike the blocking-thread
|
||
tools); each `agent()` call inside it reuses the exact same subagent spawn/
|
||
engine/fold-back machinery Phase 6 already needed for koma parity — so
|
||
`workflow_run` is a thin orchestration layer *over* individual subagent
|
||
spawns, not a reimplementation of subagent execution. Concurrency, isolation
|
||
(own `Conversation` seed, own tool allow-list, zero recursion into
|
||
`workflow_run` itself to bound blast radius), and the fail-closed risky-tool
|
||
gate are all inherited unchanged from the subagent engine design.
|
||
- **Result delivery**: `workflow_run` blocks the calling turn (like a normal
|
||
tool call) and returns a structured summary (per-agent status + final return
|
||
value) as the tool result once every stage completes; the model then
|
||
synthesizes a response from that structured result, same as any other tool
|
||
round-trip — no separate detached/nudge mode for v1 (koma's detached-task
|
||
nudge pattern is not carried over here, since workflow runs are expected to be
|
||
the user-visible unit of work, not a background fire-and-forget job).
|
||
- **TUI surface**: a dedicated **Workflow Progress panel**, auto-shown the
|
||
instant `workflow_run` starts (overlaying/pushing above the transcript region,
|
||
matching koma's overlay pattern for `Bash`/`Todo`/`MessageRewind` modes — same
|
||
`layout_chunks` reuse, new overlay content), rendering a live tree: phase →
|
||
per-agent row → status glyph (queued/running/done/error) → elapsed time.
|
||
Dismissible with Esc without killing the run (run keeps streaming in the
|
||
background; re-summon via a key binding or `$`-style panel toggle, mirroring
|
||
koma's sub-agents panel UX). On completion, the panel's final state persists
|
||
in scrollback as a collapsed summary block in the Main Execution Log, and the
|
||
panel auto-closes after a short delay unless pinned open.
|
||
- **Persistence**: each `workflow_run` call and its final structured result is
|
||
logged to the session transcript like any tool call/result pair — no separate
|
||
workflow-history store for v1.
|
||
|
||
**Reasoned Mutations & Self-Learning Quality Loop** (new — not present in koma;
|
||
added for your "tiap edit atau write harus ada parameter alasan... bg agent untuk
|
||
verifikasi kualitas... self learning" request):
|
||
|
||
- **`reason` becomes a required argument on `write` and `edit`** (koma's
|
||
reference shapes above gain one field each): `{"path", "content", "reason"}`
|
||
for `write`, `{"path", "old", "new", "replace_all"?, "reason"}` for `edit`.
|
||
Enforcement is deterministic and happens *before* any filesystem call, exactly
|
||
like the existing `arg_str` required-field pattern already used for `path`/
|
||
`content`/`old`/`new` — a missing or blank `reason` returns
|
||
`error: missing required string argument 'reason'` and the file is never
|
||
touched. This is intentionally the same shape as `tool_is_risky`/WC: a name-
|
||
based, always-on, non-negotiable check, not something the classifier or agent
|
||
mode can waive. `delete` and `bash` are explicitly NOT changed — the user's
|
||
ask was scoped to edit/write, and widening it would duplicate what the
|
||
harness/catastrophic-op guard already cover for those.
|
||
- **Every accepted call appends one line to an append-only per-session ledger**,
|
||
`<session_dir>/edits.jsonl` (same directory tier as `session.lock`/
|
||
`settings.json`, same atomic-append discipline as `model::memory::atomic_write`
|
||
uses for whole-file writes — here it's an `O_APPEND` line write, since the
|
||
ledger is never rewritten in place). Each line: `{ts, tool, path, reason,
|
||
content_sha256, bytes_delta}`. The ledger is the audit trail the reviewer
|
||
(below) reads; it is never truncated or rotated automatically in v1.
|
||
- **Auto-triggered review, not model-invoked.** When `advance_turn` reaches a
|
||
turn boundary (no more `tool_calls`, same point `Done` already fires today)
|
||
and that turn's slice of `edits.jsonl` is non-empty, the event loop — not the
|
||
model, not a tool call — spawns exactly one subagent using Phase 6's existing
|
||
engine in koma's **detached-nudge** invocation mode, seeded with a new
|
||
built-in agent-definition persona, `quality-reviewer`. Its tool allow-list is
|
||
read-only plus memory: `read, grep, glob, recall, remember`. It never sees
|
||
`write`/`edit`/`bash`, so it cannot itself mutate anything — a review pass can
|
||
only ever produce a memory note or nothing.
|
||
- Input: the turn's ledger lines (path + reason + diff-relevant content) plus
|
||
a `recall` of the project's existing `lesson`-typed memories, so the
|
||
reviewer can check "have I already said this" before writing a new one —
|
||
this is what keeps `MEMORY.md` from growing an unbounded stream of
|
||
restated observations. If an existing lesson already covers the point, the
|
||
reviewer does nothing (no duplicate `remember` call).
|
||
- Output: zero or more `remember(..., type="lesson")` calls for anything
|
||
genuinely worth carrying forward (a real convention violated, a recurring
|
||
smell, a reason that didn't match what the diff actually did), plus exactly
|
||
one short verdict string returned as the subagent's final result.
|
||
- The verdict is rendered as a **collapsed system note** in the Main
|
||
Execution Log (a new `StreamEvent`-adjacent variant, e.g. `SystemNote`,
|
||
routed the same way `Compacted`/`HarnessVerdict` already ride their own
|
||
channel so it can never interleave with live stream tokens) — it is
|
||
explicitly NOT appended to the model-visible conversation history, so
|
||
review passes cost the reviewer's own tokens only and never grow the main
|
||
agent's context.
|
||
- Because `model::memory::load_memory_index` already injects the whole
|
||
`MEMORY.md` index (now including any `lesson` entries) into every future
|
||
turn's system prompt, this is the entire feedback loop — no new injection
|
||
plumbing needed, only the new `kind` value and the auto-trigger.
|
||
- **Bounded, not free-running**: at most one reviewer subagent runs at a time
|
||
(reviews queue rather than pile up — reuses the same bounded-concurrency
|
||
primitive Phase 6 builds for `parallel()`, capacity fixed at 1 for this
|
||
queue); the trigger is skipped entirely while `AgentMode::Plan` is active
|
||
(nothing mutates in Plan mode, so `edits.jsonl` can't have a non-empty slice)
|
||
and skipped for edits made *by* a reviewer or workflow subagent itself (no
|
||
recursive review-the-reviewer chains — ledger lines carry an `origin: main |
|
||
subagent | reviewer` tag precisely so this filter is a cheap equality check,
|
||
not a heuristic).
|
||
- **Explicitly reuses, does not reinvent**: Phase 6's subagent spawn/engine/
|
||
fold-back machinery (detached-nudge mode already exists there for other
|
||
reasons); `model::memory`'s `write_memory`/`read_memory`/`slugify` primitives,
|
||
widened by one `kind` value; the existing pattern of a dedicated side-channel
|
||
for out-of-band events (`harness_rx` today, a parallel `notes_rx` for
|
||
`SystemNote`).
|
||
|
||
**Outcome-linked learning** (language-agnostic by design — zesdex itself being
|
||
a Rust project is incidental; the mechanism below never assumes Rust): the
|
||
reviewer's input is upgraded from "reasons + diff only" to "reasons + diff +
|
||
real verification result." Detection of "how do I build/test this project"
|
||
is a small ordered probe, not a hard-coded `cargo` call: a `verify_command`
|
||
in `settings.json` always wins if the user sets one; otherwise zesdex probes
|
||
for well-known project markers already visible to `dir_cache` (`Cargo.toml` →
|
||
`cargo build && cargo test`, `package.json` → its `scripts.test`/`scripts.
|
||
build` if present, `pyproject.toml`/`requirements.txt` → `pytest`/`python -m
|
||
pytest` if discoverable, `go.mod` → `go build ./... && go test ./...`, and so
|
||
on for the other common ecosystems); if no marker matches, verification is
|
||
simply skipped for that project and the reviewer runs on reasons+diff alone.
|
||
This probe table lives in one named function (mirroring the `tool_is_risky`/
|
||
catastrophic-op-guard pattern of "one place to extend"), so adding a new
|
||
ecosystem later is a one-line addition, not a redesign. The resolved command
|
||
runs on a plain `std::thread`, capped at a short timeout, and hands the
|
||
reviewer the pass/fail + trimmed output alongside the ledger. A lesson born
|
||
from an actual build/test failure is written with `confidence: verified`; a
|
||
lesson born from the reviewer's own judgement alone (no reproducible failure)
|
||
is written with `confidence: opinion`. Both render in the memory index, but
|
||
`opinion` entries are visually distinguished (dimmer/marked) in every mode
|
||
that lists memories, so verified lessons are never confused with the
|
||
reviewer's own guesses. If the verification command itself is unavailable or
|
||
misconfigured, the reviewer still runs on reasons+diff alone — this is a
|
||
strict enhancement, never a hard dependency.
|
||
|
||
**Lesson lifecycle**: every `lesson` entry gains two frontmatter fields beyond
|
||
koma's existing `name/description/type` — `confidence` (`verified | opinion`,
|
||
from outcome-linked learning above) and `last_confirmed` (a timestamp, bumped
|
||
whenever a later review cites the same lesson as still applicable). A
|
||
lightweight idle-time sweep (piggybacks on the same idle-detection the
|
||
background-bash completion nudge already uses, no new polling loop) flags —
|
||
never silently deletes — any `lesson` whose `last_confirmed` is older than a
|
||
configurable staleness window as `stale: true` in its frontmatter; a stale
|
||
lesson still gets injected into the index (so it's still visible) but is
|
||
labeled so the model and the user both know it may no longer hold. Separately,
|
||
before any new `remember(type="lesson")` call is accepted (both from the
|
||
regular reviewer and human-in-the-loop confirmation below), a cheap
|
||
recall-and-compare pass checks the new lesson's text against existing lessons
|
||
for direct contradiction (e.g. "always use X" vs. a new "never use X"); a
|
||
detected contradiction blocks the auto-save and instead surfaces both entries
|
||
side by side in the calibration toast for a human decision — contradictions
|
||
are never silently resolved by whichever write happens to land last.
|
||
|
||
**Human-in-the-loop calibration**: a `remember(type="lesson")` call from the
|
||
quality-reviewer does not go straight into `MEMORY.md`. It lands first in a
|
||
small pending queue (one more small JSON file alongside `edits.jsonl`,
|
||
`pending_lessons.json`, same atomic-write discipline) and renders as a
|
||
collapsed toast in the Main Execution Log: `Lesson baru: <description>
|
||
[keep]/[discard]`, using the same `SystemNote` side-channel Phase 7 already
|
||
adds. In Auto/Yolo mode the toast auto-resolves to `keep` after a short grace
|
||
window (matching the "unrestricted, no prompts" mandate for routine flow) but
|
||
the toast and its outcome remain visible/scrollable in the log and reversible
|
||
at any time via `forget`; in Normal/Plan mode (where a human is already
|
||
expected to be attentive) the toast waits for an explicit keypress instead of
|
||
auto-resolving. Either way, nothing about this blocks the *main* agent's own
|
||
turn — calibration only gates the reviewer's side-channel writes, never the
|
||
user-facing conversation.
|
||
|
||
**Global cross-project lesson tier**: mirrors the Builtin→Global→Session
|
||
tiering Phase 6 already builds for agent-definition files. A `lesson` memory
|
||
gains a `scope` field (`project | global`); `project` is the existing default
|
||
(`<pwd_bucket_dir>/memory/`, unchanged), `global` writes instead to a new
|
||
`~/.zesdex/memory/` store shared across every project. The reviewer defaults
|
||
every lesson to `project` scope; a lesson is only promoted to `global` when
|
||
either (a) the human-in-the-loop toast is used to explicitly promote it, or
|
||
(b) the escalation-on-repeated-violations mechanism (below) fires the same
|
||
lesson's underlying pattern across two or more distinct projects — never
|
||
silently promoted by the reviewer's own judgement alone, since a
|
||
project-specific convention wrongly promoted to global would pollute every
|
||
other project's context. `load_memory_index` (Phase 4/existing koma mechanism)
|
||
is extended to concatenate the global index ahead of the project index when
|
||
building the system prompt, so global lessons are always present and project
|
||
lessons can still override/refine them by appearing after.
|
||
|
||
**Batch/offline pattern mining** (advanced, Phase 7 scope but lower priority
|
||
within it): beyond the per-turn reviewer, an idle-triggered (same idle
|
||
detection as the lifecycle sweep) deeper pass periodically walks the full
|
||
`edits.jsonl` history (not just the latest turn's slice) plus `git log -p`
|
||
for the same file set, looking for cross-turn patterns a single-turn reviewer
|
||
structurally cannot see (e.g., "every time `parser.rs` changes, the matching
|
||
test file does not"). Runs as one more detached subagent, same allow-list and
|
||
confidence/calibration path as the per-turn reviewer — it is the same
|
||
mechanism at a longer time horizon, not a separate code path. Rate-limited to
|
||
at most once per idle window so it can never compound with the per-turn
|
||
reviewer into runaway subagent spawning.
|
||
|
||
**Escalation on repeated violations**: every accepted ledger line is checked
|
||
(cheap substring/embedding-free heuristic first — a lesson's `slug` matched
|
||
against the touched file's recent `edits.jsonl` history — deferred to the
|
||
reviewer's own judgement only when the heuristic is ambiguous) against the
|
||
existing lesson index for a *repeat* violation of something already
|
||
`remember`ed. On the third repeat of the same lesson within a configurable
|
||
window, zesdex stops treating it as routine background learning and instead
|
||
surfaces a non-collapsed, attention-grabbing note directly in the Main
|
||
Execution Log (not the quieter `SystemNote` styling used for ordinary
|
||
verdicts) — the premise being that if index-injection alone hasn't stopped
|
||
the pattern after three tries, silently trusting the system prompt again is
|
||
not working and a human should see it. This counter is also the trigger
|
||
condition that can promote a lesson from `project` to `global` scope per the
|
||
tiering feature above, when the same pattern repeats across multiple distinct
|
||
project directories rather than one.
|
||
|
||
**Lesson graduation into a lint check** (language-agnostic — the check itself
|
||
is expressed as a pattern-matching rule against the tool call's own arguments/
|
||
diff text, never a Rust-specific AST): once a lesson has stayed at
|
||
`confidence: verified` and been cited by the escalation counter above at
|
||
least once (i.e., it has already proven itself both true and still being
|
||
violated), it becomes eligible for graduation. A graduated lesson gains a
|
||
`check` frontmatter field — a small structured rule (substring/regex against
|
||
the file's new content, or a path-glob + required-companion-file rule, e.g.
|
||
"editing `parser.*` without also touching a matching `*test*` file") — and
|
||
from that point on is enforced the same way the `reason` argument is: a
|
||
deterministic, always-on check inside `write`/`edit`, evaluated *before* the
|
||
harness/classifier layer, same shared pattern as `arg_str`. A failed
|
||
graduated-lesson check does not silently block the write; it returns a
|
||
result string naming the lesson and asking the model to address it or
|
||
explicitly override with a `reason` that acknowledges the exception — so
|
||
graduation raises the bar without ever becoming an unbreakable wall. Only a
|
||
human (via `/learning`, Phase 14) or the consensus-promotion mechanism below
|
||
can graduate or de-graduate a lesson — the reviewer itself cannot
|
||
self-promote into an enforced check, for the same "no single opinion becomes
|
||
policy" reason `global` promotion is gated.
|
||
|
||
**Session-end retrospective + per-lesson provenance**: every `lesson` memory
|
||
gains a `provenance` field recording the session id + turn index + ledger
|
||
line hash it was born from (or, for a `/lesson`-authored entry below, a
|
||
literal `"user"` marker instead) — a cheap addition since all of that is
|
||
already available at the point `remember` is called, purely for later
|
||
auditability (`/learning` can jump from a lesson straight to the originating
|
||
edit). Separately, when a session ends (explicit `/quit` or the daemon's
|
||
tombstone-on-idle path, Phase 11), one more detached subagent — reusing the
|
||
identical `quality-reviewer`-style engine, not a new code path — is hand a
|
||
session-wide view (every ledger line + every lesson touched this session,
|
||
not just the last turn) and produces a single holistic summary memory
|
||
(`type: retrospective`, one per session, not one per turn) capturing things a
|
||
per-turn reviewer structurally cannot ("this session repeatedly restructured
|
||
the same module," "three different lessons this session all pointed at the
|
||
same underlying gap"). Retrospectives are additive and separate from
|
||
`lesson` entries — they are not injected into the system prompt by default
|
||
(they're a human-facing summary, browsable from `/learning`), so they cannot
|
||
bloat future turns' context the way an unbounded `lesson` count could.
|
||
|
||
**User-authored lessons, consensus-gated global promotion, token-budget
|
||
visibility**: a new `/lesson <text>` slash command lets the user author a
|
||
`lesson` memory directly, tagged `confidence: human` — a rank above both
|
||
`verified` and `opinion`, since it is authoritative by construction and
|
||
therefore exempt from the contradiction-block/calibration-toast path
|
||
entirely (a human overriding a prior automated lesson is always accepted
|
||
immediately, never queued). A companion `/lesson import <path>` command feeds
|
||
an existing convention document (`CONTRIBUTING.md`, a style guide, any
|
||
markdown file the user points at) through the same `quality-reviewer`-style
|
||
subagent to extract candidate lessons in bulk, seeding the project's memory
|
||
from day one instead of starting from zero — extracted entries still land at
|
||
`confidence: opinion` (the source document is trusted, but the *extraction*
|
||
is still an LLM judgement call) and go through the normal calibration toast
|
||
before becoming permanent. Global promotion, already gated behind either
|
||
explicit human calibration or the multi-project repeated-violation counter
|
||
(per the tiering feature above), is tightened further: an automatic
|
||
(non-`/lesson`) promotion additionally requires two independent reviewer
|
||
subagents (spawned fresh, no shared context between them — cheap since
|
||
review subagents are already isolated by design) to agree the pattern
|
||
generalizes, before it's written to `~/.zesdex/memory/`; a single reviewer's
|
||
opinion is never sufficient for a `global`-scope write. A `/lesson`-authored
|
||
entry can still be promoted to `global` directly by the user with no
|
||
consensus requirement, same "human is authoritative" exemption as above.
|
||
Finally, every reviewer/retrospective/batch-mining subagent's token usage is
|
||
tagged with a `origin: self-learning` cost-ledger category (reusing whatever
|
||
per-turn cost tracking the usage dashboard, Phase 14, already accumulates for
|
||
the main agent) and surfaced as its own line item there, so the self-learning
|
||
loop's token cost is always visible against the session total rather than
|
||
invisibly folded into "background work."
|
||
|
||
**Shadow-mode enforcement**: a lint-graduation candidate (per the graduation
|
||
mechanism above) does not go straight from `confidence: verified` to an
|
||
active `check`. It first spends a configurable trial window (N subsequent
|
||
`write`/`edit` calls against matching files, default a small fixed count) in
|
||
`shadow: true` state: the same rule evaluation runs on every matching call,
|
||
but a match only increments a would-have-fired counter on the lesson's own
|
||
frontmatter — it never returns anything to the model and never blocks
|
||
anything. Only once the trial window closes does the rule either graduate to
|
||
a real active `check` (per the existing enforcement path) or, if its
|
||
would-have-fired rate looks too indiscriminate (fires on most edits to a
|
||
matching file rather than the specific pattern it was meant to catch — a
|
||
simple ratio threshold, not a judgement call), get flagged back down to a
|
||
plain `opinion` lesson instead of ever reaching enforcement. This reuses the
|
||
exact same rule-evaluation code path enforcement already needs (Phase 7); the
|
||
only new state is the `shadow`/`trial_count` pair on the lesson and the one
|
||
branch deciding graduate-vs-demote when the window closes.
|
||
|
||
**Cross-subagent learning within one workflow run**: today, a finding one
|
||
`agent()` thunk makes only reaches the rest of the system on the *next* turn,
|
||
once the reviewer has run and the memory index has been rebuilt — too late
|
||
to help sibling thunks still executing in the *same* `workflow_run` call.
|
||
`workflow_run`'s existing execution context (Phase 6) gains one small
|
||
in-memory, run-scoped broadcast channel (not persisted, not `MEMORY.md` —
|
||
this is intentionally ephemeral and scoped to the single run, distinct from
|
||
the durable per-project lesson store): any subagent spawned via `agent()` can
|
||
call a new restricted tool, `note_finding(text)`, whose only effect is to
|
||
push `text` onto that run's shared findings list; every *other* subagent
|
||
still active in the same run receives the current findings list prepended to
|
||
its next tool-round system context, cheaply (no re-spawn, no interruption of
|
||
in-flight work — it lands on the subagent's next turn boundary like any other
|
||
context update). Findings are ephemeral to the run: they do not automatically
|
||
become `lesson` memories (that still requires the normal reviewer pass after
|
||
the workflow completes, reading the same `edits.jsonl` ledger as any other
|
||
turn) — this feature closes the *within-run* propagation gap, not the
|
||
long-term memory gap, which the rest of Phase 7 already covers.
|
||
|
||
**Adaptive review frequency**: the per-turn reviewer trigger gains a small
|
||
run-length counter — consecutive review passes that produce zero new lessons
|
||
(no `remember` calls at all, verified or opinion) increment it; any pass that
|
||
produces at least one lesson resets it to zero. Past a small threshold of
|
||
consecutive empty passes, the trigger throttles: instead of reviewing every
|
||
qualifying turn, it skips a growing number of turns between reviews (simple
|
||
backoff, not a fixed schedule), directly reducing the `self-learning`-tagged
|
||
token spend the usage dashboard now surfaces. The moment a review does find
|
||
something (or the escalation counter fires, or a graduated `check` matches),
|
||
the throttle resets to zero immediately — the system is quick to re-engage
|
||
and only slow to relax, so a real emerging problem is never delayed by a
|
||
prior quiet streak. This is the direct token-cost counterpart to escalation
|
||
(escalation raises alarm on repeated *violations*; this lowers cost on
|
||
repeated *silence*), and both read the same underlying review-history data,
|
||
just in opposite directions.
|
||
|
||
**Quality trend metrics on the usage dashboard**: alongside the `self-
|
||
learning` token-cost line item already planned, the usage dashboard (Phase
|
||
14) accumulates and charts three time series purely from data Phase 7's other
|
||
features already produce (no new instrumentation): build/test pass-rate from
|
||
outcome-linked verification runs, escalation-note count from the
|
||
repeated-violations mechanism, and per-lesson trigger count from the same
|
||
counter escalation already reads. A lesson whose trigger count is high but
|
||
whose citations never precede a drop in the escalation rate (i.e., it keeps
|
||
firing without the underlying pattern ever actually going away) is flagged
|
||
as a `forget` candidate in `/learning` — a lesson that keeps getting matched
|
||
but never changes behavior is more likely to be miscalibrated or ignored
|
||
than genuinely useful, so it's surfaced for human review rather than kept
|
||
accumulating silently.
|
||
|
||
**Structured lesson bodies**: a `lesson` memory's body gains an optional
|
||
`before`/`after` code-snippet pair alongside its existing prose description —
|
||
extracted directly from the specific ledger line / diff that triggered the
|
||
lesson (the reviewer already has this content in hand when it calls
|
||
`remember`; this is a shape change to what it passes, not new data
|
||
collection). Rendering (in `/learning`, in the collapsed verdict note, and
|
||
implicitly in whatever text gets injected into future system prompts) shows
|
||
the snippet pair when present, falling back to prose-only for lessons that
|
||
predate this field or for which no single before/after pair meaningfully
|
||
captures the point (e.g. a cross-cutting retrospective observation). This
|
||
makes a lesson concretely actionable — "change shape X to shape Y" — rather
|
||
than purely descriptive, without requiring any new tool or trigger.
|
||
|
||
**Lesson-pack export/import**: an opt-in `/lesson export <path>` command
|
||
serializes a project's (or, with a flag, the global store's) `lesson`
|
||
entries — full frontmatter including `confidence`/`scope`/`provenance`/
|
||
`check` state, structured bodies, everything — into one portable bundle file,
|
||
independent of and without requiring a shared `~/.zesdex/`. `/lesson import
|
||
<path>` on another machine (or by a teammate) merges a pack into the local
|
||
store using the exact same recall-first dedup and contradiction-detection
|
||
path new reviewer-authored lessons already go through (Phase 7's lifecycle
|
||
mechanism, reused verbatim — an imported pack is just another source of
|
||
candidate lessons, not a privileged bypass), so imported entries land in the
|
||
calibration queue exactly like any other non-`/lesson` addition unless they
|
||
came from a `/lesson`-authored source originally (in which case they keep
|
||
their `confidence: human` standing on import too, since that's a property of
|
||
the entry, not of the transport). Export/import is explicitly scoped to
|
||
`lesson`/`retrospective` memory types only — it does not touch session
|
||
transcripts, `edits.jsonl`, or anything else in the per-session directory.
|
||
|
||
## Patterns to Mirror
|
||
|
||
| Category | Source | Pattern |
|
||
|---|---|---|
|
||
| State mutation | `app/runtime/actions/*.rs`, `app/runtime/commands/*.rs` | All state mutation flows through `apply_action`/`apply_slash`; controller/view layers are read-only translators |
|
||
| Tool definition | `tool/mod.rs` `Tool` trait | Zero-size struct per tool, `name/description/parameters/run`, registered in one `all_tools()` vec, risk flag is a name-based match in one place |
|
||
| Required-argument enforcement | `tool/fs/*.rs` `arg_str` helper | Missing/blank required args fail deterministically before any I/O, via one small shared helper — the pattern zesdex's new `reason` field reuses verbatim |
|
||
| Async bridge | `app/runtime/stream/*.rs` | One `mpsc` channel per request; main loop stays synchronous; blocking tool I/O deferred to `std::thread`, not `.await`ed inline |
|
||
| Approval gating | `app/harness.rs`, `stream/tools/approval.rs` | Deterministic check → advisory background classifier → per-call intent-aware classifier, each layer independently timeout-bounded and explicit about fail-open vs fail-closed |
|
||
| Memory persistence | `model/memory.rs` | Index-of-pointers (`MEMORY.md`) + one frontmatter+body file per entry, atomic rename-over-write, hard-sanitized slugs, only the index (never full bodies) injected into the system prompt |
|
||
| IPC framing | `ipc/frame.rs` | 4-byte-BE length prefix + JSON, hard cap enforced before payload allocation |
|
||
| Daemon liveness | `ipc/server.rs`, `model/session_lock.rs` | `bind()`/`connect()` success is the liveness oracle, never a PID file; lock files store PID + `kill(pid,0)` liveness check |
|
||
| Config layout | `model/settings.rs`, `model/app_config.rs` | Global prefs in one small `config.json`; all per-session config in `settings.json` under a per-session directory |
|
||
| Sidecar protocol | `koma_sec_daemon/protocol.py` | Newline-delimited JSON frames, single shared secret handshake, single-threaded serialized dispatch, every error surfaces as a string result, never a crash |
|
||
|
||
## Files to Create (Phase 1 scope; later phases add modules incrementally)
|
||
|
||
| File | Action | Why |
|
||
|---|---|---|
|
||
| `Cargo.toml` | CREATE | Package metadata + dependency set per your spec, extracted from koma's `src-agent/Cargo.toml` |
|
||
| `src/main.rs` | CREATE | Entry point: terminal setup/teardown, tokio runtime construction, mode dispatch into the event loop |
|
||
| `src/app/mod.rs`, `src/app/state.rs`, `src/app/mode.rs` | CREATE | `AppState`, `Mode` enum (Chat + Onboard to start), core runtime state |
|
||
| `src/app/runtime/mod.rs`, `event_loop.rs`, `actions.rs`, `commands.rs`, `stream.rs` | CREATE | Synchronous main loop, `Action`/`Command` dispatch, tokio stream bridge |
|
||
| `src/app/harness.rs` | CREATE | WC/PC/TAC three-layer approval harness + catastrophic-op guard |
|
||
| `src/controller/input.rs`, `src/controller/command.rs` | CREATE | Key→Action translation, slash-command parsing |
|
||
| `src/view/mod.rs`, `view/chat.rs`, `view/theme.rs`, `view/markdown.rs`, `view/status.rs` | CREATE | Ratatui rendering: 3-region chat layout + status bar per your spec |
|
||
| `src/tool/mod.rs`, `tool/fs.rs`, `tool/search.rs`, `tool/shell.rs`, `tool/shell_filter/*.rs` | CREATE | `Tool` trait + filesystem/search/shell tool set with output filtering; `write`/`edit` schemas include required `reason` from day one |
|
||
| `src/tool/git_operator.rs`, `tool/git_worktree.rs`, `tool/git_cred.rs` | CREATE | Git tool set with destructive-op guard |
|
||
| `src/dto/chat.rs`, `dto/openrouter.rs` | CREATE | Wire types for the LLM API (OpenRouter-compatible) |
|
||
| `src/service/mod.rs`, `service/openrouter.rs` | CREATE | Streaming client, `StreamEvent`, classifier calls |
|
||
| `src/model/session.rs`, `model/settings.rs`, `model/store.rs`, `model/app_config.rs`, `model/conversation.rs` | CREATE | Session persistence, per-session settings, config |
|
||
| `src/resources.rs`, `src-misc/*.txt` | CREATE | Embedded system/personality/tool-usage/classifier prompts |
|
||
|
||
Later phases (see Tasks) add: `src/model/memory.rs` + `src/model/editlog.rs`
|
||
(reasoned-edit ledger + `lesson`-typed memory, Phase 7), `src/app/review/mod.rs`
|
||
(auto quality-review trigger, Phase 7; extended in-place by 7B/7C/7D/7E — no
|
||
new top-level module per sub-phase), `src/app/subagent/*` +
|
||
`src/app/workflow/{mod,script,engine}.rs` + `src/tool/workflow.rs` +
|
||
`src/view/workflow.rs` + `src/app/mode` workflow-panel overlay (dynamic workflow
|
||
orchestration, replaces koma's `task`, Phase 6), `src/app/bgbash/*` (background
|
||
jobs), `src/app/mcp/*` (MCP client), `src/service/oauth/*` (OAuth), `src/ipc/*`
|
||
(daemon/client split), `src/tool/internet/*` (web fetch/search/download),
|
||
`src/app/sec/*` + a `security-sidecar/` Python package (security toolkit),
|
||
`src/model/msglog.rs` (SQLite blob archive), `src/app/runtime/shortsend.rs`
|
||
(token-efficiency rail).
|
||
|
||
## Milestones, Dependency Graph, and Implementation Gates
|
||
|
||
This section turns the long phase list below into an executable roadmap. It does
|
||
not remove scope; it groups the work so implementation can proceed in stable,
|
||
reviewable increments without letting the advanced self-learning surface block
|
||
the basic product from becoming usable.
|
||
|
||
### Milestone groups
|
||
|
||
| Milestone | Phases | Outcome | Can stop here? |
|
||
|---|---|---|---|
|
||
| **M0 — Plan freeze** | Current document | Scope, risks, dependencies, and acceptance are frozen before code starts | Yes — no code written yet |
|
||
| **M1 — Native TUI shell + agent loop** | 1 → 2 → 3 | A usable single-process TUI that can stream model output, run core tools, auto-approve routine work, and still refuse catastrophic operations | Yes — first functional developer preview |
|
||
| **M2 — Durable sessions + context economy** | 4 → 5 | Sessions survive restarts, `/resume` works, long conversations stay usable via SQLite archive + short-send shaping | Yes — usable daily-driver baseline without subagents |
|
||
| **M3 — Dynamic workflow orchestration** | 6 | `workflow_run` replaces koma's `task`, supports many subagents, progress panel, `parallel`/`pipeline`, and run-scoped `note_finding` | Yes — first multi-agent release |
|
||
| **M4 — Self-learning core and growth path** | 7 → 7B → 7C → 7D → 7E | Required edit reasons, audit ledger, background reviewer, lesson lifecycle, escalation, enforced checks, retrospectives, import/export, dashboard data | Yes after any sub-phase; 7 alone is useful |
|
||
| **M5 — Background jobs + extensibility + auth** | 8 → 9 → 10 | Long-running shell jobs, MCP tools, and OAuth login are integrated without blocking the TUI | Yes — plugin/provider-ready release |
|
||
| **M6 — Detachable multi-session runtime** | 11 | Daemon/client split, attach/detach, snapshot/diff/shadow-state rendering | Yes — multi-session release |
|
||
| **M7 — External capability packs** | 12 → 13 | Authorized security toolkit and full internet/search/fetch/download stack land behind their respective opt-in controls | Yes — full agent capability release |
|
||
| **M8 — Full UI parity polish** | 14 | Settings, agents editor, MCP browser, usage dashboard, `/learning`, help, rewind, quit confirm, session hub | Yes — full koma-parity UI release |
|
||
|
||
### Dependency graph
|
||
|
||
```text
|
||
M0 plan freeze
|
||
↓
|
||
Phase 1 TUI skeleton
|
||
↓
|
||
Phase 2 streaming agent loop + core tools
|
||
↓
|
||
Phase 3 approval harness + catastrophic-op guard
|
||
↓
|
||
Phase 4 session persistence
|
||
↓
|
||
Phase 5 SQLite archive + short-send
|
||
↓
|
||
Phase 6 workflow orchestration
|
||
├─ requires Phase 2 tool loop
|
||
├─ requires Phase 3 non-interactive risky-tool gating for subagents
|
||
└─ provides subagent engine used by Phase 7+
|
||
↓
|
||
Phase 7 self-learning MVP
|
||
├─ requires Phase 6 detached subagent execution
|
||
├─ requires Phase 4 session directory for edits.jsonl
|
||
└─ provides lesson memory + review side-channel
|
||
↓
|
||
Phase 7B outcome/lifecycle/calibration/global tier
|
||
└─ requires Phase 7 lesson memory + review trigger
|
||
↓
|
||
Phase 7C escalation/adaptive/batch mining
|
||
└─ requires Phase 7B confidence/scope/lifecycle fields
|
||
↓
|
||
Phase 7D graduated checks + shadow mode
|
||
├─ requires Phase 7B lesson frontmatter fields
|
||
└─ benefits from Phase 7C repeated-violation counter
|
||
↓
|
||
Phase 7E retrospectives/provenance/user lessons/export/trends
|
||
├─ requires Phase 7B confidence/scope model
|
||
├─ requires Phase 7C counters for trends
|
||
├─ requires Phase 7D checks for graduation UI
|
||
└─ rendered fully in Phase 14 dashboard
|
||
|
||
Independent after M3/M4 baseline:
|
||
Phase 8 background bash jobs
|
||
Phase 9 MCP client
|
||
Phase 10 OAuth login
|
||
|
||
Late structural split:
|
||
Phase 11 daemon/client IPC
|
||
├─ requires stable Phase 1 renderer
|
||
├─ requires stable Phase 4 session model
|
||
└─ should happen after Phase 8 so background nudges are daemon-aware
|
||
|
||
External packs:
|
||
Phase 12 security toolkit
|
||
├─ requires Phase 3 catastrophic guard
|
||
├─ requires Phase 4 settings/session storage
|
||
└─ benefits from Phase 11 daemon lifecycle, but can be built single-process first
|
||
Phase 13 internet module
|
||
├─ requires Phase 2 tool loop
|
||
├─ requires Phase 3 web_download guard adjacency
|
||
└─ benefits from Phase 8 background job patterns for sidecar timeout handling
|
||
|
||
Final UI parity:
|
||
Phase 14 remaining UI modes
|
||
├─ depends on the modes/features it surfaces
|
||
└─ can be partially implemented earlier for `/learning` if Phase 7E needs it
|
||
```
|
||
|
||
### Implementation gates
|
||
|
||
Each gate must pass before the next milestone starts. This is stricter than
|
||
"cargo build passes"; it prevents hidden partials from accumulating.
|
||
|
||
| Gate | Required before | Criteria |
|
||
|---|---|---|
|
||
| **G1 — TUI correctness gate** | Start Phase 2 | Phase 1 renders the exact 3-region layout, handles resize, input, scroll basics, and restores terminal state after panic/quit |
|
||
| **G2 — Tool-loop correctness gate** | Start Phase 3 | Phase 2 can complete a full mocked tool-use loop, including tool-call parse, tool execution, result append, and second model call |
|
||
| **G3 — Safety/speed gate** | Start Phase 4 | Auto/Yolo routine mutations have zero prompts; catastrophic-op guard blocks every named rule even in Yolo; no classifier failure can bypass deterministic guards |
|
||
| **G4 — Persistence gate** | Start Phase 5/6 | Sessions create/list/resume cleanly, stale locks are cleared, and no session file write can corrupt partial state on crash |
|
||
| **G5 — Context gate** | Start Phase 6 | Long conversation fixture proves SQLite archive and short-send shaping reduce wire payload without destroying transcript fidelity |
|
||
| **G6 — Workflow gate** | Start Phase 7 | `workflow_run` is the only delegation surface, `parallel` cap and `pipeline` no-barrier semantics are tested, progress panel survives dismiss/resummon, and `note_finding` is ephemeral/run-scoped |
|
||
| **G7 — Self-learning MVP gate** | Start 7B | `write`/`edit` require `reason`, ledger is append-only and exact, one detached reviewer runs per edited turn, and verdicts never enter model-visible history |
|
||
| **G8 — Lesson lifecycle gate** | Start 7C | Confidence, stale, scope, calibration, contradiction blocking, and global/project memory injection all have direct tests |
|
||
| **G9 — Escalation/adaptive gate** | Start 7D | Repeated-violation escalation and adaptive review throttling both work under synthetic histories and never run reviewers concurrently |
|
||
| **G10 — Graduated-check gate** | Start 7E | Shadow-mode trials, graduated checks, override behavior, and de-graduation are tested without making any check an unbreakable blocker |
|
||
| **G11 — Learning UI data gate** | Start Phase 14 `/learning` UI | Provenance, retrospectives, import/export, trend metrics, and self-learning token-cost records are all present as model/data APIs before UI work begins |
|
||
| **G12 — Daemon split gate** | Start Phase 11 | Single-process runtime is stable through Phases 1–10; renderer state can be snapshotted without borrowing live TUI objects |
|
||
| **G13 — External capability gate** | Start Phase 12/13 | Tool gating, settings, background subprocess supervision, and side-channel notes are stable enough that security/internet sidecars cannot destabilize core chat |
|
||
| **G14 — Full parity gate** | Finish Phase 14 | Every mode has an Esc/back path, every dashboard reflects live data, and every feature-specific acceptance item below remains green |
|
||
|
||
### Recommended implementation order
|
||
|
||
For the actual coding pass, the safest order is:
|
||
|
||
1. **M1**: Phases 1–3, because they produce the first runnable agent shell and the
|
||
always-on catastrophic guard.
|
||
2. **M2**: Phase 4, then Phase 5 only if long-session support is needed before
|
||
subagents; otherwise Phase 5 can be done immediately after Phase 6.
|
||
3. **M3**: Phase 6, because every self-learning feature depends on subagent
|
||
execution and workflow orchestration.
|
||
4. **M4**: Phase 7 only, stop, test it hard, then decide whether to continue with
|
||
7B–7E immediately or ship the MVP learning loop first.
|
||
5. **M5/M6/M7/M8**: Phases 8–14 in order, except `/learning` UI slices from Phase
|
||
14 may be pulled forward after 7E if needed for manual calibration.
|
||
|
||
The important policy: **do not start Phase 7B before Phase 7 is stable, and do not
|
||
start Phase 11 before the single-process app is already boringly reliable**.
|
||
|
||
## Tasks
|
||
|
||
### Phase 1 — Skeleton: Cargo.toml, event loop, 3-panel TUI, status bar
|
||
- **Action**: Stand up the buildable skeleton exactly matching your explicit
|
||
layout spec: Main Execution Log (transcript), Input Prompt (bottom fixed-height
|
||
block), Status Bar (`[zesdex] | STATUS: AUTO-APPROVE (UNRESTRICTED) | PROTOCOL:
|
||
ZERO-STUBS ACTIVE`). Wire the synchronous main loop + tokio runtime bridge with
|
||
no LLM calls yet (a `pong`-style echo tool only), so the rendering/input/event
|
||
loop is provably correct before the agentic loop is layered on.
|
||
- **Mirror**: `app/runtime/event_loop/mod.rs` tick structure (drain stream events
|
||
→ drain harness verdicts → input poll with adaptive 8ms/100ms timeout → draw if
|
||
dirty); `view/chat/mod.rs` layout_chunks region split.
|
||
- **Validate**: `cargo build`, `cargo clippy -- -D warnings`, manual TUI smoke run
|
||
(type a line, see it echo into the log region, resize terminal, quit cleanly
|
||
restoring the terminal state).
|
||
|
||
### Phase 2 — OpenRouter-compatible client + streaming agentic loop
|
||
- **Action**: Implement `dto::chat`/`dto::openrouter` wire types, the streaming
|
||
SSE client, `StreamEvent` bridge, and `advance_turn`/`process_tools` loop
|
||
against a real filesystem/shell tool set (`read, grep, glob, write, edit,
|
||
delete, bash, dir_list, cd`). `write`/`edit` already require `reason` per the
|
||
Files-to-Create note above — enforced now, even though the review loop that
|
||
reads the ledger doesn't land until Phase 7, so the ledger has real data by
|
||
the time Phase 7 needs it. No approval gating yet — this phase proves the
|
||
model can drive tools end-to-end.
|
||
- **Mirror**: `service/openrouter.rs` streaming/classify pattern; `tool/mod.rs`
|
||
`Tool` trait and `all_tools()` registration; `MAX_AGENT_STEPS` loop cap.
|
||
- **Validate**: Integration test against a mock SSE server; a `write` call
|
||
without `reason` is rejected before any file is created; manual run against a
|
||
real OpenRouter-compatible endpoint editing a scratch file end-to-end.
|
||
|
||
### Phase 3 — Approval harness: Auto/Normal/Plan/Yolo + catastrophic-op guard
|
||
- **Action**: Implement WC (deterministic workspace containment), TAC (per-risky-
|
||
call classifier, intent-aware, mode-dependent fail-open/fail-closed), and the
|
||
`AgentMode` cycle (Auto/Normal/Plan/Yolo). Implement the **catastrophic-op
|
||
guard** as a small, explicit, always-on function independent of mode/classifier
|
||
state — modeled directly on koma's `git_operator::check_destructive` list
|
||
(force-push, `reset --hard`, `clean -f/-d/-x`, force checkout, recursive delete
|
||
outside any configured workspace root, credential-file read/exfil patterns) —
|
||
and extend it to `bash`/`delete`/`web_download` per your "keep a hard stop on
|
||
catastrophic ops" decision. Skip PC (advisory prompt classifier) as lower value
|
||
for v1; note it in follow-up.
|
||
- **Mirror**: `app/harness.rs` WC/TAC design, `git_operator.rs::check_destructive`.
|
||
- **Validate**: Unit tests per guard rule (force-push blocked without confirm,
|
||
`reset --hard` blocked, ordinary `git commit`/`write`/`edit` in Auto mode runs
|
||
with zero prompts and zero added latency vs. Phase 2 baseline).
|
||
|
||
### Phase 4 — Session model + on-disk persistence + `/resume`
|
||
- **Action**: `Session`/`Settings`/`AppConfig`/`store` (create/list/rename
|
||
sessions), per-session directory layout (`settings.json`, `messages.json`,
|
||
`session.lock`, `memory/`, `edits.jsonl`), session picker mode.
|
||
- **Mirror**: `model/session.rs`, `model/store.rs`, `model/session_lock.rs`
|
||
(PID-liveness lock, never a stale-file oracle).
|
||
- **Validate**: Kill `-9` a session mid-write, confirm lock is detected stale and
|
||
cleared on next list; `/resume` round-trips a full conversation.
|
||
|
||
### Phase 5 — SQLite blob archive + short-send token efficiency (optional-early)
|
||
- **Action**: `messages.sqlite` archive (append-only messages/blobs/summary
|
||
tables), non-destructive wire-payload compaction rail (`shortsend::shape`) so
|
||
budget models stay in-context on long sessions.
|
||
- **Mirror**: `model/msglog.rs`, `app/runtime/shortsend.rs` engage-hysteresis math.
|
||
- **Validate**: Long synthetic conversation exceeding context window compacts the
|
||
wire payload while `messages.json`/transcript stay uncompressed.
|
||
|
||
### Phase 6 — Dynamic workflow orchestration engine (replaces koma's `task`)
|
||
- **Action**: Build the subagent execution core first (isolated `Conversation`
|
||
seed, restricted tool allow-list, non-interactive engine loop with fail-closed
|
||
risky-tool gating — this part *is* still a direct mirror of koma's subagent
|
||
engine design, just not exposed as a standalone `task` tool). Layer the
|
||
`workflow_run(script, args)` tool on top with the four primitives
|
||
(`agent`/`parallel`/`pipeline`/`phase`/`log`), a bounded concurrency pool
|
||
(default 5, configurable), and the Workflow Progress panel (new `Mode`
|
||
overlay reusing `layout_chunks`, live phase→agent→status tree, dismiss/
|
||
re-summon without killing the run). Update the system prompt (`src-misc/
|
||
system-tools.txt` equivalent) to describe `workflow_run` and give the model
|
||
concrete guidance on when decomposition is worth it vs. handling inline.
|
||
Agent-definition format (frontmatter+markdown, tiered Builtin→Global→Session
|
||
merge) is still implemented so `agent()` calls can reference named personas,
|
||
and so Phase 7's built-in `quality-reviewer` persona has somewhere to live.
|
||
- **Mirror**: `app/subagent/{context,engine,event,spawn}.rs`, `model/
|
||
agent_def/*.rs` for the execution core; `view/chat/subagents.rs` +
|
||
`app/mode/mod.rs` overlay-mode pattern for the progress panel's UI shell (no
|
||
koma source for the orchestration script layer itself — that's new, specified
|
||
above under "Dynamic Workflow Orchestration").
|
||
Also add `note_finding(text)` — a restricted tool available only inside a
|
||
`workflow_run` execution context (not part of a standalone subagent's
|
||
allow-list) that pushes `text` onto a run-scoped, in-memory, non-persisted
|
||
findings broadcast; every other `agent()` thunk still active in the same run
|
||
receives the current findings list at its next tool-round boundary. This is
|
||
workflow-execution coordination, not lesson persistence, which is why it
|
||
lives here rather than in Phase 7 — a finding only becomes a durable
|
||
`lesson` if the normal Phase 7 reviewer independently decides to `remember`
|
||
it after the run, reading the same `edits.jsonl` ledger as any other turn.
|
||
- **Mirror**: `app/subagent/{context,engine,event,spawn}.rs`, `model/
|
||
agent_def/*.rs` for the execution core; `view/chat/subagents.rs` +
|
||
`app/mode/mod.rs` overlay-mode pattern for the progress panel's UI shell (no
|
||
koma source for the orchestration script layer itself, nor for
|
||
`note_finding` — both are new, specified above under "Dynamic Workflow
|
||
Orchestration").
|
||
- **Validate**: A `parallel()` workflow with 6 agent thunks confirms the 6th
|
||
queues on the concurrency cap; a `pipeline()` workflow with 3 items × 2 stages
|
||
confirms item A reaches stage 2 while item B is still on stage 1 (no barrier);
|
||
progress panel updates live during a real multi-agent run and correctly
|
||
survives Esc-dismiss-and-resummon without losing state; a trivial single-step
|
||
request in a manual smoke test is handled inline by the model without
|
||
invoking `workflow_run` at all (confirms the "dynamic, not forced" trigger
|
||
behavior is actually working, not just documented); inside a `workflow_run`
|
||
with two concurrent `agent()` thunks, a `note_finding` call from one is
|
||
observably visible to the other's next tool round, and is confirmed absent
|
||
from `MEMORY.md` after the run (ephemeral, not persisted).
|
||
|
||
### Phase 7 — Reasoned mutations + self-learning quality loop (MVP core)
|
||
|
||
Phase 7 was originally one block covering every self-learning idea explored
|
||
in planning, but that made it far larger and more tightly coupled than every
|
||
other phase in this plan — a violation of the plan's own "each phase is
|
||
independently buildable/testable" principle. It is now split into this core
|
||
phase plus four additive sub-phases (**7B**–**7E**, deliberately lettered
|
||
rather than renumbered so Phases 8–14 below are untouched). **Phase 7 alone
|
||
is the complete, working answer to your original request** — required
|
||
`reason`, an audit ledger, and an automatic background quality pass; 7B–7E
|
||
are refinements layered on top, each independently deferrable if you want a
|
||
smaller first slice of the self-learning feature specifically. Anywhere the
|
||
Risks/Acceptance tables below say "Phase 7," read it as "the 7/7B/7C/7D/7E
|
||
family" unless a sub-phase is called out explicitly.
|
||
- **Action**: Add the required `reason` argument to `write`/`edit` (schema +
|
||
deterministic rejection before any filesystem call, before any I/O); implement
|
||
the append-only `edits.jsonl` ledger with an `origin: main | subagent |
|
||
reviewer` tag on every line; add a `lesson` memory kind to `model::memory`
|
||
(`name/description/type` only for now — the richer frontmatter fields land
|
||
in 7B–7E) plus a recall-before-remember dedup convention so a duplicate
|
||
observation is never saved twice; add the built-in `quality-reviewer`
|
||
agent-definition persona (read-only + `recall`/`remember` allow-list, no
|
||
`write`/`edit`/`bash`, so a review pass can only ever produce a memory note
|
||
or nothing); wire the auto-trigger into the event loop's existing
|
||
turn-completion point (the same point `Done` already fires) so a non-empty
|
||
ledger slice for the just-finished turn spawns exactly one detached-nudge
|
||
subagent (Phase 6's engine, detached-nudge mode), capacity-1 queued if
|
||
another review is already running, skipped entirely in Plan mode and for
|
||
ledger lines whose `origin` is `subagent`/`reviewer` (no recursive
|
||
review-the-reviewer chains); add the `SystemNote` side-channel (parallel to
|
||
`harness_rx`) so the reviewer's verdict renders as a collapsed note in the
|
||
Main Execution Log without ever entering the model-visible transcript.
|
||
This is the complete MVP: reasoned mutations, an audit trail, and an
|
||
automatic background quality pass that folds observations into memory the
|
||
existing `load_memory_index` mechanism already injects every future turn.
|
||
- **Mirror**: `tool/fs/write.rs`/`edit.rs` (extend schema, reuse `arg_str`),
|
||
`model/memory.rs` (add `kind`, reuse `write_memory`/`read_memory`/
|
||
`slugify`/`atomic_write`), `app/subagent/spawn.rs` detached-nudge fold-back
|
||
(built in Phase 6), `app/runtime/event_loop/mod.rs` turn-completion hook,
|
||
`harness_rx`'s dedicated-channel pattern for the new `SystemNote` channel.
|
||
- **Validate**: A `write`/`edit` call missing `reason` is rejected before any
|
||
file is created/modified (unit test, both tools); a turn with 3 edits
|
||
produces exactly 3 ledger lines with matching reasons and `origin: main`; a
|
||
workflow subagent's own edits are tagged `origin: subagent` and do NOT
|
||
double-trigger a nested review; after a turn with edits, exactly one
|
||
reviewer subagent runs detached and the Main Execution Log gets exactly one
|
||
collapsed verdict note with zero additional turns appended to the
|
||
model-visible conversation; a deliberately duplicate lesson (same point,
|
||
reworded) is not re-saved because the reviewer's `recall`-first check finds
|
||
the existing entry; Plan mode never triggers a reviewer even after a
|
||
hypothetical mutation (defence in depth).
|
||
|
||
### Phase 7B — Outcome-linked learning, lesson lifecycle, calibration, global tier
|
||
|
||
Builds directly on Phase 7's `lesson` kind and reviewer trigger; independently
|
||
deferrable from 7C/7D/7E. See "Outcome-linked learning," "Lesson lifecycle,"
|
||
"Human-in-the-loop calibration," and "Global cross-project lesson tier" in the
|
||
digest above for full mechanism detail — summarized here as build items.
|
||
- **Action**: Extend `lesson` frontmatter with `confidence: verified |
|
||
opinion`, `last_confirmed`, `stale: bool`, and `scope: project | global`;
|
||
implement the language-agnostic build/test probe table (settings override →
|
||
ordered project-marker detection → skip if none match) as one named,
|
||
independently-extensible function, run on a plain thread and handed to the
|
||
reviewer alongside the ledger; implement the `pending_lessons.json`
|
||
calibration queue (auto-resolve-to-keep after a grace window in Auto/Yolo,
|
||
explicit keypress in Normal/Plan) with contradiction detection blocking
|
||
auto-save; implement the idle-triggered staleness sweep that flags (never
|
||
deletes) old `lesson` entries; implement the `~/.zesdex/memory/` global
|
||
store and extend `load_memory_index` to concatenate global-then-project.
|
||
- **Mirror**: `shell_filter`'s per-command-special-case pattern generalized
|
||
into the build/test probe table; `model/memory.rs` frontmatter extension;
|
||
`app/bgbash/*`'s idle-detection reused for the staleness sweep; `app/mode/*`
|
||
Builtin→Global→Session tiering pattern reused for the `project`/`global`
|
||
lesson split.
|
||
- **Validate**: A lesson born from a reproduced build/test failure is tagged
|
||
`confidence: verified` and one born from reviewer judgement alone is tagged
|
||
`opinion`, visibly distinguished; a deliberately contradictory lesson is
|
||
blocked from auto-save and surfaces both entries in the calibration toast;
|
||
in Auto mode a pending lesson auto-resolves to `keep` after the grace window
|
||
and remains `forget`-able afterward, in Normal mode it waits for an explicit
|
||
keypress; an artificially back-dated `last_confirmed` gets flagged
|
||
`stale: true` without deletion; a `scope: project` lesson does not leak into
|
||
a second project's injected index, while `scope: global` appears in both;
|
||
the build/test probe resolves correctly for at least three distinct
|
||
project-ecosystem fixtures with no ecosystem hard-coded outside the one
|
||
probe table.
|
||
|
||
### Phase 7C — Escalation, adaptive review frequency, batch pattern mining
|
||
|
||
Builds on 7B's lesson/confidence fields; independently deferrable. See
|
||
"Escalation on repeated violations," "Adaptive review frequency," and
|
||
"Batch/offline pattern mining" in the digest above.
|
||
- **Action**: Implement the repeated-violation counter that escalates to a
|
||
non-collapsed attention note (distinct styling from the routine collapsed
|
||
`SystemNote`) on the third repeat of a lesson within a configurable window,
|
||
and can promote a lesson from `project` to `global` scope when the repeat
|
||
spans multiple project directories; implement the consecutive-empty-pass
|
||
counter that throttles the per-turn review trigger's frequency and resets
|
||
to zero on any new lesson/escalation; implement the idle-triggered batch
|
||
pattern-mining pass over full `edits.jsonl` + `git log -p` history,
|
||
rate-limited to once per idle window and sharing the same capacity-1 review
|
||
queue as the per-turn reviewer.
|
||
- **Mirror**: `app/bgbash/*`'s idle trigger as the template for both the
|
||
batch-mining pass and the throttle counter's reset condition; ledger
|
||
`origin` tagging (Phase 7) reused as the cheap equality check for
|
||
repeat-detection scope.
|
||
- **Validate**: The same lesson pattern violated a third time within the
|
||
window produces a non-collapsed escalation note distinct from routine
|
||
verdicts; a synthetic run of N consecutive empty review passes measurably
|
||
reduces review frequency, and a single subsequent lesson-producing pass
|
||
resets it to reviewing every turn; batch mining never runs concurrently
|
||
with a per-turn review (shared capacity-1 queue verified under load).
|
||
|
||
### Phase 7D — Lesson graduation into enforced checks (shadow-mode gated)
|
||
|
||
Builds on 7B/7C; independently deferrable — a project can run indefinitely on
|
||
`lesson`-as-injected-text alone without ever enabling this sub-phase. See
|
||
"Lesson graduation into a lint check" and "Shadow-mode enforcement" in the
|
||
digest above.
|
||
- **Action**: Add the `check` frontmatter field (a small structured rule:
|
||
substring/regex against new content, or a path-glob + required-companion-
|
||
file rule) and its enforcement inside `write`/`edit` (same deterministic-
|
||
before-harness slot as `reason`) for graduated lessons — a failed check
|
||
never blocks the write, it returns a naming, addressable/overridable
|
||
result; gate graduation behind a mandatory shadow-mode trial window
|
||
(`shadow: bool` + `trial_count` frontmatter, same rule-evaluation path,
|
||
would-have-fired counting only, nothing surfaced to the model) with a
|
||
graduate-vs-demote decision at window close based on fire ratio; restrict
|
||
manual graduation/de-graduation to `/learning` (Phase 14) or the two-
|
||
reviewer consensus mechanism (Phase 7E) — never reviewer self-promotion.
|
||
- **Mirror**: `tool/fs/write.rs`/`edit.rs`'s `reason`-enforcement slot,
|
||
reused verbatim for `check` enforcement.
|
||
- **Validate**: A `check` candidate spends its full trial window in
|
||
`shadow: true` before ever surfacing anything to the model, and correctly
|
||
graduates or demotes at window close based on its fire ratio; a graduated
|
||
lesson's `check` fires on a violating `write`/`edit` as a named,
|
||
addressable result — never a silent or unbreakable block.
|
||
|
||
### Phase 7E — Retrospectives, provenance, user-authored lessons, consensus promotion, export/import, trend metrics
|
||
|
||
Builds on 7B (confidence/scope) and 7D (consensus gates graduation
|
||
promotion); independently deferrable — the MVP loop (Phase 7) and its
|
||
lifecycle/escalation refinements (7B/7C) are fully useful without this. See
|
||
"Session-end retrospective + per-lesson provenance," "User-authored lessons,
|
||
consensus-gated global promotion, token-budget visibility," "Quality trend
|
||
metrics on the usage dashboard," "Structured lesson bodies," and
|
||
"Lesson-pack export/import" in the digest above.
|
||
- **Action**: Add a `provenance` field to every `lesson` (session id + turn
|
||
index + ledger line hash, or `"user"` for `/lesson`-authored entries) and
|
||
the session-end retrospective trigger (reuses the daemon idle/tombstone
|
||
path, Phase 11, and the same reviewer-engine shape at session scope)
|
||
writing one `type: retrospective` memory per session, never injected into
|
||
the system prompt; add `/lesson <text>` and `/lesson import <path>` slash
|
||
commands (`confidence: human` for direct authorship, exempt from
|
||
calibration/contradiction; `opinion` for bulk-extracted imports, which DO
|
||
go through calibration); tighten automatic (non-`/lesson`) `global`
|
||
promotion to require two independently-spawned reviewer subagents to agree
|
||
before the write; tag every self-learning subagent's token spend with
|
||
`origin: self-learning` and surface it as its own usage-dashboard line item
|
||
(Phase 14), plus the three trend time series (build/test pass-rate,
|
||
escalation count, per-lesson trigger count) and the trigger-without-
|
||
improvement `forget`-candidate flag, all derived from data 7A–7D already
|
||
produce; extend a `lesson`'s body with an optional structured `before`/
|
||
`after` snippet pulled from the triggering ledger line/diff; add `/lesson
|
||
export <path>` and `/lesson import <path>` (project or `--global` scope),
|
||
the latter routed through the exact same recall-first dedup/contradiction/
|
||
calibration path as any other new lesson, restricted to `lesson`/
|
||
`retrospective` memory types only.
|
||
- **Mirror**: `app/subagent/spawn.rs` detached-nudge fold-back (Phase 6,
|
||
reused unchanged for retrospective and two-reviewer consensus spawns);
|
||
`controller/command.rs` slash-command parsing for `/lesson`; koma's
|
||
existing usage-dashboard view, extended for the new cost category and
|
||
trend charts.
|
||
- **Validate**: Exactly one `type: retrospective` memory is produced per
|
||
ended session and is absent from the next session's injected system-prompt
|
||
index; `/lesson "text"` creates a `confidence: human` entry immediately
|
||
with no calibration toast; `/lesson import` seeds `opinion`-confidence
|
||
entries that DO go through calibration; automatic global promotion is
|
||
blocked when only one of two spawned reviewers agrees, and proceeds when
|
||
both do; the usage dashboard shows a nonzero `self-learning`-tagged token
|
||
total and correctly plotted trend charts against a synthetic session
|
||
history, with a high-trigger/no-improvement lesson flagged as a `forget`
|
||
candidate; a lesson with a structured `before`/`after` pair renders it in
|
||
both `/learning` and its injected system-prompt text; `/lesson export` then
|
||
`/lesson import` on a second (empty) project reproduces the same lesson set
|
||
with matching `confidence`/`scope`/`provenance`, and a near-duplicate
|
||
import is caught by the existing dedup path.
|
||
|
||
### Phase 8 — Background bash jobs
|
||
- **Action**: `run_in_background` bash execution on a plain OS thread, job
|
||
registry, `bash_output`/`bash_kill`, idle-triggered completion nudge.
|
||
- **Mirror**: `app/bgbash/mod.rs` thread/reader-thread/status-enum design.
|
||
- **Validate**: Long-running background command (`sleep 30 && echo done`) polls
|
||
correctly mid-run and nudges the model on completion without blocking the UI.
|
||
|
||
### Phase 9 — MCP client
|
||
- **Action**: `rmcp`-based client, stdio + streamable-HTTP transports, tool
|
||
namespacing/dedup, `/mcp` settings UI, config persisted in global config.
|
||
- **Mirror**: `app/mcp/mod.rs` connect/dispatch/sync-async-bridge design.
|
||
- **Validate**: Connect to a local test MCP stdio server, confirm its tools appear
|
||
namespaced in the model's advertised tool list and dispatch correctly.
|
||
|
||
### Phase 10 — OAuth login
|
||
- **Action**: PKCE authorization-code flow with a loopback callback listener,
|
||
token cache with single-flight refresh, `/settings` OAuth submenu.
|
||
- **Mirror**: `service/oauth/{pkce,loopback,manager}.rs` design.
|
||
- **Validate**: Full browser round-trip against a real OAuth-compatible provider
|
||
in a sandboxed test app registration.
|
||
|
||
### Phase 11 — Daemon/client IPC split (detachable multi-session)
|
||
- **Action**: Unix-socket, length-prefixed JSON-frame IPC; daemon-per-session
|
||
process model; snapshot/diff/shadow-state client rendering reusing the Phase 1
|
||
view layer unmodified; CLI daemon subcommands (`status/kill/restart/clean`).
|
||
- **Mirror**: `ipc/{frame,server,client,conn}.rs`, `ipc/snapshot/*.rs`,
|
||
`app/runtime/client/*.rs`, `app/runtime/event_loop/daemon/*.rs`.
|
||
- **Validate**: Detach a running session, confirm the daemon keeps streaming in
|
||
the background; reattach and confirm the shadow state matches exactly (no
|
||
full-vs-delta divergence) after a burst of token/tool-call activity.
|
||
|
||
### Phase 12 — Security toolkit (opt-in, authorized-use only)
|
||
- **Action**: `security-sidecar/` Python package mirroring `koma_sec_daemon`
|
||
(newline-JSON stdin/stdout protocol, tool registry, sandbox timeouts, tiered
|
||
installer, health probe). Rust-side `SecDaemonManager`, `/security` panel
|
||
requiring explicit enable + an on-screen authorized-use acknowledgment before
|
||
the panel arms, `sec_*` tools harness-exempt per koma's own design (explicit
|
||
user-invoked, not silent). Tool roster ported 1:1 from the reference inventory
|
||
(http, sqlmap, nuclei, ffuf, dalfox, zap, xss-confirm, z3, sage, rsa, factordb,
|
||
lattice, hashcat, hashid, decode, jsdeobf, unmin, sourcemap, wasm, triage, rop,
|
||
pwntmpl, remote-session).
|
||
- **Mirror**: `src-security/koma_sec_daemon/*` end to end.
|
||
- **Validate**: Health probe correctly reports install status per tool; a
|
||
representative tool from each category (web/crypto/re/pwn) round-trips a real
|
||
call against a local authorized test target (e.g., a deliberately vulnerable
|
||
local CTF binary/webapp, not a live third-party system).
|
||
|
||
### Phase 13 — Internet module (search/fetch/download + full browser mode)
|
||
- **Action**: Lightweight in-process `web_fetch`/`web_search` (reqwest + html2md,
|
||
no Python dependency) plus opt-in `internet-sidecar/` Python package
|
||
(Playwright-Firefox scraping) for pages that resist the lightweight path,
|
||
invoked as a stateless one-shot subprocess per call with a hard timeout and
|
||
fallback to the lightweight path on any failure. `web_download` carries the
|
||
500 MiB cap and is in the catastrophic-op-guard-adjacent risky set (still no
|
||
approval prompt in Auto/Yolo, but still subject to the destructive-target guard
|
||
if it targets a path outside the workspace).
|
||
- **Mirror**: `tool/internet/*.rs`, `src-internet/scrapion_agent/*`.
|
||
- **Validate**: Fetch a JS-heavy page that fails the lightweight path, confirm
|
||
automatic fallback to the browser sidecar succeeds and returns clean markdown.
|
||
|
||
### Phase 14 — Remaining UI surface (settings, agents editor, MCP browse, usage
|
||
dashboard, help, effort picker, message rewind, quit confirm, session hub,
|
||
`/learning` dashboard)
|
||
- **Action**: Fill out the remaining `Mode` variants and their views/controllers
|
||
so the app matches koma's full UI surface, not just chat. Add the new
|
||
`/learning` dashboard mode (no koma equivalent — new for zesdex's Phase 7
|
||
feature set): lists every `lesson` memory (project + global, clearly
|
||
separated), sortable by trigger count / `confidence` / `stale` flag / recency,
|
||
with inline `forget`, manual `project → global` promotion, and manual
|
||
lesson-graduation/de-graduation actions reusing the same primitives Phase 7
|
||
built; a per-lesson detail view resolves its `provenance` back to the
|
||
originating session/turn and renders its structured `before`/`after`
|
||
snippet when present; a shadow-mode lesson's row shows its live
|
||
would-have-fired counter and trial-window progress; a separate list browses
|
||
`type: retrospective` memories by session; `/lesson export`/`import` are
|
||
also reachable as dashboard actions (not just slash commands), for
|
||
discoverability; the existing usage dashboard mode gains a `self-learning`-
|
||
tagged line item alongside the main agent's own token spend plus the three
|
||
quality-trend charts (build/test pass-rate, escalation count, per-lesson
|
||
trigger count) and the trigger-without-improvement `forget`-candidate
|
||
flagging (Phase 7 already tags/produces all of this data — this phase only
|
||
needs to render it).
|
||
- **Mirror**: `app/mode/*.rs`, `view/*.rs` per-mode files (`/learning` has no
|
||
koma source — built directly on Phase 7's `model::memory` extensions and the
|
||
existing `session_hub`-style list-mode view pattern for the layout shell);
|
||
koma's existing usage-dashboard view (extended, not replaced, for the new
|
||
cost category and trend charts).
|
||
- **Validate**: Manual pass through every mode transition documented in the
|
||
reference's mode state machine; no dead-end mode with no Esc/back path;
|
||
`/learning` correctly reflects a lesson's live `confidence`/`stale`/`scope`
|
||
state and a `forget` performed from the dashboard is picked up by the next
|
||
system-prompt rebuild; a lesson's provenance link correctly resolves to its
|
||
originating turn; a shadow-mode lesson's dashboard row updates live as its
|
||
trial counter increments; the usage dashboard's `self-learning` line item
|
||
and trend charts sum/plot correctly against a synthetic session history.
|
||
|
||
## Validation (run after every phase, not just at the end)
|
||
|
||
```sh
|
||
cargo build --workspace
|
||
cargo clippy --workspace -- -D warnings
|
||
cargo test --workspace
|
||
cargo run -- # manual smoke: start, drive one full tool-use turn, quit cleanly
|
||
```
|
||
|
||
## Risks
|
||
|
||
| Risk | Likelihood | Mitigation |
|
||
|---|---|---|
|
||
| "Zero comments, self-documenting" collides with genuinely non-obvious logic (e.g. cache-split marker placement, hysteresis thresholds in short-send) | Medium | Push complexity into well-named functions/types instead of comments; where a magic constant needs justification, name the constant itself descriptively rather than adding a comment |
|
||
| Security toolkit scope creep into unauthorized-use tooling | Low, given your stated authorized-use context | Phase 12 ships behind an explicit enable + on-screen acknowledgment, and the plan does not add any capability beyond the reference's own curated roster |
|
||
| "Full 1:1 clone" is a multi-week effort disguised as one `/plan` | High if treated as one PR | Plan is explicitly phased; each phase is independently buildable/testable and merges before the next starts — flag now if you want a smaller first slice |
|
||
| IPC/daemon split (Phase 11) is the highest-complexity phase and has the most failure modes (stale sockets, snapshot divergence) | Medium | Scheduled late, after the single-process TUI is fully proven, so the daemon split only has to preserve already-working behavior, not debug it simultaneously |
|
||
| Catastrophic-op guard list drifts out of sync as new risky tools are added later | Medium | Guard membership is a single named function/table (mirroring `tool_is_risky`), not scattered checks — Phase 3 establishes this as the one place to extend |
|
||
| Model over-invokes or under-invokes `workflow_run` (fan-out for trivial asks, or never fans out for genuinely decomposable work) since the trigger is judgement-based, not rule-based | Medium | Phase 6's validate step includes an explicit "trivial request stays inline" smoke test; system-prompt guidance is iterated on with real usage rather than assumed correct on first write |
|
||
| `workflow_run`'s embedded script surface (`agent`/`parallel`/`pipeline`) becomes a de facto second scripting language to maintain if scope creeps beyond the four primitives | Medium | Primitives are fixed at four for v1 (`agent`, `parallel`, `pipeline`, `phase`+`log` as cosmetic); no general control flow, no arbitrary host-function calls beyond spawning agents |
|
||
| Reviewer subagent produces low-signal or wrong "lessons" (hallucinated conventions, misjudged reasons) that then get injected into every future system prompt, actively degrading quality instead of improving it | Medium | Reviewer is read-only (can't touch files, only `remember`), every `remember` requires an explicit `recall`-first dedup pass, memory entries stay human-editable/deletable (`forget`) at any time, and the plan's acceptance bar requires the verdict note to be visible in-session so a bad lesson is easy to spot and remove, not silently compounding |
|
||
| Reasoned-edit ledger (`edits.jsonl`) grows unbounded over a long-lived project, and the reviewer queue (capacity 1) backs up during a burst of many small edits | Low for v1 scope | Ledger rotation/compaction is explicitly out of scope for v1 (flag as a fast-follow); the capacity-1 reviewer queue is a deliberate backpressure choice — reviews are best-effort and coalescing/skipping under load is acceptable, unlike the ledger itself which must never drop a line |
|
||
| A bad `global`-scope lesson pollutes every future project's system prompt, unlike a project-scope mistake which stays contained | Medium | Promotion to `global` is gated behind either explicit human calibration or the multi-project repeated-violation counter — never a single reviewer's own judgement — and `/learning` (Phase 14) gives an always-available manual `forget` path for any global lesson |
|
||
| Outcome-linked build/test verification adds real wall-clock cost to every file-touching turn, and a slow/flaky project test suite could make the reviewer trigger feel laggy or noisy | Medium | Verification runs on its own thread off the critical path (the user's turn is never blocked waiting on it), is capped at a short timeout, and degrades gracefully to reasons+diff-only review on timeout/failure/misconfiguration rather than blocking the reviewer entirely |
|
||
| Batch/offline pattern mining and the per-turn reviewer could in principle compound into runaway subagent spawning on a very active session | Low | Batch mining is explicitly rate-limited to at most once per idle window and shares the same capacity-1 review queue as the per-turn reviewer, so the two can never run concurrently |
|
||
| A graduated lesson's `check` rule is a blunt substring/glob match, not real language-aware parsing — it can misfire on a legitimate edit that superficially matches the pattern | Medium | A failed check never blocks the write outright; it returns a naming result the model can address or explicitly override via `reason`, and de-graduation is always available from `/learning` if a rule proves too blunt in practice |
|
||
| Two-reviewer consensus for automatic `global` promotion doubles the token cost of that specific path, and both reviewers could still share the same blind spot if spawned from an identical prompt | Low | The cost only applies to the rare `global`-promotion path, not routine per-turn review; the two reviewers are spawned independently (fresh context each) specifically so correlated blind spots are less likely than a single pass, though not impossible — flagged as a known limitation, not a guarantee |
|
||
| `/lesson import` bulk-extracting from a large convention document could flood the calibration queue with many pending entries at once | Low | Imported entries land at `confidence: opinion` and go through the same calibration toast as any reviewer-authored lesson — no special bypass — so the existing keep/discard/auto-resolve mechanics apply unchanged, just at higher volume for that one command |
|
||
| Shadow-mode trial window adds a delay before a genuinely useful graduated check becomes protective, during which the same real mistake could recur uncaught | Low | The trial window is a small fixed count by design (not open-ended), and the underlying `lesson` itself is still actively injected into the system prompt as `opinion`/`verified` text throughout the trial — shadow mode only gates the *hard enforcement* layer, not the softer prompt-injection signal, so protection is never fully absent during the trial |
|
||
| `note_finding`'s run-scoped broadcast could be misused as a general side-channel for workflow subagents to coordinate beyond its intended "share a quality finding" purpose, growing into an unbounded ad hoc messaging primitive | Medium | Scoped to exactly one tool, one direction (broadcast, not point-to-point), ephemeral (never persisted, dies with the run), and delivered only at existing tool-round boundaries (no new scheduling primitive) — the same "fixed primitives, no scope creep" discipline already applied to `workflow_run`'s four script primitives |
|
||
| Adaptive throttling could under-review a genuinely quiet-but-still-risky stretch of edits if the backoff schedule grows too aggressive | Low | Backoff resets to zero immediately on any new lesson/escalation/check-match, and the cap on how far it can back off is a fixed named constant (same "one place to extend" discipline as the catastrophic-op guard), so the maximum possible under-review window is always a known, bounded quantity, not unbounded drift |
|
||
| Quality trend metrics could be over-interpreted as proof the self-learning loop "works" when the underlying signals (pass-rate, escalation count) can move for unrelated reasons (e.g. the project itself just got harder) | Low | Charts are presented as raw trend data for human interpretation in `/learning`, not as an automated pass/fail verdict on the feature itself — the `forget`-candidate flag is a nudge for human review, not an automatic deletion |
|
||
| Lesson-pack export could leak project-specific naming/paths embedded in a `before`/`after` snippet when shared with a team or another machine, unlike prose-only lessons | Low | Export is opt-in and explicit per invocation (never automatic), scoped to `lesson`/`retrospective` types only, and the user reviews the exported file (plain readable frontmatter+markdown, same format as any other memory file) before sharing it — same trust model as sharing any other project file |
|
||
|
||
## Acceptance
|
||
- [ ] Phase 1 skeleton builds, runs, renders the exact 3-region layout + status
|
||
bar text specified, and restores the terminal cleanly on quit
|
||
- [ ] `Cargo.toml` matches the specified name/authors/dependency set exactly
|
||
- [ ] No `unimplemented!()`, `todo!()`, `/* implement later */`, or equivalent
|
||
placeholder in any file belonging to a "done" phase
|
||
- [ ] No `//`, `///`, or `/* */` comments anywhere in `src/`
|
||
- [ ] No `#![allow(dead_code)]` / `#![allow(unused_variables)]` or per-item
|
||
`#[allow(...)]` suppressions
|
||
- [ ] Auto/Yolo mode performs routine writes/edits/shell/git with zero approval
|
||
prompts and zero added latency
|
||
- [ ] Catastrophic-op guard blocks its named operation list even in Yolo mode,
|
||
verified by unit test per rule
|
||
- [ ] Security toolkit is inert until explicitly enabled, and the enable path
|
||
surfaces an authorized-use acknowledgment
|
||
- [ ] `workflow_run` is the sole subagent-delegation surface (no separate
|
||
single-agent `task` tool coexists with it)
|
||
- [ ] Workflow Progress panel auto-appears on `workflow_run` start, is
|
||
dismissible without killing the run, and is re-summonable
|
||
- [ ] `pipeline()` demonstrably has no inter-stage barrier and `parallel()`
|
||
demonstrably respects the concurrency cap, each covered by its own test
|
||
- [ ] `write`/`edit` calls without a non-empty `reason` are rejected before any
|
||
filesystem mutation, verified by unit test for both tools
|
||
- [ ] Every accepted `write`/`edit` produces exactly one `edits.jsonl` line with
|
||
a matching `path`/`reason`/`origin`
|
||
- [ ] Exactly one detached reviewer subagent runs per turn that touched files
|
||
(never zero when edits happened outside Plan mode, never more than one
|
||
concurrently), and its verdict appears as a collapsed system note that is
|
||
never appended to the model-visible transcript
|
||
- [ ] A reviewer's `remember(type="lesson")` calls are demonstrably dedup'd
|
||
against existing lessons via a `recall`-first check (no duplicate memory
|
||
entries from repeated similar observations)
|
||
- [ ] Lessons are labeled `confidence: verified` only when backed by a
|
||
reproduced build/test failure, `opinion` otherwise, and the two are
|
||
visibly distinguished everywhere lessons are listed
|
||
- [ ] A contradictory new lesson blocks auto-save and surfaces both entries in
|
||
the calibration toast instead of overwriting silently
|
||
- [ ] Pending lessons auto-resolve to `keep` in Auto/Yolo after a grace window
|
||
(staying visible/reversible) and require an explicit keypress in Normal
|
||
mode
|
||
- [ ] A `scope: project` lesson never leaks into a different project's
|
||
injected memory index; a `scope: global` lesson appears in every project
|
||
- [ ] A lesson pattern repeated a third time within the configured window
|
||
produces a non-collapsed escalation note distinct from routine verdicts
|
||
- [ ] `/learning` dashboard (Phase 14) lists every lesson with live
|
||
confidence/stale/scope state and supports `forget` + manual promotion
|
||
- [ ] The build/test verification probe resolves correctly for at least three
|
||
distinct project-ecosystem fixtures with no ecosystem hard-coded outside
|
||
the one named probe table (language-agnostic requirement)
|
||
- [ ] A graduated lesson's `check` fires on a violating `write`/`edit` as a
|
||
named, addressable/overridable result — never a silent or unbreakable
|
||
block
|
||
- [ ] Every `lesson` resolves its `provenance` back to a real session/turn (or
|
||
`"user"` for a `/lesson`-authored entry), verified from `/learning`
|
||
- [ ] Exactly one `type: retrospective` memory is produced per ended session
|
||
and is absent from the next session's injected system-prompt index
|
||
- [ ] `/lesson <text>` creates a `confidence: human` entry with no calibration
|
||
delay; `/lesson import <path>` seeds `opinion`-confidence entries that DO
|
||
go through calibration
|
||
- [ ] Automatic (non-`/lesson`) promotion to `global` scope requires two
|
||
independently-spawned reviewers to agree, verified by a test where only
|
||
one agrees and the promotion is blocked
|
||
- [ ] The usage dashboard (Phase 14) shows a distinct `self-learning` token-
|
||
cost line item, separate from the main agent's own spend
|
||
- [ ] A `check` candidate spends its full trial window in `shadow: true`
|
||
before ever blocking/naming anything to the model, and correctly
|
||
graduates or demotes at window close based on its fire ratio
|
||
- [ ] `note_finding` calls inside a `workflow_run` are observable by sibling
|
||
`agent()` thunks in the same run and are confirmed absent from
|
||
`MEMORY.md` afterward (ephemeral, run-scoped only)
|
||
- [ ] Consecutive empty review passes measurably reduce review frequency, and
|
||
a single subsequent lesson/escalation/check-match resets it to zero
|
||
- [ ] Usage dashboard trend charts (pass-rate, escalation count, per-lesson
|
||
trigger count) plot correctly against a synthetic session history, and
|
||
a high-trigger/no-improvement lesson is flagged as a `forget` candidate
|
||
- [ ] A lesson with a structured `before`/`after` snippet renders it in both
|
||
`/learning` and its injected system-prompt text
|
||
- [ ] `/lesson export` followed by `/lesson import` on a separate project
|
||
reproduces the same lesson set with matching confidence/scope/
|
||
provenance, and a near-duplicate import is caught by the existing dedup
|
||
path
|
||
- [ ] Each phase's own validation commands pass before the next phase starts
|
||
|
||
**WAITING FOR CONFIRMATION**: Proceed with this plan? (yes/no/modify — you can
|
||
also ask to start at a specific phase, e.g. "just do Phase 1-3 for now")
|