101 KiB
Plan: zesdex — Rust TUI AI Coding + Security Agent (koma architectural clone)
Source reference: /mnt/code/koma/ (Rust src-agent + Python src-security + Python src-internet)
Complexity: Large (multi-session, phased delivery)
Summary
zesdex is a from-scratch Rust rebuild of koma's full architecture: a ratatui TUI
agent driven by an OpenRouter-compatible LLM backend, with a rich sandboxed tool set
(filesystem, shell, git, web, memory, sub-agents), a three-layer approval harness,
MCP client support, OAuth login, a detachable daemon/client IPC split for
multi-session use, and an opt-in security-tooling module wired to a Python sidecar
daemon (mirroring koma_sec_daemon) for authorized pentesting/CTF/research use
against systems the operator owns or is explicitly authorized to test.
Four explicit deviations from the literal 1:1 clone, all already confirmed with you:
-
Approval harness stays, reshaped as an opt-in speed dial, not deleted. Auto mode is genuinely fast (WC + TAC run inline, no
y/nprompts, mutations proceed immediately) — that satisfies "unrestricted auto-approve at maximum speed." What does not get removed: the deterministic catastrophic-op guard (check_destructiveingit_operator, the recursive-delete-outside-workspace check, force-push-over-history, secret/credential-file exfil patterns inweb_download/bash). That guard is not a per-turn approval prompt — it never blocks routine work — it is a hard-coded refusal on a short, named list of irreversible operations, and it fires even in Yolo/Auto mode. Everything else (file writes, edits, ordinary shell commands, git commits/branches, web fetches) runs with zero prompts in Auto/Yolo mode, exactly as requested. -
Security toolkit ships as an explicit opt-in module (
/securitypanel, disabled by default, requires the user to explicitly enable it per koma's ownsecurity_enabled/yolo_armedgating pattern) with a startup banner and--security-installflow that states the authorized-use scope, mirroring koma's own "curated, opt-in suite... for authorized testing and research" framing.sec_*tools remain harness-exempt (per koma's own design — these are explicit user-invoked offensive actions, not silent mutations to gate), but the catastrophic-op guard's network/destructive-target patterns still apply where they'd protect the operator's own machine (e.g., asec_z3/sec_sagescript that shells out torm -rf /). -
koma's
tasksubagent primitive (Phase 6) is replaced wholesale by a dynamic workflow orchestration engine, modeled directly on theWorkflowtool this very session runs on — not on anything in koma, which has no equivalent. See "Dynamic Workflow Orchestration" in the digest below and the rewritten Phase 6. Three decisions you confirmed:- Fan-out is model-decided, not forced. Simple requests are handled inline by the main agent with its own tools, exactly as today; the model only invokes the workflow engine when it judges a task benefits from decomposition, parallelism, or independent verification passes.
- Full replacement, not a layer on top. There is no separate simple
"delegate one subagent" tool anymore — a single-agent delegation is just a
one-line workflow script (
agent(...)with noparallel/pipelinewrapper), so the workflow engine is the only delegation surface, keeping one mental model instead of two. - UI: a separate live progress-tree panel, auto-shown the instant a workflow starts (phase → agent → status), dismissible/re-summonable without killing the run, while the Main Execution Log keeps behaving exactly as specified for everything else.
-
NEW — Reasoned mutations + self-learning quality loop, added after your latest request: every
write/edittool call must carry a requiredreasonargument (rejected deterministically, before touching the filesystem, if missing/empty), so every mutation is auditable after the fact. Once an agentic turn finishes and any files were touched, zesdex automatically — no model decision involved — spawns a detached background subagent (reusing Phase 6's subagent engine in koma's existing detached-nudge mode) that reviews the touched files against their stated reasons and the surrounding code's own conventions, and folds anything worth keeping back into per-project memory as a newlesson-typed entry. Because the memory index is already injected into every future turn's system prompt (koma's existing mechanism, unchanged), this closes a real feedback loop: code quality compounds across turns instead of the agent repeating the same mistake every session. See "Reasoned Mutations & Self-Learning Quality Loop" in the digest below and the new Phase 7.Following your explicit request for concrete self-learning feature suggestions, four core additions and three advanced additions were confirmed and are folded into the same digest section and Phase 7 (advanced ones note their own phase where they land instead): outcome-linked learning (lessons backed by real
build/testresults, not just LLM opinion), lesson lifecycle (decay/staleness + contradiction detection, never silent auto- delete), human-in-the-loop calibration (a keep/discard toast before a lesson becomes permanent, auto-keep in Auto/Yolo but always visible/ reversible), and a global cross-project lesson tier (mirrors the Builtin→ Global→Session tiering Phase 6 already builds for agent definitions) as core Phase 7 scope; batch/offline pattern mining, a/learningdashboard (delivered in Phase 14 alongside the rest of the remaining UI surface), and escalation on repeated violations as the advanced layer on top.A follow-up round of suggestions added three more confirmed additions, all still Phase 7/Phase 14 scope: lesson graduation into a real lint check (a lesson that survives repeatedly at
confidence: verifiedstops being just injected text and becomes a deterministic pre-write check), a session-end retrospective with per-lesson provenance (one holistic summary per session, plus every lesson traceable back to the exact turn/ edit that produced it), and user-authored lessons with consensus-gated global promotion and token-budget visibility (/lessonto author aconfidence: humanentry directly, a style-guide import path, 2-of-2 independent-reviewer agreement required before any lesson reachesglobalscope, and the reviewer loop's own token spend surfaced on the usage dashboard). Per your explicit requirement, every one of these — and everything already in Phase 7 — is language-agnostic by design: nothing in the self-learning loop hard-codes Rust or any single ecosystem (build/test detection, lint graduation, and lesson content all generalize across whatever language a given project/workspace is written in — zesdex itself being written in Rust is incidental to this feature, not a constraint on it).A third round added six more confirmed additions, still Phase 7/Phase 14 scope: shadow-mode enforcement (a lint-graduation candidate runs silent for a trial window, logging would-have-fired counts before ever surfacing to the model, so a blunt rule is caught before it annoys anyone), cross- subagent learning within one workflow run (a finding from one
agent()thunk can reach sibling thunks still running in the sameworkflow_runcall, not just the next turn/session via memory), adaptive review frequency (consecutive empty review passes throttle the reviewer down; fresh violations un-throttle it back up — the token-spend counterpart to escalation), quality trend metrics on the usage dashboard (build/test pass-rate, escalation count, and per-lesson trigger count charted over time, with lessons that fire often but never correlate with fewer escalations flagged asforgetcandidates), structured lesson bodies (abefore/aftersnippet pulled straight from the triggering diff, alongside the prose description, so a lesson is concretely actionable, not just abstract), and lesson-pack export/import (an opt-in, portable bundle of a project's lessons for moving to another machine or sharing with a team, independent of a shared~/.zesdex/).Per your request to focus on the plan itself before any code, Phase 7 — which had grown to cover all seventeen self-learning additions across three rounds of suggestions — is now split into Phase 7 (MVP core:
reason, the ledger, the automatic reviewer) plus four independently deferrable sub-phases 7B (outcome-linking/lifecycle/calibration/global tier), 7C (escalation/adaptive frequency/batch mining), 7D (lint graduation, shadow-mode gated), and 7E (retrospectives/provenance/ user-authored/consensus/export/trend metrics), restoring the "independently buildable/testable" property every other phase in this plan already has. Phase 7 alone is a complete, working answer to your original request; 7B–7E are refinements you can build immediately after or defer indefinitely. See the rewritten Phase 7 family in Tasks below.
No comments in generated Rust source, per your original instruction — self-documenting naming instead. No stubs; every module in a phase is fully implemented before that phase is considered done. Cargo.toml matches your requested metadata exactly.
Reference Architecture Digest (from /mnt/code/koma/)
Captured here so later phases can cite it without re-reading the whole reference tree.
Workspace: Cargo workspace with a single agent binary crate (src-agent),
edition 2021, plus two out-of-process Python sidecars (src-security,
src-internet) invoked via subprocess/daemon, not compiled in.
Core dependency set (from src-agent/Cargo.toml, informs zesdex's Cargo.toml):
ratatui 0.30, tokio (rt-multi-thread, macros, sync, time, net, io-util, signal),
reqwest (json, stream, blocking, native-tls-vendored), serde/serde_json/
serde_yaml_ng, anyhow, uuid (v4, v5), dirs, futures-util,
pulldown-cmark, syntect (default-fancy), rusqlite (bundled), ignore,
regex, globset, infer, base64, sha2, libc, rmcp (MCP client:
client, transport-child-process, transport-streamable-http-client-reqwest, macros),
include_dir, plus web-scraping-adjacent crates (dom_smoothie, fast_html2md,
scraper, url, percent-encoding) for the lightweight (non-Python) web-fetch path.
MVC data flow: KeyEvent → controller/input::handle_key → Action → runtime/actions::apply_action → state mutation → view::draw. Slash commands:
controller/command::parse → Command → runtime/commands::apply_slash. The
controller layer never mutates state; apply_action/apply_slash own all
mutation and async spawns; the view is read-only.
Async/streaming bridge: one tokio::runtime::Runtime created in main.rs; the
main loop is synchronous (try_recv/poll, no .await in the hot loop). One
mpsc::unbounded_channel per request stored in active_rx; harness verdicts ride
a separate harness_rx channel so they never interleave with stream tokens.
StreamEvent variants: Token, Reasoning, Usage, ToolCalls, Done, Error, Compacted, HarnessVerdict.
Agentic loop: model streams tool_calls; on Done, advance_turn commits the
assistant message, then either ends the turn or runs process_tools → gate/execute
each call → finish_tool_round → re-invoke the model. Loop continues until no more
tool calls or MAX_AGENT_STEPS (40) is hit. No forced plan gate — tools run on the
first call unless the model explicitly requests Plan mode via plan_enter.
Tool trait: name() -> &'static str, description() -> &'static str,
parameters() -> Value (JSON Schema), run(&self, ctx: &ToolCtx, args: &Value) -> Result<String>. ToolCtx carries workspace, workspaces: Vec<PathBuf>,
dir_cache, plus (per the fuller subagent report) memory dir, MCP manager, security
manager, bash-saving flag.
Deterministic tool inventory (tool_is_risky: write | delete | edit | bash | git_operator | web_download): read, grep, glob, write, edit, delete, bash, bash_output, bash_kill, cd, dir_list, dir_cache_update, pong, remember, forget, recall, task, task_output, task_kill, todowrite, web_fetch, web_search, web_download, git_cred, git_operator, git_worktree, plan_enter, plan_ready, seqthink — plus dynamic mcp__* and sec_* tools. DEFERRED_TOOLS (blocking I/O
tools) run on a plain std::thread, never inline on the event loop or inside the
tokio runtime directly.
write/edit (koma reference, tool/fs/write.rs + tool/fs/edit.rs): write
takes path + content, always creates parent dirs, always writes/overwrites,
reindexes the dir cache. edit takes path + old + new + optional
replace_all, requires the file to exist, refuses (returns a string result, not an
error) if old is missing or non-unique unless replace_all=true. Neither carries
any notion of "why" in koma today — zesdex's reason requirement (see below) is a
net-new addition on top of this exact shape, not a koma pattern.
Three-layer harness (app/harness.rs): WC (workspace containment check,
deterministic, gates the whole turn) → PC (advisory prompt classifier, fire-and-
forget, fail-open) → TAC (per-risky-call classifier, intent-aware, mode-dependent
fail-open/fail-closed policy). CLASSIFY_TIMEOUT = 12s. Modes: Auto, Normal, Plan, Yolo (AgentMode::cycle), Yolo requires explicit arming
(yolo_armed) and bypasses TAC/y/n entirely — only deterministic guards remain.
shell_filter: post-processes bash/git_operator stdout — smart per-command
filters (cargo, git status/log/diff) tried first, static regex spec table
(npm/pip/docker/make/wget-curl) as fallback, full raw output always teed to disk
when trimmed/truncated/non-zero-exit.
git_operator/git_worktree: direct argv exec (no shell), SSH key injection
via git_cred, GIT_TERMINAL_PROMPT=0, and a named destructive-op guard
(check_destructive) requiring explicit confirm_destructive=true for force-push,
reset --hard, clean -f/-d/-x, branch -D, force checkout/switch/restore,
stash drop/clear, tag -d, update-ref -d, filter-branch, gc --prune.
Worktrees live in a shadow dir outside the repo.
Subagents: up to 5 concurrent (MAX_SUBAGENTS), agent definitions are
frontmatter+markdown files merged Builtin→Global→Session tier, isolated
Conversation seed (own system prompt, no shared history with parent), own
restricted tool allow-list (always minus task/task_output/task_kill, and
zero MCP tools), engine runs the same commit/execute/classify loop non-
interactively with fail-closed risky-tool gating (no human to ask), completion
folds back via blocking-park / chat-append / detached-nudge depending on
invocation mode. The detached-nudge mode matters most for zesdex: it lets a
subagent run fully off the critical path and only surface a synthetic-user-turn
nudge (or, per Phase 7, a collapsed system note) once the session goes idle,
without blocking whatever the user does next.
Background bash: run_in_background: true bash calls spawn a plain OS thread
(not tokio — child process needs to run outside the runtime), answer the model
immediately with a job id, poll via bash_output, kill via bash_kill
(SIGTERM to the recorded pid), completion is toast + buffered synthetic-user-turn
nudge once the session goes idle.
IPC/daemon: daemon-per-session, not daemon-per-machine — each session UUID
gets its own headless process at ~/.<app>/run/<session_id>.sock. Transport: Unix
domain socket, 4-byte-BE-length-prefixed JSON frames (FrameReader reassembles
split/coalesced reads, 64 MiB hard cap enforced at the prefix). bind() success is
the sole liveness oracle (never a PID file). Per-client connection task splits
into independent async read/write halves bridged to the synchronous daemon
loop via std::sync::mpsc (daemon loop stays sync so it shares turn-servicing
code with a hypothetical standalone/no-daemon mode). ClientRequest/DaemonFrame
message sets cover attach/detach/list/status/submit/shell/key/paste/approve/plan-
decision/new/quit/switch-foreground. Multiple clients can attach one daemon
(controller + viewers, promotion on controller disconnect). Self-exit after a
grace window once all sessions are tombstoned and no client is enrolled.
Snapshot/diff: daemon builds a pure StateSnapshot from live AppState;
per-client diff() decides needs_full (structural change → resend full
snapshot; correctness-first, "when in doubt, ask for a full snapshot") vs a
cheap Vec<StateDelta> for pure-append/status/scroll changes. Client applies
frames into a shadow AppState and reuses the same view::draw renderer,
so daemon and thin-client rendering can never diverge.
Session locking: one session.lock file per session holding the owner PID;
liveness via kill(pid, 0) (portable, not /proc parsing); stale locks
opportunistically cleared. Lock always written with the daemon's PID, never
the thin client's.
MCP: rmcp-based client, two transports (stdio child process, streamable
HTTP), tools namespaced mcp__<sanitized_server>__<tool>, Local vs Proxy
(shared global MCP daemon) backend behind one facade so multiple session-daemons
don't each spawn duplicate heavyweight MCP servers, sync↔async bridge for
Tool::run compatibility, 20s connect / 60s call timeouts, non-blocking
background connect/reconnect with a generation counter to discard stale in-flight
connects.
OAuth: PKCE for the primary provider (loopback TcpListener on a fixed local
port for the redirect, hand-parsed HTTP, CSRF state check), device-authorization
polling for a secondary provider, unverified local JWT payload decode for display
metadata only, single-flight per-provider refresh with a stale/near-expiry lead
window and an unrecoverable latch on invalid_grant.
TUI layout: header (mode indicator) → transcript (scrollable, markdown-
rendered assistant turns, plain-text user/streaming turns) → optional pending-
steer panel → model-name row → input box (dynamic height, capped at half the
frame) → status bar. Markdown rendering emits pre-wrapped Vec<Vec<Span>> so
line count matches on-screen rows exactly (needed for scroll math); syntect
lazy-singleton highlighting in fenced code blocks with hard-splitting on
over-width lines; tables/lists/blockquotes word-wrap via a shared char-level
wrapper. Theme: a Palette struct of ~12 semantic colors + a 5-stop heatmap
ramp, computed fresh per frame from config, several named palettes plus an
accent-color resolver.
Mode state machine: far more than "5 modes" — onboarding, provider-OAuth
onboarding, credentials form, session picker/hub, chat, loading splash, settings,
agents editor, MCP manager, help, effort picker, usage dashboard, message rewind,
quit confirm, security panel, background-bash panel, todo panel. Each mode is a
struct variant on one Mode enum, exactly one active at a time, stored per-session.
Security sidecar (koma_sec_daemon equivalent): long-lived Python child
process, newline-delimited JSON frames over stdin/stdout (not a socket) —
handshake with a shared token, call/health/install ops, single-threaded
serialized dispatch. Tool roster spans web (http, sqlmap, nuclei, ffuf, dalfox,
zap, xss-confirm), crypto (z3, sage, rsa, factordb, lattice/fpylll, hashcat,
hashid, generic decode), reverse-engineering (js deobfuscate/unminify, source-map
recovery, wasm decompile), and pwn (static triage, ROPgadget, pwntools remote
session, exploit-template scaffold). Sandbox = wall-clock timeout + optional
RLIMIT_AS/RLIMIT_CPU + optional (opt-in, currently-unused-upstream) bubblewrap
network unshare — no container/chroot/seccomp. Tiered installer (pip / GitHub-
release binary / gem / manual-detect-only) with a health-probe surfaced in the UI.
Internet sidecar (scrapion_agent equivalent): stateless, one-shot
subprocess per call (python -m ... --json "<url-or-query>"), Playwright-Firefox-
headless search+scrape with stealth tweaks, HTML→Markdown conversion, JSON out on
stdout. Rust side spawns it on a plain OS thread (not tokio) with a recv_timeout
guard and falls back to a raw HTTP fetch on any failure.
Dynamic Workflow Orchestration (new — not present in koma; modeled on this
session's own Workflow tool): a script-driven orchestration layer that lets the
main agent decompose a task into many subagents running in parallel, pipeline, or
phased stages, rather than koma's one-subagent-at-a-time task primitive.
- Trigger: exposed to the model as one new tool,
workflow_run(script, args). There is no separate "spawn one subagent" tool — a single delegated task is simply a one-agent()-call script, soworkflow_runis the sole delegation surface. The system prompt instructs the model to reach for it when a task benefits from decomposition/parallelism/independent verification, and to skip it (use its ownread/grep/edit/bashdirectly) for anything simple — this is model judgement, not a hard trigger rule, matching your "dynamic, model decides" choice. - Script primitives (a small embedded scripting surface, not a general
scripting language — deliberately narrow so it stays auditable and
serializable into the persisted transcript):
agent(prompt, opts?) -> AgentResult— spawn one subagent (own isolated conversation seed, own restricted tool allow-list, same non-interactive engine loop as koma's subagent engine: commit/execute/classify with fail-closed risky-tool gating since there's no human to ask).optscoverslabel,phase,modeloverride,effortoverride.parallel(thunks) -> Vec<AgentResult>— run N agent thunks concurrently, barrier: all must finish before returning. Bounded by a concurrency cap (mirrors koma'sMAX_SUBAGENTS-style ceiling, default 5, configurable).pipeline(items, stage1, stage2, ...) -> Vec<Any>— run each item through every stage independently with no barrier between stages (item A can be on stage 3 while item B is on stage 1); the default recommended shape for multi-stage fan-out because wall-clock is the slowest single chain, not the sum of slowest-per-stage.phase(title)— label subsequentagent()calls under a named phase for the progress-tree UI; purely cosmetic/organizational, no execution effect.log(message)— narrator line surfaced in the progress panel above the tree.
- Execution model:
workflow_runitself is aDEFERRED_TOOL(per the existingDEFERRED_TOOLSpattern) whose body runs on the tokio runtime (it needs real concurrency forparallel/pipeline, unlike the blocking-thread tools); eachagent()call inside it reuses the exact same subagent spawn/ engine/fold-back machinery Phase 6 already needed for koma parity — soworkflow_runis a thin orchestration layer over individual subagent spawns, not a reimplementation of subagent execution. Concurrency, isolation (ownConversationseed, own tool allow-list, zero recursion intoworkflow_runitself to bound blast radius), and the fail-closed risky-tool gate are all inherited unchanged from the subagent engine design. - Result delivery:
workflow_runblocks the calling turn (like a normal tool call) and returns a structured summary (per-agent status + final return value) as the tool result once every stage completes; the model then synthesizes a response from that structured result, same as any other tool round-trip — no separate detached/nudge mode for v1 (koma's detached-task nudge pattern is not carried over here, since workflow runs are expected to be the user-visible unit of work, not a background fire-and-forget job). - TUI surface: a dedicated Workflow Progress panel, auto-shown the
instant
workflow_runstarts (overlaying/pushing above the transcript region, matching koma's overlay pattern forBash/Todo/MessageRewindmodes — samelayout_chunksreuse, new overlay content), rendering a live tree: phase → per-agent row → status glyph (queued/running/done/error) → elapsed time. Dismissible with Esc without killing the run (run keeps streaming in the background; re-summon via a key binding or$-style panel toggle, mirroring koma's sub-agents panel UX). On completion, the panel's final state persists in scrollback as a collapsed summary block in the Main Execution Log, and the panel auto-closes after a short delay unless pinned open. - Persistence: each
workflow_runcall and its final structured result is logged to the session transcript like any tool call/result pair — no separate workflow-history store for v1.
Reasoned Mutations & Self-Learning Quality Loop (new — not present in koma; added for your "tiap edit atau write harus ada parameter alasan... bg agent untuk verifikasi kualitas... self learning" request):
reasonbecomes a required argument onwriteandedit(koma's reference shapes above gain one field each):{"path", "content", "reason"}forwrite,{"path", "old", "new", "replace_all"?, "reason"}foredit. Enforcement is deterministic and happens before any filesystem call, exactly like the existingarg_strrequired-field pattern already used forpath/content/old/new— a missing or blankreasonreturnserror: missing required string argument 'reason'and the file is never touched. This is intentionally the same shape astool_is_risky/WC: a name- based, always-on, non-negotiable check, not something the classifier or agent mode can waive.deleteandbashare explicitly NOT changed — the user's ask was scoped to edit/write, and widening it would duplicate what the harness/catastrophic-op guard already cover for those.- Every accepted call appends one line to an append-only per-session ledger,
<session_dir>/edits.jsonl(same directory tier assession.lock/settings.json, same atomic-append discipline asmodel::memory::atomic_writeuses for whole-file writes — here it's anO_APPENDline write, since the ledger is never rewritten in place). Each line:{ts, tool, path, reason, content_sha256, bytes_delta}. The ledger is the audit trail the reviewer (below) reads; it is never truncated or rotated automatically in v1. - Auto-triggered review, not model-invoked. When
advance_turnreaches a turn boundary (no moretool_calls, same pointDonealready fires today) and that turn's slice ofedits.jsonlis non-empty, the event loop — not the model, not a tool call — spawns exactly one subagent using Phase 6's existing engine in koma's detached-nudge invocation mode, seeded with a new built-in agent-definition persona,quality-reviewer. Its tool allow-list is read-only plus memory:read, grep, glob, recall, remember. It never seeswrite/edit/bash, so it cannot itself mutate anything — a review pass can only ever produce a memory note or nothing.- Input: the turn's ledger lines (path + reason + diff-relevant content) plus
a
recallof the project's existinglesson-typed memories, so the reviewer can check "have I already said this" before writing a new one — this is what keepsMEMORY.mdfrom growing an unbounded stream of restated observations. If an existing lesson already covers the point, the reviewer does nothing (no duplicateremembercall). - Output: zero or more
remember(..., type="lesson")calls for anything genuinely worth carrying forward (a real convention violated, a recurring smell, a reason that didn't match what the diff actually did), plus exactly one short verdict string returned as the subagent's final result. - The verdict is rendered as a collapsed system note in the Main
Execution Log (a new
StreamEvent-adjacent variant, e.g.SystemNote, routed the same wayCompacted/HarnessVerdictalready ride their own channel so it can never interleave with live stream tokens) — it is explicitly NOT appended to the model-visible conversation history, so review passes cost the reviewer's own tokens only and never grow the main agent's context. - Because
model::memory::load_memory_indexalready injects the wholeMEMORY.mdindex (now including anylessonentries) into every future turn's system prompt, this is the entire feedback loop — no new injection plumbing needed, only the newkindvalue and the auto-trigger.
- Input: the turn's ledger lines (path + reason + diff-relevant content) plus
a
- Bounded, not free-running: at most one reviewer subagent runs at a time
(reviews queue rather than pile up — reuses the same bounded-concurrency
primitive Phase 6 builds for
parallel(), capacity fixed at 1 for this queue); the trigger is skipped entirely whileAgentMode::Planis active (nothing mutates in Plan mode, soedits.jsonlcan't have a non-empty slice) and skipped for edits made by a reviewer or workflow subagent itself (no recursive review-the-reviewer chains — ledger lines carry anorigin: main | subagent | reviewertag precisely so this filter is a cheap equality check, not a heuristic). - Explicitly reuses, does not reinvent: Phase 6's subagent spawn/engine/
fold-back machinery (detached-nudge mode already exists there for other
reasons);
model::memory'swrite_memory/read_memory/slugifyprimitives, widened by onekindvalue; the existing pattern of a dedicated side-channel for out-of-band events (harness_rxtoday, a parallelnotes_rxforSystemNote).
Outcome-linked learning (language-agnostic by design — zesdex itself being
a Rust project is incidental; the mechanism below never assumes Rust): the
reviewer's input is upgraded from "reasons + diff only" to "reasons + diff +
real verification result." Detection of "how do I build/test this project"
is a small ordered probe, not a hard-coded cargo call: a verify_command
in settings.json always wins if the user sets one; otherwise zesdex probes
for well-known project markers already visible to dir_cache (Cargo.toml →
cargo build && cargo test, package.json → its scripts.test/scripts. build if present, pyproject.toml/requirements.txt → pytest/python -m pytest if discoverable, go.mod → go build ./... && go test ./..., and so
on for the other common ecosystems); if no marker matches, verification is
simply skipped for that project and the reviewer runs on reasons+diff alone.
This probe table lives in one named function (mirroring the tool_is_risky/
catastrophic-op-guard pattern of "one place to extend"), so adding a new
ecosystem later is a one-line addition, not a redesign. The resolved command
runs on a plain std::thread, capped at a short timeout, and hands the
reviewer the pass/fail + trimmed output alongside the ledger. A lesson born
from an actual build/test failure is written with confidence: verified; a
lesson born from the reviewer's own judgement alone (no reproducible failure)
is written with confidence: opinion. Both render in the memory index, but
opinion entries are visually distinguished (dimmer/marked) in every mode
that lists memories, so verified lessons are never confused with the
reviewer's own guesses. If the verification command itself is unavailable or
misconfigured, the reviewer still runs on reasons+diff alone — this is a
strict enhancement, never a hard dependency.
Lesson lifecycle: every lesson entry gains two frontmatter fields beyond
koma's existing name/description/type — confidence (verified | opinion,
from outcome-linked learning above) and last_confirmed (a timestamp, bumped
whenever a later review cites the same lesson as still applicable). A
lightweight idle-time sweep (piggybacks on the same idle-detection the
background-bash completion nudge already uses, no new polling loop) flags —
never silently deletes — any lesson whose last_confirmed is older than a
configurable staleness window as stale: true in its frontmatter; a stale
lesson still gets injected into the index (so it's still visible) but is
labeled so the model and the user both know it may no longer hold. Separately,
before any new remember(type="lesson") call is accepted (both from the
regular reviewer and human-in-the-loop confirmation below), a cheap
recall-and-compare pass checks the new lesson's text against existing lessons
for direct contradiction (e.g. "always use X" vs. a new "never use X"); a
detected contradiction blocks the auto-save and instead surfaces both entries
side by side in the calibration toast for a human decision — contradictions
are never silently resolved by whichever write happens to land last.
Human-in-the-loop calibration: a remember(type="lesson") call from the
quality-reviewer does not go straight into MEMORY.md. It lands first in a
small pending queue (one more small JSON file alongside edits.jsonl,
pending_lessons.json, same atomic-write discipline) and renders as a
collapsed toast in the Main Execution Log: Lesson baru: <description> [keep]/[discard], using the same SystemNote side-channel Phase 7 already
adds. In Auto/Yolo mode the toast auto-resolves to keep after a short grace
window (matching the "unrestricted, no prompts" mandate for routine flow) but
the toast and its outcome remain visible/scrollable in the log and reversible
at any time via forget; in Normal/Plan mode (where a human is already
expected to be attentive) the toast waits for an explicit keypress instead of
auto-resolving. Either way, nothing about this blocks the main agent's own
turn — calibration only gates the reviewer's side-channel writes, never the
user-facing conversation.
Global cross-project lesson tier: mirrors the Builtin→Global→Session
tiering Phase 6 already builds for agent-definition files. A lesson memory
gains a scope field (project | global); project is the existing default
(<pwd_bucket_dir>/memory/, unchanged), global writes instead to a new
~/.zesdex/memory/ store shared across every project. The reviewer defaults
every lesson to project scope; a lesson is only promoted to global when
either (a) the human-in-the-loop toast is used to explicitly promote it, or
(b) the escalation-on-repeated-violations mechanism (below) fires the same
lesson's underlying pattern across two or more distinct projects — never
silently promoted by the reviewer's own judgement alone, since a
project-specific convention wrongly promoted to global would pollute every
other project's context. load_memory_index (Phase 4/existing koma mechanism)
is extended to concatenate the global index ahead of the project index when
building the system prompt, so global lessons are always present and project
lessons can still override/refine them by appearing after.
Batch/offline pattern mining (advanced, Phase 7 scope but lower priority
within it): beyond the per-turn reviewer, an idle-triggered (same idle
detection as the lifecycle sweep) deeper pass periodically walks the full
edits.jsonl history (not just the latest turn's slice) plus git log -p
for the same file set, looking for cross-turn patterns a single-turn reviewer
structurally cannot see (e.g., "every time parser.rs changes, the matching
test file does not"). Runs as one more detached subagent, same allow-list and
confidence/calibration path as the per-turn reviewer — it is the same
mechanism at a longer time horizon, not a separate code path. Rate-limited to
at most once per idle window so it can never compound with the per-turn
reviewer into runaway subagent spawning.
Escalation on repeated violations: every accepted ledger line is checked
(cheap substring/embedding-free heuristic first — a lesson's slug matched
against the touched file's recent edits.jsonl history — deferred to the
reviewer's own judgement only when the heuristic is ambiguous) against the
existing lesson index for a repeat violation of something already
remembered. On the third repeat of the same lesson within a configurable
window, zesdex stops treating it as routine background learning and instead
surfaces a non-collapsed, attention-grabbing note directly in the Main
Execution Log (not the quieter SystemNote styling used for ordinary
verdicts) — the premise being that if index-injection alone hasn't stopped
the pattern after three tries, silently trusting the system prompt again is
not working and a human should see it. This counter is also the trigger
condition that can promote a lesson from project to global scope per the
tiering feature above, when the same pattern repeats across multiple distinct
project directories rather than one.
Lesson graduation into a lint check (language-agnostic — the check itself
is expressed as a pattern-matching rule against the tool call's own arguments/
diff text, never a Rust-specific AST): once a lesson has stayed at
confidence: verified and been cited by the escalation counter above at
least once (i.e., it has already proven itself both true and still being
violated), it becomes eligible for graduation. A graduated lesson gains a
check frontmatter field — a small structured rule (substring/regex against
the file's new content, or a path-glob + required-companion-file rule, e.g.
"editing parser.* without also touching a matching *test* file") — and
from that point on is enforced the same way the reason argument is: a
deterministic, always-on check inside write/edit, evaluated before the
harness/classifier layer, same shared pattern as arg_str. A failed
graduated-lesson check does not silently block the write; it returns a
result string naming the lesson and asking the model to address it or
explicitly override with a reason that acknowledges the exception — so
graduation raises the bar without ever becoming an unbreakable wall. Only a
human (via /learning, Phase 14) or the consensus-promotion mechanism below
can graduate or de-graduate a lesson — the reviewer itself cannot
self-promote into an enforced check, for the same "no single opinion becomes
policy" reason global promotion is gated.
Session-end retrospective + per-lesson provenance: every lesson memory
gains a provenance field recording the session id + turn index + ledger
line hash it was born from (or, for a /lesson-authored entry below, a
literal "user" marker instead) — a cheap addition since all of that is
already available at the point remember is called, purely for later
auditability (/learning can jump from a lesson straight to the originating
edit). Separately, when a session ends (explicit /quit or the daemon's
tombstone-on-idle path, Phase 11), one more detached subagent — reusing the
identical quality-reviewer-style engine, not a new code path — is hand a
session-wide view (every ledger line + every lesson touched this session,
not just the last turn) and produces a single holistic summary memory
(type: retrospective, one per session, not one per turn) capturing things a
per-turn reviewer structurally cannot ("this session repeatedly restructured
the same module," "three different lessons this session all pointed at the
same underlying gap"). Retrospectives are additive and separate from
lesson entries — they are not injected into the system prompt by default
(they're a human-facing summary, browsable from /learning), so they cannot
bloat future turns' context the way an unbounded lesson count could.
User-authored lessons, consensus-gated global promotion, token-budget
visibility: a new /lesson <text> slash command lets the user author a
lesson memory directly, tagged confidence: human — a rank above both
verified and opinion, since it is authoritative by construction and
therefore exempt from the contradiction-block/calibration-toast path
entirely (a human overriding a prior automated lesson is always accepted
immediately, never queued). A companion /lesson import <path> command feeds
an existing convention document (CONTRIBUTING.md, a style guide, any
markdown file the user points at) through the same quality-reviewer-style
subagent to extract candidate lessons in bulk, seeding the project's memory
from day one instead of starting from zero — extracted entries still land at
confidence: opinion (the source document is trusted, but the extraction
is still an LLM judgement call) and go through the normal calibration toast
before becoming permanent. Global promotion, already gated behind either
explicit human calibration or the multi-project repeated-violation counter
(per the tiering feature above), is tightened further: an automatic
(non-/lesson) promotion additionally requires two independent reviewer
subagents (spawned fresh, no shared context between them — cheap since
review subagents are already isolated by design) to agree the pattern
generalizes, before it's written to ~/.zesdex/memory/; a single reviewer's
opinion is never sufficient for a global-scope write. A /lesson-authored
entry can still be promoted to global directly by the user with no
consensus requirement, same "human is authoritative" exemption as above.
Finally, every reviewer/retrospective/batch-mining subagent's token usage is
tagged with a origin: self-learning cost-ledger category (reusing whatever
per-turn cost tracking the usage dashboard, Phase 14, already accumulates for
the main agent) and surfaced as its own line item there, so the self-learning
loop's token cost is always visible against the session total rather than
invisibly folded into "background work."
Shadow-mode enforcement: a lint-graduation candidate (per the graduation
mechanism above) does not go straight from confidence: verified to an
active check. It first spends a configurable trial window (N subsequent
write/edit calls against matching files, default a small fixed count) in
shadow: true state: the same rule evaluation runs on every matching call,
but a match only increments a would-have-fired counter on the lesson's own
frontmatter — it never returns anything to the model and never blocks
anything. Only once the trial window closes does the rule either graduate to
a real active check (per the existing enforcement path) or, if its
would-have-fired rate looks too indiscriminate (fires on most edits to a
matching file rather than the specific pattern it was meant to catch — a
simple ratio threshold, not a judgement call), get flagged back down to a
plain opinion lesson instead of ever reaching enforcement. This reuses the
exact same rule-evaluation code path enforcement already needs (Phase 7); the
only new state is the shadow/trial_count pair on the lesson and the one
branch deciding graduate-vs-demote when the window closes.
Cross-subagent learning within one workflow run: today, a finding one
agent() thunk makes only reaches the rest of the system on the next turn,
once the reviewer has run and the memory index has been rebuilt — too late
to help sibling thunks still executing in the same workflow_run call.
workflow_run's existing execution context (Phase 6) gains one small
in-memory, run-scoped broadcast channel (not persisted, not MEMORY.md —
this is intentionally ephemeral and scoped to the single run, distinct from
the durable per-project lesson store): any subagent spawned via agent() can
call a new restricted tool, note_finding(text), whose only effect is to
push text onto that run's shared findings list; every other subagent
still active in the same run receives the current findings list prepended to
its next tool-round system context, cheaply (no re-spawn, no interruption of
in-flight work — it lands on the subagent's next turn boundary like any other
context update). Findings are ephemeral to the run: they do not automatically
become lesson memories (that still requires the normal reviewer pass after
the workflow completes, reading the same edits.jsonl ledger as any other
turn) — this feature closes the within-run propagation gap, not the
long-term memory gap, which the rest of Phase 7 already covers.
Adaptive review frequency: the per-turn reviewer trigger gains a small
run-length counter — consecutive review passes that produce zero new lessons
(no remember calls at all, verified or opinion) increment it; any pass that
produces at least one lesson resets it to zero. Past a small threshold of
consecutive empty passes, the trigger throttles: instead of reviewing every
qualifying turn, it skips a growing number of turns between reviews (simple
backoff, not a fixed schedule), directly reducing the self-learning-tagged
token spend the usage dashboard now surfaces. The moment a review does find
something (or the escalation counter fires, or a graduated check matches),
the throttle resets to zero immediately — the system is quick to re-engage
and only slow to relax, so a real emerging problem is never delayed by a
prior quiet streak. This is the direct token-cost counterpart to escalation
(escalation raises alarm on repeated violations; this lowers cost on
repeated silence), and both read the same underlying review-history data,
just in opposite directions.
Quality trend metrics on the usage dashboard: alongside the self- learning token-cost line item already planned, the usage dashboard (Phase
14) accumulates and charts three time series purely from data Phase 7's other
features already produce (no new instrumentation): build/test pass-rate from
outcome-linked verification runs, escalation-note count from the
repeated-violations mechanism, and per-lesson trigger count from the same
counter escalation already reads. A lesson whose trigger count is high but
whose citations never precede a drop in the escalation rate (i.e., it keeps
firing without the underlying pattern ever actually going away) is flagged
as a forget candidate in /learning — a lesson that keeps getting matched
but never changes behavior is more likely to be miscalibrated or ignored
than genuinely useful, so it's surfaced for human review rather than kept
accumulating silently.
Structured lesson bodies: a lesson memory's body gains an optional
before/after code-snippet pair alongside its existing prose description —
extracted directly from the specific ledger line / diff that triggered the
lesson (the reviewer already has this content in hand when it calls
remember; this is a shape change to what it passes, not new data
collection). Rendering (in /learning, in the collapsed verdict note, and
implicitly in whatever text gets injected into future system prompts) shows
the snippet pair when present, falling back to prose-only for lessons that
predate this field or for which no single before/after pair meaningfully
captures the point (e.g. a cross-cutting retrospective observation). This
makes a lesson concretely actionable — "change shape X to shape Y" — rather
than purely descriptive, without requiring any new tool or trigger.
Lesson-pack export/import: an opt-in /lesson export <path> command
serializes a project's (or, with a flag, the global store's) lesson
entries — full frontmatter including confidence/scope/provenance/
check state, structured bodies, everything — into one portable bundle file,
independent of and without requiring a shared ~/.zesdex/. /lesson import <path> on another machine (or by a teammate) merges a pack into the local
store using the exact same recall-first dedup and contradiction-detection
path new reviewer-authored lessons already go through (Phase 7's lifecycle
mechanism, reused verbatim — an imported pack is just another source of
candidate lessons, not a privileged bypass), so imported entries land in the
calibration queue exactly like any other non-/lesson addition unless they
came from a /lesson-authored source originally (in which case they keep
their confidence: human standing on import too, since that's a property of
the entry, not of the transport). Export/import is explicitly scoped to
lesson/retrospective memory types only — it does not touch session
transcripts, edits.jsonl, or anything else in the per-session directory.
Patterns to Mirror
| Category | Source | Pattern |
|---|---|---|
| State mutation | app/runtime/actions/*.rs, app/runtime/commands/*.rs |
All state mutation flows through apply_action/apply_slash; controller/view layers are read-only translators |
| Tool definition | tool/mod.rs Tool trait |
Zero-size struct per tool, name/description/parameters/run, registered in one all_tools() vec, risk flag is a name-based match in one place |
| Required-argument enforcement | tool/fs/*.rs arg_str helper |
Missing/blank required args fail deterministically before any I/O, via one small shared helper — the pattern zesdex's new reason field reuses verbatim |
| Async bridge | app/runtime/stream/*.rs |
One mpsc channel per request; main loop stays synchronous; blocking tool I/O deferred to std::thread, not .awaited inline |
| Approval gating | app/harness.rs, stream/tools/approval.rs |
Deterministic check → advisory background classifier → per-call intent-aware classifier, each layer independently timeout-bounded and explicit about fail-open vs fail-closed |
| Memory persistence | model/memory.rs |
Index-of-pointers (MEMORY.md) + one frontmatter+body file per entry, atomic rename-over-write, hard-sanitized slugs, only the index (never full bodies) injected into the system prompt |
| IPC framing | ipc/frame.rs |
4-byte-BE length prefix + JSON, hard cap enforced before payload allocation |
| Daemon liveness | ipc/server.rs, model/session_lock.rs |
bind()/connect() success is the liveness oracle, never a PID file; lock files store PID + kill(pid,0) liveness check |
| Config layout | model/settings.rs, model/app_config.rs |
Global prefs in one small config.json; all per-session config in settings.json under a per-session directory |
| Sidecar protocol | koma_sec_daemon/protocol.py |
Newline-delimited JSON frames, single shared secret handshake, single-threaded serialized dispatch, every error surfaces as a string result, never a crash |
Files to Create (Phase 1 scope; later phases add modules incrementally)
| File | Action | Why |
|---|---|---|
Cargo.toml |
CREATE | Package metadata + dependency set per your spec, extracted from koma's src-agent/Cargo.toml |
src/main.rs |
CREATE | Entry point: terminal setup/teardown, tokio runtime construction, mode dispatch into the event loop |
src/app/mod.rs, src/app/state.rs, src/app/mode.rs |
CREATE | AppState, Mode enum (Chat + Onboard to start), core runtime state |
src/app/runtime/mod.rs, event_loop.rs, actions.rs, commands.rs, stream.rs |
CREATE | Synchronous main loop, Action/Command dispatch, tokio stream bridge |
src/app/harness.rs |
CREATE | WC/PC/TAC three-layer approval harness + catastrophic-op guard |
src/controller/input.rs, src/controller/command.rs |
CREATE | Key→Action translation, slash-command parsing |
src/view/mod.rs, view/chat.rs, view/theme.rs, view/markdown.rs, view/status.rs |
CREATE | Ratatui rendering: 3-region chat layout + status bar per your spec |
src/tool/mod.rs, tool/fs.rs, tool/search.rs, tool/shell.rs, tool/shell_filter/*.rs |
CREATE | Tool trait + filesystem/search/shell tool set with output filtering; write/edit schemas include required reason from day one |
src/tool/git_operator.rs, tool/git_worktree.rs, tool/git_cred.rs |
CREATE | Git tool set with destructive-op guard |
src/dto/chat.rs, dto/openrouter.rs |
CREATE | Wire types for the LLM API (OpenRouter-compatible) |
src/service/mod.rs, service/openrouter.rs |
CREATE | Streaming client, StreamEvent, classifier calls |
src/model/session.rs, model/settings.rs, model/store.rs, model/app_config.rs, model/conversation.rs |
CREATE | Session persistence, per-session settings, config |
src/resources.rs, src-misc/*.txt |
CREATE | Embedded system/personality/tool-usage/classifier prompts |
Later phases (see Tasks) add: src/model/memory.rs + src/model/editlog.rs
(reasoned-edit ledger + lesson-typed memory, Phase 7), src/app/review/mod.rs
(auto quality-review trigger, Phase 7; extended in-place by 7B/7C/7D/7E — no
new top-level module per sub-phase), src/app/subagent/* +
src/app/workflow/{mod,script,engine}.rs + src/tool/workflow.rs +
src/view/workflow.rs + src/app/mode workflow-panel overlay (dynamic workflow
orchestration, replaces koma's task, Phase 6), src/app/bgbash/* (background
jobs), src/app/mcp/* (MCP client), src/service/oauth/* (OAuth), src/ipc/*
(daemon/client split), src/tool/internet/* (web fetch/search/download),
src/app/sec/* + a security-sidecar/ Python package (security toolkit),
src/model/msglog.rs (SQLite blob archive), src/app/runtime/shortsend.rs
(token-efficiency rail).
Milestones, Dependency Graph, and Implementation Gates
This section turns the long phase list below into an executable roadmap. It does not remove scope; it groups the work so implementation can proceed in stable, reviewable increments without letting the advanced self-learning surface block the basic product from becoming usable.
Milestone groups
| Milestone | Phases | Outcome | Can stop here? |
|---|---|---|---|
| M0 — Plan freeze | Current document | Scope, risks, dependencies, and acceptance are frozen before code starts | Yes — no code written yet |
| M1 — Native TUI shell + agent loop | 1 → 2 → 3 | A usable single-process TUI that can stream model output, run core tools, auto-approve routine work, and still refuse catastrophic operations | Yes — first functional developer preview |
| M2 — Durable sessions + context economy | 4 → 5 | Sessions survive restarts, /resume works, long conversations stay usable via SQLite archive + short-send shaping |
Yes — usable daily-driver baseline without subagents |
| M3 — Dynamic workflow orchestration | 6 | workflow_run replaces koma's task, supports many subagents, progress panel, parallel/pipeline, and run-scoped note_finding |
Yes — first multi-agent release |
| M4 — Self-learning core and growth path | 7 → 7B → 7C → 7D → 7E | Required edit reasons, audit ledger, background reviewer, lesson lifecycle, escalation, enforced checks, retrospectives, import/export, dashboard data | Yes after any sub-phase; 7 alone is useful |
| M5 — Background jobs + extensibility + auth | 8 → 9 → 10 | Long-running shell jobs, MCP tools, and OAuth login are integrated without blocking the TUI | Yes — plugin/provider-ready release |
| M6 — Detachable multi-session runtime | 11 | Daemon/client split, attach/detach, snapshot/diff/shadow-state rendering | Yes — multi-session release |
| M7 — External capability packs | 12 → 13 | Authorized security toolkit and full internet/search/fetch/download stack land behind their respective opt-in controls | Yes — full agent capability release |
| M8 — Full UI parity polish | 14 | Settings, agents editor, MCP browser, usage dashboard, /learning, help, rewind, quit confirm, session hub |
Yes — full koma-parity UI release |
Dependency graph
M0 plan freeze
↓
Phase 1 TUI skeleton
↓
Phase 2 streaming agent loop + core tools
↓
Phase 3 approval harness + catastrophic-op guard
↓
Phase 4 session persistence
↓
Phase 5 SQLite archive + short-send
↓
Phase 6 workflow orchestration
├─ requires Phase 2 tool loop
├─ requires Phase 3 non-interactive risky-tool gating for subagents
└─ provides subagent engine used by Phase 7+
↓
Phase 7 self-learning MVP
├─ requires Phase 6 detached subagent execution
├─ requires Phase 4 session directory for edits.jsonl
└─ provides lesson memory + review side-channel
↓
Phase 7B outcome/lifecycle/calibration/global tier
└─ requires Phase 7 lesson memory + review trigger
↓
Phase 7C escalation/adaptive/batch mining
└─ requires Phase 7B confidence/scope/lifecycle fields
↓
Phase 7D graduated checks + shadow mode
├─ requires Phase 7B lesson frontmatter fields
└─ benefits from Phase 7C repeated-violation counter
↓
Phase 7E retrospectives/provenance/user lessons/export/trends
├─ requires Phase 7B confidence/scope model
├─ requires Phase 7C counters for trends
├─ requires Phase 7D checks for graduation UI
└─ rendered fully in Phase 14 dashboard
Independent after M3/M4 baseline:
Phase 8 background bash jobs
Phase 9 MCP client
Phase 10 OAuth login
Late structural split:
Phase 11 daemon/client IPC
├─ requires stable Phase 1 renderer
├─ requires stable Phase 4 session model
└─ should happen after Phase 8 so background nudges are daemon-aware
External packs:
Phase 12 security toolkit
├─ requires Phase 3 catastrophic guard
├─ requires Phase 4 settings/session storage
└─ benefits from Phase 11 daemon lifecycle, but can be built single-process first
Phase 13 internet module
├─ requires Phase 2 tool loop
├─ requires Phase 3 web_download guard adjacency
└─ benefits from Phase 8 background job patterns for sidecar timeout handling
Final UI parity:
Phase 14 remaining UI modes
├─ depends on the modes/features it surfaces
└─ can be partially implemented earlier for `/learning` if Phase 7E needs it
Implementation gates
Each gate must pass before the next milestone starts. This is stricter than "cargo build passes"; it prevents hidden partials from accumulating.
| Gate | Required before | Criteria |
|---|---|---|
| G1 — TUI correctness gate | Start Phase 2 | Phase 1 renders the exact 3-region layout, handles resize, input, scroll basics, and restores terminal state after panic/quit |
| G2 — Tool-loop correctness gate | Start Phase 3 | Phase 2 can complete a full mocked tool-use loop, including tool-call parse, tool execution, result append, and second model call |
| G3 — Safety/speed gate | Start Phase 4 | Auto/Yolo routine mutations have zero prompts; catastrophic-op guard blocks every named rule even in Yolo; no classifier failure can bypass deterministic guards |
| G4 — Persistence gate | Start Phase 5/6 | Sessions create/list/resume cleanly, stale locks are cleared, and no session file write can corrupt partial state on crash |
| G5 — Context gate | Start Phase 6 | Long conversation fixture proves SQLite archive and short-send shaping reduce wire payload without destroying transcript fidelity |
| G6 — Workflow gate | Start Phase 7 | workflow_run is the only delegation surface, parallel cap and pipeline no-barrier semantics are tested, progress panel survives dismiss/resummon, and note_finding is ephemeral/run-scoped |
| G7 — Self-learning MVP gate | Start 7B | write/edit require reason, ledger is append-only and exact, one detached reviewer runs per edited turn, and verdicts never enter model-visible history |
| G8 — Lesson lifecycle gate | Start 7C | Confidence, stale, scope, calibration, contradiction blocking, and global/project memory injection all have direct tests |
| G9 — Escalation/adaptive gate | Start 7D | Repeated-violation escalation and adaptive review throttling both work under synthetic histories and never run reviewers concurrently |
| G10 — Graduated-check gate | Start 7E | Shadow-mode trials, graduated checks, override behavior, and de-graduation are tested without making any check an unbreakable blocker |
| G11 — Learning UI data gate | Start Phase 14 /learning UI |
Provenance, retrospectives, import/export, trend metrics, and self-learning token-cost records are all present as model/data APIs before UI work begins |
| G12 — Daemon split gate | Start Phase 11 | Single-process runtime is stable through Phases 1–10; renderer state can be snapshotted without borrowing live TUI objects |
| G13 — External capability gate | Start Phase 12/13 | Tool gating, settings, background subprocess supervision, and side-channel notes are stable enough that security/internet sidecars cannot destabilize core chat |
| G14 — Full parity gate | Finish Phase 14 | Every mode has an Esc/back path, every dashboard reflects live data, and every feature-specific acceptance item below remains green |
Recommended implementation order
For the actual coding pass, the safest order is:
- M1: Phases 1–3, because they produce the first runnable agent shell and the always-on catastrophic guard.
- M2: Phase 4, then Phase 5 only if long-session support is needed before subagents; otherwise Phase 5 can be done immediately after Phase 6.
- M3: Phase 6, because every self-learning feature depends on subagent execution and workflow orchestration.
- M4: Phase 7 only, stop, test it hard, then decide whether to continue with 7B–7E immediately or ship the MVP learning loop first.
- M5/M6/M7/M8: Phases 8–14 in order, except
/learningUI slices from Phase 14 may be pulled forward after 7E if needed for manual calibration.
The important policy: do not start Phase 7B before Phase 7 is stable, and do not start Phase 11 before the single-process app is already boringly reliable.
Tasks
Phase 1 — Skeleton: Cargo.toml, event loop, 3-panel TUI, status bar
- Action: Stand up the buildable skeleton exactly matching your explicit
layout spec: Main Execution Log (transcript), Input Prompt (bottom fixed-height
block), Status Bar (
[zesdex] | STATUS: AUTO-APPROVE (UNRESTRICTED) | PROTOCOL: ZERO-STUBS ACTIVE). Wire the synchronous main loop + tokio runtime bridge with no LLM calls yet (apong-style echo tool only), so the rendering/input/event loop is provably correct before the agentic loop is layered on. - Mirror:
app/runtime/event_loop/mod.rstick structure (drain stream events → drain harness verdicts → input poll with adaptive 8ms/100ms timeout → draw if dirty);view/chat/mod.rslayout_chunks region split. - Validate:
cargo build,cargo clippy -- -D warnings, manual TUI smoke run (type a line, see it echo into the log region, resize terminal, quit cleanly restoring the terminal state).
Phase 2 — OpenRouter-compatible client + streaming agentic loop
- Action: Implement
dto::chat/dto::openrouterwire types, the streaming SSE client,StreamEventbridge, andadvance_turn/process_toolsloop against a real filesystem/shell tool set (read, grep, glob, write, edit, delete, bash, dir_list, cd).write/editalready requirereasonper the Files-to-Create note above — enforced now, even though the review loop that reads the ledger doesn't land until Phase 7, so the ledger has real data by the time Phase 7 needs it. No approval gating yet — this phase proves the model can drive tools end-to-end. - Mirror:
service/openrouter.rsstreaming/classify pattern;tool/mod.rsTooltrait andall_tools()registration;MAX_AGENT_STEPSloop cap. - Validate: Integration test against a mock SSE server; a
writecall withoutreasonis rejected before any file is created; manual run against a real OpenRouter-compatible endpoint editing a scratch file end-to-end.
Phase 3 — Approval harness: Auto/Normal/Plan/Yolo + catastrophic-op guard
- Action: Implement WC (deterministic workspace containment), TAC (per-risky-
call classifier, intent-aware, mode-dependent fail-open/fail-closed), and the
AgentModecycle (Auto/Normal/Plan/Yolo). Implement the catastrophic-op guard as a small, explicit, always-on function independent of mode/classifier state — modeled directly on koma'sgit_operator::check_destructivelist (force-push,reset --hard,clean -f/-d/-x, force checkout, recursive delete outside any configured workspace root, credential-file read/exfil patterns) — and extend it tobash/delete/web_downloadper your "keep a hard stop on catastrophic ops" decision. Skip PC (advisory prompt classifier) as lower value for v1; note it in follow-up. - Mirror:
app/harness.rsWC/TAC design,git_operator.rs::check_destructive. - Validate: Unit tests per guard rule (force-push blocked without confirm,
reset --hardblocked, ordinarygit commit/write/editin Auto mode runs with zero prompts and zero added latency vs. Phase 2 baseline).
Phase 4 — Session model + on-disk persistence + /resume
- Action:
Session/Settings/AppConfig/store(create/list/rename sessions), per-session directory layout (settings.json,messages.json,session.lock,memory/,edits.jsonl), session picker mode. - Mirror:
model/session.rs,model/store.rs,model/session_lock.rs(PID-liveness lock, never a stale-file oracle). - Validate: Kill
-9a session mid-write, confirm lock is detected stale and cleared on next list;/resumeround-trips a full conversation.
Phase 5 — SQLite blob archive + short-send token efficiency (optional-early)
- Action:
messages.sqlitearchive (append-only messages/blobs/summary tables), non-destructive wire-payload compaction rail (shortsend::shape) so budget models stay in-context on long sessions. - Mirror:
model/msglog.rs,app/runtime/shortsend.rsengage-hysteresis math. - Validate: Long synthetic conversation exceeding context window compacts the
wire payload while
messages.json/transcript stay uncompressed.
Phase 6 — Dynamic workflow orchestration engine (replaces koma's task)
- Action: Build the subagent execution core first (isolated
Conversationseed, restricted tool allow-list, non-interactive engine loop with fail-closed risky-tool gating — this part is still a direct mirror of koma's subagent engine design, just not exposed as a standalonetasktool). Layer theworkflow_run(script, args)tool on top with the four primitives (agent/parallel/pipeline/phase/log), a bounded concurrency pool (default 5, configurable), and the Workflow Progress panel (newModeoverlay reusinglayout_chunks, live phase→agent→status tree, dismiss/ re-summon without killing the run). Update the system prompt (src-misc/ system-tools.txtequivalent) to describeworkflow_runand give the model concrete guidance on when decomposition is worth it vs. handling inline. Agent-definition format (frontmatter+markdown, tiered Builtin→Global→Session merge) is still implemented soagent()calls can reference named personas, and so Phase 7's built-inquality-reviewerpersona has somewhere to live. - Mirror:
app/subagent/{context,engine,event,spawn}.rs,model/ agent_def/*.rsfor the execution core;view/chat/subagents.rs+app/mode/mod.rsoverlay-mode pattern for the progress panel's UI shell (no koma source for the orchestration script layer itself — that's new, specified above under "Dynamic Workflow Orchestration"). Also addnote_finding(text)— a restricted tool available only inside aworkflow_runexecution context (not part of a standalone subagent's allow-list) that pushestextonto a run-scoped, in-memory, non-persisted findings broadcast; every otheragent()thunk still active in the same run receives the current findings list at its next tool-round boundary. This is workflow-execution coordination, not lesson persistence, which is why it lives here rather than in Phase 7 — a finding only becomes a durablelessonif the normal Phase 7 reviewer independently decides torememberit after the run, reading the sameedits.jsonlledger as any other turn. - Mirror:
app/subagent/{context,engine,event,spawn}.rs,model/ agent_def/*.rsfor the execution core;view/chat/subagents.rs+app/mode/mod.rsoverlay-mode pattern for the progress panel's UI shell (no koma source for the orchestration script layer itself, nor fornote_finding— both are new, specified above under "Dynamic Workflow Orchestration"). - Validate: A
parallel()workflow with 6 agent thunks confirms the 6th queues on the concurrency cap; apipeline()workflow with 3 items × 2 stages confirms item A reaches stage 2 while item B is still on stage 1 (no barrier); progress panel updates live during a real multi-agent run and correctly survives Esc-dismiss-and-resummon without losing state; a trivial single-step request in a manual smoke test is handled inline by the model without invokingworkflow_runat all (confirms the "dynamic, not forced" trigger behavior is actually working, not just documented); inside aworkflow_runwith two concurrentagent()thunks, anote_findingcall from one is observably visible to the other's next tool round, and is confirmed absent fromMEMORY.mdafter the run (ephemeral, not persisted).
Phase 7 — Reasoned mutations + self-learning quality loop (MVP core)
Phase 7 was originally one block covering every self-learning idea explored
in planning, but that made it far larger and more tightly coupled than every
other phase in this plan — a violation of the plan's own "each phase is
independently buildable/testable" principle. It is now split into this core
phase plus four additive sub-phases (7B–7E, deliberately lettered
rather than renumbered so Phases 8–14 below are untouched). Phase 7 alone
is the complete, working answer to your original request — required
reason, an audit ledger, and an automatic background quality pass; 7B–7E
are refinements layered on top, each independently deferrable if you want a
smaller first slice of the self-learning feature specifically. Anywhere the
Risks/Acceptance tables below say "Phase 7," read it as "the 7/7B/7C/7D/7E
family" unless a sub-phase is called out explicitly.
- Action: Add the required
reasonargument towrite/edit(schema + deterministic rejection before any filesystem call, before any I/O); implement the append-onlyedits.jsonlledger with anorigin: main | subagent | reviewertag on every line; add alessonmemory kind tomodel::memory(name/description/typeonly for now — the richer frontmatter fields land in 7B–7E) plus a recall-before-remember dedup convention so a duplicate observation is never saved twice; add the built-inquality-revieweragent-definition persona (read-only +recall/rememberallow-list, nowrite/edit/bash, so a review pass can only ever produce a memory note or nothing); wire the auto-trigger into the event loop's existing turn-completion point (the same pointDonealready fires) so a non-empty ledger slice for the just-finished turn spawns exactly one detached-nudge subagent (Phase 6's engine, detached-nudge mode), capacity-1 queued if another review is already running, skipped entirely in Plan mode and for ledger lines whoseoriginissubagent/reviewer(no recursive review-the-reviewer chains); add theSystemNoteside-channel (parallel toharness_rx) so the reviewer's verdict renders as a collapsed note in the Main Execution Log without ever entering the model-visible transcript. This is the complete MVP: reasoned mutations, an audit trail, and an automatic background quality pass that folds observations into memory the existingload_memory_indexmechanism already injects every future turn. - Mirror:
tool/fs/write.rs/edit.rs(extend schema, reusearg_str),model/memory.rs(addkind, reusewrite_memory/read_memory/slugify/atomic_write),app/subagent/spawn.rsdetached-nudge fold-back (built in Phase 6),app/runtime/event_loop/mod.rsturn-completion hook,harness_rx's dedicated-channel pattern for the newSystemNotechannel. - Validate: A
write/editcall missingreasonis rejected before any file is created/modified (unit test, both tools); a turn with 3 edits produces exactly 3 ledger lines with matching reasons andorigin: main; a workflow subagent's own edits are taggedorigin: subagentand do NOT double-trigger a nested review; after a turn with edits, exactly one reviewer subagent runs detached and the Main Execution Log gets exactly one collapsed verdict note with zero additional turns appended to the model-visible conversation; a deliberately duplicate lesson (same point, reworded) is not re-saved because the reviewer'srecall-first check finds the existing entry; Plan mode never triggers a reviewer even after a hypothetical mutation (defence in depth).
Phase 7B — Outcome-linked learning, lesson lifecycle, calibration, global tier
Builds directly on Phase 7's lesson kind and reviewer trigger; independently
deferrable from 7C/7D/7E. See "Outcome-linked learning," "Lesson lifecycle,"
"Human-in-the-loop calibration," and "Global cross-project lesson tier" in the
digest above for full mechanism detail — summarized here as build items.
- Action: Extend
lessonfrontmatter withconfidence: verified | opinion,last_confirmed,stale: bool, andscope: project | global; implement the language-agnostic build/test probe table (settings override → ordered project-marker detection → skip if none match) as one named, independently-extensible function, run on a plain thread and handed to the reviewer alongside the ledger; implement thepending_lessons.jsoncalibration queue (auto-resolve-to-keep after a grace window in Auto/Yolo, explicit keypress in Normal/Plan) with contradiction detection blocking auto-save; implement the idle-triggered staleness sweep that flags (never deletes) oldlessonentries; implement the~/.zesdex/memory/global store and extendload_memory_indexto concatenate global-then-project. - Mirror:
shell_filter's per-command-special-case pattern generalized into the build/test probe table;model/memory.rsfrontmatter extension;app/bgbash/*'s idle-detection reused for the staleness sweep;app/mode/*Builtin→Global→Session tiering pattern reused for theproject/globallesson split. - Validate: A lesson born from a reproduced build/test failure is tagged
confidence: verifiedand one born from reviewer judgement alone is taggedopinion, visibly distinguished; a deliberately contradictory lesson is blocked from auto-save and surfaces both entries in the calibration toast; in Auto mode a pending lesson auto-resolves tokeepafter the grace window and remainsforget-able afterward, in Normal mode it waits for an explicit keypress; an artificially back-datedlast_confirmedgets flaggedstale: truewithout deletion; ascope: projectlesson does not leak into a second project's injected index, whilescope: globalappears in both; the build/test probe resolves correctly for at least three distinct project-ecosystem fixtures with no ecosystem hard-coded outside the one probe table.
Phase 7C — Escalation, adaptive review frequency, batch pattern mining
Builds on 7B's lesson/confidence fields; independently deferrable. See "Escalation on repeated violations," "Adaptive review frequency," and "Batch/offline pattern mining" in the digest above.
- Action: Implement the repeated-violation counter that escalates to a
non-collapsed attention note (distinct styling from the routine collapsed
SystemNote) on the third repeat of a lesson within a configurable window, and can promote a lesson fromprojecttoglobalscope when the repeat spans multiple project directories; implement the consecutive-empty-pass counter that throttles the per-turn review trigger's frequency and resets to zero on any new lesson/escalation; implement the idle-triggered batch pattern-mining pass over fulledits.jsonl+git log -phistory, rate-limited to once per idle window and sharing the same capacity-1 review queue as the per-turn reviewer. - Mirror:
app/bgbash/*'s idle trigger as the template for both the batch-mining pass and the throttle counter's reset condition; ledgerorigintagging (Phase 7) reused as the cheap equality check for repeat-detection scope. - Validate: The same lesson pattern violated a third time within the window produces a non-collapsed escalation note distinct from routine verdicts; a synthetic run of N consecutive empty review passes measurably reduces review frequency, and a single subsequent lesson-producing pass resets it to reviewing every turn; batch mining never runs concurrently with a per-turn review (shared capacity-1 queue verified under load).
Phase 7D — Lesson graduation into enforced checks (shadow-mode gated)
Builds on 7B/7C; independently deferrable — a project can run indefinitely on
lesson-as-injected-text alone without ever enabling this sub-phase. See
"Lesson graduation into a lint check" and "Shadow-mode enforcement" in the
digest above.
- Action: Add the
checkfrontmatter field (a small structured rule: substring/regex against new content, or a path-glob + required-companion- file rule) and its enforcement insidewrite/edit(same deterministic- before-harness slot asreason) for graduated lessons — a failed check never blocks the write, it returns a naming, addressable/overridable result; gate graduation behind a mandatory shadow-mode trial window (shadow: bool+trial_countfrontmatter, same rule-evaluation path, would-have-fired counting only, nothing surfaced to the model) with a graduate-vs-demote decision at window close based on fire ratio; restrict manual graduation/de-graduation to/learning(Phase 14) or the two- reviewer consensus mechanism (Phase 7E) — never reviewer self-promotion. - Mirror:
tool/fs/write.rs/edit.rs'sreason-enforcement slot, reused verbatim forcheckenforcement. - Validate: A
checkcandidate spends its full trial window inshadow: truebefore ever surfacing anything to the model, and correctly graduates or demotes at window close based on its fire ratio; a graduated lesson'scheckfires on a violatingwrite/editas a named, addressable result — never a silent or unbreakable block.
Phase 7E — Retrospectives, provenance, user-authored lessons, consensus promotion, export/import, trend metrics
Builds on 7B (confidence/scope) and 7D (consensus gates graduation promotion); independently deferrable — the MVP loop (Phase 7) and its lifecycle/escalation refinements (7B/7C) are fully useful without this. See "Session-end retrospective + per-lesson provenance," "User-authored lessons, consensus-gated global promotion, token-budget visibility," "Quality trend metrics on the usage dashboard," "Structured lesson bodies," and "Lesson-pack export/import" in the digest above.
- Action: Add a
provenancefield to everylesson(session id + turn index + ledger line hash, or"user"for/lesson-authored entries) and the session-end retrospective trigger (reuses the daemon idle/tombstone path, Phase 11, and the same reviewer-engine shape at session scope) writing onetype: retrospectivememory per session, never injected into the system prompt; add/lesson <text>and/lesson import <path>slash commands (confidence: humanfor direct authorship, exempt from calibration/contradiction;opinionfor bulk-extracted imports, which DO go through calibration); tighten automatic (non-/lesson)globalpromotion to require two independently-spawned reviewer subagents to agree before the write; tag every self-learning subagent's token spend withorigin: self-learningand surface it as its own usage-dashboard line item (Phase 14), plus the three trend time series (build/test pass-rate, escalation count, per-lesson trigger count) and the trigger-without- improvementforget-candidate flag, all derived from data 7A–7D already produce; extend alesson's body with an optional structuredbefore/aftersnippet pulled from the triggering ledger line/diff; add/lesson export <path>and/lesson import <path>(project or--globalscope), the latter routed through the exact same recall-first dedup/contradiction/ calibration path as any other new lesson, restricted tolesson/retrospectivememory types only. - Mirror:
app/subagent/spawn.rsdetached-nudge fold-back (Phase 6, reused unchanged for retrospective and two-reviewer consensus spawns);controller/command.rsslash-command parsing for/lesson; koma's existing usage-dashboard view, extended for the new cost category and trend charts. - Validate: Exactly one
type: retrospectivememory is produced per ended session and is absent from the next session's injected system-prompt index;/lesson "text"creates aconfidence: humanentry immediately with no calibration toast;/lesson importseedsopinion-confidence entries that DO go through calibration; automatic global promotion is blocked when only one of two spawned reviewers agrees, and proceeds when both do; the usage dashboard shows a nonzeroself-learning-tagged token total and correctly plotted trend charts against a synthetic session history, with a high-trigger/no-improvement lesson flagged as aforgetcandidate; a lesson with a structuredbefore/afterpair renders it in both/learningand its injected system-prompt text;/lesson exportthen/lesson importon a second (empty) project reproduces the same lesson set with matchingconfidence/scope/provenance, and a near-duplicate import is caught by the existing dedup path.
Phase 8 — Background bash jobs
- Action:
run_in_backgroundbash execution on a plain OS thread, job registry,bash_output/bash_kill, idle-triggered completion nudge. - Mirror:
app/bgbash/mod.rsthread/reader-thread/status-enum design. - Validate: Long-running background command (
sleep 30 && echo done) polls correctly mid-run and nudges the model on completion without blocking the UI.
Phase 9 — MCP client
- Action:
rmcp-based client, stdio + streamable-HTTP transports, tool namespacing/dedup,/mcpsettings UI, config persisted in global config. - Mirror:
app/mcp/mod.rsconnect/dispatch/sync-async-bridge design. - Validate: Connect to a local test MCP stdio server, confirm its tools appear namespaced in the model's advertised tool list and dispatch correctly.
Phase 10 — OAuth login
- Action: PKCE authorization-code flow with a loopback callback listener,
token cache with single-flight refresh,
/settingsOAuth submenu. - Mirror:
service/oauth/{pkce,loopback,manager}.rsdesign. - Validate: Full browser round-trip against a real OAuth-compatible provider in a sandboxed test app registration.
Phase 11 — Daemon/client IPC split (detachable multi-session)
- Action: Unix-socket, length-prefixed JSON-frame IPC; daemon-per-session
process model; snapshot/diff/shadow-state client rendering reusing the Phase 1
view layer unmodified; CLI daemon subcommands (
status/kill/restart/clean). - Mirror:
ipc/{frame,server,client,conn}.rs,ipc/snapshot/*.rs,app/runtime/client/*.rs,app/runtime/event_loop/daemon/*.rs. - Validate: Detach a running session, confirm the daemon keeps streaming in the background; reattach and confirm the shadow state matches exactly (no full-vs-delta divergence) after a burst of token/tool-call activity.
Phase 12 — Security toolkit (opt-in, authorized-use only)
- Action:
security-sidecar/Python package mirroringkoma_sec_daemon(newline-JSON stdin/stdout protocol, tool registry, sandbox timeouts, tiered installer, health probe). Rust-sideSecDaemonManager,/securitypanel requiring explicit enable + an on-screen authorized-use acknowledgment before the panel arms,sec_*tools harness-exempt per koma's own design (explicit user-invoked, not silent). Tool roster ported 1:1 from the reference inventory (http, sqlmap, nuclei, ffuf, dalfox, zap, xss-confirm, z3, sage, rsa, factordb, lattice, hashcat, hashid, decode, jsdeobf, unmin, sourcemap, wasm, triage, rop, pwntmpl, remote-session). - Mirror:
src-security/koma_sec_daemon/*end to end. - Validate: Health probe correctly reports install status per tool; a representative tool from each category (web/crypto/re/pwn) round-trips a real call against a local authorized test target (e.g., a deliberately vulnerable local CTF binary/webapp, not a live third-party system).
Phase 13 — Internet module (search/fetch/download + full browser mode)
- Action: Lightweight in-process
web_fetch/web_search(reqwest + html2md, no Python dependency) plus opt-ininternet-sidecar/Python package (Playwright-Firefox scraping) for pages that resist the lightweight path, invoked as a stateless one-shot subprocess per call with a hard timeout and fallback to the lightweight path on any failure.web_downloadcarries the 500 MiB cap and is in the catastrophic-op-guard-adjacent risky set (still no approval prompt in Auto/Yolo, but still subject to the destructive-target guard if it targets a path outside the workspace). - Mirror:
tool/internet/*.rs,src-internet/scrapion_agent/*. - Validate: Fetch a JS-heavy page that fails the lightweight path, confirm automatic fallback to the browser sidecar succeeds and returns clean markdown.
Phase 14 — Remaining UI surface (settings, agents editor, MCP browse, usage
dashboard, help, effort picker, message rewind, quit confirm, session hub,
/learning dashboard)
- Action: Fill out the remaining
Modevariants and their views/controllers so the app matches koma's full UI surface, not just chat. Add the new/learningdashboard mode (no koma equivalent — new for zesdex's Phase 7 feature set): lists everylessonmemory (project + global, clearly separated), sortable by trigger count /confidence/staleflag / recency, with inlineforget, manualproject → globalpromotion, and manual lesson-graduation/de-graduation actions reusing the same primitives Phase 7 built; a per-lesson detail view resolves itsprovenanceback to the originating session/turn and renders its structuredbefore/aftersnippet when present; a shadow-mode lesson's row shows its live would-have-fired counter and trial-window progress; a separate list browsestype: retrospectivememories by session;/lesson export/importare also reachable as dashboard actions (not just slash commands), for discoverability; the existing usage dashboard mode gains aself-learning- tagged line item alongside the main agent's own token spend plus the three quality-trend charts (build/test pass-rate, escalation count, per-lesson trigger count) and the trigger-without-improvementforget-candidate flagging (Phase 7 already tags/produces all of this data — this phase only needs to render it). - Mirror:
app/mode/*.rs,view/*.rsper-mode files (/learninghas no koma source — built directly on Phase 7'smodel::memoryextensions and the existingsession_hub-style list-mode view pattern for the layout shell); koma's existing usage-dashboard view (extended, not replaced, for the new cost category and trend charts). - Validate: Manual pass through every mode transition documented in the
reference's mode state machine; no dead-end mode with no Esc/back path;
/learningcorrectly reflects a lesson's liveconfidence/stale/scopestate and aforgetperformed from the dashboard is picked up by the next system-prompt rebuild; a lesson's provenance link correctly resolves to its originating turn; a shadow-mode lesson's dashboard row updates live as its trial counter increments; the usage dashboard'sself-learningline item and trend charts sum/plot correctly against a synthetic session history.
Validation (run after every phase, not just at the end)
cargo build --workspace
cargo clippy --workspace -- -D warnings
cargo test --workspace
cargo run -- # manual smoke: start, drive one full tool-use turn, quit cleanly
Risks
| Risk | Likelihood | Mitigation |
|---|---|---|
| "Zero comments, self-documenting" collides with genuinely non-obvious logic (e.g. cache-split marker placement, hysteresis thresholds in short-send) | Medium | Push complexity into well-named functions/types instead of comments; where a magic constant needs justification, name the constant itself descriptively rather than adding a comment |
| Security toolkit scope creep into unauthorized-use tooling | Low, given your stated authorized-use context | Phase 12 ships behind an explicit enable + on-screen acknowledgment, and the plan does not add any capability beyond the reference's own curated roster |
"Full 1:1 clone" is a multi-week effort disguised as one /plan |
High if treated as one PR | Plan is explicitly phased; each phase is independently buildable/testable and merges before the next starts — flag now if you want a smaller first slice |
| IPC/daemon split (Phase 11) is the highest-complexity phase and has the most failure modes (stale sockets, snapshot divergence) | Medium | Scheduled late, after the single-process TUI is fully proven, so the daemon split only has to preserve already-working behavior, not debug it simultaneously |
| Catastrophic-op guard list drifts out of sync as new risky tools are added later | Medium | Guard membership is a single named function/table (mirroring tool_is_risky), not scattered checks — Phase 3 establishes this as the one place to extend |
Model over-invokes or under-invokes workflow_run (fan-out for trivial asks, or never fans out for genuinely decomposable work) since the trigger is judgement-based, not rule-based |
Medium | Phase 6's validate step includes an explicit "trivial request stays inline" smoke test; system-prompt guidance is iterated on with real usage rather than assumed correct on first write |
workflow_run's embedded script surface (agent/parallel/pipeline) becomes a de facto second scripting language to maintain if scope creeps beyond the four primitives |
Medium | Primitives are fixed at four for v1 (agent, parallel, pipeline, phase+log as cosmetic); no general control flow, no arbitrary host-function calls beyond spawning agents |
| Reviewer subagent produces low-signal or wrong "lessons" (hallucinated conventions, misjudged reasons) that then get injected into every future system prompt, actively degrading quality instead of improving it | Medium | Reviewer is read-only (can't touch files, only remember), every remember requires an explicit recall-first dedup pass, memory entries stay human-editable/deletable (forget) at any time, and the plan's acceptance bar requires the verdict note to be visible in-session so a bad lesson is easy to spot and remove, not silently compounding |
Reasoned-edit ledger (edits.jsonl) grows unbounded over a long-lived project, and the reviewer queue (capacity 1) backs up during a burst of many small edits |
Low for v1 scope | Ledger rotation/compaction is explicitly out of scope for v1 (flag as a fast-follow); the capacity-1 reviewer queue is a deliberate backpressure choice — reviews are best-effort and coalescing/skipping under load is acceptable, unlike the ledger itself which must never drop a line |
A bad global-scope lesson pollutes every future project's system prompt, unlike a project-scope mistake which stays contained |
Medium | Promotion to global is gated behind either explicit human calibration or the multi-project repeated-violation counter — never a single reviewer's own judgement — and /learning (Phase 14) gives an always-available manual forget path for any global lesson |
| Outcome-linked build/test verification adds real wall-clock cost to every file-touching turn, and a slow/flaky project test suite could make the reviewer trigger feel laggy or noisy | Medium | Verification runs on its own thread off the critical path (the user's turn is never blocked waiting on it), is capped at a short timeout, and degrades gracefully to reasons+diff-only review on timeout/failure/misconfiguration rather than blocking the reviewer entirely |
| Batch/offline pattern mining and the per-turn reviewer could in principle compound into runaway subagent spawning on a very active session | Low | Batch mining is explicitly rate-limited to at most once per idle window and shares the same capacity-1 review queue as the per-turn reviewer, so the two can never run concurrently |
A graduated lesson's check rule is a blunt substring/glob match, not real language-aware parsing — it can misfire on a legitimate edit that superficially matches the pattern |
Medium | A failed check never blocks the write outright; it returns a naming result the model can address or explicitly override via reason, and de-graduation is always available from /learning if a rule proves too blunt in practice |
Two-reviewer consensus for automatic global promotion doubles the token cost of that specific path, and both reviewers could still share the same blind spot if spawned from an identical prompt |
Low | The cost only applies to the rare global-promotion path, not routine per-turn review; the two reviewers are spawned independently (fresh context each) specifically so correlated blind spots are less likely than a single pass, though not impossible — flagged as a known limitation, not a guarantee |
/lesson import bulk-extracting from a large convention document could flood the calibration queue with many pending entries at once |
Low | Imported entries land at confidence: opinion and go through the same calibration toast as any reviewer-authored lesson — no special bypass — so the existing keep/discard/auto-resolve mechanics apply unchanged, just at higher volume for that one command |
| Shadow-mode trial window adds a delay before a genuinely useful graduated check becomes protective, during which the same real mistake could recur uncaught | Low | The trial window is a small fixed count by design (not open-ended), and the underlying lesson itself is still actively injected into the system prompt as opinion/verified text throughout the trial — shadow mode only gates the hard enforcement layer, not the softer prompt-injection signal, so protection is never fully absent during the trial |
note_finding's run-scoped broadcast could be misused as a general side-channel for workflow subagents to coordinate beyond its intended "share a quality finding" purpose, growing into an unbounded ad hoc messaging primitive |
Medium | Scoped to exactly one tool, one direction (broadcast, not point-to-point), ephemeral (never persisted, dies with the run), and delivered only at existing tool-round boundaries (no new scheduling primitive) — the same "fixed primitives, no scope creep" discipline already applied to workflow_run's four script primitives |
| Adaptive throttling could under-review a genuinely quiet-but-still-risky stretch of edits if the backoff schedule grows too aggressive | Low | Backoff resets to zero immediately on any new lesson/escalation/check-match, and the cap on how far it can back off is a fixed named constant (same "one place to extend" discipline as the catastrophic-op guard), so the maximum possible under-review window is always a known, bounded quantity, not unbounded drift |
| Quality trend metrics could be over-interpreted as proof the self-learning loop "works" when the underlying signals (pass-rate, escalation count) can move for unrelated reasons (e.g. the project itself just got harder) | Low | Charts are presented as raw trend data for human interpretation in /learning, not as an automated pass/fail verdict on the feature itself — the forget-candidate flag is a nudge for human review, not an automatic deletion |
Lesson-pack export could leak project-specific naming/paths embedded in a before/after snippet when shared with a team or another machine, unlike prose-only lessons |
Low | Export is opt-in and explicit per invocation (never automatic), scoped to lesson/retrospective types only, and the user reviews the exported file (plain readable frontmatter+markdown, same format as any other memory file) before sharing it — same trust model as sharing any other project file |
Acceptance
- Phase 1 skeleton builds, runs, renders the exact 3-region layout + status bar text specified, and restores the terminal cleanly on quit
Cargo.tomlmatches the specified name/authors/dependency set exactly- No
unimplemented!(),todo!(),/* implement later */, or equivalent placeholder in any file belonging to a "done" phase - No
//,///, or/* */comments anywhere insrc/ - No
#![allow(dead_code)]/#![allow(unused_variables)]or per-item#[allow(...)]suppressions - Auto/Yolo mode performs routine writes/edits/shell/git with zero approval prompts and zero added latency
- Catastrophic-op guard blocks its named operation list even in Yolo mode, verified by unit test per rule
- Security toolkit is inert until explicitly enabled, and the enable path surfaces an authorized-use acknowledgment
workflow_runis the sole subagent-delegation surface (no separate single-agenttasktool coexists with it)- Workflow Progress panel auto-appears on
workflow_runstart, is dismissible without killing the run, and is re-summonable pipeline()demonstrably has no inter-stage barrier andparallel()demonstrably respects the concurrency cap, each covered by its own testwrite/editcalls without a non-emptyreasonare rejected before any filesystem mutation, verified by unit test for both tools- Every accepted
write/editproduces exactly oneedits.jsonlline with a matchingpath/reason/origin - Exactly one detached reviewer subagent runs per turn that touched files (never zero when edits happened outside Plan mode, never more than one concurrently), and its verdict appears as a collapsed system note that is never appended to the model-visible transcript
- A reviewer's
remember(type="lesson")calls are demonstrably dedup'd against existing lessons via arecall-first check (no duplicate memory entries from repeated similar observations) - Lessons are labeled
confidence: verifiedonly when backed by a reproduced build/test failure,opinionotherwise, and the two are visibly distinguished everywhere lessons are listed - A contradictory new lesson blocks auto-save and surfaces both entries in the calibration toast instead of overwriting silently
- Pending lessons auto-resolve to
keepin Auto/Yolo after a grace window (staying visible/reversible) and require an explicit keypress in Normal mode - A
scope: projectlesson never leaks into a different project's injected memory index; ascope: globallesson appears in every project - A lesson pattern repeated a third time within the configured window produces a non-collapsed escalation note distinct from routine verdicts
/learningdashboard (Phase 14) lists every lesson with live confidence/stale/scope state and supportsforget+ manual promotion- The build/test verification probe resolves correctly for at least three distinct project-ecosystem fixtures with no ecosystem hard-coded outside the one named probe table (language-agnostic requirement)
- A graduated lesson's
checkfires on a violatingwrite/editas a named, addressable/overridable result — never a silent or unbreakable block - Every
lessonresolves itsprovenanceback to a real session/turn (or"user"for a/lesson-authored entry), verified from/learning - Exactly one
type: retrospectivememory is produced per ended session and is absent from the next session's injected system-prompt index /lesson <text>creates aconfidence: humanentry with no calibration delay;/lesson import <path>seedsopinion-confidence entries that DO go through calibration- Automatic (non-
/lesson) promotion toglobalscope requires two independently-spawned reviewers to agree, verified by a test where only one agrees and the promotion is blocked - The usage dashboard (Phase 14) shows a distinct
self-learningtoken- cost line item, separate from the main agent's own spend - A
checkcandidate spends its full trial window inshadow: truebefore ever blocking/naming anything to the model, and correctly graduates or demotes at window close based on its fire ratio note_findingcalls inside aworkflow_runare observable by siblingagent()thunks in the same run and are confirmed absent fromMEMORY.mdafterward (ephemeral, run-scoped only)- Consecutive empty review passes measurably reduce review frequency, and a single subsequent lesson/escalation/check-match resets it to zero
- Usage dashboard trend charts (pass-rate, escalation count, per-lesson
trigger count) plot correctly against a synthetic session history, and
a high-trigger/no-improvement lesson is flagged as a
forgetcandidate - A lesson with a structured
before/aftersnippet renders it in both/learningand its injected system-prompt text /lesson exportfollowed by/lesson importon a separate project reproduces the same lesson set with matching confidence/scope/ provenance, and a near-duplicate import is caught by the existing dedup path- Each phase's own validation commands pass before the next phase starts
WAITING FOR CONFIRMATION: Proceed with this plan? (yes/no/modify — you can also ask to start at a specific phase, e.g. "just do Phase 1-3 for now")