Files
shiro-neko/TODO.md
T
asepharyana 97834e1411
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s
Merge remote-tracking branch 'refs/remotes/upstream/main'
# Conflicts:
#	ROADMAP.md
#	TODO.md
#	src/cli.tsx
#	src/config.ts
#	src/mcp.ts
#	src/permission.ts
#	src/prompt.ts
#	src/session.ts
#	src/snapshot.ts
#	src/subagent.ts
#	src/tools-extra.ts
#	src/tools.ts
#	src/ui/App.tsx
#	test/mcp.test.ts
#	test/prune.test.ts
#	test/session.test.ts
#	test/tools.test.ts
2026-09-21 20:43:35 +07:00

7.4 KiB
Raw Blame History

TODO

Next up. One item, one outcome, verifiable when done.

Longer-term direction lives in ROADMAP.md.


Now

Empty — pick from Known rough edges below.


Next

Empty.


Maintenance

All caught up.


Known rough edges

Not bugs exactly, but things that will bite someone.

  • /clear wipes the terminal scrollback. <Static> output is already committed, so clearing React state alone leaves it on screen. The escape sequence works but takes the user's earlier terminal history with it.
  • Memory has no conflict resolution. Two contradictory notes both persist and both get injected. /memory may merge them, or may keep both.
  • Windows cmd /c differs from bash -lc. A command the model writes for one shell may fail on the other. The prompt states the platform; it does not translate.
  • Permission rules gate the call, not what it does. bash with git * allowed will run a git alias that shells out to anything, and there is no sandbox around the shell. Codex solves this with OS-level isolation — Seatbelt, Landlock, a Windows equivalent — which is three platform-specific implementations and not something to half-ship.
  • The reasoning panel is per-turn, not per-step. Reasoning from an early step stays on screen through later ones until the turn ends.
  • An interrupted command's effects are unknown, and the model is told so. Nothing can know how far a half-run migration got.
  • An installed skill is a stranger's words in your system prompt. The install shows the body first and /skills records the origin, but nothing re-checks it later: a registry that changes a URL's contents affects the next install, not one already on disk.
  • A registry index is trusted for its contents, not its authorship. There are no signatures. registryUrl is the whole trust decision.

Done

Kept for one release, then deleted.

1.0.0 release batch

  • A spend ceiling (maxSpendUsd): checked before each turn, refused at 100% naming the ceiling, warns once at 80%, headless exits non-zero. Unpriced models are not enforced
  • A cheaper subagent model (subagentModel): explore resolves against it, review and worker keep the parent's, /cost splits subagent spend by model id
  • Twenty new built-in tools (41 total) in a new extra set: line edits, filesystem navigation, read-only git extensions, and code/environment reads
  • Twenty new bundled skills (29 total) plus the eleven originals deepened; all moved to src/skills-md/*.md as the Markdown source of truth, embedded at build
  • Ten new data-only plugins: safety refusals on by default (force push, pipe-to-shell, root, env credential writes) and opt-in workflow plugins (conventional commit, tests-first, small diffs, main-branch commits, git config, confirm-delete)
  • Custom slash commands from Markdown files, with $ARGUMENTS/$1 and guarded shell substitution; a custom command never shadows a built-in
  • Auto-loaded external skills, tools, and plugins from ~/.shiro-neko/<kind> and .shiro/<kind>, all data, never code; a bad file is reported and skipped
  • The welcome interface redesigned into a structured dashboard with a session banner, a grouped environment panel, and a meta bar; the input in a two-tone box with a split footer
  • The system prompt advanced: a failure-recovery loop, a delegation policy, compaction awareness
  • The release workflow's dead dry_run input wired: manual dispatch publishes only when unchecked, tag pushes always publish

Post-1.0 — Now / Next / Maintenance (landed)

  • Summarize the pruned span — prune.droppedSpan + session.summarizeDiscarded, injected as Note (retained from compacted history), budgeted 6k excerpt + 3–6 lines, one call per compaction (test/compact.test.ts)
  • Hot-reload an installed entry — Session.updateSkills/updatePlugins + cli.tsx hot-reload, pendingSkills/pendingHost + drainPendingHotReload at turn boundary (test/hot-reload.test.ts)
  • MCP without the schema tax — mcp_list / mcp_inspect / mcp_call phi meta-tools, prompt names-only, permission mcp_call + bindMcpGuard, mcpExpose=phi|direct|auto (test/mcp.test.ts)
  • Derive the tool-name lists — withMeta({ set, mutating }), setsFrom(tools) derives TOOL_SETS/MUTATING_TOOLS/DEFAULT_PERMISSIONS (test/tool-derive.test.ts)
  • Subagent parallelism — task fans out tasks: TaskSpec[] via Promise.all up to 8 (test/subagent-parallel.test.ts)
  • Undo a turn — src/snapshot.ts per-turn capture cap 100, /undo + /redo files+messages (test/undo.test.ts)
  • Pricing source + date — PRICING_VERIFIED_AT='2026-09-09' + source URLs, /cost shows pricing verified: 2026-09-09 (est.)
  • estimateTokens labelling — heuristic doc + (est.) in /cost + ~N est. in context
  • listPaths stale walk — fileChangeSeq on recordBeforeWrite/restoreFiles, @ invalidates paths on seq change
  • MUTATING_TOOLS derivation — BASE_PERMISSIONS + buildDefaults() derives from MUTATING_TOOLS via require('./tools')
  • Unknown toolSets silently dropped — unknownToolSetNames() + startup notice unknown toolSets ignored: … (test/config-toolsets.test.ts, 5028ea6)
  • @ directories — walk({ includeDirs: true }) yields src/ with trailing /, matchPaths ranks dirs before files (test/complete-dirs.test.ts, 4b4ddd0)

Session-feature batch (7 features)

  • /changes — diff the last turn's snapshot: added / modified / deleted, per absolute path, no bash effects (session.lastTurnSummary + snapshot.peek)
  • System-prompt memoization — version counters on notebook/memory/skills/plugins/tools/workspace, cached per version key, hit-rate in /cost (foundation for provider prompt caching)
  • web_search — DuckDuckGo Lite, no API key, 5 results with title/URL/snippet, SSRF-filtered through checkUrl, in the net set
  • /search <query> — full-text over saved sessions, matches transcript strings and tool-input JSON
  • Workspace file list refresh — after a turn that wrote files, re-walk at the boundary so the next prompt shows new paths
  • Per-turn spend cap — maxSpendPerTurn, aborts a step past the line with a notice
  • /fork — clone the session at the last turn boundary to a new saved session; original untouched

Project-driven workflow batch (v8)

  • Workflow policy in the system prompt when the repo tracks progress (TODO.md/ROADMAP.md/docs) — read the task list first, keep it current, spec-first, complete unit tests, verify before done

  • TODO.md + ROADMAP.md loaded as Project tracker instruction blocks (capped 6k each, git root down to cwd)

  • Once-per-session soft nudge when a turn writes files without calling todo_write (skipped when the task list was updated)

  • /workflow — status panel: TODO/ROADMAP/docs presence + line/file counts + nudge state

  • workflow.enabled / workflow.docsDir config keys, merged in config.merge

  • Docs: docs/workflow.md, ROADMAP entry, TODO done-list entry

  • Tests: test/workflow.test.ts (9 tests)

  • A visible escape hatch for the stalled-agent loop: the existing repeat guard now offers a configurable step_back recovery primitive that reads the session loop trace and steers a model circling without progress to stop and change direction