Fix the loop stalling after compaction, add an external registry

The compaction bug, which is the important one:

beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.

The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.

So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.

Registry, via `/registry [list|search|add|remove|installed]`:

Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.

Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.

Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
  red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.

538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 03:07:56 +07:00
parent 8b1895de98
commit 84c60f2022
26 changed files with 1929 additions and 149 deletions
+25 -9
View File
@@ -4,8 +4,17 @@ A plugin extends the agent in four ways: it can add tools, mark tools auto-appro
a tool call before it runs, and append to the system prompt. It can also run something after
each turn.
Plugins are compiled into the binary. Loading them from disk is deliberately not supported
yet — see [ROADMAP.md](../ROADMAP.md).
Two kinds exist, and only one can contain code:
- **Builtin** plugins are compiled into the binary and may do anything in the interface below.
- **Installed** plugins come from a registry as a JSON manifest of refusal rules. They are
data: the guard evaluating them is compiled code, identical for every install. See
[registry](registry.md).
Loading TypeScript from disk or a URL is deliberately not supported. A plugin that can block
tool calls can also lie about blocking them, and one that could execute could read every file
the agent can read. That is a sandbox problem, not a loader problem — see
[ROADMAP.md](../ROADMAP.md).
## Enabling
@@ -13,9 +22,12 @@ yet — see [ROADMAP.md](../ROADMAP.md).
{ "plugins": ["guard", "time"] }
```
That is also the default when the field is absent. `--no-plugins` disables all of them,
including the guard. `/plugins` lists what is active and reports any name that did not
resolve.
That is also the default when the field is absent, and it lists **builtin** plugins only.
Installed plugins are always active once present, because installing one was the decision to
enable it; remove it with `/registry remove <name>`.
`--no-plugins` disables everything, builtin and installed, including the guard. `/plugins`
lists what is active, marks installed entries, and reports any name that did not resolve.
## The interface
@@ -92,7 +104,11 @@ intrusive, but it is genuinely useful when a turn takes minutes.
## Writing one
Plugins live in `src/plugins-builtin.ts` and are registered in `BUILTIN_PLUGINS`.
A refusal rule is usually better as an installed manifest: no rebuild, and nothing to review.
See [registry](registry.md) for the manifest shape. Reach for a builtin only when the plugin
needs to contribute a tool or run something after a turn.
Builtin plugins live in `src/plugins-builtin.ts` and are registered in `BUILTIN_PLUGINS`.
```ts
export const noSecretsPlugin: Plugin = {
@@ -122,6 +138,6 @@ refusal it was never told about and tries to route around it.
## Ordering
Plugins run in the order they are enabled. The first `beforeToolCall` to block wins;
later hooks are not consulted. `afterTurn` runs every hook, and one throwing does not stop
the rest.
Builtin plugins run first, in the order they are enabled, then installed ones. The first
`beforeToolCall` to block wins; later hooks are not consulted. `afterTurn` runs every hook, and
one throwing does not stop the rest.