Fix the loop stalling after compaction, add an external registry

The compaction bug, which is the important one:

beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.

The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.

So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.

Registry, via `/registry [list|search|add|remove|installed]`:

Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.

Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.

Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
  red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.

538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 03:07:56 +07:00
parent 8b1895de98
commit 84c60f2022
26 changed files with 1929 additions and 149 deletions
+46
View File
@@ -112,3 +112,49 @@ test('aliases are hidden from the menu but still parse', () => {
expect(matchCommands('/lo')).toEqual([]);
expect(parseCommand('/login').type).toBe('provider');
});
test('/registry with no verb lists everything', () => {
expect(parseCommand('/registry')).toEqual({ type: 'registry', action: 'list' });
expect(parseCommand('/registry list')).toEqual({ type: 'registry', action: 'list' });
});
test('/registry installed asks for what is already here', () => {
expect(parseCommand('/registry installed')).toEqual({ type: 'registry', action: 'installed' });
});
test('/registry search carries the query', () => {
expect(parseCommand('/registry search migration')).toEqual({
type: 'registry',
action: 'search',
arg: 'migration',
});
});
test('/registry add and remove carry the name, and their aliases work', () => {
expect(parseCommand('/registry add migration')).toEqual({ type: 'registry', action: 'add', arg: 'migration' });
expect(parseCommand('/registry install migration')).toEqual({ type: 'registry', action: 'add', arg: 'migration' });
expect(parseCommand('/registry remove migration')).toEqual({ type: 'registry', action: 'remove', arg: 'migration' });
expect(parseCommand('/registry uninstall migration')).toEqual({
type: 'registry',
action: 'remove',
arg: 'migration',
});
});
test('a kind-qualified name survives parsing, since a name can be both', () => {
expect(parseCommand('/registry add plugin:review')).toEqual({
type: 'registry',
action: 'add',
arg: 'plugin:review',
});
});
test('/registry add with no name returns usage rather than fetching anything', () => {
expect(parseCommand('/registry add')).toEqual({ type: 'info', text: 'usage: /registry add <name>' });
expect(parseCommand('/registry remove')).toEqual({ type: 'info', text: 'usage: /registry remove <name>' });
expect(parseCommand('/registry search')).toEqual({ type: 'info', text: 'usage: /registry search <query>' });
});
test('a bare word after /registry is treated as a search', () => {
expect(parseCommand('/registry migration')).toEqual({ type: 'registry', action: 'search', arg: 'migration' });
});
+129
View File
@@ -2,6 +2,9 @@ import { expect, test } from 'bun:test';
import { MockLanguageModelV4, simulateReadableStream } from 'ai/test';
import type { LanguageModelV4CallOptions, LanguageModelV4StreamPart } from '@ai-sdk/provider';
import type { ModelMessage } from 'ai';
import { mkdtempSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { Session, type AgentEvent } from '../src/session';
const usage = {
@@ -173,3 +176,129 @@ test('onChange fires for every history mutation so autosave stays current', asyn
expect(snapshots).toEqual([1, 2]);
});
/** A reasoning model's turn: a reasoning item, then the tool call, both with item ids. */
const reasoningToolStep = (n: number): LanguageModelV4StreamPart[] => [
{ type: 'reasoning-start', id: `r${n}`, providerMetadata: { openai: { itemId: `rs_${n}` } } } as never,
{ type: 'reasoning-delta', id: `r${n}`, delta: 'deciding what to read next' } as never,
{ type: 'reasoning-end', id: `r${n}`, providerMetadata: { openai: { itemId: `rs_${n}` } } } as never,
{ type: 'tool-input-start', id: `c${n}`, toolName: 'read_file' },
{ type: 'tool-input-end', id: `c${n}` },
{
type: 'tool-call',
toolCallId: `c${n}`,
toolName: 'read_file',
input: JSON.stringify({ path: 'big.txt' }),
providerMetadata: { openai: { itemId: `fc_${n}` } },
} as never,
{ type: 'finish', finishReason: { unified: 'tool-calls', raw: 'tool_use' }, usage },
];
function inTempDir<T>(fn: () => Promise<T>): Promise<T> {
const orig = process.cwd();
const dir = mkdtempSync(join(tmpdir(), 'shiro-compact-'));
process.chdir(dir);
return fn().finally(() => {
process.chdir(orig);
rmSync(dir, { recursive: true, force: true });
});
}
/**
* The bug this guards against: compaction used to drop any assistant part whose
* reasoning item it had pruned. On a reasoning model that is every tool call, so
* after the first compaction the model could no longer see what it had already
* run — and kept re-running it until the step limit stopped the turn.
*/
test('a compacted turn still shows the model the tool calls it already made', async () =>
inTempDir(async () => {
await Bun.write(join(process.cwd(), 'big.txt'), 'lorem ipsum dolor sit amet\n'.repeat(1500));
const seen: LanguageModelV4CallOptions[] = [];
let call = 0;
const session = new Session({
compactThreshold: 4000,
maxSteps: 8,
model: new MockLanguageModelV4({
doStream: async (o) => {
seen.push(o);
const n = call++;
return n < 3 ? stream(reasoningToolStep(n)) : stream(text('read it three times'));
},
}),
askApproval: async () => 'deny',
});
const events: string[] = [];
for await (const ev of session.send('read big.txt a few times')) events.push(ev.type);
expect(events).toContain('compacted');
expect(events).toContain('text');
expect(events.at(-1)).toBe('done');
// Four calls, not eight: the loop ended because the model chose to, not
// because maxSteps cut it off.
expect(call).toBe(4);
const shapeOf = (o: LanguageModelV4CallOptions) =>
o.prompt
.filter((m) => m.role !== 'system')
.flatMap((m) => (Array.isArray(m.content) ? (m.content as { type: string }[]).map((p) => p.type) : ['str']));
// Every call after the first has to carry the earlier exchange.
for (const o of seen.slice(1)) {
const shape = shapeOf(o);
expect(shape).toContain('tool-call');
expect(shape).toContain('tool-result');
}
}), 20_000);
test('a compacted turn sends no assistant item reference whose reasoning was pruned', async () =>
inTempDir(async () => {
await Bun.write(join(process.cwd(), 'big.txt'), 'lorem ipsum dolor sit amet\n'.repeat(1500));
const seen: LanguageModelV4CallOptions[] = [];
let call = 0;
const session = new Session({
compactThreshold: 4000,
model: new MockLanguageModelV4({
doStream: async (o) => {
seen.push(o);
const n = call++;
return n < 3 ? stream(reasoningToolStep(n)) : stream(text('done'));
},
}),
askApproval: async () => 'deny',
});
for await (const _ of session.send('read it')) void _;
// An itemId on an assistant part is serialised as `item_reference`, which the
// responses API resolves against a stored item that depends on its reasoning
// item. Send one without that reasoning and the request is a 400. An itemId on
// a tool result is harmless: it goes out as a plain function_call_output.
for (const [i, o] of seen.entries()) {
const assistant = o.prompt.filter((m) => m.role === 'assistant');
const reasoningIds = new Set<string>();
for (const m of assistant) {
if (!Array.isArray(m.content)) continue;
for (const p of m.content as { type: string; providerOptions?: Record<string, Record<string, unknown>> }[]) {
if (p.type !== 'reasoning') continue;
for (const options of Object.values(p.providerOptions ?? {})) {
if (typeof options['itemId'] === 'string') reasoningIds.add(options['itemId']);
}
}
}
for (const m of assistant) {
if (!Array.isArray(m.content)) continue;
for (const p of m.content as { type: string; providerOptions?: Record<string, Record<string, unknown>> }[]) {
if (p.type === 'reasoning') continue;
for (const options of Object.values(p.providerOptions ?? {})) {
const id = options['itemId'];
if (typeof id !== 'string') continue;
expect(reasoningIds.size, `call ${i}: ${p.type} references ${id} with no reasoning item`).toBeGreaterThan(0);
}
}
}
}
}), 20_000);
+9
View File
@@ -21,6 +21,15 @@ export function testHooks(over: Partial<AppHooks> = {}): AppHooks {
saveSession: async () => 'saved',
instructionFiles: () => [],
listPaths: async () => [],
registry: {
list: async () => [],
installed: async () => [],
stage: async () => {
throw new Error('no registry in tests unless a test provides one');
},
install: async () => 'installed',
remove: async () => 'removed',
},
initPrompt: 'write AGENTS.md',
history: [],
recordPrompt: () => {},
+100 -26
View File
@@ -1,10 +1,24 @@
import { expect, test } from 'bun:test';
import type { ModelMessage } from 'ai';
import { dropOrphanedItems, dropOrphanedResults, prunePreservingItems } from '../src/prune';
import { detachOrphanedItems, dropOrphanedResults, prunePreservingItems } from '../src/prune';
const kinds = (messages: ModelMessage[]) =>
messages.map((m) => (Array.isArray(m.content) ? `${m.role}:${m.content.map((p) => p.type).join('+')}` : m.role));
/** Every provider itemId in a message tree, which is what the repair strips. */
const itemIds = (messages: ModelMessage[]): string[] => {
const found: string[] = [];
for (const m of messages) {
if (!Array.isArray(m.content)) continue;
for (const p of m.content as { providerOptions?: Record<string, Record<string, unknown>> }[]) {
for (const options of Object.values(p.providerOptions ?? {})) {
if (typeof options['itemId'] === 'string') found.push(options['itemId']);
}
}
}
return found;
};
/** An assistant turn as the OpenAI responses API returns it. */
const reasoningTurn = (rs: string, msg: string, text = 'answer'): ModelMessage => ({
role: 'assistant',
@@ -28,24 +42,33 @@ const toolTurn = (rs: string, call: string): ModelMessage => ({
],
});
test('a message left without its reasoning item is dropped', () => {
/**
* A part carrying an itemId is serialised as `{ type: 'item_reference', id }`,
* which depends on the stored reasoning item. Stripping the id sends the same
* content inline instead, so the turn survives without the dependency.
*/
test('a message left without its reasoning item keeps its text and loses its item id', () => {
const before = [{ role: 'user' as const, content: 'q' }, reasoningTurn('rs_1', 'msg_1')];
const after = [{ role: 'user' as const, content: 'q' }, { role: 'assistant' as const, content: [{ type: 'text' as const, text: 'answer', providerOptions: { openai: { itemId: 'msg_1' } } }] }];
const after: ModelMessage[] = [
{ role: 'user', content: 'q' },
{ role: 'assistant', content: [{ type: 'text', text: 'answer', providerOptions: { openai: { itemId: 'msg_1' } } }] },
];
const cleaned = dropOrphanedItems(before, after);
expect(JSON.stringify(cleaned)).not.toContain('msg_1');
expect(kinds(cleaned)).toEqual(['user']);
const cleaned = detachOrphanedItems(before, after);
expect(itemIds(cleaned)).toEqual([]);
expect(kinds(cleaned)).toEqual(['user', 'assistant:text']);
expect(JSON.stringify(cleaned)).toContain('answer');
});
test('a tool call left without its reasoning item is dropped too', () => {
test('a tool call left without its reasoning item survives, detached', () => {
const before = [{ role: 'user' as const, content: 'q' }, toolTurn('rs_1', 'fc_1')];
const after = [
{ role: 'user' as const, content: 'q' },
const after: ModelMessage[] = [
{ role: 'user', content: 'q' },
{
role: 'assistant' as const,
role: 'assistant',
content: [
{
type: 'tool-call' as const,
type: 'tool-call',
toolCallId: 'tc1',
toolName: 'grep',
input: { pattern: 'x' },
@@ -55,12 +78,60 @@ test('a tool call left without its reasoning item is dropped too', () => {
},
];
expect(JSON.stringify(dropOrphanedItems(before, after))).not.toContain('fc_1');
const cleaned = detachOrphanedItems(before, after);
expect(itemIds(cleaned)).toEqual([]);
// The call itself has to stay, or its result is orphaned and the model loses
// any record of what it already ran.
expect(JSON.stringify(cleaned)).toContain('tc1');
expect(kinds(cleaned)).toEqual(['user', 'assistant:tool-call']);
});
test('an empty providerOptions object is removed rather than left behind', () => {
const before = [toolTurn('rs_1', 'fc_1')];
const after: ModelMessage[] = [
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolCallId: 'tc1',
toolName: 'grep',
input: { pattern: 'x' },
providerOptions: { openai: { itemId: 'fc_1' } },
},
],
},
];
const part = (detachOrphanedItems(before, after)[0]!.content as Record<string, unknown>[])[0]!;
expect('providerOptions' in part).toBe(false);
});
test('other provider options are kept when the item id is stripped', () => {
const before: ModelMessage[] = [
{
role: 'assistant',
content: [
{ type: 'reasoning', text: 't', providerOptions: { openai: { itemId: 'rs_1' } } },
{ type: 'text', text: 'a', providerOptions: { openai: { itemId: 'msg_1', phase: 'final' } } },
],
},
];
const after: ModelMessage[] = [
{
role: 'assistant',
content: [{ type: 'text', text: 'a', providerOptions: { openai: { itemId: 'msg_1', phase: 'final' } } }],
},
];
const json = JSON.stringify(detachOrphanedItems(before, after));
expect(json).not.toContain('msg_1');
expect(json).toContain('final');
});
test('a turn whose reasoning survived is left alone', () => {
const messages = [{ role: 'user' as const, content: 'q' }, reasoningTurn('rs_1', 'msg_1')];
expect(dropOrphanedItems(messages, messages)).toEqual(messages);
expect(detachOrphanedItems(messages, messages)).toEqual(messages);
});
test('nothing is touched when no reasoning was removed', () => {
@@ -68,13 +139,13 @@ test('nothing is touched when no reasoning was removed', () => {
{ role: 'user', content: 'q' },
{ role: 'assistant', content: 'plain answer' },
];
expect(dropOrphanedItems(before, before)).toEqual(before);
expect(detachOrphanedItems(before, before)).toEqual(before);
});
test('parts with no provider itemId are always kept', () => {
test('parts with no provider itemId are returned unchanged', () => {
const before = [reasoningTurn('rs_1', 'msg_1')];
const after: ModelMessage[] = [{ role: 'assistant', content: [{ type: 'text', text: 'no item id here' }] }];
expect(dropOrphanedItems(before, after)).toEqual(after);
expect(detachOrphanedItems(before, after)).toEqual(after);
});
test('user and tool messages are never affected', () => {
@@ -83,10 +154,10 @@ test('user and tool messages are never affected', () => {
{ role: 'user', content: 'q' },
{ role: 'tool', content: [{ type: 'tool-result', toolCallId: 't1', toolName: 'grep', output: { type: 'text', value: 'hit' } }] },
];
expect(dropOrphanedItems(before, after)).toEqual(after);
expect(detachOrphanedItems(before, after)).toEqual(after);
});
test('one orphaned turn does not take a healthy one with it', () => {
test('one orphaned turn does not detach a healthy one with it', () => {
const before = [
{ role: 'user' as const, content: 'q1' },
reasoningTurn('rs_1', 'msg_1', 'old answer'),
@@ -100,14 +171,15 @@ test('one orphaned turn does not take a healthy one with it', () => {
reasoningTurn('rs_2', 'msg_2', 'new answer'),
];
const cleaned = dropOrphanedItems(before, after);
const cleaned = detachOrphanedItems(before, after);
const json = JSON.stringify(cleaned);
expect(json).not.toContain('msg_1');
expect(json).toContain('old answer');
expect(json).toContain('msg_2');
expect(json).toContain('rs_2');
});
test('prunePreservingItems leaves no orphan behind on a real prune', () => {
test('prunePreservingItems leaves no item reference behind on a real prune', () => {
const messages: ModelMessage[] = [];
for (let i = 0; i < 6; i++) {
messages.push({ role: 'user', content: `question ${i} ${'x'.repeat(3000)}` });
@@ -121,12 +193,11 @@ test('prunePreservingItems leaves no orphan behind on a real prune', () => {
emptyMessages: 'remove',
});
// Every surviving text part must either have no item id or belong to a turn
// whose reasoning also survived. Since reasoning: 'all' removes them all, no
// itemId-bearing assistant part may remain.
const survivingIds = JSON.stringify(pruned);
for (let i = 0; i < 6; i++) expect(survivingIds).not.toContain(`msg_${i}`);
// reasoning: 'all' removes every reasoning item, so no surviving part may still
// reference one. The text itself stays: that is the model's memory of the turn.
expect(itemIds(pruned)).toEqual([]);
expect(pruned.filter((m) => m.role === 'user')).toHaveLength(6);
expect(JSON.stringify(pruned)).toContain('answer 5');
});
test('prunePreservingItems is a no-op when nothing needs pruning', () => {
@@ -150,7 +221,10 @@ test('a provider other than openai is handled the same way', () => {
const after: ModelMessage[] = [
{ role: 'assistant', content: [{ type: 'text', text: 'a', providerOptions: { someProvider: { itemId: 'm1' } } }] },
];
expect(dropOrphanedItems(before, after)).toEqual([]);
const cleaned = detachOrphanedItems(before, after);
expect(itemIds(cleaned)).toEqual([]);
expect(JSON.stringify(cleaned)).toContain('"text":"a"');
});
/** The assistant tool-call plus the tool message answering it, as one exchange. */
+266
View File
@@ -0,0 +1,266 @@
import { expect, test } from 'bun:test';
import { render } from 'ink-testing-library';
import React from 'react';
import { MockLanguageModelV4, simulateReadableStream } from 'ai/test';
import { Session } from '../src/session';
import { App, createApprovalBridge, type AppHooks } from '../src/ui/App';
import { InstallPrompt, RegistryPanel, type RegistryRow } from '../src/ui/Panels';
import { testHooks } from './helpers';
const usage = { inputTokens: { total: 3, noCache: 3, cacheRead: 0, cacheWrite: 0 }, outputTokens: { total: 1 } } as any;
const model = new MockLanguageModelV4({
doStream: async () =>
({
stream: simulateReadableStream({
chunks: [
{ type: 'text-start', id: '0' },
{ type: 'text-delta', id: '0', delta: 'reply' },
{ type: 'text-end', id: '0' },
{ type: 'finish', finishReason: { unified: 'stop', raw: 'stop' }, usage },
],
chunkDelayInMs: null,
initialDelayInMs: null,
}),
}) as any,
});
const rows: RegistryRow[] = [
{ name: 'migration', kind: 'skill', description: 'Write a database migration' },
{ name: 'no-secrets', kind: 'plugin', description: 'Refuses credential writes', installed: true },
];
const wait = (ms: number) => new Promise((r) => setTimeout(r, ms));
function mount(over: Partial<AppHooks> = {}) {
const bridge = createApprovalBridge();
const session = new Session({ model, askApproval: bridge.ask });
const app = render(<App session={session} bridge={bridge} header="hdr" hooks={testHooks(over)} />);
return { app, session };
}
async function run(app: ReturnType<typeof render>, command: string, settle = 400) {
for (const ch of command) {
app.stdin.write(ch);
await wait(30);
}
app.stdin.write('\r');
await wait(settle);
}
test('the registry panel marks skill and plugin differently and flags what is installed', () => {
const app = render(<RegistryPanel rows={rows} hint="2 available" />);
const frame = app.lastFrame() ?? '';
expect(frame).toContain('migration');
expect(frame).toContain('no-secrets');
expect(frame).toContain('installed');
expect(frame).toContain('S skill P plugin');
app.unmount();
});
test('an empty registry panel says so rather than rendering an empty box', () => {
const app = render(<RegistryPanel rows={[]} />);
expect(app.lastFrame()).toContain('nothing found');
app.unmount();
});
test('the install prompt shows the body and says what a skill actually is', () => {
const app = render(
<InstallPrompt
name="migration"
kind="skill"
url="https://example.com/m.md"
preview={'Migrations live in db/migrations.'}
/>,
);
const frame = app.lastFrame() ?? '';
expect(frame).toContain('install skill "migration"?');
expect(frame).toContain('https://example.com/m.md');
expect(frame).toContain('Migrations live in db/migrations.');
expect(frame).toContain('joins your system prompt');
app.unmount();
});
test('the install prompt says a plugin is data, not code', () => {
const app = render(
<InstallPrompt name="no-secrets" kind="plugin" url="https://e.com/p.json" preview="- denies write_file" />,
);
const frame = app.lastFrame() ?? '';
expect(frame).toContain('nothing here is executed');
expect(frame).toContain('denies write_file');
app.unmount();
});
test('a long body is truncated with a count rather than flooding the screen', () => {
const preview = Array.from({ length: 40 }, (_, i) => `line ${i}`).join('\n');
const app = render(<InstallPrompt name="x" kind="skill" url="https://e.com/x.md" preview={preview} lines={5} />);
const frame = app.lastFrame() ?? '';
expect(frame).toContain('line 4');
expect(frame).not.toContain('line 20');
expect(frame).toContain('35 more lines');
app.unmount();
});
test('/registry lists the index', async () => {
const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } });
await wait(150);
await run(app, '/registry');
const frame = app.lastFrame() ?? '';
expect(frame).toContain('migration');
expect(frame).toContain('2 of 2 available');
app.unmount();
}, 20_000);
test('/registry search narrows the list', async () => {
const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } });
await wait(150);
await run(app, '/registry search migra');
const frame = app.lastFrame() ?? '';
expect(frame).toContain('migration');
expect(frame).not.toContain('no-secrets');
expect(frame).toContain('1 of 2');
app.unmount();
}, 20_000);
test('a failing index surfaces the error instead of an empty panel', async () => {
const { app } = mount({
registry: {
...testHooks().registry,
list: async () => {
throw new Error('the registry index is not valid JSON');
},
},
});
await wait(150);
await run(app, '/registry');
expect(app.lastFrame()).toContain('not valid JSON');
app.unmount();
}, 20_000);
test('/registry add stages, shows the body, and installs only on y', async () => {
const installed: string[] = [];
const { app } = mount({
registry: {
...testHooks().registry,
stage: async (name) => ({
row: { name, kind: 'skill', description: 'd' },
url: 'https://example.com/m.md',
preview: 'Migrations live in db/migrations.',
}),
install: async (name) => {
installed.push(name);
return `installed skill ${name}`;
},
},
});
await wait(150);
await run(app, '/registry add migration');
expect(app.lastFrame()).toContain('install skill "migration"?');
expect(installed).toEqual([]);
app.stdin.write('y');
await wait(400);
expect(installed).toEqual(['migration']);
expect(app.lastFrame()).toContain('installed skill migration');
app.unmount();
}, 20_000);
test('n cancels the install and nothing is written', async () => {
const installed: string[] = [];
const { app } = mount({
registry: {
...testHooks().registry,
stage: async (name) => ({
row: { name, kind: 'skill', description: 'd' },
url: 'https://example.com/m.md',
preview: 'body',
}),
install: async (name) => {
installed.push(name);
return 'installed';
},
},
});
await wait(150);
await run(app, '/registry add migration');
app.stdin.write('n');
await wait(400);
expect(installed).toEqual([]);
expect(app.lastFrame()).toContain('install cancelled: migration');
app.unmount();
}, 20_000);
test('a staging failure never reaches the confirmation prompt', async () => {
const { app } = mount({
registry: {
...testHooks().registry,
stage: async () => {
throw new Error('no registry entry named "nope"');
},
},
});
await wait(150);
await run(app, '/registry add nope');
const frame = app.lastFrame() ?? '';
expect(frame).toContain('no registry entry named');
expect(frame).not.toContain('install skill');
app.unmount();
}, 20_000);
test('/registry remove reports what it removed', async () => {
const { app } = mount({
registry: { ...testHooks().registry, remove: async (name) => `removed skill ${name}` },
});
await wait(150);
await run(app, '/registry remove migration');
expect(app.lastFrame()).toContain('removed skill migration');
app.unmount();
}, 20_000);
test('/registry installed shows what is already here', async () => {
const { app } = mount({
registry: { ...testHooks().registry, installed: async () => [rows[1]!] },
});
await wait(150);
await run(app, '/registry installed');
const frame = app.lastFrame() ?? '';
expect(frame).toContain('no-secrets');
expect(frame).toContain('1 from the registry');
app.unmount();
}, 20_000);
test('esc dismisses the registry panel', async () => {
const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } });
await wait(150);
await run(app, '/registry');
expect(app.lastFrame()).toContain('migration');
app.stdin.write('\u001B');
await wait(300);
expect(app.lastFrame()).not.toContain('S skill P plugin');
app.unmount();
}, 20_000);
+314
View File
@@ -0,0 +1,314 @@
import { afterEach, beforeEach, expect, test } from 'bun:test';
import { mkdtempSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import {
DEFAULT_INDEX_URL,
fetchIndex,
install,
loadInstalledPlugins,
manifestToPlugin,
parseIndex,
parseManifest,
pluginsDir,
searchEntries,
skillsDir,
stage,
uninstall,
type RegistryEntry,
} from '../src/registry';
let home: string;
let origHome: string | undefined;
beforeEach(() => {
origHome = process.env['SHIRO_HOME'];
home = mkdtempSync(join(tmpdir(), 'shiro-reg-'));
process.env['SHIRO_HOME'] = home;
});
afterEach(() => {
if (origHome === undefined) delete process.env['SHIRO_HOME'];
else process.env['SHIRO_HOME'] = origHome;
rmSync(home, { recursive: true, force: true });
});
const index = (over: Record<string, unknown> = {}) =>
JSON.stringify({
skills: [{ name: 'migration', description: 'Write a database migration', url: 'https://example.com/m.md' }],
plugins: [{ name: 'no-secrets', description: 'Refuses credential writes', url: 'https://example.com/p.json' }],
...over,
});
const manifest = (over: Record<string, unknown> = {}) =>
JSON.stringify({
name: 'no-secrets',
description: 'Refuses credential writes',
deny: [{ tools: ['write_file', 'edit_file'], pathPattern: '\\.env$', reason: 'refusing to write a .env file' }],
...over,
});
test('the default index is an https URL', () => {
expect(DEFAULT_INDEX_URL.startsWith('https://')).toBe(true);
});
test('an index parses into entries tagged with their kind', () => {
const entries = parseIndex(index());
expect(entries).toHaveLength(2);
expect(entries.find((e) => e.name === 'migration')?.kind).toBe('skill');
expect(entries.find((e) => e.name === 'no-secrets')?.kind).toBe('plugin');
});
test('a malformed index is reported rather than half-loaded', () => {
expect(() => parseIndex('not json')).toThrow(/not valid JSON/);
expect(() => parseIndex(JSON.stringify({ skills: [{ name: 'x' }] }))).toThrow(/malformed/);
});
test('a name that could escape its directory is rejected', () => {
for (const name of ['../evil', 'a/b', 'UPPER', '.hidden', 'with space']) {
const bad = JSON.stringify({ skills: [{ name, description: 'd', url: 'https://e.com/x.md' }] });
expect(() => parseIndex(bad), name).toThrow(/malformed/);
}
});
test('a non-http url is rejected by the schema', () => {
const bad = JSON.stringify({ skills: [{ name: 'x', description: 'd', url: 'file:///etc/passwd' }] });
expect(() => parseIndex(bad)).toThrow(/malformed/);
});
test('a duplicate name within one kind keeps the first', () => {
const dup = JSON.stringify({
skills: [
{ name: 'x', description: 'first', url: 'https://e.com/1.md' },
{ name: 'x', description: 'second', url: 'https://e.com/2.md' },
],
});
const entries = parseIndex(dup);
expect(entries).toHaveLength(1);
expect(entries[0]?.description).toBe('first');
});
test('the same name may exist as both a skill and a plugin', () => {
const both = JSON.stringify({
skills: [{ name: 'review', description: 's', url: 'https://e.com/s.md' }],
plugins: [{ name: 'review', description: 'p', url: 'https://e.com/p.json' }],
});
expect(parseIndex(both)).toHaveLength(2);
});
test('search matches name and description, case-insensitively', () => {
const entries = parseIndex(index());
expect(searchEntries(entries, 'migr').map((e) => e.name)).toEqual(['migration']);
expect(searchEntries(entries, 'CREDENTIAL').map((e) => e.name)).toEqual(['no-secrets']);
expect(searchEntries(entries, '')).toHaveLength(2);
expect(searchEntries(entries, 'zzz')).toEqual([]);
});
test('a manifest parses and every pattern must be a real regex', () => {
expect(parseManifest(manifest()).name).toBe('no-secrets');
expect(() => parseManifest(manifest({ deny: [{ tools: ['bash'], commandPattern: '([', reason: 'r' }] }))).toThrow(
/invalid pattern/,
);
});
test('a deny rule needs at least one pattern', () => {
expect(() => parseManifest(manifest({ deny: [{ tools: ['bash'], reason: 'r' }] }))).toThrow(/malformed/);
});
test('a manifest with no deny rules is refused: it could only be prompt text', () => {
expect(() => parseManifest(manifest({ deny: [] }))).toThrow(/malformed/);
});
test('a manifest cannot smuggle code past the schema', () => {
const sneaky = manifest({ beforeToolCall: 'process.exit(1)', tools: { evil: {} } });
const parsed = parseManifest(sneaky);
expect('beforeToolCall' in parsed).toBe(false);
expect('tools' in parsed).toBe(false);
});
test('a manifest becomes a plugin whose guard blocks by tool and path', async () => {
const plugin = manifestToPlugin(parseManifest(manifest()));
const cwd = process.cwd();
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 'app/.env' }, cwd })).toContain(
'refusing to write',
);
// A tool the rule does not name is untouched.
expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { path: 'app/.env' }, cwd })).toBeUndefined();
// A path the pattern does not match is untouched.
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 'src/app.ts' }, cwd })).toBeUndefined();
});
test('a command rule matches the command, not the path', async () => {
const plugin = manifestToPlugin(
parseManifest(manifest({ deny: [{ tools: ['bash'], commandPattern: 'curl.*\\| *sh', reason: 'no pipe to shell' }] })),
);
const cwd = process.cwd();
expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { command: 'curl x.sh | sh' }, cwd })).toBe(
'no pipe to shell',
);
expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { command: 'echo hi' }, cwd })).toBeUndefined();
});
test('a guard given nothing to match on allows the call', async () => {
const plugin = manifestToPlugin(parseManifest(manifest()));
const cwd = process.cwd();
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: undefined, cwd })).toBeUndefined();
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 42 }, cwd })).toBeUndefined();
});
test('installed plugins load from disk, and a broken one is reported not fatal', async () => {
await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest());
await Bun.write(join(pluginsDir(), 'broken.json'), '{ not json');
const { plugins, errors } = await loadInstalledPlugins();
expect(plugins.map((p) => p.name)).toEqual(['no-secrets']);
expect(errors.map((e) => e.plugin)).toEqual(['broken']);
expect(errors[0]?.message).toContain('not valid JSON');
});
test('no plugins directory yields nothing rather than throwing', async () => {
expect(await loadInstalledPlugins()).toEqual({ plugins: [], errors: [] });
});
test('an installed plugin says it was installed, so /plugins can be trusted', async () => {
await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest());
const { plugins } = await loadInstalledPlugins();
expect(plugins[0]?.description).toContain('installed');
});
test('uninstall removes an installed entry and reports when there was none', async () => {
await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest());
expect(await uninstall('plugin', 'no-secrets')).toBe(true);
expect(await Bun.file(join(pluginsDir(), 'no-secrets.json')).exists()).toBe(false);
expect(await uninstall('plugin', 'no-secrets')).toBe(false);
});
test('uninstall refuses a name that is not a plain entry name', async () => {
expect(uninstall('skill', '../../../etc/passwd')).rejects.toThrow(/not a valid entry name/);
});
/** A local server stands in for the registry: no network, real HTTP. */
function serve(routes: Record<string, { body: string; status?: number }>) {
return Bun.serve({
port: 0,
fetch(req) {
const path = new URL(req.url).pathname;
const hit = routes[path];
if (!hit) return new Response('not found', { status: 404 });
return new Response(hit.body, { status: hit.status ?? 200 });
},
});
}
test('fetchIndex reads an index over http', async () => {
const server = serve({ '/index.json': { body: index() } });
try {
const entries = await fetchIndex(`${server.url}index.json`);
expect(entries.map((e) => e.name).sort()).toEqual(['migration', 'no-secrets']);
} finally {
server.stop(true);
}
});
test('an index that returns an error status is reported with the status', async () => {
const server = serve({ '/index.json': { body: 'nope', status: 503 } });
try {
expect(fetchIndex(`${server.url}index.json`)).rejects.toThrow(/503/);
} finally {
server.stop(true);
}
});
const SKILL_BODY = `---
name: migration
description: Write a database migration
---
Migrations live in db/migrations and are never renumbered.`;
test('stage returns what would be written without writing it', async () => {
const server = serve({ '/m.md': { body: SKILL_BODY } });
try {
const entry: RegistryEntry = {
kind: 'skill',
name: 'migration',
description: 'Write a database migration',
url: `${server.url}m.md`,
};
const staged = await stage(entry);
expect(staged.path).toBe(join(skillsDir(), 'migration.md'));
expect(staged.preview).toContain('never renumbered');
expect(await Bun.file(staged.path).exists()).toBe(false);
} finally {
server.stop(true);
}
});
test('install writes the skill where the loader will find it', async () => {
const server = serve({ '/m.md': { body: SKILL_BODY } });
try {
const installed = await install({
kind: 'skill',
name: 'migration',
description: 'd',
url: `${server.url}m.md`,
});
expect(installed.path).toBe(join(skillsDir(), 'migration.md'));
expect(await Bun.file(installed.path).text()).toContain('name: migration');
const { loadSkills } = await import('../src/skills');
const loaded = await loadSkills(home);
const skill = loaded.find((s) => s.name === 'migration');
expect(skill?.origin).toBe('registry');
} finally {
server.stop(true);
}
});
test('a body whose name disagrees with the index is refused', async () => {
const server = serve({ '/m.md': { body: SKILL_BODY } });
try {
expect(
stage({ kind: 'skill', name: 'something-else', description: 'd', url: `${server.url}m.md` }),
).rejects.toThrow(/calls itself "migration"/);
} finally {
server.stop(true);
}
});
test('a skill with no frontmatter is refused rather than installed as prose', async () => {
const server = serve({ '/m.md': { body: 'just some text, no frontmatter' } });
try {
expect(stage({ kind: 'skill', name: 'migration', description: 'd', url: `${server.url}m.md` })).rejects.toThrow(
/frontmatter/,
);
} finally {
server.stop(true);
}
});
test('a plugin manifest is validated before it can be staged', async () => {
const server = serve({
'/good.json': { body: manifest() },
'/bad.json': { body: manifest({ deny: [{ tools: ['bash'], commandPattern: '([', reason: 'r' }] }) },
});
try {
const staged = await stage({
kind: 'plugin',
name: 'no-secrets',
description: 'd',
url: `${server.url}good.json`,
});
expect(staged.path).toBe(join(pluginsDir(), 'no-secrets.json'));
expect(staged.preview).toContain('denies write_file');
expect(
stage({ kind: 'plugin', name: 'no-secrets', description: 'd', url: `${server.url}bad.json` }),
).rejects.toThrow(/invalid pattern/);
} finally {
server.stop(true);
}
});
+34
View File
@@ -106,6 +106,40 @@ test('the status bar reports model, agent, thinking, context, and spend', () =>
app.unmount();
});
test('a context limit turns the raw token count into a percentage', () => {
const app = render(
<StatusBar
model="gpt-5"
agent="default"
thinking="medium"
contextTokens={60_000}
contextLimit={120_000}
cost="$0.10"
toolCount={14}
/>,
);
const frame = app.lastFrame() ?? '';
expect(frame).toContain('50% ctx');
expect(frame).not.toContain('60000');
app.unmount();
});
test('the context percentage is capped at 100 rather than running over', () => {
const app = render(
<StatusBar
model="m"
agent="a"
thinking="t"
contextTokens={300_000}
contextLimit={120_000}
cost="$1"
toolCount={1}
/>,
);
expect(app.lastFrame()).toContain('100% ctx');
app.unmount();
});
test('the info panel renders a markdown body', () => {
const app = render(<InfoPanel title="tools" hint="7 offered" lines={'- `read_file`\n- `bash`'} />);
const frame = app.lastFrame() ?? '';