Files
shiro-neko/docs/mcp.md
T
Muhammad Zakir Ramadhan 9b978fdbe1 Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
2026-09-03 09:26:37 +07:00

4.7 KiB

MCP

Model Context Protocol servers contribute tools. Configure them in ~/.shiro-neko/config.json and they appear alongside the builtins.

Configuration

{
  "mcpServers": {
    "fs": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-filesystem", "."]
    },
    "db": {
      "command": "python",
      "args": ["-m", "my_mcp_server"],
      "env": { "DATABASE_URL": "postgres://localhost/dev" },
      "cwd": "/home/you/tools"
    },
    "api": {
      "url": "http://localhost:3000/mcp",
      "type": "http",
      "headers": { "Authorization": "Bearer local-dev-token" }
    }
  }
}

stdio servers take command, and optionally args, env, cwd. The process is spawned at startup and closed on exit. env is merged over the inherited environment, so a server inherits your PATH unless you replace it.

Remote servers take url, and optionally type (http or sse, default http) and headers.

A token in headers sits in config.json in plain text, same as apiKey. For anything beyond a local dev token, prefer a stdio server that reads its own credential from the environment.

Startup cost

Servers connect in parallel, so the slowest one sets how long startup takes rather than the sum of them. npx -y some-server re-resolves the package on each launch; installing it and calling the binary directly is usually the difference between a noticeable wait and none.

--no-mcp skips them all, which is also the quickest way to tell whether a slow start is MCP or something else.

Naming

Tools arrive as mcp__<server>__<tool>. A server named fs exposing read_file becomes mcp__fs__read_file.

The namespace is not cosmetic. Two servers both exposing search would otherwise silently shadow each other, and the model would call one believing it was the other.

Approval

Every MCP tool requires approval on every call. They are third-party code with unknown side effects, so they are treated like bash rather than like read_file. a whitelists one tool for the session.

--yolo skips these prompts, as it does for the builtins. Plugin guards still apply.

Failure handling

A server that fails to start is reported and the session continues:

shiro-neko 0.1.0-beta.3  openai/gpt-5  session 0193ab2c
mcp: 4 tools
mcp db failed: spawn python ENOENT

Nothing else is lost — the other servers still load, the builtins still work. A missing Python interpreter should not stop you from editing a file.

--no-mcp skips them all.

Inspecting

/tools lists everything offered this turn, MCP tools included. The system prompt describes them as a group:

- mcp__api__query, mcp__fs__read_file: from MCP servers, named mcp__<server>__<tool>.
  Each needs approval; read its own description before calling.

Their individual descriptions come from the server, so that is what the model reads before calling one.

Cost

Each tool adds its name, description, and JSON schema to every request. The built-ins average 548 bytes; MCP tools vary with how verbose the server's schema is. A server exposing twenty tools costs roughly 2,750 tokens per turn, sent whether or not the model uses any of them.

MCP tools are not covered by toolSets — that budget only governs the built-ins. There is no per-server switch either, so the choice is a server or no server, and --no-mcp for all of them. If one exposes many tools you never use, a narrower server is worth finding or writing.

/tools shows the count both ways:

tools
26 offered this turn of 26 registered

A gap between the two numbers means a tool set or a read-only agent variant is withholding something. MCP tools never appear in that gap.

Writing a server

Any MCP-compliant server works. A minimal stdio one needs three methods: initialize, tools/list, and tools/call. The test suite includes one at test/fixtures/mcp-stub.ts — about 50 lines, and useful as a starting point.

The suite runs it as a real subprocess rather than mocking the transport, because the parts that break in practice are the handshake and the framing, and a mock asserts neither.

Debugging a server

A server that starts but returns nothing useful is the harder case. In order of speed:

  1. /tools — did the tools arrive at all? A server with no tools is a tools/list problem.
  2. shiro -p "call mcp__x__y with ..." --json --yolo — the exact tool-call input and tool-result output, one JSON object per line.
  3. Run the server by hand: echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | your-server. If that is wrong, nothing above it can be right.

For an HTTP server, curl -X POST $URL -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' answers the same question without shiro in the way.