Files
shiro-neko/docs/configuration.md
T
Muhammad Zakir Ramadhan 7fd578e13b Gate tool calls per command and path, not per tool name
Approval was a list of tool names: `bash` needed it, `read_file` did not. That
fails in a specific way. `bash` covers `git status` and `rm -rf` equally, so a
user working through a batch presses `a` — always allow — on the first prompt and
every later command runs unasked, including the one they would have refused. The
gate was strongest when it mattered least and gone by the time it mattered.

Rules now match the *subject* of a call: the command for `bash`, the path for a
file tool, the pattern for a search.

    "permission": {
      "bash": { "*": "ask", "git *": "allow", "rm *": "deny" },
      "edit_file": { "*": "deny", "src/generated/*": "allow" }
    }

Plain last-match-wins, with no special case for deny. An earlier version made
deny win wherever it sat, on the theory that a refusal should be impossible to
undo by accident, and it made default-deny-with-exceptions unexpressible — which
is the shape a careful user actually writes, and the same shape as `*.env` denied
while `*.env.example` is allowed. Refusals that must never be configurable stay
in the guard plugin, which runs ahead of this and which --yolo cannot reach.

Three behaviours fall out of it:

- `always` grants the pattern the tool suggests, not the tool. Approving
  `git status` runs `git log` unprompted and still asks about `npm publish`.
- `.env`, `.env.*`, and `.pem` are denied on read by default. Not gated, refused:
  a secret that reaches the context is on the wire and in the session file, and
  there is no taking it back. `.env.example` stays allowed.
- A call repeated identically three times in one turn asks even when allowed. A
  model repeating itself is not making progress, and `bash: allow` is a statement
  about which commands are safe rather than permission to loop.

The prompt now says which rule matched and what `always` would grant:

    bash wants to run
    git status --porcelain
    y allow once | a always allow bash git * | n deny

A typo in a decision string is dropped at parse time, leaving the tool on its
default. Treating an unparseable value as `allow` would mean one misspelling
silently removing the gate.

Written after surveying Claude Code, Codex, opencode, and phi. Three of the four
had already moved to per-pattern rules; the credential deny and the repeat guard
come from opencode directly. ROADMAP records what was deliberately not taken and
why — OS sandboxing needs three platform implementations and is worse than
nothing if half-built, and opencode's own docs say LSP integration is often not a
net positive.

581 tests, up from 572. The engine is a pure function tested on its own, and the
loop is tested through a real Session: an allowed pattern never prompts, a denied
one never executes, and a `.env` read leaves no secret in the transcript.
2026-09-03 11:19:53 +07:00

10 KiB

Configuration

Settings come from three places. Later wins:

  1. ~/.shiro-neko/config.json
  2. environment variables
  3. command-line flags

The config file

Written by /provider, editable by hand. Every field is optional.

{
  "provider": "openai",
  "model": "gpt-5",
  "baseURL": "https://api.openai.com/v1",
  "apiKey": "sk-...",
  "presetId": "openai",
  "agent": "default",
  "thinking": "medium",
  "maxRetries": 3,
  "plugins": ["guard", "time"],
  "toolSets": ["edit-plus", "git"],
  "permission": {
    "bash": { "*": "ask", "git *": "allow" }
  },
  "registryUrl": "https://example.com/my-registry/index.json",
  "mcpServers": {
    "fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] }
  }
}
Field Meaning
provider wire protocol: anthropic or openai. Not the vendor — Groq, OpenRouter, and Ollama all speak openai
model model id as the endpoint names it
baseURL API root. Defaults to the official endpoint for the provider
apiKey sent as Authorization: Bearer for openai, x-api-key for anthropic
presetId which preset /provider chose, so it can show what is configured
agent default variant: default, quick, deep, plan, review
thinking default level: off, low, medium, high, max
maxRetries retries per model call for transient failures. Default 3
plugins which builtin plugins to enable. Omit for ["guard", "time"]
toolSets optional tool sets beyond core: edit-plus, git. Omit for all of them. See tools
permission which calls run, ask, or are refused, matched per command or path. See permissions
registryUrl index for /registry. Omit for the default. See registry
mcpServers see MCP

Provider presets

/provider offers these. Each sets baseURL and the wire protocol for you.

Preset Protocol Endpoint Env var checked
Anthropic anthropic api.anthropic.com/v1 ANTHROPIC_API_KEY
OpenAI openai api.openai.com/v1 OPENAI_API_KEY
OpenRouter openai openrouter.ai/api/v1 OPENROUTER_API_KEY
Groq openai api.groq.com/openai/v1 GROQ_API_KEY
DeepSeek openai api.deepseek.com/v1 DEEPSEEK_API_KEY
xAI openai api.x.ai/v1 XAI_API_KEY
Ollama openai localhost:11434/v1 none, keyless
LM Studio openai localhost:1234/v1 none, keyless
Custom OpenAI-compatible openai you supply it none
Custom Anthropic-compatible anthropic you supply it none

provider is the wire protocol, not the vendor. Groq, DeepSeek, xAI, OpenRouter, Ollama, and LM Studio all speak openai; only Anthropic speaks anthropic. Two things differ between them: the auth header (Authorization: Bearer versus x-api-key), and how thinking levels map.

After the key is entered, GET /v1/models is called and the list becomes a picker. Both protocols expose that endpoint with the same data[].id shape, so one code path handles both. If the endpoint does not implement it, a preset with a known model list falls back to that; otherwise you type the model id and setup still completes.

Anything the picker offers is a model the endpoint actually reports, which is more reliable than a hard-coded list — that is why the fallback lists are short and only exist for Anthropic and OpenAI.

Cost estimates

/cost and the status bar price a turn from a table in src/pricing.ts, matched by longest prefix on the model id, so claude-sonnet-4-5-20250929 resolves via claude-sonnet-4-5. An OpenRouter-style anthropic/claude-sonnet-4-5 has its vendor prefix stripped first.

An unknown model is reported as unpriced rather than guessed:

4210 in / 88 out tokens (llama-3.3-70b is unpriced)

Two limits worth knowing. The rates are hand-entered and drift as vendors change them, so treat the figure as an estimate, not a bill. And the token counts come from the provider's usage report, while ~ctx in the status bar is JSON.stringify(messages).length / 4 — good enough to decide when to compact, wrong enough that it should not be read as a token count.

Environment variables

Variable Effect
SHIRO_PROVIDER overrides provider
SHIRO_MODEL overrides model
SHIRO_BASE_URL overrides baseURL
SHIRO_API_KEY overrides apiKey
ANTHROPIC_API_KEY used when provider is anthropic and no key is set
OPENAI_API_KEY used when provider is openai and no key is set
SHIRO_HOME relocates config, sessions, memory, history, user skills, and installs
SHIRO_INSTALL_DIR where install:local and the installers put the binary
SHIRO_REPO which GitHub repo the installers download from
SHIRO_VERSION pins the version the installers fetch

SHIRO_HOME is what the test suite uses to keep a run out of your real config. It is also the way to run two isolated setups side by side — a work profile and a personal one — since it moves every piece of state at once:

SHIRO_HOME=~/work-shiro shiro

A key on the command line ends up in your shell history and in ps. SHIRO_API_KEY in front of one command is better; /provider writing to config.json is better still.

Flags

shiro [options]
shiro -p "prompt"          headless, prints to stdout
cat file | shiro -p        prompt read from stdin
Flag Effect
-p, --print [prompt] headless mode. Needs --yolo for tool use
--json with -p, one JSON event per line
-c, --continue resume the newest session for this directory
-r, --resume <id> resume by session id or unique prefix
--agent <name> default, quick, deep, plan, review
--think <level> off, low, medium, high, max
--provider <name> anthropic or openai
--model <id> model id
--base-url <url> API root
--no-mcp skip MCP servers
--no-subagent omit the task tool
--no-instructions ignore AGENTS.md and friends
--no-skills ignore builtin, installed, and project skills
--no-plugins disable all plugins, builtin and installed, including the guard
--no-memory do not load or write project memory
--yolo skip every approval prompt
-v, --version version, bun version, platform, source or compiled
-h, --help usage

The --no-* flags exist for isolating a problem. All six together strip the agent to its built-in tools and nothing else, which answers "is this the loop or something layered on it?" in one run:

shiro --no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp

--no-plugins also disables the guard, so rm -rf becomes an ordinary approval prompt. Reasonable while debugging, not something to leave on.

An unknown value fails at startup with the valid list rather than falling back silently:

$ shiro --agent turbo
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review

Where things live

~/.shiro-neko/
  config.json                 provider, model, key, defaults
  sessions/<uuid>.json        transcripts, token counts, cost, task list
  memory/<hash>.json          durable per-project notes
  history/<hash>.json         prompt history for up-arrow recall
  skills/*.md                 your own skills
  registry/skills/*.md        skills installed with /registry
  registry/plugins/*.json     plugin manifests installed with /registry

Project files:

<project>/
  AGENTS.md                   instructions injected into the system prompt
  .shiro/skills/*.md          project skills, override user and builtin
  .shiroignore                extra ignore rules on top of .gitignore

Memory and history file names are SHA-256 prefixes of the absolute project path, because a path is not a safe filename. Two consequences: moving a project loses its memory and history, and two checkouts of the same repo at different paths keep separate ones.

OpenAI reasoning models

Newer OpenAI models reject function tools on /v1/chat/completions and require /v1/responses. For api.openai.com both are chained: a 400, 404, 405, 415, 422, or 501 on the first switches to the second, sticks for the rest of the session, and prints one notice. Retryable failures — 429 and 5xx — are left to the SDK's backoff instead.

Only those six codes qualify, because they mean "this endpoint cannot serve this request shape". A 401 is a wrong key and switching endpoints would only produce a second 401 with a more confusing message.

The switch is sticky on purpose: once an endpoint rejects the shape it rejects every later step too, so re-probing would waste a round trip per step of every turn.

Third-party endpoints get a plain chat-completions model with no fallback probe, since they do not implement /v1/responses.

The two endpoints also differ in how they carry assistant history, which is where compaction gets interesting — see memory.

Config that changes behaviour subtly

Four fields do more than they look like they do.

thinking costs money and latency on every turn, not just hard ones. off on a hard problem produces confident wrong answers; max on a rename wastes cents and seconds. The agent variants already pick sensible levels — see agents.

toolSets removes tools from the model's view entirely. If the agent stops using a tool you expected, check the startup header for which sets loaded: an unrecognised name is dropped silently, so "gti" reads as "git is off". See tools.

permission replaces a tool's defaults rather than merging with them, so { "read_file": { "*": "allow" } } also allows reading .env. Order inside a rule table decides the outcome, since later rules win. See permissions.

registryUrl is the whole trust decision for installed skills and plugins. There are no signatures, so pointing it at an index means trusting whoever controls that URL — including for whatever they publish later. See registry.