The non-streaming Anthropic path leaked OpenAI chat.completion bodies to
Claude clients when the upstream provider forces streaming (forced
SSE->JSON path): parseSSEToOpenAIResponse yields choices[] which Claude
Code cannot parse. Convert to Claude Messages shape for sourceFormat
CLAUDE in sseToJsonHandler.
Shared toClaudeMessageShape/openAICompletionToClaudeMessage moved to
translator/concerns/claudeShape.js to avoid circular import between
nonStreamingHandler and sseToJsonHandler. 5 new tests: sse-to-json
reassembly + Claude guard, plus existing anthropic-nonstream-shape
updated to import from the concern.
BUG: Anthropic client (sourceFormat=CLAUDE) hitting a claude-target provider
got a raw OpenAI chat.completion body on non-streaming requests. The
needsTranslation(CLAUDE,CLAUDE) gate is false when target===source, so the
translator never ran; a claude-transport executor replying OpenAI JSON
(opencode/big-pickle) leaked choices[]/prompt_tokens to the client, which
Anthropic SDKs cannot parse (no content[] blocks, no type:"message").
FIX: shape-aware guard toClaudeMessageShape() in nonStreamingHandler — when
sourceFormat is CLAUDE, convert any OpenAI-shape body to a proper Claude
message (type, content blocks with thinking/text/tool_use, stop_reason via
finish mapping, usage input/output tokens). Claude-shaped bodies pass through.
Also: strip <|im_end|>/<|endoftext|>/<|eot_id|> EOS sentinels from Claude
text deltas in openai-to-claude and kiro-to-claude translators — the upstream
EOS token leaks into the final text_delta (observed 'OK<|im_end|>').
Tests: anthropic-nonstream-shape.test.js (6: eos strip + shape guard incl
tool_calls→tool_use, pass-through, finish mapping). 28/28 related tests green.
The version label now renders 'v0.5.86 (<shortsha>)' when the flake build
injects NEXT_PUBLIC_GIT_SHA (git commit identity of the deployed artifact),
so any running dashboard traces to its exact source commit. Falls back to
semver-only when the env var is absent (local dev). Also fixes the profile
page repo link back to asepharyana/9router (MIBP merge target), replacing
the mhiqrambg fork label left over from the merge.
flake.nix now uses self.shortRev (git HEAD short hash) as the derivation
version, so the store path is /nix/store/<hash>-9router-<shortsha> and any
deployed build traces to its exact source commit. Removes the package.json
version bump churn; the updater's deployed/current comparison switches to
git SHA.
Migrate the Nix build + deploy source from /home/code/9router-repo (a pure
upstream mirror, now deleted) to this repo, which carries the MIBP fork
merge (freebuff, cline free models, proxy-pool fitness). flake.lock and
flake.nix copied verbatim from 9router-repo; bun.lock generated from the
merged dependency set. The cron updater (9router-update.sh) now points at
/home/code/9router with a keep-ours conflict policy.
## Features
- **Xiaomi MiMo**: server-assisted desktop login for headless/Docker deployments, five account clusters (cn/sgp/ams/ru/in), and v2.6 pro/flash/pro-ultraspeed models with dual-route (account service vs. cloud API)
- **Claude**: add Claude Opus 5.5 support
- **i18n**: translate React text rewrites via characterData mutation observer
## Fixes
- **Proxy Pools**: keep request headers intact through Vercel/Cloudflare/Deno relays (spreading a `Headers` instance yielded `{}`, dropping auth and content-type)
- **Xiaomi MiMo login**: keep the session in the httpOnly cookie only, require dashboard auth on the proxy branch, and stop forwarding authorization headers upstream
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.
- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
route when desktop credentials exist, cloud API (sk- key) otherwise;
drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
httpOnly cookie (never in the URL), proxy branch requires dashboard
auth, authorization/proxy-authorization never forwarded upstream, and
upstream Set-Cookie is not replayed onto the app origin
Preserve request headers through the Vercel/Cloudflare/Deno relay.
- proxyFetch: normalize options.headers via Object.fromEntries before
spreading into the relay headers. Spreading a Headers instance yields
{} and silently dropped every entry (auth + content-type), which is why
the same pool worked on one path and failed on another.
- vercel relay: build the forwarded header object from req.headers.entries()
instead of new Headers(req.headers), avoiding edge-runtime normalization
of casing/duplicate keys that some providers reject.
## Features
- **System One**: add `/v1/systemone` decision endpoint for Jev models (OpenCode Zen and OpenRouter lanes), wire into sidebar and Media Providers page with interactive probe testing
- **CLI Tools**: add dynamic configuration, settings APIs, and official logos for Pi, OMP, Crush, ForgeCode, Smelt, and CodeWhale
- **Analytics & Usage**: add Requests mode, provider/model breakdown charts, All Time period filter, and refined overview cards
- **Combos**: add Cursor/Claude Default presets; support bulk select/delete and bulk strategy changes (Fallback / Round Robin / Fusion)
- **Model Capabilities**: expose model capability metadata on `/v1/models` and aggregate capabilities across combo targets
- **OpenCode Zen & MiMo**: add OpenCode Zen (`opencode-zen`) provider with free-tier fingerprint; switch default vision fallback to MiMo V2.6 Flash Free
- **Qoder CN**: add `qoder-cn` provider for qoder.com.cn with OAuth flow, COSY protocol, and CN gateway routing
## Fixes
- **Translator**: map Claude `refusal` stop_reason to `content_filter` and surface explanation; strip replayed reasoning fields for Groq, Mistral, and Cerebras (#4220)
- **Antigravity**: drop requestType `agent` to avoid false 429 `RESOURCE_EXHAUSTED`; separate weekly and short-window (5-hour) quotas and deduplicate dashboard rows
- **Responses API**: report usage on `response.completed` so clients can auto-compact (#3432)
- **Hugging Face**: migrate to Inference Providers router (`router.huggingface.co`), expand image models catalog, and add STT route
- **Qoder**: prevent signed request replay (`403/103 Duplicate request`), handle code 110 billing blocks, and preserve upstream SSE error status
- **Performance**: bound usage `lastUsed` scan to a 2-day window; map large budget tokens to `max` reasoning tier
- **Docker**: publish verified multi-platform images (linux/amd64 and linux/arm64) with configurable apk build mirrors
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
- Add "all" period option to usage dashboard and chart API
- Aggregate all days in getChartData when period is "all"
- Center overview card metrics and adjust font size to prevent truncation
Co-Authored-By: Claude Code <noreply@anthropic.com>
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.
budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
- Add dedicated settings API routes for pi, omp, crush, forge, smelt, codewhale
- Integrate GenericCliToolCard with multi-model support for Pi and auto-discovery for OMP
- Register tools in cliTools catalog and all-statuses route
- Add official logos for all new CLI tools
Co-Authored-By: Claude Code <noreply@anthropic.com>
OpenRouter serves TypeSafe Jev at POST /api/v1/systemone with the same
request/response shape, so it plugs into systemoneConfig directly with
model typesafe/jev-1.13. Mark the System One media kind isNew and
render a New badge on the sidebar kind item and the Media Providers
accordion when any visible kind is new.
Co-Authored-By: Claude Code <noreply@anthropic.com>
The inline model test sent a chat-completions payload and failed with
500 on decision models. Add a systemone branch to pingModelByKind that
submits a native state+questions probe, and restore the test button
that was hidden for System One models.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Allow customizing evaluation instructions in the System One example
card, and disable the inline model probe button for System One models
since decision models do not accept chat completion probes.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Expose System One in the Media Providers sidebar accordion, set
kind to "systemone" on Jev models for ModelsCard filtering, wire
systemoneConfig into ProviderInfoCard, and configure GenericExampleCard
for interactive testing of decision models.
Co-Authored-By: Claude Code <noreply@anthropic.com>
New /v1/systemone pass-through route for Jev decision models (jev-1.13,
jev-1.13-free) on OpenCode Zen and the free lane. Follows the media-route
pattern: systemoneConfig in the registry drives URL/headers, the handler
mirrors the embeddings account-fallback + usage flow, and the dashboard
gains a System One media-provider kind. No chat-pipeline changes.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes#3432.
- Build linux/amd64 and linux/arm64 on native GitHub runners
- Assemble version manifests from platform digests and promote latest only after verification
- Add release/tag validation, manual republishing, timeouts, and health smoke tests
- Make Docker build mirrors configurable via build args and remove unnecessary runtime apk upgrades
- Update DOCKER.md documentation
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").
Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.
Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
unprefixed combo names that match Cursor/Claude client model IDs,
seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.
- Fold system prompts into the current user message. custom_system_prompt
(RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
tool_calls with finish_reason tool_calls.
- Map Composer thinking / Grok thinking_delta (field 4) into visible content
instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
encodeField in cursorProtobuf.js.
RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
intersection for tools, primary-model for reasoning fields, min
contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
Introduce AGENTS.md (root, primary agent instruction file) documenting six
hard-won fixes with explicit DO NOT / WHY, plus executable enforcement so a
future AI cannot delete or reintroduce them:
1. package-lock.json must be generated with npm 10 (Docker's npm 10.9.8).
npm 11 drops the top-level @emnapi/core + @emnapi/runtime entries npm 10
needs, breaking the tag-triggered Docker build at `npm ci` (happened on
v1.0.14). Add scripts/verify-lockfile-npm10.mjs + .npmrc + a Dockerfile
fail-fast check + a CI step + tests/unit/lockfile-npm10-guard.test.js.
Also re-fix the lockfile itself (regenerated with npm 10.9.8).
2. Tests must never write to the real ~/.9router DB (isolateDataDir).
3. Hidden providers must not leak into Usage (usageProviders !p.hidden).
4. codebuddy-intl connection test + OAuth identity.
5. Fork-only features that must survive upstream syncs.
6. Upstream sync procedure.
Each marker cross-references AGENTS.md and the covering test. CLAUDE.md now
points to AGENTS.md at the top. Verified: build ok, guard script passes,
full suite leaves the real DB count unchanged (38), 0 new regressions.
Root cause of fake connections in Usage (zed-live-*@example.com,
guard-*@example.com, zed "Account N", kimchi-nope): route-level tests
(zed-live-models, zed-native-auth) call createProviderConnection, which
persists to $DATA_DIR/db/data.sqlite. With DATA_DIR unset — the default
for `npx vitest run` — that resolved to the user's real ~/.9router DB,
appending test rows on every run. Both files documented "RUN WITH AN
ISOLATED DB" but never enforced it.
Add tests/setup/isolateDataDir.js (wired via vitest setupFiles) that
points DATA_DIR at a throwaway temp dir before src/lib/dataDir.js is
imported. Opt out with RUN_REAL=1 or an explicit DATA_DIR (used by the
*.real.test.js suites that read live credentials).
Verified: a full suite run now leaves the real DB byte-count unchanged;
new guard test tests/unit/test-data-dir-isolation.test.js locks it in.
The Usage page auto-adds every noAuth free provider so connectionless
providers (opencode) still appear. It did not filter the registry's
hidden flag, so devin-cli and mimo-free — both category:"free" with
noAuth:true and hidden:true — showed up in Usage despite having no
connection and being absent from the Providers page (which does filter
hidden).
Extract the list assembly into buildUsageProviderList (shared/utils/
usageProviders.js) and skip hidden free providers there. Behavior for
visible noAuth providers (opencode) and dedup of active connections is
unchanged; covered by tests/unit/usage-provider-list.test.js.
Two bugs on codebuddy-intl connections:
1. Test Connection always failed with "Provider test not supported":
codebuddy-intl was missing from OAUTH_TEST_CONFIG, so testOAuthConnection
bailed before probing. Add a real probe against the Keycloak realm's
userinfo endpoint (URL derived from the token's iss claim), and wire
refreshable so an expired token is rotated via refreshCodebuddyIntlToken.
2. OAuth logins were named "Account N" with no email: mapTokens returned no
identity, even though the access token is a Keycloak JWT carrying
email/name claims. Extract email + displayName in mapTokens (new shared
extractDisplayNameFromAccessToken helper) so fresh logins are named and
deduped by identity.
Also add a run-once backfill (backfillCodeBuddyIntlIdentity) invoked from
GET /api/providers and /api/providers/client to self-heal existing rows
(backfill email/displayName, rename the generic "Account N" placeholder).
Verified live: the real connection now returns valid:true and the row is
renamed to the account email.
Upstream f6e7cabe stopped workos:-prefixing opaque tokens (ClinePass API
keys like clp_...), so getClineAccessToken now only prefixes WorkOS JWTs.
The fork's test asserted the old unconditional-prefix behavior.