BUG: Anthropic client (sourceFormat=CLAUDE) hitting a claude-target provider
got a raw OpenAI chat.completion body on non-streaming requests. The
needsTranslation(CLAUDE,CLAUDE) gate is false when target===source, so the
translator never ran; a claude-transport executor replying OpenAI JSON
(opencode/big-pickle) leaked choices[]/prompt_tokens to the client, which
Anthropic SDKs cannot parse (no content[] blocks, no type:"message").
FIX: shape-aware guard toClaudeMessageShape() in nonStreamingHandler — when
sourceFormat is CLAUDE, convert any OpenAI-shape body to a proper Claude
message (type, content blocks with thinking/text/tool_use, stop_reason via
finish mapping, usage input/output tokens). Claude-shaped bodies pass through.
Also: strip <|im_end|>/<|endoftext|>/<|eot_id|> EOS sentinels from Claude
text deltas in openai-to-claude and kiro-to-claude translators — the upstream
EOS token leaks into the final text_delta (observed 'OK<|im_end|>').
Tests: anthropic-nonstream-shape.test.js (6: eos strip + shape guard incl
tool_calls→tool_use, pass-through, finish mapping). 28/28 related tests green.
Reproduce the MiMo Desktop login surface server-side so headless/Docker
deployments can link a Xiaomi account without the Desktop client. The
account session (passToken) is captured during the proxied login and
stored per connection.
- Five account clusters (cn/sgp/ams/ru/in): per-region mimo-server host
and SSO sid, unknown region falls back to sgp
- mimo-v2.6-pro/flash/pro-ultraspeed dual-route models: account-service
route when desktop credentials exist, cloud API (sk- key) otherwise;
drops obsolete mimo-x-*-preview ids
- Desktop ServiceTokenManager 2-phase handshake (single serviceLogin with
target sid, raw 64-bit nonce preserved), per-region session cache
- reasoning_effort bridged to output_config.effort; i18n runtime now
observes characterData mutations so React text rewrites get translated
- Security hardening on the login proxy: session travels only in the
httpOnly cookie (never in the URL), proxy branch requires dashboard
auth, authorization/proxy-authorization never forwarded upstream, and
upstream Set-Cookie is not replayed onto the app origin
Anthropic's API-level refusal (streaming classifier / ToS) ends the stream
with stop_reason "refusal", stop_details carrying the reason, zero output
tokens and no content blocks. Map refusal to content_filter in both
directions, surface stop_details.explanation as message text, and add
CLAUDE_STOP.REFUSAL to schema.
getUsageStats("all") shipped the entire usageHistory table to JS just to
refine lastUsed (~2s on 290K rows, on every statsEmitter update per SSE
listener). Bound the overlay to a 2-day indexed range scan; older entries
keep day-level lastUsed from usageDaily aggregates. Totals unaffected.
budgetToLevel now maps budgets > 80384 (midpoint of 32768/128000) to
"max" instead of clamping to "xhigh", so the top reasoning tier is
reachable from large budget_tokens requests.
Register mimo-v2.6-flash-free on opencode-zen (chat lane) with a v2.6
capability pattern, and switch the vision adapter default from the old
mimo-v2.5-free.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Map upstream Chat Completions usage to the Responses API shape and attach it to response.completed. Capture chunk.usage before the empty-choices guard so the usage-only trailer chunk survives, and defer completion to flushEvents() when usage is not yet known — only on the direct openai:openai-responses route, since a pivoted stream never reaches flushEvents. Fixes#3432.
- Track both weekly and 5-hour session buckets in parseWeeklyQuotaSummary,
distinguishing sliding-window limits from multi-day weekly limits
- Preserve disabled session buckets at 0% rather than dropping them when weekly limits are reached
- Target 5-hour session rows (not weekly rows) during family exhaustion reconciliation in getAntigravityUsage
- Suppress synthesized per-model duplicate rows in dashboard normalization when family summaries are present
- Add unit test coverage for multi-bucket extraction, reconciliation isolation, and dashboard deduplication
- Match code 110 (billing daily count exceeded) alongside 112/10605/pricingUrl
in isBillingBlock, parsing JSON safely and accepting numeric/string codes
- Accept numeric strings for statusCodeValue and object bodies in envelope peek
- Emit structured 403 quota error chunk instead of synthetic assistant text
when a billing envelope appears mid-stream
- Preserve upstream HTTP status in handleForcedSSEToJson when error chunk carries
a valid 400-599 status
- Add unit tests for code-110 detection, mid-stream billing envelopes, and false-positive guard
Strict OpenAI-compatible validators reject unknown assistant-message
fields: Groq 400 ("property 'reasoning_content' is unsupported"),
Mistral 422 ("extra_forbidden"), Cerebras 400 ("wrong_api_format").
Clients driving reasoning models (Hermes Agent, and anything following
the DeepSeek/Kimi convention) echo the previous turn's reasoning_content
on every assistant message, so from the second turn on every request to
these providers fails and a fallback combo silently skips them.
Add a dropMessageFields rule to paramSupport.js that strips
reasoning_content / reasoning / reasoning_details from assistant turns
for groq, mistral, and cerebras.
- Add Cursor Default / Claude Default on Dashboard -> Combos to generate
unprefixed combo names that match Cursor/Claude client model IDs,
seeded with cu/... or cc/... so those clients can route through 9Router.
- Add multi-select bulk Delete and bulk Set strategy (Fallback / Round Robin / Fusion).
- Docs and unit tests for preset builder.
Cursor-hosted models (cu/composer-2.5, cu/cursor-grok-*, cu/default) returned
HTTP 200 with an empty turn, or hung, whenever a client sent tools.
- Fold system prompts into the current user message. custom_system_prompt
(RunRequest field 8) makes AgentService return an empty turn.
- Send ModelDetails (field 3); thinking variants (Composer, Grok, *-thinking)
return an empty turn when only requested_model (field 9) is set.
- Route tool-call history and declared tool schemas through AgentService:
encode OpenAI tools into mcp_tools (field 4), decode McpArgs and emit real
tool_calls with finish_reason tool_calls.
- Map Composer thinking / Grok thinking_delta (field 4) into visible content
instead of dropping the answer with the unsigned reasoning.
- Ack request_context without echoing MCP tools (double-advertise stalls the
HTTP/2 stream) and ack kv_server_message so the run proceeds.
- Reject IDE builtin execs instead of failing the turn, so the model can
continue with MCP tools or a text answer.
- Add google.protobuf.Value / MCP encoders and a FIXED64 branch to
encodeField in cursorProtobuf.js.
RTK now compresses the source-format body before translation for cursor only:
its translator rewrites role:tool into user XML, so the post-translate pass
missed those tool results. Every other provider keeps the post-translate pass
unchanged.
- Export aggregateComboCapabilities: union for vision/audio/search/pdf,
intersection for tools, primary-model for reasoning fields, min
contextWindow, max maxOutput
- Support nested combo resolution in aggregateComboCapabilities via
comboLookup with depth guard (max 6)
- Wire capability metadata to all /v1/models entries and combos
- Show aggregated ctx/max metadata line and capability badges on combo chips
- Pattern fixes: MiMo v2.5/omni reasoning, qwen max/plus vision, minimax m2.x vision
- Sync commandcode model catalog and add openai gpt-5.5
- Add unit tests for capability patterns and combo capability aggregation
Introduce AGENTS.md (root, primary agent instruction file) documenting six
hard-won fixes with explicit DO NOT / WHY, plus executable enforcement so a
future AI cannot delete or reintroduce them:
1. package-lock.json must be generated with npm 10 (Docker's npm 10.9.8).
npm 11 drops the top-level @emnapi/core + @emnapi/runtime entries npm 10
needs, breaking the tag-triggered Docker build at `npm ci` (happened on
v1.0.14). Add scripts/verify-lockfile-npm10.mjs + .npmrc + a Dockerfile
fail-fast check + a CI step + tests/unit/lockfile-npm10-guard.test.js.
Also re-fix the lockfile itself (regenerated with npm 10.9.8).
2. Tests must never write to the real ~/.9router DB (isolateDataDir).
3. Hidden providers must not leak into Usage (usageProviders !p.hidden).
4. codebuddy-intl connection test + OAuth identity.
5. Fork-only features that must survive upstream syncs.
6. Upstream sync procedure.
Each marker cross-references AGENTS.md and the covering test. CLAUDE.md now
points to AGENTS.md at the top. Verified: build ok, guard script passes,
full suite leaves the real DB count unchanged (38), 0 new regressions.
Root cause of fake connections in Usage (zed-live-*@example.com,
guard-*@example.com, zed "Account N", kimchi-nope): route-level tests
(zed-live-models, zed-native-auth) call createProviderConnection, which
persists to $DATA_DIR/db/data.sqlite. With DATA_DIR unset — the default
for `npx vitest run` — that resolved to the user's real ~/.9router DB,
appending test rows on every run. Both files documented "RUN WITH AN
ISOLATED DB" but never enforced it.
Add tests/setup/isolateDataDir.js (wired via vitest setupFiles) that
points DATA_DIR at a throwaway temp dir before src/lib/dataDir.js is
imported. Opt out with RUN_REAL=1 or an explicit DATA_DIR (used by the
*.real.test.js suites that read live credentials).
Verified: a full suite run now leaves the real DB byte-count unchanged;
new guard test tests/unit/test-data-dir-isolation.test.js locks it in.
The Usage page auto-adds every noAuth free provider so connectionless
providers (opencode) still appear. It did not filter the registry's
hidden flag, so devin-cli and mimo-free — both category:"free" with
noAuth:true and hidden:true — showed up in Usage despite having no
connection and being absent from the Providers page (which does filter
hidden).
Extract the list assembly into buildUsageProviderList (shared/utils/
usageProviders.js) and skip hidden free providers there. Behavior for
visible noAuth providers (opencode) and dedup of active connections is
unchanged; covered by tests/unit/usage-provider-list.test.js.
Two bugs on codebuddy-intl connections:
1. Test Connection always failed with "Provider test not supported":
codebuddy-intl was missing from OAUTH_TEST_CONFIG, so testOAuthConnection
bailed before probing. Add a real probe against the Keycloak realm's
userinfo endpoint (URL derived from the token's iss claim), and wire
refreshable so an expired token is rotated via refreshCodebuddyIntlToken.
2. OAuth logins were named "Account N" with no email: mapTokens returned no
identity, even though the access token is a Keycloak JWT carrying
email/name claims. Extract email + displayName in mapTokens (new shared
extractDisplayNameFromAccessToken helper) so fresh logins are named and
deduped by identity.
Also add a run-once backfill (backfillCodeBuddyIntlIdentity) invoked from
GET /api/providers and /api/providers/client to self-heal existing rows
(backfill email/displayName, rename the generic "Account N" placeholder).
Verified live: the real connection now returns valid:true and the row is
renamed to the account email.
Upstream f6e7cabe stopped workos:-prefixing opaque tokens (ClinePass API
keys like clp_...), so getClineAccessToken now only prefixes WorkOS JWTs.
The fork's test asserted the old unconditional-prefix behavior.
Replace the retired api-inference.huggingface.co host with the Inference
Providers router (router.huggingface.co): imageConfig.modelMap resolves
Hub ids to provider-resolved ids, image-to-image models receive the
source image in inputs with the prompt under parameters.prompt, and a
new sttConfig wires the hf-inference ASR route. The image catalog grows
to 23 models, dead whisper-small is replaced by whisper-large-v3-turbo,
the unusable "language" param is dropped, and edit models declare the
edit capability so the dashboard offers a source image. Adds unit and
end-to-end coverage plus a model-id guard on custom endpoints.
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
## Features
- **Xiaomi MiMo**: merge MiMo Desktop support into `xiaomi-mimo` with dual auth (API key + Desktop/OAuth session), Preview models support, and encrypted-callback OAuth flow
- **Claude Code**: add 1M-context toggle (`[1m]` marker) and drive `CLAUDE_CODE_AUTO_COMPACT_WINDOW` directly from the dashboard
- **Models**: add DeepSeek-V4.1-Flash to DeepSeek provider, CodeBuddy-Intl, and Ollama (`deepseek-v4.1-flash:cloud`); enable `low`..`max` reasoning effort levels and vision capability for DeepSeek-V4.*
- **i18n**: integrate Persian (fa) translation
## Fixes
- **OpenCode / OpenCode Go**: resolve 403 `FreeTierError` and 429 rate limits with canonical session format, valid User-Agent, and stable upstream session reuse; force stream and declare `forceStream` for free-tier SSE aggregation; cloak decoy tools, normalize Muse Free tool choice, and strip prior reasoning items on Responses models; route Union Alpha via Messages API
- **Kiro**: preserve underscores in tool names (`mcp__server__tool`) and restore client tool names in responses; use neutral placeholder for tool-result-only turns; forward tool-result images
- **Stream**: report aborts after HTTP 200 in-band (per-format error frames) instead of closing silently
- **Command Code**: preserve images and `reasoning_effort` on `/alpha/generate`; retry transient stream errors and avoid fake stop chunks; add Quota Tracker support
- **Zed**: harden OAuth lifecycle (preserve `systemId`, renew proxy timeout), support live model resolution, and lower display priority in OAuth list
- **Antigravity**: scope cached thought signatures to model family; strip Claude Code billing headers from system prompts; sanitize Hermes system identity
- **Codex**: route bare `codex-auto-review` requests to the Codex provider (#4135)
- **Auth**: do not cool down an account for request-scoped 4xx errors
- **Usage**: improve DeepSeek credit balance display as currency credit instead of 0/total quota bar
- **Model Catalog**: scope synced catalog to gateways and declare vision capabilities for DeepSeek V4.1-Flash IDs
- Force stream:true and cloak decoy tools (bash, read) for OpenCode free tier
- Support connection testing for opencode in testUtils
- Expand error message slice limits in auth and ping to preserve workspace link
- Add concise China region link chip in provider detail page
Co-Authored-By: Claude Code <noreply@anthropic.com>
- executors/zed.js: use exact wire values (anthropic, open_ai, google, x_ai)
and strip incompatible Vertex safetySettings on the Google path
- shared/zedAuth.js: robust callback query parsing, reject garbage PKCS#1 v1.5
decryptions, and thread proxyOptions when fetching LLM tokens
- oauth: preserve systemId across authorize/register/exchange lifecycle,
renew proxy idle timeout on reuse, and ignore non-callback localhost requests
- shared/OAuthModal.js: track owned proxy in flowRef and stop at most once
- api/providers/[id]/models: add connection-scoped live Zed model resolver
- registry: unhide provider in dashboard
- tests: add unit coverage for wire format, native auth, and live models
Do not trigger account cooldown or fallback for request-scoped 4xx errors that match no account rules so healthy credentials are not locked out for context length or validation errors.
OpenCode Free returns HTTP 400 for muse-spark-1.3-contributor-free when tool_choice is non-auto. Declare forceAutoToolChoiceModels quirk and normalize explicit tool_choice to auto.
Strip prior-turn type: reasoning items and encrypted_content fields from body.input on Muse Spark Responses endpoints in OpenCode and OpenCode Go executors to avoid HTTP 400 errors across rotated proxy accounts.
Do not collapse consecutive underscores in uniqueName so mcp__server__tool is sent intact to Kiro, attach reverse map on request translation, and restore client tool names in responses.
Replace the literal 'continue' placeholder on tool-result-only user turns with 'Tool results provided.' to prevent models from treating it as a new user instruction.
Forward images inside tool_result to OpenAI and Kiro upstreams via following user messages, restore original client tool names on Kiro responses via _toolNameMap, and preserve thinking display settings across translations.
Follow-up to the canonical-session fix: with no explicit session,
every request minted a fresh x-opencode-session, and upstream free-tier
quota is accounted per session. That burns through quota and surfaces
as 429 FreeUsageLimitError with growing reset-after delays, while the
real CLI reuses one long-lived session per conversation.
- Stable canonical session per downstream identity (connectionId, else
auth-header hash, else shared default), evicted after
MEMORY_CONFIG.sessionTtlMs like the other session stores.
- Deterministic x-opencode-request per message (stable across retries,
like the CLI user message id); valid downstream ids preserved.
- 6 more unit tests (22 total).
OpenCode upstream validates free-tier requests: User-Agent must be opencode/<version> (>= 1.17.0) and x-opencode-session must match canonical ses_ format. Default OPENCODE_UA to opencode/1.18.31, generate canonical descending session IDs, provide deterministic foreign session translation, and isolate credentials per-request.
Route union-alpha to /zen/v1/messages with targetFormat claude, add anthropic-version header, and register model capabilities (vision, 262K context, 131K max output).
Derive responses-only routing from the model registry's targetFormat instead of hardcoding model checks, and strip thinking suffixes when looking up models in providerModels so variants like gpt-5.6-luna(high) are routed correctly to /responses.
- Declare deepseek-v4.1-flash and deepseek-flash as vision-capable in MODEL_CAPABILITIES
- Share installed catalogSource across route chunks via globalThis.__9rCatalogSource
- Scope catalog modality keys by provider:model to prevent cross-gateway collisions
- Upgrade catalog format to v2 with automatic rebuild of older schemas
Command Code dropped vision and ignored client effort through the router:
image blocks became "[image omitted]", HTTP image URLs were never inlined,
and reasoning_effort landed on the envelope wrapper instead of params (so the
DeepSeek family mapping remapped low -> high). The catalog also treated
deepseek/deepseek-v4.1-flash as text-only, so the vision adapter stole those
requests to another provider.
- Map OpenAI image_url / Claude image blocks (base64 or data-URI) to the
native {type:"image", image:"data:...;base64,...", mimeType} generate block.
- Add FORMATS.COMMANDCODE to TARGETS_NEED_BASE64 so remote http(s) images are
inlined by the existing SSRF-safe fetcher before translation.
- Write reasoning_effort inside params for targetFormat commandcode and pass
low|medium|high|xhigh|max through unmapped; allow it in thinkingLevels.
- Provider-scoped capabilities for commandcode/cmc: vision except the CLI
text-only denylist, thinkingFormat commandcode, so family patterns
(deepseek-v4 -> thinkingFormat deepseek, vision false) no longer win.
- Quota Tracker: whoami + billing credits/subscriptions (credits vs plan cap,
5h and weekly windows), labels from AI_PROVIDERS[].name.
A stream that stalled or lost its upstream was closed with no terminal frame
at all, so clients saw "200 OK, a few chunks, then nothing" and could not tell
a truncated reply from a finished one. The Responses passthrough path already
synthesized response.failed; every other client format got nothing.
The watchdog now hands its reason ("stream stall timeout" or "upstream
connection lost") to onAbortTerminal, and buildStreamErrorBytes frames it per
client format: OpenAI-compatible clients get data: {"error":{...}} followed by
data: [DONE], Anthropic clients get `event: error`. The error frame always
precedes [DONE] (openai-python raises APIError on any data payload carrying an
error key), and no synthetic finish_reason is ever emitted — a truncated
stream must not look like a clean stop.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Adds the Desktop-exclusive Preview models and the Xiaomi account-session
route to the existing xiaomi-mimo provider instead of a separate
xiaomi-desktop provider, so the dashboard shows one MiMo entry rather than
three overlapping ones.
Dual auth, same pattern as kimi — API key (sk-) covers the cloud API,
Desktop/OAuth adds the account session used by the Preview models:
- registry: category oauth, authModes [oauth, apikey], oauth block, the two
mimo-x-*-preview models, and the invite signupUrl
- executor: routes Preview models to the account-service route with a Cookie
session, everything else keeps the sourceFormat-matched transport
- oauth: custom ECDH encrypted-callback flow (X25519 -> SHA256 -> AES-256-GCM)
with a loopback callback proxy, plus one-click import of the local Desktop
auth.json
- usage: weekly quota from the account session
Fixes found while merging:
- the OAuth browser flow was dead: poll-status cleared the session before the
client could POST /exchange, so every exchange returned 400
- a Claude-format client was sent to /v1/chat/completions instead of the
declared /anthropic/v1/messages transport, because buildUrl ignored
runtimeTransport
- stopXiaomiMimoProxy leaked every pending session (each holding an X25519
private key) for the process lifetime
- the OAuth exchange did not persist the Desktop passToken, so the Preview
models could never work after a browser sign-in
Removes dead code: the local engine token minting (mimoEngine, never called
on the request path), the model-catalog and usage routes, engineToken/
engineUrl plumbing, and an unread top-level usage block.
Adds tests/unit/xiaomi-mimo-{executor,oauth-session,oauth-proxy}.test.js —
the provider previously had none.
Registry order is the display order for the provider page, /v1/models, and the
CLI selector, so moving the entry to the head of `models` is the whole change.
deepseek-flash keeps its id and supportedFormats — only its position moves.
Co-Authored-By: Claude Code <noreply@anthropic.com>