feat(gmw): route all LLM traffic through 9router

GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA
vision endpoint onto 9router, which runs on the same host as both services
(127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls
bypass 9router's remote-key guard.

- gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1
- drop stale 'omniroute' router references from comments/docs now that the
  active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS)

Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite
(SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers
image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the
existing Qdrant collections (no reindex needed). The GMW key is already
registered in 9router's apiKeys table.

typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
This commit is contained in:
asepharyana
2026-09-24 16:05:36 +07:00
parent 750f3aa598
commit ef7708bf7d
6 changed files with 15 additions and 13 deletions
+4 -4
View File
@@ -96,10 +96,10 @@ export const configSchema = z
.transform((v) => v === "true")
.default(false),
AI_LLM_API_KEY: z.string().optional(),
AI_LLM_BASE_URL: z
.string()
.url()
.default("http://100.121.180.82:20128/api/v1"),
// 9router — OpenAI-compatible router on this host (127.0.0.1:4014).
// Loopback on purpose: backend runs on the same machine as 9router, so no
// TLS/proxy hop is needed.
AI_LLM_BASE_URL: z.string().url().default("http://127.0.0.1:4014/v1"),
AI_LLM_MODEL: z.string().default("text"),
AI_LLM_VISION_MODEL: z.string().optional(),
AI_LLM_EMBEDDING_MODEL: z.string().optional(),
+1 -1
View File
@@ -54,7 +54,7 @@ src/
**Never** reintroduce regex/heuristic content classification.
2. **Discord tokens sanitized** before reaching LLM (`discordTokens.ts`).
3. **Semantic cache is batched** — one embed call + one Qdrant batch search.
4. **Streaming is mandatory** against the omniroute base URL.
4. **Streaming is mandatory** against the router base URL.
## AI moderation pipeline
+1 -1
View File
@@ -194,5 +194,5 @@ pipeline gauges registered by `app/metrics-collector.ts` —
numeric snowflake IDs never trigger false positives.
- **Semantic cache is batched** (one embed call + one Qdrant batch search),
not N sequential round-trips. `ensureQdrantCollection` is memoized.
- **Streaming is mandatory** against the omniroute base URL (non-stream waits for
- **Streaming is mandatory** against the router base URL (non-stream waits for
the full body and times out). `llmClient` aggregates SSE chunks.
@@ -92,7 +92,7 @@ export async function callModerationLLM(
jsonResponse: { type: "json_object" },
retries: 0,
signal,
// Router (omniroute) always streams SSE even when the
// Router always streams SSE even when the
// request omits `stream`. In non-stream mode the OpenAI SDK waits
// for the FULL body before parsing, so slow/long upstream streams
// hit the 30s/60s timeout and abort mid-generation. Streaming mode
@@ -131,7 +131,7 @@ type LLMResponseChunk = {
* `delta.content`; falls back to reasoning fields so reasoning-only models
* still produce usable aggregated text. Providers differ in the field name:
* - DeepSeek-style / Cloudflare gemma → `delta.reasoning_content`
* - mimo (via omniroute) streams reasoning in `delta.reasoning` +
* - mimo (via the router) streams reasoning in `delta.reasoning` +
* `delta.reasoning_details[].text` (content:"") — without these fallbacks
* vision aggregation came back empty ("Vision API null response").
* Exported for unit tests.
@@ -267,7 +267,7 @@ export function buildLlmParams(
reasoning: { enabled: false },
// vLLM / Qwen / litellm
chat_template_kwargs: { enable_thinking: false },
// Anthropic / Claude-format (omniroute exposes thinkingFormat
// Anthropic / Claude-format (router exposes thinkingFormat
// "claude-adaptive" / "claude-budget" on its reasoning models)
thinking: { type: "disabled" },
} as Record<string, unknown>);
@@ -149,10 +149,12 @@ export const configSchema = z
.transform((v) => v === "true")
.default(false),
AI_LLM_API_KEY: z.string().optional(),
AI_LLM_BASE_URL: z
.string()
.url()
.default("http://100.121.180.82:20128/api/v1"),
// 9router — the OpenAI-compatible router on this host (127.0.0.1:4014).
// Loopback on purpose: the gateway runs on the same machine as 9router,
// so no TLS/proxy hop is needed (and localhost bypasses 9router's
// remote-key guard). Public alias https://9router.asepharyana.my.id/v1
// works too but requires the key for every call.
AI_LLM_BASE_URL: z.string().url().default("http://127.0.0.1:4014/v1"),
AI_LLM_MODEL: z.string().default("text"),
// Vision uses the SAME router/base URL as text moderation
// (AI_LLM_BASE_URL) but a different model alias. The dedicated NVIDIA