feat(gmw): route all LLM traffic through 9router
GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA vision endpoint onto 9router, which runs on the same host as both services (127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls bypass 9router's remote-key guard. - gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1 - drop stale 'omniroute' router references from comments/docs now that the active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS) Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite (SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the existing Qdrant collections (no reindex needed). The GMW key is already registered in 9router's apiKeys table. typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
This commit is contained in:
@@ -96,10 +96,10 @@ export const configSchema = z
|
||||
.transform((v) => v === "true")
|
||||
.default(false),
|
||||
AI_LLM_API_KEY: z.string().optional(),
|
||||
AI_LLM_BASE_URL: z
|
||||
.string()
|
||||
.url()
|
||||
.default("http://100.121.180.82:20128/api/v1"),
|
||||
// 9router — OpenAI-compatible router on this host (127.0.0.1:4014).
|
||||
// Loopback on purpose: backend runs on the same machine as 9router, so no
|
||||
// TLS/proxy hop is needed.
|
||||
AI_LLM_BASE_URL: z.string().url().default("http://127.0.0.1:4014/v1"),
|
||||
AI_LLM_MODEL: z.string().default("text"),
|
||||
AI_LLM_VISION_MODEL: z.string().optional(),
|
||||
AI_LLM_EMBEDDING_MODEL: z.string().optional(),
|
||||
|
||||
@@ -54,7 +54,7 @@ src/
|
||||
**Never** reintroduce regex/heuristic content classification.
|
||||
2. **Discord tokens sanitized** before reaching LLM (`discordTokens.ts`).
|
||||
3. **Semantic cache is batched** — one embed call + one Qdrant batch search.
|
||||
4. **Streaming is mandatory** against the omniroute base URL.
|
||||
4. **Streaming is mandatory** against the router base URL.
|
||||
|
||||
## AI moderation pipeline
|
||||
|
||||
|
||||
@@ -194,5 +194,5 @@ pipeline gauges registered by `app/metrics-collector.ts` —
|
||||
numeric snowflake IDs never trigger false positives.
|
||||
- **Semantic cache is batched** (one embed call + one Qdrant batch search),
|
||||
not N sequential round-trips. `ensureQdrantCollection` is memoized.
|
||||
- **Streaming is mandatory** against the omniroute base URL (non-stream waits for
|
||||
- **Streaming is mandatory** against the router base URL (non-stream waits for
|
||||
the full body and times out). `llmClient` aggregates SSE chunks.
|
||||
|
||||
@@ -92,7 +92,7 @@ export async function callModerationLLM(
|
||||
jsonResponse: { type: "json_object" },
|
||||
retries: 0,
|
||||
signal,
|
||||
// Router (omniroute) always streams SSE even when the
|
||||
// Router always streams SSE even when the
|
||||
// request omits `stream`. In non-stream mode the OpenAI SDK waits
|
||||
// for the FULL body before parsing, so slow/long upstream streams
|
||||
// hit the 30s/60s timeout and abort mid-generation. Streaming mode
|
||||
|
||||
@@ -131,7 +131,7 @@ type LLMResponseChunk = {
|
||||
* `delta.content`; falls back to reasoning fields so reasoning-only models
|
||||
* still produce usable aggregated text. Providers differ in the field name:
|
||||
* - DeepSeek-style / Cloudflare gemma → `delta.reasoning_content`
|
||||
* - mimo (via omniroute) streams reasoning in `delta.reasoning` +
|
||||
* - mimo (via the router) streams reasoning in `delta.reasoning` +
|
||||
* `delta.reasoning_details[].text` (content:"") — without these fallbacks
|
||||
* vision aggregation came back empty ("Vision API null response").
|
||||
* Exported for unit tests.
|
||||
@@ -267,7 +267,7 @@ export function buildLlmParams(
|
||||
reasoning: { enabled: false },
|
||||
// vLLM / Qwen / litellm
|
||||
chat_template_kwargs: { enable_thinking: false },
|
||||
// Anthropic / Claude-format (omniroute exposes thinkingFormat
|
||||
// Anthropic / Claude-format (router exposes thinkingFormat
|
||||
// "claude-adaptive" / "claude-budget" on its reasoning models)
|
||||
thinking: { type: "disabled" },
|
||||
} as Record<string, unknown>);
|
||||
|
||||
@@ -149,10 +149,12 @@ export const configSchema = z
|
||||
.transform((v) => v === "true")
|
||||
.default(false),
|
||||
AI_LLM_API_KEY: z.string().optional(),
|
||||
AI_LLM_BASE_URL: z
|
||||
.string()
|
||||
.url()
|
||||
.default("http://100.121.180.82:20128/api/v1"),
|
||||
// 9router — the OpenAI-compatible router on this host (127.0.0.1:4014).
|
||||
// Loopback on purpose: the gateway runs on the same machine as 9router,
|
||||
// so no TLS/proxy hop is needed (and localhost bypasses 9router's
|
||||
// remote-key guard). Public alias https://9router.asepharyana.my.id/v1
|
||||
// works too but requires the key for every call.
|
||||
AI_LLM_BASE_URL: z.string().url().default("http://127.0.0.1:4014/v1"),
|
||||
AI_LLM_MODEL: z.string().default("text"),
|
||||
// Vision uses the SAME router/base URL as text moderation
|
||||
// (AI_LLM_BASE_URL) but a different model alias. The dedicated NVIDIA
|
||||
|
||||
Reference in New Issue
Block a user