feat(gmw): route all LLM traffic through 9router

GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA
vision endpoint onto 9router, which runs on the same host as both services
(127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls
bypass 9router's remote-key guard.

- gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1
- drop stale 'omniroute' router references from comments/docs now that the
  active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS)

Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite
(SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers
image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the
existing Qdrant collections (no reindex needed). The GMW key is already
registered in 9router's apiKeys table.

typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
This commit is contained in:
asepharyana
2026-09-24 16:05:36 +07:00
parent 750f3aa598
commit ef7708bf7d
6 changed files with 15 additions and 13 deletions
+4 -4
View File
@@ -96,10 +96,10 @@ export const configSchema = z
.transform((v) => v === "true")
.default(false),
AI_LLM_API_KEY: z.string().optional(),
AI_LLM_BASE_URL: z
.string()
.url()
.default("http://100.121.180.82:20128/api/v1"),
// 9router — OpenAI-compatible router on this host (127.0.0.1:4014).
// Loopback on purpose: backend runs on the same machine as 9router, so no
// TLS/proxy hop is needed.
AI_LLM_BASE_URL: z.string().url().default("http://127.0.0.1:4014/v1"),
AI_LLM_MODEL: z.string().default("text"),
AI_LLM_VISION_MODEL: z.string().optional(),
AI_LLM_EMBEDDING_MODEL: z.string().optional(),