GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA
vision endpoint onto 9router, which runs on the same host as both services
(127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls
bypass 9router's remote-key guard.
- gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1
- drop stale 'omniroute' router references from comments/docs now that the
active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS)
Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite
(SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers
image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the
existing Qdrant collections (no reindex needed). The GMW key is already
registered in 9router's apiKeys table.
typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
The metrics server binds METRICS_PORT (code default 9090; this host runs it
on 4018 — 4016 is occupied by another process). Docs said 4016 in three
places, which is not what the service does.
README.md was the extraction-era document (referenced winston, mock-crc.ts,
llmModerationClient.ts, indonesianTextNormalizer.ts — all long gone) and
duplicated ARCHITECTURE.md. Rewritten as a short run-the-service guide;
layout/design lives only in ARCHITECTURE.md.
MODULE_STRUCTURE.md deleted: it was a stale duplicate of ARCHITECTURE.md,
referenced by nothing but itself.
ARCHITECTURE.md updated to the post-refactor reality: app/ lifecycle split
(bootstrap/lifecycle/process-guards/metrics-collector), ai-moderation
recovery-worker + cache-prune, per-module index.ts facades, one-way
dependency rule, corrected init/shutdown/observability sections.
app/:
- bootstrap.ts 277 -> 145 lines: config guard, DB connect, client debug
logging and startup order are now named steps with a comment header
- lifecycle.ts (new): everything wired on the Discord 'ready' hook, in
explicit order (inject broadcaster -> register listeners -> start workers)
- process-guards.ts (new): SIGINT/SIGTERM/uncaughtException/unhandledRejection
in ONE place, using isTransientStreamError() instead of two duplicated
inline code lists
- metrics-collector.ts (new): AI pipeline Prometheus gauges
modules/:
- ai-moderation/index.ts + message-capture/index.ts (new): public facades so
app/ never reaches into internal files
- aiAnalyzer.ts 317 -> 146 lines: pure entry API; skip-verdict recording
extracted into recordSkip()
- recovery-worker.ts (new): stranded-message recovery + stale lane/CB pruning
- cache-prune.ts (new): 6h expired-verdict sweep, throttled + resettable
- drop 3 dead re-exports (pickBatchWithinBudget/onCircuitBreakerAlert/
getConversationKey) whose consumers import the origin files directly
No behavior change. typecheck + lint + 138 tests green; nix build OK.
- delete src/shared/index.ts fat barrel; point 10 importers at the exact
module they use (redis-channels, moderation-types, utils/pagination)
- message-capture/types.ts re-exports from shared/moderation-types directly
- shared/errors: add errorMessage() + isTransientStreamError() helpers,
replacing the repeated err-message and transient-code checks
- drop unused imports flagged by biome
No behavior change. typecheck + lint + 138 tests green.
Image messages previously blocked the whole analysis pipeline:
- conversationProcessing was a single lock per conversation; processBatch
awaited BOTH text and media worker jobs before releasing it, so a fast
text verdict sat unused until the slow vision/media batch finished
- one global LLM semaphore (AI_LLM_MAX_CONCURRENT) was shared by text and
media, so a vision backlog could starve text inference
- recovery worker gated on conversationProcessing.size
Now the queue is split into independent text/media lanes:
- conversationProcessing maps key -> Partial<Record<lane, startedAt>>;
each lane holds its own lock and frees it the moment ITS worker job
resolves (ownership-guarded clear prevents stale timers clearing newer
slots)
- two LLM semaphores: AI_LLM_MAX_CONCURRENT (text, default 8) and
AI_LLM_MEDIA_MAX_CONCURRENT (media, default 4) via
withLlmConcurrency(fn, { lane })
- batchScheduler schedules per conversation+lane (timer keys
'<key>::<lane>'); splitMessagesByLane/laneOfMessage moved to pure
analysisLanes.ts (unit-testable without Piscina)
- ai-analysis-worker batch jobs carry a lane field; per-lane active
request gauges (active_text_requests / active_media_requests)
- added tests/analysisLaneLock.test.ts (7 tests: independent lane locks,
preserving other-lane lock, clear-all, ownership guard, lane split)
Docs: ARCHITECTURE.md + AGENTS.md concurrency model updated.
typecheck/lint/test(138)/build all green.
Jev (oc/jev-1.13-free via 9router /v1/systemone) added as primary text
analyzer was underperforming. Delete the whole feature:
- jevAnalyzer.ts + its unit & live-smoke tests
- Jev-first branch in textBatchProcessor, restore pure callModerationLLM path
- AI_LLM_JEV_* config vars (zod) and .env.example entries
- @typesafe-ai/sdk dependency (+ lockfile)
Behavior: text moderation is LLM-only again, exactly as before the
Jev feature; AGENTS.md invariant 'LLM is the only judge' holds.
Show relative time, always-on username with channel context, and
collapse top emojis to two; add VIEW ALL toggle to reveal all 20
fetched reactions instead of the top 4.
* feat(gateway): Jev (System One) as primary text moderator with LLM fallback
Add TypeSafe Jev via @typesafe-ai/sdk v0.6.0 as the PRIMARY analyzer for
text-only moderation sub-batches; the existing LLM stays as the fallback
for anything Jev cannot decide confidently (per-message gate rejection or
API failure) and for media batches.
- jevAnalyzer.ts: TypeSafeClient wrapper, declarative state builder
(System One models MUST get factual state, not chat-XML — chat framing
made Jev confidently wrong on clean messages at 0.98 confidence),
message-id-keyed question builder (5 typed questions per message),
cross-consistency acceptance gate (noul↔status↔severity↔action↔category),
answer→AnalysisResult mapper, fail-open outcome.
- textBatchProcessor.ts: Jev-first per sub-batch, rejected ids + API
failures fall back to callModerationLLM; no cross-batch pollution.
- config: AI_LLM_JEV_ENABLED/API_KEY/BASE_URL/MODEL/TIMEOUT_MS/MIN_CONFIDENCE.
- .env.example: documented all 6 Jev env vars.
- tests: unit (question/state builders, gate, mapper) + live smoke
(gated behind AI_LLM_JEV_SMOKE=1) verified 4/4 accepted vs real 9router.
* fix: auto-fix code quality [skip ci]
Add tests/llmE2e.test.ts — 7 end-to-end tests driving the REAL
moderation prompt pipeline (buildSystemPrompt → XML payload → llmChat
→ parseModerationResponse) against a live model via omniroute.
Covers: clean technical content (no false positives), harassment
(flagged), username-only offenses including 'Pecinta Pria' +
sexual/provocative usernames + SARA-in-username (always warn/low,
NEVER delete — the nickname-reset path), and spam bursts.
Gated behind AI_LLM_BASE_URL + AI_LLM_API_KEY: CI (no creds) skips
the file → 216 unit tests stay green, zero LLM cost. Run locally via
pnpm test:e2e:live (scripts/run-llm-e2e.sh injects creds from bws).
Verified: 223/223 tests pass with live LLM, stability across 4 runs,
typecheck + biome clean. docs: TESTING.md. ignore .hermes/ plans.