Piscina worker threads outlive process.exit() and linger as orphaned
processes holding DB connections/locks after a deploy restart. Two live
gateways then fight over the same messages table rows (one claims
processing, the other reverts), which left messages stuck in
ai_status='processing' forever.
Destroy both worker pools before closing the DB, with a 5s fallback so
a hung vision job cannot block shutdown indefinitely.
getPendingMessagesByConversation() flips every fetched pending row to
'processing', then pickBatchWithinBudget() may stop early on the token
budget. The tail rows that did NOT make the batch were never un-claimed,
so they stayed 'processing' forever — the recovery worker reverted them
(120s) only for the next wave to re-claim them, an infinite loop of
stuck messages that never get analyzed (saw 22 rows, some recycled for
40+ minutes).
- add computeBudgetOverflowMessages() pure helper (batchBudget.ts)
- batchScheduler un-claims overflow rows back to 'pending' before
dispatching the trimmed batch
- recovery-worker now reverts stuck processing unconditionally (the old
'conversationProcessing.size > 0' guard skipped the revert when the
in-memory lock map was empty, e.g. fresh boot — exactly when stranded
rows from a previous process need rescuing)
- 3 regression tests for the overflow helper
GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA
vision endpoint onto 9router, which runs on the same host as both services
(127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls
bypass 9router's remote-key guard.
- gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1
- drop stale 'omniroute' router references from comments/docs now that the
active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS)
Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite
(SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers
image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the
existing Qdrant collections (no reindex needed). The GMW key is already
registered in 9router's apiKeys table.
typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
The metrics server binds METRICS_PORT (code default 9090; this host runs it
on 4018 — 4016 is occupied by another process). Docs said 4016 in three
places, which is not what the service does.
README.md was the extraction-era document (referenced winston, mock-crc.ts,
llmModerationClient.ts, indonesianTextNormalizer.ts — all long gone) and
duplicated ARCHITECTURE.md. Rewritten as a short run-the-service guide;
layout/design lives only in ARCHITECTURE.md.
MODULE_STRUCTURE.md deleted: it was a stale duplicate of ARCHITECTURE.md,
referenced by nothing but itself.
ARCHITECTURE.md updated to the post-refactor reality: app/ lifecycle split
(bootstrap/lifecycle/process-guards/metrics-collector), ai-moderation
recovery-worker + cache-prune, per-module index.ts facades, one-way
dependency rule, corrected init/shutdown/observability sections.
app/:
- bootstrap.ts 277 -> 145 lines: config guard, DB connect, client debug
logging and startup order are now named steps with a comment header
- lifecycle.ts (new): everything wired on the Discord 'ready' hook, in
explicit order (inject broadcaster -> register listeners -> start workers)
- process-guards.ts (new): SIGINT/SIGTERM/uncaughtException/unhandledRejection
in ONE place, using isTransientStreamError() instead of two duplicated
inline code lists
- metrics-collector.ts (new): AI pipeline Prometheus gauges
modules/:
- ai-moderation/index.ts + message-capture/index.ts (new): public facades so
app/ never reaches into internal files
- aiAnalyzer.ts 317 -> 146 lines: pure entry API; skip-verdict recording
extracted into recordSkip()
- recovery-worker.ts (new): stranded-message recovery + stale lane/CB pruning
- cache-prune.ts (new): 6h expired-verdict sweep, throttled + resettable
- drop 3 dead re-exports (pickBatchWithinBudget/onCircuitBreakerAlert/
getConversationKey) whose consumers import the origin files directly
No behavior change. typecheck + lint + 138 tests green; nix build OK.
- delete src/shared/index.ts fat barrel; point 10 importers at the exact
module they use (redis-channels, moderation-types, utils/pagination)
- message-capture/types.ts re-exports from shared/moderation-types directly
- shared/errors: add errorMessage() + isTransientStreamError() helpers,
replacing the repeated err-message and transient-code checks
- drop unused imports flagged by biome
No behavior change. typecheck + lint + 138 tests green.
Image messages previously blocked the whole analysis pipeline:
- conversationProcessing was a single lock per conversation; processBatch
awaited BOTH text and media worker jobs before releasing it, so a fast
text verdict sat unused until the slow vision/media batch finished
- one global LLM semaphore (AI_LLM_MAX_CONCURRENT) was shared by text and
media, so a vision backlog could starve text inference
- recovery worker gated on conversationProcessing.size
Now the queue is split into independent text/media lanes:
- conversationProcessing maps key -> Partial<Record<lane, startedAt>>;
each lane holds its own lock and frees it the moment ITS worker job
resolves (ownership-guarded clear prevents stale timers clearing newer
slots)
- two LLM semaphores: AI_LLM_MAX_CONCURRENT (text, default 8) and
AI_LLM_MEDIA_MAX_CONCURRENT (media, default 4) via
withLlmConcurrency(fn, { lane })
- batchScheduler schedules per conversation+lane (timer keys
'<key>::<lane>'); splitMessagesByLane/laneOfMessage moved to pure
analysisLanes.ts (unit-testable without Piscina)
- ai-analysis-worker batch jobs carry a lane field; per-lane active
request gauges (active_text_requests / active_media_requests)
- added tests/analysisLaneLock.test.ts (7 tests: independent lane locks,
preserving other-lane lock, clear-all, ownership guard, lane split)
Docs: ARCHITECTURE.md + AGENTS.md concurrency model updated.
typecheck/lint/test(138)/build all green.
Jev (oc/jev-1.13-free via 9router /v1/systemone) added as primary text
analyzer was underperforming. Delete the whole feature:
- jevAnalyzer.ts + its unit & live-smoke tests
- Jev-first branch in textBatchProcessor, restore pure callModerationLLM path
- AI_LLM_JEV_* config vars (zod) and .env.example entries
- @typesafe-ai/sdk dependency (+ lockfile)
Behavior: text moderation is LLM-only again, exactly as before the
Jev feature; AGENTS.md invariant 'LLM is the only judge' holds.
Show relative time, always-on username with channel context, and
collapse top emojis to two; add VIEW ALL toggle to reveal all 20
fetched reactions instead of the top 4.
* feat(gateway): Jev (System One) as primary text moderator with LLM fallback
Add TypeSafe Jev via @typesafe-ai/sdk v0.6.0 as the PRIMARY analyzer for
text-only moderation sub-batches; the existing LLM stays as the fallback
for anything Jev cannot decide confidently (per-message gate rejection or
API failure) and for media batches.
- jevAnalyzer.ts: TypeSafeClient wrapper, declarative state builder
(System One models MUST get factual state, not chat-XML — chat framing
made Jev confidently wrong on clean messages at 0.98 confidence),
message-id-keyed question builder (5 typed questions per message),
cross-consistency acceptance gate (noul↔status↔severity↔action↔category),
answer→AnalysisResult mapper, fail-open outcome.
- textBatchProcessor.ts: Jev-first per sub-batch, rejected ids + API
failures fall back to callModerationLLM; no cross-batch pollution.
- config: AI_LLM_JEV_ENABLED/API_KEY/BASE_URL/MODEL/TIMEOUT_MS/MIN_CONFIDENCE.
- .env.example: documented all 6 Jev env vars.
- tests: unit (question/state builders, gate, mapper) + live smoke
(gated behind AI_LLM_JEV_SMOKE=1) verified 4/4 accepted vs real 9router.
* fix: auto-fix code quality [skip ci]