getPendingMessagesByConversation() flips every fetched pending row to
'processing', then pickBatchWithinBudget() may stop early on the token
budget. The tail rows that did NOT make the batch were never un-claimed,
so they stayed 'processing' forever — the recovery worker reverted them
(120s) only for the next wave to re-claim them, an infinite loop of
stuck messages that never get analyzed (saw 22 rows, some recycled for
40+ minutes).
- add computeBudgetOverflowMessages() pure helper (batchBudget.ts)
- batchScheduler un-claims overflow rows back to 'pending' before
dispatching the trimmed batch
- recovery-worker now reverts stuck processing unconditionally (the old
'conversationProcessing.size > 0' guard skipped the revert when the
in-memory lock map was empty, e.g. fresh boot — exactly when stranded
rows from a previous process need rescuing)
- 3 regression tests for the overflow helper
GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA
vision endpoint onto 9router, which runs on the same host as both services
(127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls
bypass 9router's remote-key guard.
- gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1
- drop stale 'omniroute' router references from comments/docs now that the
active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS)
Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite
(SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers
image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the
existing Qdrant collections (no reindex needed). The GMW key is already
registered in 9router's apiKeys table.
typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
The metrics server binds METRICS_PORT (code default 9090; this host runs it
on 4018 — 4016 is occupied by another process). Docs said 4016 in three
places, which is not what the service does.
README.md was the extraction-era document (referenced winston, mock-crc.ts,
llmModerationClient.ts, indonesianTextNormalizer.ts — all long gone) and
duplicated ARCHITECTURE.md. Rewritten as a short run-the-service guide;
layout/design lives only in ARCHITECTURE.md.
MODULE_STRUCTURE.md deleted: it was a stale duplicate of ARCHITECTURE.md,
referenced by nothing but itself.
ARCHITECTURE.md updated to the post-refactor reality: app/ lifecycle split
(bootstrap/lifecycle/process-guards/metrics-collector), ai-moderation
recovery-worker + cache-prune, per-module index.ts facades, one-way
dependency rule, corrected init/shutdown/observability sections.
app/:
- bootstrap.ts 277 -> 145 lines: config guard, DB connect, client debug
logging and startup order are now named steps with a comment header
- lifecycle.ts (new): everything wired on the Discord 'ready' hook, in
explicit order (inject broadcaster -> register listeners -> start workers)
- process-guards.ts (new): SIGINT/SIGTERM/uncaughtException/unhandledRejection
in ONE place, using isTransientStreamError() instead of two duplicated
inline code lists
- metrics-collector.ts (new): AI pipeline Prometheus gauges
modules/:
- ai-moderation/index.ts + message-capture/index.ts (new): public facades so
app/ never reaches into internal files
- aiAnalyzer.ts 317 -> 146 lines: pure entry API; skip-verdict recording
extracted into recordSkip()
- recovery-worker.ts (new): stranded-message recovery + stale lane/CB pruning
- cache-prune.ts (new): 6h expired-verdict sweep, throttled + resettable
- drop 3 dead re-exports (pickBatchWithinBudget/onCircuitBreakerAlert/
getConversationKey) whose consumers import the origin files directly
No behavior change. typecheck + lint + 138 tests green; nix build OK.
- delete src/shared/index.ts fat barrel; point 10 importers at the exact
module they use (redis-channels, moderation-types, utils/pagination)
- message-capture/types.ts re-exports from shared/moderation-types directly
- shared/errors: add errorMessage() + isTransientStreamError() helpers,
replacing the repeated err-message and transient-code checks
- drop unused imports flagged by biome
No behavior change. typecheck + lint + 138 tests green.
Image messages previously blocked the whole analysis pipeline:
- conversationProcessing was a single lock per conversation; processBatch
awaited BOTH text and media worker jobs before releasing it, so a fast
text verdict sat unused until the slow vision/media batch finished
- one global LLM semaphore (AI_LLM_MAX_CONCURRENT) was shared by text and
media, so a vision backlog could starve text inference
- recovery worker gated on conversationProcessing.size
Now the queue is split into independent text/media lanes:
- conversationProcessing maps key -> Partial<Record<lane, startedAt>>;
each lane holds its own lock and frees it the moment ITS worker job
resolves (ownership-guarded clear prevents stale timers clearing newer
slots)
- two LLM semaphores: AI_LLM_MAX_CONCURRENT (text, default 8) and
AI_LLM_MEDIA_MAX_CONCURRENT (media, default 4) via
withLlmConcurrency(fn, { lane })
- batchScheduler schedules per conversation+lane (timer keys
'<key>::<lane>'); splitMessagesByLane/laneOfMessage moved to pure
analysisLanes.ts (unit-testable without Piscina)
- ai-analysis-worker batch jobs carry a lane field; per-lane active
request gauges (active_text_requests / active_media_requests)
- added tests/analysisLaneLock.test.ts (7 tests: independent lane locks,
preserving other-lane lock, clear-all, ownership guard, lane split)
Docs: ARCHITECTURE.md + AGENTS.md concurrency model updated.
typecheck/lint/test(138)/build all green.
Jev (oc/jev-1.13-free via 9router /v1/systemone) added as primary text
analyzer was underperforming. Delete the whole feature:
- jevAnalyzer.ts + its unit & live-smoke tests
- Jev-first branch in textBatchProcessor, restore pure callModerationLLM path
- AI_LLM_JEV_* config vars (zod) and .env.example entries
- @typesafe-ai/sdk dependency (+ lockfile)
Behavior: text moderation is LLM-only again, exactly as before the
Jev feature; AGENTS.md invariant 'LLM is the only judge' holds.
* feat(gateway): Jev (System One) as primary text moderator with LLM fallback
Add TypeSafe Jev via @typesafe-ai/sdk v0.6.0 as the PRIMARY analyzer for
text-only moderation sub-batches; the existing LLM stays as the fallback
for anything Jev cannot decide confidently (per-message gate rejection or
API failure) and for media batches.
- jevAnalyzer.ts: TypeSafeClient wrapper, declarative state builder
(System One models MUST get factual state, not chat-XML — chat framing
made Jev confidently wrong on clean messages at 0.98 confidence),
message-id-keyed question builder (5 typed questions per message),
cross-consistency acceptance gate (noul↔status↔severity↔action↔category),
answer→AnalysisResult mapper, fail-open outcome.
- textBatchProcessor.ts: Jev-first per sub-batch, rejected ids + API
failures fall back to callModerationLLM; no cross-batch pollution.
- config: AI_LLM_JEV_ENABLED/API_KEY/BASE_URL/MODEL/TIMEOUT_MS/MIN_CONFIDENCE.
- .env.example: documented all 6 Jev env vars.
- tests: unit (question/state builders, gate, mapper) + live smoke
(gated behind AI_LLM_JEV_SMOKE=1) verified 4/4 accepted vs real 9router.
* fix: auto-fix code quality [skip ci]
Add tests/llmE2e.test.ts — 7 end-to-end tests driving the REAL
moderation prompt pipeline (buildSystemPrompt → XML payload → llmChat
→ parseModerationResponse) against a live model via omniroute.
Covers: clean technical content (no false positives), harassment
(flagged), username-only offenses including 'Pecinta Pria' +
sexual/provocative usernames + SARA-in-username (always warn/low,
NEVER delete — the nickname-reset path), and spam bursts.
Gated behind AI_LLM_BASE_URL + AI_LLM_API_KEY: CI (no creds) skips
the file → 216 unit tests stay green, zero LLM cost. Run locally via
pnpm test:e2e:live (scripts/run-llm-e2e.sh injects creds from bws).
Verified: 223/223 tests pass with live LLM, stability across 4 runs,
typecheck + biome clean. docs: TESTING.md. ignore .hermes/ plans.
- rules.ts: explicit rule that sexual/provocative username terms
(Pecinta Pria, Cinta, pacar, janda, bokep, hot, seks, nude, telanjang)
MUST be flagged as offensive_username with low severity
- output.ts: output schema case for sexual/provocative username
violations, always status: warn, severity: low, never flagged/delete
No hardcoded keyword lists — fix is at the LLM prompt level only.
- new tinyFishSearch module: GET api.search.tinyfish.ai with X-API-Key,
maps top-3 to SearchResult shape, never throws (all failure modes -> [])
- wikipediaSearch: on wiki miss, one tinyfish attempt; hits cached 6h
under the same key so fallback latency is paid once
- termGlossary: on summary miss, top tinyfish hit becomes the definition
(persisted permanently like wiki defs); miss keeps 1h sentinel
- config: TINYFISH_API_KEY (empty = fallback disabled), ENABLED,
BASE_URL, TIMEOUT_MS, LOCATION, LANGUAGE knobs
- tests: 6 coverage for disabled/mapping/non-OK/network/bad-json
API key NOT committed — set TINYFISH_API_KEY in BWS gmw secrets.
Verified: typecheck + lint clean, 216/216 tests pass, live probe
'gubernur jawa barat' returned 3 mapped results
- attachmentUploader: downloadDiscordAttachment retries CDN timeouts
via retryWithBackoff (ATTACHMENT_RETRY_ATTEMPTS, 1s-8s backoff);
AbortError normalized so 403/404 refresh path never fires on timeouts
- attachmentUploader: uploadAttachmentToTele uses ATTACHMENT_RETRY_ATTEMPTS
instead of retries:0 (Tele 5xx under load was failing outright)
- wikipediaClient: wikipediaSummary retries once on abort/timeout
(100% of prod summary errors were aborts); termGlossary drops its
redundant second call (was up to 4 reqs/term under miss+retry)
Verified: typecheck + lint clean, 210/210 tests pass
Jockie Music (user 411916947773587456) posts now-playing embeds/spotify links
~1347 captured messages — every one consumed a moderation LLM call for zero
signal and contributed to batch timeouts. Config AI_SKIP_ANALYSIS_USER_IDS
(default=Jockie) skips them at ALL three analysis paths:
- queueMessageAnalysis entry (direct skip-result like age-restricted)
- batchScheduler processing (pre-batch filter)
- individual recovery path (no fallback spam for already-skipped authors)
Skip-result mirrors age_restricted: status=clean, flags=[skip_analysis_user],
action=none — stays visible in the dashboard, never analyzed.
The text model behind omniroute/9router consistently takes >45s on long-context
batches. At 45s every such batch fell through to the individual-fallback
queue which re-runs with its own timeout, then exhausted to ai_status=error.
75s keeps the bounded budget while letting the first-pass batch succeed.
- AI_LLM_MEDIA_ANALYSIS_TIMEOUT_MS 60s→120s + vision 60s→120s: vision model
via router regularly exceeded 60s, dropping media batches into the
individual-fallback chain then exhausting into ai_status=error.
- Qdrant upserts: retryWithRetry() wraps PUT /points with exponential
backoff (3 attempts, jitter) for transient 408/abort/ECONNRESET — the
41 six-hour 'Qdrant upsert failed — semantic entry skipped' warnings were
single-hop timeouts on a healthy-but-loaded Qdrant.
- LLM caller: on parse failure, attempt extractJson() structural repair of
the raw content (models with thinking disabled sometimes emit JSON as
plain text) before giving up and re-requesting.
C2 full scope persistence:
- getRecentConversationContext now returns per-turn guildId/channelId
- processMessage merges historical scope when current request is unscoped
- chatbot remembers server context across all 8 history exchanges (not just 3)
A4+Metrics:
- Add incrementCounterBy(name, delta, labels) to gateway-metrics
- Token-usage counters: llm_tokens_total{model, type, label} per batch
- Cache hit counters: moderation_cache_hits{type} exact/semantic-qdrant/semantic-pg
- Cache miss counters: moderation_cache_misses per batch
- Prometheus /metrics now exposes cost + cache hit-rate for dashboards
Performance:
- Parallel tool execution within each chatbot round (Promise.all)
- All tool results collected before sending to model (ordering preserved)
- Tool failure now logged with structured warning (chatbot.tools.ts)
Verified: backend tsc+biome 37/37, discord-gateway tsc+biome 210/210
Even with the prompt firewall (username vs content), the LLM can still
occasionally mis-apply a content-level zero-tolerance flag (sara /
conflict_instigation) to a message whose ONLY violation is the username
(e.g. 'matikanetanyahu'). The auto-delete eligibility check only recognized
exact offensive_username flags, so such false positives still deleted the
message.
Add a belt-and-suspenders guard in isNicknameOnlyViolation: if the flag set
is entirely username-attributable (offensive_username/sara/conflict_instigation)
AND the analysis text corroborates that the violation is username-only with
clean message content, route to nickname-reset instead of message deletion.
Adds 6 test cases covering the real matikanetanyahu scenario and the
false-positive/negative boundaries.
Add explicit FIREWALL PENILAIAN rule separating username assessment from
message content assessment. Username containing political/religious terms
(matikanetanyahu etc) is assessed as offensive_username with severity low
— never triggers sara/conflict_instigation zero tolerance for content.
Changes:
- rules.ts: Add FIREWALL section before LARANGAN BERAT; clarify each
zero-tolerance rule applies to ISI PESAN only; update hierarchy #9
- examples.ts: Fix example #11 status flagged→warn for clean username;
add example #11b with matikanetanyahu case
- output.ts: Explicit status/action mapping for username-only violations
result.score was required by zod; the LLM (gemini-3.5-flash-lite via 9router)
occasionally omits it for media batches, hard-failing the whole batch parse
('Zod validation failed: expected number, received undefined' at
results[0].score). Callers already null-coalesce (result.score ?? 0) and the
parser clampScore()s it, so requiring it only caused parse failures.
Adds regression tests: media-batch without score parses (score->0), and
score-present responses still parse with the value.
A selfbot (user token) cannot auto-detect other members' camera/share
(no VOICE_STATE_UPDATE for others, 403 on member fetch). The only
selfbot-viable path to capture another member's SCREEN SHARE is an
operator-initiated STREAM_WATCH (gateway op 20, not gated on bot-vs-user).
Add video:watch / video:unwatch Redis commands routed via the existing
command handler to startStreamWatch/stopStreamWatch, which then does the
DAVE handshake + per-burst MP4 segmentation + DB insert + Tele upload
(already implemented in streamWatchReceiver).
- new VideoHandler (command-handler/video.handler.ts)
- register video:watch / video:unwatch in handler-registry + CommandHandler
- command constants COMMAND_VIDEO_WATCH / COMMAND_VIDEO_UNWATCH
- resolve active voice channel from voice controller + client cache
- 8 unit tests (videoHandler.test.ts)
- biome fixes for pre-existing test import ordering
All green: typecheck, build, lint (174 files), 200 tests.
Dump all WS raw event types after 10s (see if GUILD_CREATE exists at all),
and log broadcaster_user_ids + sessions.voice from READY payload.
Selfbot may RESUME (skip GUILD_CREATE) — voice data may only be in READY.
Temporary diagnostic to see whether GUILD_CREATE.voice_states actually
reaches the selfbot raw listener (selfbot-v13 emits everything via
WebSocketShard Events.RAW). Will remove once root cause confirmed.