@discordjs/voice only decrypts/forwards AUDIO (opus) — its onUdpMessage drops
every non-opus RTP packet (dist/index.mjs:2068 `!== RTP_OPUS_PAYLOAD_TYPE`, 120).
Video RTP (H264 camera + screen share, plus VP8/VP9/AV1) arrives on the same UDP
socket but was silently discarded.
New videoReceiver.ts wraps receiver.onUdpMessage (like screenShareAudio.ts):
- detects video payload types (96/98/101/102/106/116/126/127),
- decrypts them with the connection secret key/encryptionMode via the SAME
receiver.parsePacket path @discordjs/voice uses for audio (so DAVE + voice
encryption are handled identically),
- depacketizes H264 to AnnexB (single NAL, STAP-A, and FU-A fragmentation),
waiting for a keyframe (SPS/PPS/IDR) before writing,
- writes a raw .h264 elementary stream per user per burst under
<RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.h264.
Attribution: a videoSSRC→user index is built from ssrcMap updates; a proximity
fallback mirrors screenShareAudio's inferScreenShareOwner. Bot's own video is
skipped.
Unit tests: tests/videoReceiver.test.ts (AnnexB start code, keyframe gating,
FU-A reassembly, orphan-fragment tolerance) — 5/5 green.
Phase A only (capture raw h264). Phase B (ffmpeg decode+mux to MP4/WebM +
persist) and Phase C (frontend playback) are follow-ups.
A server-muted/server-deafened bot can't reliably receive/record members' audio
(and definitely can't receive video/screen share). After the voice connection
is Ready, force a REST guild-members PATCH (mute:false, deaf:false) on the self
member so the bot is auto-unmuted & undeafened on every join/reconnect.
Requires MUTE_MEMBERS + DEAFEN_MEMBERS permissions (user granted). Failure is
logged as a warning and never breaks the voice join.
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.
NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
Root-cause fixes for 'banyak miss & terpotong' in the voice->recording flow:
- subscribe BEFORE collecting user metadata. receiver.speaking 'start' fires
on the FIRST opus packet, and onUdpMessage forwards frames to the
subscription only when one exists — every frame during the old
await collectUserMetadata (a Discord REST roundtrip on cache miss) was
dropped, cutting off the start of every burst. Now subscribe synchronously
(guard first, no await in between), then fetch metadata in the background
and discard the burst if the speaker turns out to be a bot.
- one segment per burst: drop the fixed 5s RECORDING_SEGMENT_MS rotation on
the OGG path, which split continuous speech mid-word/sentence. Only the
web-PCM decoder still rotates (bounds memory).
- finalize only once the underlying file has flushed to disk (wait on the
write stream 'finish'), so upload/transcode reads a complete file.
- raise AfterSilence 3000->4000ms so natural pauses (thinking, interruptions)
don't split one utterance into several recordings.
- lower the 'too short to keep' threshold 1000->300ms so brief replies
("ya", "siap") are kept instead of dropped.
All typecheck / biome(src/) / vitest (164) green.
- mediaDownloader: flip URL candidate order so discord_url is tried
before uploaded_url (uploaded_url is archive-only fallback)
- ai-analysis-worker: remove upload-pending race guard that blocked
analysis until Tele upload completed; analysis now runs immediately
on the Discord CDN URL
- batchProcessor: remove upload-pending defer/poll-backoff logic
- individualFallbackProcessor: remove upload_pending requeue loop
- batchOutcomeClassifier/fallbackResultClassifier: drop upload_pending
classification (no longer needed)
- tests: update batchOutcomeClassifier + fallbackResultClassifier tests
to reflect removed upload_pending signal
Switch GMW's AI LLM base URL from 9router (https://9router.asepharyana.my.id/v1)
to omniroute on imrnes (http://100.121.180.82:20128/api/v1).
- Update default AI_LLM_BASE_URL in discord-gateway + backend config schemas
- Update .env.example documentation
- Update all 9router references in comments/docs/tests to omniroute
- Production BWS secret gmw_ai_llm_base_url already updated
Omniroute uses /api/v1 prefix (not /v1 like 9router), so the base URL
now correctly points at the right API path for the OpenAI SDK.
Prevent recurring infinite restart loop (389x crash) caused by drizzle
re-applying already-applied migrations when public.__drizzle_migrations
tracking is empty/partial.
- 0017/0018: ADD COLUMN IF NOT EXISTS (re-run safe)
- 0019: DO-block rename that handles all prior states (server_name-only,
both columns, or server_nick-only) so it never errors or double-renames
- seedDrizzleHistory: reconcile tracked created_at to the journal's latest
'when' when the schema already reflects the latest migration, instead of
early-returning on an existing-but-empty/partial tracking table
Rename server_name (guild name) to server_nick and populate it from
the member's server-specific display name (metadata.member.displayName)
at write time. This is what the moderation dashboard should show as
TARGET — e.g. server nick 'Bandar Togel「✔ ᵛᵉʳᶦᶠᶦᵉᵈ 」' for global
username '.nichiyobi'. Backfilled 210 existing actions from messages
metadata (reset_nickname rows now show 'Sarjana .jav', 'Penindas
Minoritas', etc). Frontend TARGET shows server nick with global
username as secondary context.
Replace static OFFENSIVE_USERNAME_KEYWORDS substring matching with a
lightweight LLM call that evaluates whether a global username violates
server rules (gambling, scam, NSFW, SARA, etc). Fail-open design:
if the LLM call fails/times out, the nickname reset still completes.
After resetting an offensive server nickname to the global username,
check the global username against gambling/scam keyword list. If it
also violates, generate a random 'UserXXXXX' nickname to prevent
circumvention via offensive global usernames.
Denormalize guild name alongside username so the moderation dashboard
shows both TARGET and server even after message table purges.
Migration 0018. Frontend displays 'username · server_name' in TARGET.
Gateway (transmitter.ts):
- Auto-stop on FFmpeg crash: non-zero exit triggers stop() to prevent
silent audio loss and resource leaks
- Voice activity timeout (10s): auto-stops transmitter when no PCM
received, preventing dead-air CPU waste on backgrounded tabs
- Stderr cap (4KB): prevents unbounded memory growth in long sessions
Gateway (voice.handler.ts):
- Double-check voiceController.getStatus().connected before starting
transmitter — detects stale player state after gateway disconnect
Frontend (context.tsx):
- Force-refetch voice status on WS reconnect — UI converges in <1s
instead of waiting up to 4s for SWR poll interval
All: tsc clean, biome clean
Two bugs causing 23 spurious 'error' logs after successful deletions:
1. batchProcessor switch missing 'completed' case: partitionBatchOutcome
returns 'completed' for successful messages, but the switch only handled
'upload_pending' and 'api_failed'. Successful messages fell through to
default → re-enqueued to individual fallback → re-analyzed → re-delete
attempt → error (message already gone from Discord). Now explicitly
skips 'completed' messages.
2. isAlreadyDeletedError only caught codes 10008/404. Discord also returns
10003 (Unknown Channel) and 50001 (Missing Access) when a message or
channel is gone. Added these codes plus text-based fallback matching
'Unknown Message'/'Unknown Channel'.
Impact: eliminates ~23 redundant error logs per day + stops wasted LLM
calls re-analyzing already-processed messages.
isEligibleForAutoDelete was rejecting messages where recommendedAction
was 'warn' or 'review' — only 'delete' and 'escalate' were accepted.
This caused 28+ medium-severity flagged messages to be logged as
'not_eligible' instead of being auto-deleted.
Changes:
- deriveRecommendedAction: return 'delete' for flagged+medium severity
(previously only critical/high triggered delete; medium got 'review')
- isEligibleForAutoDelete: accept 'warn' as valid recommendedAction
alongside 'delete' and 'escalate'
Impact: ~28 pending medium-severity flagged messages + all future
'warn'-action flagged messages will now be eligible for auto-deletion.
- Switch AI_LLM_BASE_URL from omniroute.imrnes.team to 9router.asepharyana.my.id
- Keep AI_LLM_MODEL as 'text' (9router uses alias-based routing, not bare names)
- Update .env.example comments to document 9router
- Per user: multimodal stays 'multimodal' alias
API verified: curl to 9router/v1/chat/completions with model 'text'
returns HTTP 200 (OpenAI-compatible format)
- Change AI_LLM_BASE_URL default from omniroute.imrnes.team to 9router.asepharyana.my.id
- Update AI_LLM_MODEL default from 'text' to 'claude-opus-5' (bare model name
compatible with 9router/OpenAI-compatible router)
- Update .env.example and inline comments to reflect 9router
- discord-gateway config now matches backend (which already uses 9router)
Batch race guard balikin {ok:true, rows:[]} tanpa sinyal saat semua target
masih upload-pending -> processor klasifikasi semua incomplete -> fanout ke
individual queue -> di situ requeue + reschedule 250ms -> balik ke batch:
hot loop ~300ms sepanjang upload (10 siklus/3 dtk di log prod 08:13).
Fix: worker batch kini return uploadPendingIds eksplisit; classifier pure
baru (partitionBatchOutcome) partisi completed/upload_pending/incomplete/
parse_failed/api_failed; target upload-pending DEFERRED dengan poll backoff
linear (AI_ANALYSIS_UPLOAD_POLL_MS 1500 base, cap AI_ANALYSIS_MAX_UPLOAD_POLL_MS
8000), tidak pernah masuk fanout; tail shouldScheduleNext tak menimpa defer.
Test: tests/batchOutcomeClassifier.test.ts (8 kasus, pure tanpa DB/Piscina).
Analisis pertama tetap berkonteks (chat history) demi akurasi, tapi
verdict clean non-actionable (conf>=0.85) juga ditulis di bare key
tanpa konteks. Repeat teks sama di channel lain -> exact cache HIT,
bukan LLM call baru. Guard sama dgn read path; dedupe LRU per proses;
bare row tanpa embedding (tier semantic sudah global).
- Fase-1 exact-cache lookup: N query serial -> SATU query ANY($1::text[])
- Global reuse utk bare key legacy, HANYA verdict non-actionable
(clean/flagless/action=none, conf>=0.85, umur<=72h) — flagged/warn
tetap context-scoped
- Semantic cache dua-band: clean band 0.92 default, actionable tetap
0.97; di antara band -> LLM (fail-open ke akurasi)
- hit_count kini di-increment (bulk UPDATE per batch) -> hit-rate terukur
- Cache hasil wikipediaSearch di Redis (6h, hanya hasil non-kosong)
- Memoize fetchUrlSafely utk type=text (LRU 30m + in-flight dedupe)
- makeImageCacheKey strip query CDN Discord (?ex/is/hm, format/width)
-> attachment sama = satu key vision, skip re-download+re-vision
Spec: .hermes/plans/2026-08-24-ai-analysis-cache-optimization.md
Tests: +33 (cacheGuards, discordImageKeyNormalize, cacheBatchLookup)
- pickBatchWithinBudget: stop di overflow pertama (break), bukan skip —
batch tetap prefix kronologis tanpa gap analisis di tengah timeline.
Diekstrak ke batchBudget.ts (pure, estimator di-inject) + regression test.
- callModerationLLM: param opsional maxTokens; text/media caller menghitung
ceiling dari estimasi prompt (floor 2048, cap 16384) — batch kecil tak
lagi reserve window completion 16k.
- getPending/IncompleteMessagesByConversation: sort hasil UPDATE..RETURNING
by created_at ASC — Postgres tak menjamin urutan, konsumen (anchor konteks
messages[0], prefix batch) bergantung pada urutan kronologis.
- normalizeStoredStatus(): exact-hash & semantic (Qdrant/PG) cache reader
sebelumnya menipiskan 'warn' jadi 'flagged'/'clean' (type narrowing
legacy clean|flagged) — merusak gating auto-delete & label dashboard.
Kini status tersimpan dipertahankan penuh (clean/warn/flagged).
- prompts: hapus referensi <user_history> yang tak pernah di-inject,
SearXNG -> Wikipedia (sudah migrasi), referensi section yang tak ada,
typo 'secifik', dan baris list rusak '|-'.
- moderationBuilders: buang dead code buildUserProfilesBlock/
buildUserProfileRef/UserProfileEntry/buildUserHistoryXml (tanpa caller
produksi sejak context minimization) + test-nya.
- test baru: tests/storedStatusNormalization.test.ts (regresi warn).
upload.asepharyana.my.id redirect ke file mp3 tunggal; generic extractor
yt-dlp expose format ID '0' sehingga '-f bestaudio' gagal 'Requested
format is not available'. Chain bestaudio[ext=m4a]/bestaudio/best tetap
dapat m4a di YouTube dan jatuh ke 'best' untuk direct file.
SearXNG was already replaced by Wikipedia REST/Action APIs (wikipediaClient.ts).
Update comments to reflect the current implementation: term glossary now
resolves definitions via Wikipedia → Redis → Postgres cache chain, with no
SearXNG dependency.
Discord GoLive sends screen-share audio on a separate SSRC from the
user's microphone. In @discordjs/voice v0.19, VoiceReceiver.onUdpMessage
silently drops packets for SSRCs not in ssrcMap (which is only populated
from VOICE_STATE_UPDATE/VOICE_SERVER_UPDATE). This caused screen-share
audio to never trigger receiver.speaking and never reach the speakingHandler.
Fix: hookScreenShareAudio() wraps onUdpMessage to:
1. Detect incoming RTP packets with unknown SSRCs (OPRUS payload type 120)
2. Infer the owning userId by proximity to known audioSSRC
3. Clone the user's VoiceUserData into ssrcMap under the new SSRC
4. Let the original handler decrypt and forward to the subscription stream
5. Listen on ssrcMap 'create'/'update' events for video SSRC changes
Also removes the broken initial approach (polling ssrcMap which never
contains screen-share SSRCs).
Automated public weekly summary: top categories/domains/channels + coverage rate, posted to configured webhook. Uses getDatabase() direct query (no oRPC HTTP dependency), guards one-fire-per-week on restart.
Gateway @/ alias imports use no .js extension (relative imports
keep .js). The .js suffix on @/ paths caused double-extension
ERR_MODULE_NOT_FOUND (embeddingClient.js.js) at runtime.
- Persist structured verdict (flags/severity/confidence/evidence) on
moderation_actions so the public web can show WHY a message was moderated.
- Add a persistent Qdrant archive collection (gmw_message_archive); embed
every captured message at capture time (fire-and-forget, best-effort).
- Public semantic search over the archive (backend oRPC + FE toggle on the
messages view). Both features are read-only/public and fully automatic.
Migration: 0015_add_moderation_explainability.sql
- Memoize buildSystemPrompt by (mode|channelCulture); identical signatures
now reuse the ~5k-token core instead of rebuilding per sub-batch call
(textBatchProcessor rebuilt it inside the loop; a 200-msg batch re-sent
the full system prompt ~4x). Correction tail stays per-attempt (uncached).
- Hoist URL-image -> vision evidence out of the per-sub-batch loop in
textBatchProcessor: it depends only on fetched images + full target set,
so compute once per whole batch, not per sub-batch.
- Compact system instructions: collapse 3x-duplicated 'evaluate by content
alone' statements into one standalone rule; trim output.ts channel-culture
+ context framing already covered by rules.ts/system.ts; drop duplicate
programming-error-log few-shot (id 17, covered by rules AMAN list).
- Fix misleading config default: AI_LLM_BASE_URL default -> omniroute
(gateway already runs omniroute via BWS; 9router was dead/misleading).
typecheck + lint + build green.
- resetOffensiveNickname: skip when target role sits above bot
(member.manageable) instead of hammering a doomed setNickname PATCH
that Discord rejects with 50013 'Missing Permissions'. Log the
Discord error code on failure for clear diagnosis.
- llmCaller: include contentPreview (first 200 chars) in the parse-
failure warning so non-JSON LLM responses are debuggable.
- Add wikipediaClient.ts: native fetch to Wikipedia REST/Action APIs
(search + summary), no extra npm dependency.
- Extract shared Redis cache into cacheStore.ts (decoupled from search).
- Term glossary now uses wikipediaSummary for direct article lookup.
- Remove searxngSearch.ts entirely; drop SEARXNG_BASE_URL config,
add WIKIPEDIA_LANG / WIKIPEDIA_TIMEOUT_MS.
- Rename backend searxngCalls metric to webSearchCalls.
User: 'jangan ada reputasi juga' — no profile, no reputation in the prompt,
raw messages only.
- textBatchProcessor: drop initializeUserReputation fetch + <user_reputation>
tag injection (kept the minimal <message> tag + reply/reference context).
- visionAnalyzer (prepareMediaMessage): same removal.
- prompts/system.ts + prompts/output.ts: replace <user_reputation>/<user_history>
instructions with an explicit 'no per-user profile/reputation context'
note so the LLM judges purely on message content + conversation/web/location.
- mediaBatchProcessor: fix stale comment.
Trust/infraction state is STILL written to the DB (userReputationsTable) for
enforcement — only the LLM context injection is removed, so moderation
actions (mute/ban via infraction thresholds) keep working.
Net: even smaller prompts (no per-user context at all) → more messages fit
per request, and one fewer DB round-trip per unique user per sub-batch.
tsc, biome, vitest (129) all clean.
User insight: personal profile summaries bloat the prompt (less room per
request) and add a per-user DB/Redis round-trip for little moderation signal.
Only the behavioural <user_reputation> history is kept.
- textBatchProcessor: stop fetching getUserProfile; remove <user_profiles>
block + <user_profile_ref> from message tags. Keep <user_reputation>.
- mediaBatchProcessor + visionAnalyzer: same removal (profile fetch + ref).
- prompts/system.ts + prompts/output.ts: drop stale <user_profiles>/
<user_profile_ref> instructions; point LLM at <user_reputation> instead.
- aiAnalyzer: gate userProfileLearner behind AI_USER_PROFILE_LEARNING_ENABLED
(default false) — generates profiles nobody reads, pure LLM/DB waste.
- Add AI_USER_PROFILE_LEARNING_ENABLED config knob.
Net: smaller prompts (more messages fit per request), fewer DB round-trips
per sub-batch, and no background LLM calls learning unused profiles.
tsc, biome, vitest (129) all clean.
User insight: rather than many small per-batch API requests, pack many
messages into ONE request so a burst is analyzed with far fewer calls.
- AI_LLM_TEXT_BATCH_SIZE 20 -> 60 (one request now carries ~3x more messages).
- AI_ANALYSIS_MAX_TARGET_TOKENS 4000 -> 14000 (the scheduler's token-budget
gate was trimming pending messages to ~20 before they reached the sub-batch
splitter; raising it lets ~60 messages through to a single LLM call).
- AI_LLM_TEXT_ANALYSIS_TIMEOUT_MS 30000 -> 45000 (one larger call needs more
headroom; gemini-flash-lite has a 1M-token context so 14k+8k is trivial).
Net effect when ramai: a 60-message burst = 1-2 API calls instead of 3+,
less semaphore contention, faster throughput.