Root cause of why Phase A/B captured zero video: @discordjs/voice is audio-only
and never sends the gateway STREAM_WATCH signal, so Discord never forwards a
member's video RTP to the bot. Live diagnostic confirmed: while members were
sharing, audio .ogg files flowed for many users but no non-opus RTP ever
arrived.
Correct path: discord.js-selfbot-v13 ships a complete native watch/record stack.
New src/modules/voice-recording/videoRecorder.ts:
- Detects streamers via voiceState.streaming on a single idempotent
voiceStateUpdate listener.
- client.voice.joinChannel() (reuses the single session alongside
@discordjs/voice) + joinStreamConnection(userId) -> STREAM_WATCH (op 20).
- receiver.createVideoStream(userId, path) -> PacketHandler -> Recorder
(ffmpeg over UDP loopback) -> Matroska .mkv, decryption handled internally.
Wired in recorder.ts (trackChannel/untrackChannel) + bootstrap.ts
(setVideoRecorderClient/setVideoRecordingsDir). All best-effort; failures log
and never break existing voice/audio. Unit tests 6/6 (single listener, watch
handshake + mkv path, skip own video, idempotence, teardown). Full suite
20 files / 176 tests green; typecheck + build + biome clean.
UNVERIFIED live yet: needs deploy + a streamer to confirm Recorder ready +
playable .mkv.
Temporary diagnostic to answer definitively whether Discord actually delivers
video RTP to the bot when someone screen-shares / turns on camera. Both
videoReceiver and screenShareAudio rely on ssrcMap emitting videoSSRC from
voice-state updates, and NO "Video SSRC appeared"/"Screen-share video started"
lines appear even after a real share+record. This logs any RTP packet whose
payload type is not Opus (120) so we can tell: (a) video RTP IS arriving but
attribution/signaling fails, vs (b) Discord sends no video at all to a
non-signaling receiver. Remove this log once the gap is understood.
ssrcMap emits "delete" when a VoiceUserData is removed (user stops sharing /
leaves voice). Hook it to close the user's open video burst ~500ms later so the
ffmpeg mux starts as soon as they stop, instead of waiting up to 5s for the idle
sweep. No-op if no burst exists; safe on normal teardown.
Phase A captured raw .h264 streams but left them as non-playable elementary
streams. Phase B adds automatic muxing: when a video burst closes, the raw
.h264 is remuxed to a self-contained MP4 via `ffmpeg -c copy` (no re-encode,
fast) with `+faststart`, waits for the write stream to fully flush first so the
mux never reads a truncated tail, and deletes the raw .h264 on success (keeping
it on failure). Output: <RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.mp4.
muxToMp4 is exported + covered by a real-ffmpeg vitest (tests/videoReceiver.test.ts):
generates a tiny baseline h264, remuxes, asserts mp4 exists/non-empty & raw deleted
(also the 5 depacketizer tests). Full gateway suite 170/170 green, tsc + biome clean.
@discordjs/voice only decrypts/forwards AUDIO (opus) — its onUdpMessage drops
every non-opus RTP packet (dist/index.mjs:2068 `!== RTP_OPUS_PAYLOAD_TYPE`, 120).
Video RTP (H264 camera + screen share, plus VP8/VP9/AV1) arrives on the same UDP
socket but was silently discarded.
New videoReceiver.ts wraps receiver.onUdpMessage (like screenShareAudio.ts):
- detects video payload types (96/98/101/102/106/116/126/127),
- decrypts them with the connection secret key/encryptionMode via the SAME
receiver.parsePacket path @discordjs/voice uses for audio (so DAVE + voice
encryption are handled identically),
- depacketizes H264 to AnnexB (single NAL, STAP-A, and FU-A fragmentation),
waiting for a keyframe (SPS/PPS/IDR) before writing,
- writes a raw .h264 elementary stream per user per burst under
<RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.h264.
Attribution: a videoSSRC→user index is built from ssrcMap updates; a proximity
fallback mirrors screenShareAudio's inferScreenShareOwner. Bot's own video is
skipped.
Unit tests: tests/videoReceiver.test.ts (AnnexB start code, keyframe gating,
FU-A reassembly, orphan-fragment tolerance) — 5/5 green.
Phase A only (capture raw h264). Phase B (ffmpeg decode+mux to MP4/WebM +
persist) and Phase C (frontend playback) are follow-ups.
A server-muted/server-deafened bot can't reliably receive/record members' audio
(and definitely can't receive video/screen share). After the voice connection
is Ready, force a REST guild-members PATCH (mute:false, deaf:false) on the self
member so the bot is auto-unmuted & undeafened on every join/reconnect.
Requires MUTE_MEMBERS + DEAFEN_MEMBERS permissions (user granted). Failure is
logged as a warning and never breaks the voice join.
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.
NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.
NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
Root-cause fixes for 'banyak miss & terpotong' in the voice->recording flow:
- subscribe BEFORE collecting user metadata. receiver.speaking 'start' fires
on the FIRST opus packet, and onUdpMessage forwards frames to the
subscription only when one exists — every frame during the old
await collectUserMetadata (a Discord REST roundtrip on cache miss) was
dropped, cutting off the start of every burst. Now subscribe synchronously
(guard first, no await in between), then fetch metadata in the background
and discard the burst if the speaker turns out to be a bot.
- one segment per burst: drop the fixed 5s RECORDING_SEGMENT_MS rotation on
the OGG path, which split continuous speech mid-word/sentence. Only the
web-PCM decoder still rotates (bounds memory).
- finalize only once the underlying file has flushed to disk (wait on the
write stream 'finish'), so upload/transcode reads a complete file.
- raise AfterSilence 3000->4000ms so natural pauses (thinking, interruptions)
don't split one utterance into several recordings.
- lower the 'too short to keep' threshold 1000->300ms so brief replies
("ya", "siap") are kept instead of dropped.
All typecheck / biome(src/) / vitest (164) green.
- nativeBuildInputs: remove cmake, rustc, cargo, git — GMW's only native deps (@discordjs/opus, sharp) are PREBUILT, no source compile needed (cmake/rust were inherited for node-datachannel which is 9router, not GMW). python3/gnumake/gcc stay as node-gyp fallback for opus.
- pruneProd: also strip .pnpm/@types+* (pure TS decls pulled into the prod graph as real deps by discord-api-types/pg-protocol, never required at runtime) and .pnpm/opusscript@* (pure-JS fallback Opus engine that prism-media only loads IF native @discordjs/opus fails — native is always present, so opusscript is never executed).
- Verified: nix flake check OK; typecheck + biome check src/ green; runtime smoke test post-prune loads @discordjs/voice, sharp, selfbot, tiktoken, piscina and encodes a frame via native opus.
Batch race guard balikin {ok:true, rows:[]} tanpa sinyal saat semua target
masih upload-pending -> processor klasifikasi semua incomplete -> fanout ke
individual queue -> di situ requeue + reschedule 250ms -> balik ke batch:
hot loop ~300ms sepanjang upload (10 siklus/3 dtk di log prod 08:13).
Fix: worker batch kini return uploadPendingIds eksplisit; classifier pure
baru (partitionBatchOutcome) partisi completed/upload_pending/incomplete/
parse_failed/api_failed; target upload-pending DEFERRED dengan poll backoff
linear (AI_ANALYSIS_UPLOAD_POLL_MS 1500 base, cap AI_ANALYSIS_MAX_UPLOAD_POLL_MS
8000), tidak pernah masuk fanout; tail shouldScheduleNext tak menimpa defer.
Test: tests/batchOutcomeClassifier.test.ts (8 kasus, pure tanpa DB/Piscina).
Analisis pertama tetap berkonteks (chat history) demi akurasi, tapi
verdict clean non-actionable (conf>=0.85) juga ditulis di bare key
tanpa konteks. Repeat teks sama di channel lain -> exact cache HIT,
bukan LLM call baru. Guard sama dgn read path; dedupe LRU per proses;
bare row tanpa embedding (tier semantic sudah global).
- Fase-1 exact-cache lookup: N query serial -> SATU query ANY($1::text[])
- Global reuse utk bare key legacy, HANYA verdict non-actionable
(clean/flagless/action=none, conf>=0.85, umur<=72h) — flagged/warn
tetap context-scoped
- Semantic cache dua-band: clean band 0.92 default, actionable tetap
0.97; di antara band -> LLM (fail-open ke akurasi)
- hit_count kini di-increment (bulk UPDATE per batch) -> hit-rate terukur
- Cache hasil wikipediaSearch di Redis (6h, hanya hasil non-kosong)
- Memoize fetchUrlSafely utk type=text (LRU 30m + in-flight dedupe)
- makeImageCacheKey strip query CDN Discord (?ex/is/hm, format/width)
-> attachment sama = satu key vision, skip re-download+re-vision
Spec: .hermes/plans/2026-08-24-ai-analysis-cache-optimization.md
Tests: +33 (cacheGuards, discordImageKeyNormalize, cacheBatchLookup)
Two independent WS handlers (useMessagesWsSync for message_created/
updated/analyzed, and useMessagesStream for message_snapshot) both
prepend live messages to the SWR list without enforcing order.
When frames arrive out-of-order (common with batched WS delivery),
the message feed gets scrambled.
Fix: add sortMessages() helper that sorts newest-first by created_at
(the list's stored order before .reverse() for display) and apply it
in every patchLists/mutate updater: message_created, message_updated,
message_analyzed, message_snapshot, and useLoadMore page appends.
Function declaration is hoisted so useLoadMore (defined above the
helper) can use it.
The fix-imports.mjs script blindly appended '.js' to every @/ alias
import, even when the source specifier already carried a .js
extension (e.g. '@/shared/config/index.js'). This produced
'index.js.js' in the emitted dist/, causing ERR_MODULE_NOT_FOUND
at startup.
This was latent: only triggered once digestScheduler.ts (which
uses @/shared/config/index.js with explicit extension) was built.
The user-reputation removal (2a8f6d9) was also blocked by this
bug — stale binary kept crashing with 'user_reputations' query
errors because it was never redeployed.
Fix: only append .js when the @/ specifier has no existing
extension. Applied to both gateway and backend scripts.
Automated public weekly summary: top categories/domains/channels + coverage rate, posted to configured webhook. Uses getDatabase() direct query (no oRPC HTTP dependency), guards one-fire-per-week on restart.
Gateway @/ alias imports use no .js extension (relative imports
keep .js). The .js suffix on @/ paths caused double-extension
ERR_MODULE_NOT_FOUND (embeddingClient.js.js) at runtime.
- Persist structured verdict (flags/severity/confidence/evidence) on
moderation_actions so the public web can show WHY a message was moderated.
- Add a persistent Qdrant archive collection (gmw_message_archive); embed
every captured message at capture time (fire-and-forget, best-effort).
- Public semantic search over the archive (backend oRPC + FE toggle on the
messages view). Both features are read-only/public and fully automatic.
Migration: 0015_add_moderation_explainability.sql
- Memoize buildSystemPrompt by (mode|channelCulture); identical signatures
now reuse the ~5k-token core instead of rebuilding per sub-batch call
(textBatchProcessor rebuilt it inside the loop; a 200-msg batch re-sent
the full system prompt ~4x). Correction tail stays per-attempt (uncached).
- Hoist URL-image -> vision evidence out of the per-sub-batch loop in
textBatchProcessor: it depends only on fetched images + full target set,
so compute once per whole batch, not per sub-batch.
- Compact system instructions: collapse 3x-duplicated 'evaluate by content
alone' statements into one standalone rule; trim output.ts channel-culture
+ context framing already covered by rules.ts/system.ts; drop duplicate
programming-error-log few-shot (id 17, covered by rules AMAN list).
- Fix misleading config default: AI_LLM_BASE_URL default -> omniroute
(gateway already runs omniroute via BWS; 9router was dead/misleading).
typecheck + lint + build green.
- resetOffensiveNickname: skip when target role sits above bot
(member.manageable) instead of hammering a doomed setNickname PATCH
that Discord rejects with 50013 'Missing Permissions'. Log the
Discord error code on failure for clear diagnosis.
- llmCaller: include contentPreview (first 200 chars) in the parse-
failure warning so non-JSON LLM responses are debuggable.
- Add wikipediaClient.ts: native fetch to Wikipedia REST/Action APIs
(search + summary), no extra npm dependency.
- Extract shared Redis cache into cacheStore.ts (decoupled from search).
- Term glossary now uses wikipediaSummary for direct article lookup.
- Remove searxngSearch.ts entirely; drop SEARXNG_BASE_URL config,
add WIKIPEDIA_LANG / WIKIPEDIA_TIMEOUT_MS.
- Rename backend searxngCalls metric to webSearchCalls.