The bot's own VOICE_STATE_UPDATE showed server-level deaf:true — a
server-deafened member is NOT sent the streamer's audiovisual RTP by Discord,
which is the likely reason no H264 arrives despite the DAVE watch reaching
Ready. The previous fire-and-forget forceSelfServerUnmuteUndeafen() ran once
after the first Ready join and silently reverted on reconnect/restart.
- Export forceSelfServerUnmuteUndeafen from recorder.ts; re-assert it (with
read-back verification logging stillDeaf) at the START of every
startStreamWatch() before STREAM_WATCH is sent (dynamic import avoids the
recorder <-> videoRecorder <-> streamWatchReceiver module cycle).
- Re-assert it again after a successful voice reconnect.
- startStreamWatch() is now async; callers use void.
maxLen stayed 72 across 243 packets (no real H264, which is hundreds+ bytes) —
only 44-72-byte RTP packets on PT 76/72/73 arrive. Add ssrc + first-32-bytes
hex so we can identify exactly what Discord sends to the watch socket (control
packets vs stale video), which determines whether the gap is upstream routing
or whether large H264 packets are missing entirely.
Enhance watch-socket diagnostic to report distinct RTP payload types seen and
the max packet length, so we can distinguish 'only small control packets arrive
(no real H264)' from 'H264 arrives but decrypt fails'. Live already confirmed
dave=true ready=true with packets flowing but no burst — need to know if they're
tiny 52-byte control packets (PT 73) or large H264.
Instrument handleUdpMessage to log (rate-limited, first 3 then /20s) whether
video RTP packets actually reach the watch socket, and whether the DAVE session
is attached+ready and encryption key material present. Needed to diagnose why
no .h264 is written despite DAVE Ready + MLS: is the packet not arriving, or is
decrypt returning null?
CRITICAL: djs/voice stores the DAVESession wrapper at net.state.dave
(createDaveSession assigns to state.dave on op4 SessionDescription), NOT
inside connectionData. decryptVideoPacket looked up connectionData.dave which
is ALWAYS undefined -> every video packet hit '!dave?.session' guard and was
silently dropped, so no .h264/.mp4 ever got written despite the handshake
reaching Ready.
Fix: pass net.state.dave as a separate arg (the wrapper has .session ->
Davey.DAVESession) so the DAVE-layer decrypt (MediaType.VIDEO) actually runs.
Typecheck + build pass, lint clean (src/), 179/179 tests.
Live deploy 15:51 reached DAVE watch READY + completed MLS handshake on the
stream RTC (0->1->2->3->4, MLS commit processed, heartbeats alive). scan-on-
join also confirmed: 'Scanned pre-existing streamers on join watched=1'.
P4 remaining = capture actual video RTP (burst->mp4) while a stream is live.
If someone is ALREADY sharing screen / camera on when the bot joins the
channel, no voiceStateUpdate with streaming:true fires for them, so the
bot never sent STREAM_WATCH and missed their video entirely. trackChannel
now scans channel.members and starts a watch for anyone already streaming
(ignoring the bot itself and non-streamers). Idempotent: startStreamWatch
no-ops if a watch already exists. +2 tests (9/9 in videoRecorder).
djs/voice Networking emits stateChange(oldState, newState), but the watch
handler declared (newState, oldState) -- reversed. So the code-4 Ready
branch (which attaches udp.on('message') + logs 'DAVE watch READY') never
fired when entering Ready; it only fired spuriously when LEAVING Ready.
Result: full DAVE handshake completed on the watch RTC (Ready + DAVE MLA +
video stream 21029 active 1920x1080@60) but no UDP listener -> no video
captured. Swap to (oldState, newState) so enter-Ready wires the socket.
Add watch-state N->M log on every djs/voice Networking stateChange (with
hasUdp flag) so the stream-watch connection's exact progression is visible:
OpeningWs(0)->Identifying(1)->UdpHandshaking(2)->SelectingProtocol(3)->
Ready(4). Pins down where the DAVE flow stalls instead of guessing from the
absence of logs. Pairs with debug:true + watch-djs-debug.
Add debug:true to watch Networking options and wire net.on('debug') to
logger.info so djs/voice internal WS/DAVE state transitions appear in
journal. Without this, the stream-watch connection went silent after
'Streaming DAVE Networking' — no ready/error/close visible. Needed to
diagnose why the WS to stream endpoint 'c-sin14-xxx:2083' produced no
events.
Live log (14:52) showed the stream-watch flow reaching STREAM_CREATE +
STREAM_SERVER_UPDATE but then DAVE processProposals threw
'ValidationError(WrongGroupId)' -- the Davey MLS session derived the wrong
group because connectionOptions.channelId was the guild voice channel id.
Per Discord-RE StreamConnection.daveChannelId = BigInt(serverId)-1n, the
stream-watch DAVE MLS group is keyed to rtc_server_id-1, not the vc channel.
Fix: pass BigInt(serverId)-1n as channelId to the watch Networking.
This error also surface as an uncaughtException that crashed the gateway
(systemd restarted it). Correct channel id prevents it at the root.
The previous code read the sessionId from the selfbot client's voice manager
(client.voice.connection), which is no longer established since ensureSelfbotVoice
was removed — it would have sent sessionId:'none' in the watch Networking
identify and been rejected. Read the active session from the guild
@discordjs/voice connection (getVoiceConnection(guildId).state.networking...
connectionOptions.sessionId) instead.
The first streamWatchReceiver only did dave.session.decrypt(msg.subarray(12))
which skipped the outer legacy-AES layer Discord wraps around the DAVE payload
on every RTC packet (encrypt = dave.encrypt then aead_aes256_gcm with RTP
header as AAD). Port @discordjs/voice VoiceReceiver.decrypt/parsePacket
faithfully (header strip incl CSRC+extension+padding, AES-GCM auth tag, then
DAVE MediaType.VIDEO). Without this the .h264 would be garbage.
Replace the dead selfbot-v13 video path (WS 4017 DAVE). videoRecorder
now delegates to a new streamWatchReceiver that:
- sends STREAM_WATCH (op 20) on voiceState.streaming
- opens a @discordjs/voice Networking to the watch RTC (STREAM_CREATE +
STREAM_SERVER_UPDATE) with DAVE enabled
- decrypts H264 via Davey MediaType.VIDEO, depacketizes + muxes to mp4
- tears down on streaming-stop / leave / untrack
Remove ensureSelfbotVoice/createVideoStream/joinStreamConnection (dead).
recorder.ts no longer fires the futile eager selfbot join.
Supersedes the eager-selfbot connection plan: Discord now REQUIRES DAVE (E2EE,
WS 4017) on all voice RTC, and discord.js-selfbot-v13's voice stack predates
DAVE, so its video-receive path (joinChannel + joinStreamConnection +
receiver.createVideoStream) cannot authenticate. @discordjs/voice 0.19.2 exports
VoiceWebSocket/VoiceUDPSocket/DAVESession/Networking + @snazzah/davey supports
MediaType.VIDEO/Codec.H264 decrypt, so we can build a DAVE-capable stream-watch
connection. Phased plan: prototype (P2), gateway integration (P3), live verify (P4).
Video capture (camera + screen share) recorded ZERO frames because the selfbot
ClientVoiceManager.connection was created LAZILY — only when a user started
streaming. At that point the bot is already connected via @discordjs/voice, so
the selfbot re-join never gets a fresh VOICE_SERVER_UPDATE and times out with
VOICE_CONNECTION_TIMEOUT after 15s. joinStreamConnection (STREAM_WATCH) +
receiver.createVideoStream both need that selfbot VoiceConnection CONNECTED.
Fix: establish the selfbot VoiceConnection eagerly in recorder.startRecording,
BEFORE joinVoiceChannel, so it rides the bot's fresh join (Discord emits
VOICE_SERVER_UPDATE → selfbot authenticates). videoRecorder reuses the cached
connection per guild, tears it down on voice stop/destroy. Best-effort — never
blocks audio recording.
Verified: typecheck + build + biome (src/) green; 9/9 videoRecorder tests.
Root cause of why Phase A/B captured zero video: @discordjs/voice is audio-only
and never sends the gateway STREAM_WATCH signal, so Discord never forwards a
member's video RTP to the bot. Live diagnostic confirmed: while members were
sharing, audio .ogg files flowed for many users but no non-opus RTP ever
arrived.
Correct path: discord.js-selfbot-v13 ships a complete native watch/record stack.
New src/modules/voice-recording/videoRecorder.ts:
- Detects streamers via voiceState.streaming on a single idempotent
voiceStateUpdate listener.
- client.voice.joinChannel() (reuses the single session alongside
@discordjs/voice) + joinStreamConnection(userId) -> STREAM_WATCH (op 20).
- receiver.createVideoStream(userId, path) -> PacketHandler -> Recorder
(ffmpeg over UDP loopback) -> Matroska .mkv, decryption handled internally.
Wired in recorder.ts (trackChannel/untrackChannel) + bootstrap.ts
(setVideoRecorderClient/setVideoRecordingsDir). All best-effort; failures log
and never break existing voice/audio. Unit tests 6/6 (single listener, watch
handshake + mkv path, skip own video, idempotence, teardown). Full suite
20 files / 176 tests green; typecheck + build + biome clean.
UNVERIFIED live yet: needs deploy + a streamer to confirm Recorder ready +
playable .mkv.
Temporary diagnostic to answer definitively whether Discord actually delivers
video RTP to the bot when someone screen-shares / turns on camera. Both
videoReceiver and screenShareAudio rely on ssrcMap emitting videoSSRC from
voice-state updates, and NO "Video SSRC appeared"/"Screen-share video started"
lines appear even after a real share+record. This logs any RTP packet whose
payload type is not Opus (120) so we can tell: (a) video RTP IS arriving but
attribution/signaling fails, vs (b) Discord sends no video at all to a
non-signaling receiver. Remove this log once the gap is understood.
ssrcMap emits "delete" when a VoiceUserData is removed (user stops sharing /
leaves voice). Hook it to close the user's open video burst ~500ms later so the
ffmpeg mux starts as soon as they stop, instead of waiting up to 5s for the idle
sweep. No-op if no burst exists; safe on normal teardown.
Phase A captured raw .h264 streams but left them as non-playable elementary
streams. Phase B adds automatic muxing: when a video burst closes, the raw
.h264 is remuxed to a self-contained MP4 via `ffmpeg -c copy` (no re-encode,
fast) with `+faststart`, waits for the write stream to fully flush first so the
mux never reads a truncated tail, and deletes the raw .h264 on success (keeping
it on failure). Output: <RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.mp4.
muxToMp4 is exported + covered by a real-ffmpeg vitest (tests/videoReceiver.test.ts):
generates a tiny baseline h264, remuxes, asserts mp4 exists/non-empty & raw deleted
(also the 5 depacketizer tests). Full gateway suite 170/170 green, tsc + biome clean.
@discordjs/voice only decrypts/forwards AUDIO (opus) — its onUdpMessage drops
every non-opus RTP packet (dist/index.mjs:2068 `!== RTP_OPUS_PAYLOAD_TYPE`, 120).
Video RTP (H264 camera + screen share, plus VP8/VP9/AV1) arrives on the same UDP
socket but was silently discarded.
New videoReceiver.ts wraps receiver.onUdpMessage (like screenShareAudio.ts):
- detects video payload types (96/98/101/102/106/116/126/127),
- decrypts them with the connection secret key/encryptionMode via the SAME
receiver.parsePacket path @discordjs/voice uses for audio (so DAVE + voice
encryption are handled identically),
- depacketizes H264 to AnnexB (single NAL, STAP-A, and FU-A fragmentation),
waiting for a keyframe (SPS/PPS/IDR) before writing,
- writes a raw .h264 elementary stream per user per burst under
<RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.h264.
Attribution: a videoSSRC→user index is built from ssrcMap updates; a proximity
fallback mirrors screenShareAudio's inferScreenShareOwner. Bot's own video is
skipped.
Unit tests: tests/videoReceiver.test.ts (AnnexB start code, keyframe gating,
FU-A reassembly, orphan-fragment tolerance) — 5/5 green.
Phase A only (capture raw h264). Phase B (ffmpeg decode+mux to MP4/WebM +
persist) and Phase C (frontend playback) are follow-ups.
A server-muted/server-deafened bot can't reliably receive/record members' audio
(and definitely can't receive video/screen share). After the voice connection
is Ready, force a REST guild-members PATCH (mute:false, deaf:false) on the self
member so the bot is auto-unmuted & undeafened on every join/reconnect.
Requires MUTE_MEMBERS + DEAFEN_MEMBERS permissions (user granted). Failure is
logged as a warning and never breaks the voice join.
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.
NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.
NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
Root-cause fixes for 'banyak miss & terpotong' in the voice->recording flow:
- subscribe BEFORE collecting user metadata. receiver.speaking 'start' fires
on the FIRST opus packet, and onUdpMessage forwards frames to the
subscription only when one exists — every frame during the old
await collectUserMetadata (a Discord REST roundtrip on cache miss) was
dropped, cutting off the start of every burst. Now subscribe synchronously
(guard first, no await in between), then fetch metadata in the background
and discard the burst if the speaker turns out to be a bot.
- one segment per burst: drop the fixed 5s RECORDING_SEGMENT_MS rotation on
the OGG path, which split continuous speech mid-word/sentence. Only the
web-PCM decoder still rotates (bounds memory).
- finalize only once the underlying file has flushed to disk (wait on the
write stream 'finish'), so upload/transcode reads a complete file.
- raise AfterSilence 3000->4000ms so natural pauses (thinking, interruptions)
don't split one utterance into several recordings.
- lower the 'too short to keep' threshold 1000->300ms so brief replies
("ya", "siap") are kept instead of dropped.
All typecheck / biome(src/) / vitest (164) green.
- nativeBuildInputs: remove cmake, rustc, cargo, git — GMW's only native deps (@discordjs/opus, sharp) are PREBUILT, no source compile needed (cmake/rust were inherited for node-datachannel which is 9router, not GMW). python3/gnumake/gcc stay as node-gyp fallback for opus.
- pruneProd: also strip .pnpm/@types+* (pure TS decls pulled into the prod graph as real deps by discord-api-types/pg-protocol, never required at runtime) and .pnpm/opusscript@* (pure-JS fallback Opus engine that prism-media only loads IF native @discordjs/opus fails — native is always present, so opusscript is never executed).
- Verified: nix flake check OK; typecheck + biome check src/ green; runtime smoke test post-prune loads @discordjs/voice, sharp, selfbot, tiktoken, piscina and encodes a frame via native opus.
- mediaDownloader: flip URL candidate order so discord_url is tried
before uploaded_url (uploaded_url is archive-only fallback)
- ai-analysis-worker: remove upload-pending race guard that blocked
analysis until Tele upload completed; analysis now runs immediately
on the Discord CDN URL
- batchProcessor: remove upload-pending defer/poll-backoff logic
- individualFallbackProcessor: remove upload_pending requeue loop
- batchOutcomeClassifier/fallbackResultClassifier: drop upload_pending
classification (no longer needed)
- tests: update batchOutcomeClassifier + fallbackResultClassifier tests
to reflect removed upload_pending signal
Switch GMW's AI LLM base URL from 9router (https://9router.asepharyana.my.id/v1)
to omniroute on imrnes (http://100.121.180.82:20128/api/v1).
- Update default AI_LLM_BASE_URL in discord-gateway + backend config schemas
- Update .env.example documentation
- Update all 9router references in comments/docs/tests to omniroute
- Production BWS secret gmw_ai_llm_base_url already updated
Omniroute uses /api/v1 prefix (not /v1 like 9router), so the base URL
now correctly points at the right API path for the OpenAI SDK.
Root cause of 'navbar mobile tak bisa pindah halaman': Next <Link>
client-side navigation is dead app-wide. A React hydration mismatch
(#418: 'server rendered text didn't match the client') is thrown by the
SSR-seeded live feeds — relative times (formatRelativeTime(e.edited_at) /
m.created_at) computed with Date.now() render slightly differently on
server vs client, which breaks the Next client router (router.push is a
no-op). The desktop NavRail worked only because it uses plain <a href>
(hard navigation bypasses the broken router).
Fixes:
- mobile-nav.tsx: use plain <a href> (NOT Next <Link>), identical to the
working sidebar NavRail, so mobile nav always navigates regardless of
router state ('ikuti cara kerja sidebar').
- Add suppressHydrationWarning to the SSR-seeded relative-time spans so
server/client drift no longer throws #418 (EditHistory, LiveModerationFeed,
messages/results + detail rows, recordings, TermGlossary,
ChannelCultureGlossary, CategoryDrilldown).
Verified on non-prod :4024 @375px: Voice/Media/Search all navigate, no #418
in console. Plan: .hermes/plans/mobile-nav-hydration-fix.md