Video (camera + screen share) DAVE stream-watch now produces per-burst
MP4 segments instead of one long .h264 per watch:
- Detects VIDEO_SILENCE_MS (4000ms) of no H264 packets → closes the
current segment, muxes to MP4, registers in voice_recordings + uploads
to TeleUploader, then reopens for the next burst (mirrors voice AfterSilence).
- Per-watch segment counter + per-segment depacketizer reset + closing
guard + write-error swallow so races (silence close vs in-flight UDP
packet) never corrupt files or crash the gateway.
- Frontend: recordings deck renders a native <video> player for MP4 rows
(detected by filename), keeps single-playback registry across audio+video.
Gateway archive embedder now parses metadata.channel.{channelName,threadName}
from each message and stores channel_name/thread_name in the Qdrant payload.
Backend exposes them; the semantic results card renders the thread name (or
channel name) instead of a raw #snowflake, with the ID as a last-resort
fallback for legacy points. Matches the message feed's channel-label logic
(getMessageChannelLabel).
Archive payload now stores username, channel_id, guild_id, thread_id and the
real message created_at (not embed time). Backend searchArray accepts an
optional guildId and applies a Qdrant payload filter so results can be scoped
to the guild being viewed. API/frontend expose the new fields and the
semantic results card shows who said it, in which channel, and when —
turning bare text blobs into contextual results. Old points fall back to
analyzed_at and omit the new fields gracefully.
- AI_VOICE_TRANSCRIPTION_MODEL config (default whisper-1) so the model can be a provider-qualified id (openrouter/openai/whisper-1) that actually has credentials through 9router/omniroute — bare whisper-1 maps to the openai provider which has none
- response_format json (not text): 9router proxies only json/verbose_json transcription responses; text returns 400
- parse text from the json response object
- prod env updated: model=openrouter/openai/whisper-1 (still needs OpenRouter STT balance — 402 until funded)
- Normalize text before embedding (strip mentions/URLs/emoji/markdown/control chars, lowercase, truncate) on both write and query sides so vectors aren't diluted and tokens aren't wasted
- embeddingClient: retry embeddings (maxRetries 2), validate batch dimension consistency, preserve index alignment for empty-normalized texts
- archiveEmbedder: store normalized text in archive payload, skip empty-normalized content
- backend: normalize search queries, make archive search similarity threshold configurable (AI_LLM_EMBEDDING_ARCHIVE_MIN_SIMILARITY, default 0.6)
The bot's own VOICE_STATE_UPDATE showed server-level deaf:true — a
server-deafened member is NOT sent the streamer's audiovisual RTP by Discord,
which is the likely reason no H264 arrives despite the DAVE watch reaching
Ready. The previous fire-and-forget forceSelfServerUnmuteUndeafen() ran once
after the first Ready join and silently reverted on reconnect/restart.
- Export forceSelfServerUnmuteUndeafen from recorder.ts; re-assert it (with
read-back verification logging stillDeaf) at the START of every
startStreamWatch() before STREAM_WATCH is sent (dynamic import avoids the
recorder <-> videoRecorder <-> streamWatchReceiver module cycle).
- Re-assert it again after a successful voice reconnect.
- startStreamWatch() is now async; callers use void.
maxLen stayed 72 across 243 packets (no real H264, which is hundreds+ bytes) —
only 44-72-byte RTP packets on PT 76/72/73 arrive. Add ssrc + first-32-bytes
hex so we can identify exactly what Discord sends to the watch socket (control
packets vs stale video), which determines whether the gap is upstream routing
or whether large H264 packets are missing entirely.
Enhance watch-socket diagnostic to report distinct RTP payload types seen and
the max packet length, so we can distinguish 'only small control packets arrive
(no real H264)' from 'H264 arrives but decrypt fails'. Live already confirmed
dave=true ready=true with packets flowing but no burst — need to know if they're
tiny 52-byte control packets (PT 73) or large H264.
Instrument handleUdpMessage to log (rate-limited, first 3 then /20s) whether
video RTP packets actually reach the watch socket, and whether the DAVE session
is attached+ready and encryption key material present. Needed to diagnose why
no .h264 is written despite DAVE Ready + MLS: is the packet not arriving, or is
decrypt returning null?
CRITICAL: djs/voice stores the DAVESession wrapper at net.state.dave
(createDaveSession assigns to state.dave on op4 SessionDescription), NOT
inside connectionData. decryptVideoPacket looked up connectionData.dave which
is ALWAYS undefined -> every video packet hit '!dave?.session' guard and was
silently dropped, so no .h264/.mp4 ever got written despite the handshake
reaching Ready.
Fix: pass net.state.dave as a separate arg (the wrapper has .session ->
Davey.DAVESession) so the DAVE-layer decrypt (MediaType.VIDEO) actually runs.
Typecheck + build pass, lint clean (src/), 179/179 tests.
Live deploy 15:51 reached DAVE watch READY + completed MLS handshake on the
stream RTC (0->1->2->3->4, MLS commit processed, heartbeats alive). scan-on-
join also confirmed: 'Scanned pre-existing streamers on join watched=1'.
P4 remaining = capture actual video RTP (burst->mp4) while a stream is live.
If someone is ALREADY sharing screen / camera on when the bot joins the
channel, no voiceStateUpdate with streaming:true fires for them, so the
bot never sent STREAM_WATCH and missed their video entirely. trackChannel
now scans channel.members and starts a watch for anyone already streaming
(ignoring the bot itself and non-streamers). Idempotent: startStreamWatch
no-ops if a watch already exists. +2 tests (9/9 in videoRecorder).
djs/voice Networking emits stateChange(oldState, newState), but the watch
handler declared (newState, oldState) -- reversed. So the code-4 Ready
branch (which attaches udp.on('message') + logs 'DAVE watch READY') never
fired when entering Ready; it only fired spuriously when LEAVING Ready.
Result: full DAVE handshake completed on the watch RTC (Ready + DAVE MLA +
video stream 21029 active 1920x1080@60) but no UDP listener -> no video
captured. Swap to (oldState, newState) so enter-Ready wires the socket.
Add watch-state N->M log on every djs/voice Networking stateChange (with
hasUdp flag) so the stream-watch connection's exact progression is visible:
OpeningWs(0)->Identifying(1)->UdpHandshaking(2)->SelectingProtocol(3)->
Ready(4). Pins down where the DAVE flow stalls instead of guessing from the
absence of logs. Pairs with debug:true + watch-djs-debug.
Add debug:true to watch Networking options and wire net.on('debug') to
logger.info so djs/voice internal WS/DAVE state transitions appear in
journal. Without this, the stream-watch connection went silent after
'Streaming DAVE Networking' — no ready/error/close visible. Needed to
diagnose why the WS to stream endpoint 'c-sin14-xxx:2083' produced no
events.
Live log (14:52) showed the stream-watch flow reaching STREAM_CREATE +
STREAM_SERVER_UPDATE but then DAVE processProposals threw
'ValidationError(WrongGroupId)' -- the Davey MLS session derived the wrong
group because connectionOptions.channelId was the guild voice channel id.
Per Discord-RE StreamConnection.daveChannelId = BigInt(serverId)-1n, the
stream-watch DAVE MLS group is keyed to rtc_server_id-1, not the vc channel.
Fix: pass BigInt(serverId)-1n as channelId to the watch Networking.
This error also surface as an uncaughtException that crashed the gateway
(systemd restarted it). Correct channel id prevents it at the root.
The previous code read the sessionId from the selfbot client's voice manager
(client.voice.connection), which is no longer established since ensureSelfbotVoice
was removed — it would have sent sessionId:'none' in the watch Networking
identify and been rejected. Read the active session from the guild
@discordjs/voice connection (getVoiceConnection(guildId).state.networking...
connectionOptions.sessionId) instead.
The first streamWatchReceiver only did dave.session.decrypt(msg.subarray(12))
which skipped the outer legacy-AES layer Discord wraps around the DAVE payload
on every RTC packet (encrypt = dave.encrypt then aead_aes256_gcm with RTP
header as AAD). Port @discordjs/voice VoiceReceiver.decrypt/parsePacket
faithfully (header strip incl CSRC+extension+padding, AES-GCM auth tag, then
DAVE MediaType.VIDEO). Without this the .h264 would be garbage.
Replace the dead selfbot-v13 video path (WS 4017 DAVE). videoRecorder
now delegates to a new streamWatchReceiver that:
- sends STREAM_WATCH (op 20) on voiceState.streaming
- opens a @discordjs/voice Networking to the watch RTC (STREAM_CREATE +
STREAM_SERVER_UPDATE) with DAVE enabled
- decrypts H264 via Davey MediaType.VIDEO, depacketizes + muxes to mp4
- tears down on streaming-stop / leave / untrack
Remove ensureSelfbotVoice/createVideoStream/joinStreamConnection (dead).
recorder.ts no longer fires the futile eager selfbot join.
Supersedes the eager-selfbot connection plan: Discord now REQUIRES DAVE (E2EE,
WS 4017) on all voice RTC, and discord.js-selfbot-v13's voice stack predates
DAVE, so its video-receive path (joinChannel + joinStreamConnection +
receiver.createVideoStream) cannot authenticate. @discordjs/voice 0.19.2 exports
VoiceWebSocket/VoiceUDPSocket/DAVESession/Networking + @snazzah/davey supports
MediaType.VIDEO/Codec.H264 decrypt, so we can build a DAVE-capable stream-watch
connection. Phased plan: prototype (P2), gateway integration (P3), live verify (P4).