refactor: remove voice/recording/media features from frontend + prune lockfiles

- Delete pages: (dashboard)/{recordings,voice,media}/ incl. view.tsx
- Delete components/{voice,media}, hooks/{use-voice,use-recordings,use-media},
  lib/audio/ (mic-transmit, pcm-player, wav), api/{voice,recordings,media},
  types/{voice,recording,media}, lib/hash.ts
- Cut nav tiles (Active Voice Stages, Voice Recording Archive) from dashboard,
  MiniPlayer from ambient-app, hooks/types barrel exports, WS voice/media events
- Re-home guilds + textChannels to messages oRPC router (DB-derived) so the
  messages page picker keeps working; guild-picker simplified to text-only
- Clean Channel/AppConfig types of voice remnants
- Delete infra/docker/recordings/ + 11 voice/recording/video spec docs
- pnpm install: prune direct voice deps from gateway + backend lockfiles
  (prism-media/opusscript remain only as transitive discord.js deps)
This commit is contained in:
asepharyana
2026-09-23 19:45:14 +07:00
parent 33013697e0
commit ce784f8305
55 changed files with 151 additions and 5227 deletions
@@ -1,67 +0,0 @@
# Spec: Perbagus fitur Voice + Audio Playback (GMW frontend)
Tanggal: 2026-08-22 · Scope: **frontend only** (backend/gateway API sudah cukup)
## Masalah (audit)
1. Recordings: semua kartu pakai `<audio controls>` native — tampilan identik,
tidak ada indikasi which-clip-playing / loading / paused, dan N audio bisa
play bareng (overlap).
2. Media view: `thumbnailUrl` dari gateway tidak dipakai; tidak ada visual
"sedang playing" selain disc spin; queue item semua sama tanpa badge up-next.
3. Mini-player (`lib/hooks/use-media-player.tsx`) ada tapi TIDAK PERNAH
dimount → dead code, user tidak lihat status musik di halaman lain.
4. Voice page: `useMicTransmit.setVolume` + `useVoiceListen.setVolume`
tersedia tapi tak ada UI-nya; mic live tidak punya level feedback.
## Desain
### A. RecordingAudioPlayer (baru, `components/voice/recording-audio-player.tsx`)
Custom player menggantikan `<audio controls>`:
- Play/pause button (ikon berubah), spinner saat buffering (`waiting` event).
- Progress bar seekable (click-to-seek) + time label `m:ss / m:ss`.
- Waveform-ish equalizer bars saat playing (CSS animation, reduced-motion safe).
- **Single-playback**: module-level registry `activePlayers` — memainkan satu
clip otomatis pause yang lain.
- Kartu pemilik player aktif dapat highlight border signal + "Now playing" chip.
### B. Recordings view — pasang player baru
- Ganti `<audio>` → `<RecordingAudioPlayer src download_url>`.
- Highlight kartu via state lifted: `playingId` di view, callback `onPlay`.
### C. Media view polish
- Hero: thumbnail (jika `current.thumbnailUrl`) sebagai disc center image;
fallback ListMusic icon. Equalizer bars animasi CSS saat `playing`.
- Queue row pertama: badge "up next"; baris current track diberi ring signal.
- Volume read-only tetap.
### D. MiniPlayer global
- Hapus `lib/hooks/use-media-player.tsx` (dead) — ganti dengan komponen
`components/media/mini-player.tsx` yang subscribe `useMediaState` +
`useMediaWsSync` langsung (SWR cache shared antar route), mounted di
`AppFrame` bawah layar (fixed bottom, hidden di route `/media`).
- Menampilkan: thumbnail kecil/judul, tombol skip/stop, link ke /media.
### E. Voice UI
- Mic live: level meter (Equalizer bars) — mic-transmitter sudah punya worklet;
tambah `getLevel()` via AnalyserNode pada stream (simple RMS) di hook.
- Listen: volume slider (input range) wired ke `listen.setVolume`.
- Mic volume slider wired ke `mic.setVolume`.
## File touched
| File | Aksi |
|---|---|
| services/frontend/src/components/voice/recording-audio-player.tsx | new |
| services/frontend/src/app/(dashboard)/recordings/view.tsx | edit |
| services/frontend/src/app/(dashboard)/media/view.tsx | edit |
| services/frontend/src/components/media/mini-player.tsx | new |
| services/frontend/src/components/shell/ambient-app.tsx | mount MiniPlayer |
| services/frontend/src/lib/hooks/use-media-player.tsx | delete |
| services/frontend/src/hooks/use-voice.ts | tambah micLevel |
| services/frontend/src/lib/audio/mic-transmit.ts | expose analyser level |
| services/frontend/src/app/(dashboard)/voice/view.tsx | sliders + meter |
## Verifikasi
1. `pnpm lint` (biome) + `pnpm build` clean.
2. Smoke di port **4024** (BUKAN 4017) → curl 200 semua route.
3. Commit (tanpa trailer) → push → `gh run watch` → live check
https://imphnen.asepharyana.my.id/{media,recordings,voice}/ = 200.
@@ -1,82 +0,0 @@
# Spec: Recordings — filter per user + export WAV (Audacity)
## Konteks / Gejala
Halaman `services/frontend/src/app/(dashboard)/recordings` menampilkan semua
rekaman voice (deck). User ingin:
1. **Filter per orang** (tampil rekaman satu user saja).
2. **Export ke format untuk Audacity** (buka & edit rekaman di Audacity).
## Fakta saat ini (verified)
- Backend `recordings.list` (services/backend/src/orpc/router.ts:291) SUDAH
menerima `userId`/`channelId` filter → `RecordingsService.getRecent`.
- Frontend `recordingsApi.list(limit, channelId, userId, cursor)` (lib/api/recordings.ts)
sudah meneruskan `userId`. `useLoadMoreRecordings` juga sudah bawa userId.
- Tapi UI `RecordingsView` (app/(dashboard)/recordings/view.tsx) TIDAK punya
filter UI, dan `useRecordingsPage` dipanggil tanpa userId → semua tampil.
- Setiap rekaman punya `download_url` (MP3 di TeleUploader), `user_id`, `username`.
- Audacity membuka MP3/OGG tapi editing paling bersih dari WAV (uncompressed)
/ FLAC (lossless). Backend TIDAK punya ffmpeg & Nix flake backend tak include
ffmpeg → transcode server-side bukan pilihan. Browser punya codec MP3 → export
WAV via Web Audio API (client-side) adalah solusi self-contained terbaik.
## Keputusan desain
1. **Filter per user**: UI dropdown (Semua User + per user) di header halaman.
Memilih user → re-fetch `recordingsApi.list(50, undefined, userId)` (server
filter, benar untuk dataset besar + pagination). Dropdown dibangun dari
distinct `user_id`/`username` pada items yang sedang tampil.
2. **Export WAV (Audacity)**: client-side via Web Audio API.
- Per kartu: tombol "WAV" → decode `download_url` → WAV 16-bit PCM → download.
- Header: tombol "EXPORT WAV (N)" → gabung (concat) semua rekaman yang
sedang tampil (ter-filter) jadi 1 file WAV → download. Ideal untuk analisis
/ mixdown per orang.
- Implementasi di `lib/audio/wav.ts` (decode + encode + concat), tanpa dep baru.
## Perubahan
### Frontend
- **`src/lib/audio/wav.ts`** (baru):
- `decodeAudio(url: string): Promise<AudioBuffer>` — fetch arrayBuffer →
`new AudioContext().decodeAudioData`.
- `audioBufferToWav(buf: AudioBuffer, sampleRate=48000): Blob` — PCM 16-bit
interleaved, mono→stereo handling, RIFF/WAVE writer. Audacity-importable.
- `concatBuffers(buffers: AudioBuffer[]): AudioBuffer` — gabung di channel 0
(mono) dengan sample-rate max; untuk export gabungan.
- `downloadWav(blob: Blob, filename: string): void` — obj URL + <a download>.
- **`src/app/(dashboard)/recordings/view.tsx`**:
- Toolbar filter: dropdown user (built from distinct items) + tombol reset.
- State `filterUserId`; saat berubah → `recordingsApi.list(50, undefined, id)`
→ set ke SWR (key includes filter), reset pagination.
- Tombol "WAV" per kartu (disabled jika `!r.download_url`).
- Tombol "EXPORT WAV (N)" di header (disabled jika 0 item punya download_url);
concat semua items ter-filter yang punya download_url.
- Status loading saat export (spinner/disable).
- **`src/hooks/use-recordings.ts`**: `useRecordingsPage` menerima `userId?` dan
memasukkan ke key + call, supaya filter re-fetch bersih (per-user cache key).
`useRecordings`/`useLoadMoreRecordings` propagate `userId`.
- **`src/lib/types/recording.ts`**: tidak berubah (userId dari items).
### Backend / gateway
- Tidak ada perubahan. Filter & export sepenuhnya frontend.
## File yang disentuh (frontend only)
- `src/lib/audio/wav.ts` (baru)
- `src/app/(dashboard)/recordings/view.tsx`
- `src/hooks/use-recordings.ts`
## Verification
1. `cd services/frontend && pnpm typecheck` (tsc --noEmit) — 0 error.
2. `pnpm lint` (biome check src/) — exit 0.
3. `pnpm build` (next build) — hijau.
4. Manual (user): buka /recordings; pilih user di dropdown → hanya rekaman user
itu; klik WAV di kartu → file .wav ter-download & terbuka di Audacity; klik
EXPORT WAV (filtered) → satu .wav gabungan.
5. Push → CI `Build & Deploy (Nix)` (frontend job) hijau → deploy landing.
## Risiko / Trade-off
- Web Audio decode MP3 di client: butuh CORS pada download_url (TeleUploader
asepharyana.my.id — sudah same-serve/proxied, CORS ikut origin). Jika 403/CORS
gagal, error toaster + fallback manual (RAW MP3 tetap ada).
- concat gabungan = mono 48k; Audacity bisa edit per-channel nanti. Acceptable.
- Filter client (dropdown dari items yang dimuat) hanya menawarkan user yang
sudah tampil; dataset besar bisa pakai search nanti. Server filter benar untuk
yang dipilih.
@@ -1,81 +0,0 @@
# Spec: Recordings v2 — transcription, search, filters, leaderboard, sessions
## Konteks
4 fitur lanjutan untuk halaman /recordings (dipilih user):
1. Tampilkan transkripsi + search by kata kunci
2. Filter lanjutan: by channel + rentang tanggal
3. Kelompokkan klip jadi "sesi rapat" + autoplay berurutan + export satu sesi
4. Leaderboard bicara per user + ringkasan
## Fakta terverifikasi (2026-08-30)
- `voice_recordings` kolom: id, user_id, username, avatar_url, guild_id,
channel_id, channel_name, filename, size_bytes, download_url, upload_status,
upload_error, created_at, uploaded_at, **transcription** (schema
`shared/database/schema.ts:246`). Index user_id/channel_id/created_at.
- **Transcription 0/10.640** terisi prod: `AI_VOICE_TRANSCRIPTION_ENABLED`
default **false** (config/index.ts:333) & tidak diset di BWS env →
`transcribeRecording` (voiceTranscriber.ts) langsung return null.
`AI_LLM_BASE_URL` + `AI_LLM_API_KEY` SUDAH dikonfig (Whisper via router GMW).
- Transcriber hardcode `language: "en"` (voiceTranscriber.ts:34) — salah utk
ucapan campur id/en. Utk auto-detect: hapus param `language` (Whisper
auto-detect jika tidak diberikan).
- Backend `RecordingsService.getRecent` (recordings.service.ts) TIDAK select
`transcription`; SUDAH dukung filter `channelId`+`userId`+`cursor`; belum
dukung date-range & keyword search. `RecordingRow` interface juga tak punya
`transcription`.
- Tidak ada kolom `session_id` / `duration_ms` → sesi grouping = heuristik
(channel sama + gap created_at), durasi leaderboard = estimasi dari
size_bytes (MP3 128kbps: durasi_s ≈ size_bytes*8/128000).
## Keputusan desain
1. **Aktifkan transkripsi (fondasi)**: transcriber auto-detect (hapus
`language:"en"`), set secret BWS `AI_VOICE_TRANSCRIPTION_ENABLED=true`
(dibaca runtime oleh bws-exec saat service start). Rekaman BARU dapat
transkripsi. Backfill rekaman lama TIDAK dilakukan (pilih user: fokus baru;
file OGG lama kemungkinan besar sudah tidak dipakai).
2. **Backend** — perluas `getRecent`:
- select `transcription` (+ interface RecordingRow + FE type)
- filter baru: `q` (ILIKE on transcription + username), `startDate`/`endDate`
(created_at range, bigint ms)
- endpoint baru `recordings.summary`: agregasi per user → {user_id,
username, avatar_url, clips, est_duration_s, words, last_at}. `words`
dihitung dari transcription (tokenisasi spasi). Return sorted by clips.
3. **Frontend**:
- Kartu: tampilkan transkripsi (collapse/expand line-clamp) + durasi estimasi.
- Toolbar: search box (q), Select channel, date range (start/end), speaker
(sudah ada), reset filter.
- Tab/segment "Tape Deck" vs "Leaderboard": leaderboard render summary per
user + klik → filter deck by user itu.
- Sesi grouping (deck view): klip di-group jadi sesi bila channel sama &
gap antar klip < SESSION_GAP_MS (default 120s). Header sesi (channel,
waktu mulai, jumlah klip, total durasi). Autoplay tombol "Play session" &
"Export session WAV" (concat klip sesi via lib/audio/wav.ts yg sudah ada).
## Perubahan file
### Gateway (Tahap 1)
- `voiceTranscriber.ts`: hapus baris `language: "en"`.
- (secret) set `AI_VOICE_TRANSCRIPTION_ENABLED=true` via bws.
### Backend (Tahap 2)
- `recordings.service.ts`: interface + select + getRecent tambah transcription;
tambah filter q/startDate/endDate; method getSummary() untuk leaderboard.
- `orpc/router.ts`: procedur `recordings.list` schema tambah fields; prosedur
baru `recordings.summary`.
### Frontend (Tahap 3-5)
- `lib/types/recording.ts`: tambah transcription, est_duration_s opsional,
Summary type.
- `lib/api/recordings.ts`: list tambah q/startDate/endDate; + summary().
- `hooks/use-recordings.ts`: propagate filter baru ke key+call; hook
useRecordingsSummary.
- `app/(dashboard)/recordings/view.tsx`: toolbar search+channel+date, kartu
transkripsi, tab leaderboard, grouping sesi + autoplay + export sesi.
## Verification
- tiap tahap: gateway `pnpm build` + `biome check src/`; backend `pnpm build`
+ `biome check src/ tests/`; FE `pnpm build` + `biome check src/`.
- CI Build & Deploy hijau tiap tahap; deploy landing dicek via
`systemctl show gmw-<svc>.service --property=ActiveEnterTimestamp`.
- Tahap 1c: setelah deploy + rekaman baru, cek
`SELECT COUNT(*) FROM voice_recordings WHERE transcription IS NOT NULL`.
@@ -1,71 +0,0 @@
# Spec: Record video (kamera/screenshare) orang lain — WebRTC receive (Request 2)
Date: 2026-08-30. Status: Phase A + B DONE (capture → playable MP4); Phase C (UI) open.
## Why this is hard (ground truth, verified from @discordjs/voice 0.19.2 source)
`VoiceReceiver.onUdpMessage` (dist/index.mjs:2059) drops EVERY non-opus RTP
packet at line 2068: `if ((msg[1] & 127) !== RTP_OPUS_PAYLOAD_TYPE) return;`.
So video (kamera H264? actually Discord uses VP8/H264; screenshare combines with
video SSRC) is decrypted-capable but never forwarded. `receiver.parsePacket`
(2033) DOES decrypt any payload type generically (audio + video) using
`connectionData.{encryptionMode, nonceBuffer, secretKey}` — the only audio gate
is the opus check inside onUdpMessage.
=> FIX: wrap `receiver.onUdpMessage` (like screenShareAudio.ts already does for
screen-share AUDIO SSRCs): for packets whose payload type is a VIDEO type
(payload 96 VP8, 101/102 H264, 106/116/126/127 AV1, VP9 98...), call
`receiver.parsePacket(...)` myself to decrypt, then depacketize + write frames.
Delegate opus (120) to the original handler. Delegate audio to original.
## Audio already works (screenShareAudio.ts). We add VIDEO.
## Science-of-the-changes below.
## Phase A — capture + decrypt + depacketize to AnnexB h264 (THIS change)
Files (new): `src/modules/voice-recording/videoReceiver.ts`
- Hook into `recorder.startRecording` alongside `hookScreenShareAudio`.
- Wrap `receiver.onUdpMessage`:
- read ssrc = msg.readUInt32BE(8); userData = receiver.ssrcMap.get(ssrc)
- if payload type is video AND we have a "watching" subscription for that user
(videoSSRC present), decrypt via receiver.parsePacket(...), then:
- H264 (101/102 + payload 120 not): strip RTP header, reassemble FU-A
fragments into AnnexB NALs (start-code prefixed), buffer until we have
a full access unit (keyframe SPS/PPS/IDR or slices), append to a per-
user-per-burst `.h264` file.
- else delegate to original onUdpMessage.
- Watch `receiver.ssrcMap` "create"/"update" for `videoSSRC !== undefined` →
signal a video burst started for that user (like screenShareAudio does).
- Per-user video files written to `config.RECORDINGS_DIR/<uid>/video-<ts>.h264`.
- Guard: skip bot's own video (client.user.id).
Dependencies: NO new npm deps for Phase A (only crypto already in
@discordjs/voice via parsePacket + Buffer). ffmpeg-headless (already in Nix
buildInputs) used in Phase B for decode+mux.
## Phase B — decode + mux to playable MP4/WebM (DONE, commit 999c054b)
- `closeBurst` waits for the WriteStream `finish` (full flush/fd close), then
`muxToMp4(rawPath)`: `ffmpeg -f h264 -i raw.h264 -c copy -movflags +faststart
out.mp4`, deletes raw on success (>=1B mp4), keeps it on failure.
- Output: `<RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.mp4`.
- ffmpeg is on the gateway runtime PATH (pkgs.ffmpeg-headless, already in the
Nix buildInputs for the music/GoLive players).
- Test: `tests/videoReceiver.test.ts` muxToMp4 case (real ffmpeg, generates a
tiny baseline h264, asserts mp4 non-empty + raw deleted; skipped if no ffmpeg).
## Phase C — frontend playback + session grouping (follow-up)
- Backend oRPC list video files; FE video player, group by call session like audio.
## Verification
- Phase A: join voice, have a member screen-share/camera, confirm `.h264` file
grows with NAL frames + keyframes; journal shows "video burst" logs.
- Run vitest unit: RTP header strip + FU-A reassembly gives correct bytes.
## Open questions / risks
- Discord codec for camera = H264(101/102); screenshare uses H264 (101/103?)
and can also be VP8/VP9. Handle H264 first (depacketize proven), VP8/VP9 in
Phase B via ffmpeg RTP input.
- Encryption: DAVE (dave_protocol_version) adds a session layer; parsePacket
already applies daveSession.decrypt for audio — we must call the SAME
parsePacket path so DAVE/encryption is handled identically.
- ssrc↔user mapping during a broadcast: videoSSRC is in ssrcMap after the
voice state; may need the STREAM_CREATE network events to key reliably.
@@ -1,66 +0,0 @@
# Spec: Voice auto-reconnect (persistent state + rejoin on drop)
## Goal
Setelah `VoiceController.connect()` berhasil, state "sedang merekam di <guild>/<channel>"
disimpan di Postgres. Kalau gateway restart/reboot, atau koneksi voice drop tidak
disengaja (dikeluarkan/moved/server restart), gateway otomatis join ulang ke channel
yang sama.
## Requirement mapping (user's ask)
- "autoreconnect ke channel yg sama jika server restart atau reboot" → reconnect on
startup (ready handler) + keep DB record across graceful shutdown.
- "state nya persistent di db" → `voice_auto_reconnect` table.
- "rejoin jika tidak sengaja dikeluarkan" → watchdog on `Disconnected`/`Destroyed`
(kick / moved / voice server restart) → full rejoin with backoff.
- Manual leave (`/voice disconnect`, dashboard disconnect) MUST NOT rejoin.
## Design decisions
1. **New table** `voice_auto_reconnect` (dedicated, not `ui_state`):
- `guild_id` text PK
- `channel_id` text NOT NULL
- `channel_name` text
- `connected_at` bigint epoch-ms
- `updated_at` bigint epoch-ms
DAO: `voiceAutoReconnectRepo.ts` — `upsert(record)`, `list()`, `delete(guildId)`.
2. **Write on connect**: `VoiceController.connect()` → after `startRecording` success →
`upsert`. IDEMPOTENT (upsert per guild).
3. **Clear on manual leave**: `handleVoiceDisconnect` (all) + `handleVoiceDisconnectGuild`
pass `clearPersisted: true`. Graceful shutdown `disconnect()` keeps the record.
4. **Rejoin on startup**: bootstrap `ready` → `await voiceController.autoReconnect()`
(list persisted → connect each, non-fatal on failure).
5. **Rejoin on unexpected drop**: monitor per-connection; on `Disconnected`/`Destroyed`
with `!intentional` → schedule rejoin `connect(guildId, persistedChannelId)` with
backoff (min 2s, max 30s, max 5 attempts). Track `rejoinAttempts`, reset on success.
6. **Intentional flag**: `disconnectGuild(guildId, { clearPersisted?, intentional? })`.
- shutdown `disconnect()` → `{ intentional: true, clearPersisted: false }`.
- manual `disconnect()` (dashboard) → `{ clearPersisted: true }`, sets intentional.
- manual `disconnectGuild` → `{ clearPersisted: true }`, sets intentional.
The monitor checks `intentional` before rejoin; `clearPersisted` only deletes the row.
## Files touched
- `services/discord-gateway/src/shared/database/schema.ts` — add `pgVoiceAutoReconnectTable`
+ types.
- `services/discord-gateway/src/shared/database/voiceAutoReconnectRepo.ts` (NEW) — DAO.
- `services/discord-gateway/src/modules/voice-recording/voiceController.ts` — upsert on
connect; monitor + rejoin; `autoReconnect()`; `disconnect/disconnectGuild` opts.
- `services/discord-gateway/src/modules/command-handler/voice.handler.ts` — manual
disconnect/disconnectGuild pass `clearPersisted: true`.
- `services/discord-gateway/src/app/bootstrap.ts` — call `voiceController.autoReconnect()`
in `ready`.
- `services/discord-gateway/drizzle/migrations/0020_add_voice_auto_reconnect.sql` +
`meta/_journal.json` entry (apply manually per gmw-ops).
## Edge cases
- Channel deleted / guild lost while persisted → `connect()` throws (channel not found)
→ log + delete persisted row (don't retry forever).
- Rejoin attempts exhausted → keep row (so next restart retries) + log.
- Multiple guilds: per-guild monitor, per-guild persisted row.
- Graceful shutdown order: shutdown sets intentional=true (so no rejoin during teardown)
but keeps row.
## Verification
- `pnpm typecheck && pnpm build && pnpm lint` in `services/discord-gateway`.
- Apply migration `0020` manually; verify table exists.
- CI `Build & Deploy (Nix)` green; gateway deploy lands.
- Manual: connect via dashboard → check `voice_auto_reconnect` row; simulate drop →
confirm rejoin; manual disconnect → row cleared.
@@ -1,166 +0,0 @@
# Spec: Perbaiki alur voice → recording (miss & terpotong)
## Konteks & Gejala
User melaporkan alur voice sampai recording **banyak miss** (audio tidak tercatat)
dan **terpotong** (satu alur bicara kebelah jadi beberapa segmen / audio putus di
tengah). Ini domain `services/discord-gateway/src/modules/voice-recording/`.
Pipeline per user yang mulai bicara (speaking "start"):
```
receiver.speaking "start" → speakingHandler(userId)
├─ await collectUserMetadata(...) ← roundtrip API, subscribe tertunda
├─ receiver.subscribe(userId, {end: AfterSilence, duration: 3000ms}) → audioStream
├─ attach data/end/error handlers → audioStream.pipe(PacketFilter) → oggPacketStream
├─ SegmentManager.open() → OggLogicalBitstream → file .ogg
├─ data: SegmentManager.rotateIfNeeded (rotasi 5s) + decoder.write (web PCM tho
└─ end: SegmentManager.close() → segmen finish → finalizeSegment upload + transkrip
```
## Root cause (dari pembacaan kode — justifikasi di bawah)
### A. MISS bagian awal bicara — subscribe tertunda (utama)
`speakingHandler.ts:53` melakukan `await collectUserMetadata(...)` SEBELUM
`receiver.subscribe`. `collectUserMetadata` (metadata.ts:30) pada cold path
(cache miss) melakukan `client.users.fetch` + `guild.members.fetch` roundtrip
Discord API (ratusan ms–detik). Selama await, seluruh opus awal bocor → awal
kalimat hilang. Cache menghilangkan ini untuk user yang pernah ter-record, tapi
user baru/evict (cache max 200) kena setiap kali.
### B. Double-subscribe race
Guard `receiver.subscriptions.has(userId)` di `speakingHandler.ts:63` diletakkan
SETELAH `await collectUserMetadata`. Dua event "start" cepat keduanya melewati
guard (belum subscribe) → dua subscription → audio terbelah/ganda per user.
### C. TERPOTONG di jeda bicara — AfterSilence 3000ms
`AUDIO_STREAM_SILENCE_DURATION_MS=3000`. Setelah 3s diam, stream auto-`end` →
`SegmentManager.close` → segmen baru saat bicara lagi. Jeda normal (berpikir,
interupsi) memecah 1 alur bicara jadi beberapa segmen/file. Ini source "terpotong".
Segmen pendek hasil jeda <1s juga DIBUANG di `finalizeSegment` (MIN_DURATION_MS=1000)
→ miss kata singkat ("ya", "siap").
### D. Rotasi segmen 5s di tengah bicara
`RECORDING_SEGMENT_MS=5000`: `rotateIfNeeded` menutup bitstream & membuka baru
setiap 5s walau bicara kontinu. Pipenya di-re-wire di dalam handler data →
window drop kecil + continuity file pecah (bukan masalah besar, tapi berkontribusi).
### E. Tidak ada sinkronisasi "stop" speaking & stream "end" flaky
Handler hanya listen "start"; mengandalkan `AfterSilence` untuk emit "end".
Bug @discordjs/voice yang dikenal: `AfterSilence` bisa TIDAK emit "end" saat
koneksi gagal/teardown → segmen menggantung & tidak pernah finalize/upload
(recording "hilang"). Tidak ada watchdog.
## Scope
Hanya `services/discord-gateway/src/modules/voice-recording/` (+ config index bila
perlu default baru). Tidak menyentuh playback (player.ts), transmitter (voice dari
browser → Discord, arah berlawanan), muxer (konsolidasi akhir), atau screen-share
video (hanya audio SSRC via `hookScreenShareAudio` sudah ada & dibiarkan).
## Perubahan
### 1. Speak-before-metadata: subscribe LEBIH DULU, metadata paralel
`recorder/speakingHandler.ts`:
- Pindahkan `receiver.subscribe` + pipeline setup ke ATAS, SEGERA di handler,
sebelum `collectUserMetadata`.
- Jalankan `collectUserMetadata` secara paralel non-blocking; gunakan metadata
cache untuk registrasi segmen saat finalize.
- Pertahankan guard skip bot/user (bot bisa dicek dari `client.users.cache` /
`client.user.id` tanpa await) SEBELUM subscribe — jangan tunggu fetch user.
Detail konkret:
```
async handler(userId):
if (userId === client.user?.id) return;
if (receiver.subscriptions.has(userId)) return; // guard kini di DEPAN, tanpa await
// (belum tahu bot? gunakan cache user; subscribe dulu biar nggak miss)
clone = subscribe(userId, {AfterSilence, duration}) // TANPA await metadata
setup pipeline (data/end/error, pipe, open segment)
collectUserMetadata(...).then(meta => {
if (meta.bot) { drain & close subscription (jangan simpan) }
else { registrasi ulang metadata utk segmen aktif }
})
```
Karena listener `start` dibuang untuk bot, harus tutup subscription bot tanpa
menyimpan segmen (buang hasil). Pakai `receiver.subscriptions.get(userId)?.destroy()`.
### 2. Selesaikan race double-subscribe (bagian dari #1)
Guard `receiver.subscriptions.has(userId)` diletakkan SINCRON di awal (sebelum
await). Karena `subscribe` sinkron dan `subscriptions` terisi sinkron saat
dipanggil, event "start" kedua yang tiba setelah subscribe akan melihat
subscription aktif → di-skip. Tidak ada await antara guard & subscribe.
### 3. Naikkan AfterSilence + tail-length → kurangi "terpotong"
`shared/config/index.ts`:
- `AUDIO_STREAM_SILENCE_DURATION_MS` default 3000 → 4000 (beri ruang jeda
alami; Discord packet 20ms, 4s masih wajar, tidak membengkak file).
Opsional via env override di production (tidak wajib komit env).
### 4. Segmen berorientasi "burst bicara" daripada rotasi jam
`recorder/segment.ts` + `recorder/speakingHandler.ts`:
- Hapus/lepas rotasi SEGMEN berbasis waktu (RECORDING_SEGMENT_MS). Alih-alih,
satu segmen = satu burst bicara (buka di "start", tutup di "end"/AfterSilence
end). Ini menghilangkan pemecahan di tengah kalimat.
- `RECORDING_SEGMENT_MS` tetap dipakai untuk rotasi decoder web-PCM (broadcast
live), di mana segmen besar bisa menunda frame — biarkan seperti ada.
JADI: `SegmentManager.rotateIfNeeded` TIDAK lagi dipanggil pada jalur OGG
recording; decoder rotate tetap dijalankan.
Catatan: dengan satu segmen per burst, ukuran file ~ durasi bicara. File panjang
dibutuhkan transkrip & transcode; tidak ada batas keras yang perlu di-override.
Watchdog di #5 membatasi durasi menggantung.
### 5. Watchdog end-of-burst & teardown recovery
`recorder/speakingHandler.ts`:
- Setelah subscribe, arm timer watchdog (mis. `config twin`/hitung) yang menutup
segmen jika `AfterSilence` tidak emit "end" dalam X detik setelah "stop"
speaking — atau, lebih sederhana & robust: dengarkan BOTH stream "end" DAN
timer dari `receiver.speaking` "stop" (hingga @discordjs/voice meng-klaim
AfterSilence). Bila "stop" fire, mulai countdown kecil (mis. 500ms) lalu
`segmentManager.close` + `decoder.destroy` + destroy subscription jika stream
belum "end".
- Ini menutup A: segmen menggantung → jadi pasti finalize & upload.
Implementasi: subscriptionStream (audioStream) + track milik per-user di
Map<userId, {audioStream, segmentManager, decoder, timer}>; handler "stop"
menjadwalkan finalize.
### 6. Naikkan/ambil MIN segmen duration lebih rendah
`recorder/segmentFinalizer.ts`: MIN_DURATION_MS 1000 → 300ms. Kata pendek
("ya", "siap") tetap tersimpan. GUI biarkan.
## File yang disentuh
- `src/modules/voice-recording/recorder/speakingHandler.ts` (utama: subscribe
first, guard depan, watchdog stop, hapus rotasi segmen dari jalur OGG)
- `src/modules/voice-recording/recorder/segment.ts` (opsional: API close/open,
pertahankan rotate untuk decoder tapi tak dipakai jalur OGG)
- `src/modules/voice-recording/recorder/streamSetup.ts` (kecil: backfill subscribe
supaya return subscription utk cleanup/destroy bot)
- `src/modules/voice-recording/recorder/segmentFinalizer.ts` (MIN_DURATION)
- `src/shared/config/index.ts` (default AfterSilence 4000)
## Yang TIDAK disentuh
- `transmitter.ts` (arah browser→Discord, bukan recording)
- `player.ts`, `mediaSource.ts`, `screenShareAudio.ts` (hook SSRC sudah benar)
- `muxer.ts` (konsolidasi akhir tetap jalan)
- Flake/deps/bundle
## Verification
1. `cd services/discord-gateway && pnpm typecheck` (tsc --noEmit) — 0 error.
2. `pnpm lint` (biome check src/) — exit 0.
3. `pnpm build` (tsc → dist/).
4. Unit test baru (vitest, kalau infra tes ada):
- subscribe terjadi tanpa await metadata (spy urutan panggilan)
- double-start skip lewat guard sinkron
- "stop" → watchdog finalize segmen walau stream tidak "end"
- MIN_DURATION 300ms menyimpan kata pendek
5. Smoke/CI: `nix flake check` (eval). Push → CI `Build & Deploy (Nix)` hijau →
deploy landing (`systemctl show gmw-discord-gateway.service --property=ActiveEnterTimestamp`).
6. Runtime manual (user): join voice, bicara dengan jeda >3s, pastikan 1 alur
kontinu = 1 segmen utuh (bukan 3), dan awal kata tidak hilang.
## Risiko / Trade-off
- Subscribe-before-metadata: burst bot akan dikumpulkan sesaat lalu dibuang
(cost kecil: buang segmen). Lebih baik miss bot daripada miss user.
- AfterSilence naik: file lebih panjang sedikit saat jeda; upload/transkrip
timeout (transcodeToMp3 30s) tetap aman.
- Satu segmen per burst: tidak ada rotasi paksa → durasi segmen = durasi bicara
(bisa menit). Transkrip & transcode tetap ok. Watchdog batasi menggantung.
@@ -1,177 +0,0 @@
# Spec: Receive Others' Screen-Share/Camera Video Under DAVE — Build a DAVE-capable Stream-Watch Connection
Status: **P1–P3 DONE + 4th CRITICAL FIX deployed (f1a7b0c2); DAVE Ready + MLS handshake CONFIRMED live; P4 = waiting on active streamer to confirm video-burst→mp4**
Date: 2026-08-31
Author: Hermes
Related: `docs/specs/2026-08-31_video-receive-eager-selfbot-connection-spec.md` (superseded by this)
`docs/specs/2026-08-31_video-receive-phaseC-spec.md` (Phase C build, selfbot path — dead)
## Problem / Ground truth (established from live logs 2026-08-31)
GMW must record OTHER members' screen-share + camera video in a voice channel it
records. Audio works (via `@discordjs/voice` 0.19.2 negotiating DAVE). Video does
not. Verified live: the selfbot path (`discord.js-selfbot-v13` eager `joinChannel`
→ `joinStreamConnection` → `receiver.createVideoStream`) authenticates but Discord
closes the connection with WS code **4017 "E2EE/DAVE protocol required"** (5x →
`VOICE_CONNECTION_ATTEMPTS_EXCEEDED`). Root cause: **Discord now REQUIRES DAVE
(E2EE) on every voice RTC, and `discord.js-selfbot-v13`'s voice stack predates
DAVE** (identify has no `max_dave_protocol_version`, no MLS handshake). The selfbot
path is dead, cannot be repaired. Full details: skill `gmw-ops` →
`references/video-receive-and-unmute.md` §4.
Facts:
- Watching a stream = a **SEPARATE RTC connection**, not the guild audio socket:
gateway `STREAM_WATCH` (op 20) → Discord replies `STREAM_CREATE` (rtc_server_id)
+ `STREAM_SERVER_UPDATE` (separate token+endpoint) → client opens its own voice
WS+UDP to that endpoint (`StreamConnectionReadonly` in selfbot). The watched
video never rides the @discordjs/voice guild socket.
- The stream-watch RTC ALSO requires DAVE (same 4017 mechanism).
- `@snazzah/davey` (bundled with @discordjs/voice 0.19.2) supports
`MediaType.VIDEO` + `Codec.H264` decrypt — DAVE machinery CAN decrypt H264 video.
- No off-the-shelf DAVE-capable video-RECEIVE path exists. Closing references:
- **Discord-RE/Discord-video-stream** (fork of `@dank074/discord-video-stream`,
master 2026-08-28): full DAVE in `src/client/voice/BaseMediaConnection.ts`
(Davey `DAVESession` init via `initDave`, MLS key-package / proposals /
commit / welcome / transitions; `WebRtcConnWrapper` encrypts audio/video via
`daveSession.encrypt(MediaType.VIDEO, codec, …)`). BUT it is STREAMING only
(send). No STREAM_WATCH / receive.
- `@discordjs/voice`: full DAVE receive for AUDIO only; `DAVESession.decrypt`
hardcodes `MediaType.AUDIO` (dist ~line 892); `onUdpMessage` drops non-opus;
no STREAM_WATCH.
- `discord.js-selfbot-v13`: video receive but no DAVE.
## Goal
Replace the dead selfbot receive path with a **DAVE-capable stream-watch voice
connection**: on detecting a member `voiceState.streaming`, send `STREAM_WATCH`,
connect a DAVE-authenticated RTC to the stream endpoint, decrypt incoming H264
RTP (`MediaType.VIDEO`), reassemble via the existing `H264Depacketizer`, mux to a
playable container. Reuse every tested building block already in the repo.
## Strategy decision (default A; B as fallback) — de-risk in Phase 2
Two implementation routes. Decide by Phase 2 prototype result.
### Strategy A — extend @discordjs/voice's tested native stack (PREFERRED, lighter)
Reuse @discordjs/voice 0.19.2 internals (already a runtime dep, already DAVE-tested
for audio):
- Drive a connection to the stream endpoint using djs/voice's `VoiceWebSocket` +
`VoiceUDPSocket` + `DAVESession` (the same classes that work for the guild
connection — they take arbitrary endpoint/token/session).
- Send the voice identify with `max_dave_protocol_version`, complete the DAVE
handshake (Davey), then on receipt of a video RTP packet call
`daveSession.decrypt(userId, MediaType.VIDEO, packet)` (Davey exposes
`MediaType.VIDEO` + `Codec.H264`) — djs/voice's hardcoded `AUDIO` is the only
blocker, fix by invoking Davey directly with `MediaType.VIDEO` for video SSRCs.
- STREAM_WATCH sent via the existing selfbot `client.ws.broadcast` (cheap, works —
it needs no selfbot voice connection).
- Feed decrypted H264 → `H264Depacketizer` → `.h264` → `muxToMp4` (both already
in `videoReceiver.ts`, unit-tested).
- No new runtime deps. Risk: relies on non-exported djs/voice internals (reachable
via `as any`, as the existing `parsePacket` usage shows).
### Strategy B — port Discord-RE's BaseMediaConnection (heavier, more self-contained)
Port `BaseMediaConnection.ts` DAVE handling + `WebRtcConnWrapper` into a
receive/watch connection in the gateway. Deps: requires `@lng2004/node-datachannel`
(new native WebRTC dep) + `@snazzah/davey` (already available). More code, more
risk (native dep in Nix store), but a clean-room receive path decoupled from
djs/voice internals. Use only if A proves infeasible.
## Files touched (Strategy A shape)
- `services/discord-gateway/src/modules/voice-recording/streamWatchReceiver.ts`
(NEW): DAVE stream-watch connection wrapper. Owns, per watched user:
- `sendStreamWatch(client, streamKey)` (via `client.ws.broadcast({op:20,
d:{stream_key}})`), stream_key = `guild:<gid>:<chid>:<uid>`.
- collects STREAM_CREATE (rtc_server_id) + STREAM_SERVER_UPDATE (token+
endpoint) via `client.on('raw')` match on stream_key.
- builds a djs/voice-style connection to `<endpoint>` with the received token/
session; completes DAVE handshake.
- `onUdpMessage` wrapper: for video payload types, decrypt with
`daveSession.decrypt(userId, MediaType.VIDEO, buf)`, depacketize, write.
- teardown on STREAM_DELETE / user stops streaming / leave / channel untrack.
- `recorder.ts`: wire `trackChannel`/`untrackChannel` already exist; ensure the
selfbot *eager voice connection* attempt is REMOVED (it only 4017-spams logs) —
but KEEP `client.ws.broadcast` availability for STREAM_WATCH.
- `videoRecorder.ts`: remove the dead selfbot `joinChannel`/`joinStreamConnection`
calls; keep the `voiceStateUpdate` streaming detection + teardown bookkeeping as
the entry point; delegate the actual receive to `streamWatchReceiver`.
- `videoReceiver.ts`: keep `H264Depacketizer` + `muxToMp4` (reused). The
guild-socket `hookVideoReceiver` can be removed or left inert.
- Tests: `tests/streamWatchReceiver.test.ts` (DAVE-handshake stub, RTP decrypt path
with a mocked Davey, STREAM_WATCH packet shape); keep `tests/videoReceiver.test.ts`.
## Phases (each independently verifiable)
1. **Phase 1 (this session): spec + source reconnaissance.** Confirm djs/voice
internals are reachable (VoiceWebSocket/VoiceUDPSocket/DAVESession exports &
shapes), confirm Davey `MediaType.VIDEO` decrypt signature, confirm how a raw
stream-watch connection's identify/select-protocol flows. Verify the selfbot
`streamKey` format + `raw` STREAM_CREATE/SERVER_UPDATE payload. GATE: accurate
spec + no unknowns blocking A.
2. **Phase 2: de-risk prototype.** Standalone script (not in the gateway) that:
logs into the same selfbot token, joins a real voice channel, sends STREAM_WATCH
for a live streamer, receives STREAM_CREATE/SERVER_UPDATE, and attempts a
DAVE-authenticated connect + receive of ≥1 H264 packet to prove the path before
any gateway integration. GATE: at least one decrypted H264 NAL captured in the
lab.
3. **Phase 3: gateway integration** per files-touched. GATE: typecheck + build +
biome + unit tests green; CI deploy ok.
4. **Phase 4: live verify.** With a real streamer in a recorded channel: journal
shows `STREAM_WATCH sent`, `DAVE ready`, `Video burst opened`, and a playable
`.h264`/`.mp4`/`.mkv` on disk. GATE: playable file with real video content.
## Risks / open questions
- Does Discord require the stream-watch connection to use the SAME session_id as
the bot's active voice session, or a fresh one? (Affects identify.) Resolve in P2.
- Which video codec does Discord actually send for camera vs GoLive (H264 likely,
but VP9/AV1 possible) — the depacketizer only handles H264. P2 measures the
payload type live; add depacketizers for other codecs only if observed.
- djs/voice `DAVESession`/`VoiceUDPSocket` reachability via `as any` must be
confirmed against the installed 0.19.2 build (P1).
- The separate stream RTC may need `selectProtocol`/SDP even for receive-only; the
Discord-RE SDP shows a `m=video ... inactive` section. Follow the same shape.
## Verification (overall)
- Per-phase gates above.
- No regression: audio recording + message capture still work after changes.
- `pnpm typecheck && pnpm build && pnpm lint` green in discord-gateway.
- Commit + push; CI `Build & Deploy (Nix)` green; live streamer produces a file.
## Phase 1 findings (CONFIRMED 2026-08-31, Strategy A feasible)
- `@discordjs/voice` 0.19.2 dist/index.mjs PUBLICLY exports exactly the primitives
needed: `DAVESession`, `Networking`, `NetworkingStatusCode`, `VoiceConnection`,
`VoiceReceiver`, `VoiceUDPSocket`, `VoiceWebSocket`, `SSRCMap`,
`RTP_OPUS_PAYLOAD_TYPE` (export block ~3143). So a stream-watch connection can be
built OUTSIDE the lib using these constructors — no `as any` needed for the heavy
lifting.
- `Networking` child wiring (~line 1364-1484): `new VoiceWebSocket('wss://' +
endpoint + '?v=8', debug)`; on WS open send Identify `{op, d:{server_id,
user_id, session_id, token, max_dave_protocol_version: getMaxProtocolVersion()}}`;
`createDaveSession(protocolVersion)` → `new DAVESession(protocolVersion, userId,
channelId, {decryptionFailureTolerance})` then `.reinit()`; UDP via
`new VoiceUDPSocket({ip, port})` after Ready gives modes + ssrc.
- `DAVESession` wraps `@snazzah/davey` `Davey.DAVESession(protocolVersion, userId,
channelId)`; on network packets it calls `this.session.decrypt(userId,
Davey.MediaType.AUDIO, packet)` — hardcoded AUDIO (line ~892). For video we call
Davey directly with `MediaType.VIDEO` + `Codec.H264`.
- `@snazzah/davey` MediaType enum: AUDIO=0, VIDEO=1; Codec H264=4; methods
`decrypt(mediaType, codec, packet): Buffer` + `encrypt(...)`. Confirmed in davey
index.d.ts.
- Stream key format (selfbot VoiceConnection.js ~1240): `guild:<gid>:<chid>:<uid>`
for guild channels; `STREAM_WATCH` = gateway op 20, `d:{stream_key}`;
`sendSignalScreenshare` = `client.ws.broadcast({op:20,d:{stream_key}})`. Replies
come as gateway `raw` events `STREAM_CREATE` (`d.rtc_server_id`) +
`STREAM_SERVER_UPDATE` (`d.token`, `d.endpoint`); selfbot routes them to the
stream connection via `client.on('raw')` matching `d.stream_key`, setting
`setSessionId(sessionId)` + `setTokenAndEndpoint(token, endpoint)` (Watch case in
`StreamConnectionReadonly`).
- Discord-RE reference for the identify SDP: stream connections send a `m=video`
section with `a=inactive` (receive-only-ish) + standard DAVE/VoiceOpCodes
(op 0 identify, op 2 select protocol incl. `max_dave_protocol_version`). See
`BaseMediaConnection.handleProtocolAck` + `initDave`.
- Decision: proceed with **Strategy A**. Selfbot code to REMOVE: the eager
`ensureSelfbotVoice` join + `joinStreamConnection`/`receiver.createVideoStream`
in `videoRecorder.ts` (proven dead — 4017). Keep the `voiceStateUpdate` streaming
detection + bookkeeping; swap the receive plumbing to a new `streamWatchReceiver`
driven by a djs/voice-style connection. STREAM_WATCH itself still sent via the
selfbot `client.ws.broadcast` (needs only the WS, not a selfbot voice conn).
- OPEN (resolve in Phase 2 lab): (a) whether the stream connection's identify must
use the bot's ACTIVE voice session_id or a fresh one; (b) actual video codec Discord
sends for camera vs GoLive (measure payload type live; H264 assumed, VP9/AV1 possible
→ add depacketizers only if observed).
@@ -1,93 +0,0 @@
# Spec: Fix Video Capture — Eagerly Establish the Selfbot Voice Connection at Join Time
Status: PLANNED
Date: 2026-08-31
Author: Hermes
Related: `docs/specs/2026-08-31_video-receive-phaseC-spec.md` (Phase C build, made Option A this fix)
## Symptom (from live logs, 2026-08-31 ~12:34)
A user was actively screen-sharing + on camera in the recorded voice channel.
The gateway recorded MANY users' audio (.ogg) fine, but video capture produced
nothing. The only video signal in `journalctl -u gmw-discord-gateway` was:
```
[VOICE (guild:2)]: Sending voice state update: {"self_mute":false,...,"flags":2}
[VOICE] received voice state update: {member hunterz ...} # OTHER user, not bot
[VOICE] connection? true, guild session channel
[VOICE (guild:2)]: Setting sessionId <S> (stored as "undefined")
[VOICE (guild:2)]: Authenticated with sessionId <S> # debug print only
[VOICE (guild:2)]: Authenticate failed - VOICE_CONNECTION_TIMEOUT # +15s
video-recorder: userId=..., "Connection not established within 15 seconds."
```
## Root cause (verified against discord.js-selfbot-v13 3.7.1 source)
The gateway records audio via `@discordjs/voice` (`joinVoiceChannel` + adapter).
Video receive lives on the SEPARATE selfbot `ClientVoiceManager.connection`
(a singleton `VoiceConnection`). `videoRecorder.ts` currently calls
`client.voice.joinChannel(channel)` LAZILY — only when a `voiceStateUpdate`
shows `newState.streaming === true`.
At that moment the bot is ALREADY connected to the channel via @discordjs/voice.
A selfbot `joinChannel` then does `VoiceConnection.authenticate()` →
`sendVoiceStateUpdate()`, and waits for a fresh `VOICE_SERVER_UPDATE`
(`setTokenAndEndpoint`) + `VOICE_STATE_UPDATE` (`setSessionId`) to reach
`checkAuthenticated()` (needs token+endpoint+sessionId). Because the bot is
already in an established voice session, Discord does NOT emit a new
`VOICE_SERVER_UPDATE` for the lazy selfbot re-join → token/endpoint never set →
15s `VOICE_CONNECTION_TIMEOUT`.
This is fatal to video: `joinStreamConnection(userId)` (STREAM_WATCH op 20) and
`receiver.createVideoStream(userId, out)` (Recorder/ffmpeg) BOTH live on the
parent selfbot `VoiceConnection` and require it `CONNECTED` (its own voice
WS+UDP socket feeds `PacketHandler.push`, authenticated with
`authentication.secret_key`).
## Fix — establish the selfbot connection eagerly, at voice-join time
The selfbot `VoiceConnection` must exist and be `CONNECTED` before any streamer
appears. Establish it once, synchronously alongside the @discordjs/voice join in
`recorder.startRecording`, so it rides the bot's FRESH voice join — when Discord
DOES emit VOICE_SERVER_UPDATE. Then cache it and let `videoRecorder` reuse it.
Ordering: in `startRecording`, after the @discordjs/voice `joinVoiceChannel`
returns (and retries) — fire `ensureSelfbotVoice(channel)` best-effort:
1. `await client.voice.joinChannel(channel, { selfMute:false, selfDeaf:false,
selfVideo:false })` (rejects ~VOICE_CONNECTION_TIMEOUT on failure → log +
return null; do NOT block audio).
2. Cache the returned selfbot `VoiceConnection` keyed by guildId.
3. Wire teardown: on `recorder` voice stop / destroyed → `untrackChannel` +
destroy the cached selfbot connection (`disconnect()`).
`videoRecorder.startVideoRecording` then uses the cached selfbot connection:
- If cached & `status === CONNECTED` → use it.
- Else → fall back to a lazy `joinChannel` (still best-effort).
## The two-connection coexistence risk (must verify live)
@discordjs/voice (audio) and the selfbot `VoiceConnection` (video) each open
their OWN low-level voice WS+UDP on the same session. The spec's original
open-question flagged this. Mitigations:
- Clear logging: `Selfbot voice connected (guild=...)`, plus a periodic
`djs/voice status` log so we can confirm audio stays `READY` while the selfbot
connection is up.
- If Discord kicks/breaks the audio connection, logs will show
@discordjs/voice `Disconnected`/reconnect churn — we detect and pivot.
## Files touched
- `services/discord-gateway/src/modules/voice-recording/videoRecorder.ts`:
add `ensureSelfbotVoice(channel)` (return cached/connected), use it in
`startVideoRecording`, add `destroyGuildSelfbotVoice(guildId)`,
richer status logging.
- `services/discord-gateway/src/modules/voice-recording/recorder.ts`: call
`ensureSelfbotVoice(channel)` after `joinVoiceChannel` (best-effort);
call `destroyGuildSelfbotVoice` on voice stop/destroy.
- Tests: `tests/videoRecorder.test.ts` (update to assert eager-connection reuse
+ status gating).
## Verification
1. `pnpm typecheck` + `pnpm build` + biome clean (discord-gateway).
2. Tests green.
3. Commit + push; CI `Build & Deploy (Nix)` green, service restarts.
4. LIVE (deploy): join a channel with the bot → journal shows
`Selfbot voice connected` (parent CONNECTED). When a member streams →
`Sender signal screenshare` / `Video recorder ready` + a `.mkv` under
`<RECORDINGS_DIR>/<uid>/video-*.mkv`; playable via ffmpeg. Confirm audio
recording still flows (no djs/voice reconnect churn).
@@ -1,97 +0,0 @@
# Spec: Record Other Users' Video (Camera / Screen Share) — Phase C
Status: PLANNED (not built)
Date: 2026-08-31
Author: Hermes
Related: `docs/specs/2026-08-30_video-record-receive-spec.md` (Phase A/B — raw UDP hook, superseded for receive)
## TL;DR — what changed vs Phase A/B
Phase A/B (commit `999c054b` etc.) hooked `@discordjs/voice`'s UDP socket to capture non-opus RTP and
depacketize H264 → mp4. **It captured ZERO video** because `@discordjs/voice` never authorizes the bot to
receive others' video (no STREAM_WATCH). This spec replaces that approach with the **native, selfbot-lib
receive path**, which is battle-tested and does the authorization + decryption + ffmpeg muxing for us.
## Ground truth (verified in discord.js-selfbot-v13 3.7.1 source)
1. `ClientVoiceManager.joinChannel(channel, config)` → a **selfbot `VoiceConnection`** with
`.receiver` (`VoiceReceiver` → `PacketHandler`). [ClientVoiceManager.js:102-118]
2. `VoiceConnection.receiver` is created in the constructor. [VoiceConnection.js:140]
3. `VoiceReceiver.createVideoStream(user, output)` → `PacketHandler.makeVideoStream` → **`Recorder`**
(ffmpeg that muxes H264+Opus RTP over UDP → **Matroska (.mkv)**). [Receiver.js, Recorder.js]
4. `PacketHandler` routes: video RTP → Recorder UDP 65506, opus RTP → UDP 65510; decodes all via
`connection.authentication.{secret_key, mode}` (supports `aead_aes256_gcm_rtpsize` and
`aead_xchacha20_poly1305_rtpsize` = DAVE-compatible). [PacketHandler.js:115-155, 195-240]
5. `StreamConnectionReadonly.joinStreamConnection(userId)` + `sendSignalScreenshare()` sends
gateway op `STREAM_WATCH` so Discord actually forwards the streamer's RTP to us. [VoiceConnection.js:1100-1240]
6. `VoiceState.streaming` = `data.self_stream ?? false` — lets us detect a streamer on voice state update. [VoiceState.js:94]
## Problem / the crux
The gateway's voice today is **`@discordjs/voice`** (audio + music + GoLive-send). The selfbot-lib
video-receive path lives on the **selfbot-lib `VoiceConnection`** — a separate voice stack. Two options:
### Option A (RECOMMENDED): Parallel selfbot video-watch connection
Keep `@discordjs/voice` for everything it does today. Add a **second, selfbot-lib voice connection**
to the same channel whose ONLY job is to watch + record others' video.
- Pros: zero regression risk to audio/music/screenshare-send; uses native `createVideoStream` → mk4.
- Cons: two voice connections for the same bot user in one channel. Need to verify Discord tolerates it
(real selfbots like Discord-RE do exactly this for multi-stream). The selfbot lib's `joinChannel`
reuses `ClientVoiceManager.connection` (it's a singleton) — see caveat below.
### Option B: Migrate primary voice to selfbot lib
Make the selfbot `VoiceConnection` THE voice layer (it also does audio via `receiver.createStream`).
- Pros: one connection; video+audio unified.
- Cons: large refactor; high regression risk to the entire existing audio/music/GoLive stack. NOT chosen now.
## CAVEAT — ClientVoiceManager.connection is a singleton
`ClientVoiceManager.connection` is a single `VoiceConnection`. The gateway's `@discordjs/voice` adapter and
the selfbot lib both drive the same client voice state. Need to verify whether `client.voice.joinChannel()`
can coexist with the active `@discordjs/voice` session, or whether we must create the selfbot VoiceConnection
manually / re-use the existing voice state. This is the #1 technical risk to validate in the spike before
committing to Option A.
## Implementation plan (Option A)
### 1. Streamer detector (new: `modules/voice-recording/videoRecorder.ts`)
- Listen to voice state updates (`client.on('voiceStateUpdate')` or the existing voice-state hook).
- When `voiceState.streaming === true` for a member in the bot's channel → candidate to record.
- Skip bot's own user id (unless we also want self-video; default skip).
### 2. Watch + record wiring
- Ensure a selfbot-lib `VoiceConnection` exists for the channel (spike: `client.voice.joinChannel(channel)`,
fallback: build a `VoiceConnection` directly from the existing voice auth).
- `await selfbotVoiceConn.joinStreamConnection(userId)` → STREAM_WATCH op 20.
- `const recorder = selfbotVoiceConn.receiver.createVideoStream(userId, outPath)` where outPath points under
`<RECORDINGS_DIR>/<uid>/video-<streamKey>-<ts>.mkv` (Recorder outputs MKV natively).
- On `recorder.on('ready')` → mark recording; `recorder.on('closed')` → finalize.
- Transcript later: MKV → mp4 via ffmpeg (Phase B `muxToMp4` can accept mkv) for dashboard playback.
### 3. Teardown
- When `voiceState.streaming === false` / user leaves / channel emptied → `recorder.destroy()`,
`selfbotVoiceConn.streamWatchConnection.delete(userId)` / `sendStopScreenshare()`.
### 4. Frontend (Phase UI, later)
- oRPC/backend list `.mkv` per call session + FE `<video>` player (mirror audio recordings UI).
## Files touched
- `services/discord-gateway/src/modules/voice-recording/videoRecorder.ts` (new)
- `services/discord-gateway/src/modules/voice-recording/recorder.ts` (wire streamer detector on voice join)
- Possibly `voiceController.ts` (voice state update subscription)
- Tests: `tests/videoRecorder.test.ts` (mock selfbot VoiceConnection + Recorder)
## Verification
1. `pnpm typecheck` + `pnpm build` + biome clean in discord-gateway.
2. Unit: Recorder wiring + streamer detection with mocked VoiceConnection.
3. Live (deploy): user shares screen → journal shows `STREAM_WATCH` sent + `Recorder ready` + `.mkv` file
appears under recordings dir; playable via ffmpeg.
4. CI Build & Deploy (Nix) green.
## Open questions for spike (before full build)
- [ ] Can `client.voice.joinChannel()` run alongside the active `@discordjs/voice` session, or does the
singleton `ClientVoiceManager.connection` collide / tear down the existing audio connection?
- [ ] Does the selfbot `VoiceConnection` need the bot's `video: true` flag in IDENTIFY to receive video
(it advertises `streams` in IDENTIFY — see BaseMediaConnection/identify vs selfbot VoiceConnection)?
- [ ] Does `Recorder` (spawns system ffmpeg, UDP loopback on 65506/65510) work in the Nix store runtime
(ffmpeg-headless on PATH confirmed; UDP loopback fine)?
@@ -1,33 +0,0 @@
# Video Recording Splitting — Like Voice Recording
## Goal
Camera + screen share (stream watch) recording should split into per-burst
segments just like voice recording does — each time a streamer pauses/stops
and resumes, a new MP4 segment is created and registered in the DB + uploaded.
## Voice Recording Model (to replicate)
1. `receiver.speaking.start` → new OGG segment per burst
2. AfterSilence (4000ms) → stream "end" → segment finalized + uploaded
3. Each segment → DB insert → OGG→MP3 transcode → upload → update DB
4. File stored as `<userId>/<startTime>.ogg` + `.json`
## Video Recording Splitting
1. DAVE video RTP → depacketize H264 → write to current segment .h264
2. Silence detection: no H264 packets for 4000ms → close segment → flush →
mux to MP4 → insert DB record → upload → start new segment on next packet
3. Each segment: `<userId>/video-<channelId>-<startTime>.h264` → `.mp4`
4. DB: reuse `voice_recordings` table (filename indicates video, e.g. `video-XXX-1234.mp4`)
5. Upload: MP4 to TeleUploader (no transcode needed — MP4 plays everywhere)
## Files Modified
- `services/discord-gateway/src/modules/voice-recording/streamWatchReceiver.ts`
— Main change: silence-based splitting + DB registration + upload
## Constants
- `VIDEO_SILENCE_MS = 4000` (matches voice AfterSilence)
- `VIDEO_MIN_SEGMENT_MS = 1000` (skip segments <1s — avoid noise)
## Verification
- `pnpm typecheck` in `services/discord-gateway`
- `pnpm build` (dist/ is the deployed artifact)
- Push → CI deploy → live test with a streamer
@@ -1,81 +0,0 @@
# Spec: Selfbot-Viable Video Capture — manual screen-share watch command (Phase D)
Status: PLANNED (not yet built)
Date: 2026-09-02
Author: Hermes
Related: `docs/specs/2026-08-31_video-receive-phaseC-spec.md` (auto-receive, superseded
for selfbot), `gmw-ops/references/selfbot-presence-detection-limits.md`,
`gmw-ops/references/discord-voice-fork-video-receive.md`
## TL;DR — the decisive finding (verified live 2026-09-02)
User insists on keeping the **selfbot** (no bot-token migration). Live diagnostics prove
a selfbot CANNOT auto-detect other members' camera/share because:
- It never receives `VOICE_STATE_UPDATE` for other members (only its own).
- `guild.members.fetch()` → 403, `GET /channels/{id}/voice-states` → 404.
- No `GUILD_CREATE`, no `READY.broadcaster_user_ids` presence.
- `scanExistingStreamers` + `handleVoiceStateUpdate` (the only two `startStreamWatch`
triggers) are therefore both **dead on a selfbot**.
- No manual watch command exists today, so even on-demand capture is impossible.
→ The ONE selfbot-viable path is a **manual, operator-initiated STREAM_WATCH** on a
member known to be screen-sharing. Gateway op 20 (STREAM_WATCH) is **NOT gated on
bot-vs-user**; the DAVE handshake to Ready+MLS was already verified live in earlier
sessions. The receive/mux/segment/upload pipeline (`streamWatchReceiver.ts`) is already
built and only lacks a real streamer to produce its first `.mp4`.
Camera-of-others is NOT viable on a selfbot even with `unknown-ssrc` fallback:
`@discordjs/voice` `parsePacket` calls `daveSession.decrypt(packet, userId)` keyed per
REAL userId (vendor fork dist/index.js:2143), so a fake id selects no MLS decryptor →
garbage, not H264. (The uncommitted `unknown-ssrc` change was reverted this session.)
Selfbot CAN capture the OWNER's own video (its own VOICE_STATE_UPDATE + fork op12
videoSSRC are attributable), but `videoRecorder.ts` hard-skips its own id — parameterized
self-capture is a follow-up, not the default.
## Goal
Add a **manual watch command** so an operator can say "record <member>'s screen share"
and the gateway `startStreamWatch`s that member → DAVE watch → per-burst `.mp4` segments
(mirroring voice silence split) → upload → DB `voice_recordings` → dashboard `<video>`.
This is the only form of OTHER-member video capture a selfbot can deliver, and it is
genuinely buildable with the existing receive pipeline.
## Scope / files
Gateway (`services/discord-gateway`):
- New command type `VIDEO_WATCH` + handler in `command-handler/` (dedicated
`video.handler.ts`), routed via `createHandlerRegistry`.
- Handler resolves a VoiceChannel (from persisted `voice_auto_reconnect` / active
connections) + target memberId from the command payload, calls
`startStreamWatch(channel, memberId)` (already exported).
- Idempotent (startStreamWatch early-returns if a watch exists); a `VIDEO_UNWATCH`
command calls `stopStreamWatch(guildId, userId)`.
- Reply: success/failure via the standard `CommandReply` publish.
Backend (`services/backend`):
- oRPC procedure (or the existing command bridge) that publishes a `VIDEO_WATCH`
command to `backend:command` with `{ guildId, channelId, userId }`. Reuse the same
bridge the FE already uses for voice commands.
Frontend (`services/frontend`):
- A "Video Watch" control: pick a voice member + a "Record screen" button → calls the
backend procedure. Shows live status (watching / recording / segments uploaded).
(Each layer optional independently; gateway alone gives a Redis-testable path.)
## Verification
1. `pnpm typecheck` + `pnpm build` + `biome check src/` green in discord-gateway.
2. Unit test: handler publishes reply + calls startStreamWatch with the right args
(mock the module).
3. Live: operator invokes `!videorec <member>` while that member screen-shares →
journal shows `Sending STREAM_WATCH` → `STREAM_CREATE` → `DAVE watch READY` → `Video
burst opened` → `Video muxed to mp4` → a `video-*.mp4` appears under
`<recordingsDir>/<uid>/` and a `video-%` row lands in `voice_recordings`.
4. `Build & Deploy (Nix)` CI green.
## Out of scope (documented dead ends on selfbot)
- Auto camera/share capture of OTHER members (impossible at detection layer).
- Camera-of-others via `unknown-ssrc` (DAVE decrypt needs real userId).
- Bot-token migration (user declined).