Files
GMW/.hermes/plans/2026-09-01_video-recording-splitting.md
T
asepharyana 4ec9685194 feat(gateway): split video recording into silence-based segments like voice
Video (camera + screen share) DAVE stream-watch now produces per-burst
MP4 segments instead of one long .h264 per watch:
- Detects VIDEO_SILENCE_MS (4000ms) of no H264 packets → closes the
  current segment, muxes to MP4, registers in voice_recordings + uploads
  to TeleUploader, then reopens for the next burst (mirrors voice AfterSilence).
- Per-watch segment counter + per-segment depacketizer reset + closing
  guard + write-error swallow so races (silence close vs in-flight UDP
  packet) never corrupt files or crash the gateway.
- Frontend: recordings deck renders a native <video> player for MP4 rows
  (detected by filename), keeps single-playback registry across audio+video.
2026-09-01 21:51:44 +07:00

1.6 KiB

Video Recording Splitting — Like Voice Recording

Goal

Camera + screen share (stream watch) recording should split into per-burst segments just like voice recording does — each time a streamer pauses/stops and resumes, a new MP4 segment is created and registered in the DB + uploaded.

Voice Recording Model (to replicate)

  1. receiver.speaking.start → new OGG segment per burst
  2. AfterSilence (4000ms) → stream "end" → segment finalized + uploaded
  3. Each segment → DB insert → OGG→MP3 transcode → upload → update DB
  4. File stored as <userId>/<startTime>.ogg + .json

Video Recording Splitting

  1. DAVE video RTP → depacketize H264 → write to current segment .h264
  2. Silence detection: no H264 packets for 4000ms → close segment → flush → mux to MP4 → insert DB record → upload → start new segment on next packet
  3. Each segment: <userId>/video-<channelId>-<startTime>.h264 → .mp4
  4. DB: reuse voice_recordings table (filename indicates video, e.g. video-XXX-1234.mp4)
  5. Upload: MP4 to TeleUploader (no transcode needed — MP4 plays everywhere)

Files Modified

  • services/discord-gateway/src/modules/voice-recording/streamWatchReceiver.ts — Main change: silence-based splitting + DB registration + upload

Constants

  • VIDEO_SILENCE_MS = 4000 (matches voice AfterSilence)
  • VIDEO_MIN_SEGMENT_MS = 1000 (skip segments <1s — avoid noise)

Verification

  • pnpm typecheck in services/discord-gateway
  • pnpm build (dist/ is the deployed artifact)
  • Push → CI deploy → live test with a streamer