Commit Graph
26 Commits
Author SHA1 Message Date
asepharyana e5304fde29 feat(gateway): add manual video-watch command for selfbot screen-share capture
A selfbot (user token) cannot auto-detect other members' camera/share
(no VOICE_STATE_UPDATE for others, 403 on member fetch). The only
selfbot-viable path to capture another member's SCREEN SHARE is an
operator-initiated STREAM_WATCH (gateway op 20, not gated on bot-vs-user).

Add video:watch / video:unwatch Redis commands routed via the existing
command handler to startStreamWatch/stopStreamWatch, which then does the
DAVE handshake + per-burst MP4 segmentation + DB insert + Tele upload
(already implemented in streamWatchReceiver).

- new VideoHandler (command-handler/video.handler.ts)
- register video:watch / video:unwatch in handler-registry + CommandHandler
- command constants COMMAND_VIDEO_WATCH / COMMAND_VIDEO_UNWATCH
- resolve active voice channel from voice controller + client cache
- 8 unit tests (videoHandler.test.ts)
- biome fixes for pre-existing test import ordering

All green: typecheck, build, lint (174 files), 200 tests.
2026-09-02 10:12:03 +07:00
asepharyana 4ec9685194 feat(gateway): split video recording into silence-based segments like voice
Video (camera + screen share) DAVE stream-watch now produces per-burst
MP4 segments instead of one long .h264 per watch:
- Detects VIDEO_SILENCE_MS (4000ms) of no H264 packets → closes the
  current segment, muxes to MP4, registers in voice_recordings + uploads
  to TeleUploader, then reopens for the next burst (mirrors voice AfterSilence).
- Per-watch segment counter + per-segment depacketizer reset + closing
  guard + write-error swallow so races (silence close vs in-flight UDP
  packet) never corrupt files or crash the gateway.
- Frontend: recordings deck renders a native <video> player for MP4 rows
  (detected by filename), keeps single-playback registry across audio+video.
2026-09-01 21:51:44 +07:00
asepharyana 12cc956329 feat(gateway): implement separate Piscina pools for text and media analysis to optimize processing 2026-08-31 22:59:27 +07:00
asepharyana 026c66a03a docs(gateway): record 4th critical fix — DAVE session at net.state.dave (f1a7b0c2) 2026-08-31 16:17:56 +07:00
asepharyana dea284357c docs(gateway): DAVE watch Ready + MLS handshake confirmed live; P4 = video burst only
Live deploy 15:51 reached DAVE watch READY + completed MLS handshake on the
stream RTC (0->1->2->3->4, MLS commit processed, heartbeats alive). scan-on-
join also confirmed: 'Scanned pre-existing streamers on join watched=1'.
P4 remaining = capture actual video RTP (burst->mp4) while a stream is live.
2026-08-31 15:54:04 +07:00
asepharyana f1da40691f docs(gateway): mark DAVE stream-watch P1-P3 done, P4 blocked on live streamer 2026-08-31 14:09:06 +07:00
asepharyana 0995c2db81 docs(gateway): spec + Phase-1 recon for DAVE-capable stream-watch video receive
Supersedes the eager-selfbot connection plan: Discord now REQUIRES DAVE (E2EE,
WS 4017) on all voice RTC, and discord.js-selfbot-v13's voice stack predates
DAVE, so its video-receive path (joinChannel + joinStreamConnection +
receiver.createVideoStream) cannot authenticate. @discordjs/voice 0.19.2 exports
VoiceWebSocket/VoiceUDPSocket/DAVESession/Networking + @snazzah/davey supports
MediaType.VIDEO/Codec.H264 decrypt, so we can build a DAVE-capable stream-watch
connection. Phased plan: prototype (P2), gateway integration (P3), live verify (P4).
2026-08-31 12:59:33 +07:00
asepharyana 6738279bad feat(gateway): record others' video via native selfbot watch/receive (Phase C)
Root cause of why Phase A/B captured zero video: @discordjs/voice is audio-only
and never sends the gateway STREAM_WATCH signal, so Discord never forwards a
member's video RTP to the bot. Live diagnostic confirmed: while members were
sharing, audio .ogg files flowed for many users but no non-opus RTP ever
arrived.

Correct path: discord.js-selfbot-v13 ships a complete native watch/record stack.
New src/modules/voice-recording/videoRecorder.ts:
- Detects streamers via voiceState.streaming on a single idempotent
  voiceStateUpdate listener.
- client.voice.joinChannel() (reuses the single session alongside
  @discordjs/voice) + joinStreamConnection(userId) -> STREAM_WATCH (op 20).
- receiver.createVideoStream(userId, path) -> PacketHandler -> Recorder
  (ffmpeg over UDP loopback) -> Matroska .mkv, decryption handled internally.

Wired in recorder.ts (trackChannel/untrackChannel) + bootstrap.ts
(setVideoRecorderClient/setVideoRecordingsDir). All best-effort; failures log
and never break existing voice/audio. Unit tests 6/6 (single listener, watch
handshake + mkv path, skip own video, idempotence, teardown). Full suite
20 files / 176 tests green; typecheck + build + biome clean.

UNVERIFIED live yet: needs deploy + a streamer to confirm Recorder ready +
playable .mkv.
2026-08-30 23:26:59 +07:00
asepharyana feca8a476c docs(plans): mark video receive Phase A+B done (capture + mp4 mux) 2026-08-30 21:50:16 +07:00
asepharyana 419946f38e feat(gateway): record others' video (camera/screen share) — Phase A capture
@discordjs/voice only decrypts/forwards AUDIO (opus) — its onUdpMessage drops
every non-opus RTP packet (dist/index.mjs:2068 `!== RTP_OPUS_PAYLOAD_TYPE`, 120).
Video RTP (H264 camera + screen share, plus VP8/VP9/AV1) arrives on the same UDP
socket but was silently discarded.

New videoReceiver.ts wraps receiver.onUdpMessage (like screenShareAudio.ts):
- detects video payload types (96/98/101/102/106/116/126/127),
- decrypts them with the connection secret key/encryptionMode via the SAME
  receiver.parsePacket path @discordjs/voice uses for audio (so DAVE + voice
  encryption are handled identically),
- depacketizes H264 to AnnexB (single NAL, STAP-A, and FU-A fragmentation),
  waiting for a keyframe (SPS/PPS/IDR) before writing,
- writes a raw .h264 elementary stream per user per burst under
  <RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.h264.

Attribution: a videoSSRC→user index is built from ssrcMap updates; a proximity
fallback mirrors screenShareAudio's inferScreenShareOwner. Bot's own video is
skipped.

Unit tests: tests/videoReceiver.test.ts (AnnexB start code, keyframe gating,
FU-A reassembly, orphan-fragment tolerance) — 5/5 green.

Phase A only (capture raw h264). Phase B (ffmpeg decode+mux to MP4/WebM +
persist) and Phase C (frontend playback) are follow-ups.
2026-08-30 21:20:32 +07:00
asepharyana 441b5282ef feat(gateway): capture NSFW/age-restricted messages but skip AI analysis + exclude from public archive
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.

NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
2026-08-30 20:19:34 +07:00
asepharyana 50967f6481 feat(gateway): capture NSFW/age-restricted messages but skip AI analysis + exclude from public archive
Previously isAgeRestrictedMessage() early-returned in messageCreate/messageUpdate,
so NSFW-channel messages were never stored at all. Now they are captured like
any message (visible in dashboard), while the existing age-restricted skip path
(queueMessageAnalysis -> buildAgeRestrictedSkipResult) marks them clean with flag
age_restricted WITHOUT calling the LLM.

NSFW content is also deliberately kept OUT of the Qdrant public semantic-search
archive (archiveMessageEmbedded skips when isAgeRestricted), so it can't be found
via public web search. No schema change needed (metadata already carries channel.nsfw).
2026-08-30 20:13:35 +07:00
asepharyana 7dfb4035b7 feat(gateway): persistent voice auto-reconnect — rejoin same channel after restart/reboot or unexpected drop 2026-08-30 14:03:12 +07:00
asepharyana a3a91aa2ce feat(voice): auto-detect speech language for transcription (drop forced 'en')
Whisper previously hardcoded language:en, mis-transcribing id/en-mixed
speech. Omitting 'language' makes Whisper auto-detect. Paired with enabling
AI_VOICE_TRANSCRIPTION_ENABLED (BWS secret gmw_ai_voice_transcription_enabled
= true) so new recordings are transcribed.

Spec: .hermes/plans/2026-08-30_recordings-v2-features-spec.md
2026-08-30 13:28:17 +07:00
asepharyana 75ff9274b3 feat(frontend): filter recordings per-speaker + export WAV for Audacity
- recordings page: speaker filter dropdown (server-side userId filter via
  existing recordings.list proc) + clear-filter + pagination reset on change
- WAV export per card: decode download_url via Web Audio, re-encode 16-bit PCM
  WAV (Audacity-ready) client-side, no server/ffmpeg needed
- EXPORT WAV (N) header button: concatenate visible (filtered) recordings
  into a single mono WAV for mixdown/analysis
- hooks: per-filter SWR cache key (recordings/<userId|all>); delete & live
  WS-sync invalidate every filter cache via key matcher
- new lib/audio/wav.ts: decodeAudio / audioBufferToWav / concatAudioBuffers /
  downloadBlob
2026-08-30 12:59:46 +07:00
asepharyana e3016a858a fix(voice-recording): stop missing start-of-burst audio & mid-burst splits
Root-cause fixes for 'banyak miss & terpotong' in the voice->recording flow:

- subscribe BEFORE collecting user metadata. receiver.speaking 'start' fires
  on the FIRST opus packet, and onUdpMessage forwards frames to the
  subscription only when one exists — every frame during the old
  await collectUserMetadata (a Discord REST roundtrip on cache miss) was
  dropped, cutting off the start of every burst. Now subscribe synchronously
  (guard first, no await in between), then fetch metadata in the background
  and discard the burst if the speaker turns out to be a bot.
- one segment per burst: drop the fixed 5s RECORDING_SEGMENT_MS rotation on
  the OGG path, which split continuous speech mid-word/sentence. Only the
  web-PCM decoder still rotates (bounds memory).
- finalize only once the underlying file has flushed to disk (wait on the
  write stream 'finish'), so upload/transcode reads a complete file.
- raise AfterSilence 3000->4000ms so natural pauses (thinking, interruptions)
  don't split one utterance into several recordings.
- lower the 'too short to keep' threshold 1000->300ms so brief replies
  ("ya", "siap") are kept instead of dropped.

All typecheck / biome(src/) / vitest (164) green.
2026-08-30 12:23:03 +07:00
asepharyana 14bd20f072 fix(frontend): mobile navbar navigation + SSR hydration mismatch
Root cause of 'navbar mobile tak bisa pindah halaman': Next <Link>
client-side navigation is dead app-wide. A React hydration mismatch
(#418: 'server rendered text didn't match the client') is thrown by the
SSR-seeded live feeds — relative times (formatRelativeTime(e.edited_at) /
m.created_at) computed with Date.now() render slightly differently on
server vs client, which breaks the Next client router (router.push is a
no-op). The desktop NavRail worked only because it uses plain <a href>
(hard navigation bypasses the broken router).

Fixes:
- mobile-nav.tsx: use plain <a href> (NOT Next <Link>), identical to the
  working sidebar NavRail, so mobile nav always navigates regardless of
  router state ('ikuti cara kerja sidebar').
- Add suppressHydrationWarning to the SSR-seeded relative-time spans so
  server/client drift no longer throws #418 (EditHistory, LiveModerationFeed,
  messages/results + detail rows, recordings, TermGlossary,
  ChannelCultureGlossary, CategoryDrilldown).

Verified on non-prod :4024 @375px: Voice/Media/Search all navigate, no #418
in console. Plan: .hermes/plans/mobile-nav-hydration-fix.md
2026-08-28 09:57:49 +07:00
asepharyana 842610b1af fix(ai-moderation): bedah delay attachment ~330s -> target <20s
Root cause (trace msg 1541417073245290638):
- Race-guard upload-pending balik results:[] diperlakukan sbg SUKSES
  -> row yatam 'processing' sampai cleanup 300s mengembalikan
- Vision gagal 3x utk GIF besar (SSE truncation) tanpa fallback

Fix:
- Sinyal eksplisit uploadPending dari worker race guard
- Classifier murni classifyIndividualWorkerResult(): upload_pending ->
  requeue pending + reschedule segera (250ms), bukan error palsu;
  empty-results ok:true kini error transien (bug silent-success mati)
- llmVision fallback stream:false sekali saat SSE truncation
- Safety-net cleanup stuck processing 300s -> 120s
2026-08-24 22:05:28 +07:00
asepharyana 1accfd9390 perf(ai-moderation): naikkan cache hit dgn guard akurasi
- Fase-1 exact-cache lookup: N query serial -> SATU query ANY($1::text[])
- Global reuse utk bare key legacy, HANYA verdict non-actionable
  (clean/flagless/action=none, conf>=0.85, umur<=72h) — flagged/warn
  tetap context-scoped
- Semantic cache dua-band: clean band 0.92 default, actionable tetap
  0.97; di antara band -> LLM (fail-open ke akurasi)
- hit_count kini di-increment (bulk UPDATE per batch) -> hit-rate terukur
- Cache hasil wikipediaSearch di Redis (6h, hanya hasil non-kosong)
- Memoize fetchUrlSafely utk type=text (LRU 30m + in-flight dedupe)
- makeImageCacheKey strip query CDN Discord (?ex/is/hm, format/width)
  -> attachment sama = satu key vision, skip re-download+re-vision

Spec: .hermes/plans/2026-08-24-ai-analysis-cache-optimization.md
Tests: +33 (cacheGuards, discordImageKeyNormalize, cacheBatchLookup)
2026-08-24 18:51:23 +07:00
asepharyana 5f42c17caa feat(frontend): tema monokrom + sidebar animasi game-menu + gate ringan mobile 2026-08-24 17:35:51 +07:00
asepharyana 16becd5340 perf(ai): kontiguitas batch budget + max_tokens dinamis + urutan kronologis RETURNING
- pickBatchWithinBudget: stop di overflow pertama (break), bukan skip —
  batch tetap prefix kronologis tanpa gap analisis di tengah timeline.
  Diekstrak ke batchBudget.ts (pure, estimator di-inject) + regression test.
- callModerationLLM: param opsional maxTokens; text/media caller menghitung
  ceiling dari estimasi prompt (floor 2048, cap 16384) — batch kecil tak
  lagi reserve window completion 16k.
- getPending/IncompleteMessagesByConversation: sort hasil UPDATE..RETURNING
  by created_at ASC — Postgres tak menjamin urutan, konsumen (anchor konteks
  messages[0], prefix batch) bergantung pada urutan kronologis.
2026-08-22 17:19:41 +07:00
asepharyana 4e0c21d86c feat(frontend): perbagus voice & audio playback UX
- Recordings: custom RecordingAudioPlayer (play/pause, buffering spinner,
  click-to-seek, time label, eq bars, single-playback antar kartu) +
  highlight kartu now-playing
- Media: thumbnail di disc hero + queue row, equalizer saat playing,
  badge 'up next', label Paused vs Now playing
- MiniPlayer global di AppFrame (fixed bottom-right, hidden on /media)
  menggantikan use-media-player.tsx dead provider (dihapus)
- Voice: mic level meter live (AnalyserNode RMS) + slider mic/listen volume
2026-08-22 12:53:18 +07:00
asepharyana 00e8d68ce5 feat(gmw): public features #7-14 — scam domains, top channels, hourly heatmap, category drill-down, coverage stats, channel culture glossary, term KB, edit history
ALSO fixes: dashboard.repository still JOINed dropped user_reputations table (listUsers/getUserDetail crash).
2026-08-18 20:45:12 +07:00
asepharyana 9b3134d767 feat(gmw): public features #2-#6 — live moderation feed, toxic topic trends, channel timeline, CSV export, activity heatmap
- Live Moderation Feed: gateway publishes discord:moderation:action (Redis) → backend WS emits moderation_action → public web shows realtime stream.
- Toxic Topic Trends: backend moderation.trends aggregates categories/severity/action_type (read-only) → SVG bar + donut.
- Channel Timeline: messages view gets Feed/Timeline toggle with date-grouped separators.
- CSV Export: client-side downloadCsv for moderation actions (no backend write scope).
- Activity Heatmap: backend messages.activity (per-hour volume by channel) → pure-SVG grid.

User reputation deliberately excluded — no such feature exists in the codebase.
All read-only / public-facing / fully automatic per project rules.
2026-08-18 17:43:02 +07:00
asepharyana 1ae19074ee feat(gmw): moderation explainability + semantic message search
- Persist structured verdict (flags/severity/confidence/evidence) on
  moderation_actions so the public web can show WHY a message was moderated.
- Add a persistent Qdrant archive collection (gmw_message_archive); embed
  every captured message at capture time (fire-and-forget, best-effort).
- Public semantic search over the archive (backend oRPC + FE toggle on the
  messages view). Both features are read-only/public and fully automatic.

Migration: 0015_add_moderation_explainability.sql
2026-08-18 15:11:01 +07:00
asepharyana 3c2c1c3b15 Add Puppeteer scripts for navigation testing and debugging
- Created nav-debug.cjs to log anchor tags and simulate clicks on the Voice navigation link, capturing click events and page navigation.
- Added nav-test.cjs to test the Voice link click and log the URL at various intervals, capturing any page errors.
- Introduced nav-test2.cjs to check the presence of specific elements on the /voice/ page and log any console errors.
- Implemented nav-test4019.cjs to monitor network requests and responses related to the Voice navigation, verifying button presence and click functionality.
2026-08-15 19:23:17 +07:00