Commit Graph
421 Commits
Author SHA1 Message Date
asepharyana ef41898a60 feat: add sexual/provocative username detection to LLM prompt rules
- rules.ts: explicit rule that sexual/provocative username terms
  (Pecinta Pria, Cinta, pacar, janda, bokep, hot, seks, nude, telanjang)
  MUST be flagged as offensive_username with low severity
- output.ts: output schema case for sexual/provocative username
  violations, always status: warn, severity: low, never flagged/delete

No hardcoded keyword lists — fix is at the LLM prompt level only.
2026-09-18 18:11:15 +07:00
asepharyana 5c133f7302 feat(gateway): tinyfish web search as fallback when wikipedia misses
- new tinyFishSearch module: GET api.search.tinyfish.ai with X-API-Key,
  maps top-3 to SearchResult shape, never throws (all failure modes -> [])
- wikipediaSearch: on wiki miss, one tinyfish attempt; hits cached 6h
  under the same key so fallback latency is paid once
- termGlossary: on summary miss, top tinyfish hit becomes the definition
  (persisted permanently like wiki defs); miss keeps 1h sentinel
- config: TINYFISH_API_KEY (empty = fallback disabled), ENABLED,
  BASE_URL, TIMEOUT_MS, LOCATION, LANGUAGE knobs
- tests: 6 coverage for disabled/mapping/non-OK/network/bad-json

API key NOT committed — set TINYFISH_API_KEY in BWS gmw secrets.

Verified: typecheck + lint clean, 216/216 tests pass, live probe
'gubernur jawa barat' returned 3 mapped results
2026-09-14 00:44:38 +07:00
asepharyana 9685fa02be perf(gateway): retry transient attachment + wikipedia failures
- attachmentUploader: downloadDiscordAttachment retries CDN timeouts
  via retryWithBackoff (ATTACHMENT_RETRY_ATTEMPTS, 1s-8s backoff);
  AbortError normalized so 403/404 refresh path never fires on timeouts
- attachmentUploader: uploadAttachmentToTele uses ATTACHMENT_RETRY_ATTEMPTS
  instead of retries:0 (Tele 5xx under load was failing outright)
- wikipediaClient: wikipediaSummary retries once on abort/timeout
  (100% of prod summary errors were aborts); termGlossary drops its
  redundant second call (was up to 4 reqs/term under miss+retry)

Verified: typecheck + lint clean, 210/210 tests pass
2026-09-14 00:13:56 +07:00
asepharyana 023217b260 feat(moderation): skip AI analysis for music bots via AI_SKIP_ANALYSIS_USER_IDS
Jockie Music (user 411916947773587456) posts now-playing embeds/spotify links
~1347 captured messages — every one consumed a moderation LLM call for zero
signal and contributed to batch timeouts. Config AI_SKIP_ANALYSIS_USER_IDS
(default=Jockie) skips them at ALL three analysis paths:
- queueMessageAnalysis entry (direct skip-result like age-restricted)
- batchScheduler processing (pre-batch filter)
- individual recovery path (no fallback spam for already-skipped authors)

Skip-result mirrors age_restricted: status=clean, flags=[skip_analysis_user],
action=none — stays visible in the dashboard, never analyzed.
2026-09-09 23:38:32 +07:00
asepharyana db4e84f057 fix(moderation): text batch timeout 45s→75s — router text model regularly exceeds 45s
The text model behind omniroute/9router consistently takes >45s on long-context
batches. At 45s every such batch fell through to the individual-fallback
queue which re-runs with its own timeout, then exhausted to ai_status=error.
75s keeps the bounded budget while letting the first-pass batch succeed.
2026-09-09 21:38:16 +07:00
asepharyana f85af3952a fix(moderation): audit fixes — media/vision timeout 120s, Qdrant retry w/ backoff, JSON repair in LLM caller
- AI_LLM_MEDIA_ANALYSIS_TIMEOUT_MS 60s→120s + vision 60s→120s: vision model
  via router regularly exceeded 60s, dropping media batches into the
  individual-fallback chain then exhausting into ai_status=error.
- Qdrant upserts: retryWithRetry() wraps PUT /points with exponential
  backoff (3 attempts, jitter) for transient 408/abort/ECONNRESET — the
  41 six-hour 'Qdrant upsert failed — semantic entry skipped' warnings were
  single-hop timeouts on a healthy-but-loaded Qdrant.
- LLM caller: on parse failure, attempt extractJson() structural repair of
  the raw content (models with thinking disabled sometimes emit JSON as
  plain text) before giving up and re-requesting.
2026-09-09 21:29:01 +07:00
asepharyana 276062fdd4 feat(ai): scope persistence, parallel tools, Prometheus token/cache counters
C2 full scope persistence:
- getRecentConversationContext now returns per-turn guildId/channelId
- processMessage merges historical scope when current request is unscoped
- chatbot remembers server context across all 8 history exchanges (not just 3)

A4+Metrics:
- Add incrementCounterBy(name, delta, labels) to gateway-metrics
- Token-usage counters: llm_tokens_total{model, type, label} per batch
- Cache hit counters: moderation_cache_hits{type} exact/semantic-qdrant/semantic-pg
- Cache miss counters: moderation_cache_misses per batch
- Prometheus /metrics now exposes cost + cache hit-rate for dashboards

Performance:
- Parallel tool execution within each chatbot round (Promise.all)
- All tool results collected before sending to model (ordering preserved)
- Tool failure now logged with structured warning (chatbot.tools.ts)

Verified: backend tsc+biome 37/37, discord-gateway tsc+biome 210/210
2026-09-04 15:32:03 +07:00
asepharyana 85204ca6f0 fix(moderation): enforcement safety net — username-only offense never auto-deleted
Even with the prompt firewall (username vs content), the LLM can still
occasionally mis-apply a content-level zero-tolerance flag (sara /
conflict_instigation) to a message whose ONLY violation is the username
(e.g. 'matikanetanyahu'). The auto-delete eligibility check only recognized
exact offensive_username flags, so such false positives still deleted the
message.

Add a belt-and-suspenders guard in isNicknameOnlyViolation: if the flag set
is entirely username-attributable (offensive_username/sara/conflict_instigation)
AND the analysis text corroborates that the violation is username-only with
clean message content, route to nickname-reset instead of message deletion.

Adds 6 test cases covering the real matikanetanyahu scenario and the
false-positive/negative boundaries.
2026-09-03 20:39:17 +07:00
asepharyana 68fbaa631a fix(moderation): prevent username-only violations from triggering content-level zero tolerance
Add explicit FIREWALL PENILAIAN rule separating username assessment from
message content assessment. Username containing political/religious terms
(matikanetanyahu etc) is assessed as offensive_username with severity low
— never triggers sara/conflict_instigation zero tolerance for content.

Changes:
- rules.ts: Add FIREWALL section before LARANGAN BERAT; clarify each
  zero-tolerance rule applies to ISI PESAN only; update hierarchy #9
- examples.ts: Fix example #11 status flagged→warn for clean username;
  add example #11b with matikanetanyahu case
- output.ts: Explicit status/action mapping for username-only violations
2026-09-03 20:05:37 +07:00
asepharyana 955da396c7 debug(stream-watch): add per-packet decrypt diagnostics — why VIDEO-PKT fires but no mp4? 2026-09-02 19:24:25 +07:00
asepharyana ccfdbc860e revert(gateway): source defaults kembali ke imrnes (100.121.180.82) — outage usai 2026-09-02 16:02:29 +07:00
asepharyana e6aa9af283 fix(gateway): make moderation score optional — LLM omits it in media batches
result.score was required by zod; the LLM (gemini-3.5-flash-lite via 9router)
occasionally omits it for media batches, hard-failing the whole batch parse
('Zod validation failed: expected number, received undefined' at
results[0].score). Callers already null-coalesce (result.score ?? 0) and the
parser clampScore()s it, so requiring it only caused parse failures.
Adds regression tests: media-batch without score parses (score->0), and
score-present responses still parse with the value.
2026-09-02 13:15:05 +07:00
asepharyana 4c38d53972 fix(gateway): local infra defaults + qdrant collection retry-on-failure
- AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1 (was imrnes :20128/api/v1)
- QDRANT_URL fallback -> http://127.0.0.1:6333 (was imrnes :6333)
- ensureQdrantCollection: reset memoised promise on failure so a mid-way
  recreate abort (DELETE done, PUT failed) does not leave the collection
  permanently missing until process restart
- tests: qdrantEnsure.test.ts (3 cases: retry-on-failure, recreate, idempotent)
2026-09-02 12:57:54 +07:00
asepharyana e5304fde29 feat(gateway): add manual video-watch command for selfbot screen-share capture
A selfbot (user token) cannot auto-detect other members' camera/share
(no VOICE_STATE_UPDATE for others, 403 on member fetch). The only
selfbot-viable path to capture another member's SCREEN SHARE is an
operator-initiated STREAM_WATCH (gateway op 20, not gated on bot-vs-user).

Add video:watch / video:unwatch Redis commands routed via the existing
command handler to startStreamWatch/stopStreamWatch, which then does the
DAVE handshake + per-burst MP4 segmentation + DB insert + Tele upload
(already implemented in streamWatchReceiver).

- new VideoHandler (command-handler/video.handler.ts)
- register video:watch / video:unwatch in handler-registry + CommandHandler
- command constants COMMAND_VIDEO_WATCH / COMMAND_VIDEO_UNWATCH
- resolve active voice channel from voice controller + client cache
- 8 unit tests (videoHandler.test.ts)
- biome fixes for pre-existing test import ordering

All green: typecheck, build, lint (174 files), 200 tests.
2026-09-02 10:12:03 +07:00
asepharyana 43594af3c8 fix: CI biome errors + enhanced GUILD_CREATE/READY voice_states diagnostic
- live-speaker.ts: unused var [id] in clearAllSpeakers → [_]
- migrate.ts: useTemplate string concat → template literal
- streamWatchReceiver: remove unused watchKey param from closeCurrentSegment
- videoRecorder: log all raw WS event types + READY sessions.voice + broadcaster_user_ids
2026-09-02 02:31:32 +07:00
asepharyana fdcd53d149 debug(video): enhanced diag — READY sessions.voice + broadcaster_user_ids + all raw types
Dump all WS raw event types after 10s (see if GUILD_CREATE exists at all),
and log broadcaster_user_ids + sessions.voice from READY payload.
Selfbot may RESUME (skip GUILD_CREATE) — voice data may only be in READY.
2026-09-02 02:26:26 +07:00
asepharyana bc66babd45 debug(video): log raw GUILD_CREATE/READY/GUILD_MEMBERS_CHUNK shape
Temporary diagnostic to see whether GUILD_CREATE.voice_states actually
reaches the selfbot raw listener (selfbot-v13 emits everything via
WebSocketShard Events.RAW). Will remove once root cause confirmed.
2026-09-02 02:10:22 +07:00
asepharyana ac32179bb1 fix(video): detect pre-existing voice users from GUILD_CREATE.voice_states
Real root cause of 'tidak mendeteksi user yg sudah ada di voice':

The selfbot's discord.js-selfbot-v13 GUILD_CREATE handler only sends
GUILD_SUBSCRIPTIONS_BULK — it DROPS d.voice_states from the payload.
So when the gateway (re)starts, every user who was ALREADY in the voice
channel (with camera or screen-share on) is invisible: channel.members is
empty, and REST fallbacks don't work for user tokens (verified live:
GET /channels/{id}/voice-states -> 404, GET /guilds/{id}/members -> 403).

Fix: register our own raw listener for GUILD_CREATE, capture
d.voice_states per guild (buffer it), and consume it in
scanExistingStreamers when the channel gets tracked (voice join happens
after GUILD_CREATE). This is the ONLY user-token-compatible source of
'who is in voice with video right now'.

- videoRecorder: +pendingGuildCreateVoiceStates buffer, +raw GUILD_CREATE
  listener, Path C consumes buffered states (self_video/self_stream,
  matching channel, skip self) before the REST fallback
- scanExistingStreamers: Path A cache -> Path C GUILD_CREATE -> Path D
  REST members.fetch() best-effort
- +1 test: GUILD_CREATE voice_states buffer -> watch camera+screen-share,
  skip self/no-video/other-channel (192/192 pass)
- typecheck + build + biome clean (1 pre-existing warning)
2026-09-02 02:03:31 +07:00
asepharyana cee33c0d1f fix(video): detect pre-existing voice members after gateway restart
Root cause: scanExistingStreamers only read channel.members, which is
EMPTY after a gateway restart because the selfbot's guild member cache
hasn't been populated yet. So any user who was ALREADY on camera /
screen-sharing when the bot (re)joined was never detected → no video.

Fix: when the member cache is empty, fall back to guild.members.fetch()
(REST GET /guilds/{id}/members — user-token compatible; the bot-only
GET /channels/{id}/voice-states returns 404 for selfbots, verified live)
which populates member.voice states, then re-scan channel.members for
streaming/selfVideo.

- scanExistingStreamers is now async; trackChannel fire-and-forgets it
- logs source=cache vs source=rest-members for observability
- +1 test: cold-start channel.members empty → REST fetch → watch camera user
- 191/191 tests pass, typecheck + build + biome clean
2026-09-02 01:13:50 +07:00
asepharyana dd9f2f6bbc fix(voice): self-undeafen/self-unmute bot instantly on server mute/deafen
Previously forceSelfServerUnmuteUndeafen only ran on video-watch attempts and
after voice reconnects — so when an admin server-muted or server-deafened the
bot, it stayed muted/deafened for minutes (or forever if no streamer came on).

Add registerSelfVoiceStateGuard: a voiceStateUpdate listener that detects the
bot's own serverMute/serverDeaf transition to true and immediately re-issues
mute:false,deaf:false. Wired at client-ready in bootstrap.ts. Best-effort,
idempotent, never blocks the gateway.
2026-09-02 00:30:57 +07:00
asepharyana 7122258d7c fix(voice): read SSRCMap internal map via 'map' (not '_map')
videoReceiver.getSsrcInternalMap read asAny._map, but @discordjs/voice
0.19.x exposes the SSRC map as the public field 'map'. So inferVideoOwner
(e.g. matching a video RTP SSRC to a user via audio-SSRC proximity) always
returned null -> every video packet was silently dropped after decryption.
Accept both 'map' and '_map'. Unblocks camera/screen-share attribute when
op12 does not carry a videoSSRC stream entry.
2026-09-02 00:10:02 +07:00
asepharyana 7c21629404 fix(voice): DAVE video decrypt fallback chain (VIDEO→AUDIO→passthrough)
streamWatchReceiver: DAVE decrypt for screen-share video packets was
returning null every time (VIDEO-PKT diag showed packets arriving but
no 'Video segment opened'). Three possible causes:
1. No VIDEO decryptor in MLS group (audio-only handshake)
2. GoLive stream tags video packets as AUDIO
3. Screen-share payloads are unencrypted above the AES layer

Fix: try MediaType.VIDEO first, then MediaType.AUDIO, then passthrough
(legacy-decrypted payload as-is). This covers all three modes without
breaking audio recording.
2026-09-01 23:54:07 +07:00
asepharyana 1367a2257f fix(gateway): detect camera-only video (selfVideo) for stream watch
Previously only member.voice.streaming (screen share, self_stream) triggered
video capture. Discord reports camera via self_video:true, so camera-only
users were never watched. Now scanExistingStreamers and handleVoiceStateUpdate
start a watch when streaming OR selfVideo is set, and stop it when both clear.
2026-09-01 22:26:38 +07:00
asepharyana 4ec9685194 feat(gateway): split video recording into silence-based segments like voice
Video (camera + screen share) DAVE stream-watch now produces per-burst
MP4 segments instead of one long .h264 per watch:
- Detects VIDEO_SILENCE_MS (4000ms) of no H264 packets → closes the
  current segment, muxes to MP4, registers in voice_recordings + uploads
  to TeleUploader, then reopens for the next burst (mirrors voice AfterSilence).
- Per-watch segment counter + per-segment depacketizer reset + closing
  guard + write-error swallow so races (silence close vs in-flight UDP
  packet) never corrupt files or crash the gateway.
- Frontend: recordings deck renders a native <video> player for MP4 rows
  (detected by filename), keeps single-playback registry across audio+video.
2026-09-01 21:51:44 +07:00
asepharyana fa72fe03cd feat(archive): show real channel/thread names in semantic search UI
Gateway archive embedder now parses metadata.channel.{channelName,threadName}
from each message and stores channel_name/thread_name in the Qdrant payload.
Backend exposes them; the semantic results card renders the thread name (or
channel name) instead of a raw #snowflake, with the ID as a last-resort
fallback for legacy points. Matches the message feed's channel-label logic
(getMessageChannelLabel).
2026-09-01 19:27:56 +07:00
asepharyana 26a690943d feat(archive): rich metadata in semantic search — username/channel/guild context + guild filter
Archive payload now stores username, channel_id, guild_id, thread_id and the
real message created_at (not embed time). Backend searchArray accepts an
optional guildId and applies a Qdrant payload filter so results can be scoped
to the guild being viewed. API/frontend expose the new fields and the
semantic results card shows who said it, in which channel, and when —
turning bare text blobs into contextual results. Old points fall back to
analyzed_at and omit the new fields gracefully.
2026-09-01 18:56:14 +07:00
asepharyana c704fbf7a5 fix(gateway): make voice transcription model configurable + router-compatible
- AI_VOICE_TRANSCRIPTION_MODEL config (default whisper-1) so the model can be a provider-qualified id (openrouter/openai/whisper-1) that actually has credentials through 9router/omniroute — bare whisper-1 maps to the openai provider which has none
- response_format json (not text): 9router proxies only json/verbose_json transcription responses; text returns 400
- parse text from the json response object
- prod env updated: model=openrouter/openai/whisper-1 (still needs OpenRouter STT balance — 402 until funded)
2026-09-01 18:25:14 +07:00
asepharyana 7ef86c81ca feat(ai): audit + harden embedding pipeline
- Normalize text before embedding (strip mentions/URLs/emoji/markdown/control chars, lowercase, truncate) on both write and query sides so vectors aren't diluted and tokens aren't wasted
- embeddingClient: retry embeddings (maxRetries 2), validate batch dimension consistency, preserve index alignment for empty-normalized texts
- archiveEmbedder: store normalized text in archive payload, skip empty-normalized content
- backend: normalize search queries, make archive search similarity threshold configurable (AI_LLM_EMBEDDING_ARCHIVE_MIN_SIMILARITY, default 0.6)
2026-09-01 18:01:47 +07:00
asepharyana 0e31aa06b8 feat(gateway): refactor term extraction and scoring logic into textSignals.ts for reuse 2026-08-31 22:59:27 +07:00
asepharyana 12cc956329 feat(gateway): implement separate Piscina pools for text and media analysis to optimize processing 2026-08-31 22:59:27 +07:00
asepharyana 2b6a1eca19 fix(gateway): re-assert server-undeafen+unmute before every video watch
The bot's own VOICE_STATE_UPDATE showed server-level deaf:true — a
server-deafened member is NOT sent the streamer's audiovisual RTP by Discord,
which is the likely reason no H264 arrives despite the DAVE watch reaching
Ready. The previous fire-and-forget forceSelfServerUnmuteUndeafen() ran once
after the first Ready join and silently reverted on reconnect/restart.

- Export forceSelfServerUnmuteUndeafen from recorder.ts; re-assert it (with
  read-back verification logging stillDeaf) at the START of every
  startStreamWatch() before STREAM_WATCH is sent (dynamic import avoids the
  recorder <-> videoRecorder <-> streamWatchReceiver module cycle).
- Re-assert it again after a successful voice reconnect.
- startStreamWatch() is now async; callers use void.
2026-08-31 19:16:56 +07:00
asepharyana 9606187861 feat(gateway): hexdump+ssrc of watch UDP packets
maxLen stayed 72 across 243 packets (no real H264, which is hundreds+ bytes) —
only 44-72-byte RTP packets on PT 76/72/73 arrive. Add ssrc + first-32-bytes
hex so we can identify exactly what Discord sends to the watch socket (control
packets vs stale video), which determines whether the gap is upstream routing
or whether large H264 packets are missing entirely.
2026-08-31 16:47:27 +07:00
asepharyana b423d21b23 feat(gateway): aggregate VIDEO-PKT diag — distinct PTs + maxLen
Enhance watch-socket diagnostic to report distinct RTP payload types seen and
the max packet length, so we can distinguish 'only small control packets arrive
(no real H264)' from 'H264 arrives but decrypt fails'. Live already confirmed
dave=true ready=true with packets flowing but no burst — need to know if they're
tiny 52-byte control packets (PT 73) or large H264.
2026-08-31 16:40:36 +07:00
asepharyana 266e149233 feat(gateway): add VIDEO-PKT diagnostic logging to stream-watch UDP handler
Instrument handleUdpMessage to log (rate-limited, first 3 then /20s) whether
video RTP packets actually reach the watch socket, and whether the DAVE session
is attached+ready and encryption key material present. Needed to diagnose why
no .h264 is written despite DAVE Ready + MLS: is the packet not arriving, or is
decrypt returning null?
2026-08-31 16:29:17 +07:00
asepharyana f1a7b0c2a1 fix(gateway): resolve DAVE session from net.state.dave, not connectionData
CRITICAL: djs/voice stores the DAVESession wrapper at net.state.dave
(createDaveSession assigns to state.dave on op4 SessionDescription), NOT
inside connectionData. decryptVideoPacket looked up connectionData.dave which
is ALWAYS undefined -> every video packet hit '!dave?.session' guard and was
silently dropped, so no .h264/.mp4 ever got written despite the handshake
reaching Ready.

Fix: pass net.state.dave as a separate arg (the wrapper has .session ->
Davey.DAVESession) so the DAVE-layer decrypt (MediaType.VIDEO) actually runs.
Typecheck + build pass, lint clean (src/), 179/179 tests.
2026-08-31 16:08:55 +07:00
asepharyana 87f1f8be8d feat(gateway): detect pre-existing streamers on bot voice join
If someone is ALREADY sharing screen / camera on when the bot joins the
channel, no voiceStateUpdate with streaming:true fires for them, so the
bot never sent STREAM_WATCH and missed their video entirely. trackChannel
now scans channel.members and starts a watch for anyone already streaming
(ignoring the bot itself and non-streamers). Idempotent: startStreamWatch
no-ops if a watch already exists. +2 tests (9/9 in videoRecorder).
2026-08-31 15:48:20 +07:00
asepharyana 4405647b33 fix(gateway): swap stateChange arg order so Ready actually attaches UDP
djs/voice Networking emits stateChange(oldState, newState), but the watch
handler declared (newState, oldState) -- reversed. So the code-4 Ready
branch (which attaches udp.on('message') + logs 'DAVE watch READY') never
fired when entering Ready; it only fired spuriously when LEAVING Ready.
Result: full DAVE handshake completed on the watch RTC (Ready + DAVE MLA +
video stream 21029 active 1920x1080@60) but no UDP listener -> no video
captured. Swap to (oldState, newState) so enter-Ready wires the socket.
2026-08-31 15:42:12 +07:00
asepharyana 4534b7a17d feat(gateway): log full watch Networking state transitions
Add watch-state N->M log on every djs/voice Networking stateChange (with
hasUdp flag) so the stream-watch connection's exact progression is visible:
OpeningWs(0)->Identifying(1)->UdpHandshaking(2)->SelectingProtocol(3)->
Ready(4). Pins down where the DAVE flow stalls instead of guessing from the
absence of logs. Pairs with debug:true + watch-djs-debug.
2026-08-31 15:36:05 +07:00
asepharyana b3e350d40f feat(gateway): enable djs/voice debug logging on watch Networking
Add debug:true to watch Networking options and wire net.on('debug') to
logger.info so djs/voice internal WS/DAVE state transitions appear in
journal. Without this, the stream-watch connection went silent after
'Streaming DAVE Networking' — no ready/error/close visible. Needed to
diagnose why the WS to stream endpoint 'c-sin14-xxx:2083' produced no
events.
2026-08-31 15:23:57 +07:00
asepharyana 75050bc088 fix(gateway): use rtc_server_id-1 as watch DAVE MLS channelId (WrongGroupId)
Live log (14:52) showed the stream-watch flow reaching STREAM_CREATE +
STREAM_SERVER_UPDATE but then DAVE processProposals threw
'ValidationError(WrongGroupId)' -- the Davey MLS session derived the wrong
group because connectionOptions.channelId was the guild voice channel id.
Per Discord-RE StreamConnection.daveChannelId = BigInt(serverId)-1n, the
stream-watch DAVE MLS group is keyed to rtc_server_id-1, not the vc channel.
Fix: pass BigInt(serverId)-1n as channelId to the watch Networking.

This error also surface as an uncaughtException that crashed the gateway
(systemd restarted it). Correct channel id prevents it at the root.
2026-08-31 14:54:32 +07:00
asepharyana 374fd5a9c2 fix(gateway): use guild voice sessionId for watch RTC identify
The previous code read the sessionId from the selfbot client's voice manager
(client.voice.connection), which is no longer established since ensureSelfbotVoice
was removed — it would have sent sessionId:'none' in the watch Networking
identify and been rejected. Read the active session from the guild
@discordjs/voice connection (getVoiceConnection(guildId).state.networking...
connectionOptions.sessionId) instead.
2026-08-31 14:03:00 +07:00
asepharyana 77c8454bb2 fix(gateway): replicate dual-layer DAVE+LTS decrypt for watch video RTP
The first streamWatchReceiver only did dave.session.decrypt(msg.subarray(12))
which skipped the outer legacy-AES layer Discord wraps around the DAVE payload
on every RTC packet (encrypt = dave.encrypt then aead_aes256_gcm with RTP
header as AAD). Port @discordjs/voice VoiceReceiver.decrypt/parsePacket
faithfully (header strip incl CSRC+extension+padding, AES-GCM auth tag, then
DAVE MediaType.VIDEO). Without this the .h264 would be garbage.
2026-08-31 13:55:58 +07:00
asepharyana 0ea76a8373 feat(gateway): DAVE-capable stream-watch video receive (Phase D)
Replace the dead selfbot-v13 video path (WS 4017 DAVE). videoRecorder
now delegates to a new streamWatchReceiver that:
- sends STREAM_WATCH (op 20) on voiceState.streaming
- opens a @discordjs/voice Networking to the watch RTC (STREAM_CREATE +
  STREAM_SERVER_UPDATE) with DAVE enabled
- decrypts H264 via Davey MediaType.VIDEO, depacketizes + muxes to mp4
- tears down on streaming-stop / leave / untrack

Remove ensureSelfbotVoice/createVideoStream/joinStreamConnection (dead).
recorder.ts no longer fires the futile eager selfbot join.
2026-08-31 13:46:24 +07:00
asepharyana 6aeebe7826 fix(gateway): establish selfbot voice connection eagerly so video (camera/screen-share) capture works
Video capture (camera + screen share) recorded ZERO frames because the selfbot
ClientVoiceManager.connection was created LAZILY — only when a user started
streaming. At that point the bot is already connected via @discordjs/voice, so
the selfbot re-join never gets a fresh VOICE_SERVER_UPDATE and times out with
VOICE_CONNECTION_TIMEOUT after 15s. joinStreamConnection (STREAM_WATCH) +
receiver.createVideoStream both need that selfbot VoiceConnection CONNECTED.

Fix: establish the selfbot VoiceConnection eagerly in recorder.startRecording,
BEFORE joinVoiceChannel, so it rides the bot's fresh join (Discord emits
VOICE_SERVER_UPDATE → selfbot authenticates). videoRecorder reuses the cached
connection per guild, tears it down on voice stop/destroy. Best-effort — never
blocks audio recording.

Verified: typecheck + build + biome (src/) green; 9/9 videoRecorder tests.
2026-08-31 12:44:36 +07:00
asepharyana 6738279bad feat(gateway): record others' video via native selfbot watch/receive (Phase C)
Root cause of why Phase A/B captured zero video: @discordjs/voice is audio-only
and never sends the gateway STREAM_WATCH signal, so Discord never forwards a
member's video RTP to the bot. Live diagnostic confirmed: while members were
sharing, audio .ogg files flowed for many users but no non-opus RTP ever
arrived.

Correct path: discord.js-selfbot-v13 ships a complete native watch/record stack.
New src/modules/voice-recording/videoRecorder.ts:
- Detects streamers via voiceState.streaming on a single idempotent
  voiceStateUpdate listener.
- client.voice.joinChannel() (reuses the single session alongside
  @discordjs/voice) + joinStreamConnection(userId) -> STREAM_WATCH (op 20).
- receiver.createVideoStream(userId, path) -> PacketHandler -> Recorder
  (ffmpeg over UDP loopback) -> Matroska .mkv, decryption handled internally.

Wired in recorder.ts (trackChannel/untrackChannel) + bootstrap.ts
(setVideoRecorderClient/setVideoRecordingsDir). All best-effort; failures log
and never break existing voice/audio. Unit tests 6/6 (single listener, watch
handshake + mkv path, skip own video, idempotence, teardown). Full suite
20 files / 176 tests green; typecheck + build + biome clean.

UNVERIFIED live yet: needs deploy + a streamer to confirm Recorder ready +
playable .mkv.
2026-08-30 23:26:59 +07:00
asepharyana 260ecabb3c debug(gateway): log non-opus RTP payload types on the voice socket
Temporary diagnostic to answer definitively whether Discord actually delivers
video RTP to the bot when someone screen-shares / turns on camera. Both
videoReceiver and screenShareAudio rely on ssrcMap emitting videoSSRC from
voice-state updates, and NO "Video SSRC appeared"/"Screen-share video started"
lines appear even after a real share+record. This logs any RTP packet whose
payload type is not Opus (120) so we can tell: (a) video RTP IS arriving but
attribution/signaling fails, vs (b) Discord sends no video at all to a
non-signaling receiver. Remove this log once the gap is understood.
2026-08-30 22:44:17 +07:00
asepharyana 5241f9784a feat(gateway): close video burst promptly when a user's SSRC is removed
ssrcMap emits "delete" when a VoiceUserData is removed (user stops sharing /
leaves voice). Hook it to close the user's open video burst ~500ms later so the
ffmpeg mux starts as soon as they stop, instead of waiting up to 5s for the idle
sweep. No-op if no burst exists; safe on normal teardown.
2026-08-30 21:51:14 +07:00
asepharyana 999c054bb2 feat(gateway): auto-mux recorded video to playable MP4 (Phase B)
Phase A captured raw .h264 streams but left them as non-playable elementary
streams. Phase B adds automatic muxing: when a video burst closes, the raw
.h264 is remuxed to a self-contained MP4 via `ffmpeg -c copy` (no re-encode,
fast) with `+faststart`, waits for the write stream to fully flush first so the
mux never reads a truncated tail, and deletes the raw .h264 on success (keeping
it on failure). Output: <RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.mp4.

muxToMp4 is exported + covered by a real-ffmpeg vitest (tests/videoReceiver.test.ts):
generates a tiny baseline h264, remuxes, asserts mp4 exists/non-empty & raw deleted
(also the 5 depacketizer tests). Full gateway suite 170/170 green, tsc + biome clean.
2026-08-30 21:49:40 +07:00
asepharyana 80e9b913c1 chore(gateway): drop unused H264 SINGLE_NAL constant 2026-08-30 21:21:07 +07:00
asepharyana 419946f38e feat(gateway): record others' video (camera/screen share) — Phase A capture
@discordjs/voice only decrypts/forwards AUDIO (opus) — its onUdpMessage drops
every non-opus RTP packet (dist/index.mjs:2068 `!== RTP_OPUS_PAYLOAD_TYPE`, 120).
Video RTP (H264 camera + screen share, plus VP8/VP9/AV1) arrives on the same UDP
socket but was silently discarded.

New videoReceiver.ts wraps receiver.onUdpMessage (like screenShareAudio.ts):
- detects video payload types (96/98/101/102/106/116/126/127),
- decrypts them with the connection secret key/encryptionMode via the SAME
  receiver.parsePacket path @discordjs/voice uses for audio (so DAVE + voice
  encryption are handled identically),
- depacketizes H264 to AnnexB (single NAL, STAP-A, and FU-A fragmentation),
  waiting for a keyframe (SPS/PPS/IDR) before writing,
- writes a raw .h264 elementary stream per user per burst under
  <RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.h264.

Attribution: a videoSSRC→user index is built from ssrcMap updates; a proximity
fallback mirrors screenShareAudio's inferScreenShareOwner. Bot's own video is
skipped.

Unit tests: tests/videoReceiver.test.ts (AnnexB start code, keyframe gating,
FU-A reassembly, orphan-fragment tolerance) — 5/5 green.

Phase A only (capture raw h264). Phase B (ffmpeg decode+mux to MP4/WebM +
persist) and Phase C (frontend playback) are follow-ups.
2026-08-30 21:20:32 +07:00