feat: Implement video capture improvements and mobile navigation fixes

- Added eager selfbot voice connection establishment to ensure video capture works seamlessly during voice channel joins.
- Introduced a manual video watch command for selfbots to allow operators to initiate screen recording of other members' streams.
- Enhanced video recording functionality to split recordings into segments based on user activity, similar to voice recordings.
- Fixed mobile navigation issues by extending the navbar to include all items and ensuring responsive form controls.
- Standardized error handling across frontend components to improve user experience during failures.
This commit is contained in:
asepharyana
2026-09-11 18:56:28 +07:00
parent 0eb089d9dc
commit b3418ad799
26 changed files with 0 additions and 0 deletions
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,101 @@
# GMW — Fitur Publik Lanjutan (#2–#6) Implementation Plan
> **For Hermes:** Implement task-by-task. Build + lint + typecheck each service
> after its changes. Deploy via push to main (CI handles Nix build + systemd).
> Hard constraint (user 2026-08-18): public read-only web, fully automatic,
> rules in code, NO admin endpoints, NO shadow mode, NO per-channel web config.
> **EXPLICITLY EXCLUDED: User Reputation / Strike History** (user: "hapus
> sepenuhnya fitur user reputation" — it was never built; do not add it).
## Existing infra to reuse (verified)
- **WS**: backend `ws/server.ts` broadcasts JSON `{type,data,timestamp}` to
frontendClients. Backend `ws/redis-bridge.ts` subscribes Redis channels
listed in `DISCORD_CHANNEL_TO_WS_EVENT` (backend `shared/redis-channels.ts`)
and re-emits as WS events. FE `src/lib/ws` auto-reconnect typed client.
- **Gateway → Redis**: `EventBroadcaster` + `RedisEventPublisher` (
`discord-gateway/src/modules/event-broadcaster`). Publish via
`eventBroadcaster.publish(EventChannels.X, payload)`.
- **Moderation data**: `moderation_actions` table (now has explainability
cols). `moderation.repository.listActions` returns rows. `ModerationAction`
FE type at `frontend/src/lib/types/moderation.ts`.
- **Messages**: `messages.list` / `getMessagesByChannel` (backend oRPC +
repository). FE `messagesApi` + `useMessages`.
- **Charts**: NO chart lib installed. Use **pure SVG/CSS** (consistent with
repo; avoid new deps).
- **CSV**: client-side Blob download, no backend.
## Task 1 — Live Moderation Feed (#2)
**Gateway**: add `MODERATION_ACTION: "discord:moderation:action"` to
`redis-channels.ts` (shared) + `EventChannels.MODERATION_ACTION` in
`eventTypes.ts`. In `moderationActionsDb.createModerationAction`, after insert,
publish `eventBroadcaster.publish(EventChannels.MODERATION_ACTION, actionRow)`.
**Backend**: add `DISCORD_MODERATION_ACTION` constant + map
`[DISCORD_MODERATION_ACTION]: "moderation_action"` in `DISCORD_CHANNEL_TO_WS_EVENT`.
**FE**: in `src/lib/ws`, subscribe to `moderation_action`; add `useLiveModeration`
hook (SWR-style with WS push, capped buffer ~50). Add `<LiveModerationFeed>`
client component on `/moderation` page (top of list, animated new-row).
Risk: gateway publish at every action (already async insert) — fire-and-forget,
wrap in try/catch. Verify WS event reaches FE via `wscat`/curl or log.
## Task 2 — Toxic Topic Trends (#3)
**Backend**: add `moderation.trends` oRPC. Query `moderation_actions` grouped
by `categories` (jsonb text[]) over last 30 days, count per category + severity
breakdown. Also `action_type` distribution. Return
`{ categories: {name,count}[], severities: {level,count}[], actions: {type,count}[] }`.
Map jsonb array in SQL (use `unnest` or parse in JS). Reuse `getDatabase`.
**FE**: `useModerationTrends` hook + `<TopicTrends>` SVG bar chart (top 10
categories) + severity donut (SVG arcs). Place on `/moderation` as a panel.
## Task 3 — Channel Timeline / Replay (#4)
Reuse existing `messages.list` (guildId) + `getMessagesByChannel`. Add a
**Timeline tab** to `/messages` that groups messages by date (client-side
bucket from `created_at`). Load-more via cursor. No new backend (existing
`messagesRouter.list` already supports guildId+limit+cursor). If needed, add
`messages.timeline` aggregation (count per day) — but keep simple: client
groups fetched rows. Verify existing endpoint returns enough history.
## Task 4 — Export CSV (#5)
**FE only**. `lib/csv.ts` `toCsv(rows, columns)` + `downloadCsv(filename, csv)`.
Add "Export CSV" button on `/moderation` (exports current actions) and
`/messages` (exports current list). Pure client-side, read-only. No backend.
## Task 5 — Activity Heatmap (#6)
**Backend**: add `messages.activity` oRPC: per-channel message count grouped by
hour-of-day (0–23) over last 14 days. Return
`{ channels: {channelId, name, byHour: number[24]}[], max }`. Use SQL
`EXTRACT(hour from ...)` + group by channel. Channel name from
`message.metadata->'channel'->>'channelName'`.
**FE**: `useMessageActivity` hook + `<ActivityHeatmap>` SVG grid (channels ×
24h, color intensity = count/max). Place on `/messages` or `/dashboard`.
## Verification checklist
- [ ] `pnpm typecheck && pnpm lint && pnpm build` green for gateway, backend, frontend
- [ ] Backend `/trpc/moderation/trends` returns categories/severities/actions
- [ ] Backend `/trpc/messages/activity` returns byHour grids
- [ ] WS `moderation_action` received by FE (log or visible live row)
- [ ] No admin/write endpoint added; all public read-only
- [ ] No User Reputation code anywhere (grep "reputation|strike|reputasi")
- [ ] Deploy via push; all 3 services `running`; moderation + messages pages load
## Files touched (summary)
- gateway: `shared/redis-channels.ts`, `event-broadcaster/eventTypes.ts`,
`event-broadcaster/eventBroadcaster.ts`, `message-capture/moderationActionsDb.ts`
- backend: `shared/redis-channels.ts`, `orpc/router.ts`,
`modules/moderation/moderation.service.ts` (+repository),
`modules/messages/messages.service.ts` (+repository, +schema)
- frontend: `lib/ws/*`, `hooks/use-moderation.ts`, `hooks/use-messages.ts`,
`lib/csv.ts`, `lib/types/*`, `app/(dashboard)/moderation/view.tsx`,
`app/(dashboard)/messages/view.tsx`, new components under `components/`
## Status: COMPLETE (deployed + verified)
- Commit 9b3134d: features #2–#6 (live feed, trends, timeline, CSV export, heatmap)
- Commit 2a8f6d9: user reputation feature fully removed (643 deletions, no trace in src/tests)
- Migration 0016 applied: user_reputations DROPPED (DB verified: false)
- All 3 services active (gateway + backend restarted 18:29, frontend running)
- Gateway typecheck/lint/test(117 passed); backend typecheck/lint/build; FE lint/build — all GREEN
## Verification
- moderation/stats WS returns data (32 actions) → WS adapter works
- DB: user_reputations gone; moderation_actions explainability cols present
- Live Feed: gateway publishes discord:moderation:action → backend WS (same path as guild_member_*)
- Trends/Activity: backend router procedures registered (typecheck+tsc), same WS adapter
@@ -0,0 +1,119 @@
# GMW — Fitur Publik Lanjutan #7–#15 + Bug Fix Reputation Removal
> **For Hermes:** Implement task-by-task. Build + lint + typecheck each service after its
> changes. Deploy via push to main (CI handles Nix build + systemd). Apply any new
> drizzle migration MANUALLY (systemd does NOT run migrations).
> Hard constraint (user): public read-only web, fully automatic, rules in code,
> NO admin endpoints, NO shadow mode, NO per-user reputation aggregation.
## Bug fix discovered during planning (MUST do first)
`services/backend/src/modules/dashboard/dashboard.repository.ts` still references
`pgUserReputationsTable` (import line 8; JOINs at lines 173 + 457) — that table was
DROPPED in migration `0016`. `dashboard.listUsers` / `dashboard.userDetail` will
**crash at runtime** (undefined table). Remove the import + the `r.*` join columns
(`trust_score`, `clean_message_streak`, `total_infractions`) from both queries.
This is a regression introduced by the reputation removal commit.
## Features to implement (#7–#15)
All reuse existing infra: `moderation_actions`, `messages`, `channel_cultures`,
`term_glossary_cache`, `ai_analysis_runs`, `message_edits`, gateway cron (for #15),
WS (proven Live Feed pattern), oRPC over WS (proven), pure-SVG charts (no libs).
| # | Feature | Data source | Surface |
|---|---------|-------------|---------|
| 7 | Flagged Link / Scam Domain Reporter | regex URL from `moderation_actions.content`/`evidence` | `/moderation` |
| 8 | Top Flagged Channels | join `moderation_actions.message_id`→`messages.channel_id` | `/moderation` |
| 9 | Moderation Heatmap by Hour | `moderation_actions.created_at` hour-of-day | `/moderation` |
| 10 | Flag Category Drill-down | `moderation_actions.categories` (reuse Trends) | `/moderation` FE-only |
| 11 | Channel Culture Glossary | `channel_cultures` (exists) | new `/channels` panel |
| 12 | Term Knowledge Base | `term_glossary_cache` (exists) | new `/glossary` panel |
| 13 | Edit/Evasion Tracker | `message_edits` (exists) | `/messages` |
| 14 | Auto-mod Coverage Stats | `ai_analysis_runs` (exists) | `/moderation` metric tiles |
| 15 | Weekly Digest (auto, cron) | aggregate #7/#8/#9 → Discord via gateway cron | gateway cron + `/moderation` |
## Architecture per layer
### Backend (oRPC, `services/backend/src`)
- New repository methods (add to existing repos, follow `getTrends` SQL style):
- `moderation.repository.ts`:
- `getTopFlaggedDomains(days)` — `regexp_matches(content,'https?://([^/\s]+)')` on
`moderation_actions WHERE created_at>=since`, group by host, COUNT, order DESC LIMIT 20.
- `getTopFlaggedChannels(days)` — join `moderation_actions a` LEFT JOIN `messages m`
ON `m.id=a.message_id`, group by `m.channel_id`, COUNT, order DESC LIMIT 15.
Channel name via `m.metadata::jsonb->'channel'->>'channelName'`.
- `getHourlyModeration(days)` — `EXTRACT(HOUR FROM to_timestamp(created_at/1000))`
group by hour, COUNT, severity breakdown. (24 rows)
- `getFlaggedByCategory(days, category)` — list actions where `categories` contains
`category` (reuse `listActions` filter or new query), for drill-down #10.
- `getCoverage(days)` — from `ai_analysis_runs`: total runs, status breakdown
(clean/flagged/warn/error/pending), coverage % = (analyzed)/(captured in window).
- `dashboard.repository.ts` (or new `knowledge.repository.ts`):
- `listChannelCultures(limit, search?)` — `channel_cultures` rows (channel_id,
guild_id, channel_name from messages metadata, culture_summary, last_analyzed_at).
- `listGlossary(limit, search?)` — `term_glossary_cache` (term, definition, source_url,
resolved_at, hit_count) order by hit_count DESC.
- `messages.repository.ts`:
- `getEditHistory(limit, channelId?)` — `message_edits` join `messages` for
old_content + channel + username + edited_at, order DESC LIMIT.
- `moderation.service.ts` / `dashboard.service.ts` / `messages.service.ts`: thin wrappers.
- `orpc/router.ts`: add procedures (follow `trends` shape):
- `moderation.topDomains`, `moderation.topChannels`, `moderation.byHour`,
`moderation.byCategory` (input `{days,category}`), `moderation.coverage`.
- `dashboard.channelCultures`, `dashboard.glossary`.
- `messages.editHistory`.
### Frontend (`services/frontend/src`)
- `lib/types/moderation.ts`: add `FlaggedDomain`, `FlaggedChannel`, `HourlyModeration`,
`ModerationCoverage` interfaces.
- `lib/types/index.ts` (+ message.ts): add `ChannelCultureRow`, `GlossaryRow`, `EditHistoryRow`.
- `lib/api/moderation.ts`: add `topDomains`, `topChannels`, `byHour`, `byCategory`, `coverage`.
- `lib/api/dashboard.ts` (or messages.ts): add `channelCultures`, `glossary`, `editHistory`.
- `lib/api/server.ts`: add SSR seed fetchers (follow `getModerationStats`).
- `hooks/use-moderation.ts`: add `useTopDomains`, `useTopChannels`, `useHourlyModeration`,
`useByCategory`, `useCoverage`. `hooks/use-dashboard.ts`/`use-messages.ts`: add culture/glossary/edit hooks. `hooks/index.ts`: export all.
- New components (pure SVG/CSS, reuse `GlassPanel`/`SectionHeader`/`Badge`/`Donut`):
- `components/ScamDomains.tsx`, `components/TopChannels.tsx`, `components/ModerationHeatmap.tsx`,
`components/CoverageTiles.tsx`, `components/ChannelCultureGlossary.tsx`,
`components/TermGlossary.tsx`, `components/EditHistory.tsx`.
- Wire into `app/(dashboard)/moderation/view.tsx` (grid col-span-2/3/5 as space allows)
and `app/(dashboard)/messages/view.tsx` (EditHistory panel) and new route pages
`app/(dashboard)/channels/page.tsx` + `app/(dashboard)/glossary/page.tsx` with
matching `view.tsx` (follow existing page→view SSR pattern; check `app/(dashboard)/dashboard/page.tsx`).
- Export CSV buttons reuse `lib/csv.ts` `downloadCsv` (client-side) for domains/channels/edits.
### Gateway (#15 Weekly Digest)
- Add a cron/interval in `services/discord-gateway` (check existing scheduler pattern —
search `setInterval`/`cron` in `src`). On a 7-day cadence, query backend oRPC
(`dashboard.activity`, `moderation.trends`, `moderation.topChannels`) — OR compute
directly via a shared repository — and post a formatted summary to the monitor guild
channel (via existing `discordClient.channels.send` helper). Fully automatic, no UI.
## Files touched (summary)
- backend: `modules/moderation/{repository,service}.ts`, `modules/dashboard/{repository,service}.ts`,
`modules/messages/{repository,service}.ts`, `orpc/router.ts`, `shared/index.ts` (if new tables),
`lib/types/*` (FE)
- frontend: `lib/api/*`, `lib/types/*`, `hooks/*`, `components/*`, `app/(dashboard)/*`
- gateway: new digest scheduler + (none if reuse backend) maybe `shared/redis-channels.ts`
## Constraints / pitfalls (from gmw-ops skill)
- `created_at` is bigint epoch-MS — compare with `<`/`>`, do NOT divide by 1000 in SQL.
- Pure SVG only — frontend has ZERO chart libs.
- `Badge` Tone = signal|amber|vermilion|neutral (no "rose").
- Frontend WS import is `@/lib/ws/context`; method `on` not `subscribe`.
- Commit author `asepharyana`, no Co-Authored-By.
- Rebuild `dist/` after gateway changes; apply drizzle migrations manually.
## Verification
- Per service: `pnpm typecheck && pnpm lint && pnpm build` green.
- Gateway: `pnpm test` (117+ pass).
- Live: `moderation/stats` WS returns data (proves adapter); new procedures registered
(typecheck = proof). `systemctl show` new ActiveEnterTimestamp after deploy.
- DB: confirm `channel_cultures`/`term_glossary_cache`/`message_edits`/`ai_analysis_runs`
have rows before relying on them (some may be empty → components handle empty state).
## Execution order
1. Bug fix dashboard.repository (reputation JOIN) — deploy-safe.
2. Backend repositories + service + router (#7,#8,#9,#14 dashboard; #11,#12; #13).
3. FE types + api + hooks + components + wire (#7,#8,#9,#10,#11,#12,#13,#14).
4. Gateway #15 digest (if scheduler exists) — verify via log, not UI.
5. Build/lint all 3 services; commit; push; monitor CI; apply migrations; verify live.
@@ -0,0 +1,111 @@
# AI Analysis Flow — Audit & Optimization (discord-gateway)
**Goal:** Analisis alur AI analysis end-to-end, temukan bug/inconsistency yang merusak kualitas verdict, lalu perbaiki root cause-nya.
## Scope
- `services/discord-gateway/src/modules/ai-moderation/**`
- Tidak menyentuh chatbot backend / frontend.
## Alur saat ini (hasil tracing)
```
message capture → aiAnalyzer.queueMessageAnalysis(messageId)
→ batchScheduler.scheduleConversationAnalysis(conversationKey) [debounce 250ms, CB gate]
→ messageStore.getPendingMessagesByConversation(≤200)
→ skipAgeRestrictedMessages
→ pickBatchWithinBudget(14000 tokens, 50/msg)
→ processBatch [Piscina worker, ≤4 threads]
→ ai-analysis-worker.processBatch
→ getConversationContextBefore(20 msgs) + attachments
→ attachment-upload race guard (pending upload → skip)
→ runModerationAnalysis
→ Phase 1: exact-hash cache (PG text_analysis_cache, per channel/thread)
→ Phase 2: semantic cache (embedTexts → Qdrant batch search; PG fallback)
→ split text-only vs media
→ runTextOnlyBatch: URL fetch + wiki search + glossary (paralel)
→ dedup short messages → sub-batches (60/sub-batch)
→ vision evidence utk URL images (hoisted, 15s cap per image)
→ callModerationLLM per sub-batch (stream:true, retries 3, JSON parse + correction retry)
→ runMediaBatch: download → vision per image (cache LRU→DB→live, lock) → 1 LLM call
→ setCachedTextModeration (PG + Qdrant upsert w/ embedding)
→ normalizeResult (confidence clamp, fallback analysis)
→ updateMessagesAIAnalysisBulk → broadcast + scheduleAutoDelete
→ recovery worker tiap 10s: pending keys → re-schedule; incomplete → individual fallback queue
→ individual fallback: 1 msg = 1 worker job (context + full LLM)
→ cache prune tiap 6 jam (PG expired + Qdrant expired points)
```
## Temuan audit (ranked)
### F1 — Cache hit menghapus status "warn" (BUG AKURASI)
`moderationOrchestrator.ts` Phase-2 semantic hit & PG-fallback memetakan status via
`parseQdrantVerdict`: storedStatus bukan "warn"/"flagged" → dipaksa "clean".
TAPI exact-hash lookup (`getCachedTextModeration`, textCacheStore.ts:288-295) lebih parah:
hanya menerima "clean"|"flagged" — **"warn" jatuh ke branch flags.length===0 ? clean : flagged**
→ warn dengan flags=["conflict_instigation"] dibaca sebagai FLAGGED.
Efek: auto-delete eligibility (butuh recommendedAction delete/escalate + severity list) salah baca;
dashboard menampilkan flagged padahal verdict asli warn. Root cause: type narrowing legacy
(`status: "clean" | "flagged"`) tidak diupdate ketika "warn" ditambahkan ke schema.
### F2 — Exact-cache key mengabaikan edit (BUG EVASION)
Key = sha256(content)+context. Pesan yang DIEDIT (`edited_content`) menghasilkan hash berbeda,
tapi verdict lama utk konten pre-edit tetap hidup; lebih penting: pesan edited="true" adalah sinyal
evasion di prompt, sedangkan cache bisa menyajikan verdict dari konten lama jika content sama.
(Minor, tapi konsistensi: `resolveIsEdited` ada di prompt, tidak ada di cache key.)
### F3 — `pickBatchWithinBudget` skip-bukan-break (LATENSI/KUALITAS)
Loop `if (usedTokens + msgTokens <= maxTokens) {push}` — pesan BESAR di tengah list dilewati
dan iterasi lanjut mencoba msg berikutnya. Efek: batch berisi "lubang" (msg pending tetap pending,
dianalisis di gelombang berikutnya = LLM call tambahan). Ini by-design tolerable, tapi ada bug halus:
pesan >budget tunggal tidak pernah masuk (scheduler sudah punya fallback slice(0,1), OK).
Keputusan: biarkan (bukan bug nyata), catat saja.
### F4 — `callModerationLLM` max_tokens 16384 hardcoded (COST)
Sub-batch 60 pesan × output ~150 token/pesan ≈ 9k token cukup; 16k aman. Biarkan.
### F5 — Dead code builder user-profile/reputation
`buildUserProfilesBlock`, `buildUserProfileRef`, `UserProfileEntry` di moderationBuilders.ts
tidak dipakai lagi sejak context minimization (hanya tests). `<user_history>` juga tak pernah
di-inject (rules masih menyebutnya — misleading bagi model). Bersihkan referensi prompt.
### F6 — rules.ts menyebut `<user_history>` yang tidak pernah ada di payload
Model diberi instruksi tentang blok yang tak pernah muncul → pemborosan token + potensi
kelakuan aneh ("menunggu" data yang tak ada). Hapus/ubah kalimat.
### F7 — system.ts "Blok Data" menyebut `<term_glossary> (SearXNG)` — STALE
Sumber sudah Wikipedia. Komentar kode & teks prompt menyebut SearXNG. Perbaiki teks (kecil).
### F8 — output.ts typo "secifik", baris tabel `-|-` rusak
Kualitas prompt: typo + markdown table broken (`||-`) di beberapa baris. Rapikan.
### F9 — llmCaller parse-error correction tail hanya di SYSTEM
Correction tail ditambahkan ke system prompt; provider caching fine, tapi preview invalid
content (800 char) ikut SYSTEM — ok. Skip.
### F10 — `getLlmSemaphore` race kecil saat config berubah di tengah flight
Non-issue praktis (config statis per proses). Skip.
## Keputusan perbaikan (yang dieksekusi sekarang)
1. **F1 (utama):** normalisasi status di SATU tempat — `normalizeStoredStatus()` di
textCacheStore.ts yang menerima clean/warn/flagged; pakai di getCachedTextModeration
DAN parseQdrantVerdict; perluas return types ke union penuh. Orchestrator tinggal pakai.
2. **F6+F7+F8:** bersihkan stale references di prompts (user_history, SearXNG, typo).
3. **F5:** hapus dead builders + test-nya (biome/tsc yang jaga).
4. Regression test untuk F1 (vitest): warn tersimpan → warn terbaca (exact + qdrant path).
## Files touched
- services/discord-gateway/src/modules/ai-moderation/textCacheStore.ts (F1)
- services/discord-gateway/src/modules/ai-moderation/moderationOrchestrator.ts (type only)
- services/discord-gateway/src/modules/ai-moderation/prompts/rules.ts (F6)
- services/discord-gateway/src/modules/ai-moderation/prompts/system.ts (F7)
- services/discord-gateway/src/modules/ai-moderation/prompts/output.ts (F8)
- services/discord-gateway/src/modules/ai-moderation/moderationBuilders.ts (F5)
- services/discord-gateway/tests/contextEnrichment.test.ts (F5 test cleanup + F1 regression test baru)
## Verification
```
cd services/discord-gateway
npx tsc --noEmit
npx biome check --diagnostic-level=error .
npx vitest run
```
Semua harus hijau sebelum commit. Deploy via GHA (push main) — user konfirmasi belakangan.
@@ -0,0 +1,33 @@
# Optimisasi "non-issue" AI analysis pipeline
## Scope
Dua item yang sebelumnya dinyatakan non-issue, kini dioptimalkan + 1 bug ordering
yang ditemukan saat menelusuri:
1. **pickBatchWithinBudget: skip → break.** Pesan diurutkan `created_at ASC`
oleh DB. Setelah budget habis, pesan berikutnya pasti lebih besar/lebih kecil
arbitrer — skip-then-take menghasilkan batch non-kontigu (ada gap analisis
di tengah timeline). Ubah jadi stop at first overflow (break) supaya prefix
kronologis utuh; sisanya otomatis diambil gelombang berikutnya
(`shouldScheduleNext` sudah selalu true setelah sukses).
2. **max_tokens dinamis.** Hard-coded 16384 di llmCaller.ts → parameter
opsional `maxTokens?`; default tetap 16384. Caller text/media batch pass
nilai berbasis ukuran prompt (tiktoken) dengan floor/ceiling.
3. **Bug ordering UPDATE..RETURNING (bonus).** messagesAnalysis.ts
`getPendingMessagesByConversation`: SELECT ids di-order `created_at ASC`
tapi UPDATE...RETURNING tanpa ORDER BY → urutan rows balik tidak
terjamin. Konsumen pakai messages[0] sebagai anchor konteks
(beforeCreatedAt) dan pickBatchWithinBudget asumsi urutan. Fix: re-sort in
JS by created_at (stable) sebelum return.
## Files touched
- src/modules/ai-moderation/batchProcessor.ts — break bukan skip; test baru.
- src/modules/ai-moderation/llmCaller.ts — param maxTokens.
- src/modules/ai-moderation/textBatchProcessor.ts / mediaBatchProcessor.ts —
hitung token prompt & pass maxTokens.
- src/modules/message-capture/messagesAnalysis.ts — sort hasil RETURNING.
- tests/batchBudget.test.ts — baru.
## Verification
cd services/discord-gateway && bun run typecheck && bun run lint && bun run test
lalu commit+push, watch GHA, restart service via deploy pipeline.
@@ -0,0 +1,67 @@
# Spec: Perbagus fitur Voice + Audio Playback (GMW frontend)
Tanggal: 2026-08-22 · Scope: **frontend only** (backend/gateway API sudah cukup)
## Masalah (audit)
1. Recordings: semua kartu pakai `<audio controls>` native — tampilan identik,
tidak ada indikasi which-clip-playing / loading / paused, dan N audio bisa
play bareng (overlap).
2. Media view: `thumbnailUrl` dari gateway tidak dipakai; tidak ada visual
"sedang playing" selain disc spin; queue item semua sama tanpa badge up-next.
3. Mini-player (`lib/hooks/use-media-player.tsx`) ada tapi TIDAK PERNAH
dimount → dead code, user tidak lihat status musik di halaman lain.
4. Voice page: `useMicTransmit.setVolume` + `useVoiceListen.setVolume`
tersedia tapi tak ada UI-nya; mic live tidak punya level feedback.
## Desain
### A. RecordingAudioPlayer (baru, `components/voice/recording-audio-player.tsx`)
Custom player menggantikan `<audio controls>`:
- Play/pause button (ikon berubah), spinner saat buffering (`waiting` event).
- Progress bar seekable (click-to-seek) + time label `m:ss / m:ss`.
- Waveform-ish equalizer bars saat playing (CSS animation, reduced-motion safe).
- **Single-playback**: module-level registry `activePlayers` — memainkan satu
clip otomatis pause yang lain.
- Kartu pemilik player aktif dapat highlight border signal + "Now playing" chip.
### B. Recordings view — pasang player baru
- Ganti `<audio>` → `<RecordingAudioPlayer src download_url>`.
- Highlight kartu via state lifted: `playingId` di view, callback `onPlay`.
### C. Media view polish
- Hero: thumbnail (jika `current.thumbnailUrl`) sebagai disc center image;
fallback ListMusic icon. Equalizer bars animasi CSS saat `playing`.
- Queue row pertama: badge "up next"; baris current track diberi ring signal.
- Volume read-only tetap.
### D. MiniPlayer global
- Hapus `lib/hooks/use-media-player.tsx` (dead) — ganti dengan komponen
`components/media/mini-player.tsx` yang subscribe `useMediaState` +
`useMediaWsSync` langsung (SWR cache shared antar route), mounted di
`AppFrame` bawah layar (fixed bottom, hidden di route `/media`).
- Menampilkan: thumbnail kecil/judul, tombol skip/stop, link ke /media.
### E. Voice UI
- Mic live: level meter (Equalizer bars) — mic-transmitter sudah punya worklet;
tambah `getLevel()` via AnalyserNode pada stream (simple RMS) di hook.
- Listen: volume slider (input range) wired ke `listen.setVolume`.
- Mic volume slider wired ke `mic.setVolume`.
## File touched
| File | Aksi |
|---|---|
| services/frontend/src/components/voice/recording-audio-player.tsx | new |
| services/frontend/src/app/(dashboard)/recordings/view.tsx | edit |
| services/frontend/src/app/(dashboard)/media/view.tsx | edit |
| services/frontend/src/components/media/mini-player.tsx | new |
| services/frontend/src/components/shell/ambient-app.tsx | mount MiniPlayer |
| services/frontend/src/lib/hooks/use-media-player.tsx | delete |
| services/frontend/src/hooks/use-voice.ts | tambah micLevel |
| services/frontend/src/lib/audio/mic-transmit.ts | expose analyser level |
| services/frontend/src/app/(dashboard)/voice/view.tsx | sliders + meter |
## Verifikasi
1. `pnpm lint` (biome) + `pnpm build` clean.
2. Smoke di port **4024** (BUKAN 4017) → curl 200 semua route.
3. Commit (tanpa trailer) → push → `gh run watch` → live check
https://imphnen.asepharyana.my.id/{media,recordings,voice}/ = 200.
@@ -0,0 +1,127 @@
# Spec: Optimasi AI Analysis GMW — Naikkan Cache Hit Tanpa Kehilangan Akurasi
Tanggal: 2026-08-24 · Repo: `~/GMW` (branch `main`) · Service: `services/discord-gateway`
## Latar & Evidence (audit 2026-08-24)
State produksi:
- Qdrant `gmw_text_moderation`: **1.550 poin, status green** (vectors size 2048, Cosine).
- PG `text_analysis_cache`: 1.634 row `user_moderation`, 277 `vision_llm`; **sum(hit_count) = 0** →
hit-rate tidak pernah terukur.
- Embedding aktif (`AI_LLM_EMBEDDING_MODEL` set, Nemotron-embed, dim 2048), `AI_LLM_EMBEDDING_MIN_SIMILARITY`
tidak diset di BWS → default **0.97** (sangat konservatif).
- Messages: 9.375 total; 643 status `error` (banyak retry), 49 pending.
Temuan audit alur (`moderationOrchestrator.ts` → `textCacheStore.ts` → `qdrantClient.ts`,
`textBatchProcessor.ts`, `urlFetcher.ts`, `wikipediaClient.ts`, `visionAnalyzer.ts`):
| # | Temuan | Dampak |
|---|--------|--------|
| F1 | Exact-hash cache key menyertakan context (channel/thread) → teks sama di channel lain selalu miss | Killer hit-rate #1 |
| F2 | Semantic tier TIDAK memfilter context (Qdrant payload tak punya context) — sudah global tapi hanya aman krn sim 0.97 ketat | Inkonsisten dgn exact tier |
| F3 | Phase-1 lookup loop `await getCachedTextModeration(key)` per pesan → N round-trip PgBouncer per batch (60 msg = 60 query serial) | Latensi + beban DB |
| F4 | Verdict actionable (flagged/warn) dan clean sama-sama boleh di-serve semantic; toleransi akurasi beda | Risiko akurasi |
| F5 | `hit_count` tidak pernah di-increment oleh reader manapun | Hit-rate tak terukur |
| F6 | `wikipediaSearch()` (blok `<web_searches>`) tanpa cache — re-fetch tiap batch utk query sama | Latensi + spam ke WP |
| F7 | `fetchUrlSafely()` tanpa cache — link sama di batch berikutnya di-download lagi penuh | Latensi + bandwidth |
| F8 | Vision cache key dari data-URL base64 hasil resize → attachment sama via jalur berbeda (URL vs embed) = key beda → re-download + re-vision | Duplikasi kerja vision |
Non-goals: mengubah pipeline enforcement (auto-mute/ban trust-store writes), mengubah prompt
kebijakan moderasi, mengubah model/embedding provider.
## Desain
Semua perubahan degrade gracefully — cache gagal → perilaku lama (LLM). Akurasi dilindungi
asimetris: **hemat boleh untuk verdict non-actionable, konservatif untuk yang memicu aksi.**
### D1 — Cache metrics (F5)
- `textCacheStore.getCachedTextModeration()`: saat hit valid, increment `hit_count`
(`UPDATE ... SET hit_count = hit_count + 1`) fire-and-forget (`.catch(()=>{})`), jangan blokir return.
- Log info periodik ringkas di orchestrator sudah ada ("User moderation cache applied") — cukup.
### D2 — Batched exact-cache lookup (F3)
- Fungsi baru `getCachedTextModerations(keys: string[]): Promise<Map<string, StoredModerationVerdict>>`
di `textCacheStore.ts`: **satu** `SELECT ... WHERE text = ANY($1)` (chunk 200 key/query),
parse + `normalizeStoredStatus` per row (reuse helper existing).
- Orchestrator fase-1: kumpulkan semua key unik → satu call batched → distribusi hasil.
- Semantik identik dengan loop lama (row expired/error-artifact tetap miss); hanya jumlah round-trip
yang turun N→1.
### D3 — Global exact reuse untuk verdict non-actionable (F1)
- Key scoped-context TETAP ditulis (kompatibel, invalidasi moderator tetap presisi).
- Reader tambahan: kalau key `<ctx>:<hash>` miss, coba key legacy global `text_mod:<hash>` (bare).
- Guard akurasi (WAJIB semua terpenuhi):
- `status === "clean"` DAN `flags.length === 0`;
- `confidence >= AI_CACHE_GLOBAL_REUSE_MIN_CONFIDENCE` (default 0.85);
- `recommendedAction === "none"`;
- umur entry ≤ `AI_CACHE_GLOBAL_REUSE_MAX_AGE_H` (default 72h) — cek `analyzed_at`.
- Flag baru `policyVersion: "cached-global-clean-2026-08"` supaya terlacak di dashboard/log.
- Verdict flagged/warn TETAP context-scoped (tidak pernah lintas channel).
### D4 — Semantic dua-band similarity (F2+F4)
- Config baru: `AI_LLM_EMBEDDING_MIN_SIMILARITY_ACTIONABLE` default **0.97** (perilaku lama),
`AI_LLM_EMBEDDING_MIN_SIMILARITY_CLEAN` default **0.92**, keduanya coerce number 0..1.
- Satu Qdrant batch search pakai threshold RENDAH (0.92). Per hit, klasifikasi ulang:
- verdict non-actionable (clean, no flags, action=none): terima jika `score >= CLEAN_BAND`;
- verdict actionable (warn/flagged atau flags ada / action != none): terima hanya jika
`score >= ACTIONABLE_BAND` (0.97 — persis gate lama);
- di antara dua band → buang hit, pesan lanjut ke LLM (fail-open ke akurasi).
- Legacy PG fallback path: filter serupa di `findSimilarTextModeration` via parameter band.
### D5 — Cache Wikipedia search (F6)
- `wikipediaClient.wikipediaSearch(query)`: cek `cacheGet(makeCacheKey("wikisearch", q))` dulu;
miss → fetch (timeout existing) → sukses & hasil non-kosong → `cacheSet(..., TTL 6h)`.
Hasil kosong TIDAK di-cache (biar retry nanti). Redis down → langsung fetch (no-op cache).
### D6 — Cache URL text fetch (F7)
- `urlFetcher.fetchUrlSafely(url)`: wrapper async memoize in-process LRU (max 500, TTL 30 menit)
untuk `type === "text"` saja (image tetap selalu fresh-download karena dipakai sbg bukti vision
+ buffer besar; error tidak di-cache).
- Import `LRUCache` dari `lru-cache` (sudah dep gateway).
### D7 — Unified vision cache key (F8)
- `makeImageCacheKey(imageUrl)` di `textCacheStore.ts`: sebelum hash, strip query Discord CDN
(`?ex=&is=&hm=` signed tokens, `format/width/height/size`) — regex `(\?[^#]*)$` dibuang bila host
CDN discord (`cdn.discordapp.com`, `media.discordapp.net`, `images-ext-*.discordapp.net`);
URL non-Discord: hash full URL seperti sekarang.
- Efek: attachment sama yang lolos lewat jalur embed vs inline vs re-fetch dgn token beda → SATU
entry cache → skip download+vision kedua kali. Data-URL base64 tetap di-hash apa adanya.
## File yang disentuh
1. `src/shared/config/index.ts` — 3 config baru (D3×2, D4×2 — total 4 nilai, 3 baris zod + deskripsi).
2. `src/modules/ai-moderation/textCacheStore.ts` — hit_count inc (D1), batched getter (D2),
global-reuse guard helper (D3), image-key normalize (D7).
3. `src/modules/ai-moderation/moderationOrchestrator.ts` — pakai batched getter (D2),
global bare-key fallback (D3), dua-band semantic accept (D4).
4. `src/modules/ai-moderation/qdrantClient.ts` — `searchQdrantBatch` menerima threshold rendah
(sudah parametrik — mungkin tanpa perubahan; verifikasi).
5. `src/modules/ai-moderation/wikipediaClient.ts` — cache layer (D5).
6. `src/modules/ai-moderation/urlFetcher.ts` — LRU text-fetch memoize (D6).
## Schema/type changes
- Tidak ada migrasi DB (kolom `hit_count`, `analyzed_at`, `expires_at` sudah ada).
- Tidak ada perubahan kontrak WS/oRPC/frontend.
- Type baru: none public; internal `StoredModerationVerdict` dipakai ulang.
## Verification
1. Unit tests baru (`tests/`):
- `cacheBatchLookup.test.ts`: batched getter — hit/miss/expired/error-artifact mapping,
chunking >200 keys (mock executeAll), hit_count increment called.
- `globalReuseGuard.test.ts`: guard menerima clean+conf≥0.85+action none+umur ≤72h;
menolak flagged/warn/conf rendah/action≠none/stale.
- `semanticBands.test.ts`: clean @0.93 diterima, flagged @0.93 ditolak, flagged @0.98 diterima.
- `imageKeyNormalize.test.ts`: URL Discord dgn/ex token → key sama; non-Discord beda query → beda.
2. Gate service: `pnpm typecheck && pnpm exec biome check --diagnostic-level=error . && pnpm exec vitest run`.
3. Deploy via GHA (`git push origin main`) → watch `Build & Deploy (Nix)` → verifikasi
`systemctl show gmw-discord-gateway -p ActiveEnterTimestamp` baru.
4. Runtime probe pasca-deploy: journalctl level 30 normal; beberapa jam kemudian
`SELECT sum(hit_count) FROM text_analysis_cache WHERE source='user_moderation'` > 0 membuktikan
metrics jalan; log "User moderation cache applied" menunjukkan hits>0 pada traffic ramai.
## Rollback
Semua fitur behind config defaults yang mempertahankan perilaku lama pada nilai konservatif;
rollback = redeploy commit sebelumnya (tanpa migrasi DB, tanpa state eksternal).
@@ -0,0 +1,56 @@
# Spec: Perbaiki Delay Attachment 162s→<20s (GMW AI Analysis)
Tanggal: 2026-08-24 · Repo `~/GMW` · Service discord-gateway
## Evidence (audit produksi)
Klaster pesan attachment delay ~330–400 detik. Trace pesan `1541417073245290638` (.gif):
19:01:08 dibuat → 19:01:09 batch incomplete → fan-out individual → **guard upload-pending
mengembalikan `results:[]`** → diperalakukan sukses (`complete ... (undefined)`) → row
tertahan `ai_status='processing'` **tanpa penanggung jawab** → 19:06:12 cleanup mengembalikan
ke `pending` (tepat 300s) → baru dianalisis. Plus vision gagal 3× utk GIF besar
("Stream ended before producing a non-ping SSE event") → degradasi teks.
## Root causes
- **A (fatal)**: `individualFallbackProcessor.processIndividualFallback` memperlakukan
`ok:true + results:[]` sebagai sukses. Race-guard upload di `ai-analysis-worker.processIndividual`
sengaja balik `results:[]` (desain lama) → pesan yatim `processing` sampai cleanup 300s.
- **B**: `llmVision` hanya mencoba `stream:true`; kegagalan SSE truncation pada gambar besar
= 3 retry sia-sia (semua jalur sama) → bukti media hilang.
- **C**: safety-net cleanup 300s terlalu lambat sbg satu-satunya pemulih `processing`.
## Fix
1. **F1 — sinyal eksplisit upload-pending**: `IndividualOkResponse` + field opsional
`uploadPending?: boolean`. Worker set `uploadPending:true` saat race guard kena.
2. **F2 — processor menangani 3 kondisi** via helper murni baru
`classifyIndividualWorkerResult(result): "success" | "upload_pending" | "incomplete" | "error"`
(modul baru `fallbackResultClassifier.ts`, zero-dep agar mudah dites):
- `upload_pending` → tulis ulang row ke `pending` (pola sama dgn revert apiFailed di
batchProcessor) + broadcast + **re-schedule analisis percakapan segera**
(dynamic import batchScheduler, pola anti-siklus yg sudah ada) → retry dalam ~250ms
begitu upload beres. Bukan error, tidak naikkan CB counter.
- `incomplete` (flags analysis_incomplete) → perilaku lama (exhausted path).
- `error` / `results kosong tanpa penjelasan` → throw transien (retry oleh recovery),
BUKAN sukses palsu. Log "(undefined)" hilang.
3. **F3 — vision non-stream fallback**: di `llmVision`, jika error match
`/Stream ended before producing a non-ping SSE|stream ended/i` → coba SEKALI lagi dengan
`stream:false` (router agregasi penuh; timeout tetap 60s). Konversi hard-fail jadi sukses.
4. **F4 — turunkan safety net**: default `revertStuckProcessingMessages` 300000 → 120000 ms.
## File disentuh
- `src/modules/ai-moderation/fallbackResultClassifier.ts` (BARU, pure)
- `src/modules/ai-moderation/ai-analysis-worker.ts` (tipe + set flag uploadPending)
- `src/modules/ai-moderation/individualFallbackProcessor.ts` (konsumsi classifier + reschedule)
- `src/modules/ai-moderation/llmClient.ts` (fallback non-stream di llmVision)
- `src/modules/message-capture/messagesCleanup.ts` (default 120s)
## Verifikasi
- Test baru `tests/fallbackResultClassifier.test.ts` (4 klasifikasi + edge kosong).
- Gate: tsc --noEmit, biome error-level, vitest run semua hijau.
- Deploy GHA sukses; pasca-deploy: pesan attachment baru p50 < 20s
(`SELECT percentile_cont(0.5) ... WHERE metadata attachments>0 AND created_at > deploy`),
tidak ada lagi "complete ... (undefined)".
File diff suppressed because one or more lines are too long
@@ -0,0 +1,39 @@
# GMW FE — Monokrom Hitam-Putih + Sidebar Ala Menu Game + Ringan di Mobile
Tanggal: 2026-08-24 · Basis: `eda5c75` (shell usable hasil revert)
## Tujuan
1. Tema **monokrom murni** (hitam-putih, tanpa warna) di dark & light.
2. Sidebar (desktop NavRail + mobile dock) beranimasi **ala menu game** — corner
brackets, sweep, stagger masuk, marker segitiga.
3. **Ringan di mobile**: matikan WebGL ambient di layar kecil, kurangi biaya
blur/backdrop, animasi transform/opacity saja.
## Non-goals
- Tidak menyentuh backend, endpoint, hooks/data-flow, struktur route.
- Tidak menambah dependensi baru (CSS murni untuk semua animasi).
## File yang disentuh
| File | Perubahan |
|---|---|
| `src/app/globals.css` | Token mono (dark+light): signal/amber/vermilion → skala putih-abu; `.glass` blur adaptif; kelas baru `.game-nav-item` (bracket ::before/::after, sweep, stagger via `--i`), `.game-frame` (panel sudut terpotong + garis tergambar), keyframes `sweep-x`, `draw-line`, `nav-in`; media query `<md`: blur 18→8px, hambat animasi berat |
| `src/components/shell/nav-rail.tsx` | Item pakai `.game-nav-item` + `style={{'--i': n}}`; marker aktif jadi segitiga ▸ putih; hapus box-shadow glow besar (ganti sweep) |
| `src/components/shell/mobile-nav.tsx` | Dock mono: tab aktif = bar atas putih + sweep sekali; target sentuh ≥44px; hapus glow blob |
| `src/components/shell/topbar.tsx` | Aksen mono + `.game-frame` pada container (cek markup dulu) |
| `src/components/ambient/ambient-canvas.tsx` | Early-return WebGL bila `(pointer: coarse)` / lebar <768 / `saveData` / core ≤4; fallback statik CSS tetap |
| `src/components/ambient/status/signal tone` (`SIGNAL_RGB`) | Semua tone jadi grayscale (putih; intensitas beda per tone) |
| `src/app/(dashboard)/dashboard/view.tsx` | Hero + kartu metrik pakai `.game-frame`/cut-corner sebagai showcase |
## Keputusan desain
- **Full monokrom termasuk danger**: flag/moderation tidak lagi merah —
ditandai badge putih-di-atlas-hitam inversi + pulse. Kalau user kangen merah,
tinggal isi ulang `--color-vermilion`.
- Semua animasi hanya `transform`/`opacity` (compositor-friendly), hormati
`prefers-reduced-motion` (sudah ada kill-switch global).
## Verifikasi (gerbang)
1. `tsc --noEmit` bersih; biome 0 error 0 warning.
2. `pnpm build` sukses; smoke lokal 4024 → 9 route 200.
3. Push → GHA "Build & Deploy (Nix)" hijau → live 9×200.
4. Visual check live: desktop (rail game-menu terlihat) + cek rule mobile
(media query & gate kode) — screenshot disimpan.
@@ -0,0 +1,56 @@
# Spec: Simpan pesan NSFW tanpa analisis AI (Request 1)
Date: 2026-08-30
## Goal
Saat ini GMW **skip capture** untuk pesan di channel age-restricted/NSFW
(`messageCapture.ts` baris 291 & 310 memanggil `isAgeRestrictedMessage` lalu
`return`). User ingin pesan NSFW tetap **disimpan** ke database (jadi terlihat
di dashboard), tetapi **tidak dianalisis AI** (tidak dipanggil LLM).
## Behavior target
1. Pesan NSFW/age-restricted di `messageCreate` → DISIMPAN (capture normal).
2. Pesan NSFW di `messageUpdate`/`messageDelete` → tetap diproses seperti pesan
biasa (edit/deleted state tercatat).
3. AI analysis TIDAK berjalan untuk pesan NSFW. Sudah ada jalur yang benar:
`queueMessageAnalysis` → `isAgeRestrictedMessage(message)` →
`buildAgeRestrictedSkipResult()` yang menulis `status=clean` + flag
`age_restricted` + `action=none` TANPA memanggil LLM. Jalur ini sudah ada dan
dipakai di `aiAnalyzer.queueMessageAnalysis` (baris 72-84) dan
`batchProcessor.skipAgeRestrictedMessages`. Jadi tinggal membuka capture.
4. Konten NSFW TIDAK masuk ke Qdrant public archive (semantic search publik).
`archiveMessageEmbedded` harus di-guard untuk pesan age-restricted.
## Files touched
- `services/discord-gateway/src/modules/message-capture/messageCapture.ts`
- Hapus guard `isAgeRestrictedMessage` di `messageCreate` (baris 291) dan
`messageUpdate` (baris 310). Jangan hapus guard `isExcludedThread`,
`isBotExcludedChannel`, `shouldCaptureForAnyTarget`.
- `services/discord-gateway/src/modules/message-capture/archiveEmbedder.ts`
- `archiveMessageEmbedded` terima flag/cek metadata age-restricted → skip
embed untuk NSFW.
- Call-site di `messageCapture.ts` `captureMessage` (baris 218) pass isNSFW
atau cek dulu.
## Schema/type changes
- TIDAK ada perubahan DB schema. Metadata pesan sudah membawa `channel.nsfw`
(`getMessageLocation`). Tidak perlu kolom baru — flag `age_restricted` sudah
ditulis ke `ai_moderation_flags` lewat skip-result.
## Verification
- `pnpm typecheck`, `pnpm lint` (biome), `pnpm build` di `services/discord-gateway`.
- Unit test: pastikan ada test untuk age-restricted skip (sudah ada
`tests/conversationContext.test.ts` ref nsfw; cek apakah ada test sesuai).
- Deploy via GHA push; verify pesan NSFW muncul di DB dan `ai_status` = clean
dengan flag age_restricted, dan TIDAK ada panggilan LLM (log llm-caller tidak
menampilkan id pesan NSFW).
## Status
- Request 1: DONE & DEPLOYED (commit 50967f64, CI green, gateway restart 20:17 WIB).
- Request 2 (video): dicatat, belum dikerjakan — lihat section di bawah.
## Request 2 (video) — NOT in this change
Fitur record video kamera/screenshare orang lain. @discordjs/voice hanya
mendukung **audio** receive. Video orang lain butuh WebRTC viewer baru
(SDP answer, decrypt H264/VP8, decode frame, mux MP4/WebM). Diluar scope
changeset ini; dicatat untuk desain lanjutan.
@@ -0,0 +1,82 @@
# Spec: Recordings — filter per user + export WAV (Audacity)
## Konteks / Gejala
Halaman `services/frontend/src/app/(dashboard)/recordings` menampilkan semua
rekaman voice (deck). User ingin:
1. **Filter per orang** (tampil rekaman satu user saja).
2. **Export ke format untuk Audacity** (buka & edit rekaman di Audacity).
## Fakta saat ini (verified)
- Backend `recordings.list` (services/backend/src/orpc/router.ts:291) SUDAH
menerima `userId`/`channelId` filter → `RecordingsService.getRecent`.
- Frontend `recordingsApi.list(limit, channelId, userId, cursor)` (lib/api/recordings.ts)
sudah meneruskan `userId`. `useLoadMoreRecordings` juga sudah bawa userId.
- Tapi UI `RecordingsView` (app/(dashboard)/recordings/view.tsx) TIDAK punya
filter UI, dan `useRecordingsPage` dipanggil tanpa userId → semua tampil.
- Setiap rekaman punya `download_url` (MP3 di TeleUploader), `user_id`, `username`.
- Audacity membuka MP3/OGG tapi editing paling bersih dari WAV (uncompressed)
/ FLAC (lossless). Backend TIDAK punya ffmpeg & Nix flake backend tak include
ffmpeg → transcode server-side bukan pilihan. Browser punya codec MP3 → export
WAV via Web Audio API (client-side) adalah solusi self-contained terbaik.
## Keputusan desain
1. **Filter per user**: UI dropdown (Semua User + per user) di header halaman.
Memilih user → re-fetch `recordingsApi.list(50, undefined, userId)` (server
filter, benar untuk dataset besar + pagination). Dropdown dibangun dari
distinct `user_id`/`username` pada items yang sedang tampil.
2. **Export WAV (Audacity)**: client-side via Web Audio API.
- Per kartu: tombol "WAV" → decode `download_url` → WAV 16-bit PCM → download.
- Header: tombol "EXPORT WAV (N)" → gabung (concat) semua rekaman yang
sedang tampil (ter-filter) jadi 1 file WAV → download. Ideal untuk analisis
/ mixdown per orang.
- Implementasi di `lib/audio/wav.ts` (decode + encode + concat), tanpa dep baru.
## Perubahan
### Frontend
- **`src/lib/audio/wav.ts`** (baru):
- `decodeAudio(url: string): Promise<AudioBuffer>` — fetch arrayBuffer →
`new AudioContext().decodeAudioData`.
- `audioBufferToWav(buf: AudioBuffer, sampleRate=48000): Blob` — PCM 16-bit
interleaved, mono→stereo handling, RIFF/WAVE writer. Audacity-importable.
- `concatBuffers(buffers: AudioBuffer[]): AudioBuffer` — gabung di channel 0
(mono) dengan sample-rate max; untuk export gabungan.
- `downloadWav(blob: Blob, filename: string): void` — obj URL + <a download>.
- **`src/app/(dashboard)/recordings/view.tsx`**:
- Toolbar filter: dropdown user (built from distinct items) + tombol reset.
- State `filterUserId`; saat berubah → `recordingsApi.list(50, undefined, id)`
→ set ke SWR (key includes filter), reset pagination.
- Tombol "WAV" per kartu (disabled jika `!r.download_url`).
- Tombol "EXPORT WAV (N)" di header (disabled jika 0 item punya download_url);
concat semua items ter-filter yang punya download_url.
- Status loading saat export (spinner/disable).
- **`src/hooks/use-recordings.ts`**: `useRecordingsPage` menerima `userId?` dan
memasukkan ke key + call, supaya filter re-fetch bersih (per-user cache key).
`useRecordings`/`useLoadMoreRecordings` propagate `userId`.
- **`src/lib/types/recording.ts`**: tidak berubah (userId dari items).
### Backend / gateway
- Tidak ada perubahan. Filter & export sepenuhnya frontend.
## File yang disentuh (frontend only)
- `src/lib/audio/wav.ts` (baru)
- `src/app/(dashboard)/recordings/view.tsx`
- `src/hooks/use-recordings.ts`
## Verification
1. `cd services/frontend && pnpm typecheck` (tsc --noEmit) — 0 error.
2. `pnpm lint` (biome check src/) — exit 0.
3. `pnpm build` (next build) — hijau.
4. Manual (user): buka /recordings; pilih user di dropdown → hanya rekaman user
itu; klik WAV di kartu → file .wav ter-download & terbuka di Audacity; klik
EXPORT WAV (filtered) → satu .wav gabungan.
5. Push → CI `Build & Deploy (Nix)` (frontend job) hijau → deploy landing.
## Risiko / Trade-off
- Web Audio decode MP3 di client: butuh CORS pada download_url (TeleUploader
asepharyana.my.id — sudah same-serve/proxied, CORS ikut origin). Jika 403/CORS
gagal, error toaster + fallback manual (RAW MP3 tetap ada).
- concat gabungan = mono 48k; Audacity bisa edit per-channel nanti. Acceptable.
- Filter client (dropdown dari items yang dimuat) hanya menawarkan user yang
sudah tampil; dataset besar bisa pakai search nanti. Server filter benar untuk
yang dipilih.
@@ -0,0 +1,81 @@
# Spec: Recordings v2 — transcription, search, filters, leaderboard, sessions
## Konteks
4 fitur lanjutan untuk halaman /recordings (dipilih user):
1. Tampilkan transkripsi + search by kata kunci
2. Filter lanjutan: by channel + rentang tanggal
3. Kelompokkan klip jadi "sesi rapat" + autoplay berurutan + export satu sesi
4. Leaderboard bicara per user + ringkasan
## Fakta terverifikasi (2026-08-30)
- `voice_recordings` kolom: id, user_id, username, avatar_url, guild_id,
channel_id, channel_name, filename, size_bytes, download_url, upload_status,
upload_error, created_at, uploaded_at, **transcription** (schema
`shared/database/schema.ts:246`). Index user_id/channel_id/created_at.
- **Transcription 0/10.640** terisi prod: `AI_VOICE_TRANSCRIPTION_ENABLED`
default **false** (config/index.ts:333) & tidak diset di BWS env →
`transcribeRecording` (voiceTranscriber.ts) langsung return null.
`AI_LLM_BASE_URL` + `AI_LLM_API_KEY` SUDAH dikonfig (Whisper via router GMW).
- Transcriber hardcode `language: "en"` (voiceTranscriber.ts:34) — salah utk
ucapan campur id/en. Utk auto-detect: hapus param `language` (Whisper
auto-detect jika tidak diberikan).
- Backend `RecordingsService.getRecent` (recordings.service.ts) TIDAK select
`transcription`; SUDAH dukung filter `channelId`+`userId`+`cursor`; belum
dukung date-range & keyword search. `RecordingRow` interface juga tak punya
`transcription`.
- Tidak ada kolom `session_id` / `duration_ms` → sesi grouping = heuristik
(channel sama + gap created_at), durasi leaderboard = estimasi dari
size_bytes (MP3 128kbps: durasi_s ≈ size_bytes*8/128000).
## Keputusan desain
1. **Aktifkan transkripsi (fondasi)**: transcriber auto-detect (hapus
`language:"en"`), set secret BWS `AI_VOICE_TRANSCRIPTION_ENABLED=true`
(dibaca runtime oleh bws-exec saat service start). Rekaman BARU dapat
transkripsi. Backfill rekaman lama TIDAK dilakukan (pilih user: fokus baru;
file OGG lama kemungkinan besar sudah tidak dipakai).
2. **Backend** — perluas `getRecent`:
- select `transcription` (+ interface RecordingRow + FE type)
- filter baru: `q` (ILIKE on transcription + username), `startDate`/`endDate`
(created_at range, bigint ms)
- endpoint baru `recordings.summary`: agregasi per user → {user_id,
username, avatar_url, clips, est_duration_s, words, last_at}. `words`
dihitung dari transcription (tokenisasi spasi). Return sorted by clips.
3. **Frontend**:
- Kartu: tampilkan transkripsi (collapse/expand line-clamp) + durasi estimasi.
- Toolbar: search box (q), Select channel, date range (start/end), speaker
(sudah ada), reset filter.
- Tab/segment "Tape Deck" vs "Leaderboard": leaderboard render summary per
user + klik → filter deck by user itu.
- Sesi grouping (deck view): klip di-group jadi sesi bila channel sama &
gap antar klip < SESSION_GAP_MS (default 120s). Header sesi (channel,
waktu mulai, jumlah klip, total durasi). Autoplay tombol "Play session" &
"Export session WAV" (concat klip sesi via lib/audio/wav.ts yg sudah ada).
## Perubahan file
### Gateway (Tahap 1)
- `voiceTranscriber.ts`: hapus baris `language: "en"`.
- (secret) set `AI_VOICE_TRANSCRIPTION_ENABLED=true` via bws.
### Backend (Tahap 2)
- `recordings.service.ts`: interface + select + getRecent tambah transcription;
tambah filter q/startDate/endDate; method getSummary() untuk leaderboard.
- `orpc/router.ts`: procedur `recordings.list` schema tambah fields; prosedur
baru `recordings.summary`.
### Frontend (Tahap 3-5)
- `lib/types/recording.ts`: tambah transcription, est_duration_s opsional,
Summary type.
- `lib/api/recordings.ts`: list tambah q/startDate/endDate; + summary().
- `hooks/use-recordings.ts`: propagate filter baru ke key+call; hook
useRecordingsSummary.
- `app/(dashboard)/recordings/view.tsx`: toolbar search+channel+date, kartu
transkripsi, tab leaderboard, grouping sesi + autoplay + export sesi.
## Verification
- tiap tahap: gateway `pnpm build` + `biome check src/`; backend `pnpm build`
+ `biome check src/ tests/`; FE `pnpm build` + `biome check src/`.
- CI Build & Deploy hijau tiap tahap; deploy landing dicek via
`systemctl show gmw-<svc>.service --property=ActiveEnterTimestamp`.
- Tahap 1c: setelah deploy + rekaman baru, cek
`SELECT COUNT(*) FROM voice_recordings WHERE transcription IS NOT NULL`.
@@ -0,0 +1,71 @@
# Spec: Record video (kamera/screenshare) orang lain — WebRTC receive (Request 2)
Date: 2026-08-30. Status: Phase A + B DONE (capture → playable MP4); Phase C (UI) open.
## Why this is hard (ground truth, verified from @discordjs/voice 0.19.2 source)
`VoiceReceiver.onUdpMessage` (dist/index.mjs:2059) drops EVERY non-opus RTP
packet at line 2068: `if ((msg[1] & 127) !== RTP_OPUS_PAYLOAD_TYPE) return;`.
So video (kamera H264? actually Discord uses VP8/H264; screenshare combines with
video SSRC) is decrypted-capable but never forwarded. `receiver.parsePacket`
(2033) DOES decrypt any payload type generically (audio + video) using
`connectionData.{encryptionMode, nonceBuffer, secretKey}` — the only audio gate
is the opus check inside onUdpMessage.
=> FIX: wrap `receiver.onUdpMessage` (like screenShareAudio.ts already does for
screen-share AUDIO SSRCs): for packets whose payload type is a VIDEO type
(payload 96 VP8, 101/102 H264, 106/116/126/127 AV1, VP9 98...), call
`receiver.parsePacket(...)` myself to decrypt, then depacketize + write frames.
Delegate opus (120) to the original handler. Delegate audio to original.
## Audio already works (screenShareAudio.ts). We add VIDEO.
## Science-of-the-changes below.
## Phase A — capture + decrypt + depacketize to AnnexB h264 (THIS change)
Files (new): `src/modules/voice-recording/videoReceiver.ts`
- Hook into `recorder.startRecording` alongside `hookScreenShareAudio`.
- Wrap `receiver.onUdpMessage`:
- read ssrc = msg.readUInt32BE(8); userData = receiver.ssrcMap.get(ssrc)
- if payload type is video AND we have a "watching" subscription for that user
(videoSSRC present), decrypt via receiver.parsePacket(...), then:
- H264 (101/102 + payload 120 not): strip RTP header, reassemble FU-A
fragments into AnnexB NALs (start-code prefixed), buffer until we have
a full access unit (keyframe SPS/PPS/IDR or slices), append to a per-
user-per-burst `.h264` file.
- else delegate to original onUdpMessage.
- Watch `receiver.ssrcMap` "create"/"update" for `videoSSRC !== undefined` →
signal a video burst started for that user (like screenShareAudio does).
- Per-user video files written to `config.RECORDINGS_DIR/<uid>/video-<ts>.h264`.
- Guard: skip bot's own video (client.user.id).
Dependencies: NO new npm deps for Phase A (only crypto already in
@discordjs/voice via parsePacket + Buffer). ffmpeg-headless (already in Nix
buildInputs) used in Phase B for decode+mux.
## Phase B — decode + mux to playable MP4/WebM (DONE, commit 999c054b)
- `closeBurst` waits for the WriteStream `finish` (full flush/fd close), then
`muxToMp4(rawPath)`: `ffmpeg -f h264 -i raw.h264 -c copy -movflags +faststart
out.mp4`, deletes raw on success (>=1B mp4), keeps it on failure.
- Output: `<RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.mp4`.
- ffmpeg is on the gateway runtime PATH (pkgs.ffmpeg-headless, already in the
Nix buildInputs for the music/GoLive players).
- Test: `tests/videoReceiver.test.ts` muxToMp4 case (real ffmpeg, generates a
tiny baseline h264, asserts mp4 non-empty + raw deleted; skipped if no ffmpeg).
## Phase C — frontend playback + session grouping (follow-up)
- Backend oRPC list video files; FE video player, group by call session like audio.
## Verification
- Phase A: join voice, have a member screen-share/camera, confirm `.h264` file
grows with NAL frames + keyframes; journal shows "video burst" logs.
- Run vitest unit: RTP header strip + FU-A reassembly gives correct bytes.
## Open questions / risks
- Discord codec for camera = H264(101/102); screenshare uses H264 (101/103?)
and can also be VP8/VP9. Handle H264 first (depacketize proven), VP8/VP9 in
Phase B via ffmpeg RTP input.
- Encryption: DAVE (dave_protocol_version) adds a session layer; parsePacket
already applies daveSession.decrypt for audio — we must call the SAME
parsePacket path so DAVE/encryption is handled identically.
- ssrc↔user mapping during a broadcast: videoSSRC is in ssrcMap after the
voice state; may need the STREAM_CREATE network events to key reliably.
@@ -0,0 +1,66 @@
# Spec: Voice auto-reconnect (persistent state + rejoin on drop)
## Goal
Setelah `VoiceController.connect()` berhasil, state "sedang merekam di <guild>/<channel>"
disimpan di Postgres. Kalau gateway restart/reboot, atau koneksi voice drop tidak
disengaja (dikeluarkan/moved/server restart), gateway otomatis join ulang ke channel
yang sama.
## Requirement mapping (user's ask)
- "autoreconnect ke channel yg sama jika server restart atau reboot" → reconnect on
startup (ready handler) + keep DB record across graceful shutdown.
- "state nya persistent di db" → `voice_auto_reconnect` table.
- "rejoin jika tidak sengaja dikeluarkan" → watchdog on `Disconnected`/`Destroyed`
(kick / moved / voice server restart) → full rejoin with backoff.
- Manual leave (`/voice disconnect`, dashboard disconnect) MUST NOT rejoin.
## Design decisions
1. **New table** `voice_auto_reconnect` (dedicated, not `ui_state`):
- `guild_id` text PK
- `channel_id` text NOT NULL
- `channel_name` text
- `connected_at` bigint epoch-ms
- `updated_at` bigint epoch-ms
DAO: `voiceAutoReconnectRepo.ts` — `upsert(record)`, `list()`, `delete(guildId)`.
2. **Write on connect**: `VoiceController.connect()` → after `startRecording` success →
`upsert`. IDEMPOTENT (upsert per guild).
3. **Clear on manual leave**: `handleVoiceDisconnect` (all) + `handleVoiceDisconnectGuild`
pass `clearPersisted: true`. Graceful shutdown `disconnect()` keeps the record.
4. **Rejoin on startup**: bootstrap `ready` → `await voiceController.autoReconnect()`
(list persisted → connect each, non-fatal on failure).
5. **Rejoin on unexpected drop**: monitor per-connection; on `Disconnected`/`Destroyed`
with `!intentional` → schedule rejoin `connect(guildId, persistedChannelId)` with
backoff (min 2s, max 30s, max 5 attempts). Track `rejoinAttempts`, reset on success.
6. **Intentional flag**: `disconnectGuild(guildId, { clearPersisted?, intentional? })`.
- shutdown `disconnect()` → `{ intentional: true, clearPersisted: false }`.
- manual `disconnect()` (dashboard) → `{ clearPersisted: true }`, sets intentional.
- manual `disconnectGuild` → `{ clearPersisted: true }`, sets intentional.
The monitor checks `intentional` before rejoin; `clearPersisted` only deletes the row.
## Files touched
- `services/discord-gateway/src/shared/database/schema.ts` — add `pgVoiceAutoReconnectTable`
+ types.
- `services/discord-gateway/src/shared/database/voiceAutoReconnectRepo.ts` (NEW) — DAO.
- `services/discord-gateway/src/modules/voice-recording/voiceController.ts` — upsert on
connect; monitor + rejoin; `autoReconnect()`; `disconnect/disconnectGuild` opts.
- `services/discord-gateway/src/modules/command-handler/voice.handler.ts` — manual
disconnect/disconnectGuild pass `clearPersisted: true`.
- `services/discord-gateway/src/app/bootstrap.ts` — call `voiceController.autoReconnect()`
in `ready`.
- `services/discord-gateway/drizzle/migrations/0020_add_voice_auto_reconnect.sql` +
`meta/_journal.json` entry (apply manually per gmw-ops).
## Edge cases
- Channel deleted / guild lost while persisted → `connect()` throws (channel not found)
→ log + delete persisted row (don't retry forever).
- Rejoin attempts exhausted → keep row (so next restart retries) + log.
- Multiple guilds: per-guild monitor, per-guild persisted row.
- Graceful shutdown order: shutdown sets intentional=true (so no rejoin during teardown)
but keeps row.
## Verification
- `pnpm typecheck && pnpm build && pnpm lint` in `services/discord-gateway`.
- Apply migration `0020` manually; verify table exists.
- CI `Build & Deploy (Nix)` green; gateway deploy lands.
- Manual: connect via dashboard → check `voice_auto_reconnect` row; simulate drop →
confirm rejoin; manual disconnect → row cleared.
@@ -0,0 +1,166 @@
# Spec: Perbaiki alur voice → recording (miss & terpotong)
## Konteks & Gejala
User melaporkan alur voice sampai recording **banyak miss** (audio tidak tercatat)
dan **terpotong** (satu alur bicara kebelah jadi beberapa segmen / audio putus di
tengah). Ini domain `services/discord-gateway/src/modules/voice-recording/`.
Pipeline per user yang mulai bicara (speaking "start"):
```
receiver.speaking "start" → speakingHandler(userId)
├─ await collectUserMetadata(...) ← roundtrip API, subscribe tertunda
├─ receiver.subscribe(userId, {end: AfterSilence, duration: 3000ms}) → audioStream
├─ attach data/end/error handlers → audioStream.pipe(PacketFilter) → oggPacketStream
├─ SegmentManager.open() → OggLogicalBitstream → file .ogg
├─ data: SegmentManager.rotateIfNeeded (rotasi 5s) + decoder.write (web PCM tho
└─ end: SegmentManager.close() → segmen finish → finalizeSegment upload + transkrip
```
## Root cause (dari pembacaan kode — justifikasi di bawah)
### A. MISS bagian awal bicara — subscribe tertunda (utama)
`speakingHandler.ts:53` melakukan `await collectUserMetadata(...)` SEBELUM
`receiver.subscribe`. `collectUserMetadata` (metadata.ts:30) pada cold path
(cache miss) melakukan `client.users.fetch` + `guild.members.fetch` roundtrip
Discord API (ratusan ms–detik). Selama await, seluruh opus awal bocor → awal
kalimat hilang. Cache menghilangkan ini untuk user yang pernah ter-record, tapi
user baru/evict (cache max 200) kena setiap kali.
### B. Double-subscribe race
Guard `receiver.subscriptions.has(userId)` di `speakingHandler.ts:63` diletakkan
SETELAH `await collectUserMetadata`. Dua event "start" cepat keduanya melewati
guard (belum subscribe) → dua subscription → audio terbelah/ganda per user.
### C. TERPOTONG di jeda bicara — AfterSilence 3000ms
`AUDIO_STREAM_SILENCE_DURATION_MS=3000`. Setelah 3s diam, stream auto-`end` →
`SegmentManager.close` → segmen baru saat bicara lagi. Jeda normal (berpikir,
interupsi) memecah 1 alur bicara jadi beberapa segmen/file. Ini source "terpotong".
Segmen pendek hasil jeda <1s juga DIBUANG di `finalizeSegment` (MIN_DURATION_MS=1000)
→ miss kata singkat ("ya", "siap").
### D. Rotasi segmen 5s di tengah bicara
`RECORDING_SEGMENT_MS=5000`: `rotateIfNeeded` menutup bitstream & membuka baru
setiap 5s walau bicara kontinu. Pipenya di-re-wire di dalam handler data →
window drop kecil + continuity file pecah (bukan masalah besar, tapi berkontribusi).
### E. Tidak ada sinkronisasi "stop" speaking & stream "end" flaky
Handler hanya listen "start"; mengandalkan `AfterSilence` untuk emit "end".
Bug @discordjs/voice yang dikenal: `AfterSilence` bisa TIDAK emit "end" saat
koneksi gagal/teardown → segmen menggantung & tidak pernah finalize/upload
(recording "hilang"). Tidak ada watchdog.
## Scope
Hanya `services/discord-gateway/src/modules/voice-recording/` (+ config index bila
perlu default baru). Tidak menyentuh playback (player.ts), transmitter (voice dari
browser → Discord, arah berlawanan), muxer (konsolidasi akhir), atau screen-share
video (hanya audio SSRC via `hookScreenShareAudio` sudah ada & dibiarkan).
## Perubahan
### 1. Speak-before-metadata: subscribe LEBIH DULU, metadata paralel
`recorder/speakingHandler.ts`:
- Pindahkan `receiver.subscribe` + pipeline setup ke ATAS, SEGERA di handler,
sebelum `collectUserMetadata`.
- Jalankan `collectUserMetadata` secara paralel non-blocking; gunakan metadata
cache untuk registrasi segmen saat finalize.
- Pertahankan guard skip bot/user (bot bisa dicek dari `client.users.cache` /
`client.user.id` tanpa await) SEBELUM subscribe — jangan tunggu fetch user.
Detail konkret:
```
async handler(userId):
if (userId === client.user?.id) return;
if (receiver.subscriptions.has(userId)) return; // guard kini di DEPAN, tanpa await
// (belum tahu bot? gunakan cache user; subscribe dulu biar nggak miss)
clone = subscribe(userId, {AfterSilence, duration}) // TANPA await metadata
setup pipeline (data/end/error, pipe, open segment)
collectUserMetadata(...).then(meta => {
if (meta.bot) { drain & close subscription (jangan simpan) }
else { registrasi ulang metadata utk segmen aktif }
})
```
Karena listener `start` dibuang untuk bot, harus tutup subscription bot tanpa
menyimpan segmen (buang hasil). Pakai `receiver.subscriptions.get(userId)?.destroy()`.
### 2. Selesaikan race double-subscribe (bagian dari #1)
Guard `receiver.subscriptions.has(userId)` diletakkan SINCRON di awal (sebelum
await). Karena `subscribe` sinkron dan `subscriptions` terisi sinkron saat
dipanggil, event "start" kedua yang tiba setelah subscribe akan melihat
subscription aktif → di-skip. Tidak ada await antara guard & subscribe.
### 3. Naikkan AfterSilence + tail-length → kurangi "terpotong"
`shared/config/index.ts`:
- `AUDIO_STREAM_SILENCE_DURATION_MS` default 3000 → 4000 (beri ruang jeda
alami; Discord packet 20ms, 4s masih wajar, tidak membengkak file).
Opsional via env override di production (tidak wajib komit env).
### 4. Segmen berorientasi "burst bicara" daripada rotasi jam
`recorder/segment.ts` + `recorder/speakingHandler.ts`:
- Hapus/lepas rotasi SEGMEN berbasis waktu (RECORDING_SEGMENT_MS). Alih-alih,
satu segmen = satu burst bicara (buka di "start", tutup di "end"/AfterSilence
end). Ini menghilangkan pemecahan di tengah kalimat.
- `RECORDING_SEGMENT_MS` tetap dipakai untuk rotasi decoder web-PCM (broadcast
live), di mana segmen besar bisa menunda frame — biarkan seperti ada.
JADI: `SegmentManager.rotateIfNeeded` TIDAK lagi dipanggil pada jalur OGG
recording; decoder rotate tetap dijalankan.
Catatan: dengan satu segmen per burst, ukuran file ~ durasi bicara. File panjang
dibutuhkan transkrip & transcode; tidak ada batas keras yang perlu di-override.
Watchdog di #5 membatasi durasi menggantung.
### 5. Watchdog end-of-burst & teardown recovery
`recorder/speakingHandler.ts`:
- Setelah subscribe, arm timer watchdog (mis. `config twin`/hitung) yang menutup
segmen jika `AfterSilence` tidak emit "end" dalam X detik setelah "stop"
speaking — atau, lebih sederhana & robust: dengarkan BOTH stream "end" DAN
timer dari `receiver.speaking` "stop" (hingga @discordjs/voice meng-klaim
AfterSilence). Bila "stop" fire, mulai countdown kecil (mis. 500ms) lalu
`segmentManager.close` + `decoder.destroy` + destroy subscription jika stream
belum "end".
- Ini menutup A: segmen menggantung → jadi pasti finalize & upload.
Implementasi: subscriptionStream (audioStream) + track milik per-user di
Map<userId, {audioStream, segmentManager, decoder, timer}>; handler "stop"
menjadwalkan finalize.
### 6. Naikkan/ambil MIN segmen duration lebih rendah
`recorder/segmentFinalizer.ts`: MIN_DURATION_MS 1000 → 300ms. Kata pendek
("ya", "siap") tetap tersimpan. GUI biarkan.
## File yang disentuh
- `src/modules/voice-recording/recorder/speakingHandler.ts` (utama: subscribe
first, guard depan, watchdog stop, hapus rotasi segmen dari jalur OGG)
- `src/modules/voice-recording/recorder/segment.ts` (opsional: API close/open,
pertahankan rotate untuk decoder tapi tak dipakai jalur OGG)
- `src/modules/voice-recording/recorder/streamSetup.ts` (kecil: backfill subscribe
supaya return subscription utk cleanup/destroy bot)
- `src/modules/voice-recording/recorder/segmentFinalizer.ts` (MIN_DURATION)
- `src/shared/config/index.ts` (default AfterSilence 4000)
## Yang TIDAK disentuh
- `transmitter.ts` (arah browser→Discord, bukan recording)
- `player.ts`, `mediaSource.ts`, `screenShareAudio.ts` (hook SSRC sudah benar)
- `muxer.ts` (konsolidasi akhir tetap jalan)
- Flake/deps/bundle
## Verification
1. `cd services/discord-gateway && pnpm typecheck` (tsc --noEmit) — 0 error.
2. `pnpm lint` (biome check src/) — exit 0.
3. `pnpm build` (tsc → dist/).
4. Unit test baru (vitest, kalau infra tes ada):
- subscribe terjadi tanpa await metadata (spy urutan panggilan)
- double-start skip lewat guard sinkron
- "stop" → watchdog finalize segmen walau stream tidak "end"
- MIN_DURATION 300ms menyimpan kata pendek
5. Smoke/CI: `nix flake check` (eval). Push → CI `Build & Deploy (Nix)` hijau →
deploy landing (`systemctl show gmw-discord-gateway.service --property=ActiveEnterTimestamp`).
6. Runtime manual (user): join voice, bicara dengan jeda >3s, pastikan 1 alur
kontinu = 1 segmen utuh (bukan 3), dan awal kata tidak hilang.
## Risiko / Trade-off
- Subscribe-before-metadata: burst bot akan dikumpulkan sesaat lalu dibuang
(cost kecil: buang segmen). Lebih baik miss bot daripada miss user.
- AfterSilence naik: file lebih panjang sedikit saat jeda; upload/transkrip
timeout (transcodeToMp3 30s) tetap aman.
- Satu segmen per burst: tidak ada rotasi paksa → durasi segmen = durasi bicara
(bisa menit). Transkrip & transcode tetap ok. Watchdog batasi menggantung.
@@ -0,0 +1,177 @@
# Spec: Receive Others' Screen-Share/Camera Video Under DAVE — Build a DAVE-capable Stream-Watch Connection
Status: **P1–P3 DONE + 4th CRITICAL FIX deployed (f1a7b0c2); DAVE Ready + MLS handshake CONFIRMED live; P4 = waiting on active streamer to confirm video-burst→mp4**
Date: 2026-08-31
Author: Hermes
Related: `.hermes/plans/2026-08-31_video-receive-eager-selfbot-connection-spec.md` (superseded by this)
`.hermes/plans/2026-08-31_video-receive-phaseC-spec.md` (Phase C build, selfbot path — dead)
## Problem / Ground truth (established from live logs 2026-08-31)
GMW must record OTHER members' screen-share + camera video in a voice channel it
records. Audio works (via `@discordjs/voice` 0.19.2 negotiating DAVE). Video does
not. Verified live: the selfbot path (`discord.js-selfbot-v13` eager `joinChannel`
→ `joinStreamConnection` → `receiver.createVideoStream`) authenticates but Discord
closes the connection with WS code **4017 "E2EE/DAVE protocol required"** (5x →
`VOICE_CONNECTION_ATTEMPTS_EXCEEDED`). Root cause: **Discord now REQUIRES DAVE
(E2EE) on every voice RTC, and `discord.js-selfbot-v13`'s voice stack predates
DAVE** (identify has no `max_dave_protocol_version`, no MLS handshake). The selfbot
path is dead, cannot be repaired. Full details: skill `gmw-ops` →
`references/video-receive-and-unmute.md` §4.
Facts:
- Watching a stream = a **SEPARATE RTC connection**, not the guild audio socket:
gateway `STREAM_WATCH` (op 20) → Discord replies `STREAM_CREATE` (rtc_server_id)
+ `STREAM_SERVER_UPDATE` (separate token+endpoint) → client opens its own voice
WS+UDP to that endpoint (`StreamConnectionReadonly` in selfbot). The watched
video never rides the @discordjs/voice guild socket.
- The stream-watch RTC ALSO requires DAVE (same 4017 mechanism).
- `@snazzah/davey` (bundled with @discordjs/voice 0.19.2) supports
`MediaType.VIDEO` + `Codec.H264` decrypt — DAVE machinery CAN decrypt H264 video.
- No off-the-shelf DAVE-capable video-RECEIVE path exists. Closing references:
- **Discord-RE/Discord-video-stream** (fork of `@dank074/discord-video-stream`,
master 2026-08-28): full DAVE in `src/client/voice/BaseMediaConnection.ts`
(Davey `DAVESession` init via `initDave`, MLS key-package / proposals /
commit / welcome / transitions; `WebRtcConnWrapper` encrypts audio/video via
`daveSession.encrypt(MediaType.VIDEO, codec, …)`). BUT it is STREAMING only
(send). No STREAM_WATCH / receive.
- `@discordjs/voice`: full DAVE receive for AUDIO only; `DAVESession.decrypt`
hardcodes `MediaType.AUDIO` (dist ~line 892); `onUdpMessage` drops non-opus;
no STREAM_WATCH.
- `discord.js-selfbot-v13`: video receive but no DAVE.
## Goal
Replace the dead selfbot receive path with a **DAVE-capable stream-watch voice
connection**: on detecting a member `voiceState.streaming`, send `STREAM_WATCH`,
connect a DAVE-authenticated RTC to the stream endpoint, decrypt incoming H264
RTP (`MediaType.VIDEO`), reassemble via the existing `H264Depacketizer`, mux to a
playable container. Reuse every tested building block already in the repo.
## Strategy decision (default A; B as fallback) — de-risk in Phase 2
Two implementation routes. Decide by Phase 2 prototype result.
### Strategy A — extend @discordjs/voice's tested native stack (PREFERRED, lighter)
Reuse @discordjs/voice 0.19.2 internals (already a runtime dep, already DAVE-tested
for audio):
- Drive a connection to the stream endpoint using djs/voice's `VoiceWebSocket` +
`VoiceUDPSocket` + `DAVESession` (the same classes that work for the guild
connection — they take arbitrary endpoint/token/session).
- Send the voice identify with `max_dave_protocol_version`, complete the DAVE
handshake (Davey), then on receipt of a video RTP packet call
`daveSession.decrypt(userId, MediaType.VIDEO, packet)` (Davey exposes
`MediaType.VIDEO` + `Codec.H264`) — djs/voice's hardcoded `AUDIO` is the only
blocker, fix by invoking Davey directly with `MediaType.VIDEO` for video SSRCs.
- STREAM_WATCH sent via the existing selfbot `client.ws.broadcast` (cheap, works —
it needs no selfbot voice connection).
- Feed decrypted H264 → `H264Depacketizer` → `.h264` → `muxToMp4` (both already
in `videoReceiver.ts`, unit-tested).
- No new runtime deps. Risk: relies on non-exported djs/voice internals (reachable
via `as any`, as the existing `parsePacket` usage shows).
### Strategy B — port Discord-RE's BaseMediaConnection (heavier, more self-contained)
Port `BaseMediaConnection.ts` DAVE handling + `WebRtcConnWrapper` into a
receive/watch connection in the gateway. Deps: requires `@lng2004/node-datachannel`
(new native WebRTC dep) + `@snazzah/davey` (already available). More code, more
risk (native dep in Nix store), but a clean-room receive path decoupled from
djs/voice internals. Use only if A proves infeasible.
## Files touched (Strategy A shape)
- `services/discord-gateway/src/modules/voice-recording/streamWatchReceiver.ts`
(NEW): DAVE stream-watch connection wrapper. Owns, per watched user:
- `sendStreamWatch(client, streamKey)` (via `client.ws.broadcast({op:20,
d:{stream_key}})`), stream_key = `guild:<gid>:<chid>:<uid>`.
- collects STREAM_CREATE (rtc_server_id) + STREAM_SERVER_UPDATE (token+
endpoint) via `client.on('raw')` match on stream_key.
- builds a djs/voice-style connection to `<endpoint>` with the received token/
session; completes DAVE handshake.
- `onUdpMessage` wrapper: for video payload types, decrypt with
`daveSession.decrypt(userId, MediaType.VIDEO, buf)`, depacketize, write.
- teardown on STREAM_DELETE / user stops streaming / leave / channel untrack.
- `recorder.ts`: wire `trackChannel`/`untrackChannel` already exist; ensure the
selfbot *eager voice connection* attempt is REMOVED (it only 4017-spams logs) —
but KEEP `client.ws.broadcast` availability for STREAM_WATCH.
- `videoRecorder.ts`: remove the dead selfbot `joinChannel`/`joinStreamConnection`
calls; keep the `voiceStateUpdate` streaming detection + teardown bookkeeping as
the entry point; delegate the actual receive to `streamWatchReceiver`.
- `videoReceiver.ts`: keep `H264Depacketizer` + `muxToMp4` (reused). The
guild-socket `hookVideoReceiver` can be removed or left inert.
- Tests: `tests/streamWatchReceiver.test.ts` (DAVE-handshake stub, RTP decrypt path
with a mocked Davey, STREAM_WATCH packet shape); keep `tests/videoReceiver.test.ts`.
## Phases (each independently verifiable)
1. **Phase 1 (this session): spec + source reconnaissance.** Confirm djs/voice
internals are reachable (VoiceWebSocket/VoiceUDPSocket/DAVESession exports &
shapes), confirm Davey `MediaType.VIDEO` decrypt signature, confirm how a raw
stream-watch connection's identify/select-protocol flows. Verify the selfbot
`streamKey` format + `raw` STREAM_CREATE/SERVER_UPDATE payload. GATE: accurate
spec + no unknowns blocking A.
2. **Phase 2: de-risk prototype.** Standalone script (not in the gateway) that:
logs into the same selfbot token, joins a real voice channel, sends STREAM_WATCH
for a live streamer, receives STREAM_CREATE/SERVER_UPDATE, and attempts a
DAVE-authenticated connect + receive of ≥1 H264 packet to prove the path before
any gateway integration. GATE: at least one decrypted H264 NAL captured in the
lab.
3. **Phase 3: gateway integration** per files-touched. GATE: typecheck + build +
biome + unit tests green; CI deploy ok.
4. **Phase 4: live verify.** With a real streamer in a recorded channel: journal
shows `STREAM_WATCH sent`, `DAVE ready`, `Video burst opened`, and a playable
`.h264`/`.mp4`/`.mkv` on disk. GATE: playable file with real video content.
## Risks / open questions
- Does Discord require the stream-watch connection to use the SAME session_id as
the bot's active voice session, or a fresh one? (Affects identify.) Resolve in P2.
- Which video codec does Discord actually send for camera vs GoLive (H264 likely,
but VP9/AV1 possible) — the depacketizer only handles H264. P2 measures the
payload type live; add depacketizers for other codecs only if observed.
- djs/voice `DAVESession`/`VoiceUDPSocket` reachability via `as any` must be
confirmed against the installed 0.19.2 build (P1).
- The separate stream RTC may need `selectProtocol`/SDP even for receive-only; the
Discord-RE SDP shows a `m=video ... inactive` section. Follow the same shape.
## Verification (overall)
- Per-phase gates above.
- No regression: audio recording + message capture still work after changes.
- `pnpm typecheck && pnpm build && pnpm lint` green in discord-gateway.
- Commit + push; CI `Build & Deploy (Nix)` green; live streamer produces a file.
## Phase 1 findings (CONFIRMED 2026-08-31, Strategy A feasible)
- `@discordjs/voice` 0.19.2 dist/index.mjs PUBLICLY exports exactly the primitives
needed: `DAVESession`, `Networking`, `NetworkingStatusCode`, `VoiceConnection`,
`VoiceReceiver`, `VoiceUDPSocket`, `VoiceWebSocket`, `SSRCMap`,
`RTP_OPUS_PAYLOAD_TYPE` (export block ~3143). So a stream-watch connection can be
built OUTSIDE the lib using these constructors — no `as any` needed for the heavy
lifting.
- `Networking` child wiring (~line 1364-1484): `new VoiceWebSocket('wss://' +
endpoint + '?v=8', debug)`; on WS open send Identify `{op, d:{server_id,
user_id, session_id, token, max_dave_protocol_version: getMaxProtocolVersion()}}`;
`createDaveSession(protocolVersion)` → `new DAVESession(protocolVersion, userId,
channelId, {decryptionFailureTolerance})` then `.reinit()`; UDP via
`new VoiceUDPSocket({ip, port})` after Ready gives modes + ssrc.
- `DAVESession` wraps `@snazzah/davey` `Davey.DAVESession(protocolVersion, userId,
channelId)`; on network packets it calls `this.session.decrypt(userId,
Davey.MediaType.AUDIO, packet)` — hardcoded AUDIO (line ~892). For video we call
Davey directly with `MediaType.VIDEO` + `Codec.H264`.
- `@snazzah/davey` MediaType enum: AUDIO=0, VIDEO=1; Codec H264=4; methods
`decrypt(mediaType, codec, packet): Buffer` + `encrypt(...)`. Confirmed in davey
index.d.ts.
- Stream key format (selfbot VoiceConnection.js ~1240): `guild:<gid>:<chid>:<uid>`
for guild channels; `STREAM_WATCH` = gateway op 20, `d:{stream_key}`;
`sendSignalScreenshare` = `client.ws.broadcast({op:20,d:{stream_key}})`. Replies
come as gateway `raw` events `STREAM_CREATE` (`d.rtc_server_id`) +
`STREAM_SERVER_UPDATE` (`d.token`, `d.endpoint`); selfbot routes them to the
stream connection via `client.on('raw')` matching `d.stream_key`, setting
`setSessionId(sessionId)` + `setTokenAndEndpoint(token, endpoint)` (Watch case in
`StreamConnectionReadonly`).
- Discord-RE reference for the identify SDP: stream connections send a `m=video`
section with `a=inactive` (receive-only-ish) + standard DAVE/VoiceOpCodes
(op 0 identify, op 2 select protocol incl. `max_dave_protocol_version`). See
`BaseMediaConnection.handleProtocolAck` + `initDave`.
- Decision: proceed with **Strategy A**. Selfbot code to REMOVE: the eager
`ensureSelfbotVoice` join + `joinStreamConnection`/`receiver.createVideoStream`
in `videoRecorder.ts` (proven dead — 4017). Keep the `voiceStateUpdate` streaming
detection + bookkeeping; swap the receive plumbing to a new `streamWatchReceiver`
driven by a djs/voice-style connection. STREAM_WATCH itself still sent via the
selfbot `client.ws.broadcast` (needs only the WS, not a selfbot voice conn).
- OPEN (resolve in Phase 2 lab): (a) whether the stream connection's identify must
use the bot's ACTIVE voice session_id or a fresh one; (b) actual video codec Discord
sends for camera vs GoLive (measure payload type live; H264 assumed, VP9/AV1 possible
→ add depacketizers only if observed).
@@ -0,0 +1,93 @@
# Spec: Fix Video Capture — Eagerly Establish the Selfbot Voice Connection at Join Time
Status: PLANNED
Date: 2026-08-31
Author: Hermes
Related: `.hermes/plans/2026-08-31_video-receive-phaseC-spec.md` (Phase C build, made Option A this fix)
## Symptom (from live logs, 2026-08-31 ~12:34)
A user was actively screen-sharing + on camera in the recorded voice channel.
The gateway recorded MANY users' audio (.ogg) fine, but video capture produced
nothing. The only video signal in `journalctl -u gmw-discord-gateway` was:
```
[VOICE (guild:2)]: Sending voice state update: {"self_mute":false,...,"flags":2}
[VOICE] received voice state update: {member hunterz ...} # OTHER user, not bot
[VOICE] connection? true, guild session channel
[VOICE (guild:2)]: Setting sessionId <S> (stored as "undefined")
[VOICE (guild:2)]: Authenticated with sessionId <S> # debug print only
[VOICE (guild:2)]: Authenticate failed - VOICE_CONNECTION_TIMEOUT # +15s
video-recorder: userId=..., "Connection not established within 15 seconds."
```
## Root cause (verified against discord.js-selfbot-v13 3.7.1 source)
The gateway records audio via `@discordjs/voice` (`joinVoiceChannel` + adapter).
Video receive lives on the SEPARATE selfbot `ClientVoiceManager.connection`
(a singleton `VoiceConnection`). `videoRecorder.ts` currently calls
`client.voice.joinChannel(channel)` LAZILY — only when a `voiceStateUpdate`
shows `newState.streaming === true`.
At that moment the bot is ALREADY connected to the channel via @discordjs/voice.
A selfbot `joinChannel` then does `VoiceConnection.authenticate()` →
`sendVoiceStateUpdate()`, and waits for a fresh `VOICE_SERVER_UPDATE`
(`setTokenAndEndpoint`) + `VOICE_STATE_UPDATE` (`setSessionId`) to reach
`checkAuthenticated()` (needs token+endpoint+sessionId). Because the bot is
already in an established voice session, Discord does NOT emit a new
`VOICE_SERVER_UPDATE` for the lazy selfbot re-join → token/endpoint never set →
15s `VOICE_CONNECTION_TIMEOUT`.
This is fatal to video: `joinStreamConnection(userId)` (STREAM_WATCH op 20) and
`receiver.createVideoStream(userId, out)` (Recorder/ffmpeg) BOTH live on the
parent selfbot `VoiceConnection` and require it `CONNECTED` (its own voice
WS+UDP socket feeds `PacketHandler.push`, authenticated with
`authentication.secret_key`).
## Fix — establish the selfbot connection eagerly, at voice-join time
The selfbot `VoiceConnection` must exist and be `CONNECTED` before any streamer
appears. Establish it once, synchronously alongside the @discordjs/voice join in
`recorder.startRecording`, so it rides the bot's FRESH voice join — when Discord
DOES emit VOICE_SERVER_UPDATE. Then cache it and let `videoRecorder` reuse it.
Ordering: in `startRecording`, after the @discordjs/voice `joinVoiceChannel`
returns (and retries) — fire `ensureSelfbotVoice(channel)` best-effort:
1. `await client.voice.joinChannel(channel, { selfMute:false, selfDeaf:false,
selfVideo:false })` (rejects ~VOICE_CONNECTION_TIMEOUT on failure → log +
return null; do NOT block audio).
2. Cache the returned selfbot `VoiceConnection` keyed by guildId.
3. Wire teardown: on `recorder` voice stop / destroyed → `untrackChannel` +
destroy the cached selfbot connection (`disconnect()`).
`videoRecorder.startVideoRecording` then uses the cached selfbot connection:
- If cached & `status === CONNECTED` → use it.
- Else → fall back to a lazy `joinChannel` (still best-effort).
## The two-connection coexistence risk (must verify live)
@discordjs/voice (audio) and the selfbot `VoiceConnection` (video) each open
their OWN low-level voice WS+UDP on the same session. The spec's original
open-question flagged this. Mitigations:
- Clear logging: `Selfbot voice connected (guild=...)`, plus a periodic
`djs/voice status` log so we can confirm audio stays `READY` while the selfbot
connection is up.
- If Discord kicks/breaks the audio connection, logs will show
@discordjs/voice `Disconnected`/reconnect churn — we detect and pivot.
## Files touched
- `services/discord-gateway/src/modules/voice-recording/videoRecorder.ts`:
add `ensureSelfbotVoice(channel)` (return cached/connected), use it in
`startVideoRecording`, add `destroyGuildSelfbotVoice(guildId)`,
richer status logging.
- `services/discord-gateway/src/modules/voice-recording/recorder.ts`: call
`ensureSelfbotVoice(channel)` after `joinVoiceChannel` (best-effort);
call `destroyGuildSelfbotVoice` on voice stop/destroy.
- Tests: `tests/videoRecorder.test.ts` (update to assert eager-connection reuse
+ status gating).
## Verification
1. `pnpm typecheck` + `pnpm build` + biome clean (discord-gateway).
2. Tests green.
3. Commit + push; CI `Build & Deploy (Nix)` green, service restarts.
4. LIVE (deploy): join a channel with the bot → journal shows
`Selfbot voice connected` (parent CONNECTED). When a member streams →
`Sender signal screenshare` / `Video recorder ready` + a `.mkv` under
`<RECORDINGS_DIR>/<uid>/video-*.mkv`; playable via ffmpeg. Confirm audio
recording still flows (no djs/voice reconnect churn).
@@ -0,0 +1,97 @@
# Spec: Record Other Users' Video (Camera / Screen Share) — Phase C
Status: PLANNED (not built)
Date: 2026-08-31
Author: Hermes
Related: `.hermes/plans/2026-08-30_video-record-receive-spec.md` (Phase A/B — raw UDP hook, superseded for receive)
## TL;DR — what changed vs Phase A/B
Phase A/B (commit `999c054b` etc.) hooked `@discordjs/voice`'s UDP socket to capture non-opus RTP and
depacketize H264 → mp4. **It captured ZERO video** because `@discordjs/voice` never authorizes the bot to
receive others' video (no STREAM_WATCH). This spec replaces that approach with the **native, selfbot-lib
receive path**, which is battle-tested and does the authorization + decryption + ffmpeg muxing for us.
## Ground truth (verified in discord.js-selfbot-v13 3.7.1 source)
1. `ClientVoiceManager.joinChannel(channel, config)` → a **selfbot `VoiceConnection`** with
`.receiver` (`VoiceReceiver` → `PacketHandler`). [ClientVoiceManager.js:102-118]
2. `VoiceConnection.receiver` is created in the constructor. [VoiceConnection.js:140]
3. `VoiceReceiver.createVideoStream(user, output)` → `PacketHandler.makeVideoStream` → **`Recorder`**
(ffmpeg that muxes H264+Opus RTP over UDP → **Matroska (.mkv)**). [Receiver.js, Recorder.js]
4. `PacketHandler` routes: video RTP → Recorder UDP 65506, opus RTP → UDP 65510; decodes all via
`connection.authentication.{secret_key, mode}` (supports `aead_aes256_gcm_rtpsize` and
`aead_xchacha20_poly1305_rtpsize` = DAVE-compatible). [PacketHandler.js:115-155, 195-240]
5. `StreamConnectionReadonly.joinStreamConnection(userId)` + `sendSignalScreenshare()` sends
gateway op `STREAM_WATCH` so Discord actually forwards the streamer's RTP to us. [VoiceConnection.js:1100-1240]
6. `VoiceState.streaming` = `data.self_stream ?? false` — lets us detect a streamer on voice state update. [VoiceState.js:94]
## Problem / the crux
The gateway's voice today is **`@discordjs/voice`** (audio + music + GoLive-send). The selfbot-lib
video-receive path lives on the **selfbot-lib `VoiceConnection`** — a separate voice stack. Two options:
### Option A (RECOMMENDED): Parallel selfbot video-watch connection
Keep `@discordjs/voice` for everything it does today. Add a **second, selfbot-lib voice connection**
to the same channel whose ONLY job is to watch + record others' video.
- Pros: zero regression risk to audio/music/screenshare-send; uses native `createVideoStream` → mk4.
- Cons: two voice connections for the same bot user in one channel. Need to verify Discord tolerates it
(real selfbots like Discord-RE do exactly this for multi-stream). The selfbot lib's `joinChannel`
reuses `ClientVoiceManager.connection` (it's a singleton) — see caveat below.
### Option B: Migrate primary voice to selfbot lib
Make the selfbot `VoiceConnection` THE voice layer (it also does audio via `receiver.createStream`).
- Pros: one connection; video+audio unified.
- Cons: large refactor; high regression risk to the entire existing audio/music/GoLive stack. NOT chosen now.
## CAVEAT — ClientVoiceManager.connection is a singleton
`ClientVoiceManager.connection` is a single `VoiceConnection`. The gateway's `@discordjs/voice` adapter and
the selfbot lib both drive the same client voice state. Need to verify whether `client.voice.joinChannel()`
can coexist with the active `@discordjs/voice` session, or whether we must create the selfbot VoiceConnection
manually / re-use the existing voice state. This is the #1 technical risk to validate in the spike before
committing to Option A.
## Implementation plan (Option A)
### 1. Streamer detector (new: `modules/voice-recording/videoRecorder.ts`)
- Listen to voice state updates (`client.on('voiceStateUpdate')` or the existing voice-state hook).
- When `voiceState.streaming === true` for a member in the bot's channel → candidate to record.
- Skip bot's own user id (unless we also want self-video; default skip).
### 2. Watch + record wiring
- Ensure a selfbot-lib `VoiceConnection` exists for the channel (spike: `client.voice.joinChannel(channel)`,
fallback: build a `VoiceConnection` directly from the existing voice auth).
- `await selfbotVoiceConn.joinStreamConnection(userId)` → STREAM_WATCH op 20.
- `const recorder = selfbotVoiceConn.receiver.createVideoStream(userId, outPath)` where outPath points under
`<RECORDINGS_DIR>/<uid>/video-<streamKey>-<ts>.mkv` (Recorder outputs MKV natively).
- On `recorder.on('ready')` → mark recording; `recorder.on('closed')` → finalize.
- Transcript later: MKV → mp4 via ffmpeg (Phase B `muxToMp4` can accept mkv) for dashboard playback.
### 3. Teardown
- When `voiceState.streaming === false` / user leaves / channel emptied → `recorder.destroy()`,
`selfbotVoiceConn.streamWatchConnection.delete(userId)` / `sendStopScreenshare()`.
### 4. Frontend (Phase UI, later)
- oRPC/backend list `.mkv` per call session + FE `<video>` player (mirror audio recordings UI).
## Files touched
- `services/discord-gateway/src/modules/voice-recording/videoRecorder.ts` (new)
- `services/discord-gateway/src/modules/voice-recording/recorder.ts` (wire streamer detector on voice join)
- Possibly `voiceController.ts` (voice state update subscription)
- Tests: `tests/videoRecorder.test.ts` (mock selfbot VoiceConnection + Recorder)
## Verification
1. `pnpm typecheck` + `pnpm build` + biome clean in discord-gateway.
2. Unit: Recorder wiring + streamer detection with mocked VoiceConnection.
3. Live (deploy): user shares screen → journal shows `STREAM_WATCH` sent + `Recorder ready` + `.mkv` file
appears under recordings dir; playable via ffmpeg.
4. CI Build & Deploy (Nix) green.
## Open questions for spike (before full build)
- [ ] Can `client.voice.joinChannel()` run alongside the active `@discordjs/voice` session, or does the
singleton `ClientVoiceManager.connection` collide / tear down the existing audio connection?
- [ ] Does the selfbot `VoiceConnection` need the bot's `video: true` flag in IDENTIFY to receive video
(it advertises `streams` in IDENTIFY — see BaseMediaConnection/identify vs selfbot VoiceConnection)?
- [ ] Does `Recorder` (spawns system ffmpeg, UDP loopback on 65506/65510) work in the Nix store runtime
(ffmpeg-headless on PATH confirmed; UDP loopback fine)?
@@ -0,0 +1,33 @@
# Video Recording Splitting — Like Voice Recording
## Goal
Camera + screen share (stream watch) recording should split into per-burst
segments just like voice recording does — each time a streamer pauses/stops
and resumes, a new MP4 segment is created and registered in the DB + uploaded.
## Voice Recording Model (to replicate)
1. `receiver.speaking.start` → new OGG segment per burst
2. AfterSilence (4000ms) → stream "end" → segment finalized + uploaded
3. Each segment → DB insert → OGG→MP3 transcode → upload → update DB
4. File stored as `<userId>/<startTime>.ogg` + `.json`
## Video Recording Splitting
1. DAVE video RTP → depacketize H264 → write to current segment .h264
2. Silence detection: no H264 packets for 4000ms → close segment → flush →
mux to MP4 → insert DB record → upload → start new segment on next packet
3. Each segment: `<userId>/video-<channelId>-<startTime>.h264` → `.mp4`
4. DB: reuse `voice_recordings` table (filename indicates video, e.g. `video-XXX-1234.mp4`)
5. Upload: MP4 to TeleUploader (no transcode needed — MP4 plays everywhere)
## Files Modified
- `services/discord-gateway/src/modules/voice-recording/streamWatchReceiver.ts`
— Main change: silence-based splitting + DB registration + upload
## Constants
- `VIDEO_SILENCE_MS = 4000` (matches voice AfterSilence)
- `VIDEO_MIN_SEGMENT_MS = 1000` (skip segments <1s — avoid noise)
## Verification
- `pnpm typecheck` in `services/discord-gateway`
- `pnpm build` (dist/ is the deployed artifact)
- Push → CI deploy → live test with a streamer
@@ -0,0 +1,81 @@
# Spec: Selfbot-Viable Video Capture — manual screen-share watch command (Phase D)
Status: PLANNED (not yet built)
Date: 2026-09-02
Author: Hermes
Related: `.hermes/plans/2026-08-31_video-receive-phaseC-spec.md` (auto-receive, superseded
for selfbot), `gmw-ops/references/selfbot-presence-detection-limits.md`,
`gmw-ops/references/discord-voice-fork-video-receive.md`
## TL;DR — the decisive finding (verified live 2026-09-02)
User insists on keeping the **selfbot** (no bot-token migration). Live diagnostics prove
a selfbot CANNOT auto-detect other members' camera/share because:
- It never receives `VOICE_STATE_UPDATE` for other members (only its own).
- `guild.members.fetch()` → 403, `GET /channels/{id}/voice-states` → 404.
- No `GUILD_CREATE`, no `READY.broadcaster_user_ids` presence.
- `scanExistingStreamers` + `handleVoiceStateUpdate` (the only two `startStreamWatch`
triggers) are therefore both **dead on a selfbot**.
- No manual watch command exists today, so even on-demand capture is impossible.
→ The ONE selfbot-viable path is a **manual, operator-initiated STREAM_WATCH** on a
member known to be screen-sharing. Gateway op 20 (STREAM_WATCH) is **NOT gated on
bot-vs-user**; the DAVE handshake to Ready+MLS was already verified live in earlier
sessions. The receive/mux/segment/upload pipeline (`streamWatchReceiver.ts`) is already
built and only lacks a real streamer to produce its first `.mp4`.
Camera-of-others is NOT viable on a selfbot even with `unknown-ssrc` fallback:
`@discordjs/voice` `parsePacket` calls `daveSession.decrypt(packet, userId)` keyed per
REAL userId (vendor fork dist/index.js:2143), so a fake id selects no MLS decryptor →
garbage, not H264. (The uncommitted `unknown-ssrc` change was reverted this session.)
Selfbot CAN capture the OWNER's own video (its own VOICE_STATE_UPDATE + fork op12
videoSSRC are attributable), but `videoRecorder.ts` hard-skips its own id — parameterized
self-capture is a follow-up, not the default.
## Goal
Add a **manual watch command** so an operator can say "record <member>'s screen share"
and the gateway `startStreamWatch`s that member → DAVE watch → per-burst `.mp4` segments
(mirroring voice silence split) → upload → DB `voice_recordings` → dashboard `<video>`.
This is the only form of OTHER-member video capture a selfbot can deliver, and it is
genuinely buildable with the existing receive pipeline.
## Scope / files
Gateway (`services/discord-gateway`):
- New command type `VIDEO_WATCH` + handler in `command-handler/` (dedicated
`video.handler.ts`), routed via `createHandlerRegistry`.
- Handler resolves a VoiceChannel (from persisted `voice_auto_reconnect` / active
connections) + target memberId from the command payload, calls
`startStreamWatch(channel, memberId)` (already exported).
- Idempotent (startStreamWatch early-returns if a watch exists); a `VIDEO_UNWATCH`
command calls `stopStreamWatch(guildId, userId)`.
- Reply: success/failure via the standard `CommandReply` publish.
Backend (`services/backend`):
- oRPC procedure (or the existing command bridge) that publishes a `VIDEO_WATCH`
command to `backend:command` with `{ guildId, channelId, userId }`. Reuse the same
bridge the FE already uses for voice commands.
Frontend (`services/frontend`):
- A "Video Watch" control: pick a voice member + a "Record screen" button → calls the
backend procedure. Shows live status (watching / recording / segments uploaded).
(Each layer optional independently; gateway alone gives a Redis-testable path.)
## Verification
1. `pnpm typecheck` + `pnpm build` + `biome check src/` green in discord-gateway.
2. Unit test: handler publishes reply + calls startStreamWatch with the right args
(mock the module).
3. Live: operator invokes `!videorec <member>` while that member screen-shares →
journal shows `Sending STREAM_WATCH` → `STREAM_CREATE` → `DAVE watch READY` → `Video
burst opened` → `Video muxed to mp4` → a `video-*.mp4` appears under
`<recordingsDir>/<uid>/` and a `video-%` row lands in `voice_recordings`.
4. `Build & Deploy (Nix)` CI green.
## Out of scope (documented dead ends on selfbot)
- Auto camera/share capture of OTHER members (impossible at detection layer).
- Camera-of-others via `unknown-ssrc` (DAVE decrypt needs real userId).
- Bot-token migration (user declined).
+61
View File
@@ -0,0 +1,61 @@
# GMW — Fix mobile navbar "tidak bisa pindah halaman" (root cause)
## Context / symptom
- User (phone, mobile viewport <md): pressing the bottom mobile nav items does NOT
navigate ("Masih sama, walau sudah ditekan tidak pindah halaman").
- Previous fix (commits ae41f64e, 4d0bbed5) repositioned the chatbot FAB above the
nav and expanded nav to all 7 items — but the bug persisted.
## Root cause (verified live 2026-08-28)
Reproduced on the LIVE site (`imphnen.asepharyana.my.id`) with Playwright @375px:
1. **Next `<Link>` client-side navigation is dead app-wide.** Clicking a mobile nav
`<Link>` fires the click (event reaches document, `inLink=true`,
`defaultPrevented=false`) but **the URL never changes**. `window.next.router.push('/voice')`
also does nothing. The desktop NavRail navigates only because it uses a **plain
`<a href>`** (full browser navigation bypasses the broken router).
2. **A React hydration failure (#418 "server rendered text didn't match the client")**
is thrown on live pages — the only console error. Cause: relative-time text
(`formatRelativeTime(e.edited_at)` / `formatRelativeTime(m.created_at)` in the
SSR-seeded messages/edit-history feed uses `Date.now()`; server and client
render slightly different text → hydration mismatch → React re-renders, and the
Next client router ends up non-functional.
## Why desktop "worked" / mobile didn't
- `NavRail` = plain `<a href>` → **hard** navigation (works).
- `MobileNav` = Next `<Link>` → **client** navigation → dead router.
User's instruction "ikuti cara kerja sidebar" = make mobile nav behave like the
sidebar (plain anchors).
## Changes
### 1. `src/components/shell/mobile-nav.tsx`
Replace Next `<Link>` with a **plain `<a href>`** (mirrors NavRail). Keep:
`usePathname` for active-state styling + `scrollIntoView` for snap-to-active.
Drop the now-unused `Link` import (biome import-order — run biome check --write).
### 2. Hydration root cause — `formatRelativeTime` in SSR-seeded feeds
Add `suppressHydrationWarning` to the timestamp elements in the SSR-seeded
components that render live-relative text so server/client text drift no longer
throws #418:
- `src/components/EditHistory.tsx` (line ~111 `edited {formatRelativeTime(...)}`)
- `src/app/(dashboard)/messages/view.tsx` MessageRow channel/time (line ~580) +
semantic row (line ~364)
- `src/components/LiveModerationFeed.tsx` (line ~139)
- (recordings/view.tsx, TermGlossary, ChannelCultureGlossary, CategoryDrilldown,
chatbot — chatbot is client-after-mount; add where SSR-seeded.)
NOTE: `suppressHydrationWarning` is safe for these single-text spans; the values
are cosmetic and self-correct on the next interval/render.
## Verification (non-prod, NEVER 4017)
- `pnpm format` + `pnpm lint` (biome) + `pnpm build` clean.
- Smoke on :4024 (fresh `next dev`): mobile click on nav item actually navigates
(URL changes) AND no React #418 in console.
- After push: `gh run watch`, then live `imphnen.asepharyana.my.id` @375px: nav tap
navigates, console free of #418.
## Files
- services/frontend/src/components/shell/mobile-nav.tsx
- services/frontend/src/components/EditHistory.tsx
- services/frontend/src/app/(dashboard)/messages/view.tsx
- services/frontend/src/components/LiveModerationFeed.tsx
- (others only if #418 persists)