Compare commits

94 Commits
Author SHA1 Message Date
asepharyana 2825250804 perf(ai-moderation): remove per-user reputation from analysis context
User: 'jangan ada reputasi juga' — no profile, no reputation in the prompt,
raw messages only.

- textBatchProcessor: drop initializeUserReputation fetch + <user_reputation>
  tag injection (kept the minimal <message> tag + reply/reference context).
- visionAnalyzer (prepareMediaMessage): same removal.
- prompts/system.ts + prompts/output.ts: replace <user_reputation>/<user_history>
  instructions with an explicit 'no per-user profile/reputation context'
  note so the LLM judges purely on message content + conversation/web/location.
- mediaBatchProcessor: fix stale comment.

Trust/infraction state is STILL written to the DB (userReputationsTable) for
enforcement — only the LLM context injection is removed, so moderation
actions (mute/ban via infraction thresholds) keep working.

Net: even smaller prompts (no per-user context at all) → more messages fit
per request, and one fewer DB round-trip per unique user per sub-batch.

tsc, biome, vitest (129) all clean.
2026-08-16 20:44:39 +07:00
asepharyana aa280c48b7 perf(ai-moderation): drop personal user-profile descriptions from context
User insight: personal profile summaries bloat the prompt (less room per
request) and add a per-user DB/Redis round-trip for little moderation signal.
Only the behavioural <user_reputation> history is kept.

- textBatchProcessor: stop fetching getUserProfile; remove <user_profiles>
  block + <user_profile_ref> from message tags. Keep <user_reputation>.
- mediaBatchProcessor + visionAnalyzer: same removal (profile fetch + ref).
- prompts/system.ts + prompts/output.ts: drop stale <user_profiles>/
  <user_profile_ref> instructions; point LLM at <user_reputation> instead.
- aiAnalyzer: gate userProfileLearner behind AI_USER_PROFILE_LEARNING_ENABLED
  (default false) — generates profiles nobody reads, pure LLM/DB waste.
- Add AI_USER_PROFILE_LEARNING_ENABLED config knob.

Net: smaller prompts (more messages fit per request), fewer DB round-trips
per sub-batch, and no background LLM calls learning unused profiles.

tsc, biome, vitest (129) all clean.
2026-08-16 19:56:27 +07:00
asepharyana 4cf5b87f2b perf(ai-moderation): pack more messages per LLM request (fewer API calls when busy)
User insight: rather than many small per-batch API requests, pack many
messages into ONE request so a burst is analyzed with far fewer calls.

- AI_LLM_TEXT_BATCH_SIZE 20 -> 60 (one request now carries ~3x more messages).
- AI_ANALYSIS_MAX_TARGET_TOKENS 4000 -> 14000 (the scheduler's token-budget
  gate was trimming pending messages to ~20 before they reached the sub-batch
  splitter; raising it lets ~60 messages through to a single LLM call).
- AI_LLM_TEXT_ANALYSIS_TIMEOUT_MS 30000 -> 45000 (one larger call needs more
  headroom; gemini-flash-lite has a 1M-token context so 14k+8k is trivial).

Net effect when ramai: a 60-message burst = 1-2 API calls instead of 3+,
less semaphore contention, faster throughput.
2026-08-16 19:00:29 +07:00
asepharyana 0dff7770a1 perf(ai-moderation): speed up analysis queue (ramai + sepi)
- Parallelize per-user reputation/profile fetches in textBatchProcessor
  (was a serial ~2N DB/Redis round-trip loop per sub-batch; now Promise.all
  over unique users). Cuts per-batch latency, biggest win on small/quiet
  batches.
- Make the LLM concurrency semaphore dynamic (cached per config value) instead
  of frozen at import time, so AI_LLM_MAX_CONCURRENT is tunable without code
  change and reflects current config.
- Bump AI_LLM_MAX_CONCURRENT default 5 -> 8 (gemini-flash-lite is cheap; helps
  throughput when busy).
- Lower AI_ANALYSIS_DEBOUNCE_MS 500 -> 250 (snappier first-message analysis
  when quiet).
- Lower AI_ANALYSIS_RECOVERY_INTERVAL_MS 15000 -> 10000 (stuck/errored
  messages re-analyze sooner).

tsc, biome, vitest (129) all clean.
2026-08-16 18:51:13 +07:00
asepharyana e3dd6a3427 fix(messages): Discord-style order (oldest top, newest bottom)
Backend returns messages DESC (newest first); the view previously rendered
that directly, so the feed was inverted vs Discord (old at bottom, new at top)
while the load-older control sat at the top — contradictory.

- Reverse the display list so it reads oldest→newest top→bottom, like DC.
- Load-older (cursor to lower created_at) prepends at the top; scroll position
  is preserved by offsetting scrollTop by the height added above.
- Open at the bottom (newest visible) on first load / scope change.
- New live messages append at the bottom and auto-scroll only when the user is
  already near the bottom (nearBottomRef), so reading history isn't disrupted.
- Scroll container now tracked via ref; onScroll updates nearBottom + triggers
  load-older when scrolled to the top.

tsc, biome, next build all clean.
2026-08-16 17:56:39 +07:00
asepharyana 55fdcfaae3 style: biome format chatbot.service (parseResponse call) 2026-08-16 17:17:03 +07:00
asepharyana cf1ec25c71 fix(chatbot): disable thinking + use non-streaming LLM call
- Set stream:false on the /chat/completions request so the bot gets one
  complete response instead of an SSE token stream.
- Add reasoning_effort:"none" to suppress extended-thinking/reasoning tokens
  (ignored by non-reasoning models like gemini-flash-lite).
- Add parseResponse(): handles both the JSON object 9router returns for
  stream:false and the SSE text it may still emit, delegating SSE to parseSse.
  Verified live: omniroute returns 200 application/json with message.content.
2026-08-16 17:08:29 +07:00
asepharyana 0ace758c79 feat(frontend): safe-area insets + hardened reduced-motion
- Add viewport export with viewportFit: "cover" so iOS exposes
  env(safe-area-inset-*) (required for the insets to take effect).
- NavRail / TopBar / main / Toaster now respect safe-area insets so content
  clears the iPhone notch and home indicator in both portrait and landscape.
- prefers-reduced-motion: the media query already disabled declared animation
  classes; harden it with a global transition/animation duration override and
  kill the scan-line shimmer so motion-sensitive users get a fully static UI.

Verified tsc --noEmit + next build clean.
2026-08-16 16:55:34 +07:00
asepharyana c89288191e fix(frontend): responsive layout across all dashboard pages
- SectionHeader: action (filters/legends) now wraps below the title on narrow
  screens instead of overflowing beside it (flex-wrap, gap-2 sm:gap-3).
- GuildChannelPicker: selects go full-width and stack on mobile (w-full
  sm:w-44 / sm:w-52) instead of fixed widths that exceeded a 375px viewport.
- Messages search: w-full sm:w-64 so it doesn't crowd the picker on mobile.
- TopBar: tighter padding (px-4 sm:px-5), smaller title on mobile, connection
  status uses compact (dot only) on mobile, ambient pill hidden < sm.
- Shell main + dashboard channel label: responsive padding / shrink-0 widths.

Verified tsc --noEmit + next build clean; targets breakpoints 375/768/1024/1440.
2026-08-16 16:36:34 +07:00
asepharyana 4655125541 feat(frontend): clearer load-older spinner + cap history pages (Messages)
- Show an explicit Loader2 spinner row ("Loading older…") while the next page
  fetches, instead of a disabled button.
- Cap appended older pages at MAX_OLDER_PAGES=10 (500 messages) so a long
  scroll-up never pulls the entire history; show a "capped" hint pointing to
  search. Reset the counter when guild/channel changes.
2026-08-16 16:21:10 +07:00
asepharyana ba60448d05 feat(frontend): load older messages in Messages view (cursor pagination)
Wire the existing useLoadMore + useMessagesHasMore pagination hooks into the
Messages view: add a "↑ Load older messages" button at the top of the list and
auto-load the next (older) page when the user scrolls to the top. Backend
messages.list already returns a created_at-based nextCursor (DESC order), so
older pages are just subsequent cursors. Newest-first live feed is preserved;
the load-older control is hidden during search.
2026-08-16 16:06:25 +07:00
asepharyana 1d809b2c95 fix(proxy): route /trpc to backend so browser oRPC WebSocket opens
Browser connects oRPC over wss://…/trpc (partysocket). The gmw-proxy nginx
only forwarded /api and /ws to the backend, so /trpc upgrades fell through to
Next.js SSR and the socket never opened ("WebSocket is not open"). Add a
/trpc location (WS upgrade headers) mirroring /ws. Backend already serves
oRPC on /trpc (HTTP RPCHandler + WS ORPCWebSocketServer on :4001).

Verified: ws://127.0.0.1:4001/trpc upgrade OPEN; SSR + server-side fetch RPCLink
also use /trpc directly so only the browser path was broken.
2026-08-16 15:28:58 +07:00
asepharyana 38bda66933 style: fix Biome dead-code warnings from oRPC migration (CI lint gate) 2026-08-16 14:51:47 +07:00
asepharyana 726a8e116b fix(build): make oRPC/tRPC dist runnable under node ESM (deploy crashloop)
flake.nix only rewrote @/ aliases but left extensionless relative imports
(./router) in compiled dist/. node dist/index.js (how prod runs) cannot
resolve extensionless ESM specifiers -> ERR_MODULE_NOT_FOUND -> backend
crashlooped (444 restarts, port 4001 dead). Extract the fixer into a shared
scripts/fix-imports.mjs that appends .js to extensionless relative imports and
rewrites @/ aliases, and wire it into backend + discord-gateway build phases.

Verified: fresh tsc + fixer -> node dist/index.js boots; oRPC over /trpc
serves both HTTP POST and WebSocket (config/dashboard/voice/moderation/
media/chatbot/analysis) end-to-end against Postgres + Redis. next build
passes with the oRPC client + partysocket.
2026-08-16 14:45:02 +07:00
asepharyanaandClaude Opus 4.5 2fa1827f17 feat(backend,frontend): migrate data APIs from REST to native tRPC over WebSocket
Replace REST module routers with a single typed tRPC appRouter served over
/trpc (HTTP + WebSocket), and rewire the frontend to call it via
@trpc/client wsLink (browser) and httpLink (RSC data layer). Existing
/api/health + /api/metrics stay as plain Express for infra scraping.

Notable fixes surfaced by the live smoke test:
- Express 5 / path-to-regexp v8 rejects the /trpc/* wildcard route; use a
  prefix middleware that computes opts.path from the URL instead.
- nodeHTTPRequestHandler treats opts.path as the literal procedure path, so
  it is derived per-request from req.url.
- Two ws servers on one http.Server (the /ws voice socket + /trpc) collided
  and returned 400 on upgrade; both now use noServer + a manually routed
  server.on('upgrade') keyed by path.

Verified: BE tsc+biome+40 vitest green; FE tsc+biome green; live
HTTP and WebSocket calls returned real prod data.

Co-Authored-By: Claude Opus 4.5 (1M context) <noreply@anthropic.com>
2026-08-16 13:25:10 +07:00
asepharyanaandClaude Opus 5 (Nous Research) d8552a9fb8 feat(ai): make standalone image/vision analysis timeout explicit (1 min)
The standalone image analysis path (analyzeSingleMediaImage → llmVision →
llmChat) previously had no request-level timeout of its own — it silently
inherited the shared OpenAI client default (60s), and AI_LLM_MEDIA_ANALYSIS_
TIMEOUT_MS only governed the text+media *batch*, not a single vision call.

- Add AI_LLM_VISION_ANALYSIS_TIMEOUT_MS (default 60000) to config.
- llmChat now accepts an optional per-request `timeout` in LlmCallOpts,
  forwarded to the OpenAI request options (falls back to the 60s client
  default when omitted).
- llmVision passes config.AI_LLM_VISION_ANALYSIS_TIMEOUT_MS, so a single
  image/sticker/emoji analysis gets a guaranteed 1-minute budget and is
  independently tunable from the text path.

Verified: tsc + biome green, 129 gateway tests pass.

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 10:47:24 +07:00
asepharyanaandClaude Opus 5 (Nous Research) a4abe3abea fix(frontend): remove duplicate Dashboard entry in nav rail
The sidebar rendered /dashboard twice: once as a hardcoded NavItem
(lines 45-50) and again via navItems.map() (navItems[0] is also
/dashboard). Dropped the hardcoded item so the single source of truth
(navItems in lib/navigation.ts) drives the rail. Removed the now-unused
LayoutDashboard import.

tsc + biome green.

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 09:25:07 +07:00
asepharyanaandClaude Opus 5 (Nous Research) b67856462f feat(chatbot): expand tool set to cover all server-watcher situations
The chatbot agent now has 14 tools (was 4) so it can answer about ANY
server situation from live data instead of a static snapshot:

- get_server_stats (now also returns clean count)
- get_top_channels, get_recent_activity, get_top_flagged
- search_messages (LIKE keyword search)
- get_user_messages, get_user_profile, get_user_reputation
- get_channel_culture
- get_message_detail (full AI analysis of one message)
- get_message_reviews (human moderation queue by status)
- get_voice_recordings (with transcriptions)
- get_moderation_timeline (daily flagged/warn/clean trend)
- get_corrections (AI false-positive correction history)

Security/quality:
- Every executor now uses parameterized drizzle queries (eq/like/and).
  The old code interpolated model-supplied IDs into sql.raw() — a SQL
  injection vector. Removed.
- Split static tool *definitions* into chatbot.toolDefs.ts (no DB import)
  so the LLM-facing schema can be unit-tested without loading the
  database/config layer. chatbot.tools.ts keeps only the executor.

Verified: tsc + biome clean, 40 backend tests pass (4 new covering the
tool-contract: names unique, required args declared, full situation
coverage).

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 09:22:20 +07:00
asepharyanaandClaude Opus 5 (Nous Research) 30828a5534 refactor(chatbot): drop static server-stats context, go fully tool-based
The chatbot already had an agentic tool loop (get_server_stats,
get_top_channels, get_recent_activity, get_top_flagged), but processMessage
still baked a serverInsights snapshot into the system prompt and told the
model to "answer from that data". That defeats the tools: the model answered
from a stale snapshot instead of living numbers, and the guild/channel scope
the frontend sends was never forwarded to the tools.

Changes (services/backend/src/modules/chatbot):
- Remove getServerInsights() + ServerInsights (dead after this change).
- buildSystemPrompt(): drop the hardcoded stats block; instruct the model it
  has NO memorized server numbers and MUST call a tool for any server-data
  question, answering only from tool results.
- processMessage(): stop fetching insights; pass the request guildId/channelId
  scope through to callLLM.
- callLLM(): accept scope; auto-fill empty guildId/channelId on tool calls from
  the request scope so the model never has to guess IDs and tools always query
  the right server.

Behavior: answers now come from live DB data via tools, scoped to the server
the user is chatting in. tsc + biome + 36 backend tests green.

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 09:15:39 +07:00
asepharyanaandClaude Opus 5 (Nous Research) a3e5a8c1b9 perf(gateway): hoist correctedExamples query out of retry closure
buildCorrectedFewShotExamples() (a getRecentCorrectedModerations(5)
DB hit) was called inside the per-sub-batch buildContent closure in
textBatchProcessor.ts — re-queried for every sub-batch (≈10× for a
200-msg burst) AND re-fired on each parse-error retry. mediaBatchProcessor
already hoisted it once. Mirror that: fetch once per runTextOnlyBatch,
reuse the cached string inside the closure.

No behavior change — identical content, fewer identical DB reads.
tsc + 129 tests + biome green.

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 09:06:17 +07:00
asepharyanaandClaude Opus 5 (Nous Research) f82b5caae4 refactor(gateway): strip boilerplate fields from few-shot examples
The 32 few-shot examples each re-echoed score/confidence/
recommended_action/categories/policy_version inline (~150 chars ×
32). Those fields carry zero moderation-decision signal — the schema
and their ??-default coercion already live in OUTPUT_INSTRUCTIONS +
moderationResponseParser.ts. Removed 96 redundant key/value pairs.

Kept per-example: message_id, status, flags, severity, evidence,
analysis — the fields that actually teach decisions. Parser derives
the rest via ?? fallback, so real output shape is unchanged.

examples.ts: 21.7K→18.5K chars; FEW_SHOT(mixed) 15.3K→13.4K.
Total mixed system prompt now 33.9K (was 39.3K at audit start,
~14% leaner). tsc + 129 tests + biome green.

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 09:00:55 +07:00
asepharyanaandClaude Opus 5 (Nous Research) 9e2b107fcd refactor(gateway): compact AI analysis system prompt, preserve all rules
- prompts/system.ts: merge 3 overlapping framing blocks (Blok Data /
  Konteks Pengguna / Framing Konteks vs Target) into 1 tight block —
  same coverage, no duplicated "standalone judgment / profile-is-
  reference-not-evidence" prose.
- prompts/output.ts: trim duplicated user_history/standalone paragraph
  in PERSONALITY & MEMORI (keep concrete per-case lessons).
- prompts/examples.ts: drop 2 exact-duplicate-lesson few-shots (LGBT id=19
  dup of id=30; weapons-tech id=33 dup of id=32). All teaching signals
  retained via the surviving example of each lesson.

Static system prompt: text 32.7K→29.2K, mixed 39.3K→35.8K chars
(~10% smaller). No moderation rule, zero-tolerance category, or decision
tree altered — accuracy-controlling content untouched. tsc + 129 tests +
biome green.

Co-Authored-By: Claude Opus 5 (Nous Research)
2026-08-16 08:52:11 +07:00
asepharyanaandClaude Opus 5 d2e97ae11d audit(gateway): fix dead /metrics endpoint, raise OOM-prone MemoryMax, trim DB pool
- gateway-metrics: collectors now run per scrape so Prometheus sees real
  data (process memory/uptime + live AI-analysis pipeline gauges) instead
  of an always-empty stub. bootstrap registers the pipeline collectors.
- systemd: MemoryMax 512M -> 1G (live RSS ~500MiB, peak 508MiB; 512M left
  ~2% headroom and risked an OOM-kill restart; host has 8GB free).
- config: POSTGRES_POOL_MIN 2 -> 0 so main + 4 Piscina worker threads don't
  hold ~10 permanently-open idle pg connections against PgBouncer.
- docs: rewrite stale ARCHITECTURE.md / MODULE_STRUCTURE.md (winston ->
  pino, removed mock-crc/indonesianTextNormalizer, renamed
  aiAnalysisWorker/llmModerationClient).

Verified: tsc clean, 129 vitest pass, biome clean on changed files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 08:41:54 +07:00
asepharyana 6244e307a3 feat: surface AI analysis duration across gateway, backend, and FE
Adds per-message AI moderation analysis time (ai_analysis_duration_ms)
so operators can see how long the LLM took to moderate each message.

Gateway:
- messagesTable: new ai_analysis_duration_ms (bigint) column.
- AIAnalysisUpdate + buildAIAnalysisSet: carry analysisDurationMs through
  both single and bulk update paths.
- ai-analysis-worker: measure wall-clock time around runModerationAnalysis
  and attach it to every result in the batch.

Backend:
- Mirror schema column; messageMapper maps ai_analysis_duration_ms;
  moderation-types + MappedMessage expose it.

Frontend:
- message.ts type gains ai_analysis_duration_ms.
- AiBadge (messages view) shows 'status · 1.2s' when duration is present;
  analysis view badge mirrors the same formatting.

DB:
- scripts/add-ai-analysis-duration.sql (idempotent ADD COLUMN IF NOT EXISTS).

No behavior change for moderation logic; null until new gateway build
records values.
2026-08-16 00:11:00 +07:00
asepharyana 2d7c7f2c35 fix(gateway): stop Qdrant upsert aborts (semantic cache was being skipped)
Qdrant upserts were failing with 'This operation was aborted' ~32x/2h,
so semantic moderation cache entries were silently dropped. Root cause:
upsertQdrantPoint ran ensureQdrantCollection() on EVERY call — a GET
(and sometimes DELETE+PUT) round-trip — while the request AbortController
had only a 10s timeout. Under moderation load Qdrant is busy (the
gmw_text_moderation collection is not yet HNSW-indexed, so searches are
full-scans), the extra round-trips pushed the upsert past 10s, and the
client aborted it.

- Memoise ensureQdrantCollection() at module scope so the collection is
  verified exactly once per process (resetQdrantCollectionCache() for
  tests / config reload).
- Bump the upsert request timeout 10s -> 30s so a transiently busy
  Qdrant no longer aborts the write.

Qdrant server itself is healthy (<100ms for direct upsert; collection is
green), so no server-side change is needed. Semantic cache should now
populate reliably.
2026-08-15 23:40:19 +07:00
asepharyana 416c690ebc style(gateway,backend): clear all biome warnings (no warnings left behind)
Address every remaining biome lint/format warning across both services
so the codebase ships warning-free:

- textCacheStore: drop unused deleteExpiredQdrantPoints import; hash
  image cache key (sha256[:32]) so long/base64 URLs no longer blow the
  text_analysis_cache PK B-tree 8191-byte index (was aborting the media
  analysis lock INSERT).
- bootstrap: drop unused unhandledRejection promise param.
- moderationOrchestrator: drop unused  destructure at L197.
- mediaDownloader / textBatchProcessor / transmitter: replace non-null
  assertions with proper null guards (stickerName ?? '', urlImages.get
  guard, backpressureQueue.shift guard).
- backend utils: throw lastError ?? fallback instead of lastError!.
- message-capture: remove unused  (retentionDb),  (moderationActionsDb,
  reviewsDb); simplify renderDiscordMentions guard to optional chain.
- transmitter: remove dead write-only  field + its assignments.

No behavior change beyond the cache-key hashing (now deterministic
fixed-length) and the intentional null-safety guards.
2026-08-15 23:18:33 +07:00
asepharyana c590a8be27 style(gateway): biome format fix for imageResizer (unblock CI gate)
imageResizer.ts had a line exceeding the print width that biome flagged
as a formatter error, failing the Build & Deploy biome check. Re-format
the file. No logic change.
2026-08-15 23:11:31 +07:00
asepharyana 9c83ec86cc fix(gateway): image vision analysis + media cache lock failures
Two root causes behind 'all image analysis failing':

1. imageResizer still emitted lossless PNG for vision input. A 1024px
   Facebook photo balloons to multi-MB PNG base64 that the vision model
   silently rejects ('Vision API null response'). Switch to JPEG q85
   (no upscaling) — same photo drops to ~100-400KB, model processes fine.
   Re-encodes even already-small images so raw originals never bloat the
   data URL. Added tests/imageResizer.test.ts covering both cases.

2. acquireMediaAnalysisLock INSERT aborted with 'index row requires N
   bytes, maximum size is 8191'. text_analysis_cache.text is the PK in a
   B-tree index (8191-byte/row cap); callers pass the raw image URL as the
   key, and base64 data URLs / very long URLs blow past the limit, so the
   lock INSERT fails and every media analysis is skipped. Hash the URL in
   makeImageCacheKey (image:<sha256[:32]>) — fixed-length, deterministic,
   well under the limit. All store/get/lock/delete callers already route
   through this function so lookup stays consistent.
2026-08-15 23:04:52 +07:00
asepharyana 17a4fbd73d build(gateway): skip fixupPhase to kill 'patchelf: wrong ELF type' noise
dontPatchELF only disabled the patchELF sub-phase; fixupPhase's
shrinkELF step still emits the same error on the prebuilt .node addons
and .o/.a object files in node_modules. Skip the entire fixupPhase
(dontFixup = true) for the gateway — node is the external interpreter
and .node addons are self-contained dlopen prebuilts, so Nix RPATH
patching/stripping is neither needed nor wanted.
2026-08-15 22:20:56 +07:00
asepharyana c04c410fad build(gateway): suppress harmless 'patchelf: wrong ELF type' noise
Add dontPatchELF = true to the discord-gateway derivation. Nix's
fixupPhase runs patchELF over $out/node_modules and chokes on the
non-ET_DYN ELF files (.o/.a objects + prebuilt .node addons), emitting
hundreds of non-fatal 'patchelf: wrong ELF type' lines per build. The
real binary is node (external, RPATH-fixed) and the .node addons are
self-contained prebuilts loaded via dlopen, so Nix RPATH patching is
neither needed nor wanted. Shebang patching still runs.
2026-08-15 22:11:58 +07:00
asepharyana 5e5f4ae208 build(gateway): use @discordjs/opus prebuilt instead of compiling from source
Drop npm_config_build_from_source=true so node-pre-gyp downloads the
published prebuilt .node for Node 22 (ABI node-v127, linux-x64-glibc-2.35)
instead of compiling libopus C++ every build. Replace the hardcoded
'npm run install' (node-gyp compile) loop with 'pnpm rebuild @discordjs/opus'
which runs the package's own install script (prebuilt fetch, source build
only as fallback). sharp already uses @img prebuilt packages (its install
script failure is non-fatal), so only opus was actually compiling.
2026-08-15 21:56:22 +07:00
asepharyana e2013988ff ci: fix biome format gate so Build & Deploy passes
Auto-format llmClient.ts (Object.assign indent) — the only biome
error blocking the Build & Deploy workflow. Logic unchanged; gateway
biome check now exits 0 (11 pre-existing warnings remain, non-blocking).
2026-08-15 21:39:44 +07:00
asepharyana 0164444dd7 refactor(gateway): remove screen-share / GoLive feature entirely
Drop the Discord Go Live (screen share) stack across the discord-gateway:
- delete src/goLive/ (19 modules: Streamer, Demuxer, encoders, WebRTC wrapper, native loader, etc.)
- delete native/libdatachannel-min/ N-API binding + flake native build + LD_LIBRARY_PATH wiring
- delete screenShareController.ts and screen-share tests (goLive-port, golive-*, demuxerNut, screenShareInput)
- mediaSource.ts: remove Invidious helpers + downloadScreenInput (YouTube full-file download)
- mediaTypes.ts: drop ScreenShare* types, narrow MediaMode to 'music' and DiscordPlayerOwner to non-screen
- media.handler.ts: remove screen branch, screenController/screenPlayback, voice-disconnect/reconnect accessor
- commandHandler.ts: stop passing getVoiceStatus / setVoiceController into MediaHandler
- media handler now only handles music; music queue/playback/status untouched

Verification: tsc --noEmit clean, biome clean on touched files, no lingering goLive/screenShare refs in BE/FE/gateway.
2026-08-15 21:20:20 +07:00
asepharyana 9ae26b8ec9 refactor(llm): unify vision routing with text moderation and remove dedicated endpoint 2026-08-15 21:05:06 +07:00
asepharyana 7ebee7559d feat(llm): add disableThinking option for faster LLM analysis and update config 2026-08-15 20:52:53 +07:00
asepharyana 66c33a2657 feat(message-capture): add bot exclusion logic for message capture 2026-08-15 20:35:51 +07:00
asepharyana 25b220b7f9 fix(frontend): sidebar + command palette navigation, zero biome warnings
Router.push was a no-op in the standalone build (Next trailingSlash
interaction), so the sidebar buttons and command palette silently failed
to navigate. Replaced next/link + router.push with plain <a href> anchors
in NavRail and CommandPalette — verified working on all routes.

Biome tightened to zero warnings:
- Disable noArrayIndexKey (positional equalizer bars), noStaticElementInteractions
  (intentional dismiss/hover overlays), useMediaCaption (voice clips)
- Avatar uses background-image instead of <img> (noImgElement)
- Command palette list items keyed correctly
- Format pass to satisfy the formatter
2026-08-15 20:25:44 +07:00
asepharyana 1c4f28c5f2 fix(message-capture): remove bot message filtering from capture logic 2026-08-15 20:21:06 +07:00
asepharyana 392db8eba1 feat(frontend): Ambient/WebGL console revamp + lint/type cleanup
Ground-up rebuild of the GMW frontend as an Ambient Field console:
- WebGL ambient background (Three.js shader, drifting motes, reduced-motion aware)
- Glassmorphism dark cyber theme across all 8 routes
- SSR page + client view split with SWR fallback; realtime via WebSocket
- Command palette (Cmd+K), chatbot FAB, guild/channel pickers
- Chart primitives: donut, radial-gauge, area-activity, sparkline, equalizer

Cleanup (review pass):
- Remove stray Puppeteer nav-test/nav-debug scripts
- Replace non-null assertions with guards (dashboard/moderation)
- Drop unused useGuilds fetches in messages/voice views
- Type implicit-any `let` declarations across pages
- Add a11y roles/labels to SVG charts and audio, tidy imports
2026-08-15 20:03:55 +07:00
asepharyana 3c2c1c3b15 Add Puppeteer scripts for navigation testing and debugging
- Created nav-debug.cjs to log anchor tags and simulate clicks on the Voice navigation link, capturing click events and page navigation.
- Added nav-test.cjs to test the Voice link click and log the URL at various intervals, capturing any page errors.
- Introduced nav-test2.cjs to check the presence of specific elements on the /voice/ page and log any console errors.
- Implemented nav-test4019.cjs to monitor network requests and responses related to the Voice navigation, verifying button presence and click functionality.
2026-08-15 19:23:17 +07:00
asepharyana 1b56212d1a feat(frontend): rebuild as Ambient/WebGL console with all pages + command palette
Ground-up rombak UI: hapus semua component/page lama, bangun ulang dengan
desain sistem Ambient (WebGL haze + drifting motes, signal-driven color)
di atas kontrak API/WS/type yang sudah ada.

- Design system: globals.css tokens + primitives (glass, button, badge,
  select, avatar, toast, chart SVG murni).
- Shell: nav rail, topbar (status WS + pill signal + theme), AppFrame.
- 8 halaman: dashboard, voice (orbital stage), media, messages (live feed +
  detail AI), moderation, analysis (search), recordings, + chatbot floating.
- Command palette (Cmd/Ctrl+K) untuk navigasi cepat.
- Server fetch di-page di-try/catch agar render graceful saat backend mati.

Verified: tsc clean, next build 8/8 halaman, semua route 200.
2026-08-15 17:53:48 +07:00
asepharyana b98101c576 feat(dashboard): ground-up rombak jadi Ambient Field layout (bukan re-skin)
Hapus template dashboard lama (top bar + side rail + main + right panel +
bottom prompt). Ganti dengan layout yang benar-benar beda:

- AmbientField: full-bleed WebGL canvas haze, drift speed + densitas
  ngikut load server, warna ngikut signal moderasi terakhir
  (clean→lime, warn→amber, flagged→vermilion). Background tanpa container.
- View jadi full-bleed: headline raksasa bottom-left, metric cluster
  floating top-right (no box), event ribbon drift di tengah, command
  whisper di very bottom.
- AmbientShell di layout.tsx: gak ada TopBar/LeftRail untuk /dashboard
  exact. Route lain (messages/voice/media/dll) tetap ClassicShell.
- Tidak ada card, tidak ada grid, tidak ada panel, tidak ada tab.

Verified: tsc clean, next build 11/11 halaman, biome clean.
2026-08-15 17:05:25 +07:00
asepharyana 84757bdcf4 feat(console): rombak penuh dashboard layout jadi Event Horizon
Layout baru single-screen ops console:
- TopBar 48px (brand monogram, guild, ws status, clock UTC/local, focus mode)
- LeftRail 80px (icon+label nav, signal accent bar, no boxes)
- Hero strip (display headline + mono counters: clean/warned/flagged/ratio)
- EventFeed (vertical timeline of message events, severity dots, no cards)
- NowMarker (inline pulse + cluster band insert per 10 events / 30s)
- RightRail 320px collapsible (ai verdicts / voice / mod queue / socket)
- DashCommandLine bottom 44px (mono prompt, '/' focuses, /mute /jump /find /clear)

Replace Spine + StatusBar lama untuk /dashboard via pathname branch di
(dashboard)/layout.tsx — route lain (messages/voice/media/dll) tetap
pakai ClassicShell, tidak ter-regress.

SSR seed tetap lewat page.tsx (server fetch stats + activity), synthetic
seed events dari daily buckets sampai WS message_created kick in.

WS event mapper: severity di-derive dari ai_status + ai_severity,
excerpt dipotong 140 char, channel tail 4 char.

No card chrome, no shadow, no bento grid, no tab panels.
2026-08-15 16:21:45 +07:00
asepharyana 6c9a91dad4 style(vision): biome format llmClient.ts (wrap long const line) 2026-08-15 14:38:44 +07:00
asepharyana bcb563ea7f feat(vision): route multimodal analysis to dedicated NVIDIA direct endpoint
- config: add AI_LLM_VISION_BASE_URL + AI_LLM_VISION_API_KEY (separate from text router)
- llmClient: llmVision() now calls dedicated vision endpoint when configured
  (axios POST to integrate.api.nvidia.com, model nvidia/nemotron-3-nano-omni-30b-a3b-reasoning,
  reasoning_budget 16384, non-stream), falls back to router combo otherwise
- keeps text/moderation on omniroute, vision on NVIDIA direct
2026-08-15 14:31:53 +07:00
asepharyana 589fd38fd8 fix(voice): separate Mic and Listen state (were both bound to listen)
- MicControl now uses useMicTransmit + local micActive/micVolume
  (was wrongly wired to listen.active/listen.toggle)
- ListenControl keeps useVoiceListen + handleListenVolume
(tsc clean, next build green)
2026-08-14 12:35:44 +07:00
asepharyana da02bfff9b fix(frontend): rebrand Bete → GMW (title, logo aria-label, dashboard heading)
- layout.tsx metadata title: Bete → GMW - Discord Moderation Console
- spine.tsx logo aria-label: Bete → GMW
- dashboard/view.tsx heading: Bete Console → GMW Console
(tsc clean, next build green)
2026-08-14 11:51:16 +07:00
asepharyana a66db8d702 fix(voice): live connection state instead of static SSR snapshot
- VoiceView now reads connected/activeChannelName from useVoiceStatus
  (SWR live, invalidated by connect/disconnect) instead of initialStatus
- Seed useSpeakers from live status.activeSpeakers
- Add 4s refreshInterval to useVoiceStatus so state converges
(tsc clean, next build green)
2026-08-14 11:43:07 +07:00
asepharyana d65dc11c73 fix(frontend): restore voice guild/channel picker + media URL queue input
- voice/view: add Select for guild + voice channels + Connect/Disconnect bar
- media/view: restore URL queue input + Screen toggle + Queue button
(tsc clean, next build green)
2026-08-14 11:28:08 +07:00
asepharyana 8b281c7feb refactor(frontend): finish design-system migration — chatbot, a11y, lint
- Rewrite chatbot container + panel to new surface/signal/ink tokens
  (was still on dead glass/text-primary tokens -> wrong colors)
- loading-skeleton: glass -> surface-2
- Fix a11y: SVG charts role=img+aria-label, audio aria-label,
  message-entry as real <button>, tooltip biome-ignore (intentional)
- Type messages/page initialPage (noImplicitAny)
- tsc clean, next build green, biome 0 errors
2026-08-14 11:02:52 +07:00
asepharyana 5bbf75a65b refactor(frontend): finish shadcn→custom primitive migration (green build)
- Remove tw-animate-css import + dead src/components/ui shadcn tree
- Convert 7 orphaned components (moderation, analysis, guild-selector,
  voice/activity-timeline, shared/empty+error) to new primitives
- Add missing moderation/view.tsx; analysis uses SearchPanel directly
- globals.css now uses new signal-driven ops-console tokens
- tsc --noEmit clean, next build green (11 routes), local smoke 200
2026-08-14 10:49:44 +07:00
asepharyana 5816e94a63 fix(goLive): remove syncStream — synthetic PTS timebases make A/V sync deadlock
Symptom: video plays ~1s then freezes. BaseMediaStream sync logic:
- video _pts advances 33.3ms/frame (timeBase 1/fps), audio _pts advances
  20ms/packet (timeBase 1/48000) — two synthetic frame-index timebases that
  never share a clock.
- If audio starts late (ffmpeg audio init / Ogg header), ptsDelta = video-audio
  stays positive → isAhead() true → video loops 'await sleep(frametime) while
  isAhead()' → video freezes. Downchain: vPipe fills → proc.stdout paused →
  demuxer emits ~15fps (log: 30 frames per 2s).

Upstream dank sets syncStream because node-av provides REAL PTS from NUT in a
consistent timebase. Our raw-h264 demuxer has no real PTS; per-stream sleep-PTS
pacing alone keeps both at 1000ms/s, which is correct without a shared clock.
Re-enable sync only if real PTS is added.
2026-08-13 19:06:36 +07:00
asepharyana 11f2ad5f23 fix(goLive): kill 4.3s backlog — HWM2 pipes + wire A/V sync (dank-faithful)
Lag root cause: vPipe/aPipe were objectMode PassThrough HWM 128 → the pipe
held up to 128 frames ≈ 4.3s of video before backpressure reached the encoder.
The viewer was watching a 4+ second stale backlog.

Fixes (both faithful to @dank074/discord-video-stream):
1. vPipe/aPipe HWM 2 — at most ~1-2 frames in flight (~66ms @ 30fps), so the
   writeFrame() backpressure pauses ffmpeg stdout almost immediately and the
   whole chain (encoder → NUT → demuxer → vPipe → BaseMediaStream → WebRTC)
   runs at the sender's real pace, exactly like dank's 'resume &&= vPipe.write'.
2. Wire vStream.syncStream = aStream — audio is the master clock; video
   sleeps/wakes on ptsDelta like upstream newApi.js. Prevents A/V drift under
   variable encoder throughput.
2026-08-13 18:34:28 +07:00
asepharyana 6e188f81d6 refactor(goLive): revert to dank-faithful demuxer — no custom pacing clock
Per user direction ('pakai dank sebagai referensi karena itu yg berhasil'):
drop the custom setInterval/tail-drop emission clock entirely. The demuxer
now writes each access unit straight to vPipe with a monotonic PTS and lets
BaseMediaStream (ported 1:1 from @dank074) handle pacing via sleep-PTS + A/V
sync, exactly like the upstream library. The custom clocks were the source of
the blank tile (IDR delivery race) and the lag (head-drop watching 10s-old
frames).

Adds proper backpressure: pause ffmpeg stdout when vPipe.write() returns
false, resume on drain — mirrors dank's 'resume &&= vPipe.write(packet)' so the
encoder self-throttles to the WebRTC sender's real pace instead of bursting.
2026-08-13 18:12:34 +07:00
asepharyana 7c376ea66a fix(goLive): keep IDR in own slot so decoder always has a reference (was blank)
The tail-drop rewrite let a P-frame supersede a pending keyframe before the
emit tick fired, so the decoder never received an IDR → blank GoLive tile.
Give keyframes their own slot (pendingKey) that P-frames cannot steal, and
only emit a P-frame once at least one IDR has been shown (haveReference).
IDR is always emitted first when present so the reference re-establishes.
2026-08-13 17:43:16 +07:00
asepharyana 8ee32b8df8 fix(goLive): tail-drop emitter clock — always show the freshest frame, never lag
The Node token-bucket pacer used HEAD-drop (emit frames in arrival order,
drop newer ones when over budget). Under the encoder's ~330fps burst (ffmpeg
-re does not reliably throttle YouTube-DASH webm), the viewer was watching
frames ~10s behind live → frozen / 'patah-patah' video while audio (not
rate-limited) played current = desync.

Replace it with a steady setInterval emission clock at videoFps: each tick
emits exactly ONE frame — the NEWEST buffered one — and discards everything
older (tail-drop). At most one frame is ever held, so no backlog and no lag;
the emit clock (not the encoder rate) defines playback speed. Keyframes are
never superseded so the decoder keeps getting IDRs. Audio stays in sync.
2026-08-13 17:32:18 +07:00
asepharyana c285a4c813 fix(voice): copy cookies to temp before yt-dlp + fall back to Invidious on cookie/permission errors
yt-dlp 2026.07.04 rewrites the --cookies file on close. Handing it the
root-owned /etc/.../ytcookies.txt (not writable by the gmw service user)
caused PermissionError -> exit 1 on every screen-share download attempt.

- buildCookieArgs on-disk branch now copies the system cookie file into a
  per-run temp file (like the env branch) so write-back lands somewhere we
  own; unreadable -> anonymous.
- resolveInputWithRetry Invidious fallback regex now also matches
  permission|EACCES|cookie, so a cookie failure triggers the link-alternative
  (no-auth Invidious mirror) path instead of failing all retries.
- adds regression test asserting the original cookie path is never passed to yt-dlp
2026-08-13 17:16:05 +07:00
asepharyana f156fc0c9e fix(goLive): download screen-share media to file before play (not live pipe)
The live pipe (yt-dlp -o - -> ffmpeg) delivers data at network speed with
unreliable PTS, which defeats ffmpeg -re and made x264 -r 30 force-duplicate
held frames -> ~1fps video (the patah-patah symptom). Per user suggestion,
download the FULL clip to a temp file first (downloadScreenInput), then feed
that FILE PATH to prepareStream. String inputs already get -re, so the
encoder now paces cleanly at 1x against a monotonic-PTS file — proven
reliable in local tests (vs the live pipe which always bursted). Temp file
is removed on stream end / stop.

- getDirectScreenInput -> downloadScreenInput (returns file path)
- resolveInputWithRetry now awaits a completed file + retries on failure
- screenShareController.stops/cleanup removes the per-run tmpdir
- screenShareInput.test.ts updated to the file-download contract
2026-08-13 16:42:41 +07:00
asepharyana 89f1097729 fix(goLive): add -re throttle at encoder for screen-share pipe input
Previous code only added ffmpeg -re when input was a string URL. Screen
share passes a Readable pipe (yt-dlp merge -> stdout) delivered at network
speed (bursts + stalls). Without -re the encoder slurps it instantly and,
when the merge stalls, x264 -r 30 force-duplicates the last held frame
~30x -> viewer sees ~1fps while WebRTC still paces 30fps. Add -re for all
inputs so the encoder paces at the stream's native PTS rate and emits a
fresh picture every frame.
2026-08-13 16:18:51 +07:00
asepharyana df24c756a0 fix(goLive): token-bucket pacing + pin biome rules so CI passes
- Demuxer.ts: deterministic token-bucket video pacing (replace unreliable ffmpeg -re which did not throttle the live multi-stage pipe — demuxer emitted ~240fps vs 30fps sender, 100k+ frame backlog, frozen video). Surplus non-key frames dropped; keyframes forced through; audio on fd3 unaffected.
- biome.json: pin noExplicitAny/noUnused* to off/warn. Biome 2.5.x (drifted via --no-frozen-lockfile) promotes these to errors and was failing the CI gate on pre-existing backend code unrelated to this change. Restores the warn-level behavior the config schema 2.2.0 expects.
2026-08-13 14:25:02 +07:00
asepharyana add31d3561 fix(goLive): deterministic token-bucket video pacing at demuxer (replace unreliable -re)
Root cause (3rd iteration): ffmpeg '-re' on the demuxer does NOT reliably
throttle a multi-stage live pipe (merge ffmpeg -> encoder x264 -> NUT ->
demuxer). In production the demuxer still emitted ~240fps while the WebRTC
sender consumed 30fps, building a 100k+ frame backlog (observed: frames=197490
vs sent #24600, ~8.4 min in). The sender always emitted the OLDEST buffered
frame -> video frozen ~10 min behind live, while audio (tiny, jitter-buffer
recovered) stayed smooth. Local file/pipe tests showed -re working (30fps)
but the live YouTube/WebM pipeline did not — -re is not trustworthy here.

Fix: enforce 1x video output with a token-bucket limiter in the demuxer
(Node side), independent of ffmpeg. Capacity = 1s of frames, refill 1 token
per 1000/fps ms. Surplus non-key frames are DROPPED (never buffered) so the
sender always emits the newest frame; keyframes are forced through even over
budget so the decoder keeps a fresh IDR. The limiter does NOT stall the ffmpeg
process (unlike the earlier proc.stdout pause), so audio on fd3 keeps flowing.

Verified: tsc --noEmit clean.
2026-08-13 14:17:35 +07:00
asepharyana 6e7c4901c9 fix(goLive): pace demuxer with -re + bounded frame-drop (was: audio patah, 8s lag)
Root cause (revisited): the previous gate paused proc.stdout when vPipe was
full. That stalled the SAME ffmpeg process that also writes audio on fd3, so
audio stuttered; and the ~8s backlog already built never drained → permanent
lag. Symptom: 'video still lags bad, now audio also choppy'.

Fix:
- spawn demuxer ffmpeg with -re for stream (pipe) input. Verified locally:
  a 5s NUT clip demuxes in 0.088s without -re (57x burst) vs 4.539s with -re
  (real-time). -re throttles the input read, which back-pressures the whole
  upstream chain (encoder x264 -> merge ffmpeg -> yt-dlp) through OS pipes,
  pinning production at 1x. No unbounded backlog.
- drop oldest queued frame when vPipe readableLength >= 30 (transient sender
  stall guard) instead of pausing stdout — keeps video fresh and audio intact.
- removed gateSource/sourcePaused entirely.

Audio and video now pace together at 1x; video is the newest frame, not an
8-second-old one.
2026-08-13 12:41:49 +07:00
asepharyana 60faaa9304 fix(goLive): backpressure-throttle screen-share pipeline to 1x (video freezes while audio plays)
Root cause: prepareStream's ffmpeg consumed a YouTube VOD at download/CPU
speed (~10x real-time), so the demuxer buffered a huge frame backlog.
The sender paces at 30fps but always emitted the OLDEST buffered frames, so
the viewer saw frozen/laggy video while audio (tiny, jitter-buffer
recoverable) stayed smooth. That is exactly the 'video stuck, voice normal'
symptom reported live.

Fix: propagate vPipe backpressure UP to the demuxer's ffmpeg stdout — when
the sender can't keep up, pause the source, which stalls the demuxer and
back-pressures the encoder, pinning the whole pipeline to 1x. Also add a
realtime (-re) option for file/URL inputs (no-op for the streaming path,
which is what screen share uses).

Verified: 10s test clip encodes in 1.8s without -re vs 9.5s with it; tsc --noEmit clean.
2026-08-13 11:59:33 +07:00
asepharyana 5505983dbd fix(goLive): retry VIDEO(op12)/SPEAKING(op5) opcodes until ws OPEN — broken shared-screen video
Root cause: BaseMediaConnection.sendOpcode is a silent no-op when
ws.readyState !== OPEN. In GoLive, playStream() calls setVideoAttributes(true)
+ setSpeaking(true) the instant createStream() resolves (right after
SELECT_PROTOCOL_ACK), but the StreamConnection WebSocket can still be in
CONNECTING for a few ms — so op 12 (VIDEO, activating the video SSRC) was
silently DROPPED every session. Empirically verified: 0 ops 12/5 ever logged
across the entire journal, yet 10k+ video frames were sent and audio played
(audio SSRC is activated via the VoiceConnection handshake, independent of
GoLive op 12). Discord's media server thus received video RTP on video_ssrc
but was never told to forward it → black/broken shared-screen video with
working voice.

sendOpcodeWhenOpen retries up to ~2s for ws OPEN instead of dropping. Also
emits a=fmtp:101 packetization-mode=1;profile-level-id=42e01f in the answer
SDP (H264 FU-A fragments require packetization-mode=1 to reassemble).

Also removes pre-existing noNonNullAssertion lint (biome 2.5.8 now errors)
that was blocking the deploy CI.
2026-08-13 01:44:40 +07:00
asepharyana 3d57e9c102 fix(goLive): retry VIDEO(op12)/SPEAKING(op5) opcodes until ws OPEN — broken shared screen video
Root cause: BaseMediaConnection.sendOpcode is a silent no-op when
ws.readyState !== OPEN. In GoLive, playStream() calls
setVideoAttributes(true) + setSpeaking(true) the instant createStream()
resolves (right after SELECT_PROTOCOL_ACK), but the StreamConnection WebSocket
can still be in CONNECTING for a few ms — so op 12 (VIDEO, enabling the video
SSRC) was silently DROPPED every session. Empirically verified: 0 ops 12/5 ever
logged across the entire journal, yet 10k+ video frames were sent and audio
played (audio SSRC is activated via the VoiceConnection handshake, independent
of GoLive op 12). Discord's media server thus received video RTP on video_ssrc
but was never told to forward it → black/broken shared-screen video with
working voice.

sendOpcodeWhenOpen retries up to ~2s for ws OPEN instead of dropping. Also
keeps the H264 packetization-mode=1 answer-SVP (defensive SDP correctness).

Also fix: emit a=fmtp:101 packetization-mode=1;profile-level-id=42e01f in the
answer SDP — H264 FU-A fragments require packetization-mode=1 to reassemble.
2026-08-13 00:24:53 +07:00
asepharyana d3cb5f6756 refactor: rombak cache AI analisis image — pakai CDN URL langsung, hapus phash+sha
- Cache key image = CDN URL (query params stripped), bukan SHA data URL
  → re-analysis SAME attachment selalu cache-hit, berbeda attachment tidak kolisi
- Hapus perceptual hash (imghash dep + phash get/upsert/compute) sepenuhnya
- Hapus makeImageCacheKey hashing, ganti makeImageCacheKey yang return CDN URL
- textCacheStore, visionAnalyzer, mediaCache, mediaAnalysisClient updated
- imghash dependency removed from package.json
- Purge 82 stale cache rows (image: + phash:) dari DB
2026-08-12 22:31:08 +07:00
asepharyana 37787cc4f0 fix: prevent false positive moderation on physics/tech discussions
- Add examples for technical discussions (kinetic energy, drone weapon
  engineering, physics simulations) that should be marked clean
- System rule: physics/engineering topics (kinetik, gravitasi, energi,
  drone, senjata, drone warfare, CAD, CNC, 3D printing, robotics, aerospace)
  are safe when in technical context — flag only if explicit threat
- Riwayat pengguna dengan pelanggaran sebelumnya tidak memengaruhi
  penilaian pesan bersih yang terpisah dan tidak mengandung pelanggaran
2026-08-12 22:12:11 +07:00
asepharyana f849a87f2f fix: remove user history injection to prevent false positive moderation
- Removed getUserRecentInfractions usage in textBatchProcessor.ts and visionAnalyzer.ts
- Removed buildUserHistoryXml import and calls
- Messages are now evaluated standalone, not influenced by past violations in other channels
- Updated moderation prompts with clearer instructions about user_history usage
- Fixes issue where benign messages like 'tubuh manusia vs gravitasi' were incorrectly flagged due to carryover from previous drone weapons discussion

The user history context was causing the LLM to interpret unrelated current messages
as threats because it conflated them with past violations. Now each message is judged
on its own merit with only channel-specific context.
2026-08-12 20:48:26 +07:00
asepharyana deb5dedf2c test(gateway): add regression test untuk makeImageCacheKey collision
Verifies that two data URLs sharing the first 128 chars (same MIME prefix
+ identical base64 header — the real-world scenario that caused ALL images
to reuse the same cached vision analysis) produce DIFFERENT cache keys
under the fixed full-dataURL hashing, whereas the old 128-char-prefix
approach would collide. Also includes consistency + prefix tests.
2026-08-12 19:34:15 +07:00
asepharyana 3b221823e7 feat(gateway): add observability logging for vision cache hits/misses
Add debug logging to trace cacheKey + messageId + content length on
every vision cache HIT and MISS, so we can detect if the vision model
returns duplicate analysis for different images (provider issue vs
cache collision). Includes the phash on cache miss (new analysis cached).

Follow-up to 9f7ce7d which fixed makeImageCacheKey to hash full data
URL instead of just first 128 chars (root cause of all images sharing
the same cached 'konten judi' verdict due to hash collision).
2026-08-12 19:22:56 +07:00
asepharyana 9f7ce7dbd5 fix(gateway): hash full image data URL for cache key to prevent collision
Root cause: makeImageCacheKey() only hashed the first 128 chars of the
data URL. Since all resized images use the same MIME prefix
('data:image/png;base64,') + identical base64 header bytes, nearly every
image got the same 16-char hash → 'image:<same-hash>' → all images reused
the first cached vision analysis (often a gambling-detection verdict).

Fix: hash the entire data URL instead of just the prefix. Verified
114 stale 'image:' entries + 745 stale 'phash:' entries purged from prod
DB. tsc --noEmit clean, 133 tests pass.
2026-08-12 18:28:19 +07:00
asepharyana bd292fdf3d feat(gateway): Invidious fallback for YouTube 403 in screen share
YouTube blocks anon + cookies terbind ke IP browser (403 download).
Auto-rewrite youtube.com -> yewtu.be/invidious mirror saat cookies gagal.
- mediaSource: export isYoutubeWatchUrl/toInvidiousUrl/INVIDIOUS_INSTANCES
- screenShareController: resolveInputWithRetry tries Invidious instances on 403
2026-08-12 18:19:04 +07:00
asepharyana 7a7f433988 feat(gateway): read YouTube cookies from BWS env (gmw_yt_downloader_cookies) fallback to on-disk file
bws-exec exposes the BWS secret as env GMW_YT_DOWNLOADER_COOKIES.
Materialize to temp Netscape file (yt-dlp --cookies needs a path).
Falls back to /etc/gmw-discord-gateway/ytcookies.txt written by deploy.
2026-08-12 17:50:07 +07:00
asepharyana 84c5c36672 feat(gateway): YouTube cookies support for yt-dlp screen share + music
YouTube now blocks anonymous embeds (403 'Sign in to confirm you're not a
bot'). Resolve with account cookies via --cookies.

- mediaSource: buildCookieArgs() reads GMW_YT_COOKIES_PATH (default
  /etc/gmw-discord-gateway/ytcookies.txt) and injects --cookies into
  resolveMediaUrl + getDirectScreenInput + extractMediaInfo. Falls back
  to anon if file missing (graceful 403, not crash).
- bws-exec now writes cookies file from BWS secret gmw_yt_downloader_cookies
  on service start (systemd ConfigFile).
2026-08-12 17:46:05 +07:00
asepharyana c5898f7cf0 fix(gateway): crash safety on screen-share input timeout + proper error serialization
Root cause of "langsung left": YouTube bot-block/403 on u_c1tRmj7E4 (live
stream, LOGIN_REQUIRED) made yt-dlp timeout in resolveInputWithRetry (12s).
The timeout handler did cleanup() (removing once() listeners) THEN
tee.destroy(new Error(...)) — the PassThrough emitted 'error' with NO
listener left → unhandled stream 'error' event → uncaughtException →
gracefulShutdown → bot left voice.

Fix:
- resolveInputWithRetry: tee.destroy() silently after cleanup (error carried
  in the rejection only); add permanent no-op tee.on('error') safety.
- prepareStream: output.on('error') no-op so ffmpeg spawn failure before
  playStream attaches a demux listener never crashes the gateway.
- bootstrap: serialize uncaughtException/ClientError/DB errors with
  {err, errorMsg, stack} (pino only serializes the 'err' magic key — the old
  {error: err} key printed {} so crashes were invisible).
2026-08-12 16:59:41 +07:00
asepharyana 354e378e74 fix(gateway): output NUT (not raw h264) so Demuxer re-splits video+audio correctly
Revert 392bc35: streaming raw h264 video + opus on separate pipes broke
because prepareStream.output (pipe:1) feeds the Demuxer, but the opus
pipe:3 was never attached to the Demuxer's input — so for audio-capable
streams the Demuxer saw format=h264 (video-only) and emitted -an,
dropping audio RTP.

Correct design (from f1aa08c): prepareStream muxes video+audio into NUT
on a SINGLE pipe:1. The Demuxer then spawns a child ffmpeg that
demuxes NUT → -f h264 pipe:1 (pure AnnexB, start-code scan sees real
IDR type 5) + -f opus pipe:3 (Ogg Opus via createOggOpusDemux). The
start-code parser never touches NUT framing — it runs on the child
ffmpeg's clean h264 stdout.
2026-08-12 16:09:34 +07:00
asepharyana 392bc35a0d fix(gateway): output raw H264+Opus (not NUT) so Demuxer parses NAL keyframes correctly
Root cause: prepareStream muxed video+audio into a NUT container on pipe:1.
The Demuxer scans pipe:1 for AnnexB start codes (00 00 01) to split NAL
units into access units and classify keyframes (nal_type 5). NUT container
framing bytes sat in the stream and were scanned as NALs — NAL type 0
(NUT header) instead of 5 (IDR) → every frame classified key=false →
Discord decoder never got a decodable frame → static/black GoLive tile.

Fix: output raw H264 AnnexB on pipe:1 (demuxer target) and Ogg Opus on
fd3/pipe:3 for audio. NUT is only needed for *input* parsing (single
pipe carries both streams); output is demuxed into separate raw streams.
2026-08-12 15:39:47 +07:00
asepharyana 196cb1d3af fix(gateway): await audio stream line before demux resolve — audio RTP was dropped by metadata race
The demuxer resolved as soon as the VIDEO init line arrived on ffmpeg stderr.
With live NUT input the audio init line ('Stream #0:1: Audio: opus') lands in a
LATER stderr chunk (NUT info-stream packets are read incrementally from the
pipe), so `return { audio: aInfo }` captured undefined → playStream skipped
AudioStream → zero audio RTP on the audio SSRC → Discord showed a static
GoLive tile even though the NUT carried opus audio.

Fix:
- wait for BOTH video and audio init lines (when audio is expected) before
  resolving demux metadata, with a 3s timeout fallback
- default aInfo to opus/48kHz when withAudio instead of undefined, so the
  audio stream is always exposed even if the metadata line races the return
2026-08-12 14:45:06 +07:00
asepharyana 00fc852a32 feat(rtp-capture): add two-peer RTP capture test for H264 frame transmission 2026-08-12 14:26:54 +07:00
asepharyana d9f5592e6e feat(glossary): persist resolved definitions in Postgres + harden live SearXNG lookups
- Add term_glossary_cache table + migration 0014: resolved definitions are
  stored permanently (definitions rarely change); misses stay ephemeral in
  Redis/LRU with 1h TTL so transient failures get retried
- Lookup flow: LRU -> Redis -> Postgres (permanent) -> live SearXNG; DB hits
  re-warm the fast caches; stale Redis miss sentinels no longer shadow DB
- Rate-limit-aware live lookups: concurrency 2 + stagger, retry once on empty
  results, strict definition filter (Wikipedia preferred, rejects
  disambiguation/ads/translate-homepages)
- Make SEARXNG_BASE_URL configurable via env (default unchanged)
2026-08-12 14:22:22 +07:00
asepharyana f1aa08cdf6 fix(gateway): deliver audio + per-IDR SPS/PPS in GoLive screen share
Screen share showed a single frozen frame: the GoLive pipeline sent video
only (-"-an", h264 muxer cannot carry audio) so the audio SSRC never
transmitted and Discord kept the stream in thumbnail state.

- prepareStream: mux NUT when includeAudio (h264 muxer drops audio) and
  return the actual container format
- Demuxer: support NUT input with a second output pipe (fd3) carrying
  Ogg Opus; parse OGG pages into opus frames (20ms, 48kHz) emitted as
  GoLiveFrames; fix metadata parsing that dropped the audio stream line
  when it arrived in a later stderr chunk (early parsedMeta return)
- playStream: pipe audio.stream into AudioStream → RTP on the audio SSRC
- Encoders: -x264-params repeat-headers=1 → SPS/PPS inline before EVERY
  IDR (NUT remux drops container extradata; also enables PLI recovery)
- screenShareController: includeAudio true
- tests: demuxerNut.test.ts — OGG parser unit test + real ffmpeg NUT
  integration (video access units + parsed opus frames)
2026-08-12 14:20:29 +07:00
asepharyana f70a92880e feat(glossary): implement term glossary for LLM moderation with caching and extraction logic 2026-08-12 13:44:52 +07:00
asepharyana 88b13225cd fix(gateway): stream screen-share input from yt-dlp stdout — no more raw-URL 403
Second root cause (2026-08-12): even with yt-dlp http_headers forwarded,
YouTube still returns 403 when a signed DASH URL from --dump-single-json is
fetched raw by ffmpeg/curl on some videos (verified on fONoh7Pc6VU: curl
with the EXACT headers got 403; yt-dlp's own downloader succeeded). The
signature is tied to the extracting client context (po_token/visitor), not
just UA/IP.

Fix: getDirectScreenInput now spawns 'yt-dlp -o -' and returns its stdout
as a Readable — the same mechanism resolveMediaUrl already uses for music.
yt-dlp handles auth, cookies and transient retries internally. Merge
fragments go to /tmp/gmw-ytdlp-tmp (Nix store CWD is read-only → EACCES).
Removed resolveScreenInput + mergeScreenStreams (dead code).

Controller resolveInputWithRetry unchanged: tees the stream, waits for the
first byte (12s), retries with a fresh yt-dlp run up to 3x on error/EOF/
timeout, and destroys stuck inputs (EPIPE) so no process leaks.

Tests: rewritten for streaming (yt-dlp emits bytes; fail mode = exit 8
without stdout → stream must terminate with zero bytes).
2026-08-12 13:23:59 +07:00
asepharyana 67ab289caa fix(gateway): fail-fast + retry on screen share merge failure (black tile zombie)
Root cause (2026-08-12 11:50 test): merge ffmpeg hit a transient YouTube
403 and exited code 8 BEFORE prepareStream attached its input listeners
(voice release+join takes ~10s). The input's end/error events fired into
the void, the encoder stdin never received EOF, demux resolved with
fallback 0x0 metadata, setSpeaking fired anyway → stream 'started' with
zero frames for 8+ minutes (black tile, both ffmpeg processes hung).

Fixes:
- mediaSource: pass yt-dlp http_headers (UA/referer) to the merge ffmpeg
  via -headers to suppress transient 403s; destroy the returned stream
  with an error when the merge exits non-zero before producing bytes.
- screenShareController: resolveInputWithRetry — tee the merge stream and
  wait for the first readable byte (12s timeout) before proceeding; on
  error/EOF/timeout retry the whole resolution with a FRESH yt-dlp run
  (signed DASH URLs expire fast) up to 3 attempts. Stuck merges get
  EPIPE via input.destroy() so no process leaks per attempt.
- prepareStream: race guard — if the input already ended/destroyed before
  listeners attach, EOF the encoder stdin immediately; first-frame
  watchdog in playStream rejects 'started but nothing flowing' after 10s
  instead of resolving with a silent black stream.

Tests: +2 (merge-fail zero-byte terminal state, -headers forwarding).
2026-08-12 12:17:43 +07:00
asepharyana ef9e243609 fix(goLive): encode H264 baseline to match SDP profile-level-id (black tile)
SDP offer advertises profile-level-id=42e01f (constrained baseline) but
x264 encoded the default High profile — Discord's receiver configures its
decoder from the negotiated profile, so the High-profile bitstream failed
to decode → black GoLive tile despite valid access units + correct RTP
timestamps (fixed in 42a503c).

- Add -profile:v baseline to H264 encoder options (matches @dank074's
  proven config; SPS now 6742c01e → profile_idc=66 baseline, aligns with
  the 42e01f fmtp).
- Default x264 tune film → zerolatency (no lookahead — correct for live
  GoLive; @dank074 uses it).
- Update goLive-port test to assert baseline + zerolatency.
2026-08-12 11:37:04 +07:00
asepharyana 42a503c206 fix(goLive): demux access-unit grouping + correct RTP timestamps (black tile)
Demuxer emitted each AnnexB NAL as its own WebRTC frame (SPS/PPS/SEI
separate from slices) with a near-zero timestamp delta (duration=1 in a
1/90000 timebase → RTP +1/frame instead of +3000 @30fps). Discord's H264
receiver never receives a complete decodable access unit → black GoLive
tile despite frames flowing.

- Group NALs into access units: buffer param-set/SEI NALs, flush one
  frame per slice with preceding parameter sets (AnnexB start codes kept
  so the H264RtpPacketizer finds NAL boundaries).
- Timestamp each frame at the video frame rate: duration=1, timeBase
  1/fps → BaseMediaStream frametime=1000/fps ms → RTP +clockRate/fps
  (3000 @ 30fps/90kHz) and correct pacing.
- Thread explicit frameRate from playStream options (raw H264 has no
  timing info; ffmpeg guesses 25fps on stderr).
- Strengthen golive-demux-live-e2e: validates every frame has a slice,
  no bare param-set frames, keyframes carry SPS/PPS, timeBase 1/30.
2026-08-12 11:13:00 +07:00
asepharyana 652974e23a fix(goLive): gateway crash on screenshare stop — unhandledRejection during teardown
Test 00:32 confirmed the video pipeline WORKS (1410 frames @ 1280x720 sent,
ready=true, camera off) but the gateway crashed at stream stop:
unhandledRejection → graceful shutdown → systemd restart (bot offline).

Root cause candidates (both were fire-and-forget promises without .catch):
- BaseMediaConnection.setProtocols().then(...) — rejects when the PC is
  closed while setProtocols is in flight (stream teardown)
- void webRtcConn.createOffer().then(...) — rejects when the PC closes
  while the offer is still gathering

Fixes:
- .catch on both promise chains (log + continue; teardown is expected)
- unhandledRejection handler now treats transient stream errors (EPIPE,
  ERR_STREAM_DESTROYED, ERR_STREAM_WRITE_AFTER_END, ECONNRESET) like
  uncaughtException already does — warn + continue instead of shutting
  down the whole gateway. Non-transient rejections still log + shutdown
  (with String(reason) so the detail actually shows).
2026-08-12 00:36:53 +07:00
asepharyana 407e003399 fix(goLive): black screen root cause — h264 muxer can't carry audio; disable self_video camera
ROOT CAUSE of empty GoLive tile (finally): prepareStream ran with
includeAudio: true + output -f h264. The h264 muxer cannot mux audio
('h264 muxer does not support any stream of type audio') → header write
fails -22 → stdout empty → Demuxer ffmpeg 'Invalid data found when
processing input' → 0 frames → black tile. Reproduced locally end-to-end
(13s backpressure delay + prepareStream + demux).

Fixes:
- screenShareController: includeAudio: false (video-only GoLive; demux
  path never delivers audio anyway)
- Demuxer: pin input format -f h264 for stream inputs (raw AnnexB H264
  has no magic header → auto-detect unreliable on delayed pipes)
- Streamer.signalStream: self_video: false — stop flipping on the bot's
  camera in Discord (user request; screen share ≠ camera)

Verified: local repro now emits 644 frames 1280x720 (was 0); tsc/biome/
vitest all green.
2026-08-12 00:25:01 +07:00
asepharyana 968a43b0f4 debug(goLive): instrument frame pipeline — demux spawn/stderr/frames, playStream resolve, sendVideoFrame drop/send
Tile kosong meski STREAM_CREATE handshake penuh (22:18-22:19 retest):
- Demuxer logs spawn args, ffmpeg stderr errors, frame count every 30
- playStream logs createStream resolved + demux done + setPacketizer
- sendVideoFrame logs DROPPED (ready/track) + sent frame count
2026-08-11 22:31:31 +07:00
asepharyana 91c7a67d2f fix(goLive): stream demux directly instead of spool-to-file (empty screen share)
Root cause of 'tile appears but content empty': demux() spooled the live
NUT/H264 input to a temp file and awaited stream 'finish' — but the merge
ffmpeg output never ends during playback, so demux deadlocked, no probe,
no transcode, 0 frames sent.

- Demuxer: pipe input straight into ffmpeg stdin (-i pipe:0), parse NAL
  frames live from stdout; parse video metadata from ffmpeg stderr with a
  1.5s race (fall back to H264 defaults). No spool, no await-end.
- screenShareController: pass width/height/frameRate (1280x720@30) to
  playStream — matches the prepareStream encode settings, so setVideoAttributes
  gets real dimensions even when ffmpeg can't report metadata on an open pipe.
- Add tests/golive-demux-live-e2e.ts: proves frames flow while input is
  still open (regression test for the deadlock).
2026-08-11 21:43:16 +07:00
asepharyana 8615383829 fix(goLive): STREAM_CREATE handshake — self_video voice state + retry + send instrumentation
- signalStream: flip voice state to self_video:true/self_deaf:false before
  STREAM_CREATE (Discord silently ignores the request while video disabled)
- createStream: attach dispatch listeners before first signal (race), clean
  up listeners on timeout, retry STREAM_CREATE every 3s up to 4 attempts
  (upstream issue #217/#219 — Discord randomly drops the request)
- sendOpcode: direct [goLive:Streamer] log bypassing bootstrap debug filter
  (proves op 18 is actually broadcast)
2026-08-11 20:45:00 +07:00
asepharyana 10d7ecd405 fix(gateway): EPIPE crash on media stop — stream error handlers + no shutdown on transient stream errors 2026-08-11 20:10:05 +07:00
asepharyana ff554fcff2 fix(goLive): instrument voice/stream handshake + createStream timeout (12s) 2026-08-11 20:00:33 +07:00
asepharyana f8b253ba5e merge: libdatachannel-min GoLive stack (build -86pct, node_modules -1.09GB) 2026-08-11 19:07:57 +07:00
285 changed files with 11199 additions and 21897 deletions
@@ -0,0 +1,418 @@
# GMW Frontend — Greenfield Rebuild (Visual-System Overhaul + Custom UI + Motion/3D)
> **For Hermes:** Execute with `subagent-driven-development` (one fresh subagent per task, two-stage review). Each task is 13 min, atomic, independently verifiable, committed after each. Reuse `src/lib/api/*`, `src/lib/ws/*`, `src/lib/types/*`, hooks verbatim. Never invent endpoints.
**Goal:** Rebuild the GMW Discord-automod dashboard frontend from scratch — drop shadcn/ui + glass/teal/purple aesthetic entirely, replace with a custom, distinctive design system where every page has its own visual metaphor (no uniform bordered-card grid), and push the presentation layer with **Framer Motion (`motion`) choreography + signature Three.js scenes**. Re-integrate to the EXISTING backend API + WebSocket contract (do NOT touch the backend).
**Architecture:** Next.js 16 App Router (SSR) + React 19 + TS strict + Tailwind v4. Keep the *plumbing* (data contract), rebuild the *skin + primitives + motion*. Each route = one self-contained `page.tsx` (server fetch + client view in the same file via a `"use client"` sibling export). Custom SVG charts (no recharts). Custom micro-primitives (no shadcn/base-ui). New token system in `globals.css`. **Motion:** `motion/react` app-wide for page transitions, spring micro-interactions, layout animation. **3D:** raw `three` (no react-three-fiber — leaner) in exactly TWO signature scenes, lazy-loaded client-only with graceful fallback. Deployed unchanged via existing flake + `gmw-proxy` nginx (`:4009` → Next `:4017`).
**Tech Stack:** next@16, react@19, tailwindcss@4 (`@import "tailwindcss"`), `next/font/google` (Bricolage Grotesque + Inter + JetBrains Mono), `swr` (data revalidation), `lucide-react` (icons only), `motion` (Framer Motion successor — `motion/react`), `three` + `@types/three` (signature scenes only), `clsx` + `tailwind-merge`. **Removed:** `@shadcn/react`, `@base-ui/react`, `recharts`, `shadcn` CLI, `cmdk`, `sonner`, `react-day-picker`, `embla-carousel-react`, `react-resizable-panels`, `input-otp`, all 50 `components/ui/*`.
---
## 0. Design System (the creative core — read before coding)
**Persona:** Controlled UX Designer + a tactical "ops console" voice. Material honesty: hierarchy via **scale/weight/tonal blocks**, NOT borders/shadows. Per `frontend-design` skill principle — spend the boldness in ONE signature place per page, keep the rest disciplined. Motion is choreography, not confetti: **one orchestrated moment per page**, everything else quiet.
### Palette (warm, signal-driven — NO teal/cyan/purple/blue gradients)
Light mode (`:root`):
```
--canvas: oklch(0.96 0.012 80) /* warm off-white, not cream */
--surface: oklch(0.92 0.014 80) /* tonal block, replaces bordered card */
--surface-2: oklch(0.88 0.016 80)
--ink: oklch(0.22 0.02 70) /* primary text */
--ink-soft: oklch(0.46 0.02 70) /* secondary text */
--hairline: oklch(0.22 0.02 70 / 0.10) /* structural rules ONLY, sparse */
--signal: oklch(0.78 0.17 125) /* lime — OK / live / primary accent */
--signal-ink: oklch(0.20 0.03 70) /* text ON signal */
--amber: oklch(0.80 0.15 70) /* WARN */
--vermilion: oklch(0.62 0.21 25) /* FLAGGED / destructive */
--ring: var(--signal)
```
Dark mode (`.dark`, default theme per `next-themes`):
```
--canvas: oklch(0.13 0.015 70) /* warm charcoal, not blue-black */
--surface: oklch(0.18 0.02 70)
--surface-2: oklch(0.23 0.022 70)
--ink: oklch(0.93 0.01 75)
--ink-soft: oklch(0.62 0.02 75)
--hairline: oklch(1 0 0 / 0.09)
--signal: oklch(0.88 0.18 125)
--signal-ink: oklch(0.18 0.03 70)
--amber: oklch(0.85 0.15 70)
--vermilion: oklch(0.68 0.21 25)
```
Three semantic signals reused everywhere: **lime = OK/live, amber = warn, vermilion = flag/danger**. This kills the purple-accent + teal-primary monotony.
### Typography (3 roles, deliberate pairing — not "Inter everywhere")
- **Display:** `Bricolage Grotesque` (700800) — characterful grotesque for headers/big numbers.
- **Body/UI:** `Inter` (400600).
- **Data/label:** `JetBrains Mono` (500/700) — all stats, timestamps, channel IDs, metrics.
Load all three via `next/font/google` with CSS variables (keep current `--font-inter`/`--font-jetbrains-mono` names + add `--font-display`).
### Layout & signature
- **No `card` with border.** Use tonal `--surface` blocks with generous radius (`--r: 14px`) and internal padding; separate blocks with whitespace + sparse hairlines only where structurally meaningful.
- **Signature element = "scan-tick":** a 1px animated pulse line (CSS keyframe `scan`) that marks every live/section header — NOT a card outline. Global canvas carries a faint warm dot-grid texture (low opacity) instead of the current bluish dotted radial-gradient.
- **Nav = left "spine":** vertical rail of icon nodes joined by a hairline; active node gets a `--signal` dot + label reveal. Collapses to a bottom tab-bar < 768px (CSS only, no JS sidebar primitive).
- **Header = "status bar":** connection state (WS dot), guild selector, live clock — mono font, reads like an instrument readout.
## 0.1 Motion & 3D Layer (the "lebih kreatif" addition)
### Motion rules (from `motion/react`)
- **Page transitions:** one shared `RouteTransition` in `(dashboard)/layout.tsx``AnimatePresence mode="popLayout"` + `motion.div key={pathname}` (fade + 8px rise + slight blur-out, ~220ms, `easeOut`). Consistent everywhere, zero per-page boilerplate.
- **Enter choreography (per page, ONE signature moment):** staggered rise-in for the hero/ticker group using `staggerChildren` variants; afterwards, quiet springs for hover/tap (`scale: 1.03` on interactive blocks, `whileTap` on buttons).
- **Layout animation:** `layout` prop on list items (messages rows, recording rows, queue) so add/remove/filter reflows smoothly; `layoutId` for shared-element transitions (ticker → detail modal on dashboard).
- **Live pulse:** `motion` drives the severity ticks / speaker rings with springs, not CSS `transition` alone.
- **`useReducedMotion()`** (from `motion/react`) gates ALL heavy motion; CSS `@media (prefers-reduced-motion: reduce)` additionally kills `scan`/`spin-disc` keyframes. Accessibility floor, non-negotiable.
- **No scroll-jacking, no marquee loops, no per-element confetti.** One moment per page. (`frontend-design` skill: "Satu momen orkestrasi biasanya lebih mengena daripada efek tersebar.")
### 3D rules (raw `three`, no R3F — lean bundle)
- Exactly **two** scenes, chosen because they carry real data meaning: **Dashboard hero** (`SignalField` — a particle field whose pulse density reflects live activity) and **Voice page** (`OrbField` — speakers as glowing orbs whose height/ring radius reacts to who is speaking). Everything else stays 2D/motion.
- **Lazy + client-only:** `next/dynamic(() => import("./SignalField"), { ssr: false, loading: () => <StaticFallback/> })`. Three ships in its own chunk, loaded only on those two routes.
- **WebGL guard:** if `!window.WebGLRenderingContext` or context creation fails → render the static SVG/CSS fallback (a stylized 2D version of the same visual). Never blank.
- **Perf guardrails:** `dpr: [1, 1.75]`, `powerPreference: "high-performance"`, `antialias: true`; RAF loop paused on `document.hidden`; `dispose()` geometries/materials on unmount; particle count capped by `navigator.hardwareConcurrency` + viewport (`Math.min(900, w*h/2000)`).
- **Style:** warm palette ONLY — signal-lime particles, amber/vermilion for flag/warn states; fog + soft additive blending for glow (no harsh white lights, no metallic PBR).
- **Interactivity:** subtle pointer parallax (camera lerp toward cursor) + gentle idle rotation. No drag/drop, no raycasting menus.
### Per-page metaphor (kills monotony — each page feels different)
| Route | Metaphor | Signature visual | Motion / 3D |
|---|---|---|---|
| `/dashboard` | Live ops overview | **3D signal particle field** hero + asymmetric ticker blocks + radial moderation gauge | **3D SignalField** (reacts to activity), staggered ticker rise-in, gauge draws on mount |
| `/messages` | Transcript | Left channel **timeline spine**; right = flowing message entries with left **severity tick** (no bordered cards); search = command palette | `layout` on message rows, spring severity ticks, palette types in |
| `/voice` | Stage | **3D orb field** of speakers + equalizer rings; activity = horizontal **session ribbon** | **3D OrbField** (speaking → orb rises + ring pulses), session ribbon draws sequentially |
| `/media` | Turntable | Rotating **disc** now-playing; queue = borderless list | CSS 3D disc spin (spring on play/pause), queue `layout` reflow, progress bar springs |
| `/recordings` | Tape library | Rows with **waveform thumbnail** (custom SVG from duration) | Waveform bars spring on hover; new upload animates in (AnimatePresence) |
| `/moderation` | Security log | Vertical **event flow** with status nodes (dot + line), not a table of cards | Nodes pulse on live action; timeline draws in sequence |
| `/analysis` | Query console | Terminal-style search panel | Typing cursor + results stagger |
---
## 1. What to KEEP (reuse verbatim — correct + integrates to backend)
- `src/lib/api/server.ts` — 11 server fetchers (`getDashboardStats`, `getActivity`, `getMediaStatus`, `getConfig`, `getModerationStats/Actions`, `getGuilds`, `getVoiceStatus`, `getRecordings`, `getMessages`). **No change.**
- `src/lib/api/client.ts``apiRequest` + `ApiError`. **No change.**
- `src/lib/ws/*``connection.ts`, `context.tsx`, `types.ts` (17 typed events: `message_created/updated/deleted/analyzed`, `voice_*`, `media_state`, `voice_pcm_data` binary). **No change.**
- `src/lib/types/*` — all interfaces. **No change.**
- `src/lib/format.ts`, `src/lib/utils.ts` (`cn`). **No change.**
- `src/lib/navigation.ts``navItems`. Keep but extend icons/labels if needed.
- Hooks: `src/hooks/*` (use-dashboard, use-media, use-messages, use-voice, use-moderation, use-recordings, use-guilds, use-config, use-chatbot-user, use-action, use-mobile), `src/lib/hooks/use-mounted.ts`. **Reuse** (verify no bad imports into deleted barrels).
- Feature logic kept but **reskinned**: `src/components/media/music-player.tsx`, `src/components/voice/*`, `src/components/chatbot/*`, `src/components/messages/*`, `src/components/recordings/*`, `src/components/moderation/moderation-section.tsx`, `src/components/analysis/search-panel.tsx`, `src/components/dashboard/*` (charts → rewritten as custom SVG).
## 2. What to DELETE
- `src/components/ui/*` (all 50 shadcn primitives).
- `components.json`, `@shadcn/react` + `@base-ui/react` + `shadcn` deps.
- `recharts` (replace with custom SVG chart helpers in `src/components/charts/`).
- `src/app/globals.css` → rewrite (no `@import "shadcn/tailwind.css"`, no `--color-primary` teal, no `.glass*`, no `.text-gradient`/`.gradient-border` teal/purple, no bluish body texture).
- `src/components/layout/app-sidebar.tsx` (shadcn Sidebar) → replace with custom `Spine` nav.
- All `page.tsx`+`view.tsx` pairs → merge into single `page.tsx` per route.
## 3. What to BUILD (new)
- New `globals.css` (tokens above + utilities + keyframes `scan`, `eq`, `fade-up`, `spin-disc`).
- `src/components/primitives/` — minimal custom: `Button`, `Input`, `Select`, `Dialog`, `Tooltip`, `Badge`, `Progress`, `Avatar`, `Skeleton`, `Toast`, `Sheet`.
- `src/components/motion/``RouteTransition.tsx`, `Stagger.tsx`, `variants.ts`.
- `src/components/three/``SignalField.tsx`, `OrbField.tsx`, `WebGLGuard.tsx`, `StaticFallback.tsx`, `useThreeScene.ts`.
- `src/components/charts/``Sparkline`, `AreaActivity`, `RadialGauge`, `SessionRibbon`, `Waveform`.
- `src/components/layout/``Spine.tsx`, `StatusBar.tsx`, `ThemeToggle.tsx` (reskin).
- 7 merged `page.tsx` files (one per route) implementing the metaphors above.
- `src/app/layout.tsx` — root (fonts + ThemeProvider + Toaster), `src/app/(dashboard)/layout.tsx` — providers + Spine + StatusBar + RouteTransition + MiniPlayer + Chatbot, `src/app/page.tsx` → redirect `/dashboard`.
---
## 4. Target File & Folder Structure (authoritative)
Everything below `services/frontend/src/` is the new tree. `REWRITE` replaces existing; `NEW` creates; `DELETE` removes. The `app/` route tree collapses `page.tsx`+`view.tsx` into single `page.tsx` files containing BOTH server fetch (default export) and client view (`"use client"` named export in same file).
```
services/frontend/
├─ package.json REWRITE (drop shadcn/base-ui/recharts; add motion, three, @types/three)
├─ pnpm-workspace.yaml REWRITE (onlyBuiltDependencies: keep build list minimal)
├─ components.json DELETE (shadcn registry config — no longer used)
├─ next.config.ts KEEP (output: standalone, trailingSlash, images.unoptimized)
├─ tsconfig.json KEEP (paths "@/*" → src/*, strict)
├─ postcss.config.mjs KEEP (@tailwindcss/postcss)
├─ biome.json KEEP
└─ src/
├─ app/
│ ├─ layout.tsx REWRITE (3 fonts + ThemeProvider + custom Toaster; rm sonner)
│ ├─ globals.css REWRITE (new token system §0; rm glass/teal/purple)
│ ├─ page.tsx KEEP (redirect → /dashboard/)
│ └─ (dashboard)/
│ ├─ layout.tsx REWRITE (providers + Spine + StatusBar + RouteTransition + MiniPlayer + Chatbot; rm shadcn Sidebar)
│ ├─ dashboard/ page.tsx REWRITE (server fetch + <DashboardView/> client; 3D SignalField hero)
│ ├─ messages/ page.tsx REWRITE (server seeds + <MessagesView/>; spine + severity ticks)
│ ├─ voice/ page.tsx REWRITE (server seeds + <VoiceView/>; 3D OrbField hero)
│ ├─ media/ page.tsx REWRITE (server seeds + <MediaView/>; turntable disc)
│ ├─ recordings/ page.tsx REWRITE (server seeds + <RecordingsView/>; waveform rows)
│ ├─ moderation/ page.tsx REWRITE (server seeds + <ModerationView/>; event-flow)
│ └─ analysis/ page.tsx REWRITE (client <AnalysisView/>; query console)
├─ components/
│ ├─ ui/ DELETE (all 50 shadcn primitives)
│ ├─ primitives/ NEW (Button, Input, Select, Dialog, Tooltip, Badge, Progress, Avatar, Skeleton, Toast, Sheet, index.ts)
│ ├─ motion/ NEW (variants.ts, Stagger.tsx, RouteTransition.tsx)
│ ├─ three/ NEW (WebGLGuard, SignalField, OrbField, StaticFallback, useThreeScene)
│ ├─ charts/ NEW (Sparkline, AreaActivity, RadialGauge, SessionRibbon, Waveform)
│ ├─ layout/ REWRITE (Spine NEW, StatusBar NEW, ThemeToggle REWRITE, app-sidebar DELETE)
│ ├─ dashboard/ REWRITE (stat-card DELETE; activity-chart/hourly/top-channels/moderation-donut/users/channels/reactions REWRITE)
│ ├─ messages/ REWRITE (message-card DELETE; message-list/detail/detail-view/ai-status-badge/ai-analysis-panel/attachments-grid/lightbox/search-overlay REWRITE)
│ ├─ voice/ REWRITE (voice-connection-card/connection-card/microphone-card DELETE; speaker-waveform/active-speakers-panel/activity-timeline/mic-control/listen-control REWRITE)
│ ├─ media/ REWRITE (music-player, mini-player REWRITE)
│ ├─ recordings/ REWRITE (recording-card DELETE; recording-player REWRITE)
│ ├─ moderation/ REWRITE (moderation-section REWRITE)
│ ├─ analysis/ REWRITE (search-panel REWRITE)
│ ├─ chatbot/ REWRITE (chatbot-container, chat-panel REWRITE; chatbot-context, index KEEP)
│ └─ shared/ REWRITE (empty-state, error-state, loading-skeleton, error-boundary, guild-selector REWRITE; index KEEP)
├─ hooks/ KEEP (verify no bad imports)
├─ lib/
│ ├─ api/ KEEP (server.ts, client.ts, index.ts)
│ ├─ ws/ KEEP (connection.ts, context.tsx, types.ts, ws-hook.ts)
│ ├─ types/ KEEP (all interfaces)
│ ├─ hooks/ KEEP (use-media-player.tsx, use-mounted.ts)
│ ├─ audio/ KEEP (voice PCM decode helpers if present)
│ ├─ format.ts KEEP
│ ├─ utils.ts KEEP (cn)
│ └─ navigation.ts KEEP
└─ (public assets) KEEP
```
### 4.1 Single-file page pattern (mandatory)
```tsx
// server component (default export) — runs on the server, fetches initial data
import { getX, getY } from "@/lib/api/server";
import { XView } from "./page"; // self-import of the named client export
export default async function Page() {
const [a, b] = await Promise.allSettled([getX(), getY()]);
return <XView initialA={a.status === "fulfilled" ? a.value : undefined}
initialB={b.status === "fulfilled" ? b.value : undefined} />;
}
// client component (named export) — hydrated, takes initialData as SWR fallback
"use client";
export function XView({ initialA, initialB }: Props) {
const { data } = useX(initialA); // SWR fallbackData = initialA
}
```
Self-referencing the named export keeps the file single-artifact while satisfying Next's RSC boundary (default = server, named = client). Tabs live inside `XView`.
### 4.2 Import rules (lint gate)
- No `@/components/ui/*` (deleted) — all UI via `@/components/primitives`.
- `three` only imported inside `src/components/three/*`; pages import those via `next/dynamic({ ssr: false })`.
- `motion` imported from `motion/react` only.
- All data: `@/lib/api/server` (server) / `@/lib/api/client` (client) — never invented endpoints.
---
## TASKS (granular — every file is its own task)
### PHASE 0 — Dependency surgery
- **T0.1** Edit `package.json`: remove `dependencies["@base-ui/react"]`.
- **T0.2** Remove `dependencies["@shadcn/react"]`.
- **T0.3** Remove `dependencies["shadcn"]`.
- **T0.4** Remove `dependencies["recharts"]`.
- **T0.5** Remove `dependencies["cmdk"]`.
- **T0.6** Remove `dependencies["sonner"]`.
- **T0.7** Remove `dependencies["react-day-picker"]`.
- **T0.8** Remove `dependencies["embla-carousel-react"]`.
- **T0.9** Remove `dependencies["react-resizable-panels"]`.
- **T0.10** Remove `dependencies["input-otp"]`.
- **T0.11** Add `dependencies["motion"]: "^12.0.0"`, `dependencies["three"]: "^0.180.0"`, `devDependencies["@types/three"]: "^0.180.0"`.
- **T0.12** `rm -f pnpm-lock.yaml && pnpm install` (regenerate lockfile).
- **T0.13** Verify `pnpm ls recharts @shadcn/react @base-ui/react` → empty; `pnpm ls motion three @types/three` → present.
- **T0.14** `cat pnpm-workspace.yaml`: confirm `onlyBuiltDependencies` keeps needed native builds, no broken shadcn postinstall.
### PHASE 1 — Design tokens (globals.css)
- **T1.1** Rewrite `@theme { }` head: light `:root` palette from §0 (canvas/surface/ink/ink-soft/hairline/signal/signal-ink/amber/vermilion/ring).
- **T1.2** Add radius tokens `--r: 14px`, `--r-panel: 12px`, `--r-control: 8px`, `--r-pill: 9999px`.
- **T1.3** Add `--font-display` token; keep `--font-sans`/`--font-mono`.
- **T1.4** Add `.dark { }` override block with §0 dark values (warm charcoal).
- **T1.5** Replace `@layer base body` bg: warm dot-grid `radial-gradient(oklch(0.45 0.03 70 / 0.05) 1px, transparent 1px)` + faint warm glow; remove old bluish radial layers.
- **T1.6** Retint scrollbar thumb to warm `oklch(0.4 0.02 70 / 0.2)`; keep `::selection` signal-tinted.
- **T1.7** Delete `.glass`, `.glass-elevated`, `.glass-intense`, `.dark .glass-intense` utilities.
- **T1.8** Delete `.text-gradient` and `.gradient-border`.
- **T1.9** Add `.surface` utility (bg var(--surface), radius var(--r), padding).
- **T1.10** Add `.scan-tick` (1px animated pulse line, keyframe `scan`).
- **T1.11** Add `.ticker`, `.pill`, `.mono` utilities.
- **T1.12** Add keyframes `scan`, `eq`, `fade-up`, `spin-disc` (keep used existing ones if still referenced).
- **T1.13** Add `@media (prefers-reduced-motion: reduce)` kill switch for scan/eq/spin-disc/pulse-ring/shimmer.
- **T1.14** Remove `@import "shadcn/tailwind.css";` (line 3); verify nothing else depends on shadcn CSS vars.
- **T1.15** Verify `grep -c "0.52 0.17 215\|0.55 0.2 280" src/app/globals.css``0`.
- **T1.16** `pnpm biome check src/app/globals.css` → no errors.
### PHASE 2 — Root layout + fonts
- **T2.1** In `layout.tsx` add `Bricolage_Grotesque` (`variable: "--font-display"`, subsets `["latin"]`, `display: "swap"`).
- **T2.2** Apply `inter.variable`, `jetbrainsMono.variable`, `bricolage.variable` to `<html>`.
- **T2.3** Remove `import { Toaster } from "@/components/ui/sonner"`.
- **T2.4** Comment out `<Toaster />` temporarily (re-enabled after T3.10).
- **T2.5** Keep `suppressHydrationWarning`, `ThemeProvider` (defaultTheme dark, enableSystem false).
- **T2.6** `npx tsc --noEmit` (fonts only; rest may still error until primitives exist).
### PHASE 3 — Custom primitives (replace 50 shadcn ui)
- **T3.1** `primitives/Button.tsx`: `motion.button`, variants `primary`/`ghost`/`danger`, `cn()` merge, `cursor-pointer`, `focus-visible:ring-2 ring-signal`, `whileTap` scale 0.97 gated by `useReducedMotion()`.
- **T3.2** `primitives/Input.tsx`: native `<input>`, `bg-surface`, `rounded-[var(--r-control)]`, `mono` prop.
- **T3.3** `primitives/Select.tsx`: native `<select>` styled, `bg-surface`.
- **T3.4** `primitives/Dialog.tsx`: native `<dialog>` + `showModal()`, warm `::backdrop`, `AnimatePresence`, `onClose`.
- **T3.5** `primitives/Tooltip.tsx`: CSS group-hover popover.
- **T3.6** `primitives/Badge.tsx`: tonal pill, `variant``bg-{tone}/15 text-{tone}` (signal/amber/vermilion/neutral).
- **T3.7** `primitives/Progress.tsx`: SVG track + `motion` fill, `value`/`max`, signal color.
- **T3.8** `primitives/Avatar.tsx`: `<img>` + initials fallback, signal bg, size prop.
- **T3.9** `primitives/Skeleton.tsx`: shimmer block (signal-tinted), `aria-hidden`.
- **T3.10** `primitives/Toast.tsx`: `ToastProvider` context + portal, `useToast()`, motion slide-in, auto-dismiss.
- **T3.11** `primitives/Sheet.tsx`: mobile drawer (`translate-x` spring), overlay, `open`/`onClose`.
- **T3.12** `primitives/index.ts` re-export all 11.
- **T3.13** Re-enable `<Toaster />` in `layout.tsx` (T2.4).
- **T3.14** Verify `npx tsc --noEmit` on primitives; `grep -rl "@/components/ui/" src/components/primitives` → empty.
### PHASE 4 — Motion foundation
- **T4.1** `motion/variants.ts`: export `spring`, `ease`, `fadeUp`, `stagger` (per §0.1).
- **T4.2** `motion/Stagger.tsx`: `StaggerGroup` + `StaggerItem` (`"use client"`).
- **T4.3** `motion/RouteTransition.tsx`: `"use client"`, `usePathname`, `AnimatePresence mode="popLayout"`, reduced-motion fallback to plain `<div>`.
- **T4.4** Verify `npx tsc --noEmit` on motion; `motion/react` import resolves.
### PHASE 5 — Custom SVG charts (replace recharts)
- **T5.1** `charts/Sparkline.tsx`: `<svg>` polyline from `points:number[]`, signal stroke, no axes.
- **T5.2** `charts/AreaActivity.tsx`: filled `<path>` area, low-opacity signal gradient, `pathLength` draw gated by reduced-motion.
- **T5.3** `charts/RadialGauge.tsx`: `<circle>` arc `stroke-dasharray`, center mono label.
- **T5.4** `charts/SessionRibbon.tsx`: horizontal segments per speaker duration.
- **T5.5** `charts/Waveform.tsx`: bars from deterministic seed, spring scaleY on hover.
- **T5.6** Verify `npx tsc --noEmit` on charts; no `recharts` import.
### PHASE 6 — Three.js foundation (lazy, guarded)
- **T6.1** `three/WebGLGuard.tsx`: `"use client"`, detect webgl2/webgl, render `children` or `fallback`.
- **T6.2** `three/useThreeScene.ts`: shared hook — renderer init (`dpr:[1,1.75]`, `powerPreference`), RAF with `document.hidden` pause, `dispose()` on unmount, resize observer.
- **T6.3** `three/SignalField.tsx`: `Points` BufferGeometry (~min(900, w*h/2000)), additive blend, signal-lime, fog, idle rotation + sine drift, `activity` prop, pointer parallax. Uses `useThreeScene`.
- **T6.4** `three/OrbField.tsx`: per-speaker `Sphere`, y-scale + ring lerp to speaking, tones signal/idle/vermilion.
- **T6.5** `three/StaticFallback.tsx`: 2D SVG/CSS silhouette for both scenes.
- **T6.6** Verify `grep -rl "from \"three\"" src | grep -v "components/three"` → empty; `npx tsc --noEmit` on three.
### PHASE 7 — Layout shell
- **T7.1** `layout/Spine.tsx`: `"use client"`, vertical rail from `navItems`, icon node + hairline, active signal dot + label reveal (motion spring), `max-md:` bottom tab-bar.
- **T7.2** `layout/StatusBar.tsx`: `"use client"`, page title + WS status dot (motion pulse) + `GuildSelector` + live clock (mono) + `ThemeToggle`.
- **T7.3** `layout/ThemeToggle.tsx`: restyle, keep `next-themes` logic, motion icon swap.
- **T7.4** Rewrite `(dashboard)/layout.tsx`: keep SWRConfig/WsProvider/MediaPlayerProvider/ChatbotProvider + sync functions verbatim; swap `AppSidebar``Spine`, header→`StatusBar`, wrap children in `RouteTransition`; remove `SidebarInset`/`SidebarTrigger`/`Separator`; keep MiniPlayer+ChatbotContainer.
- **T7.5** Delete `layout/app-sidebar.tsx`.
- **T7.6** Verify `grep -rl "components/ui/sidebar\|app-sidebar" src` → empty; `npx tsc --noEmit`.
### PHASE 8 — Dashboard page + components
- **T8.1** Rewrite `dashboard/page.tsx`: default async `getDashboardStats`+`getActivity``<DashboardView>`; named `"use client"` view with useStats/useActivity, 3D hero + tickers + tabs.
- **T8.2** Add `WebGLGuard`+`SignalField` hero with `activity` ratio; overlay headline (Bricolage) + `<RadialGauge>`.
- **T8.3** Build asymmetric ticker row with `StaggerGroup` + 4 `.surface` blocks (mono number + label + `<Sparkline>`); inline (replaces stat-card).
- **T8.4** Delete `dashboard/stat-card.tsx`.
- **T8.5** Reskin `dashboard/activity-chart.tsx``charts/AreaActivity` (daily).
- **T8.6** Reskin `dashboard/hourly-activity-chart.tsx``charts/AreaActivity` (hourly).
- **T8.7** Reskin `dashboard/top-channels-chart.tsx``charts/` + `.surface`.
- **T8.8** Reskin `dashboard/moderation-donut.tsx``charts/RadialGauge`.
- **T8.9** Reskin `dashboard/users-section.tsx``.surface`.
- **T8.10** Reskin `dashboard/channels-section.tsx``.surface`.
- **T8.11** Reskin `dashboard/reactions-section.tsx``.surface`.
- **T8.12** Verify `grep -rl "components/ui/card" src/app/\(dashboard\)/dashboard src/components/dashboard` → empty; `npx tsc --noEmit`.
### PHASE 9 — Messages page + components
- **T9.1** Rewrite `messages/page.tsx`: default `getMessages(guildId)`(+channels) → `<MessagesView>`; client spine + entries.
- **T9.2** Delete `messages/message-card.tsx`.
- **T9.3** Rewrite `messages/message-list.tsx`: left timeline spine + right severity-tick entries (`surface` + `border-l-2` lime/amber/vermilion, motion spring tick), `layout` reflow.
- **T9.4** Rewrite `messages/message-detail.tsx``.surface`.
- **T9.5** Rewrite `messages/message-detail-view.tsx``.surface` pane.
- **T9.6** Rewrite `messages/ai-status-badge.tsx``primitives/Badge`.
- **T9.7** Rewrite `messages/ai-analysis-panel.tsx``.surface`.
- **T9.8** Rewrite `messages/attachments-grid.tsx``.surface` grid.
- **T9.9** Rewrite `messages/lightbox.tsx``primitives/Dialog`.
- **T9.10** Rewrite `messages/search-overlay.tsx` → console palette, type-in animation, `primitives/Dialog`.
- **T9.11** Verify no `components/ui/card` in messages tree; `npx tsc --noEmit`.
### PHASE 10 — Voice page + components
- **T10.1** Rewrite `voice/page.tsx`: default `getVoiceStatus()``<VoiceView>`; client `WebGLGuard`+`OrbField` hero + ribbon + tabs.
- **T10.2** Delete `voice/voice-connection-card.tsx`, `connection-card.tsx`, `microphone-card.tsx`.
- **T10.3** Rewrite `voice/speaker-waveform.tsx` → SVG ring / eq bars.
- **T10.4** Rewrite `voice/active-speakers-panel.tsx``.surface`.
- **T10.5** Rewrite `voice/activity-timeline.tsx``charts/SessionRibbon`.
- **T10.6** Rewrite `voice/mic-control.tsx``primitives/Button`.
- **T10.7** Rewrite `voice/listen-control.tsx``primitives/Button`.
- **T10.8** Verify `npx tsc --noEmit`; no `components/ui/card` in voice tree.
### PHASE 11 — Media page + components
- **T11.1** Rewrite `media/page.tsx`: default `getMediaStatus()``<MediaView>`; client turntable disc + transport + queue.
- **T11.2** Rewrite `media/music-player.tsx`: CSS-3D disc (spin-disc, pause when not playing, spring on play/pause), mono meta, `primitives/Button` transport, `.surface` queue rows with `layout`.
- **T11.3** Rewrite `media/mini-player.tsx` → compact `.surface`.
- **T11.4** Verify `npx tsc --noEmit`; no `components/ui/card` in media tree.
### PHASE 12 — Recordings page + components
- **T12.1** Rewrite `recordings/page.tsx`: default `getRecordings(50)``<RecordingsView>`; client rows `AnimatePresence`+`layout`, live `voice_recording_uploaded` prepend.
- **T12.2** Delete `recordings/recording-card.tsx`.
- **T12.3** Rewrite `recordings/recording-player.tsx``.surface` row + `charts/Waveform` + `primitives/Button`/`Dialog` play/delete.
- **T12.4** Verify `npx tsc --noEmit`.
### PHASE 13 — Moderation + Analysis pages
- **T13.1** Rewrite `moderation/page.tsx`: default `getModerationStats/Actions``<ModerationView>`; client vertical event-flow, live actions pulse-in.
- **T13.2** Rewrite `moderation/moderation-section.tsx` → event-flow, no Card/table.
- **T13.3** Rewrite `analysis/page.tsx`: client `<AnalysisView/>` terminal console (`primitives/Input` mono + blinking caret, `.surface` results staggered).
- **T13.4** Rewrite `analysis/search-panel.tsx` → terminal style.
- **T13.5** Verify `npx tsc --noEmit`.
### PHASE 14 — Chatbot + shared + final cleanup
- **T14.1** Rewrite `chatbot/chatbot-container.tsx``.surface`, keep drag/minimize.
- **T14.2** Rewrite `chatbot/chat-panel.tsx` → bubbles via `AnimatePresence`, `primitives/*`.
- **T14.3** Keep `chatbot/chatbot-context.tsx` + `index.ts`.
- **T14.4** Rewrite `shared/empty-state.tsx` → tonal.
- **T14.5** Rewrite `shared/error-state.tsx`.
- **T14.6** Rewrite `shared/loading-skeleton.tsx``primitives/Skeleton`.
- **T14.7** Rewrite `shared/error-boundary.tsx`.
- **T14.8** Rewrite `shared/guild-selector.tsx``primitives/Select`.
- **T14.9** `rm -rf src/components/ui && rm -f components.json`.
- **T14.10** `grep -rn "components/ui/\|@shadcn\|@base-ui\|recharts" src` → MUST be empty.
- **T14.11** `grep -rl "from \"recharts\"\|@base-ui\|@shadcn" src` → empty (double-check).
- **T14.12** Verify `npx tsc --noEmit` across whole `src`.
### PHASE 15 — Build + lint gate
- **T15.1** `cd services/frontend && npx tsc --noEmit` → 0 errors.
- **T15.2** `pnpm biome check` → fix all issues (no `any` in new files).
- **T15.3** `pnpm build` (standalone) → success, emits `.next/standalone/server.js`.
- **T15.4** Inspect `.next/static/chunks/` for three-heavy chunk loaded only on dashboard/voice; confirm NOT in `/dashboard/` initial SSR HTML.
- **T15.5** Confirm `pnpm-lock.yaml` present (reproducible flake install).
- **T15.6** `grep -c "0.52 0.17 215\|0.55 0.2 280" .next/static/css/*.css` → 0.
### PHASE 16 — Local runtime smoke test (no prod)
- **T16.1** Start local standalone: `GMW_BACKEND_URL=http://127.0.0.1:4001 PORT=4017 node .next/standalone/server.js &`.
- **T16.2** `curl -s -o /dev/null -w "%{http_code}"` for all 7 routes → 200.
- **T16.3** `curl /dashboard/ | grep -o "Bricolage\|signal\|surface"` → present.
- **T16.4** `curl /_next/static/css/*.css | grep "0.52 0.17 215\|0.55 0.2 280"` → empty.
- **T16.5** Headless browser `/dashboard/`+`/voice/` WebGL on: no console errors, `<canvas>` present, `<StaticFallback/>` NOT rendered.
- **T16.6** Same pages WebGL off: `<StaticFallback/>` renders, no crash.
- **T16.7** Kill local server. Do NOT touch prod unit.
- **T16.8** Confirm 7 routes 200 + no console errors in T16.5/16.6.
### PHASE 17 — Flake + staging deploy
- **T17.1** Inspect `flake.nix` frontend drv: `filterSource` ignores `out/.next/node_modules`; `pnpm-lock.yaml` included.
- **T17.2** `nix build .#gmw-frontend --impure --sandbox-off` → succeeds.
- **T17.3** `nix copy` frontend drv to VPS into staging profile (test port e.g. 4217).
- **T17.4** Create/adjust staging systemd unit with `PORT=4217` exported BEFORE `node server.js` (LIDM PORT bug).
- **T17.5** `sudo systemctl restart gmw-frontend-staging`; `curl` staging → 200.
- **T17.6** Browser-check staging `/dashboard/`+`/voice/` (3D visible, fallback test).
- **T17.7** Verify staging CSS has no old teal/purple; new design renders.
### PHASE 18 — Production swap (CONFIRM WITH USER FIRST)
- **T18.1** STOP — send staging screenshots/URL; await explicit approval before touching prod.
- **T18.2** On approval: `nix-env --profile /nix/var/nix/profiles/gmw-frontend --set <new-drv>`.
- **T18.3** Confirm `gmw-frontend.service` exports `PORT=4017` before exec.
- **T18.4** `sudo systemctl restart gmw-frontend`.
- **T18.5** `curl` all 7 routes on `https://imphnen.asepharyana.my.id` → 200.
- **T18.6** Browser verify prod: new design, old teal gone, 3D scenes render.
- **T18.7** `journalctl -u gmw-frontend -f` 5 min; confirm WS reconnect + live features.
- **T18.8** Notify user with before/after notes; keep rollback plan (`nix-env --set <previous>; systemctl restart`).
---
## Risks / Trade-offs
- **Scope:** 7 pages + charts + primitives + motion + 2 three scenes + layout. Big but mechanical; each task is isolated (~90 atomic tasks).
- **Bundle weight:** `three` adds ~150KB gz but ONLY on dashboard/voice routes (lazy chunk, `ssr:false`). `motion` ~35KB gz app-wide — acceptable.
- **WebGL compatibility:** covered by `WebGLGuard` + static fallback. Old devices / strict privacy browsers never blank.
- **Motion excess:** risk of "AI-generated" scattered animation. Guard: one signature moment per page, shared variants, reduced-motion gates.
- **Feature regressions:** Voice PCM playback, media transport, chatbot drag — logic preserved, only skin changes. Smoke test (T16) catches SSR breaks; live WS/3D needs real backend (staging T17).
- **Removed deps:** dropping `recharts`/`sonner`/`cmdk` means rewriting charts + toasts + search palette — accounted for in Phases 3/5/9.
- **Next standalone PORT bug:** `server.js` may not read `PORT` — ensure unit exports `PORT=4017` before exec (T17.4/T18.3).
- **three + React 19:** raw three avoids R3F compat surface; lifecycle (dispose + RAF) handled in T6.2.
- **next-themes:** keep (light/dark toggle); default dark.
## Open questions (answer before T18)
- Q1: Deploy to prod now or staging-only first? (Recommend staging + screenshot review.)
- Q2: Keep `react-day-picker`/`embla` if any page still needs them? (Plan assumes no — verify in T14.10 grep.)
- Q3: Any brand name/wordmark change from "Discord Automod"? (Keep "Bete" identity unless told.)
- Q4: 3D depth — full 3D scenes on dashboard+voice as specced, or also a 3D accent on media (disc)? (Default: dashboard+voice only; media disc stays CSS 3D.)
+9 -3
View File
@@ -35,12 +35,15 @@
"suspicious": {
"noUnknownAtRules": "off",
"useIterableCallbackReturn": "off",
"noArrayIndexKey": "warn"
"noArrayIndexKey": "off",
"noExplicitAny": "off"
},
"a11y": {
"useSemanticElements": "off",
"useButtonType": "off",
"noAutofocus": "off"
"noAutofocus": "off",
"useMediaCaption": "off",
"noStaticElementInteractions": "off"
},
"performance": {
"noImgElement": "warn"
@@ -50,7 +53,10 @@
},
"correctness": {
"noInvalidUseBeforeDeclaration": "off",
"noUnusedFunctionParameters": "warn"
"noUnusedFunctionParameters": "warn",
"noUnusedVariables": "warn",
"noUnusedImports": "warn",
"noUnusedPrivateClassMembers": "warn"
}
},
"domains": {
+28 -104
View File
@@ -11,11 +11,6 @@
let
pkgs = import nixpkgs { inherit system; };
# libdatachannel for the GoLive N-API binding. nixpkgs 0.24.1 is built
# against this host's glibc and ships both lib + dev headers, so the
# binding links cleanly inside the Nix sandbox (no manual cmake build).
libdatachannel = pkgs.libdatachannel;
# Source filter: `path:` literals do NOT respect .gitignore by default,
# so a dirty local out/ (stale chunks from previous builds) leaks into
# the sandbox. Filter out build artifacts explicitly.
@@ -53,8 +48,8 @@
export GIT_SSL_CAINFO=${pkgs.cacert}/etc/ssl/certs/ca-bundle.crt
export NIX_SSL_CERT_FILE=${pkgs.cacert}/etc/ssl/certs/ca-bundle.crt
# pnpm uses node-gyp for native addons provide build tools
export npm_config_build_from_source=true
# pnpm uses node-gyp for native addons provide build tools (kept for
# the rare case a prebuilt is unavailable and it falls back to compile).
export CPPFLAGS="-I${pkgs.lib.getDev pkgs.openssl}/include"
export LDFLAGS="-L${pkgs.lib.getLib pkgs.openssl}/lib"
@@ -73,7 +68,7 @@
# NOTE: do NOT use `pnpm install --prod` here — it collapses the
# public-hoist dir (.pnpm/node_modules) that runtime peer resolution
# relies on (e.g. @lng2004/node-datachannel and @seydx/node-av-linux-x64
# are only reachable through it), silently breaking voice/screenshare.
# are only reachable through it), silently breaking voice.
# Instead we keep the full install's symlink layout and only prune
# orphaned package dirs + broken symlinks.
# Must run AFTER tsc (typescript is a devDep) and after native builds.
@@ -106,31 +101,8 @@
buildPhase = pnpmInstall + ''
echo "=== Compiling TypeScript ==="
npx tsc 2>&1
echo "=== Fixing @/ path aliases to relative paths ==="
node -e "
const fs = require('fs');
const path = require('path');
let count = 0;
function walk(dir) {
if (!fs.existsSync(dir)) return;
for (const e of fs.readdirSync(dir, {withFileTypes: true})) {
const p = path.join(dir, e.name);
if (e.isDirectory()) walk(p);
else if (e.name.endsWith('.js')) {
const c = fs.readFileSync(p, 'utf8');
const pat = /from\s+['\"]@\/([^'\"]+)['\"]/g;
const n = c.replace(pat, (m, p1) => {
const target = path.join('dist', p1) + '.js';
const rel = path.relative(path.dirname(p), target);
return 'from \"' + (rel.startsWith('.') ? rel : './' + rel) + '\"';
});
if (n !== c) { fs.writeFileSync(p, n); count++; }
}
}
}
walk('dist');
console.log('Fixed ' + count + ' files');
"
echo "=== Fixing @/ path aliases + extensionless relative imports for node ESM ==="
node scripts/fix-imports.mjs
echo "=== Build complete ==="
'' + pruneProd;
@@ -167,8 +139,7 @@ WRAPPER
pkgs.pkg-config
pkgs.openssl
pkgs.openssl.dev
libdatachannel.dev # rtc/rtc.hpp headers for the GoLive binding
pkgs.git # libdatachannel FetchContent clones from GitHub
pkgs.git # for any FetchContent-based deps during native builds
pkgs.cacert
];
@@ -181,65 +152,35 @@ WRAPPER
# do NOT let stdenv run its own cmake configure phase on the source.
dontUseCmakeConfigure = true;
# The gateway bundles native node_modules (.node addons plus .o/.a
# object files left in prebuilt dirs). stdenv's fixupPhase walks
# $out/node_modules and runs patchELF + shrinkELF over every ELF it
# finds, choking on the non-ET_DYN files (.o/.a) and the prebuilt
# .node addons — emitting hundreds of harmless "patchelf: wrong ELF
# type" lines per build. The real binary is node (external, already
# RPATH-fixed in its own derivation) and the .node addons are
# self-contained prebuilts loaded via dlopen, so Nix's fixup pass is
# neither needed nor wanted here. Skip it entirely.
dontFixup = true;
buildPhase = pnpmInstall + ''
echo "=== Building native voice deps ==="
# pnpm rebuild aborts on the first failing package and runs scripts
# from the wrong cwd build each native dep explicitly with its own
# install script. Each failure is tolerated (|| true); the packages
# that matter (opus) are verified at runtime.
for pkg in \
node_modules/.pnpm/@discordjs+opus@*/node_modules/@discordjs/opus
do
if [ -d "$pkg" ]; then
echo "--- native build: $pkg ---"
(cd "$pkg" && npm run install 2>&1 || true)
fi
done
echo "=== Building libdatachannel-min N-API binding ==="
# The GoLive screen-share stack uses a minimal N-API binding
# (native/libdatachannel-min) over nixpkgs libdatachannel.
(
cd native/libdatachannel-min
# binding.gyp resolves include/lib from env (LDC_INCLUDE = .dev
# include root, LDC_LIB = lib output dir, NAPI_INCLUDE =
# node-addon-api include root).
NAPI_INCLUDE=$(find ../../node_modules/.pnpm -maxdepth 3 \
-type d -path "*node_modules/node-addon-api" | head -1)
echo "NAPI_INCLUDE=$NAPI_INCLUDE"
LDC_INCLUDE=${libdatachannel.dev} LDC_LIB=${libdatachannel.out}/lib/libdatachannel.so.0.24.1 \
NAPI_INCLUDE=$NAPI_INCLUDE \
npx node-gyp rebuild 2>&1 || true
ls -la build/Release/datachannel_min.node 2>/dev/null \
&& echo "libdatachannel-min binding OK: $(stat -c%s build/Release/datachannel_min.node) bytes" \
|| echo "WARN: libdatachannel-min binding build FAILED (screen share disabled)"
)
# @discordjs/opus ships prebuilt binaries for Node 22 (ABI node-v127,
# linux-x64-glibc-2.35) node-pre-gyp downloads the prebuilt .node
# instead of compiling C++ from source. With build_from_source unset
# (above), `pnpm rebuild` runs the package's own install script which
# fetches the matching prebuilt; it only falls back to a source build
# if the download fails. This keeps voice working without a per-build
# native compile.
echo "=== Rebuilding @discordjs/opus (prebuilt download) ==="
pnpm rebuild @discordjs/opus 2>&1 || true
echo "=== Compiling TypeScript ===="
npx tsc 2>&1
echo "=== Fixing @/ path aliases to relative paths ==="
node -e "
const fs = require('fs');
const path = require('path');
let count = 0;
function walk(dir) {
if (!fs.existsSync(dir)) return;
for (const e of fs.readdirSync(dir, {withFileTypes: true})) {
const p = path.join(dir, e.name);
if (e.isDirectory()) walk(p);
else if (e.name.endsWith('.js')) {
const c = fs.readFileSync(p, 'utf8');
const pat = /from\s+['\"]@\/([^'\"]+)['\"]/g;
const n = c.replace(pat, (m, p1) => {
const target = path.join('dist', p1) + '.js';
const rel = path.relative(path.dirname(p), target);
return 'from \"' + (rel.startsWith('.') ? rel : './' + rel) + '\"';
});
if (n !== c) { fs.writeFileSync(p, n); count++; }
}
}
}
walk('dist');
console.log('Fixed ' + count + ' files');
"
echo "=== Fixing @/ path aliases + extensionless relative imports for node ESM ==="
node scripts/fix-imports.mjs
echo "=== Build complete ==="
'' + pruneProd;
@@ -247,22 +188,6 @@ WRAPPER
mkdir -p $out/lib/gmw-discord-gateway
cp -r dist node_modules package.json tsconfig.json $out/lib/gmw-discord-gateway/
# GoLive native binding loadNative resolves it relative to
# dist/goLive/native.js, i.e. <root>/native/libdatachannel-min/
# build/Release/datachannel_min.node; libdatachannel .so must sit
# next to it and be on LD_LIBRARY_PATH at runtime.
mkdir -p $out/lib/gmw-discord-gateway/native/libdatachannel-min/build/Release
cp native/libdatachannel-min/build/Release/datachannel_min.node \
$out/lib/gmw-discord-gateway/native/libdatachannel-min/build/Release/ 2>/dev/null || true
mkdir -p $out/lib/gmw-discord-gateway/native/libdatachannel-min/build/ldc
cp -rL native/libdatachannel-min/build/ldc/libdatachannel.so* \
$out/lib/gmw-discord-gateway/native/libdatachannel-min/build/ldc/ 2>/dev/null || true
# If the binding failed to build, screen share is simply disabled
# the gateway itself must still start.
if [ ! -f $out/lib/gmw-discord-gateway/native/libdatachannel-min/build/Release/datachannel_min.node ]; then
echo "WARN: datachannel_min.node missing GoLive screen share disabled in this build"
fi
# Also include drizzle migrations if they exist
cp -r drizzle $out/lib/gmw-discord-gateway/ 2>/dev/null || true
@@ -271,7 +196,6 @@ WRAPPER
#!${pkgs.runtimeShell}
cd $out/lib/gmw-discord-gateway
export PATH=${pkgs.ffmpeg-headless}/bin:${pkgs.yt-dlp}/bin:\$PATH
export LD_LIBRARY_PATH=${libdatachannel.out}/lib:\$LD_LIBRARY_PATH
exec ${nodejs}/bin/node dist/index.js
WRAPPER
chmod +x $out/bin/gmw-discord-gateway
+19
View File
@@ -61,6 +61,25 @@ http {
proxy_send_timeout 86400s;
}
# ── Backend oRPC (structured data RPCs over WebSocket + HTTP POST)
# Browser reaches this via partysocket (wss://…/trpc); SSR/RSC uses
# the fetch RPCLink (POST /trpc). Same path, same backend handler:
# oRPC's RPCHandler (HTTP) + ORPCWebSocketServer (WS) on :4001. ──
location ^~ /trpc {
proxy_pass http://gmw_backend$uri$is_args$args;
proxy_http_version 1.1;
# Upgrade headers required for the WebSocket transport; harmless for POST.
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_buffering off;
proxy_read_timeout 86400s;
proxy_send_timeout 86400s;
}
# ── Next.js build assets — immutable, edge/shareable ───────────
location ^~ /_next/static/ {
proxy_pass http://gmw_next$uri$is_args$args;
+10
View File
@@ -0,0 +1,10 @@
-- Migration: add ai_analysis_duration_ms to messages
-- Tracks how long the AI moderation LLM call took, per message (ms).
-- Idempotent: safe to re-run.
--
-- Run against the production GMW database, e.g.:
-- PGPASSWORD=*** psql -h 100.121.180.82 -p 6432 -U asephs -d dcbot \
-- -f scripts/add-ai-analysis-duration.sql
ALTER TABLE "messages"
ADD COLUMN IF NOT EXISTS "ai_analysis_duration_ms" BIGINT;
+2 -1
View File
@@ -15,6 +15,7 @@
},
"dependencies": {
"@discordjs/voice": "^0.19.2",
"@orpc/server": "1.15.0",
"axios": "^1.16.1",
"dotenv": "^17.4.2",
"drizzle-orm": "^0.45.2",
@@ -31,10 +32,10 @@
"@biomejs/biome": "latest",
"@types/express": "^5.0.6",
"@types/node": "^25.9.0",
"@types/pg": "^8.20.0",
"@types/ws": "^8.18.1",
"tsx": "^4.22.2",
"typescript": "^5.9.3",
"@types/pg": "^8.20.0",
"vitest": "latest"
}
}
+2735
View File
File diff suppressed because it is too large Load Diff
+49
View File
@@ -0,0 +1,49 @@
// Rewrite import specifiers in the compiled dist/ so the output runs under
// plain `node dist/index.js` (native ESM, no bundler / no tsx).
//
// Background: tsconfig uses moduleResolution:"bundler", so `tsc` emits BARE
// relative specifiers WITHOUT extensions (e.g. `import "./router"`) and leaves
// the `@/*` path-alias imports untouched. Node's native ESM resolver rejects
// extensionless relative specifiers and knows nothing about the `@/` alias, so
// the emitted dist/ crashes at startup (`ERR_MODULE_NOT_FOUND`). This script
// fixes both:
// 1. `@/foo` -> relative path to dist/foo.js
// 2. `./foo` / `../foo` -> `./foo.js` / `../foo.js` (append .js)
// Already-extensioned relative imports (.js/.json/.node/.mjs/.cjs) and bare
// package specifiers are left untouched (idempotent).
import { readFileSync, writeFileSync, existsSync, readdirSync } from "node:fs";
import { join, relative, dirname } from "node:path";
let count = 0;
function walk(dir) {
if (!existsSync(dir)) return;
for (const e of readdirSync(dir, { withFileTypes: true })) {
const p = join(dir, e.name);
if (e.isDirectory()) walk(p);
else if (e.name.endsWith(".js")) {
const c = readFileSync(p, "utf8");
const pat = /from\s+['"]([^'"]+)['"]/g;
const n = c.replace(pat, (m, spec) => {
if (spec.startsWith("@/")) {
const target = join("dist", spec.slice(2)) + ".js";
let rel = relative(dirname(p), target);
if (!rel.startsWith(".")) rel = "./" + rel;
return `from "${rel}"`;
}
if (
(spec.startsWith("./") || spec.startsWith("../")) &&
!/\.(js|json|node|mjs|cjs)$/.test(spec)
) {
return `from "${spec}.js"`;
}
return m;
});
if (n !== c) {
writeFileSync(p, n);
count++;
}
}
}
}
walk("dist");
console.log(`Fixed ${count} import specifiers in dist/`);
+37 -24
View File
@@ -1,3 +1,5 @@
import { onError } from "@orpc/server";
import { RPCHandler } from "@orpc/server/node";
import express, {
type Express,
type NextFunction,
@@ -6,20 +8,17 @@ import express, {
} from "express";
import helmet from "helmet";
import { createChildLogger } from "@/shared/logger/index";
import { createAnalysisRouter } from "../modules/analysis/index.js";
import { createChatbotRouter } from "../modules/chatbot/index.js";
import { createConfigRouter } from "../modules/config/index.js";
import { createDashboardRouter } from "../modules/dashboard/index.js";
import { createHealthRouter } from "../modules/health/index.js";
import { createMediaRouter } from "../modules/media/index.js";
import { createMessagesRouter } from "../modules/messages/index.js";
import { createModerationRouter } from "../modules/moderation/index.js";
import { createRecordingsRouter } from "../modules/recordings/index.js";
import { createUiStateRouter } from "../modules/ui-state/index.js";
import { createVoiceRouter } from "../modules/voice/index.js";
import { appRouter } from "../orpc/router";
import { errorHandler } from "../shared/middlewares/index.js";
// Auth removed — dashboard is public
// Auth removed — dashboard is public.
// All data APIs (dashboard, messages, moderation, media, voice, recordings,
// analysis, chatbot, config, ui-state) now flow over oRPC, served on TWO
// transports sharing the /trpc path:
// - WebSocket (browser live RPCs) — see orpc/ws.ts
// - HTTP POST (server-side / RSC fetch) — handled below
// Only infra endpoints (health, prometheus metrics) remain plain HTTP.
const logger = createChildLogger("http.app");
@@ -33,7 +32,7 @@ export function createHttpApp(): Express {
}),
);
// Body parsing
// Body parsing (still needed for any JSON POST; oRPC is WS/HTTP-based)
app.use(express.json());
app.use(express.urlencoded({ extended: true }));
@@ -59,24 +58,38 @@ export function createHttpApp(): Express {
next();
});
// All routes are public
// Infra-only HTTP endpoints
app.use("/api", createHealthRouter());
app.use("/api", createConfigRouter());
app.use("/api", createDashboardRouter());
app.use("/api", createMessagesRouter());
app.use("/api", createAnalysisRouter());
app.use("/api", createChatbotRouter());
app.use("/api", createRecordingsRouter());
app.use("/api", createUiStateRouter());
app.use("/api", createMediaRouter());
app.use("/api", createVoiceRouter());
app.use("/api", createModerationRouter());
// oRPC over HTTP (server-side / RSC fetch). The same appRouter the browser
// reaches over the /trpc WebSocket. oRPC's node RPCHandler writes the full
// response itself; if no procedure matched we fall through to the 404 below.
const orpcHandler = new RPCHandler(appRouter, {
interceptors: [onError((error) => logger.error({ error }, "oRPC error"))],
});
app.use((req: Request, res: Response, next: NextFunction) => {
if (!req.path.startsWith("/trpc")) {
next();
return;
}
orpcHandler
.handle(req, res, { prefix: "/trpc", context: {} })
.then(({ matched }) => {
if (!matched) next();
})
.catch((err: unknown) => {
logger.error({ err }, "oRPC HTTP handler failed");
if (!res.headersSent) res.status(500).json({ error: "INTERNAL" });
});
});
// 404 handler
app.use((_req: Request, res: Response) => {
res.status(404).json({
error: "NOT_FOUND",
message: "Endpoint not found",
message:
"Endpoint not found — data APIs are served over /trpc (WebSocket/HTTP)",
});
});
+4 -2
View File
@@ -1,5 +1,6 @@
import { createServer, type Server } from "node:http";
import { createChildLogger } from "@/shared/logger/index";
import { createORPCWebSocketServer } from "../orpc/ws.js";
import { config } from "../shared/config/index.js";
import { initializeDatabase } from "../shared/database/index.js";
import { startRedisBridge } from "../ws/redis-bridge.js";
@@ -16,8 +17,9 @@ export async function startHttpServer(): Promise<Server> {
const server = createServer(app);
// Attach WebSocket server to the same HTTP server
createWebSocketServer(server);
// Attach WebSocket servers to the same HTTP server
createWebSocketServer(server); // /ws — voice PCM + gateway events
createORPCWebSocketServer(server); // /trpc — structured data RPCs
// Start Redis pub/sub bridge to forward discord-gateway events to WS clients
await startRedisBridge();
@@ -1,27 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { analysisService } from "./analysis.service.js";
const logger = createChildLogger("analysis.routes");
export function createAnalysisRouter(): Router {
const router = express.Router();
// GET /api/analysis/search
router.get(
"/analysis/search",
asyncHandler(async (req: Request, res: Response) => {
const q = (req.query.q as string) || "";
const channelId = (req.query.channelId as string) || undefined;
const limit = Number(req.query.limit) || 20;
logger.debug({ q, channelId, limit }, "Analysis search requested");
const result = await analysisService.search({ q, channelId, limit });
res.json(result);
}),
);
return router;
}
@@ -1 +0,0 @@
export { createAnalysisRouter } from "./analysis.routes.js";
@@ -1,96 +0,0 @@
import type { Request, Response } from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { chatbotService } from "./chatbot.service.js";
const logger = createChildLogger("chatbot.controller");
interface AuthenticatedRequest extends Request {
userId?: string;
}
/**
* Resolve the actor id for a request. Frontend (no-login) sends a per-device
* UUID via X-User-Id so chat history stays isolated per visitor; a registered
* auth middleware userId takes precedence when present.
*/
function resolveUserId(req: Request): string {
const authId = (req as AuthenticatedRequest).userId;
if (authId) return authId;
const header = (req.headers["x-user-id"] as string | undefined)?.trim();
return header || "anonymous";
}
export const handleChatbotChat = asyncHandler(
async (req: Request, res: Response) => {
const { message, context } = req.body as {
message: string;
context?: Record<string, unknown>;
};
// Validate required fields
if (!message || typeof message !== "string") {
return res.status(400).json({
error: "INVALID_INPUT",
message: "Message is required and must be a string",
});
}
// Get user ID from X-User-Id header (no-login device uuid) or auth
const userId = resolveUserId(req);
logger.debug(
{ userId, messageLength: message.length, context },
"Received chatbot chat message",
);
// Process message & generate response
const response = await chatbotService.processMessage(
message,
context,
userId,
);
// Save conversation to database
await chatbotService.saveConversation({
userId,
userMessage: message,
botResponse: response,
context,
timestamp: new Date(),
});
logger.info({ userId }, "Chatbot chat processed successfully");
res.status(200).json({
response,
timestamp: new Date().toISOString(),
});
},
);
export const getChatbotHistory = asyncHandler(
async (req: Request, res: Response) => {
const userId = resolveUserId(req);
const limit = Math.min(parseInt(req.query.limit as string, 10) || 50, 100);
const history = await chatbotService.getChatHistory(userId, limit);
res.status(200).json({
history,
total: history.length,
});
},
);
export const clearChatbotHistory = asyncHandler(
async (req: Request, res: Response) => {
const userId = resolveUserId(req);
await chatbotService.clearChatHistory(userId);
res.status(200).json({
message: "Chat history cleared successfully",
});
},
);
@@ -1,6 +1,6 @@
import { and, desc, eq, type SQL, sql } from "drizzle-orm";
import { desc, eq } from "drizzle-orm";
import { getDatabase } from "../../shared/database/index.js";
import { pgChatbotMessagesTable, pgMessagesTable } from "../../shared/index.js";
import { pgChatbotMessagesTable } from "../../shared/index.js";
import { createChildLogger } from "../../shared/logger/index.js";
const logger = createChildLogger("chatbot.repository");
@@ -31,13 +31,6 @@ export interface ChatbotHistoryRow {
created_at: string;
}
export interface ServerInsights {
total_messages: number;
active_users: number;
flagged: number;
warned: number;
}
export class ChatbotRepository {
async saveConversation(input: SaveConversationInput): Promise<void> {
const db = getDatabase();
@@ -83,56 +76,6 @@ export class ChatbotRepository {
"Chat history cleared",
);
}
async getServerInsights(
guildId?: string,
channelId?: string,
): Promise<ServerInsights> {
try {
const db = getDatabase();
const conditions: SQL[] = [];
if (guildId) {
conditions.push(eq(pgMessagesTable.guild_id, guildId));
}
if (channelId) {
conditions.push(eq(pgMessagesTable.channel_id, channelId));
}
const where = conditions.length > 0 ? and(...conditions) : undefined;
const [result] = await db
.select({
total_messages: sql<number>`COUNT(*)::int`,
active_users: sql<number>`COUNT(DISTINCT ${pgMessagesTable.user_id})::int`,
flagged: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'flagged')::int`,
warned: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'warn')::int`,
})
.from(pgMessagesTable)
.where(where);
const insights = result ?? {
total_messages: 0,
active_users: 0,
flagged: 0,
warned: 0,
};
logger.debug({ guildId, channelId, insights }, "Server insights fetched");
return insights;
} catch (error) {
logger.warn(
{ error, guildId, channelId },
"Failed to load server insights",
);
return {
total_messages: 0,
active_users: 0,
flagged: 0,
warned: 0,
};
}
}
}
export const chatbotRepository = new ChatbotRepository();
@@ -1,18 +0,0 @@
import express, { type Router } from "express";
import { validateBody } from "../../shared/middlewares/index.js";
import {
clearChatbotHistory,
getChatbotHistory,
handleChatbotChat,
} from "./chatbot.controller.js";
import { chatRequestSchema } from "./chatbot.schema.js";
export function createChatbotRouter(): Router {
const router = express.Router();
router.post("/chat", validateBody(chatRequestSchema), handleChatbotChat);
router.get("/chat/history", getChatbotHistory);
router.delete("/chat/history", clearChatbotHistory);
return router;
}
@@ -6,7 +6,8 @@ import type {
SaveConversationInput,
} from "./chatbot.repository.js";
import { chatbotRepository } from "./chatbot.repository.js";
import { executeTool, tools } from "./chatbot.tools.js";
import { tools } from "./chatbot.toolDefs.js";
import { executeTool } from "./chatbot.tools.js";
const logger = createChildLogger("chatbot.service");
@@ -17,22 +18,26 @@ class ChatbotService {
userId: string,
): Promise<string> {
logger.info(
{ userId, messageLength: message.length },
{ userId, messageLength: message.length, context },
"processMessage called",
);
const recentContext = await this.getRecentConversationContext(userId);
const serverInsights = await chatbotRepository.getServerInsights(
context?.guildId,
context?.channelId,
);
// Scope the agent to the server/channel the user is chatting in. We no
// longer bake server stats into the prompt — the model must pull current
// data via tools (see buildSystemPrompt), so it always answers from live
// numbers instead of a stale snapshot.
const scope = {
guildId: context?.guildId,
channelId: context?.channelId,
};
// Build LLM messages
const systemPrompt = this.buildSystemPrompt(serverInsights);
const systemPrompt = this.buildSystemPrompt(scope);
const conversationHistory = this.buildHistoryMessages(recentContext);
const llmResponse = await this.callLLM(
systemPrompt,
conversationHistory,
message,
scope,
);
return llmResponse;
@@ -66,27 +71,29 @@ class ChatbotService {
]);
}
private buildSystemPrompt(insights: {
total_messages: number;
active_users: number;
flagged: number;
warned: number;
private buildSystemPrompt(scope: {
guildId?: string;
channelId?: string;
}): string {
return `Kamu lagi ngobrol sama chatbot Discord Watcher — temen ngobrol yang tau keadaan server.
const scopeLine = scope.guildId
? `- Scope: kamu menjawab soal server/guild id="${scope.guildId}"${scope.channelId ? `, channel id="${scope.channelId}"` : ""}.`
: "- Scope: tidak ada guild spesifik — jawab umum soal server ini.";
return `Kamu adalah chatbot Discord Watcher — temen ngobrol yang tau keadaan server, dan kamu PUNYA AKSES ke data server lewat tools.
Data server saat ini:
- Pesan: ${insights.total_messages}
- User aktif: ${insights.active_users}
- Flagged: ${insights.flagged}
- Warning: ${insights.warned}
${scopeLine}
ATURAN PENTING — JANGAN PAKAI KONTEKS STATIS:
- Kamu TIDAK punya hafalan soal angka server (jumlah pesan, user aktif, flagged, dll). JANGAN tebak atau karang angka.
- Untuk SEMUA pertanyaan soal data server (jumlah pesan, user aktif, channel ramai, aktivitas terbaru, pesan di-flag), WAJIB panggil tool yang sesuai (get_server_stats, get_top_channels, get_recent_activity, get_top_flagged). Jawab HANYA dari hasil tool.
- Tool otomatis di-scope ke guild/channel di atas — kalau argumen guildId/channelId kosong, biarkan kosong (sudah otomatis ter-isi). Jangan isi ID yang kamu tebak.
- Kalau tool balas error atau kosong, bilang aja data lagi ga ketemu, jangan karang.
Gaya ngobrol:
- Santai, hangat, kayak ngobrol sama temen
- Pake Bahasa Indonesia sehari-hari, ga perlu kaku
- Sesekali pake emoji wajar aja, ga berlebihan
- Kalo ditanya sesuatu yang kamu tau dari data server, jawab pake data itu
- Kalo ga tau atau ga nyambung, bilang aja terus tanya balik biar ngobrolnya jalan
- Jangan sebut "rule", "instruksi", "prompt" atau apapun soal cara kamu berpikir
- Kalo ditanya di luar data server dan kamu ga tau, bilang aja terus tanya balik biar ngobrolnya jalan
- Jangan sebut "rule", "instruksi", "prompt", "tool", atau apapun soal cara kamu berpikir
- Biasa aja, ga usaha lucu-lucu amat — natural`;
}
@@ -106,6 +113,7 @@ Gaya ngobrol:
systemPrompt: string,
history: Array<{ role: "user" | "assistant"; content: string }>,
userMessage: string,
scope: { guildId?: string; channelId?: string },
): Promise<string> {
const apiKey = config.AI_LLM_API_KEY;
const baseUrl = config.AI_LLM_BASE_URL;
@@ -150,7 +158,13 @@ Gaya ngobrol:
tool_choice: "auto",
max_tokens: 600,
temperature: 0.4,
stream: true,
// Non-streaming: request a single complete response. 9router may
// still emit SSE even with stream:false, so the parser below
// handle both raw-JSON and SSE bodies.
stream: false,
// Disable extended thinking / reasoning tokens so the bot answers
// directly (ignored by non-reasoning models).
reasoning_effort: "none",
},
{
headers: {
@@ -158,14 +172,16 @@ Gaya ngobrol:
"Content-Type": "application/json",
},
timeout: 45_000,
// 9router returns SSE even without stream:true; force stream:true
// in the body and read the raw SSE text.
responseType: "text",
},
);
// Parse SSE `data:` lines → content + tool_calls.
const { content, toolCalls } = this.parseSse(response.data as string);
// Parse the body into content + tool_calls. 9router may return either
// a single JSON object (stream:false honored) or SSE text (stream
// implied) — parseResponse handles both.
const { content, toolCalls } = this.parseResponse(
response.data as string,
);
logger.debug(
{
@@ -190,9 +206,19 @@ Gaya ngobrol:
},
],
});
// Auto-scope: if the model omitted guildId/channelId, fill them
// from the request scope so tools query the right server without
// the model having to guess IDs.
const scopedArgs = { ...tc.args };
if (scope.guildId && scopedArgs.guildId == null) {
scopedArgs.guildId = scope.guildId;
}
if (scope.channelId && scopedArgs.channelId == null) {
scopedArgs.channelId = scope.channelId;
}
let result = "";
try {
result = await executeTool(tc.name, tc.args);
result = await executeTool(tc.name, scopedArgs);
} catch (e) {
result = `Tool error: ${(e as Error).message}`;
}
@@ -224,6 +250,60 @@ Gaya ngobrol:
}
}
/**
* Parse an LLM HTTP body into content + tool_calls. Handles both shapes
* 9router can return: a single JSON object (stream:false honored) or SSE
* text (stream implied). For SSE we delegate to parseSse.
*/
private parseResponse(body: string): {
content: string;
toolCalls: Array<{
id: string;
name: string;
arguments: string;
args: Record<string, unknown>;
}>;
} {
const trimmed = body.trim();
// Non-streaming response: a single JSON object.
if (trimmed.startsWith("{")) {
try {
const json = JSON.parse(trimmed) as {
choices?: Array<{
message?: {
content?: string | null;
tool_calls?: Array<{
id?: string;
type?: string;
function?: { name?: string; arguments?: string };
}>;
};
delta?: unknown;
}>;
};
const msg = json.choices?.[0]?.message;
// If the router returned SSE-style shape under `choices[].delta`
// (rare), fall through to the SSE parser.
if (msg) {
const content = msg.content ?? "";
const toolCalls = (msg.tool_calls ?? []).map((tc, i) => {
const id = tc.id || `tool_${i}_${Date.now()}`;
return {
id,
name: tc.function?.name ?? "",
arguments: tc.function?.arguments ?? "",
args: this.safeJsonParse(tc.function?.arguments ?? ""),
};
});
return { content: content.trim(), toolCalls };
}
} catch {
// Not valid JSON after all — treat as SSE below.
}
}
return this.parseSse(body);
}
/**
* Parse an SSE stream body into accumulated content + any tool_calls.
* 9router (and most OpenAI-compatible routers) emit `data: {json}` lines
@@ -0,0 +1,270 @@
/**
* Static tool *definitions* for the chatbot LLM (OpenAI function-calling
* format). Kept separate from the executor (chatbot.tools.ts) so the schema
* the model depends on can be imported without pulling in the database /
* config layer.
*
* The chatbot is a server-watcher agent: it can answer about ANY server
* situation — activity, moderation queue, specific users, channels, voice
* recordings, AI correction history, and trends over time — by calling these
* tools, which the executor implements against real tables.
*/
export interface ToolDef {
type: "function";
function: {
name: string;
description: string;
parameters: {
type: "object";
properties: Record<string, unknown>;
required?: string[];
};
};
}
export const tools: ToolDef[] = [
{
type: "function",
function: {
name: "get_server_stats",
description:
"Ambil statistik ringkas server/guild: total pesan, user aktif, jumlah pesan flagged, warn, dan clean. Panggil untuk jawab pertanyaan umum soal kondisi server. guildId/channelId otomatis ter-isi dari scope; kosongkan untuk semua data.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
channelId: { type: "string", description: "ID channel (opsional)." },
},
},
},
},
{
type: "function",
function: {
name: "get_top_channels",
description:
"Ambil daftar channel paling aktif (jumlah pesan terbanyak). Panggil untuk 'channel mana paling ramai' atau aktivitas per-channel.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
limit: {
type: "number",
description: "Jumlah channel teratas (default 5, max 10).",
},
},
},
},
},
{
type: "function",
function: {
name: "get_recent_activity",
description:
"Ambil pesan terbaru di server: siapa, di channel mana, jam berapa, isinya. Panggil untuk 'lagi ngapain' / aktivitas terbaru.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
channelId: { type: "string", description: "ID channel (opsional)." },
limit: {
type: "number",
description: "Jumlah pesan terakhir (default 5, max 20).",
},
},
},
},
},
{
type: "function",
function: {
name: "get_top_flagged",
description:
"Ambil pesan dengan ai_status flagged (beserta alasan, severity, analysis). Panggil untuk bahas pesan bermasalah / kerjaan moderator.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
channelId: { type: "string", description: "ID channel (opsional)." },
limit: { type: "number", description: "Jumlah pesan (default 5)." },
},
},
},
},
{
type: "function",
function: {
name: "search_messages",
description:
"Cari pesan berdasarkan kata kunci di isi pesan (case-insensitive, LIKE). Untuk 'ada yang bahas X gak?' / temukan topik tertentu. Hindari kata terlalu umum.",
parameters: {
type: "object",
properties: {
query: {
type: "string",
description: "Kata kunci pencarian (wajib).",
},
guildId: { type: "string", description: "ID server (opsional)." },
channelId: { type: "string", description: "ID channel (opsional)." },
limit: { type: "number", description: "Jumlah hasil (default 5)." },
},
required: ["query"],
},
},
},
{
type: "function",
function: {
name: "get_user_messages",
description:
"Ambil pesan terbaru dari satu user tertentu (user_id), opsional di-scope ke guild/channel. Untuk 'chat si A gimana akhir-akhir ini?' — butuh user_id.",
parameters: {
type: "object",
properties: {
userId: { type: "string", description: "ID user (wajib)." },
guildId: { type: "string", description: "ID server (opsional)." },
channelId: { type: "string", description: "ID channel (opsional)." },
limit: { type: "number", description: "Jumlah pesan (default 10)." },
},
required: ["userId"],
},
},
},
{
type: "function",
function: {
name: "get_user_profile",
description:
"Ambil ringkasan profil AI dari seorang user (pola perilaku, gaya bicara) dari tabel user_profiles. Untuk 'siapa si A?' / konteks perilaku. Butuh user_id.",
parameters: {
type: "object",
properties: {
userId: { type: "string", description: "ID user (wajib)." },
guildId: { type: "string", description: "ID server (opsional)." },
},
required: ["userId"],
},
},
},
{
type: "function",
function: {
name: "get_user_reputation",
description:
"Ambil skor trust, jumlah infraction, dan streak pesan bersih seorang user dari user_reputations. Untuk 'berapa trust score si A?' / riwayat pelanggaran. Butuh user_id.",
parameters: {
type: "object",
properties: {
userId: { type: "string", description: "ID user (wajib)." },
guildId: { type: "string", description: "ID server (opsional)." },
},
required: ["userId"],
},
},
},
{
type: "function",
function: {
name: "get_channel_culture",
description:
"Ambil ringkasan norma/slang channel dari tabel channel_cultures (AI-generated). Untuk 'norma channel ini gimana?' / konteks sebelum nge-flag. Butuh channel_id.",
parameters: {
type: "object",
properties: {
channelId: { type: "string", description: "ID channel (wajib)." },
},
required: ["channelId"],
},
},
},
{
type: "function",
function: {
name: "get_message_detail",
description:
"Ambil 1 pesan lengkap beserta hasil analisis AI-nya (status, flags, score, severity, kategori, analysis, recommended action). Untuk jelasin keputusan moderasi pada pesan tertentu. Butuh message_id.",
parameters: {
type: "object",
properties: {
messageId: { type: "string", description: "ID pesan (wajib)." },
},
required: ["messageId"],
},
},
},
{
type: "function",
function: {
name: "get_message_reviews",
description:
"Ambil antrean review moderasi manual (message_reviews) berdasarkan status: pending/approved/rejected/escalated. Untuk 'ada review moderasi pending?' / cek kerjaan human moderator. guildId otomatis ter-isi.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
status: {
type: "string",
description:
"Status review: pending / approved / rejected / escalated (opsional, default semua).",
},
limit: { type: "number", description: "Jumlah (default 10)." },
},
},
},
},
{
type: "function",
function: {
name: "get_voice_recordings",
description:
"Ambil rekaman suara terbaru (voice_recordings): user, channel, transkripsi, status upload. Untuk 'ada rekaman suara terbaru?' / cek transkripsi. Bisa di-scope ke user_id atau channel_id.",
parameters: {
type: "object",
properties: {
userId: { type: "string", description: "Filter user (opsional)." },
channelId: {
type: "string",
description: "Filter channel (opsional).",
},
guildId: { type: "string", description: "ID server (opsional)." },
limit: { type: "number", description: "Jumlah (default 10)." },
},
},
},
},
{
type: "function",
function: {
name: "get_moderation_timeline",
description:
"Ambil tren harian: per hari, jumlah total pesan vs flagged vs warn vs clean. Untuk 'minggu ini pelanggaran naik?' / lihat tren moderasi. guildId otomatis ter-isi.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
channelId: { type: "string", description: "ID channel (opsional)." },
days: {
type: "number",
description: "Jumlah hari ke belakang (default 14, max 60).",
},
},
},
},
},
{
type: "function",
function: {
name: "get_corrections",
description:
"Ambil riwayat koreksi false-positive AI (corrected_moderations): pesan yang awalnya di-flag tapi dikoreksi manusia, beserta alasannya. Untuk 'AI pernah salah nge-flag apa aja?' / audit akurasi moderasi.",
parameters: {
type: "object",
properties: {
guildId: { type: "string", description: "ID server (opsional)." },
limit: { type: "number", description: "Jumlah (default 10)." },
},
},
},
},
];
@@ -1,116 +1,25 @@
import { sql } from "drizzle-orm";
import { and, desc, eq, like, sql } from "drizzle-orm";
import { getDatabase } from "../../shared/database/index.js";
import {
pgChannelCulturesTable,
pgCorrectedModerationsTable,
pgMessageReviewsTable,
pgMessagesTable,
pgUserProfilesTable,
pgUserReputationsTable,
pgVoiceRecordingsTable,
} from "../../shared/index.js";
/**
* Tools the chatbot LLM can call. Definitions describe the schema to the
* model; the executor implements each one against the real database.
* This turns the chatbot from "blind stats guesser" into an agent that
* pulls real, current server data on demand.
* Executor for the chatbot's server-watcher tools. The tool *definitions*
* live in chatbot.toolDefs.ts (no DB import); this file implements each one
* against the real database.
*
* All queries use parameterized drizzle operators (eq/like/and) — never string
* interpolation into raw SQL — so model-supplied arguments cannot inject SQL.
*/
export type ToolResult = string;
/** JSON schema for a tool definition (OpenAI function-calling format). */
export interface ToolDef {
type: "function";
function: {
name: string;
description: string;
parameters: {
type: "object";
properties: Record<string, unknown>;
required?: string[];
};
};
}
export const tools: ToolDef[] = [
{
type: "function",
function: {
name: "get_server_stats",
description:
"Ambil statistik ringkas server/guild saat ini: total pesan, user aktif, jumlah pesan flagged, dan jumlah warning. Panggil ini untuk menjawab pertanyaan umum tentang kondisi server. Opsional fill guild_id untuk scope ke guild tertentu, channel_id untuk scope ke channel.",
parameters: {
type: "object",
properties: {
guildId: {
type: "string",
description: "ID guild/server (opsional). Kosongkan = semua data.",
},
channelId: {
type: "string",
description: "ID channel (opsional).",
},
},
},
},
},
{
type: "function",
function: {
name: "get_top_channels",
description:
"Ambil daftar channel paling aktif (jumlah pesan terbanyak) di server. Panggil buat jawab 'channel mana paling ramai' atau aktivitas per-channel.",
parameters: {
type: "object",
properties: {
guildId: {
type: "string",
description: "ID server (opsional).",
},
limit: {
type: "number",
description: "Jumlah channel teratas (default 5, max 10).",
},
},
},
},
},
{
type: "function",
function: {
name: "get_recent_activity",
description:
"Ambil aktivitas/pesan terbaru di server: siapa yang baru ngomong, di channel mana, jam berapa. Panggil buat jawaban soal 'lagi ngapain' / aktivitas terbaru di server.",
parameters: {
type: "object",
properties: {
guildId: {
type: "string",
description: "ID server (opsional).",
},
limit: {
type: "number",
description: "Jumlah pesan terakhir (default 5).",
},
},
},
},
},
{
type: "function",
function: {
name: "get_top_flagged",
description:
"Ambil pesan yang paling sering di-flag atau kena warning. Panggil buat jawab soal pesan bermasalah / moderator.",
parameters: {
type: "object",
properties: {
guildId: {
type: "string",
description: "ID server (opsional).",
},
limit: {
type: "number",
description: "Jumlah pesan (default 5).",
},
},
},
},
},
];
/** Executes a tool call against the real DB and returns a readable result. */
export async function executeTool(
name: string,
@@ -122,9 +31,11 @@ export async function executeTool(
typeof args.channelId === "string" && args.channelId
? args.channelId
: undefined;
const userId =
typeof args.userId === "string" && args.userId ? args.userId : undefined;
const limitRaw =
typeof args.limit === "number" ? args.limit : Number(args.limit) || 5;
const limit = Math.min(Math.max(1, Math.round(limitRaw)), 10);
const limit = Math.min(Math.max(1, Math.round(limitRaw)), 20);
try {
switch (name) {
@@ -133,9 +44,48 @@ export async function executeTool(
case "get_top_channels":
return await topChannels(guildId, limit);
case "get_recent_activity":
return await recentActivity(guildId, limit);
return await recentActivity(guildId, channelId, limit);
case "get_top_flagged":
return await topFlagged(guildId, limit);
return await topFlagged(guildId, channelId, limit);
case "search_messages":
return await searchMessages(
String(args.query ?? ""),
guildId,
channelId,
limit,
);
case "get_user_messages":
return await userMessages(userId, guildId, channelId, limit);
case "get_user_profile":
return await userProfile(userId, guildId);
case "get_user_reputation":
return await userReputation(userId, guildId);
case "get_channel_culture":
return await channelCulture(
typeof args.channelId === "string" ? args.channelId : undefined,
);
case "get_message_detail":
return await messageDetail(
typeof args.messageId === "string" ? args.messageId : undefined,
);
case "get_message_reviews":
return await messageReviews(
guildId,
typeof args.status === "string" ? args.status : undefined,
limit,
);
case "get_voice_recordings":
return await voiceRecordings(userId, channelId, guildId, limit);
case "get_moderation_timeline":
return await moderationTimeline(
guildId,
channelId,
typeof args.days === "number"
? Math.min(Math.max(1, args.days), 60)
: 14,
);
case "get_corrections":
return await corrections(guildId, limit);
default:
return `Unknown tool: ${name}`;
}
@@ -145,6 +95,23 @@ export async function executeTool(
}
}
// ── Query helpers ──────────────────────────────────────────
function scopeMessages(
guildId?: string,
channelId?: string,
): ReturnType<typeof and> | undefined {
const conds = [];
if (guildId) conds.push(eq(pgMessagesTable.guild_id, guildId));
if (channelId) conds.push(eq(pgMessagesTable.channel_id, channelId));
return conds.length ? and(...conds) : undefined;
}
/** Escape LIKE wildcards so user input can't break the pattern. */
function likePattern(q: string): string {
return q.replace(/[\\%_]/g, (c) => `\\${c}`);
}
// ── Tool executors ──────────────────────────────────────────
async function serverStats(
@@ -152,81 +119,330 @@ async function serverStats(
channelId?: string,
): Promise<string> {
const db = getDatabase();
const conditions: string[] = [];
if (guildId) conditions.push(`guild_id = '${guildId}'`);
if (channelId) conditions.push(`channel_id = '${channelId}'`);
const cond = conditions.length ? `WHERE ${conditions.join(" AND ")}` : "";
const [result] = await db
.select({
total_messages: sql<number>`COUNT(*)::int`,
active_users: sql<number>`COUNT(DISTINCT ${pgMessagesTable.user_id})::int`,
flagged: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'flagged')::int`,
warned: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'warn')::int`,
clean: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'clean')::int`,
})
.from(pgMessagesTable)
.where(scopeMessages(guildId, channelId));
const result = await db.execute(
sql.raw(
`SELECT COUNT(*)::int AS total_messages,
COUNT(DISTINCT user_id)::int AS active_users,
COUNT(*) FILTER (WHERE ai_status = 'flagged')::int AS flagged,
COUNT(*) FILTER (WHERE ai_status = 'warn')::int AS warned
FROM messages ${cond}`,
),
);
const rows =
(result as unknown as { rows: Record<string, unknown>[] }).rows ?? [];
const r = rows[0] ?? {};
return JSON.stringify({
total_messages: r.total_messages ?? 0,
active_users: r.active_users ?? 0,
flagged: r.flagged ?? 0,
warned: r.warned ?? 0,
});
const r = result ?? {
total_messages: 0,
active_users: 0,
flagged: 0,
warned: 0,
clean: 0,
};
return JSON.stringify(r);
}
async function topChannels(guildId?: string, limit = 5): Promise<string> {
const db = getDatabase();
const conditions: string[] = [];
if (guildId) conditions.push(`guild_id = '${guildId}'`);
const cond = conditions.length ? `WHERE ${conditions.join(" AND ")}` : "";
const result = await db.execute(
sql.raw(
`SELECT channel_id,
COUNT(*)::int AS count
FROM messages ${cond}
GROUP BY channel_id
ORDER BY count DESC
LIMIT ${limit}`,
),
);
const rows = (result as unknown as { rows: unknown[] }).rows ?? [];
return JSON.stringify(rows.slice(0, limit));
const rows = await db
.select({
channel_id: pgMessagesTable.channel_id,
count: sql<number>`COUNT(*)::int`,
})
.from(pgMessagesTable)
.where(scopeMessages(guildId))
.groupBy(pgMessagesTable.channel_id)
.orderBy(desc(sql`COUNT(*)`))
.limit(limit);
return JSON.stringify(rows);
}
async function recentActivity(guildId?: string, limit = 5): Promise<string> {
async function recentActivity(
guildId?: string,
channelId?: string,
limit = 5,
): Promise<string> {
const db = getDatabase();
const conditions: string[] = [];
if (guildId) conditions.push(`guild_id = '${guildId}'`);
const cond = conditions.length ? `WHERE ${conditions.join(" AND ")}` : "";
const result = await db.execute(
sql.raw(
`SELECT username, content, channel_id, created_at
FROM messages ${cond}
ORDER BY created_at DESC
LIMIT ${limit}`,
),
);
return JSON.stringify((result as unknown as { rows: unknown[] }).rows ?? []);
const rows = await db
.select({
id: pgMessagesTable.id,
username: pgMessagesTable.username,
user_id: pgMessagesTable.user_id,
channel_id: pgMessagesTable.channel_id,
content: pgMessagesTable.content,
created_at: pgMessagesTable.created_at,
ai_status: pgMessagesTable.ai_status,
})
.from(pgMessagesTable)
.where(scopeMessages(guildId, channelId))
.orderBy(desc(pgMessagesTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
async function topFlagged(guildId?: string, limit = 5): Promise<string> {
async function topFlagged(
guildId?: string,
channelId?: string,
limit = 5,
): Promise<string> {
const db = getDatabase();
const conditions = ["ai_status IN ('flagged', 'warn')"];
if (guildId) conditions.push(`guild_id = '${guildId}'`);
const cond = `WHERE ${conditions.join(" AND ")}`;
const result = await db.execute(
sql.raw(
`SELECT username, content, channel_id, ai_status, created_at
FROM messages ${cond}
ORDER BY created_at DESC
LIMIT ${limit}`,
),
);
return JSON.stringify((result as unknown as { rows: unknown[] }).rows ?? []);
const rows = await db
.select({
id: pgMessagesTable.id,
username: pgMessagesTable.username,
channel_id: pgMessagesTable.channel_id,
content: pgMessagesTable.content,
ai_status: pgMessagesTable.ai_status,
ai_severity: pgMessagesTable.ai_severity,
ai_moderation_flags: pgMessagesTable.ai_moderation_flags,
ai_analysis: pgMessagesTable.ai_analysis,
created_at: pgMessagesTable.created_at,
})
.from(pgMessagesTable)
.where(
and(
scopeMessages(guildId, channelId),
eq(pgMessagesTable.ai_status, "flagged"),
),
)
.orderBy(desc(pgMessagesTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
async function searchMessages(
query: string,
guildId?: string,
channelId?: string,
limit = 5,
): Promise<string> {
const db = getDatabase();
if (!query.trim()) return JSON.stringify({ error: "query kosong" });
const rows = await db
.select({
id: pgMessagesTable.id,
username: pgMessagesTable.username,
channel_id: pgMessagesTable.channel_id,
content: pgMessagesTable.content,
created_at: pgMessagesTable.created_at,
ai_status: pgMessagesTable.ai_status,
})
.from(pgMessagesTable)
.where(
and(
scopeMessages(guildId, channelId),
like(pgMessagesTable.content, `%${likePattern(query)}%`),
),
)
.orderBy(desc(pgMessagesTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
async function userMessages(
userId?: string,
guildId?: string,
channelId?: string,
limit = 10,
): Promise<string> {
const db = getDatabase();
if (!userId) return JSON.stringify({ error: "userId wajib" });
const conds = [eq(pgMessagesTable.user_id, userId)];
if (guildId) conds.push(eq(pgMessagesTable.guild_id, guildId));
if (channelId) conds.push(eq(pgMessagesTable.channel_id, channelId));
const rows = await db
.select({
id: pgMessagesTable.id,
channel_id: pgMessagesTable.channel_id,
content: pgMessagesTable.content,
created_at: pgMessagesTable.created_at,
ai_status: pgMessagesTable.ai_status,
})
.from(pgMessagesTable)
.where(and(...conds))
.orderBy(desc(pgMessagesTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
async function userProfile(userId?: string, guildId?: string): Promise<string> {
const db = getDatabase();
if (!userId) return JSON.stringify({ error: "userId wajib" });
const conds = [eq(pgUserProfilesTable.user_id, userId)];
if (guildId) conds.push(eq(pgUserProfilesTable.guild_id, guildId));
const rows = await db
.select({
user_id: pgUserProfilesTable.user_id,
guild_id: pgUserProfilesTable.guild_id,
profile_summary: pgUserProfilesTable.profile_summary,
last_analyzed_at: pgUserProfilesTable.last_analyzed_at,
})
.from(pgUserProfilesTable)
.where(and(...conds))
.limit(1);
return JSON.stringify(rows[0] ?? { error: "profil tidak ditemukan" });
}
async function userReputation(
userId?: string,
guildId?: string,
): Promise<string> {
const db = getDatabase();
if (!userId) return JSON.stringify({ error: "userId wajib" });
const conds = [eq(pgUserReputationsTable.user_id, userId)];
if (guildId) conds.push(eq(pgUserReputationsTable.guild_id, guildId));
const rows = await db
.select({
user_id: pgUserReputationsTable.user_id,
guild_id: pgUserReputationsTable.guild_id,
trust_score: pgUserReputationsTable.trust_score,
clean_message_streak: pgUserReputationsTable.clean_message_streak,
total_infractions: pgUserReputationsTable.total_infractions,
last_infraction_at: pgUserReputationsTable.last_infraction_at,
})
.from(pgUserReputationsTable)
.where(and(...conds))
.limit(1);
return JSON.stringify(rows[0] ?? { error: "reputasi tidak ditemukan" });
}
async function channelCulture(channelId?: string): Promise<string> {
const db = getDatabase();
if (!channelId) return JSON.stringify({ error: "channelId wajib" });
const rows = await db
.select({
channel_id: pgChannelCulturesTable.channel_id,
culture_summary: pgChannelCulturesTable.culture_summary,
last_analyzed_at: pgChannelCulturesTable.last_analyzed_at,
})
.from(pgChannelCulturesTable)
.where(eq(pgChannelCulturesTable.channel_id, channelId))
.limit(1);
return JSON.stringify(rows[0] ?? { error: "culture tidak ditemukan" });
}
async function messageDetail(messageId?: string): Promise<string> {
const db = getDatabase();
if (!messageId) return JSON.stringify({ error: "messageId wajib" });
const rows = await db
.select({
id: pgMessagesTable.id,
guild_id: pgMessagesTable.guild_id,
channel_id: pgMessagesTable.channel_id,
user_id: pgMessagesTable.user_id,
username: pgMessagesTable.username,
content: pgMessagesTable.content,
created_at: pgMessagesTable.created_at,
ai_status: pgMessagesTable.ai_status,
ai_moderation_flags: pgMessagesTable.ai_moderation_flags,
ai_moderation_score: pgMessagesTable.ai_moderation_score,
ai_severity: pgMessagesTable.ai_severity,
ai_categories: pgMessagesTable.ai_categories,
ai_analysis: pgMessagesTable.ai_analysis,
ai_recommended_action: pgMessagesTable.ai_recommended_action,
ai_confidence: pgMessagesTable.ai_confidence,
})
.from(pgMessagesTable)
.where(eq(pgMessagesTable.id, messageId))
.limit(1);
return JSON.stringify(rows[0] ?? { error: "pesan tidak ditemukan" });
}
async function messageReviews(
guildId?: string,
status?: string,
limit = 10,
): Promise<string> {
const db = getDatabase();
const conds = [];
if (guildId) conds.push(eq(pgMessageReviewsTable.guild_id, guildId));
if (status) conds.push(eq(pgMessageReviewsTable.status, status as never));
const rows = await db
.select({
id: pgMessageReviewsTable.id,
message_id: pgMessageReviewsTable.message_id,
reviewer_id: pgMessageReviewsTable.reviewer_id,
status: pgMessageReviewsTable.status,
notes: pgMessageReviewsTable.notes,
created_at: pgMessageReviewsTable.created_at,
reviewed_at: pgMessageReviewsTable.reviewed_at,
})
.from(pgMessageReviewsTable)
.where(conds.length ? and(...conds) : undefined)
.orderBy(desc(pgMessageReviewsTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
async function voiceRecordings(
userId?: string,
channelId?: string,
guildId?: string,
limit = 10,
): Promise<string> {
const db = getDatabase();
const conds = [];
if (userId) conds.push(eq(pgVoiceRecordingsTable.user_id, userId));
if (channelId) conds.push(eq(pgVoiceRecordingsTable.channel_id, channelId));
if (guildId) conds.push(eq(pgVoiceRecordingsTable.guild_id, guildId));
const rows = await db
.select({
id: pgVoiceRecordingsTable.id,
username: pgVoiceRecordingsTable.username,
channel_name: pgVoiceRecordingsTable.channel_name,
filename: pgVoiceRecordingsTable.filename,
size_bytes: pgVoiceRecordingsTable.size_bytes,
upload_status: pgVoiceRecordingsTable.upload_status,
transcription: pgVoiceRecordingsTable.transcription,
created_at: pgVoiceRecordingsTable.created_at,
})
.from(pgVoiceRecordingsTable)
.where(conds.length ? and(...conds) : undefined)
.orderBy(desc(pgVoiceRecordingsTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
async function moderationTimeline(
guildId?: string,
channelId?: string,
days = 14,
): Promise<string> {
const db = getDatabase();
const day = sql<string>`to_char(to_timestamp(${pgMessagesTable.created_at} / 1000), 'YYYY-MM-DD')`;
const rows = await db
.select({
day,
total: sql<number>`COUNT(*)::int`,
flagged: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'flagged')::int`,
warned: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'warn')::int`,
clean: sql<number>`COUNT(*) FILTER (WHERE ${pgMessagesTable.ai_status} = 'clean')::int`,
})
.from(pgMessagesTable)
.where(
and(
scopeMessages(guildId, channelId),
// only the last N days
sql`${pgMessagesTable.created_at} >= extract(epoch FROM now() - (${days} || ' days')::interval) * 1000`,
),
)
.groupBy(day)
.orderBy(day);
return JSON.stringify(rows);
}
async function corrections(_guildId?: string, limit = 10): Promise<string> {
const db = getDatabase();
const rows = await db
.select({
id: pgCorrectedModerationsTable.id,
message_id: pgCorrectedModerationsTable.message_id,
original_flags: pgCorrectedModerationsTable.original_flags,
corrected_flags: pgCorrectedModerationsTable.corrected_flags,
correction_notes: pgCorrectedModerationsTable.correction_notes,
content_snippet: pgCorrectedModerationsTable.content_snippet,
created_at: pgCorrectedModerationsTable.created_at,
})
.from(pgCorrectedModerationsTable)
.orderBy(desc(pgCorrectedModerationsTable.created_at))
.limit(limit);
return JSON.stringify(rows);
}
@@ -1 +0,0 @@
export { createChatbotRouter } from "./chatbot.routes.js";
@@ -1,28 +0,0 @@
import type { Router } from "express";
import express from "express";
import { config } from "../../shared/config/index.js";
export function createConfigRouter(): Router {
const router = express.Router();
// GET /api/config
router.get("/config", (_req, res) => {
res.json({
monitorGuildId: config.MONITOR_GUILD_ID || null,
webserverPort: config.WEBSERVER_PORT,
nodeEnv: config.NODE_ENV,
backlogSyncHours: config.BACKLOG_SYNC_HOURS,
backlogSyncBatchSize: config.BACKLOG_SYNC_BATCH_SIZE,
retentionMessagesDays: config.RETENTION_MESSAGES_DAYS,
retentionAttachmentsDays: config.RETENTION_ATTACHMENTS_DAYS,
retentionVoiceDays: config.RETENTION_VOICE_DAYS,
autoDeleteFlaggedEnabled: config.AUTO_DELETE_FLAGGED_ENABLED,
aiAnalysisEnabled: config.AI_ANALYSIS_ENABLED,
voiceGuildId: config.VOICE_GUILD_ID || null,
voiceChannelId: config.VOICE_CHANNEL_ID || null,
logLevel: config.LOG_LEVEL,
});
});
return router;
}
@@ -1 +0,0 @@
export { createConfigRouter } from "./config.routes.js";
@@ -1,111 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { dashboardService } from "./dashboard.service.js";
const logger = createChildLogger("dashboard.routes");
export function createDashboardRouter(): Router {
const router = express.Router();
// GET /api/dashboard/stats — aggregated server statistics
router.get(
"/dashboard/stats",
asyncHandler(async (_req: Request, res: Response) => {
logger.debug("Fetching dashboard stats");
const stats = await dashboardService.getStats();
res.json(stats);
}),
);
// GET /api/dashboard/activity?days=14 — message volume over time
router.get(
"/dashboard/activity",
asyncHandler(async (req: Request, res: Response) => {
const days = Math.min(Math.max(Number(req.query.days) || 14, 1), 90);
const activity = await dashboardService.getActivity(days);
res.json(activity);
}),
);
// GET /api/dashboard/users — paginated user list with profiles
router.get(
"/dashboard/users",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 20;
const cursor =
typeof req.query.cursor === "string" ? req.query.cursor : undefined;
const search =
typeof req.query.search === "string" ? req.query.search : undefined;
const result = await dashboardService.listUsers({
limit,
cursor,
search,
});
res.json(result);
}),
);
// GET /api/dashboard/users/:userId — single user detail
router.get(
"/dashboard/users/:userId",
asyncHandler(async (req: Request, res: Response) => {
const userId = String(req.params.userId);
const detail = await dashboardService.getUserDetail(userId);
res.json(detail);
}),
);
// GET /api/dashboard/channels — paginated channel list with culture summaries
router.get(
"/dashboard/channels",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 20;
const search =
typeof req.query.search === "string" ? req.query.search : undefined;
const guildId =
typeof req.query.guild_id === "string" ? req.query.guild_id : undefined;
const result = await dashboardService.listChannels({
limit,
search,
guildId,
});
res.json(result);
}),
);
// GET /api/dashboard/channels/:channelId — single channel detail
router.get(
"/dashboard/channels/:channelId",
asyncHandler(async (req: Request, res: Response) => {
const channelId = String(req.params.channelId);
const detail = await dashboardService.getChannelDetail(channelId);
res.json(detail);
}),
);
// GET /api/dashboard/reactions — top reacted messages
router.get(
"/dashboard/reactions",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 20;
const reactions = await dashboardService.getTopReactions(limit);
res.json(reactions);
}),
);
// GET /api/dashboard/reactors — top users by reactions given
router.get(
"/dashboard/reactors",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 20;
const reactors = await dashboardService.getTopReactors(limit);
res.json(reactors);
}),
);
return router;
}
@@ -1 +0,0 @@
export { createDashboardRouter } from "./dashboard.routes.js";
@@ -1 +0,0 @@
export { createMediaRouter } from "./media.routes.js";
@@ -1,71 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler, validateBody } from "../../shared/middlewares/index.js";
import { mediaLoopSchema, mediaQueueSchema } from "./media.schema.js";
import { getStatus, queue, setLoop, skip, stop } from "./media.service.js";
const logger = createChildLogger("media.routes");
export function createMediaRouter(): Router {
const router = express.Router();
// GET /api/media/status
router.get(
"/media/status",
asyncHandler(async (_req: Request, res: Response) => {
logger.debug("Media status requested");
const status = await getStatus();
res.json(status);
}),
);
// POST /api/media/queue
router.post(
"/media/queue",
validateBody(mediaQueueSchema),
asyncHandler(async (req: Request, res: Response) => {
const { source, mode } = req.body as {
source: string;
mode: "music" | "screen";
};
logger.debug({ source, mode }, "Media queue requested");
const state = await queue(source, mode);
res.json(state);
}),
);
// POST /api/media/skip
router.post(
"/media/skip",
asyncHandler(async (_req: Request, res: Response) => {
logger.debug("Media skip requested");
const state = await skip();
res.json(state);
}),
);
// POST /api/media/stop
router.post(
"/media/stop",
asyncHandler(async (_req: Request, res: Response) => {
logger.debug("Media stop requested");
const state = await stop();
res.json(state);
}),
);
// POST /api/media/loop
router.post(
"/media/loop",
validateBody(mediaLoopSchema),
asyncHandler(async (req: Request, res: Response) => {
const { loop } = req.body as { loop: boolean };
logger.debug({ loop }, "Media loop requested");
const state = await setLoop(loop);
res.json(state);
}),
);
return router;
}
@@ -1 +0,0 @@
export { createMessagesRouter } from "./messages.routes.js";
@@ -1,74 +0,0 @@
import type { Request, Response } from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { messageQuerySchema } from "./messages.schema.js";
import { messagesService } from "./messages.service.js";
const logger = createChildLogger("messages.controller");
export const handleListMessages = asyncHandler(
async (req: Request, res: Response) => {
const query = messageQuerySchema.parse(req.query);
logger.debug({ query }, "Handling list messages request");
const result = await messagesService.listMessages(query);
res.json(result);
},
);
export const handleGetMessagesByChannel = asyncHandler(
async (req: Request, res: Response) => {
if (!req.params.channelId) {
res.status(400).json({ error: "Missing route parameter: channelId" });
return;
}
const channelId = req.params.channelId as string;
const query = messageQuerySchema.parse(req.query);
logger.debug({ channelId, query }, "Handling get messages by channel");
const result = await messagesService.getMessagesByChannel(channelId, query);
res.json(result);
},
);
export const handleGetMessageById = asyncHandler(
async (req: Request, res: Response) => {
if (!req.params.id) {
res.status(400).json({ error: "Missing route parameter: id" });
return;
}
const id = req.params.id as string;
logger.debug({ id }, "Handling get message by ID");
const result = await messagesService.getMessageById(id);
res.json(result);
},
);
export const handleGetImageMessages = asyncHandler(
async (req: Request, res: Response) => {
const guildId = req.query.guildId as string | undefined;
if (!guildId) {
res.status(400).json({ error: "Missing query parameter: guildId" });
return;
}
const limit = Number(req.query.limit) || 50;
logger.debug({ guildId, limit }, "Handling get image messages");
const result = await messagesService.getImageMessages(guildId, limit);
res.json(result);
},
);
export const handleGetAttachmentsByChannel = asyncHandler(
async (req: Request, res: Response) => {
if (!req.params.channelId) {
res.status(400).json({ error: "Missing route parameter: channelId" });
return;
}
const channelId = req.params.channelId as string;
const query = messageQuerySchema.parse(req.query);
logger.debug({ channelId, query }, "Handling get attachments by channel");
const result = await messagesService.getAttachmentsByChannel(
channelId,
query,
);
res.json(result);
},
);
@@ -77,12 +77,11 @@ export class MessagesRepository {
// Exclude spam threads (NULL-safe: non-thread messages are kept)
if (EXCLUDED_THREAD_IDS.length > 0) {
conditions.push(
or(
isNull(pgMessagesTable.thread_id),
notInArray(pgMessagesTable.thread_id, EXCLUDED_THREAD_IDS),
)!,
const excludeThreads = or(
isNull(pgMessagesTable.thread_id),
notInArray(pgMessagesTable.thread_id, EXCLUDED_THREAD_IDS),
);
if (excludeThreads) conditions.push(excludeThreads);
}
const where = conditions.length > 0 ? and(...conditions) : undefined;
@@ -150,12 +149,11 @@ export class MessagesRepository {
// Exclude spam threads (NULL-safe)
if (EXCLUDED_THREAD_IDS.length > 0) {
conditions.push(
or(
isNull(pgMessagesTable.thread_id),
notInArray(pgMessagesTable.thread_id, EXCLUDED_THREAD_IDS),
)!,
const excludeThreads = or(
isNull(pgMessagesTable.thread_id),
notInArray(pgMessagesTable.thread_id, EXCLUDED_THREAD_IDS),
);
if (excludeThreads) conditions.push(excludeThreads);
}
const rows = await db
@@ -316,12 +314,13 @@ export class MessagesRepository {
like(pgAttachmentsTable.type, "image/%"),
// Exclude spam threads (NULL-safe for non-thread messages)
...(EXCLUDED_THREAD_IDS.length > 0
? [
or(
? (() => {
const excludeThreads = or(
isNull(pgAttachmentsTable.thread_id),
notInArray(pgAttachmentsTable.thread_id, EXCLUDED_THREAD_IDS),
)!,
]
);
return excludeThreads ? [excludeThreads] : [];
})()
: []),
),
)
@@ -1,51 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import {
handleGetAttachmentsByChannel,
handleGetImageMessages,
handleGetMessageById,
handleGetMessagesByChannel,
handleListMessages,
} from "./messages.controller.js";
import { messagesService } from "./messages.service.js";
const logger = createChildLogger("messages.routes");
export function createMessagesRouter(): Router {
const router = express.Router();
// GET /api/messages/images - Get messages with image attachments
// MUST be registered BEFORE /messages/:channelId so "images" is not
// captured as a channelId param.
router.get("/messages/images", handleGetImageMessages);
// GET /api/messages - List messages
router.get("/messages", handleListMessages);
// GET /api/messages/:channelId - Get messages by channel
router.get("/messages/:channelId", handleGetMessagesByChannel);
// GET /api/messages/:channelId/attachments - Get attachments by channel
router.get("/messages/:channelId/attachments", handleGetAttachmentsByChannel);
// GET /api/messages/detail/:id - Get single message by ID
// (uses /detail/ prefix to avoid collision with :channelId route above)
router.get("/messages/detail/:id", handleGetMessageById);
// GET /api/review - Get flagged/warned messages for review
router.get(
"/review",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 20;
const channelId = (req.query.channelId as string) || undefined;
const rows = await messagesService.getReviewMessages(channelId, limit);
logger.debug({ limit, channelId }, "Review query executed");
res.json({ results: rows, limit, cursor: null });
}),
);
return router;
}
@@ -1 +0,0 @@
export { createModerationRouter } from "./moderation.routes.js";
@@ -1,43 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "../../shared/logger/index.js";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { moderationService } from "./moderation.service.js";
const logger = createChildLogger("moderation.routes");
export function createModerationRouter(): Router {
const router = express.Router();
// GET /api/moderation/stats — moderation action summary
router.get(
"/moderation/stats",
asyncHandler(async (_req: Request, res: Response) => {
const stats = await moderationService.getStats();
res.json(stats);
}),
);
// GET /api/moderation/actions — paginated moderation action log
router.get(
"/moderation/actions",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 50;
const status = req.query.status as string | undefined;
const actionType = req.query.actionType as string | undefined;
const cursor = req.query.cursor as string | undefined;
const result = await moderationService.listActions({
limit,
status,
actionType,
cursor: cursor ? Number(cursor) : undefined,
});
logger.debug({ count: result.data.length }, "Moderation actions listed");
res.json(result);
}),
);
return router;
}
@@ -1 +0,0 @@
export { createRecordingsRouter } from "./recordings.routes.js";
@@ -1,41 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { recordingsService } from "./recordings.service.js";
const logger = createChildLogger("recordings.routes");
export function createRecordingsRouter(): Router {
const router = express.Router();
// GET /api/recordings
router.get(
"/recordings",
asyncHandler(async (req: Request, res: Response) => {
const limit = Number(req.query.limit) || 50;
const channelId = req.query.channelId as string | undefined;
const userId = req.query.userId as string | undefined;
const cursor = req.query.cursor as string | undefined;
logger.debug({ limit, channelId, userId, cursor }, "Fetching recordings");
const result = await recordingsService.getRecent(limit, {
channelId,
userId,
cursor,
});
res.json(result);
}),
);
// DELETE /api/recordings/:id
router.delete(
"/recordings/:id",
asyncHandler(async (req: Request, res: Response) => {
const id = req.params.id as string;
await recordingsService.deleteById(id);
res.json({ ok: true });
}),
);
return router;
}
@@ -1 +0,0 @@
export { createUiStateRouter } from "./ui-state.routes.js";
@@ -1,34 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { uiStateService } from "./ui-state.service.js";
const logger = createChildLogger("ui-state.routes");
export function createUiStateRouter(): Router {
const router = express.Router();
// GET /api/ui-state
router.get(
"/ui-state",
asyncHandler(async (_req: Request, res: Response) => {
logger.debug("Fetching UI state");
const state = await uiStateService.getState();
res.json(state);
}),
);
// POST /api/ui-state
router.post(
"/ui-state",
asyncHandler(async (req: Request, res: Response) => {
const updates = req.body as Record<string, unknown>;
logger.debug({ keys: Object.keys(updates) }, "Updating UI state");
const result = await uiStateService.updateState(updates);
res.json(result);
}),
);
return router;
}
@@ -1 +0,0 @@
export { createVoiceRouter } from "./voice.routes.js";
@@ -1,45 +0,0 @@
import type { Request, Response } from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler } from "../../shared/middlewares/index.js";
import { publishCommandNoReply } from "../../shared/redis/index.js";
import type { ConnectVoiceInput, VoiceCommandInput } from "./voice.schema.js";
import {
connectVoice,
disconnectVoice,
getVoiceStatus,
} from "./voice.service.js";
const logger = createChildLogger("voice.controller");
export const handleGetVoiceStatus = asyncHandler(
async (_req: Request, res: Response) => {
const status = await getVoiceStatus();
res.json(status);
},
);
export const handleConnectVoice = asyncHandler(
async (req: Request, res: Response) => {
const { guildId, channelId } = req.body as ConnectVoiceInput;
logger.debug({ guildId, channelId }, "Connecting to voice channel");
const status = await connectVoice(guildId, channelId);
res.json(status);
},
);
export const handleDisconnectVoice = asyncHandler(
async (_req: Request, res: Response) => {
logger.debug("Disconnecting from voice");
const status = await disconnectVoice();
res.json(status);
},
);
export const handleVoiceCommand = asyncHandler(
async (req: Request, res: Response) => {
const { command } = req.body as VoiceCommandInput;
logger.debug({ command }, "Publishing voice command");
await publishCommandNoReply(command);
res.json({ success: true, command });
},
);
@@ -1,80 +0,0 @@
import type { Request, Response, Router } from "express";
import express from "express";
import { createChildLogger } from "@/shared/logger/index";
import { asyncHandler, validateBody } from "../../shared/middlewares/index.js";
import {
handleConnectVoice,
handleDisconnectVoice,
handleGetVoiceStatus,
handleVoiceCommand,
} from "./voice.controller.js";
import { connectVoiceSchema, voiceCommandSchema } from "./voice.schema.js";
import {
getGuilds,
getTextChannels,
getVoiceChannels,
} from "./voice.service.js";
const logger = createChildLogger("voice.routes");
export function createVoiceRouter(): Router {
const router = express.Router();
// ── Guilds ──────────────────────────────────────────────────────────────
// GET /api/guilds
router.get(
"/guilds",
asyncHandler(async (_req: Request, res: Response) => {
logger.debug("Fetching guilds");
const guilds = await getGuilds();
res.json(guilds);
}),
);
// GET /api/guilds/:guildId/channels
router.get(
"/guilds/:guildId/channels",
asyncHandler(async (req: Request, res: Response) => {
const guildId = req.params.guildId as string;
logger.debug({ guildId }, "Fetching text channels");
const channels = await getTextChannels(guildId);
res.json(channels);
}),
);
// GET /api/guilds/:guildId/voice-channels
router.get(
"/guilds/:guildId/voice-channels",
asyncHandler(async (req: Request, res: Response) => {
const guildId = req.params.guildId as string;
logger.debug({ guildId }, "Fetching voice channels");
const channels = await getVoiceChannels(guildId);
res.json(channels);
}),
);
// ── Voice connection ────────────────────────────────────────────────────
// GET /api/voice/status
router.get("/voice/status", handleGetVoiceStatus);
// POST /api/voice/connect
router.post(
"/voice/connect",
validateBody(connectVoiceSchema),
handleConnectVoice,
);
// POST /api/voice/disconnect
router.post("/voice/disconnect", handleDisconnectVoice);
// POST /api/voice/command — send arbitrary voice command (transmit start/stop)
router.post(
"/voice/command",
validateBody(voiceCommandSchema),
handleVoiceCommand,
);
return router;
}
+343
View File
@@ -0,0 +1,343 @@
import { os } from "@orpc/server";
import { z } from "zod";
import { analysisService } from "../modules/analysis/analysis.service";
import { chatRequestSchema } from "../modules/chatbot/chatbot.schema";
import { chatbotService } from "../modules/chatbot/chatbot.service";
// ── Service imports ──────────────────────────────────────────────
import { dashboardService } from "../modules/dashboard/dashboard.service";
import {
mediaLoopSchema,
mediaQueueSchema,
} from "../modules/media/media.schema";
import {
getStatus,
queue,
setLoop,
skip,
stop,
} from "../modules/media/media.service";
import { messageQuerySchema } from "../modules/messages/messages.schema";
import { messagesService } from "../modules/messages/messages.service";
import { moderationService } from "../modules/moderation/moderation.service";
import { recordingsService } from "../modules/recordings/recordings.service";
import { uiStateService } from "../modules/ui-state/ui-state.service";
import {
connectVoice,
disconnectVoice,
getGuilds,
getTextChannels,
getVoiceChannels,
getVoiceStatus,
} from "../modules/voice/voice.service";
import { config } from "../shared/config/index";
import { publishCommandNoReply } from "../shared/redis/index";
// ── Dashboard ────────────────────────────────────────────────────
const dashboardRouter = {
stats: os.handler(() => dashboardService.getStats()),
activity: os
.input(
z.object({ days: z.coerce.number().int().min(1).max(90).default(14) }),
)
.handler(({ input }) => dashboardService.getActivity(input.days)),
users: os
.input(
z.object({
limit: z.coerce.number().int().positive().default(20),
cursor: z.string().optional(),
search: z.string().optional(),
}),
)
.handler(({ input }) =>
dashboardService.listUsers({
limit: input.limit,
cursor: input.cursor,
search: input.search,
}),
),
userDetail: os
.input(z.object({ userId: z.string() }))
.handler(({ input }) => dashboardService.getUserDetail(input.userId)),
channels: os
.input(
z.object({
limit: z.coerce.number().int().positive().default(20),
search: z.string().optional(),
guildId: z.string().optional(),
}),
)
.handler(({ input }) =>
dashboardService.listChannels({
limit: input.limit,
search: input.search,
guildId: input.guildId,
}),
),
channelDetail: os
.input(z.object({ channelId: z.string() }))
.handler(({ input }) => dashboardService.getChannelDetail(input.channelId)),
reactions: os
.input(z.object({ limit: z.coerce.number().int().positive().default(20) }))
.handler(({ input }) => dashboardService.getTopReactions(input.limit)),
reactors: os
.input(z.object({ limit: z.coerce.number().int().positive().default(20) }))
.handler(({ input }) => dashboardService.getTopReactors(input.limit)),
};
// ── Messages ─────────────────────────────────────────────────────
const messagesRouter = {
list: os
.input(messageQuerySchema)
.handler(({ input }) => messagesService.listMessages(input)),
byChannel: os
.input(
z.object({
channelId: z.string(),
query: messageQuerySchema,
}),
)
.handler(({ input }) =>
messagesService.getMessagesByChannel(input.channelId, input.query),
),
detail: os
.input(z.object({ id: z.string() }))
.handler(({ input }) => messagesService.getMessageById(input.id)),
images: os
.input(
z.object({
guildId: z.string(),
limit: z.coerce.number().int().positive().default(50),
}),
)
.handler(({ input }) =>
messagesService.getImageMessages(input.guildId, input.limit),
),
attachmentsByChannel: os
.input(
z.object({
channelId: z.string(),
query: messageQuerySchema,
}),
)
.handler(({ input }) =>
messagesService.getAttachmentsByChannel(input.channelId, input.query),
),
review: os
.input(
z.object({
limit: z.coerce.number().int().positive().default(20),
channelId: z.string().optional(),
}),
)
.handler(async ({ input }) => {
const rows = await messagesService.getReviewMessages(
input.channelId,
input.limit,
);
return { results: rows, limit: input.limit, cursor: null };
}),
};
// ── Moderation ───────────────────────────────────────────────────
const moderationRouter = {
stats: os.handler(() => moderationService.getStats()),
actions: os
.input(
z.object({
limit: z.coerce.number().int().positive().default(50),
status: z.string().optional(),
actionType: z.string().optional(),
cursor: z.coerce.number().int().optional(),
}),
)
.handler(({ input }) =>
moderationService.listActions({
limit: input.limit,
status: input.status,
actionType: input.actionType,
cursor: input.cursor,
}),
),
};
// ── Media ────────────────────────────────────────────────────────
const mediaRouter = {
status: os.handler(() => getStatus()),
queue: os.input(mediaQueueSchema).handler(async ({ input }) => {
await queue(input.source, input.mode);
return getStatus();
}),
skip: os.handler(async () => {
await skip();
return getStatus();
}),
stop: os.handler(async () => {
await stop();
return getStatus();
}),
loop: os.input(mediaLoopSchema).handler(async ({ input }) => {
await setLoop(input.loop);
return getStatus();
}),
};
// ── Voice ─────────────────────────────────────────────────────────
const voiceRouter = {
guilds: os.handler(() => getGuilds()),
textChannels: os
.input(z.object({ guildId: z.string() }))
.handler(({ input }) => getTextChannels(input.guildId)),
voiceChannels: os
.input(z.object({ guildId: z.string() }))
.handler(({ input }) => getVoiceChannels(input.guildId)),
status: os.handler(() => getVoiceStatus()),
connect: os
.input(z.object({ guildId: z.string(), channelId: z.string() }))
.handler(async ({ input }) => {
await connectVoice(input.guildId, input.channelId);
return getVoiceStatus();
}),
disconnect: os.handler(async () => {
await disconnectVoice();
return getVoiceStatus();
}),
command: os
.input(z.object({ command: z.string().min(1) }))
.handler(async ({ input }) => {
await publishCommandNoReply(input.command);
return { success: true, command: input.command };
}),
};
// ── Recordings ───────────────────────────────────────────────────
const recordingsRouter = {
list: os
.input(
z.object({
limit: z.coerce.number().int().positive().default(50),
channelId: z.string().optional(),
userId: z.string().optional(),
cursor: z.string().optional(),
}),
)
.handler(({ input }) =>
recordingsService.getRecent(input.limit, {
channelId: input.channelId,
userId: input.userId,
cursor: input.cursor,
}),
),
delete: os.input(z.object({ id: z.string() })).handler(async ({ input }) => {
await recordingsService.deleteById(input.id);
return { ok: true };
}),
};
// ── Analysis (search) ──────────────────────────────────────────────
const analysisRouter = {
search: os
.input(
z.object({
q: z.string().default(""),
channelId: z.string().optional(),
limit: z.coerce.number().int().positive().default(20),
}),
)
.handler(({ input }) =>
analysisService.search({
q: input.q,
channelId: input.channelId,
limit: input.limit,
}),
),
};
// ── Chatbot ───────────────────────────────────────────────────────
const chatbotRouter = {
chat: os
.input(
chatRequestSchema.extend({
// Per-device actor id; the old REST layer used an X-User-Id header.
// Anonymous sessions use a stable "anonymous" id.
userId: z.string().optional(),
}),
)
.handler(async ({ input }) => {
const userId = input.userId ?? "anonymous";
const response = await chatbotService.processMessage(
input.message,
input.context,
userId,
);
await chatbotService.saveConversation({
userId,
userMessage: input.message,
botResponse: response,
context: input.context,
timestamp: new Date(),
});
return { response, timestamp: new Date().toISOString() };
}),
history: os
.input(
z.object({
limit: z.coerce.number().int().positive().max(100).default(50),
userId: z.string().optional(),
}),
)
.handler(async ({ input }) => {
const userId = input.userId ?? "anonymous";
const history = await chatbotService.getChatHistory(userId, input.limit);
return { history, total: history.length };
}),
clearHistory: os
.input(z.object({ userId: z.string().optional() }))
.handler(async ({ input }) => {
const userId = input.userId ?? "anonymous";
await chatbotService.clearChatHistory(userId);
return { ok: true };
}),
};
// ── Config (public dashboard config snapshot) ──────────────────────
const configRouter = {
get: os.handler(() => ({
monitorGuildId: config.MONITOR_GUILD_ID || null,
webserverPort: config.WEBSERVER_PORT,
nodeEnv: config.NODE_ENV,
backlogSyncHours: config.BACKLOG_SYNC_HOURS,
backlogSyncBatchSize: config.BACKLOG_SYNC_BATCH_SIZE,
retentionMessagesDays: config.RETENTION_MESSAGES_DAYS,
retentionAttachmentsDays: config.RETENTION_ATTACHMENTS_DAYS,
retentionVoiceDays: config.RETENTION_VOICE_DAYS,
autoDeleteFlaggedEnabled: config.AUTO_DELETE_FLAGGED_ENABLED,
aiAnalysisEnabled: config.AI_ANALYSIS_ENABLED,
voiceGuildId: config.VOICE_GUILD_ID || null,
voiceChannelId: config.VOICE_CHANNEL_ID || null,
logLevel: config.LOG_LEVEL,
})),
};
// ── UI State ──────────────────────────────────────────────────────
const uiStateRouter = {
get: os.handler(() => uiStateService.getState()),
update: os
.input(z.record(z.string(), z.unknown()))
.handler(({ input }) => uiStateService.updateState(input)),
};
// ── Root router ───────────────────────────────────────────────────
export const appRouter = {
dashboard: dashboardRouter,
messages: messagesRouter,
moderation: moderationRouter,
media: mediaRouter,
voice: voiceRouter,
recordings: recordingsRouter,
analysis: analysisRouter,
chatbot: chatbotRouter,
config: configRouter,
uiState: uiStateRouter,
};
export type AppRouter = typeof appRouter;
+43
View File
@@ -0,0 +1,43 @@
import type { IncomingMessage, Server } from "node:http";
import type { Duplex } from "node:stream";
import { onError } from "@orpc/server";
import { RPCHandler } from "@orpc/server/ws";
import { WebSocketServer } from "ws";
import { createChildLogger } from "@/shared/logger/index";
import { appRouter } from "./router";
const logger = createChildLogger("orpc.ws");
/**
* Attach the oRPC WebSocket handler to the shared HTTP server, on a path
* SEPARATE from the voice/binary WebSocket (`/ws`). All structured data RPCs
* (dashboard, messages, moderation, media, voice control, recordings,
* analysis, chatbot, config, ui-state) flow over this `/trpc` socket; the
* `/ws` socket is left untouched for Discord PCM audio + gateway events.
*
* We use `noServer` + a manual `upgrade` router (instead of
* `new WebSocketServer({ server, path: "/trpc" })`) because two `ws` servers
* mounted with the `server` option on the SAME http.Server both register
* `upgrade` listeners, and `ws`'s path-guarded listener can reject (400) the
* other server's path. Routing the upgrade ourselves by URL keeps `/trpc`
* and `/ws` fully isolated.
*/
export function createORPCWebSocketServer(server: Server): WebSocketServer {
const handler = new RPCHandler(appRouter, {
interceptors: [
onError((error) => logger.error({ error }, "oRPC WS error")),
],
});
const wss = new WebSocketServer({ noServer: true, perMessageDeflate: false });
server.on("upgrade", (req: IncomingMessage, socket: Duplex, head: Buffer) => {
if (!req.url?.startsWith("/trpc")) return; // let the /ws server handle it
wss.handleUpgrade(req, socket, head, (ws) => {
handler.upgrade(ws, { context: {} });
});
});
logger.info({ path: "/trpc" }, "oRPC WebSocket server attached");
return wss;
}
@@ -62,6 +62,9 @@ export const pgMessagesTable = pgTable(
enum: ["none", "monitor", "warn", "review", "delete", "escalate"],
}),
ai_analyzed_at: pgBigint("ai_analyzed_at", { mode: "number" }),
ai_analysis_duration_ms: pgBigint("ai_analysis_duration_ms", {
mode: "number",
}),
ai_error: pgText("ai_error"),
},
(table) => ({
@@ -80,6 +80,7 @@ export interface MessageRecord {
ai_confidence?: number | null;
ai_recommended_action?: AIRecommendedAction | null;
ai_analyzed_at?: number | null;
ai_analysis_duration_ms?: number | null;
ai_error?: string | null;
}
+1 -1
View File
@@ -108,5 +108,5 @@ export async function retryWithBackoff<T>(
});
}
}
throw lastError!;
throw lastError ?? new Error("Request failed after all retries");
}
@@ -24,6 +24,7 @@ export interface MappedMessage {
ai_confidence: number | null;
ai_recommended_action: string | null;
ai_analyzed_at: number | null;
ai_analysis_duration_ms: number | null;
ai_error: string | null;
is_reply: boolean | null;
is_forward: boolean | null;
@@ -58,6 +59,8 @@ export function mapMessageRow(row: Record<string, unknown>): MappedMessage {
ai_confidence: (row.ai_confidence as number | null) ?? null,
ai_recommended_action: (row.ai_recommended_action as string | null) ?? null,
ai_analyzed_at: (row.ai_analyzed_at as number | null) ?? null,
ai_analysis_duration_ms:
(row.ai_analysis_duration_ms as number | null) ?? null,
ai_error: (row.ai_error as string | null) ?? null,
is_reply: row.is_reply === null ? null : Boolean(row.is_reply),
is_forward: row.is_forward === null ? null : Boolean(row.is_forward),
+14 -2
View File
@@ -1,4 +1,5 @@
import type { Server } from "node:http";
import type { IncomingMessage, Server } from "node:http";
import type { Duplex } from "node:stream";
import { WebSocket, WebSocketServer } from "ws";
import { config } from "../shared/config/index.js";
import { BACKEND_COMMAND, BACKEND_VOICE_TRANSMIT } from "../shared/index.js";
@@ -96,9 +97,20 @@ export function createWebSocketServer(server: Server): WebSocketServer {
const frontendClients = new Set<WebSocket>();
const gatewayClients = new Set<WebSocket>();
const wss = new WebSocketServer({ server, path: "/ws" });
const wss = new WebSocketServer({ noServer: true, perMessageDeflate: true });
_wss = wss;
// Manual upgrade routing: without this, two `ws` servers bound to the same
// http.Server via the `server` option both register `upgrade` listeners and
// the path-guarded one destructively rejects the other's path (400). We own
// the upgrade event and dispatch by URL instead.
server.on("upgrade", (req: IncomingMessage, socket: Duplex, head: Buffer) => {
if (!req.url?.startsWith("/ws")) return;
wss.handleUpgrade(req, socket, head, (ws) => {
wss.emit("connection", ws, req);
});
});
// Map-based dispatcher for JSON WebSocket message types
const jsonHandlers = new Map<string, MessageHandler>();
@@ -0,0 +1,57 @@
import { describe, expect, it } from "vitest";
import { tools } from "../src/modules/chatbot/chatbot.toolDefs.js";
const names = tools.map((t) => t.function.name);
describe("chatbot tool definitions", () => {
it("exposes a stable, non-empty tool set", () => {
expect(tools.length).toBeGreaterThanOrEqual(10);
expect(new Set(names).size).toBe(names.length); // no dup names
});
it("every tool declares a name, description, and object parameters", () => {
for (const t of tools) {
expect(t.type).toBe("function");
expect(typeof t.function.name).toBe("string");
expect(t.function.description.length).toBeGreaterThan(10);
expect(t.function.parameters.type).toBe("object");
}
});
it("required-only tools declare required args", () => {
const byName = new Map(tools.map((t) => [t.function.name, t]));
for (const [name, required] of [
["search_messages", "query"],
["get_user_messages", "userId"],
["get_user_profile", "userId"],
["get_user_reputation", "userId"],
["get_channel_culture", "channelId"],
["get_message_detail", "messageId"],
] as const) {
const tool = byName.get(name);
expect(tool, `missing tool ${name}`).toBeDefined();
expect(tool?.function.parameters.required).toContain(required);
}
});
it("covers the core server-watcher situations", () => {
for (const required of [
"get_server_stats",
"get_top_channels",
"get_recent_activity",
"get_top_flagged",
"search_messages",
"get_user_messages",
"get_user_profile",
"get_user_reputation",
"get_channel_culture",
"get_message_detail",
"get_message_reviews",
"get_voice_recordings",
"get_moderation_timeline",
"get_corrections",
]) {
expect(names, `missing ${required}`).toContain(required);
}
});
});
+121 -150
View File
@@ -1,172 +1,143 @@
# Discord Gateway — Architecture
Pure event-driven microservice (no HTTP server). Captures Discord
messages/voice/attachments/reactions/threads/presence, runs LLM-based AI
moderation, and publishes everything to Redis pub/sub for the backend to
consume. The backend serves the HTTP/WS API to the frontend.
> NOTE: this doc is the source of truth for the module layout. The older
> `MODULE_STRUCTURE.md` was stale (referenced `winston`, `mock-crc.ts`,
> `indonesianTextNormalizer.ts`, and `aiAnalysisWorker.ts`/`llmModerationClient.ts`
> which were renamed/merged). If they disagree, this file wins.
## Top-level layout
```
services/discord-gateway/
├── src/
│ ├── index.ts # Entry point → initializeDiscordGateway()
│ ├── app/
│ │ ├── bootstrap.ts # Discord Gateway initialization (no HTTP server)
│ │ ── shutdown.ts # Graceful shutdown handler
│ │ ├── bootstrap.ts # Wires client, DB, Redis, workers, schedulers
│ │ ── shutdown.ts # Graceful shutdown (SIGINT/SIGTERM + transient errors)
│ │ └── retention.ts # Expired-record cleanup scheduler
│ ├── shared/
│ │ ├── config/
│ │ │ └── config.ts # Environment configuration (Zod validated)
│ │ ├── database/
│ │ │ ── schema.ts # Drizzle ORM schema
│ │ │ ├── drizzle.ts # Database connection
│ │ │ ├── migrate.ts # Migration runner
│ │ │ └── voiceRecordingRepo.ts
│ │ ├── errors/
│ │ │ └── errors.ts # Custom error classes
│ │ ├── logger/
│ │ │ ├── logger.ts # Winston logger wrapper
│ │ └── serialization.ts # Log value serialization
├── utils/
│ └── retry.ts # Retry with backoff utility
── discord/
└── clientOptions.ts # Discord.js client configuration
├── modules/
├── message-capture/ # Modular MVC: Message capture & storage
│ │ ├── messageCapture.ts # Controller: Discord event listeners
│ │ ├── messageStore.ts # Repository: Database operations
│ │ ├── messageMetadata.ts # Service: Message metadata extraction
│ │ ├── types.ts # Domain types
│ │ └── index.ts # Module exports
│ │ ├── ai-moderation/ # Modular MVC: AI analysis & moderation
│ │ │ ├── aiAnalyzer.ts # Controller: Analysis orchestration
│ │ │ ├── llmModerationClient.ts # Service: LLM API client
│ │ │ ├── aiAnalysisWorker.ts # Service: Worker pool management
│ │ │ ├── indonesianTextNormalizer.ts # Service: Text normalization
│ │ │ ├── moderationPrompt.ts # Service: Prompt generation
│ │ │ └── index.ts # Module exports
│ │ ├── voice-recording/ # Modular MVC: Voice recording & streaming
│ │ │ ├── voiceController.ts # Controller: Voice connection management
│ │ │ ├── recorder.ts # Service: Recording orchestration
│ │ │ ├── recorder/
│ │ │ │ ├── audioStream.ts # Service: Audio stream subscription
│ │ │ │ ├── decoder.ts # Service: Opus decoding
│ │ │ │ ├── segment.ts # Service: OGG segment rotation
│ │ │ │ ├── metadata.ts # Service: Segment metadata
│ │ │ │ ├── sessionRecording.ts # Service: Session management
│ │ │ │ └── uploader.ts # Service: Segment upload
│ │ │ └── index.ts # Module exports
│ │ ├── attachment-upload/ # Modular MVC: Attachment handling
│ │ │ ├── attachmentUploader.ts # Service: Upload orchestration
│ │ │ ├── imageResizer.ts # Service: Image resizing
│ │ │ └── index.ts # Module exports
│ │ └── event-broadcaster/ # Event-driven: Redis pub/sub
│ │ ├── eventBroadcaster.ts # Service: Event publishing
│ │ ├── eventTypes.ts # Domain: Event type definitions
│ │ └── index.ts # Module exports
│ ├── mock-crc.ts # CRC polyfill for discord.js
│ └── index.ts # Service entry point
├── package.json # Service dependencies
└── tsconfig.json # TypeScript configuration
│ │ ├── config/ # Zod-validated env (index.ts = schema+loader)
│ │ ├── database/ # Drizzle ORM + pg Pool + migrations
│ │ │ ├── init.ts drizzle.ts pool.ts migrate.ts migrateCli.ts
│ │ │ ── schema/ # messages, cache, voice, analytics, meta
│ │ ├── logger/ # pino wrapper + createChildLogger()
│ │ ├── errors/ # AppError / ConfigError / AudioError ...
│ │ ├── utils/ # retry, pagination
│ │ ├── discord/clientOptions.ts # discord.js-selfbot-v13 client options
│ │ ├── uploader.ts # Shared attachment upload helper
│ │ ├── redis-channels.ts # Redis channel-name constants
│ │ └── moderation-types.ts # Shared AI analysis domain types
└── modules/
├── message-capture/ # Discord event listeners + DB store
├── ai-moderation/ # LLM moderation pipeline (see below)
── voice-recording/ # Voice connect + Opus→OGG recording
└── recorder/ # decoder, segment, session, uploader, oggCrc
├── voice-pcm-ws/ # Real-time PCM → backend WebSocket (bypasses Redis)
├── attachment-upload/ # Download + (sharp) resize + upload
├── event-broadcaster/ # RedisEventPublisher + EventBroadcaster
├── command-handler/ # Redis-subscribed backend→gateway commands
├── reaction-tracking/ thread-tracking/ user-presence/
├── channel-topic/ guild-member-events/
└── gateway-metrics/ # Prometheus /metrics endpoint (port 4016)
```
## Architecture Patterns
## AI moderation pipeline (`ai-moderation/`)
### Modular MVC Structure
Each module follows Controller-Service-Repository pattern:
- **Controller**: Discord event listeners (messageCapture, aiAnalyzer, voiceController)
- **Service**: Business logic (messageStore, llmModerationClient, recorder)
- **Repository**: Data access (messageStore, voiceRecordingRepo)
LLM-only judge — no regex/heuristic classification. One orchestrator call
handles a whole batch (text + media split internally, parallel paths).
### Event-Driven Design
- **Redis Pub/Sub**: All events published to Redis channels
- **Event Channels**:
- `discord:message:created` — New message captured
- `discord:message:updated` — Message edited
- `discord:message:deleted` — Message deleted
- `discord:message:analyzed` — AI analysis complete
- `discord:attachment:created` — Attachment detected
- `discord:attachment:uploaded` — Attachment uploaded to storage
- `discord:voice:started` — Voice recording started
- `discord:voice:stopped` — Voice recording stopped
- `discord:voice:uploaded` — Voice segment uploaded
- `discord:analysis:queue_status` — Analysis queue status update
- `aiAnalyzer.ts` — public API: `queueMessageAnalysis`, `getAnalysisQueueStatus`,
`startPendingAIAnalysisWorker` (recovery worker + cache-prune).
- `batchScheduler.ts` — per-conversation debounce → `processBatch`.
- `batchProcessor.ts` — batch lock/circuit-breaker, fans failed targets to
individual fallback.
- `individualFallbackProcessor.ts` — one-message-at-a-time retry path, own CB.
- `conversationState.ts` / `circuitBreaker.ts` — per-conversation state,
Piscina `workerPool`, `getConversationKey`.
- `ai-analysis-worker.ts` — Piscina entry point (`batch` / `individual` jobs).
Runs `runModerationAnalysis` off the main thread.
- `moderationOrchestrator.ts` — exact-hash cache → batched semantic (Qdrant)
cache → LLM. Text and media paths run in parallel.
- `textBatchProcessor.ts` / `mediaBatchProcessor.ts` — actual LLM calls
(one call per sub-batch, not per message).
- `llmClient.ts` — central OpenAI-compatible chat client (streaming, retries,
thinking-disable injection). `visionAnalyzer.ts` / `mediaAnalysisClient.ts`
share the same router/base URL (different model alias for vision).
- `embeddingClient.ts` + `qdrantClient.ts` — semantic cache (one embed call +
one batched Qdrant search for all uncached targets).
- `textCacheStore.ts` / `channelCultureStore.ts` / `userProfileStore.ts` /
`userReputationStore.ts` — caches & learned per-channel/user state.
### Shared Infrastructure
- **Config**: Zod-validated environment variables
- **Logger**: Winston logger with context support
- **Database**: Drizzle ORM with PostgreSQL
- **Errors**: Custom error classes with codes and status codes
- **Utils**: Retry logic with exponential backoff
### Concurrency model
### No HTTP Server
- Discord Gateway service is **event-driven only**
- No Express, WebSocket, or HTTP routes
- All communication via Redis pub/sub
- Backend service consumes events and serves HTTP API
- Main thread owns the LLM semaphore (`AI_LLM_MAX_CONCURRENT`, default 5) via
`llmClient.withLlmConcurrency`.
- Piscina pool (`PISCINA_MAX_THREADS`, default 4) runs the heavy LLM work off
the event loop; **each worker thread initializes its own pg Pool** (min 0,
grows to `POSTGRES_POOL_MAX`). See "Memory & connections" below.
## Initialization Flow
## Memory & DB connections
1. Load environment config (Zod validation)
2. Initialize database connection
3. Run pending migrations
4. Create Discord client with optimized cache settings
5. Initialize Redis event broadcaster
6. Register Discord event listeners (messageCapture, aiAnalyzer)
7. Login to Discord
8. Listen for graceful shutdown signals (SIGINT, SIGTERM)
`MemoryMax=1G` (raised from 512M — live RSS sits at ~500 MiB, peak 508 MiB,
so 512M left ~2% headroom and risked an OOM-kill restart). Host has 8 GB free.
## Graceful Shutdown
`POSTGRES_POOL_MIN=0` (default). The gateway = main process + up to 4 Piscina
worker threads, each with its own pg Pool. With min:0 the pools stay empty
until a query runs and drop idle clients afterward, instead of holding
`(1 main + 4 workers) × 2 = 10` permanently-open idle connections against
PgBouncer. The pool still grows on demand up to `POSTGRES_POOL_MAX`.
On shutdown signal:
1. Close database connection
2. Disconnect from voice channels
3. Close Redis connection
4. Destroy Discord client
5. Exit process
## Event channels (Redis pub/sub)
## Dependencies
`discord:message:{created,updated,deleted,analyzed}`,
`discord:attachment:{created,uploaded}`,
`discord:voice:{started,stopped,uploaded,active_user,pcm,analyzed}`,
`discord:analysis:queue_status`,
`discord:reaction:{added,removed}`,
`discord:thread:{created,deleted,updated}`,
`discord:channel_topic:updated`,
`discord:presence:updated`,
`discord:guild_member:{added,removed}`.
See `src/shared/redis-channels.ts` for the canonical names.
**Core Discord**:
- discord.js-selfbot-v13
- @discordjs/voice
- @discordjs/opus
## Initialization flow
**Audio Processing**:
- prism-media (Opus encoding/decoding)
- opusscript (Opus fallback)
- sharp (Image resizing)
1. Validate env (Zod). Refuse to start if `AI_ANALYSIS_ENABLED` but no key.
2. `AUTO_MIGRATE_ON_STARTUP` → run pending Drizzle migrations.
3. `initializeDatabase()` (pg Pool, min 0).
4. Create discord.js-selfbot-v13 client; register listeners on `ready`.
5. Start `gmw-discord-gateway` metrics server (port `METRICS_PORT`, default 4016).
6. `client.login(token)`.
**Data & Config**:
- drizzle-orm (ORM)
- pg (PostgreSQL driver)
- zod (Config validation)
- ioredis (Redis client)
## Graceful shutdown
**Logging & Utilities**:
- winston (Structured logging)
- p-retry (Retry logic)
- p-limit (Concurrency limiting)
- piscina (Worker pool)
`SIGINT`/`SIGTERM` (and uncaught transient stream errors: EPIPE / ECONNRESET /
ERR_STREAM_DESTROYED / ERR_STREAM_WRITE_AFTER_END are treated as non-fatal):
stop metrics → stop muxer → disconnect voice → close PCM WS → close Redis →
close command handler → close DB → destroy client → exit.
## Event Flow Example
## Observability
### Message Capture Flow
1. Discord emits `messageCreate` event
2. `messageCapture.ts` listener receives event
3. Extract metadata (user, channel, content, timestamp)
4. `messageStore.ts` inserts into database
5. `eventBroadcaster.messageCreated()` publishes to Redis
6. Backend service subscribes to `discord:message:created` channel
7. Backend processes and stores in its own database
Prometheus scrapes `127.0.0.1:4016/metrics` (`bete_*` prefix). Collectors run
per-scrape and expose: process memory/uptime, and (when AI analysis is on) live
pipeline gauges — `ai_analysis_queued_conversations`,
`ai_analysis_active_batch_requests`, `ai_analysis_active_individual_requests`,
`ai_analysis_individual_in_flight`, `ai_analysis_individual_circuit_breaker_active`,
`ai_analysis_worker_threads`, `ai_analysis_worker_threads_active`.
### Voice Recording Flow
1. `voiceController.connect()` joins voice channel
2. `recorder.ts` subscribes to user audio streams
3. For each speaking user:
- Create audio stream subscription
- Decode Opus packets to PCM
- Rotate OGG segments (5s default)
- Collect user metadata
4. On silence (3s):
- Finalize segment
- Create metadata JSON
- Upload segment to storage
- Publish `discord:voice:uploaded` event
5. Backend service receives event and indexes recording
## Key invariants (do not break)
## No Breaking Changes
- Original `src/` remains untouched for now
- Discord Gateway is a **new service** in `services/discord-gateway/`
- Can run alongside existing monolith during transition
- Backend service will consume Redis events
- Frontend continues to use Backend HTTP API
- **LLM is the only judge.** Failed LLM → `status:"error"` + recovery retry.
Never reintroduce regex/heuristic content classification.
- **Discord tokens are sanitized** (`discordTokens.ts`: `<:emoji:id>`
`[emoji:name]`, `<@id>` `@user`, etc.) before content reaches the LLM, so
numeric snowflake IDs never trigger false positives.
- **Semantic cache is batched** (one embed call + one Qdrant batch search),
not N sequential round-trips. `ensureQdrantCollection` is memoized.
- **Streaming is mandatory** against the 9router base URL (non-stream waits for
the full body and times out). `llmClient` aggregates SSE chunks.
+60 -388
View File
@@ -1,408 +1,80 @@
# Discord Gateway Service - Module Structure
# Discord Gateway Service Module Structure
## Complete Directory Tree
> Kept as a compact module map. For the authoritative layout, design
> decisions, and invariants, see `ARCHITECTURE.md`. This file was rewritten
> on 2026-08-16 to fix stale references (`winston` → pino,
> `mock-crc.ts`/`indonesianTextNormalizer.ts` removed,
> `aiAnalysisWorker.ts``ai-analysis-worker.ts`,
> `llmModerationClient.ts``llmClient.ts`).
## Top-level
```
services/discord-gateway/
├── src/
│ ├── app/
│ ├── bootstrap.ts
└── Initializes Discord client, database, Redis broadcaster
│ │ Registers event listeners, handles graceful shutdown
── shutdown.ts
── Graceful shutdown handler for SIGINT/SIGTERM/exceptions
├── shared/
├── config/
│ └── config.ts
│ │ └── Zod-validated environment configuration
│ │ - Discord token, database URL, Redis URL
│ - AI LLM settings, recording parameters
│ - Attachment upload settings, retention policies
│ │ │
│ │ ├── database/
│ │ │ ├── schema.ts
│ │ │ │ └── Drizzle ORM schema definitions
│ │ │ ├── drizzle.ts
│ │ │ │ └── PostgreSQL connection and initialization
│ │ │ ├── migrate.ts
│ │ │ │ └── Database migration runner
│ │ │ ├── migrateCli.ts
│ │ │ │ └── CLI for programmatic migrations
│ │ │ ├── voiceRecordingRepo.ts
│ │ │ │ └── Voice recording repository
│ │ │ └── migrations/
│ │ │ └── Database migration files
│ │ │
│ │ ├── errors/
│ │ │ └── errors.ts
│ │ │ └── Custom error classes
│ │ │ - AppError (base)
│ │ │ - ConfigError
│ │ │ - AudioError
│ │ │ - VoiceConnectionError
│ │ │ - ValidationError
│ │ │
│ │ ├── logger/
│ │ │ ├── logger.ts
│ │ │ │ └── Winston logger wrapper with context support
│ │ │ └── serialization.ts
│ │ │ └── Log value serialization utilities
│ │ │
│ │ ├── utils/
│ │ │ └── retry.ts
│ │ │ └── Retry with exponential backoff utility
│ │ │
│ │ └── discord/
│ │ └── clientOptions.ts
│ │ └── Discord.js client configuration
│ │
│ ├── modules/
│ │ │
│ │ ├── message-capture/
│ │ │ ├── messageCapture.ts
│ │ │ │ └── CONTROLLER: Discord event listeners
│ │ │ │ - messageCreate, messageUpdate, messageDelete
│ │ │ │ - Validates capture target, publishes events
│ │ │ │
│ │ │ ├── messageStore.ts
│ │ │ │ └── REPOSITORY: Database CRUD operations
│ │ │ │ - upsertMessageForCapture
│ │ │ │ - updateMessageAsEdited
│ │ │ │ - updateMessageAsDeleted
│ │ │ │ - insertAttachment
│ │ │ │ - getMessageById
│ │ │ │
│ │ │ ├── messageMetadata.ts
│ │ │ │ └── SERVICE: Message metadata extraction
│ │ │ │ - getMessageMetadata
│ │ │ │ - getMessageLocation
│ │ │ │ - getDisplayContent
│ │ │ │
│ │ │ ├── types.ts
│ │ │ │ └── Domain types
│ │ │ │ - MessageRecord
│ │ │ │ - AttachmentRecord
│ │ │ │ - VoiceSegmentRecord
│ │ │ │ - AIStatus, AISeverity, AIRecommendedAction
│ │ │ │
│ │ │ └── index.ts
│ │ └── Module exports
│ │
│ │ ├── ai-moderation/
│ │ │ ├── aiAnalyzer.ts
│ │ │ │ └── CONTROLLER: Analysis orchestration
│ │ │ │ - startPendingAIAnalysisWorker
│ │ │ │ - queueMessageAnalysis
│ │ │ │ - Manages analysis queue and worker pool
│ │ │ │
│ │ │ ├── llmModerationClient.ts
│ │ │ │ └── SERVICE: LLM API integration
│ │ │ │ - Calls LLM for text/image moderation
│ │ │ │ - Parses responses, handles errors
│ │ │ │ - Retry logic with backoff
│ │ │ │
│ │ │ ├── aiAnalysisWorker.ts
│ │ │ │ └── SERVICE: Worker pool management
│ │ │ │ - Piscina worker pool for parallel analysis
│ │ │ │ - Conversation context batching
│ │ │ │
│ │ │ ├── indonesianTextNormalizer.ts
│ │ │ │ └── SERVICE: Text preprocessing
│ │ │ │ - Normalize Indonesian text
│ │ │ │ - Handle diacritics, abbreviations
│ │ │ │
│ │ │ ├── moderationPrompt.ts
│ │ │ │ └── SERVICE: Prompt generation
│ │ │ │ - Generate LLM prompts for moderation
│ │ │ │ - Include context and policy
│ │ │ │
│ │ │ └── index.ts
│ │ └── Module exports
│ │
│ │ ├── voice-recording/
│ │ │ ├── voiceController.ts
│ │ │ │ └── CONTROLLER: Voice connection management
│ │ │ │ - connect(guildId, channelId)
│ │ │ │ - disconnect()
│ │ │ │ - listGuilds(), listVoiceChannels()
│ │ │ │ - getStatus()
│ │ │ │
│ │ │ ├── recorder.ts
│ │ │ │ └── SERVICE: Recording orchestration
│ │ │ │ - startRecording(client, channel)
│ │ │ │ - stopRecording(guildId)
│ │ │ │ - Manages active recording sessions
│ │ │ │
│ │ │ ├── recorder/
│ │ │ │ ├── audioStream.ts
│ │ │ │ │ └── SERVICE: Audio stream subscription
│ │ │ │ │ - subscribeToAudioStream
│ │ │ │ │ - Opus packet handling
│ │ │ │ │
│ │ │ │ ├── decoder.ts
│ │ │ │ │ └── SERVICE: Opus decoding
│ │ │ │ │ - OpusDecoder class
│ │ │ │ │ - Decode Opus to PCM
│ │ │ │ │ - Rotation and cooldown logic
│ │ │ │ │
│ │ │ │ ├── segment.ts
│ │ │ │ │ └── SERVICE: OGG segment rotation
│ │ │ │ │ - SegmentManager class
│ │ │ │ │ - Rotate segments (5s default)
│ │ │ │ │ - Write OGG files
│ │ │ │ │
│ │ │ │ ├── metadata.ts
│ │ │ │ │ └── SERVICE: Segment metadata
│ │ │ │ │ - collectUserMetadata
│ │ │ │ │ - createSegmentMetadata
│ │ │ │ │ - User info, roles, timestamps
│ │ │ │ │
│ │ │ │ ├── sessionRecording.ts
│ │ │ │ │ └── SERVICE: Session management
│ │ │ │ │ - createRecordingSession
│ │ │ │ │ - finalizeRecordingSession
│ │ │ │ │ - Track active sessions
│ │ │ │ │
│ │ │ │ └── uploader.ts
│ │ │ │ └── SERVICE: Segment upload
│ │ │ │ - uploadRecordingSegment
│ │ │ │ - Upload to external storage
│ │ │ │ - Retry logic
│ │ │ │
│ │ │ └── index.ts
│ │ └── Module exports
│ │
│ │ ├── attachment-upload/
│ │ │ ├── attachmentUploader.ts
│ │ │ │ └── SERVICE: Upload orchestration
│ │ │ │ - processAttachmentUpload
│ │ │ │ - Download from Discord
│ │ │ │ - Upload to external storage
│ │ │ │ - Retry with backoff
│ │ │ │
│ │ │ ├── imageResizer.ts
│ │ │ │ └── SERVICE: Image processing
│ │ │ │ - resizeImage
│ │ │ │ - Resize to max dimension
│ │ │ │ - Preserve aspect ratio
│ │ │ │
│ │ │ └── index.ts
│ │ └── Module exports
│ │
│ │ └── event-broadcaster/
│ │ ├── eventBroadcaster.ts
│ │ │ └── SERVICE: Redis pub/sub publisher
│ │ │ - EventBroadcaster class
│ │ │ - RedisEventPublisher class
│ │ │ - Publish to Redis channels
│ │ │ - Methods:
│ │ │ - messageCreated()
│ │ │ - messageUpdated()
│ │ │ - messageDeleted()
│ │ │ - messageAnalyzed()
│ │ │ - attachmentCreated()
│ │ │ - attachmentUploaded()
│ │ │ - voiceRecordingStarted()
│ │ │ - voiceRecordingStopped()
│ │ │ - voiceRecordingUploaded()
│ │ │ - analysisQueueStatus()
│ │ │
│ │ ├── eventTypes.ts
│ │ │ └── Domain types
│ │ │ - DiscordGatewayEvent interface
│ │ │ - EventChannels constants
│ │ │ - Event channel names
│ │ │
│ │ └── index.ts
│ └── Module exports
│ ├── mock-crc.ts
│ │ └── CRC polyfill for discord.js compatibility
│ │
│ └── index.ts
│ └── Service entry point
│ - Initialize Discord Gateway
│ - Handle startup errors
├── ARCHITECTURE.md
│ └── Detailed architecture documentation
├── README.md
│ └── Complete service documentation
├── MODULE_STRUCTURE.md
│ └── This file - module structure reference
└── package.json
└── Service dependencies and scripts
│ ├── index.ts # Entry point
├── app/ # bootstrap, shutdown, retention
├── shared/ # config, database, logger, errors, utils, discord, uploader
└── modules/
── message-capture/ # Discord listeners + DB store + metadata
── ai-moderation/ # LLM moderation pipeline (largest module)
├── voice-recording/ # Voice connect + Opus→OGG recording (+ recorder/)
├── voice-pcm-ws/ # Real-time PCM → backend WebSocket
├── attachment-upload/ # Download + sharp resize + upload
├── event-broadcaster/ # RedisEventPublisher + EventBroadcaster
├── command-handler/ # Backend→gateway Redis commands
├── reaction-tracking/ thread-tracking/ user-presence/
├── channel-topic/ guild-member-events/
└── gateway-metrics/ # Prometheus /metrics (port 4016)
├── tests/ # Vitest suites (129 tests)
├── drizzle/ # Drizzle migration SQL + journal
├── ARCHITECTURE.md README.md package.json tsconfig.json vitest.config.ts
```
## Module Responsibilities
## Module responsibilities (summary)
### message-capture
**Purpose**: Capture Discord messages (create, update, delete)
**Pattern**: Controller-Service-Repository
- **Controller** (messageCapture.ts): Listens to Discord events
- **Service** (messageMetadata.ts): Extracts metadata
- **Repository** (messageStore.ts): Database operations
- **Events Published**:
- `discord:message:created`
- `discord:message:updated`
- `discord:message:deleted`
Captures `messageCreate`/`messageUpdate`/`messageDelete`, extracts metadata,
stores to Postgres, publishes to Redis. ControllerServiceRepository split:
`messageCapture.ts` (listener) → `messageStore.ts` (DB) + `messageMetadata.ts`
(service).
### ai-moderation
**Purpose**: Analyze messages with LLM for moderation
**Pattern**: Controller-Service-Service-Service
- **Controller** (aiAnalyzer.ts): Orchestrates analysis workflow
- **Service** (llmModerationClient.ts): LLM API integration
- **Service** (aiAnalysisWorker.ts): Worker pool management
- **Service** (indonesianTextNormalizer.ts): Text preprocessing
- **Service** (moderationPrompt.ts): Prompt generation
- **Events Published**:
- `discord:message:analyzed`
- `discord:analysis:queue_status`
LLM-only moderation. Entry: `aiAnalyzer.ts` (`queueMessageAnalysis`,
`startPendingAIAnalysisWorker`, `getAnalysisQueueStatus`). Scheduling:
`batchScheduler.ts``batchProcessor.ts` (batch lock + circuit breaker) →
`individualFallbackProcessor.ts` (per-message retry). Heavy work runs in the
Piscina pool via `ai-analysis-worker.ts` (jobs `batch` / `individual`).
Orchestration/caching: `moderationOrchestrator.ts` (exact hash → batched
semantic Qdrant → LLM), `textBatchProcessor.ts` / `mediaBatchProcessor.ts`
(one LLM call per sub-batch), `llmClient.ts` (central streaming client),
`embeddingClient.ts` + `qdrantClient.ts` (semantic cache), plus
`channelCultureStore.ts` / `userProfileStore.ts` / `userReputationStore.ts`.
### voice-recording
**Purpose**: Record voice channel audio
**Pattern**: Controller-Service-SubServices
- **Controller** (voiceController.ts): Voice connection management
- **Service** (recorder.ts): Recording orchestration
- **Sub-services** (recorder/*): Audio processing pipeline
- audioStream.ts: Opus packet subscription
- decoder.ts: Opus to PCM decoding
- segment.ts: OGG file rotation
- metadata.ts: User metadata collection
- sessionRecording.ts: Session lifecycle
- uploader.ts: Segment upload
- **Events Published**:
- `discord:voice:started`
- `discord:voice:stopped`
- `discord:voice:uploaded`
`voiceController.ts` (connect/disconnect/list) + `recorder.ts` (orchestration)
+ `recorder/` (decoder, segment, session, uploader, oggCrc). Publishes
`discord:voice:*` events. Real-time audio also streamed via `voice-pcm-ws`.
### attachment-upload
**Purpose**: Upload message attachments to external storage
**Pattern**: Service-Service
- **Service** (attachmentUploader.ts): Upload orchestration
- **Service** (imageResizer.ts): Image processing
- **Events Published**:
- `discord:attachment:created`
- `discord:attachment:uploaded`
`attachmentUploader.ts` (download → upload to storage) + `imageResizer.ts`
(sharp resize). Emits `discord:attachment:*`.
### event-broadcaster
**Purpose**: Publish events to Redis pub/sub
**Pattern**: Service-Domain
- **Service** (eventBroadcaster.ts): Redis publisher
- **Domain** (eventTypes.ts): Event type definitions
- **Channels**:
- discord:message:* (message events)
- discord:attachment:* (attachment events)
- discord:voice:* (voice events)
- discord:analysis:* (analysis events)
`RedisEventPublisher` (ioredis publish) + `EventBroadcaster` (typed methods).
Channel names in `src/shared/redis-channels.ts`.
## Shared Infrastructure
### gateway-metrics
`metrics.ts` Prometheus HTTP server on `METRICS_PORT` (4016). Collectors run
per scrape; live pipeline gauges registered in `bootstrap.ts`.
### config
- Zod-validated environment variables
- Type-safe configuration access
- Sensible defaults
## Shared infrastructure
- **config** — Zod schema in `shared/config/index.ts` (single source of truth).
- **database** — Drizzle ORM over `pg`; pool `min:0` (`shared/config`).
- **logger**`pino` wrapper, `createChildLogger()` for context loggers.
- **errors**`AppError` hierarchy (`ConfigError`, `AudioError`, …).
### database
- Drizzle ORM schema
- PostgreSQL connection
- Migration management
- Voice recording repository
### logger
- Winston logger wrapper
- Context-aware logging
- Log serialization utilities
### errors
- Custom error classes
- Error codes and HTTP status codes
- Proper error hierarchy
### utils
- Retry with exponential backoff
- Configurable retry parameters
### discord
- Discord.js client configuration
- Cache optimization
- Partial handling
## Event Flow
```
Discord Events
message-capture (Controller)
messageStore (Repository) → PostgreSQL
eventBroadcaster (Service)
Redis Pub/Sub
Backend Service (Subscriber)
HTTP API / WebSocket
Frontend Application
```
## No HTTP Server
- ✅ No Express
- ✅ No WebSocket server
- ✅ No HTTP routes
- ✅ No middleware
- ✅ Pure event-driven service
## Graceful Shutdown
1. Close PostgreSQL connection
2. Disconnect from voice channels
3. Close Redis connection
4. Destroy Discord client
5. Exit process
## Dependencies
**Discord**:
- discord.js-selfbot-v13
- @discordjs/voice
- @discordjs/opus
**Audio**:
- prism-media
- opusscript
- sharp
**Data**:
- drizzle-orm
- pg
- zod
- ioredis
**Logging**:
- winston
- p-retry
- p-limit
- piscina
## Summary
The Discord Gateway service is a **pure event-driven microservice** that:
- Captures Discord messages, voice, and attachments
- Performs AI moderation analysis
- Publishes events to Redis pub/sub
- Has no HTTP server or WebSocket
- Follows Modular MVC pattern
- Maintains clean module boundaries
- Provides type-safe configuration
- Includes structured logging
- Handles graceful shutdown
The service is designed to run alongside the Backend service, which consumes Redis events and serves the HTTP API to the Frontend.
## Notes
- No HTTP server (other than the metrics endpoint). Pure event-driven.
- `MODULE_STRUCTURE.md` is intentionally a sketch; `ARCHITECTURE.md` is the
detailed reference. When they diverge, `ARCHITECTURE.md` wins.
@@ -0,0 +1,9 @@
CREATE TABLE IF NOT EXISTS "term_glossary_cache" (
"term" text PRIMARY KEY NOT NULL,
"definition" text NOT NULL,
"source_url" text DEFAULT '' NOT NULL,
"resolved_at" bigint NOT NULL,
"hit_count" integer DEFAULT 0 NOT NULL
);
--> statement-breakpoint
CREATE INDEX IF NOT EXISTS "idx_term_glossary_cache_resolved_at" ON "term_glossary_cache" USING btree ("resolved_at");
@@ -99,6 +99,13 @@
"when": 1785551832190,
"tag": "0013_rename_mascot_chat_to_chatbot",
"breakpoints": true
},
{
"idx": 14,
"version": "7",
"when": 1785621600000,
"tag": "0014_add_term_glossary_cache",
"breakpoints": true
}
]
}
@@ -1,3 +0,0 @@
node_modules/
build/
package-lock.json
@@ -1,526 +0,0 @@
// libdatachannel-min — minimal N-API binding to libdatachannel.
// Exposes ONLY what GMW GoLive needs:
// PeerConnection (offer/answer, ICE, SDP), DataChannel (signaling),
// Track send (added in media phase).
// Built against libdatachannel 0.24.0 (built from source in /tmp/ldc-build).
#include <napi.h>
#include <rtc/rtc.hpp>
#include <functional>
#include <memory>
#include <string>
#include <variant>
using namespace Napi;
namespace {
std::string stateToString(rtc::PeerConnection::State s) {
switch (s) {
case rtc::PeerConnection::State::New: return "new";
case rtc::PeerConnection::State::Connecting: return "connecting";
case rtc::PeerConnection::State::Connected: return "connected";
case rtc::PeerConnection::State::Disconnected: return "disconnected";
case rtc::PeerConnection::State::Failed: return "failed";
case rtc::PeerConnection::State::Closed: return "closed";
default: return "unknown";
}
}
std::string binaryToString(const rtc::binary& data) {
// rtc::binary is std::vector<std::byte> in libdatachannel >= 0.21
std::string msg(data.size(), '\0');
for (size_t i = 0; i < data.size(); i++) {
msg[i] = static_cast<char>(data[i]);
}
return msg;
}
// Holds a Napi::Promise::Deferred so it can be moved into TSFN lambdas
// without invalid copies (node-addon-api 8.x Deferred is not movable).
struct DeferredHolder {
Promise::Deferred deferred;
explicit DeferredHolder(Promise::Deferred d) : deferred(d) {}
};
class DataChannelWrap : public Napi::ObjectWrap<DataChannelWrap> {
public:
static Function Init(Napi::Env env) {
Function func = DefineClass(env, "DataChannel", {
InstanceMethod("send", &DataChannelWrap::Send),
InstanceMethod("isOpen", &DataChannelWrap::IsOpen),
InstanceMethod("close", &DataChannelWrap::Close),
InstanceMethod("onMessage", &DataChannelWrap::OnMessage),
InstanceMethod("onOpen", &DataChannelWrap::OnOpen),
});
dcConstructor = Napi::Persistent(func);
return func;
}
// Create a JS wrapper (calls the JS constructor, returns instance).
static Object NewInstance(Napi::Env env) {
return dcConstructor.New({});
}
DataChannelWrap(const Napi::CallbackInfo& info)
: Napi::ObjectWrap<DataChannelWrap>(info) {}
void Init(std::shared_ptr<rtc::DataChannel> dc) {
dc_ = dc;
dc_->onMessage([this](rtc::message_variant data) {
std::string msg;
if (std::holds_alternative<rtc::binary>(data)) {
msg = binaryToString(std::get<rtc::binary>(data));
} else {
msg = std::get<std::string>(data);
}
if (msgCb_) {
msgCb_->BlockingCall([msg](Napi::Env env, Function cb) {
cb.Call({String::New(env, msg)});
});
}
});
dc_->onOpen([this]() {
if (openCb_) {
openCb_->BlockingCall([](Napi::Env env, Function cb) {
cb.Call({});
});
}
});
}
private:
static FunctionReference dcConstructor;
std::shared_ptr<rtc::DataChannel> dc_;
std::shared_ptr<ThreadSafeFunction> msgCb_;
std::shared_ptr<ThreadSafeFunction> openCb_;
void Send(const Napi::CallbackInfo& info) {
std::string msg = info[0].As<String>().Utf8Value();
if (dc_) dc_->send(msg);
}
Napi::Value IsOpen(const Napi::CallbackInfo& info) {
bool open = dc_ && dc_->isOpen();
return Boolean::New(info.Env(), open);
}
void Close(const Napi::CallbackInfo& info) {
if (dc_) dc_->close();
}
void OnMessage(const Napi::CallbackInfo& info) {
Function cb = info[0].As<Function>();
msgCb_ = std::make_shared<ThreadSafeFunction>(
ThreadSafeFunction::New(info.Env(), cb, "dc-message", 0, 1));
}
void OnOpen(const Napi::CallbackInfo& info) {
Function cb = info[0].As<Function>();
openCb_ = std::make_shared<ThreadSafeFunction>(
ThreadSafeFunction::New(info.Env(), cb, "dc-open", 0, 1));
}
};
class TrackWrap : public Napi::ObjectWrap<TrackWrap> {
public:
static Function Init(Napi::Env env) {
Function func = DefineClass(env, "Track", {
InstanceMethod("send", &TrackWrap::Send),
InstanceMethod("isOpen", &TrackWrap::IsOpen),
InstanceMethod("close", &TrackWrap::Close),
InstanceMethod("setPacketizer", &TrackWrap::SetPacketizer),
InstanceMethod("sendFrame", &TrackWrap::SendFrame),
InstanceMethod("addTimestamp", &TrackWrap::AddTimestamp),
});
trackConstructor = Napi::Persistent(func);
return func;
}
static Object NewInstance(Napi::Env env) {
return trackConstructor.New({});
}
TrackWrap(const Napi::CallbackInfo& info)
: Napi::ObjectWrap<TrackWrap>(info) {}
void Init(std::shared_ptr<rtc::Track> track, Napi::Env env) {
track_ = track;
(void)env;
}
private:
static FunctionReference trackConstructor;
std::shared_ptr<rtc::Track> track_;
std::shared_ptr<rtc::RtpPacketizationConfig> rtpConfig_;
void Send(const Napi::CallbackInfo& info) {
Buffer<uint8_t> buf = info[0].As<Buffer<uint8_t>>();
if (!track_) return;
rtc::binary data(buf.Length());
for (size_t i = 0; i < buf.Length(); i++) data[i] = (std::byte)buf[i];
try {
track_->send(data);
} catch (const std::exception& e) {
fprintf(stderr, "[binding] track.send THREW: %s\n", e.what());
}
}
// setPacketizer(kind, ssrc, payloadType, clockRate, playoutDelayId,
// playoutDelayMin, playoutDelayMax)
// kind: "audio" | "h264" | "h265" | "av1"
// Builds the media-handler chain (packetizer → RTCP SR → NACK → pacing for
// video) exactly like @dank074's WebRtcWrapper does via node-datachannel.
void SetPacketizer(const Napi::CallbackInfo& info) {
Napi::Env env = info.Env();
if (!track_) throw Error::New(env, "track closed");
std::string kind = info[0].As<String>().Utf8Value();
uint32_t ssrc = info[1].As<Number>().Uint32Value();
uint8_t pt = (uint8_t)info[2].As<Number>().Uint32Value();
uint32_t clockRate = info[3].As<Number>().Uint32Value();
uint8_t playoutDelayId = (uint8_t)info[4].As<Number>().Uint32Value();
uint16_t playoutDelayMin = (uint16_t)info[5].As<Number>().Uint32Value();
uint16_t playoutDelayMax = (uint16_t)info[6].As<Number>().Uint32Value();
try {
auto cfg = std::make_shared<rtc::RtpPacketizationConfig>(
ssrc, "", pt, clockRate);
cfg->playoutDelayId = playoutDelayId;
cfg->playoutDelayMin = playoutDelayMin;
cfg->playoutDelayMax = playoutDelayMax;
std::shared_ptr<rtc::MediaHandler> handler;
if (kind == "audio") {
handler = std::make_shared<rtc::OpusRtpPacketizer>(cfg);
} else if (kind == "h264") {
handler = std::make_shared<rtc::H264RtpPacketizer>(
rtc::NalUnit::Separator::StartSequence, cfg);
} else if (kind == "h265") {
handler = std::make_shared<rtc::H265RtpPacketizer>(
rtc::NalUnit::Separator::StartSequence, cfg);
} else if (kind == "av1") {
handler = std::make_shared<rtc::AV1RtpPacketizer>(
rtc::AV1RtpPacketizer::Packetization::Obu, cfg);
} else {
throw std::runtime_error("unknown packetizer kind: " + kind);
}
handler->addToChain(std::make_shared<rtc::RtcpSrReporter>(cfg));
handler->addToChain(std::make_shared<rtc::RtcpNackResponder>());
if (kind != "audio") {
handler->addToChain(std::make_shared<rtc::PacingHandler>(
25.0 * 1000 * 1000, std::chrono::milliseconds(1)));
}
track_->setMediaHandler(handler);
rtpConfig_ = cfg;
} catch (const std::exception& e) {
fprintf(stderr, "[binding] setPacketizer THREW: %s\n", e.what());
throw Error::New(env, e.what());
}
}
// sendFrame(buffer) — sends an ENCODED frame (AnnexB H264 / raw opus /
// OBU AV1). The media-handler chain packetizes it into RTP.
void SendFrame(const Napi::CallbackInfo& info) {
Buffer<uint8_t> buf = info[0].As<Buffer<uint8_t>>();
if (!track_) return;
rtc::binary data(buf.Length());
for (size_t i = 0; i < buf.Length(); i++) data[i] = (std::byte)buf[i];
try {
track_->send(data);
} catch (const std::exception& e) {
fprintf(stderr, "[binding] track.sendFrame THREW: %s\n", e.what());
}
}
// addTimestamp(delta) — advances the packetizer RTP timestamp by delta
// (clock-rate units). Called by JS after each frame, matching the
// node-datachannel contract (WebRtcWrapper does the same increment).
void AddTimestamp(const Napi::CallbackInfo& info) {
uint32_t delta = info[0].As<Number>().Uint32Value();
if (rtpConfig_) rtpConfig_->timestamp += delta;
}
Napi::Value IsOpen(const Napi::CallbackInfo& info) {
bool open = track_ && track_->isOpen();
return Boolean::New(info.Env(), open);
}
void Close(const Napi::CallbackInfo& info) {
if (track_) track_->close();
}
void OnStateChange(const Napi::CallbackInfo& info) {
// libdatachannel Track has no state-change callback; kept for API parity.
(void)info;
}
};
class PeerConnectionWrap : public Napi::ObjectWrap<PeerConnectionWrap> {
public:
static Function Init(Napi::Env env) {
Function func = DefineClass(env, "PeerConnection", {
InstanceMethod("state", &PeerConnectionWrap::State),
InstanceMethod("createOffer", &PeerConnectionWrap::CreateOffer),
InstanceMethod("createAnswer", &PeerConnectionWrap::CreateAnswer),
InstanceMethod("setRemoteDescription",
&PeerConnectionWrap::SetRemoteDescription),
InstanceMethod("close", &PeerConnectionWrap::Close),
InstanceMethod("onStateChange", &PeerConnectionWrap::OnStateChange),
InstanceMethod("createDataChannel", &PeerConnectionWrap::CreateDataChannel),
InstanceMethod("onDataChannel", &PeerConnectionWrap::OnDataChannel),
InstanceMethod("addTrack", &PeerConnectionWrap::AddTrack),
});
return func;
}
PeerConnectionWrap(const Napi::CallbackInfo& info)
: Napi::ObjectWrap<PeerConnectionWrap>(info) {
Napi::Env env = info.Env();
if (!info[0].IsObject()) {
throw TypeError::New(env, "config object required");
}
Object config = info[0].As<Object>();
rtc::Configuration rtcConfig;
if (config.Has("iceServers")) {
Array servers = config.Get("iceServers").As<Array>();
for (uint32_t i = 0; i < servers.Length(); i++) {
std::string url = servers.Get(i).As<String>().Utf8Value();
rtcConfig.iceServers.emplace_back(url);
}
}
pc_ = std::make_shared<rtc::PeerConnection>(rtcConfig);
// IMPORTANT: register description/gathering callbacks HERE (constructor),
// BEFORE any createDataChannel call. libdatachannel only fires
// onLocalDescription for negotiations that start AFTER the callback is
// registered — if createDataChannel runs first, the offer callback never
// fires (verified in C++ spike: test3 vs test2).
pc_->onLocalDescription([this](rtc::Description desc) {
latestLocalDesc_ = std::string(desc);
fprintf(stderr, "[binding] trickle desc, %zu bytes\n",
latestLocalDesc_.size());
});
pc_->onGatheringStateChange([this](rtc::PeerConnection::GatheringState gs) {
fprintf(stderr, "[binding] gathering state: %d\n", (int)gs);
if (gs == rtc::PeerConnection::GatheringState::Complete) {
// Use the getter — it returns the FULL SDP including candidates after
// gathering (trickle callbacks only carry the initial fragment).
auto ld = pc_->localDescription();
if (ld) {
latestLocalDesc_ = std::string(*ld);
fprintf(stderr, "[binding] final desc, %zu bytes\n",
latestLocalDesc_.size());
}
resolvePendingLocalDesc_();
}
});
}
private:
std::shared_ptr<rtc::PeerConnection> pc_;
std::shared_ptr<ThreadSafeFunction> stateCb_;
std::shared_ptr<ThreadSafeFunction> dcCb_;
std::string latestLocalDesc_;
std::shared_ptr<DeferredHolder> pendingDescDeferred_;
std::shared_ptr<ThreadSafeFunction> pendingDescTsfn_;
void resolvePendingLocalDesc_() {
if (!pendingDescDeferred_ || !pendingDescTsfn_) return;
auto holder = pendingDescDeferred_;
auto tsfn = pendingDescTsfn_;
pendingDescDeferred_.reset();
pendingDescTsfn_.reset();
std::string sdp = latestLocalDesc_;
tsfn->BlockingCall([sdp, holder](Napi::Env e, Function) {
holder->deferred.Resolve(String::New(e, sdp));
});
}
Napi::Value State(const Napi::CallbackInfo& info) {
return String::New(info.Env(),
pc_ ? stateToString(pc_->state()) : "closed");
}
// createOffer() -> Promise<string> — sets local description, waits for
// ICE gathering to complete (so candidates are in the SDP), resolves SDP.
Napi::Value CreateOffer(const Napi::CallbackInfo& info) {
Napi::Env env = info.Env();
auto holder = std::make_shared<DeferredHolder>(Promise::Deferred::New(env));
if (!pc_) {
holder->deferred.Reject(Error::New(env, "peer closed").Value());
return holder->deferred.Promise();
}
// createDataChannel already triggers negotiation in libdatachannel 0.24 —
// if gathering already completed, resolve immediately from the cached SDP.
if (!latestLocalDesc_.empty()) {
auto tsfn = std::make_shared<ThreadSafeFunction>(ThreadSafeFunction::New(
env, Function::New(env, [](const CallbackInfo&) {}), "desc", 0, 1));
std::string sdp = latestLocalDesc_;
tsfn->BlockingCall([sdp, holder](Napi::Env e, Function) {
holder->deferred.Resolve(String::New(e, sdp));
});
return holder->deferred.Promise();
}
if (pendingDescDeferred_) {
pendingDescDeferred_->deferred.Reject(
Error::New(env, "previous negotiation still pending").Value());
}
pendingDescDeferred_ = holder;
pendingDescTsfn_ = std::make_shared<ThreadSafeFunction>(
ThreadSafeFunction::New(env, Function::New(env, [](const CallbackInfo&) {}),
"desc", 0, 1));
fprintf(stderr, "[binding] calling setLocalDescription(Offer)\n");
try {
pc_->setLocalDescription(rtc::Description::Type::Offer);
fprintf(stderr, "[binding] setLocalDescription returned OK\n");
} catch (const std::exception& e) {
pendingDescDeferred_.reset();
fprintf(stderr, "[binding] setLocalDescription THREW: %s\n", e.what());
throw Error::New(env, e.what());
}
return holder->deferred.Promise();
}
// createAnswer(offerSdp: string) -> Promise<string>
Napi::Value CreateAnswer(const Napi::CallbackInfo& info) {
Napi::Env env = info.Env();
std::string offer = info[0].As<String>().Utf8Value();
auto holder = std::make_shared<DeferredHolder>(Promise::Deferred::New(env));
if (!pc_) {
holder->deferred.Reject(Error::New(env, "peer closed").Value());
return holder->deferred.Promise();
}
if (pendingDescDeferred_) {
pendingDescDeferred_->deferred.Reject(
Error::New(env, "previous negotiation still pending").Value());
}
pendingDescDeferred_ = holder;
pendingDescTsfn_ = std::make_shared<ThreadSafeFunction>(
ThreadSafeFunction::New(env, Function::New(env, [](const CallbackInfo&) {}),
"desc", 0, 1));
try {
pc_->setRemoteDescription(
rtc::Description(offer, rtc::Description::Type::Offer));
fprintf(stderr, "[binding] answer: setRemoteDescription OK\n");
// libdatachannel 0.24 AUTO-GENERATES the answer when a remote offer is
// applied (verified in C++ spike test8/9: B desc type=Answer fires
// immediately with a=setup:active). Calling setLocalDescription() again
// would OVERWRITE it with a role=actpass SDP, which A rejects with
// "Illegal role actpass in remote answer description". So we do NOT call
// setLocalDescription here — we just wait for gathering complete and
// resolve with the auto-generated answer. This also matches @dank074's
// Discord voice flow.
} catch (const std::exception& e) {
pendingDescDeferred_.reset();
fprintf(stderr, "[binding] answer THREW: %s\n", e.what());
holder->deferred.Reject(Error::New(env, e.what()).Value());
}
return holder->deferred.Promise();
}
void SetRemoteDescription(const Napi::CallbackInfo& info) {
std::string sdp = info[0].As<String>().Utf8Value();
std::string type = info[1].As<String>().Utf8Value();
rtc::Description::Type t = (type == "answer")
? rtc::Description::Type::Answer
: rtc::Description::Type::Offer;
if (pc_) pc_->setRemoteDescription(rtc::Description(sdp, t));
}
void Close(const Napi::CallbackInfo& info) {
if (pc_) pc_->close();
}
void OnStateChange(const Napi::CallbackInfo& info) {
Function cb = info[0].As<Function>();
stateCb_ = std::make_shared<ThreadSafeFunction>(
ThreadSafeFunction::New(info.Env(), cb, "pc-state", 0, 1));
std::shared_ptr<rtc::PeerConnection> pc = pc_;
pc->onStateChange([this](rtc::PeerConnection::State state) {
if (stateCb_) {
std::string s = stateToString(state);
stateCb_->BlockingCall([s](Napi::Env env, Function cb) {
cb.Call({String::New(env, s)});
});
}
});
}
Napi::Value CreateDataChannel(const Napi::CallbackInfo& info) {
Napi::Env env = info.Env();
std::string label = info[0].As<String>().Utf8Value();
fprintf(stderr, "[binding] createDataChannel(%s)\n", label.c_str());
auto dc = pc_->createDataChannel(label);
Object obj = DataChannelWrap::NewInstance(env);
DataChannelWrap::Unwrap(obj)->Init(dc);
return obj;
}
Napi::Value AddTrack(const Napi::CallbackInfo& info) {
Napi::Env env = info.Env();
std::string mid = info[0].As<String>().Utf8Value();
std::string kind = info[1].As<String>().Utf8Value();
if (!pc_) throw Error::New(env, "peer closed");
fprintf(stderr, "[binding] addTrack(%s, %s) start\n", mid.c_str(), kind.c_str());
try {
std::shared_ptr<rtc::Track> track;
if (kind == "audio") {
// Opus payload type 120 (matches @dank074 CodecPayloadType.opus)
auto desc = rtc::Description::Audio(mid);
desc.addOpusCodec(120);
track = pc_->addTrack(desc);
} else {
// All video codecs with their payload types, matching WebRtcWrapper:
// H264 101/102, H265 103/104, VP8 105/106, VP9 107/108, AV1 109/110
auto desc = rtc::Description::Video(mid);
desc.addH264Codec(101);
desc.addRtxCodec(102, 101, 90000);
desc.addH265Codec(103);
desc.addRtxCodec(104, 103, 90000);
desc.addVP8Codec(105);
desc.addRtxCodec(106, 105, 90000);
desc.addVP9Codec(107);
desc.addRtxCodec(108, 107, 90000);
desc.addAV1Codec(109);
desc.addRtxCodec(110, 109, 90000);
track = pc_->addTrack(desc);
}
Object obj = TrackWrap::NewInstance(env);
TrackWrap::Unwrap(obj)->Init(track, env);
return obj;
} catch (const std::exception& e) {
fprintf(stderr, "[binding] addTrack THREW: %s\n", e.what());
throw Error::New(env, e.what());
}
}
void OnDataChannel(const Napi::CallbackInfo& info) {
Function cb = info[0].As<Function>();
dcCb_ = std::make_shared<ThreadSafeFunction>(
ThreadSafeFunction::New(info.Env(), cb, "dc", 0, 1));
std::shared_ptr<rtc::PeerConnection> pc = pc_;
pc->onDataChannel([this](std::shared_ptr<rtc::DataChannel> dc) {
if (dcCb_) {
auto dcPtr = dc;
dcCb_->BlockingCall([dcPtr](Napi::Env env, Function cb) {
Object obj = DataChannelWrap::NewInstance(env);
DataChannelWrap::Unwrap(obj)->Init(dcPtr);
cb.Call({obj});
});
}
});
}
};
Object InitAll(Napi::Env env, Object exports) {
exports.Set("PeerConnection", PeerConnectionWrap::Init(env));
exports.Set("DataChannel", DataChannelWrap::Init(env));
exports.Set("Track", TrackWrap::Init(env));
return exports;
}
NODE_API_MODULE(libdatachannel_min, InitAll)
// Definition for the static constructor references.
FunctionReference DataChannelWrap::dcConstructor;
FunctionReference TrackWrap::trackConstructor;
} // namespace
@@ -1,21 +0,0 @@
{
"targets": [
{
"target_name": "libdatachannel_min",
"sources": ["binding.cpp"],
"include_dirs": [
"<!(node -e \"console.log(process.env.NAPI_INCLUDE || (() => { try { return require('node-addon-api').include; } catch { return '/nonexistent'; } })())\")",
"<!(node -e \"const s=process.env.LDC_INCLUDE||'/nix/store/39a85gpfjqy3h3k8jwrwh7m9yc3inqw7-source';console.log(s+'/include')\")"
],
"libraries": [
"<!(node -e \"console.log(process.env.LDC_LIB || '/tmp/ldc-build/libdatachannel.so.0.24.0')\")"
],
"cflags": ["-std=c++17", "-fexceptions"],
"cflags_cc": ["-std=c++17", "-fexceptions"],
"defines": ["NAPI_CPP_EXCEPTIONS"],
"conditions": [
["OS=='linux'", { "cflags": ["-fvisibility=hidden"] }]
]
}
]
}
@@ -1,3 +0,0 @@
// libdatachannel-min — JS entry.
const native = require("./build/Release/datachannel_min.node");
module.exports = native;
@@ -1,17 +0,0 @@
{
"name": "libdatachannel-min",
"version": "0.1.0",
"description": "Minimal N-API binding to libdatachannel — PeerConnection, DataChannel, ICE, SDP (+ media tracks for GoLive)",
"main": "index.js",
"gypfile": true,
"scripts": {
"build": "node-gyp rebuild",
"test": "node test-handshake.js"
},
"dependencies": {
"node-addon-api": "^8.3.0"
},
"devDependencies": {
"node-gyp": "^11.5.0"
}
}
@@ -1,82 +0,0 @@
// Phase 0 spike: prove the minimal binding can do a full WebRTC handshake
// (offer/answer + ICE + DataChannel) between two local PeerConnections.
"use strict";
const { PeerConnection } = require("./build/Release/datachannel_min.node");
function log(...args) {
console.log("[spike]", ...args);
}
async function main() {
const pcA = new PeerConnection({ iceServers: [] });
const pcB = new PeerConnection({ iceServers: [] });
const stateLog = [];
pcA.onStateChange((s) => {
stateLog.push(`A:${s}`);
log("A state:", s);
});
pcB.onStateChange((s) => {
stateLog.push(`B:${s}`);
log("B state:", s);
});
// B waits for incoming DataChannel
const received = new Promise((resolve) => {
pcB.onDataChannel((dc) => {
log("B got incoming DataChannel");
dc.onOpen(() => log("B DataChannel open"));
dc.onMessage((msg) => {
log("B received message:", msg);
dc.send("pong from B");
resolve(msg);
});
});
});
// A creates an outgoing DataChannel
const dcA = pcA.createDataChannel("test");
dcA.onOpen(() => {
log("A DataChannel open — sending hello");
dcA.send("hello from A");
});
dcA.onMessage((msg) => {
log("A received reply:", msg);
});
// Offer/answer dance
log("A createOffer...");
const offer = await pcA.createOffer();
log("Offer SDP bytes:", offer.length);
log("B createAnswer...");
const answer = await pcB.createAnswer(offer);
log("Answer SDP bytes:", answer.length);
const setupMatch = answer.match(/a=setup:(\S+)/);
log("Answer setup role:", setupMatch ? setupMatch[1] : "NONE");
pcA.setRemoteDescription(answer, "answer");
// Wait for message roundtrip
const msg = await Promise.race([
received,
new Promise((_, rej) => setTimeout(() => rej(new Error("TIMEOUT waiting for datachannel message")), 15000)),
]);
log("ROUNDTRIP OK — B got:", msg);
log("States:", stateLog.join(" | "));
const aState = pcA.state();
const bState = pcB.state();
log("Final states — A:", aState, "B:", bState);
pcA.close();
pcB.close();
if (msg !== "hello from A") throw new Error("wrong message");
if (aState !== "connected" && aState !== "disconnected") throw new Error("A not connected: " + aState);
log("SPIKE PASSED ✅");
}
main().catch((e) => {
console.error("SPIKE FAILED:", e.message);
process.exit(1);
});
@@ -1,80 +0,0 @@
// Verify setPacketizer + sendFrame: two peers connect, audio+video tracks
// packetize real encoded frames (opus + AnnexB H264), RTP flows without crash.
"use strict";
const { PeerConnection } = require("./build/Release/datachannel_min.node");
function sleep(ms) { return new Promise((r) => setTimeout(r, ms)); }
async function main() {
const pcA = new PeerConnection({ iceServers: [] });
const pcB = new PeerConnection({ iceServers: [] });
const aAudio = pcA.addTrack("0", "audio");
const aVideo = pcA.addTrack("1", "video");
pcB.addTrack("0", "audio");
pcB.addTrack("1", "video");
let states = { a: "", b: "" };
pcA.onStateChange((s) => (states.a = s));
pcB.onStateChange((s) => (states.b = s));
// A: offer (createDataChannel not needed — tracks trigger negotiation)
const offer = await pcA.createOffer();
pcB.setRemoteDescription(offer, "offer");
const answer = await pcB.createAnswer(offer);
pcA.setRemoteDescription(answer, "answer");
// Wait for connected
for (let i = 0; i < 50; i++) {
if (states.a === "connected" && states.b === "connected") break;
await sleep(100);
}
console.log("[pkt] states:", states.a, states.b);
if (states.a !== "connected" || states.b !== "connected") {
console.log("PKT TEST FAILED: not connected");
process.exit(1);
}
// Setup packetizers on A (sender)
aAudio.setPacketizer("audio", 1234, 120, 48000, 5, 0, 1);
aVideo.setPacketizer("h264", 5678, 101, 90000, 5, 0, 10);
// Fake opus frame (20ms @48kHz stereo — payload can be any bytes)
const opusFrame = Buffer.alloc(160);
for (let i = 0; i < 160; i++) opusFrame[i] = i & 0xff;
// Fake AnnexB H264 frame: SPS + PPS + IDR slice
const sps = Buffer.from([0x00, 0x00, 0x00, 0x01, 0x67, 0x42, 0xc0, 0x1e, 0xd9, 0x01, 0x40, 0x7e]);
const pps = Buffer.from([0x00, 0x00, 0x00, 0x01, 0x68, 0xce, 0x3c, 0x80]);
const idr = Buffer.from([0x00, 0x00, 0x00, 0x01, 0x65, 0x88, 0x84, 0x00, 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07]);
const h264Frame = Buffer.concat([sps, pps, idr]);
// Send 10 audio frames (20ms each) + 3 video frames (33ms each)
for (let i = 0; i < 10; i++) {
aAudio.sendFrame(opusFrame);
aAudio.addTimestamp(960); // 20ms @ 48kHz
}
for (let i = 0; i < 3; i++) {
aVideo.sendFrame(h264Frame);
aVideo.addTimestamp(3000); // 33ms @ 90kHz
}
await sleep(500);
console.log("[pkt] after send: states:", states.a, states.b);
console.log("[pkt] audio track open:", aAudio.isOpen(), "| video track open:", aVideo.isOpen());
const ok = states.a === "connected" && aAudio.isOpen() && aVideo.isOpen();
console.log(ok ? "PKT TEST PASSED" : "PKT TEST FAILED");
pcA.close();
pcB.close();
process.exit(ok ? 0 : 1);
}
main().catch((e) => {
console.error("[pkt] FAILED:", e.message);
process.exit(1);
});
setTimeout(() => {
console.error("[pkt] TIMEOUT");
process.exit(1);
}, 25000);
@@ -1,33 +0,0 @@
// Verify addTrack produces SDP with audio+video media sections.
"use strict";
const { PeerConnection } = require("./build/Release/datachannel_min.node");
const pc = new PeerConnection({ iceServers: [] });
const audioTrack = pc.addTrack("0", "audio");
const videoTrack = pc.addTrack("1", "video");
pc.onStateChange((s) => console.log("[test-track] state:", s));
pc.createOffer().then((sdp) => {
const hasAudio = /^m=audio\s/m.test(sdp);
const hasVideo = /^m=video\s/m.test(sdp);
const audioPts = sdp.match(/a=rtpmap:(\d+) opus/g) || [];
const videoPts = sdp.match(/a=rtpmap:(\d+) H264/g) || [];
console.log("[test-track] SDP bytes:", sdp.length);
console.log("[test-track] m=audio:", hasAudio, "| m=video:", hasVideo);
console.log("[test-track] opus pt:", audioPts, "| H264 pt:", videoPts);
console.log("[test-track] audio track send ok:", typeof audioTrack.send === "function");
console.log("[test-track] video track send ok:", typeof videoTrack.send === "function");
const ok = hasAudio && hasVideo && audioPts.length > 0 && videoPts.length > 0;
console.log(ok ? "TRACK TEST PASSED" : "TRACK TEST FAILED");
pc.close();
process.exit(ok ? 0 : 1);
}).catch((e) => {
console.error("[test-track] FAILED:", e.message);
process.exit(1);
});
setTimeout(() => {
console.error("[test-track] TIMEOUT");
process.exit(1);
}, 20000);
-1
View File
@@ -28,7 +28,6 @@
"discord.js-selfbot-v13": "^3.7.1",
"dotenv": "^17.4.2",
"drizzle-orm": "^0.45.2",
"imghash": "^1.1.4",
"ioredis": "^5.11.0",
"libsodium-wrappers": "^0.8.4",
"lru-cache": "^11.5.1",
@@ -0,0 +1,49 @@
// Rewrite import specifiers in the compiled dist/ so the output runs under
// plain `node dist/index.js` (native ESM, no bundler / no tsx).
//
// Background: tsconfig uses moduleResolution:"bundler", so `tsc` emits BARE
// relative specifiers WITHOUT extensions (e.g. `import "./router"`) and leaves
// the `@/*` path-alias imports untouched. Node's native ESM resolver rejects
// extensionless relative specifiers and knows nothing about the `@/` alias, so
// the emitted dist/ crashes at startup (`ERR_MODULE_NOT_FOUND`). This script
// fixes both:
// 1. `@/foo` -> relative path to dist/foo.js
// 2. `./foo` / `../foo` -> `./foo.js` / `../foo.js` (append .js)
// Already-extensioned relative imports (.js/.json/.node/.mjs/.cjs) and bare
// package specifiers are left untouched (idempotent).
import { readFileSync, writeFileSync, existsSync, readdirSync } from "node:fs";
import { join, relative, dirname } from "node:path";
let count = 0;
function walk(dir) {
if (!existsSync(dir)) return;
for (const e of readdirSync(dir, { withFileTypes: true })) {
const p = join(dir, e.name);
if (e.isDirectory()) walk(p);
else if (e.name.endsWith(".js")) {
const c = readFileSync(p, "utf8");
const pat = /from\s+['"]([^'"]+)['"]/g;
const n = c.replace(pat, (m, spec) => {
if (spec.startsWith("@/")) {
const target = join("dist", spec.slice(2)) + ".js";
let rel = relative(dirname(p), target);
if (!rel.startsWith(".")) rel = "./" + rel;
return `from "${rel}"`;
}
if (
(spec.startsWith("./") || spec.startsWith("../")) &&
!/\.(js|json|node|mjs|cjs)$/.test(spec)
) {
return `from "${spec}.js"`;
}
return m;
});
if (n !== c) {
writeFileSync(p, n);
count++;
}
}
}
}
walk("dist");
console.log(`Fixed ${count} import specifiers in dist/`);
+100 -6
View File
@@ -3,7 +3,11 @@ import { inArray, lt } from "drizzle-orm";
import type { NodePgDatabase } from "drizzle-orm/node-postgres";
import { ConfigError, DatabaseError } from "@/shared/errors/index";
import { createChildLogger } from "@/shared/logger/index";
import { startPendingAIAnalysisWorker } from "../modules/ai-moderation/aiAnalyzer.js";
import {
getAnalysisQueueStatus,
startPendingAIAnalysisWorker,
} from "../modules/ai-moderation/aiAnalyzer.js";
import { workerPool } from "../modules/ai-moderation/circuitBreaker.js";
import { registerChannelTopicCapture } from "../modules/channel-topic/index.js";
import { CommandHandler } from "../modules/command-handler/commandHandler.js";
import {
@@ -11,6 +15,8 @@ import {
RedisEventPublisher,
} from "../modules/event-broadcaster/index.js";
import {
registerCollector,
setGauge,
startMetricsServer,
stopMetricsServer,
} from "../modules/gateway-metrics/index.js";
@@ -222,7 +228,10 @@ export async function initializeDiscordGateway() {
await initializeDatabase();
logger.info("PostgreSQL database initialized");
} catch (err) {
logger.error({ error: err }, "Failed to initialize database");
logger.error(
{ err, errorMsg: err instanceof Error ? err.message : String(err) },
"Failed to initialize database",
);
throw new DatabaseError(
`Database initialization failed: ${err instanceof Error ? err.message : String(err)}`,
);
@@ -267,7 +276,10 @@ export async function initializeDiscordGateway() {
});
client.on("error", (err) => {
logger.error({ error: err }, "Client error");
logger.error(
{ err, errorMsg: err instanceof Error ? err.message : String(err) },
"Client error",
);
});
process.on("SIGINT", () => {
@@ -279,15 +291,97 @@ export async function initializeDiscordGateway() {
});
process.on("uncaughtException", (err) => {
logger.error({ error: err }, "Uncaught exception");
const code =
typeof (err as NodeJS.ErrnoException).code === "string"
? (err as NodeJS.ErrnoException).code
: "";
// Transient stream-teardown errors (voice stop/disconnect races, child
// process stdin closed while we still write) are NOT fatal — crashing the
// gateway on EPIPE takes the whole bot offline mid-music. Log + continue.
if (
code === "EPIPE" ||
code === "ERR_STREAM_DESTROYED" ||
code === "ERR_STREAM_WRITE_AFTER_END" ||
code === "ECONNRESET"
) {
logger.warn(
{ error: err },
"Uncaught transient stream error — continuing",
);
return;
}
logger.error(
{
err,
errorMsg: err instanceof Error ? err.message : String(err),
stack: err?.stack,
},
"Uncaught exception",
);
gracefulShutdown("uncaughtException");
});
process.on("unhandledRejection", (reason, promise) => {
logger.error({ reason, promise }, "Unhandled rejection");
process.on("unhandledRejection", (reason) => {
const err =
reason instanceof Error ? reason : new Error(String(reason ?? "unknown"));
const code = (err as NodeJS.ErrnoException).code ?? "";
// Same transient-teardown policy as uncaughtException: a rejection that
// fires while a stream is being torn down (EPIPE after ffmpeg stdin
// closes, write-after-destroy, socket reset) must NOT take the whole
// gateway offline. Log detail + continue. Everything else still shuts
// down so real bugs surface.
if (
code === "EPIPE" ||
code === "ERR_STREAM_DESTROYED" ||
code === "ERR_STREAM_WRITE_AFTER_END" ||
code === "ECONNRESET"
) {
logger.warn(
{ error: err },
"Unhandled rejection transient stream error — continuing",
);
return;
}
logger.error({ error: err, reason: String(reason) }, "Unhandled rejection");
gracefulShutdown("unhandledRejection");
});
// ── Metrics: register live pipeline collectors before starting server ──
// These refresh on every scrape so Prometheus sees real AI-analysis
// queue depth, concurrency, and DB pool state instead of an empty stub.
registerCollector(() => {
if (!config.AI_ANALYSIS_ENABLED) return;
try {
const status = getAnalysisQueueStatus();
setGauge("ai_analysis_queued_conversations", status.queuedConversations);
setGauge("ai_analysis_active_batch_requests", status.activeRequests);
setGauge(
"ai_analysis_active_individual_requests",
status.activeIndividualRequests,
);
setGauge(
"ai_analysis_individual_in_flight",
status.individualInFlightCount,
);
setGauge(
"ai_analysis_individual_circuit_breaker_active",
status.individualCircuitBreakerActive ? 1 : 0,
);
if (typeof status.lastError === "string") {
setGauge("ai_analysis_last_error_present", status.lastError ? 1 : 0);
}
const pool = workerPool as unknown as {
_poolState?: { size: number; active: number };
};
if (pool._poolState) {
setGauge("ai_analysis_worker_threads", pool._poolState.size);
setGauge("ai_analysis_worker_threads_active", pool._poolState.active);
}
} catch (err) {
logger.warn({ error: String(err) }, "AI metrics collector failed");
}
});
// Start metrics server
startMetricsServer();
@@ -1,155 +0,0 @@
/**
* AnnexB bitstream reader/writer (RBSP + emulation prevention) ported
* from @dank074/discord-video-stream AnnexBBitstreamReaderWriter.js.
*/
export class AnnexBBitstreamReader {
private _buffer: Uint8Array;
private _byteOffset = 0;
private _bitOffset = 0;
constructor(buffer: Uint8Array) {
this._buffer = buffer;
}
readBits(count: number): number {
if (count === 0) return 0;
let result = 0;
while (count > 0) {
if (this._byteOffset >= this._buffer.length) {
throw new Error("Bad byte offset");
}
if (
this._bitOffset === 0 &&
this._byteOffset >= 2 &&
this._buffer[this._byteOffset - 2] === 0 &&
this._buffer[this._byteOffset - 1] === 0 &&
this._buffer[this._byteOffset] === 3
) {
// Skip over emulation prevention
this._byteOffset++;
}
if (this._bitOffset === 0 && count >= 8) {
result = (result << 8) | this._buffer[this._byteOffset++];
count -= 8;
} else {
const numBitsToRead = Math.min(count, 8 - this._bitOffset);
const mask = (1 << numBitsToRead) - 1;
const newBits =
(this._buffer[this._byteOffset] >>
(8 - this._bitOffset - numBitsToRead)) &
mask;
result = (result << numBitsToRead) | newBits;
count -= numBitsToRead;
this._bitOffset += numBitsToRead;
if (this._bitOffset === 8) {
this._bitOffset = 0;
this._byteOffset++;
}
}
}
return result;
}
readUnsigned(bits: number): number {
return this.readBits(bits);
}
readSigned(bits: number): number {
const unsigned = this.readUnsigned(bits);
if (unsigned & (1 << (bits - 1))) return unsigned - (1 << bits);
return unsigned;
}
readUnsignedExpGolomb(): number {
let leading0 = 0;
while (this.readBits(1) === 0) leading0++;
return (1 << leading0) + this.readBits(leading0) - 1;
}
readSignedExpGolomb(): number {
const unsigned = this.readUnsignedExpGolomb();
if (unsigned % 2 === 0) return unsigned / -2;
return (unsigned + 1) / 2;
}
}
export class AnnexBBitstreamWriter {
private _arr: number[] = [];
private _pendingByte = 0;
private _bitOffset = 0;
toBuffer(): Buffer {
return Buffer.from(this._arr);
}
flush(): void {
// Emulation prevention: insert 0x03 before 00 00
if (
this._pendingByte <= 3 &&
this._arr[this._arr.length - 1] === 0 &&
this._arr[this._arr.length - 2] === 0
) {
this._arr.push(3);
}
this._arr.push(this._pendingByte);
this._pendingByte = 0;
this._bitOffset = 0;
}
writeBits(bits: number, count: number): void {
while (count > 0) {
if (this._bitOffset === 0) {
if (count >= 8) {
this._pendingByte = (bits >> (count - 8)) & 0xff;
count -= 8;
this.flush();
} else {
const mask = (1 << count) - 1;
this._pendingByte |= (bits & mask) << (8 - count);
this._bitOffset = count;
count = 0;
}
} else {
const numBitsToWrite = Math.min(8 - this._bitOffset, count);
const bitsToWrite =
(bits >> (count - numBitsToWrite)) & ((1 << numBitsToWrite) - 1);
this._pendingByte |=
bitsToWrite << (8 - this._bitOffset - numBitsToWrite);
count -= numBitsToWrite;
this._bitOffset += numBitsToWrite;
if (this._bitOffset === 8) {
this._bitOffset = 0;
this.flush();
}
}
}
}
writeUnsigned(num: number, count: number): void {
if (num < 0) throw new Error("Expected a non-negative number");
this.writeBits(num, count);
}
writeSigned(num: number, count: number): void {
if (count <= 0) return;
if (count > 32) throw new Error("writeSigned supports up to 32 bits");
const mask =
count === 32 ? 0xffffffff >>> 0 : (((1 << count) >>> 0) - 1) >>> 0;
const unsigned = (num & mask) >>> 0;
this.writeBits(unsigned, count);
}
writeUnsignedExpGolomb(num: number): void {
if (num < 0) throw new Error("Expected a non-negative number");
num++;
const bitCount = 32 - Math.clz32(num >>> 0);
this.writeBits(0, bitCount - 1);
this.writeBits(num, bitCount);
}
writeSignedExpGolomb(num: number): void {
if (num < 0) this.writeUnsignedExpGolomb(-2 * num);
else this.writeUnsignedExpGolomb(2 * num - 1);
}
}
@@ -1,112 +0,0 @@
/**
* H264/H265 NAL helpers ported from @dank074/discord-video-stream
* AnnexBHelper.js. Only the H264 parts are used by GoLive (H264 encoder),
* H265 constants kept for completeness of the port.
*/
export enum H264NalUnitTypes {
Unspecified = 0,
CodedSliceNonIDR = 1,
CodedSlicePartitionA = 2,
CodedSlicePartitionB = 3,
CodedSlicePartitionC = 4,
CodedSliceIdr = 5,
SEI = 6,
SPS = 7,
PPS = 8,
AccessUnitDelimiter = 9,
EndOfSequence = 10,
EndOfStream = 11,
FillerData = 12,
SEIExtenstion = 13,
PrefixNalUnit = 14,
SubsetSPS = 15,
}
export enum H265NalUnitTypes {
TRAIL_N = 0,
TRAIL_R = 1,
TSA_N = 2,
TSA_R = 3,
STSA_N = 4,
STSA_R = 5,
RADL_N = 6,
RADL_R = 7,
RASL_N = 8,
RASL_R = 9,
RSV_VCL_N10 = 10,
RSV_VCL_R11 = 11,
RSV_VCL_N12 = 12,
RSV_VCL_R13 = 13,
RSV_VCL_N14 = 14,
RSV_VCL_R15 = 15,
BLA_W_LP = 16,
BLA_W_RADL = 17,
BLA_N_LP = 18,
IDR_W_RADL = 19,
IDR_N_LP = 20,
CRA_NUT = 21,
RSV_IRAP_VCL22 = 22,
RSV_IRAP_VCL23 = 23,
RSV_VCL24 = 24,
RSV_VCL25 = 25,
RSV_VCL26 = 26,
RSV_VCL27 = 27,
RSV_VCL28 = 28,
RSV_VCL29 = 29,
RSV_VCL30 = 30,
RSV_VCL31 = 31,
VPS_NUT = 32,
SPS_NUT = 33,
PPS_NUT = 34,
AUD_NUT = 35,
EOS_NUT = 36,
EOB_NUT = 37,
FD_NUT = 38,
PREFIX_SEI_NUT = 39,
SUFFIX_SEI_NUT = 40,
}
export const H264Helpers = {
getUnitType(frame: Uint8Array): number {
return frame[0] & 0x1f;
},
splitHeader(frame: Uint8Array): [Uint8Array, Uint8Array] {
return [frame.subarray(0, 1), frame.subarray(1)];
},
isAUD(unitType: number): boolean {
return unitType === H264NalUnitTypes.AccessUnitDelimiter;
},
};
export const H265Helpers = {
getUnitType(frame: Uint8Array): number {
return (frame[0] >> 1) & 0x3f;
},
splitHeader(frame: Uint8Array): [Uint8Array, Uint8Array] {
return [frame.subarray(0, 2), frame.subarray(2)];
},
isAUD(unitType: number): boolean {
return unitType === H265NalUnitTypes.AUD_NUT;
},
};
export const startCode3 = Buffer.from([0, 0, 1]);
/** Split an AnnexB bitstream into NAL units (start codes stripped). */
export function splitNalu(buf: Buffer): Buffer[] {
let temp: Buffer | null = buf;
const nalus: Buffer[] = [];
while (temp?.byteLength) {
let pos: number = temp.indexOf(startCode3);
let length = 3;
if (pos > 0 && temp[pos - 1] === 0) {
pos--;
length++;
}
const nalu = pos === -1 ? temp : temp.subarray(0, pos);
temp = pos === -1 ? null : temp.subarray(pos + length);
if (nalu.byteLength) nalus.push(nalu);
}
return nalus;
}
@@ -1,20 +0,0 @@
/**
* AudioStream feeds encoded opus frames into the WebRTC connection.
* Ported from @dank074/discord-video-stream AudioStream.js.
*/
import { BaseMediaStream } from "./BaseMediaStream.js";
import type { WebRtcConnWrapper } from "./WebRtcWrapper.js";
export class AudioStream extends BaseMediaStream {
_conn: WebRtcConnWrapper;
constructor(conn: WebRtcConnWrapper, noSleep = false) {
super("audio", noSleep);
this._conn = conn;
}
async _sendFrame(frame: Buffer, frametime: number): Promise<void> {
this._conn.sendAudioFrame(frame, frametime);
}
}
@@ -1,594 +0,0 @@
/**
* Base media connection for Discord GoLive ported from
* @dank074/discord-video-stream BaseMediaConnection.js.
*
* Owns the voice WebSocket (identify/select_protocol/heartbeat/resume),
* SDP negotiation against Discord's media server, DAVE E2E voice
* (via @snazzah/davey), and speaking/video attribute signaling.
*/
import { randomUUID } from "node:crypto";
import { EventEmitter } from "node:events";
import Davey from "@snazzah/davey";
import { CodecPayloadType } from "./CodecPayloadType.js";
import type { NativePeerConnection } from "./native.js";
import { isNativeAvailable } from "./native.js";
import { STREAMS_SIMULCAST } from "./utils.js";
import { VoiceOpCodes, VoiceOpCodesBinary } from "./VoiceOpCodes.js";
import { WebRtcConnWrapper } from "./WebRtcWrapper.js";
export interface MediaConnectionStatus {
hasSession: boolean;
hasToken: boolean;
started: boolean;
resuming: boolean;
}
export interface VideoAttribute {
fps: number;
width: number;
height: number;
}
export interface StreamerLike {
opts: Record<string, unknown>;
}
export class BaseMediaConnection extends EventEmitter {
interval: ReturnType<typeof setInterval> | null = null;
guildId: string | null = null;
channelId: string;
botId: string;
ws: WebSocket | null = null;
status: MediaConnectionStatus;
server: string | null = null; // websocket url
token: string | null = null;
session_id: string | null = null;
protected _webRtcWrapper: WebRtcConnWrapper;
_webRtcParams: {
address: string;
port: number;
audioSsrc: number;
videoSsrc: number;
rtxSsrc: number;
supportedEncryptionModes: string[];
} | null = null;
protected _closed = false;
ready: ((conn: WebRtcConnWrapper) => void) | null;
protected _streamer: StreamerLike;
protected _sequenceNumber = -1;
protected _daveSession: Davey.DAVESession | null = null;
protected _connectedUsers = new Set<string>();
protected _daveProtocolVersion = 0;
protected _davePendingTransitions = new Map<number, number>();
protected _daveDowngraded = false;
constructor(
streamer: StreamerLike,
guildId: string | null,
botId: string,
channelId: string,
callback: ((conn: WebRtcConnWrapper) => void) | null,
) {
super();
this._streamer = streamer;
this.status = {
hasSession: false,
hasToken: false,
started: false,
resuming: false,
};
this.guildId = guildId;
this.channelId = channelId;
this.botId = botId;
this.ready = callback;
this._webRtcWrapper = new WebRtcConnWrapper(this);
}
get type(): "guild" | "call" {
return this.guildId ? "guild" : "call";
}
get webRtcConn(): WebRtcConnWrapper {
return this._webRtcWrapper;
}
get webRtcParams(): BaseMediaConnection["_webRtcParams"] {
return this._webRtcParams;
}
get streamer(): StreamerLike {
return this._streamer;
}
/** daveChannelId — overridden in VoiceConnection (channelId) and StreamConnection (serverId - 1n). */
get daveChannelId(): string {
throw new Error("daveChannelId not implemented");
}
stop(): void {
this._closed = true;
this._webRtcWrapper.close();
this.ws?.close();
}
setSession(session_id: string): void {
this.session_id = session_id;
this.status.hasSession = true;
this.start();
}
setTokens(server: string, token: string): void {
this.token = token;
this.server = server;
this.status.hasToken = true;
this.start();
}
start(): void {
if (this.status.hasSession && this.status.hasToken) {
if (this.status.started) return;
this.status.started = true;
this.ws = new WebSocket(`wss://${this.server}/?v=8`);
this.ws.binaryType = "arraybuffer";
this.ws.addEventListener("open", () => {
if (this.status.resuming) {
this.status.resuming = false;
this.resume();
} else {
this.identify();
}
});
this.ws.addEventListener("error", (err) => {
console.error(err);
});
this.ws.addEventListener("close", (e) => {
const wasStarted = this.status.started;
this.interval && clearInterval(this.interval);
this.status.started = false;
const canResume = e.code === 4015 || e.code < 4000;
if (canResume && wasStarted) {
this.status.resuming = true;
this.start();
} else {
this._closed = true;
this._webRtcWrapper?.close();
}
});
this.setupEvents();
}
}
handleReady(d: {
ip: string;
port: number;
ssrc: number;
streams: { ssrc: number; rtx_ssrc: number }[];
modes: string[];
}): void {
// we hardcoded STREAMS_SIMULCAST, which will always be array of 1
const stream = d.streams[0];
this._webRtcParams = {
address: d.ip,
port: d.port,
audioSsrc: d.ssrc,
videoSsrc: stream.ssrc,
rtxSsrc: stream.rtx_ssrc,
supportedEncryptionModes: d.modes,
};
}
async handleProtocolAck(d: {
sdp?: string;
dave_protocol_version?: number;
}): Promise<void> {
if (!("sdp" in d)) throw new Error("Only WebRTC connections are allowed");
this._daveProtocolVersion = d.dave_protocol_version ?? 0;
this.initDave();
// Discord's SDP is garbage — generate our own from its pieces
let ip = "";
let port = "";
let iceUsername = "";
let icePassword = "";
let fingerprint = "";
let candidate = "";
for (const line of (d.sdp ?? "").split("\n")) {
if (line.startsWith("c=")) ip = line;
else if (line.startsWith("a=rtcp")) port = line.split(":")[1];
else if (line.startsWith("a=ice-ufrag")) iceUsername = line;
else if (line.startsWith("a=ice-pwd")) icePassword = line;
else if (line.startsWith("a=fingerprint")) fingerprint = line;
else if (line.startsWith("a=candidate")) candidate = line;
}
const audioPayloadType = CodecPayloadType.opus.payload_type;
const audioSection = `
m=audio ${port} UDP/TLS/RTP/SAVPF ${audioPayloadType}
${ip}
a=extmap:1 urn:ietf:params:rtp-hdrext:ssrc-audio-level
a=extmap:3 http://www.ietf.org/id/draft-holmer-rmcat-transport-wide-cc-extensions-01
a=setup:passive
a=mid:0
a=maxptime:60
a=inactive
${iceUsername}
${icePassword}
${fingerprint}
${candidate}
a=rtcp-mux
a=rtpmap:${audioPayloadType} opus/48000/2
a=fmtp:${audioPayloadType} minptime=10;useinbandfec=1;usedtx=1
a=rtcp-fb:${audioPayloadType} transport-cc
a=rtcp-fb:${audioPayloadType} nack
a=ice-lite
`.trim();
const videoPayloads = Object.values(CodecPayloadType).filter(
(el) => el.type === "video",
);
const videoPayloadTypes = videoPayloads.flatMap((el) => [
el.payload_type,
el.rtx_payload_type ?? 0,
]);
const videoSection = `
m=video ${port} UDP/TLS/RTP/SAVPF ${videoPayloadTypes.join(" ")}
${ip}
a=extmap:2 http://www.webrtc.org/experiments/rtp-hdrext/abs-send-time
a=extmap:3 http://www.ietf.org/id/draft-holmer-rmcat-transport-wide-cc-extensions-01
a=extmap:14 urn:ietf:params:rtp-hdrext:toffset
a=extmap:13 urn:3gpp:video-orientation
a=extmap:5 http://www.webrtc.org/experiments/rtp-hdrext/playout-delay
a=setup:passive
a=mid:1
a=inactive
${iceUsername}
${icePassword}
${fingerprint}
${candidate}
a=rtcp-mux
a=ice-lite
`.trim();
const videoRtpMap = videoPayloads
.flatMap((el) => [
`a=rtpmap:${el.payload_type} ${el.name}/90000`,
`a=rtpmap:${el.rtx_payload_type} rtx/90000`,
`a=fmtp:${el.rtx_payload_type} apt=${el.payload_type}`,
`a=rtcp-fb:${el.payload_type} ccm fir`,
`a=rtcp-fb:${el.payload_type} nack`,
`a=rtcp-fb:${el.payload_type} nack pli`,
`a=rtcp-fb:${el.payload_type} goog-remb`,
`a=rtcp-fb:${el.payload_type} transport-cc`,
])
.join("\n");
this._webRtcWrapper.webRtcConn?.setRemoteDescription(
[audioSection, videoSection, videoRtpMap].join("\n"),
"answer",
);
this.emit("select_protocol_ack");
}
initDave(): void {
if (this._daveProtocolVersion) {
if (this._daveSession) {
this._daveSession.reinit(
this._daveProtocolVersion,
this.botId,
this.daveChannelId,
);
} else {
this._daveSession = new Davey.DAVESession(
this._daveProtocolVersion,
this.botId,
this.daveChannelId,
);
}
this.sendOpcodeBinary(
VoiceOpCodesBinary.MLS_KEY_PACKAGE,
this._daveSession.getSerializedKeyPackage(),
);
} else if (this._daveSession) {
this._daveSession.reset();
this._daveSession.setPassthroughMode(true, 10);
}
}
processInvalidCommit(transitionId: number): void {
this.sendOpcode(VoiceOpCodes.MLS_INVALID_COMMIT_WELCOME, {
transition_id: transitionId,
});
this.initDave();
}
executePendingTransition(transitionId: number): void {
const newVersion = this._davePendingTransitions.get(transitionId);
if (newVersion === undefined) {
console.error("Unrecognized transition ID", { transitionId });
return;
}
const oldVersion = this._daveProtocolVersion;
this._daveProtocolVersion = newVersion;
if (oldVersion !== newVersion && newVersion === 0) {
// Downgraded
this._daveDowngraded = true;
} else if (transitionId > 0 && this._daveDowngraded) {
this._daveDowngraded = false;
this._daveSession?.setPassthroughMode(true, 10);
}
this._davePendingTransitions.delete(transitionId);
}
setupEvents(): void {
this.ws?.addEventListener("message", async (e) => {
if (e.data instanceof ArrayBuffer) {
this.handleBinaryMessages(Buffer.from(e.data));
return;
}
const { op, d, seq } = JSON.parse(e.data as string) as {
op: number;
// biome-ignore lint/suspicious/noExplicitAny: Discord voice WS payload is dynamically typed
d: any;
seq?: number;
};
if (seq) this._sequenceNumber = seq;
if (op === VoiceOpCodes.READY) {
this.handleReady(d);
this.setProtocols().then(() => this.ready?.(this._webRtcWrapper));
this.setVideoAttributes(false);
} else if (op >= 4000) {
console.error(`${this.constructor.name} connection error`, d);
} else if (op === VoiceOpCodes.HELLO) {
this.setupHeartbeat(d.heartbeat_interval);
} else if (op === VoiceOpCodes.SELECT_PROTOCOL_ACK) {
await this.handleProtocolAck(d);
} else if (op === VoiceOpCodes.SPEAKING) {
// ignore speaking updates
} else if (op === VoiceOpCodes.HEARTBEAT_ACK) {
// ignore heartbeat acknowledgements
} else if (op === VoiceOpCodes.RESUMED) {
this.status.started = true;
} else if (op === VoiceOpCodes.CLIENTS_CONNECT) {
d.user_ids.forEach((id: string) => {
this._connectedUsers.add(id);
});
} else if (op === VoiceOpCodes.CLIENT_DISCONNECT) {
this._connectedUsers.delete(d.user_id);
} else if (op === VoiceOpCodes.DAVE_PREPARE_TRANSITION) {
this._davePendingTransitions.set(d.transition_id, d.protocol_version);
if (d.transition_id === 0) {
this.executePendingTransition(d.transition_id);
} else {
if (d.protocol_version === 0) {
this._daveSession?.setPassthroughMode(true, 120);
}
this.sendOpcode(VoiceOpCodes.DAVE_TRANSITION_READY, {
transition_id: d.transition_id,
});
}
} else if (op === VoiceOpCodes.DAVE_EXECUTE_TRANSITION) {
this.executePendingTransition(d.transition_id);
} else if (op === VoiceOpCodes.DAVE_PREPARE_EPOCH) {
if (d.epoch === 1) {
this._daveProtocolVersion = d.protocol_version;
this.initDave();
}
}
});
}
handleBinaryMessages(msg: Buffer): void {
this._sequenceNumber = msg.readUint16BE(0);
const op = msg.readUint8(2);
switch (op) {
case VoiceOpCodesBinary.MLS_EXTERNAL_SENDER: {
this._daveSession?.setExternalSender(msg.subarray(3));
break;
}
case VoiceOpCodesBinary.MLS_PROPOSALS: {
const optype = msg.readUint8(3);
if (!this._daveSession) break;
const { commit, welcome } = this._daveSession.processProposals(
optype,
msg.subarray(4),
[...this._connectedUsers],
);
if (commit) {
this.sendOpcodeBinary(
VoiceOpCodesBinary.MLS_COMMIT_WELCOME,
welcome ? Buffer.concat([commit, welcome]) : commit,
);
}
break;
}
case VoiceOpCodesBinary.MLS_ANNOUNCE_COMMIT_TRANSITION: {
const transitionId = msg.readUInt16BE(3);
try {
this._daveSession?.processCommit(msg.subarray(5));
if (transitionId) {
this._davePendingTransitions.set(
transitionId,
this._daveProtocolVersion,
);
this.sendOpcode(VoiceOpCodes.DAVE_TRANSITION_READY, {
transition_id: transitionId,
});
}
} catch (e) {
console.debug("MLS commit errored", e);
this.processInvalidCommit(transitionId);
}
break;
}
case VoiceOpCodesBinary.MLS_WELCOME: {
const transitionId = msg.readUInt16BE(3);
try {
this._daveSession?.processWelcome(msg.subarray(5));
if (transitionId) {
this._davePendingTransitions.set(
transitionId,
this._daveProtocolVersion,
);
this.sendOpcode(VoiceOpCodes.DAVE_TRANSITION_READY, {
transition_id: transitionId,
});
}
} catch (e) {
console.debug("MLS welcome errored", e);
this.processInvalidCommit(transitionId);
}
break;
}
}
}
get daveReady(): boolean {
return !!this._daveProtocolVersion && !!this._daveSession?.ready;
}
get daveSession(): Davey.DAVESession | null {
return this._daveSession;
}
setupHeartbeat(interval: number): void {
if (this.interval) {
clearInterval(this.interval);
}
this.interval = setInterval(() => {
try {
this.sendOpcode(VoiceOpCodes.HEARTBEAT, {
t: Date.now(),
seq_ack: this._sequenceNumber,
});
} catch {
/* ignore */
}
}, interval);
}
sendOpcode(code: number, data: unknown): void {
if (this.ws?.readyState !== WebSocket.OPEN) return;
this.ws.send(JSON.stringify({ op: code, d: data }));
}
sendOpcodeBinary(code: number, data: Uint8Array): void {
if (this.ws?.readyState !== WebSocket.OPEN) return;
const buf = Buffer.allocUnsafe(data.length + 1);
buf.writeUInt8(code);
Buffer.from(data).copy(buf, 1);
this.ws.send(buf);
}
/** serverId — overridden in VoiceConnection (guildId ?? channelId) and StreamConnection (rtc_server_id). */
get serverId(): string | null {
throw new Error("serverId not implemented");
}
/** identifies with media server with credentials */
identify(): void {
if (!this.serverId) throw new Error("Server ID is null or empty");
if (!this.session_id) throw new Error("Session ID is null or empty");
if (!this.token) throw new Error("Token is null or empty");
this.sendOpcode(VoiceOpCodes.IDENTIFY, {
server_id: this.serverId,
user_id: this.botId,
session_id: this.session_id,
token: this.token,
video: true,
streams: STREAMS_SIMULCAST,
max_dave_protocol_version: Davey.DAVE_PROTOCOL_VERSION ?? 0,
});
}
resume(): void {
if (!this.serverId) throw new Error("Server ID is null or empty");
if (!this.session_id) throw new Error("Session ID is null or empty");
if (!this.token) throw new Error("Token is null or empty");
this.sendOpcode(VoiceOpCodes.RESUME, {
server_id: this.serverId,
session_id: this.session_id,
token: this.token,
seq_ack: this._sequenceNumber,
});
}
/** Sets protocols and ip data used for video and audio (vp8 video, opus audio). */
async setProtocols(): Promise<void> {
if (!this._webRtcParams) throw new Error("WebRTC parameters not set");
if (!isNativeAvailable()) {
throw new Error(
"libdatachannel-min native binding not built — cannot start GoLive",
);
}
const reconnect = () => {
const webRtcConn = this._webRtcWrapper.initWebRtc();
webRtcConn.onStateChange((state) => {
if (state === "closed" && !this._closed) reconnect();
});
this._webRtcWrapper.onLocalDescription = (sdp) => {
const rtc_connection_id = randomUUID();
this.sendOpcode(VoiceOpCodes.SELECT_PROTOCOL, {
protocol: "webrtc",
codecs: Object.values(CodecPayloadType),
data: sdp,
sdp,
rtc_connection_id,
});
};
// createOffer (binding resolves full SDP incl. candidates after gathering)
void webRtcConn.createOffer().then((sdp) => {
this._webRtcWrapper.onLocalDescription?.(sdp);
});
};
reconnect();
return new Promise((resolve) => {
this.once("select_protocol_ack", () => resolve());
});
}
setVideoAttributes(enabled: boolean, attr?: VideoAttribute): void {
if (!this._webRtcParams) throw new Error("WebRTC parameters not set");
const { audioSsrc, videoSsrc, rtxSsrc } = this._webRtcParams;
if (!enabled) {
this.sendOpcode(VoiceOpCodes.VIDEO, {
audio_ssrc: audioSsrc,
video_ssrc: 0,
rtx_ssrc: 0,
streams: [],
});
} else {
if (!attr) throw new Error("Need to specify video attributes");
this.sendOpcode(VoiceOpCodes.VIDEO, {
audio_ssrc: audioSsrc,
video_ssrc: videoSsrc,
rtx_ssrc: rtxSsrc,
streams: [
{
type: "video",
rid: "100",
ssrc: videoSsrc,
active: true,
quality: 100,
rtx_ssrc: rtxSsrc,
// hardcode the max bitrate because we don't really know anyway
max_bitrate: 10000 * 1000,
max_framerate: enabled ? attr.fps : 0,
max_resolution: {
type: "fixed",
width: attr.width,
height: attr.height,
},
},
],
});
}
}
/** Set speaking status */
setSpeaking(speaking: boolean): void {
if (!this._webRtcParams) throw new Error("WebRTC connection not ready");
this.sendOpcode(VoiceOpCodes.SPEAKING, {
delay: 0,
speaking: speaking ? 1 : 0,
ssrc: this._webRtcParams.audioSsrc,
});
}
}
export type { NativePeerConnection };
@@ -1,175 +0,0 @@
/**
* BaseMediaStream pacing/sync for GoLive frames. Ported from
* @dank074/discord-video-stream BaseMediaStream.js, minus node-av's
* AVFrame (frames are plain objects here) and debug-level (uses the GMW
* logger instead).
*/
import { Writable } from "node:stream";
import { setTimeout as sleep } from "node:timers/promises";
export interface GoLiveFrame {
data: Buffer | null;
pts: number;
duration: number;
timeBase: { num: number; den: number };
free?: () => void;
}
export class BaseMediaStream extends Writable {
_pts: number | undefined;
_syncTolerance = 20;
_noSleep: boolean;
_startTime: number | undefined;
_startPts: number | undefined;
_sync = true;
_syncStream: BaseMediaStream | undefined;
_type: string;
constructor(type: string, noSleep = false) {
super({ objectMode: true, highWaterMark: 0 });
this._type = type;
this._noSleep = noSleep;
}
get sync(): boolean {
return this._sync;
}
set sync(val: boolean) {
this._sync = val;
}
get syncStream(): BaseMediaStream | undefined {
return this._syncStream;
}
set syncStream(stream: BaseMediaStream | undefined) {
if (stream !== undefined && this === stream.syncStream) {
throw new Error("Cannot sync 2 streams with eachother");
}
this._syncStream = stream;
}
get noSleep(): boolean {
return this._noSleep;
}
set noSleep(val: boolean) {
this._noSleep = val;
if (!val) this.resetTimingCompensation();
}
get pts(): number | undefined {
return this._pts;
}
get syncTolerance(): number {
return this._syncTolerance;
}
set syncTolerance(n: number) {
if (n < 0) return;
this._syncTolerance = n;
}
async _sendFrame(_frame: Buffer, _frametime: number): Promise<void> {
throw new Error("Not implemented");
}
ptsDelta(): number | undefined {
if (this.pts !== undefined && this.syncStream?.pts !== undefined) {
return this.pts - this.syncStream.pts;
}
return undefined;
}
isAhead(): boolean {
const delta = this.ptsDelta();
return (
this.syncStream?.writableEnded === false &&
delta !== undefined &&
delta > this.syncTolerance
);
}
isBehind(): boolean {
const delta = this.ptsDelta();
return (
this.syncStream?.writableEnded === false &&
delta !== undefined &&
delta < -this.syncTolerance
);
}
resetTimingCompensation(): void {
this._startTime = this._startPts = undefined;
}
async _write(
frame: GoLiveFrame,
_encoding: BufferEncoding,
callback: (error?: Error | null) => void,
): Promise<void> {
const { data, pts, duration, timeBase } = frame;
if (!data) {
frame.free?.();
callback();
return;
}
const frametime = (Number(duration) / timeBase.den) * timeBase.num * 1000;
const start_sendFrame = performance.now();
await this._sendFrame(Buffer.from(data), frametime);
const end_sendFrame = performance.now();
this._pts = (Number(pts) / timeBase.den) * timeBase.num * 1000;
this.emit("pts", this._pts);
const sendTime = end_sendFrame - start_sendFrame;
const ratio = sendTime / frametime;
if (ratio > 1) {
// Frame takes longer to send than its frametime — warn once per 100
if (
this._lastWarnedRatio === undefined ||
ratio > this._lastWarnedRatio
) {
this._lastWarnedRatio = ratio;
}
}
this._startTime ??= start_sendFrame;
this._startPts ??= this._pts;
const sleepMs = Math.max(
0,
this._pts -
this._startPts +
frametime -
(end_sendFrame - this._startTime),
);
if (this._noSleep || sleepMs === 0) {
callback(null);
} else if (this.sync && this.isBehind()) {
// Stream is behind — don't sleep for this frame
this.resetTimingCompensation();
callback(null);
} else if (this.sync && this.isAhead()) {
// Stream is ahead — wait until the sync stream catches up
do {
await sleep(frametime);
} while (this.sync && this.isAhead());
this.resetTimingCompensation();
callback(null);
} else {
await sleep(sleepMs);
callback(null);
}
frame.free?.();
}
_lastWarnedRatio: number | undefined;
_destroy(
error: Error | null,
callback: (error?: Error | null) => void,
): void {
super._destroy(error, callback);
this.syncStream = undefined;
}
}
@@ -1,71 +0,0 @@
/** Payload types for Discord GoLive media — ported from @dank074/discord-video-stream. */
export interface CodecPayloadTypeEntry {
name: string;
type: "audio" | "video";
clockRate: number;
priority: number;
payload_type: number;
rtx_payload_type?: number;
encode?: boolean;
decode?: boolean;
}
export const CodecPayloadType: Record<string, CodecPayloadTypeEntry> = {
opus: {
name: "opus",
type: "audio",
clockRate: 48000,
priority: 1000,
payload_type: 120,
},
H264: {
name: "H264",
type: "video",
clockRate: 90000,
priority: 1000,
payload_type: 101,
rtx_payload_type: 102,
encode: true,
decode: true,
},
H265: {
name: "H265",
type: "video",
clockRate: 90000,
priority: 1000,
payload_type: 103,
rtx_payload_type: 104,
encode: true,
decode: true,
},
VP8: {
name: "VP8",
type: "video",
clockRate: 90000,
priority: 1000,
payload_type: 105,
rtx_payload_type: 106,
encode: true,
decode: true,
},
VP9: {
name: "VP9",
type: "video",
clockRate: 90000,
priority: 1000,
payload_type: 107,
rtx_payload_type: 108,
encode: true,
decode: true,
},
AV1: {
name: "AV1",
type: "video",
clockRate: 90000,
priority: 1000,
payload_type: 109,
rtx_payload_type: 110,
encode: true,
decode: true,
},
};
@@ -1,377 +0,0 @@
/**
* Lightweight demuxer replaces node-av's LibavDemuxer for GoLive.
*
* Spawns ffmpeg to remux input into H264 AnnexB on stdout (video only
* screen share doesn't need to mux audio into the demuxer; audio goes
* separately). This replaces the 114MB node-av binary with a plain ffmpeg
* spawn.
*
* Each video "frame" emitted is a complete NAL sequence terminated by a
* keyframe boundary (IDR). Audio is not extracted here for GoLive with
* audio, the NUT mux + full demuxer would be needed; screen share audio is
* handled via a separate ffmpeg instance (see getDirectScreenInput).
*/
import { spawn } from "node:child_process";
import { randomUUID } from "node:crypto";
import { createWriteStream, existsSync, readdirSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { PassThrough } from "node:stream";
/**
* Resolve ffmpeg/ffprobe binary. Prefers explicit env override, then PATH,
* then a Nix-store ffmpeg-headless (the GMW flake provides it in the service
* profile, but dev shells / tests may not have it on PATH).
*/
function resolveBin(name: "ffmpeg"): string {
const override = process.env.FFMPEG_PATH;
if (override && existsSync(override)) return override;
// Nix store scan: <store>/<hash>-ffmpeg-headless-*/bin/<name>
const store = "/nix/store";
if (existsSync(store)) {
const entries = readdirSync(store);
for (const entry of entries) {
if (!entry.includes("ffmpeg-headless-")) continue;
const candidate = join(store, entry, "bin", name);
if (existsSync(candidate)) return candidate;
}
}
return name; // fall back to PATH
}
const FFMPEG = resolveBin("ffmpeg");
export const AVCodecID = {
AV_CODEC_ID_H264: 27,
AV_CODEC_ID_HEVC: 173,
AV_CODEC_ID_VP8: 139,
AV_CODEC_ID_VP9: 167,
AV_CODEC_ID_AV1: 225,
AV_CODEC_ID_OPUS: 86019,
} as const;
export type AVCodecID = (typeof AVCodecID)[keyof typeof AVCodecID];
export const AV_PKT_FLAG_KEY = 1;
export interface Frame {
data: Buffer | null;
pts: number;
duration: number;
timeBase: { num: number; den: number };
flags: number;
streamIndex: number;
free(): void;
}
export interface DemuxedStream {
codec: number;
codecName: string;
width: number;
height: number;
framerate_num: number;
framerate_den: number;
sample_rate: number;
stream: PassThrough;
}
/**
* Probe a media file for stream info using ffmpeg's stderr (the
* ffmpeg-headless Nix package ships ffmpeg but not ffprobe). Returns
* stream descriptors in the same shape ffprobe -show_streams would.
*/
export async function probeStreams(
url: string,
): Promise<Array<Record<string, unknown>>> {
return new Promise((resolve, reject) => {
const proc = spawn(FFMPEG, [
"-hide_banner",
"-loglevel",
"info",
"-i",
url,
"-f",
"null",
"-",
]);
let stderr = "";
proc.stderr.on("data", (d: Buffer) => (stderr += d.toString()));
proc.on("close", () => {
// Parse "Stream #0:0: Video: h264 (High), yuv420p, 640x360, 30 fps"
const streams: Array<Record<string, unknown>> = [];
const re = /Stream #0:(\d+): (Video|Audio): ([^,]+)/g;
let m: RegExpExecArray | null;
// biome-ignore lint/suspicious/noAssignInExpressions: regex loop idiom
while ((m = re.exec(stderr)) !== null) {
const [full, idx, kind, codecRaw] = m;
void full;
const codecName = codecRaw.split(" ")[0].toLowerCase();
const stream: Record<string, unknown> = {
index: Number(idx),
codec_type: kind.toLowerCase(),
codec_name: codecName,
width: 0,
height: 0,
r_frame_rate: "0/1",
sample_rate: 0,
};
// dimensions: "640x360"
const dim = /(\d{2,5})x(\d{2,5})/.exec(stderr.slice(m.index));
if (dim) {
stream.width = Number(dim[1]);
stream.height = Number(dim[2]);
}
// fps: "30 fps" or "29.97 fps"
const fps = /(\d+(?:\.\d+)?) fps/.exec(stderr.slice(m.index));
if (fps) {
const v = Number(fps[1]);
stream.r_frame_rate = `${Math.round(v * 1000)}/1000`;
}
// sample rate for audio: "48000 Hz"
const sr = /(\d+) Hz/.exec(stderr.slice(m.index));
if (sr) stream.sample_rate = Number(sr[1]);
streams.push(stream);
}
resolve(streams);
});
proc.on("error", (err) => reject(err));
});
}
/**
* Demux input (URL string or readable stream) into video frames on a
* PassThrough. Uses ffmpeg -f h264 -c copy for video-only AnnexB output.
* Returns stream info + the video pipe. Audio is not extracted (GoLive
* screen share sends silence / uses Discord's mixed audio).
*/
export async function demux(
input: string | PassThrough,
_opts: { format: string },
): Promise<{
video: DemuxedStream | undefined;
audio: DemuxedStream | undefined;
close: () => void;
}> {
const _label = randomUUID();
const vPipe = new PassThrough({ objectMode: true, highWaterMark: 128 });
const aPipe = new PassThrough({ objectMode: true, highWaterMark: 128 });
// For stream input, spool to a temp file first so ffprobe can inspect it
// (ffprobe needs a seekable file; pipes can't be re-read). The stream is
// fully consumed before ffmpeg starts — acceptable for screen-share
// sources which are already fully buffered by yt-dlp in practice.
let spoolPath: string | null = null;
const cleanupSpool = () => {
if (spoolPath) {
import("node:fs").then(({ unlink }) => unlink(spoolPath!, () => {}));
spoolPath = null;
}
};
let effectiveInput: string;
if (typeof input === "string") {
effectiveInput = input;
} else {
spoolPath = join(tmpdir(), `golive-demux-${_label}.h264`);
const ws = createWriteStream(spoolPath);
await new Promise<void>((resolve, reject) => {
input.pipe(ws);
input.on("error", reject);
ws.on("finish", resolve);
ws.on("error", reject);
});
effectiveInput = spoolPath;
}
// Probe for codec + dimensions
let streams: Array<Record<string, unknown>> = [];
try {
streams = await probeStreams(effectiveInput);
} catch (_e) {
// probe failed (e.g. raw h264 without container) — infer h264 default
streams = [];
}
const v = streams.find((s) => s.codec_type === "video");
const a = streams.find((s) => s.codec_type === "audio");
let vInfo: DemuxedStream | undefined;
let aInfo: DemuxedStream | undefined;
if (v) {
const codecName = (v.codec_name as string) ?? "h264";
const rFrame = (v.r_frame_rate as string) ?? "0/1";
const [num, den] = rFrame.split("/").map((n) => Number(n));
vInfo = {
codec:
AVCodecID[
(codecName.toUpperCase() as keyof typeof AVCodecID) ??
"AV_CODEC_ID_H264"
] ?? AVCodecID.AV_CODEC_ID_H264,
codecName,
width: (v.width as number) ?? 0,
height: (v.height as number) ?? 0,
framerate_num: num ?? 0,
framerate_den: den ?? 1,
sample_rate: 0,
stream: vPipe,
};
} else {
// Probe failed (e.g. raw AnnexB h264 input) — still emit frames on the
// video pipe; playStream infers dimensions from the first frame.
vInfo = {
codec: AVCodecID.AV_CODEC_ID_H264,
codecName: "h264",
width: 0,
height: 0,
framerate_num: 0,
framerate_den: 1,
sample_rate: 0,
stream: vPipe,
};
}
if (a) {
const codecName = (a.codec_name as string) ?? "opus";
aInfo = {
codec:
AVCodecID[
(codecName.toUpperCase() as keyof typeof AVCodecID) ??
"AV_CODEC_ID_OPUS"
],
codecName,
width: 0,
height: 0,
framerate_num: 0,
framerate_den: 0,
sample_rate: Number(a.sample_rate) ?? 0,
stream: aPipe,
};
}
// Spawn ffmpeg — extract raw video (AnnexB for H264) to stdout
const args: string[] = [
"-hide_banner",
"-loglevel",
"error",
"-i",
effectiveInput,
"-c:v",
"copy",
"-an", // no audio in this minimal demuxer
"-f",
"h264",
"pipe:1",
];
const proc = spawn(FFMPEG, args, { stdio: ["ignore", "pipe", "pipe"] });
// Scan stdout for NAL units. Each NAL unit (between start codes) is one frame
// payload. We emit them individually; the packetizer chain handles FU-A.
let videoBuf = Buffer.alloc(0);
let frameCount = 0;
const emitFrame = (nal: Uint8Array, isKeyFrame: boolean) => {
vPipe.write({
data: Buffer.from(nal),
pts: frameCount,
duration: 1,
timeBase: { num: 1, den: 90000 },
flags: isKeyFrame ? AV_PKT_FLAG_KEY : 0,
streamIndex: 0,
free: () => {},
});
frameCount++;
};
if (proc.stdout) {
proc.stdout.on("data", (chunk: Buffer) => {
videoBuf = Buffer.concat([videoBuf, chunk]);
// Find start codes (00 00 01 or 00 00 00 01) and split NALs
let start = 0;
// If buffer starts with zeros, that's the first start code — emit from there
while (start < videoBuf.length) {
let scPos = -1;
for (let i = start + 1; i < videoBuf.length - 2; i++) {
if (
videoBuf[i] === 0 &&
videoBuf[i + 1] === 0 &&
videoBuf[i + 2] === 1
) {
scPos = i + 3;
break;
}
}
if (scPos === -1) break;
// Emit the NAL from `start` to `scPos` (but skip the start code bytes at `start`)
if (start < scPos) {
let nalStart = start;
// Skip start code bytes for the NAL itself (00 00 01)
if (
videoBuf[nalStart] === 0 &&
videoBuf[nalStart + 1] === 0 &&
videoBuf[nalStart + 2] === 1
) {
nalStart += 3;
} else if (
nalStart + 3 < scPos &&
videoBuf[nalStart] === 0 &&
videoBuf[nalStart + 1] === 0 &&
videoBuf[nalStart + 2] === 0 &&
videoBuf[nalStart + 3] === 1
) {
nalStart += 4;
}
const nal = videoBuf.subarray(nalStart, scPos);
// Trim trailing zero bytes (from start code overlap)
let end = nal.length;
while (end > 0 && nal[end - 1] === 0) end--;
if (end > 0) {
const nalTrimmed = nal.subarray(0, end);
const isIdr = (nalTrimmed[0] & 0x1f) === 5; // IDR
emitFrame(nalTrimmed, isIdr);
}
}
// Skip the 00 00 01 at scPos-3 to find next
start = scPos;
// But the next start code needs at least 3 bytes
if (start > videoBuf.length - 3) break;
}
// Keep remaining bytes (potential partial NAL or start code)
if (start > 0 && start < videoBuf.length) {
videoBuf = videoBuf.subarray(start);
} else if (videoBuf.length > 4) {
// No full NAL found, but avoid unbounded growth
// Keep a sliding window
videoBuf = videoBuf.subarray(videoBuf.length - 3);
}
});
proc.stdout.on("end", () => {
if (videoBuf.length > 0) {
let end = videoBuf.length;
while (end > 0 && videoBuf[end - 1] === 0) end--;
if (end > 0) emitFrame(videoBuf.subarray(0, end), false);
}
vPipe.end();
aPipe.end();
});
}
if (proc.stderr) {
proc.stderr.on("data", () => {
/* errors swallowed */
});
}
proc.on("close", () => {
vPipe.end();
aPipe.end();
});
const close = () => {
proc.kill("SIGTERM");
vPipe.end();
aPipe.end();
cleanupSpool();
};
return { video: vInfo, audio: aInfo, close };
}
@@ -1,51 +0,0 @@
/**
* Lightweight encoders config ported from @dank074/discord-video-stream
* encoders/software.js. Only software (libx264) is needed for GoLive.
*/
export interface EncoderSettings {
name: string;
options: string[];
outFilters?: string[];
globalOptions?: string[];
}
export interface EncoderSet {
H264: EncoderSettings;
H265: EncoderSettings;
VP8: EncoderSettings;
VP9: EncoderSettings;
AV1: EncoderSettings;
}
/** Software x264 encoder. Matches @dank074's software() defaults. */
export function software(
opts: {
x264?: { preset?: string; tune?: string };
x265?: { preset?: string; tune?: string };
} = {},
): () => EncoderSet {
const { x264, x265 } = opts;
const { preset: x264Preset = "superfast", tune: x264Tune = "film" } =
x264 ?? {};
const { preset: x265Preset = "superfast", tune: x265Tune } = x265 ?? {};
return () => ({
H264: {
name: "libx264",
options: ["-forced-idr 1", `-tune ${x264Tune}`, `-preset ${x264Preset}`],
},
H265: {
name: "libx265",
options: [
"-forced-idr 1",
...(x265Tune ? [`-tune ${x265Tune}`] : []),
`-preset ${x265Preset}`,
],
},
VP8: { name: "libvpx", options: ["-deadline 20000"] },
VP9: { name: "libvpx-vp9", options: ["-deadline 20000"] },
AV1: { name: "libsvtav1", options: [] },
});
}
export const Encoders = { software };
@@ -1,41 +0,0 @@
/** Discord gateway opcodes used by Streamer — ported from @dank074/discord-video-stream. */
export enum GatewayOpCodes {
DISPATCH = 0,
HEARTBEAT = 1,
IDENTIFY = 2,
PRESENCE_UPDATE = 3,
VOICE_STATE_UPDATE = 4,
VOICE_SERVER_PING = 5,
RESUME = 6,
RECONNECT = 7,
REQUEST_GUILD_MEMBERS = 8,
INVALID_SESSION = 9,
HELLO = 10,
HEARTBEAT_ACK = 11,
CALL_CONNECT = 13,
GUILD_SUBSCRIPTIONS = 14,
LOBBY_CONNECT = 15,
LOBBY_DISCONNECT = 16,
LOBBY_VOICE_STATES_UPDATE = 17,
STREAM_CREATE = 18,
STREAM_DELETE = 19,
STREAM_WATCH = 20,
STREAM_PING = 21,
STREAM_SET_PAUSED = 22,
REQUEST_GUILD_APPLICATION_COMMANDS = 24,
EMBEDDED_ACTIVITY_LAUNCH = 25,
EMBEDDED_ACTIVITY_CLOSE = 26,
EMBEDDED_ACTIVITY_UPDATE = 27,
REQUEST_FORUM_UNREADS = 28,
REMOTE_COMMAND = 29,
GET_DELETED_ENTITY_IDS_NOT_MATCHING_HASH = 30,
REQUEST_SOUNDBOARD_SOUNDS = 31,
SPEED_TEST_CREATE = 32,
SPEED_TEST_DELETE = 33,
REQUEST_LAST_MESSAGES = 34,
SEARCH_RECENT_MEMBERS = 35,
REQUEST_CHANNEL_STATUSES = 36,
GUILD_SUBSCRIPTIONS_BULK = 37,
GUILD_CHANNELS_RESYNC = 38,
REQUEST_CHANNEL_MEMBER_COUNT = 39,
}
@@ -1,291 +0,0 @@
/**
* H264 SPS VUI rewriter ported from @dank074/discord-video-stream
* SPSVUIRewriter.js. Rewrites the SPS so Discord's receiver applies
* bitstream restrictions (max_num_reorder_frames=0, max_dec_frame_buffering
* bounded) required for low-latency GoLive decode.
*/
import {
AnnexBBitstreamReader,
AnnexBBitstreamWriter,
} from "./AnnexBBitstreamReaderWriter.js";
export function rewriteSPSVUI(buffer: Uint8Array): Buffer {
const reader = new AnnexBBitstreamReader(buffer.subarray(1));
const writer = new AnnexBBitstreamWriter();
const readBit = (n = 1) => reader.readBits(n);
const writeBit = (v: number, n = 1) => writer.writeBits(v, n);
const readU = (n: number) => reader.readUnsigned(n);
const writeU = (v: number, n: number) => writer.writeUnsigned(v, n);
const readUE = () => reader.readUnsignedExpGolomb();
const writeUE = (v: number) => writer.writeUnsignedExpGolomb(v);
const readSE = () => reader.readSignedExpGolomb();
const writeSE = (v: number) => writer.writeSignedExpGolomb(v);
// Rewrite the NAL header
writeU(buffer[0], 8);
const profile_idc = readU(8);
writeU(profile_idc, 8);
const constraint_flags = readU(8);
writeU(constraint_flags, 8);
const level_idc = readU(8);
writeU(level_idc, 8);
const seq_parameter_set_id = readUE();
writeUE(seq_parameter_set_id);
// If profile in high profiles, additional fields
const highProfiles = new Set([
100, 110, 122, 244, 44, 83, 86, 118, 128, 138, 144,
]);
if (highProfiles.has(profile_idc)) {
const chroma_format_idc = readUE();
writeUE(chroma_format_idc);
if (chroma_format_idc === 3) {
const separate_colour_plane_flag = readBit(1);
writeBit(separate_colour_plane_flag, 1);
}
const bit_depth_luma_minus8 = readUE();
writeUE(bit_depth_luma_minus8);
const bit_depth_chroma_minus8 = readUE();
writeUE(bit_depth_chroma_minus8);
const qpprime_y_zero_transform_bypass_flag = readBit(1);
writeBit(qpprime_y_zero_transform_bypass_flag, 1);
const seq_scaling_matrix_present_flag = readBit(1);
writeBit(seq_scaling_matrix_present_flag, 1);
if (seq_scaling_matrix_present_flag) {
const scalingCount = chroma_format_idc !== 3 ? 8 : 12;
for (let i = 0; i < scalingCount; i++) {
const seq_scaling_list_present_flag = readBit(1);
writeBit(seq_scaling_list_present_flag, 1);
if (seq_scaling_list_present_flag) {
const size = i < 6 ? 16 : 64;
let lastScale = 8;
let nextScale = 8;
for (let j = 0; j < size; j++) {
const delta = readSE();
writeSE(delta);
nextScale = (lastScale + delta + 256) % 256;
if (nextScale !== 0) lastScale = nextScale;
}
}
}
}
}
const log2_max_frame_num_minus4 = readUE();
writeUE(log2_max_frame_num_minus4);
const pic_order_cnt_type = readUE();
writeUE(pic_order_cnt_type);
if (pic_order_cnt_type === 0) {
const log2_max_pic_order_cnt_lsb_minus4 = readUE();
writeUE(log2_max_pic_order_cnt_lsb_minus4);
} else if (pic_order_cnt_type === 1) {
const delta_pic_order_always_zero_flag = readBit(1);
writeBit(delta_pic_order_always_zero_flag, 1);
const offset_for_non_ref_pic = readSE();
writeSE(offset_for_non_ref_pic);
const offset_for_top_to_bottom_field = readSE();
writeSE(offset_for_top_to_bottom_field);
const num_ref_frames_in_pic_order_cnt_cycle = readUE();
writeUE(num_ref_frames_in_pic_order_cnt_cycle);
for (let i = 0; i < num_ref_frames_in_pic_order_cnt_cycle; i++) {
const offset_for_ref_frame = readSE();
writeSE(offset_for_ref_frame);
}
}
const max_num_ref_frames = readUE();
writeUE(max_num_ref_frames);
const gaps_in_frame_num_value_allowed_flag = readBit(1);
writeBit(gaps_in_frame_num_value_allowed_flag, 1);
const pic_width_in_mbs_minus1 = readUE();
writeUE(pic_width_in_mbs_minus1);
const pic_height_in_map_units_minus1 = readUE();
writeUE(pic_height_in_map_units_minus1);
const frame_mbs_only_flag = readBit(1);
writeBit(frame_mbs_only_flag, 1);
if (frame_mbs_only_flag === 0) {
const mb_adaptive_frame_field_flag = readBit(1);
writeBit(mb_adaptive_frame_field_flag, 1);
}
const direct_8x8_inference_flag = readBit(1);
writeBit(direct_8x8_inference_flag, 1);
const frame_cropping_flag = readBit(1);
writeBit(frame_cropping_flag, 1);
if (frame_cropping_flag) {
const frame_crop_left_offset = readUE();
writeUE(frame_crop_left_offset);
const frame_crop_right_offset = readUE();
writeUE(frame_crop_right_offset);
const frame_crop_top_offset = readUE();
writeUE(frame_crop_top_offset);
const frame_crop_bottom_offset = readUE();
writeUE(frame_crop_bottom_offset);
}
// https://webrtc.googlesource.com/src/+/5f2c9278f35e47ff72eb191669d473b7400c9f3e/common_video/h264/sps_vui_rewriter.cc#283
function addBitstreamRestriction() {
// motion_vectors_over_pic_boundaries_flag: u(1) — Default is 1 when not present.
writeBit(1, 1);
// max_bytes_per_pic_denom: ue(v) — Default is 2 when not present.
writeUE(2);
// max_bits_per_mb_denom: ue(v) — Default is 1 when not present.
writeUE(1);
// log2_max_mv_length_horizontal / vertical — both default to 16.
writeUE(16);
writeUE(16);
// IMPORTANT: max_num_reorder_frames must be 0 for low latency.
writeUE(0);
writeUE(max_num_ref_frames);
}
const vui_parameters_present_flag = readBit(1);
writeBit(1, 1);
// If no VUI exists, write one
if (!vui_parameters_present_flag) {
// aspect_ratio_info_present_flag, overscan_info_present_flag. Both u(1).
writeBit(0, 2);
// video_signal_type_present_flag, u(1) — write 0, ignore color space.
writeBit(0, 1);
// chroma_loc_info_present_flag, timing_info_present_flag,
// nal_hrd_parameters_present_flag, vcl_hrd_parameters_present_flag,
// pic_struct_present_flag — all u(1)
writeBit(0, 5);
// bitstream_restriction_flag: u(1)
writeBit(1, 1);
addBitstreamRestriction();
} else {
// VUI parsing and copying
const aspect_ratio_info_present_flag = readBit(1);
writeBit(aspect_ratio_info_present_flag, 1);
if (aspect_ratio_info_present_flag) {
const aspect_ratio_idc = readU(8);
writeU(aspect_ratio_idc, 8);
if (aspect_ratio_idc === 255) {
const sar_width = readU(16);
writeU(sar_width, 16);
const sar_height = readU(16);
writeU(sar_height, 16);
}
}
const overscan_info_present_flag = readBit(1);
writeBit(overscan_info_present_flag, 1);
if (overscan_info_present_flag) {
const overscan_appropriate_flag = readBit(1);
writeBit(overscan_appropriate_flag, 1);
}
// Read the video signal type, but don't copy it
const video_signal_type_present_flag = readBit(1);
writeBit(0, 1);
if (video_signal_type_present_flag) {
readBit(3); // _video_format
readBit(1); // _video_full_range_flag
const colour_description_present_flag = readBit(1);
if (colour_description_present_flag) {
readU(8); // _colour_primaries
readU(8); // _transfer_characteristics
readU(8); // _matrix_coeffs
}
}
const chroma_loc_info_present_flag = readBit(1);
writeBit(chroma_loc_info_present_flag, 1);
if (chroma_loc_info_present_flag) {
const chroma_sample_loc_type_top_field = readUE();
writeUE(chroma_sample_loc_type_top_field);
const chroma_sample_loc_type_bottom_field = readUE();
writeUE(chroma_sample_loc_type_bottom_field);
}
const timing_info_present_flag = readBit(1);
writeBit(timing_info_present_flag, 1);
if (timing_info_present_flag) {
const num_units_in_tick = readU(32);
writeU(num_units_in_tick, 32);
const time_scale = readU(32);
writeU(time_scale, 32);
const fixed_frame_rate_flag = readBit(1);
writeBit(fixed_frame_rate_flag, 1);
}
const nal_hrd_parameters_present_flag = readBit(1);
writeBit(nal_hrd_parameters_present_flag, 1);
if (nal_hrd_parameters_present_flag) {
// hrd_parameters()
const cpb_cnt_minus1 = readUE();
writeUE(cpb_cnt_minus1);
const bit_rate_scale = readBit(4);
writeBit(bit_rate_scale, 4);
const cpb_size_scale = readBit(4);
writeBit(cpb_size_scale, 4);
for (let i = 0; i <= cpb_cnt_minus1; i++) {
const bit_rate_value_minus1 = readUE();
writeUE(bit_rate_value_minus1);
const cpb_size_value_minus1 = readUE();
writeUE(cpb_size_value_minus1);
const cbr_flag = readBit(1);
writeBit(cbr_flag, 1);
}
const initial_cpb_removal_delay_length_minus1 = readBit(5);
writeBit(initial_cpb_removal_delay_length_minus1, 5);
const cpb_removal_delay_length_minus1 = readBit(5);
writeBit(cpb_removal_delay_length_minus1, 5);
const dpb_output_delay_length_minus1 = readBit(5);
writeBit(dpb_output_delay_length_minus1, 5);
const time_offset_length = readBit(5);
writeBit(time_offset_length, 5);
}
const vcl_hrd_parameters_present_flag = readBit(1);
writeBit(vcl_hrd_parameters_present_flag, 1);
if (vcl_hrd_parameters_present_flag) {
// hrd_parameters()
const cpb_cnt_minus1 = readUE();
writeUE(cpb_cnt_minus1);
const bit_rate_scale = readBit(4);
writeBit(bit_rate_scale, 4);
const cpb_size_scale = readBit(4);
writeBit(cpb_size_scale, 4);
for (let i = 0; i <= cpb_cnt_minus1; i++) {
const bit_rate_value_minus1 = readUE();
writeUE(bit_rate_value_minus1);
const cpb_size_value_minus1 = readUE();
writeUE(cpb_size_value_minus1);
const cbr_flag = readBit(1);
writeBit(cbr_flag, 1);
}
const initial_cpb_removal_delay_length_minus1 = readBit(5);
writeBit(initial_cpb_removal_delay_length_minus1, 5);
const cpb_removal_delay_length_minus1 = readBit(5);
writeBit(cpb_removal_delay_length_minus1, 5);
const dpb_output_delay_length_minus1 = readBit(5);
writeBit(dpb_output_delay_length_minus1, 5);
const time_offset_length = readBit(5);
writeBit(time_offset_length, 5);
}
if (nal_hrd_parameters_present_flag || vcl_hrd_parameters_present_flag) {
const low_delay_hrd_flag = readBit(1);
writeBit(low_delay_hrd_flag, 1);
}
const pic_struct_present_flag = readBit(1);
writeBit(pic_struct_present_flag, 1);
const bitstream_restriction_flag = readBit(1);
writeBit(1, 1);
if (!bitstream_restriction_flag) {
addBitstreamRestriction();
} else {
const motion_vectors_over_pic_boundaries_flag = readBit(1);
writeBit(motion_vectors_over_pic_boundaries_flag, 1);
const max_bytes_per_pic_denom = readUE();
writeUE(max_bytes_per_pic_denom);
const max_bits_per_mb_denom = readUE();
writeUE(max_bits_per_mb_denom);
const log2_max_mv_length_horizontal = readUE();
writeUE(log2_max_mv_length_horizontal);
const log2_max_mv_length_vertical = readUE();
writeUE(log2_max_mv_length_vertical);
readUE(); // _num_reorder_frames
writeUE(0);
readUE(); // _max_dec_frame_buffering
writeUE(max_num_ref_frames);
}
}
writeBit(1, 1); // rbsp_stop_one_bit
writer.flush();
return writer.toBuffer();
}
@@ -1,45 +0,0 @@
/**
* StreamConnection GoLive stream connection (screen share).
* Ported from @dank074/discord-video-stream StreamConnection.js.
*/
import { BaseMediaConnection } from "./BaseMediaConnection.js";
import { VoiceOpCodes } from "./VoiceOpCodes.js";
export class StreamConnection extends BaseMediaConnection {
_streamKey: string | null = null;
_serverId: string | null = null;
setSpeaking(speaking: boolean): void {
if (!this.webRtcParams) throw new Error("WebRTC connection not ready");
this.sendOpcode(VoiceOpCodes.SPEAKING, {
delay: 0,
speaking: speaking ? 2 : 0,
ssrc: this.webRtcParams.audioSsrc,
});
}
get daveChannelId(): string {
if (this._serverId === null) {
throw new Error("Server ID not set (this shouldn't happen)");
}
const channelId = BigInt(this._serverId) - 1n;
return channelId.toString();
}
get serverId(): string | null {
return this._serverId;
}
set serverId(id: string | null) {
this._serverId = id;
}
get streamKey(): string | null {
return this._streamKey;
}
set streamKey(value: string | null) {
this._streamKey = value;
}
}
@@ -1,280 +0,0 @@
/**
* Streamer gateway-level GoLive controller. Ported from
* @dank074/discord-video-stream Streamer.js.
*
* Drives the Discord gateway (VOICE_STATE_UPDATE, STREAM_CREATE, ...) and
* hands back a VoiceConnection / StreamConnection once the media server
* session is ready.
*/
import { EventEmitter } from "node:events";
import { GatewayOpCodes } from "./GatewayOpCodes.js";
import type { NativePeerConnection } from "./native.js";
import { StreamConnection } from "./StreamConnection.js";
import { generateStreamKey, parseStreamKey } from "./utils.js";
import { VoiceConnection } from "./VoiceConnection.js";
import type { WebRtcConnWrapper } from "./WebRtcWrapper.js";
/** Minimal surface of a discord.js-selfbot-v13 client used by Streamer. */
export interface StreamerClientLike {
user: { id: string; username?: string } | null;
token: string | null;
on(
event: "raw",
listener: (packet: { t: string; d: unknown }) => void,
): unknown;
ws: {
broadcast(data: { op: number; d: unknown }): void;
};
guilds?: {
// biome-ignore lint/suspicious/noExplicitAny: discord.js-selfbot client shape is dynamic
fetch(id: string): Promise<any>;
};
}
/** Minimal channel shape accepted by joinVoiceChannel. */
export interface VoiceChannelLike {
id: string;
type: string;
guildId?: string | null;
}
export class Streamer {
_voiceConnection: VoiceConnection | null = null;
_client: StreamerClientLike;
_gatewayEmitter = new EventEmitter();
constructor(client: StreamerClientLike) {
this._client = client;
// listen for gateway dispatch events
this.client.on("raw", (packet) => {
this._gatewayEmitter.emit(packet.t, packet.d);
});
}
get client(): StreamerClientLike {
return this._client;
}
get opts(): Record<string, unknown> {
return {};
}
get voiceConnection(): VoiceConnection | null {
return this._voiceConnection;
}
sendOpcode(code: number, data: unknown): void {
this.client.ws.broadcast({ op: code, d: data });
}
joinVoiceChannel(channel: VoiceChannelLike): Promise<WebRtcConnWrapper> {
let guildId: string | null = null;
if (
channel.type === "GUILD_STAGE_VOICE" ||
channel.type === "GUILD_VOICE"
) {
guildId = channel.guildId ?? null;
}
return this.joinVoice(guildId, channel.id);
}
/**
* Joins a voice channel and resolves with the WebRtcConnWrapper when the
* media session is ready.
*/
joinVoice(
guild_id: string | null,
channel_id: string,
): Promise<WebRtcConnWrapper> {
return new Promise((resolve, reject) => {
if (!this.client.user) {
reject(new Error("Client not logged in"));
return;
}
const user_id = this.client.user.id;
const voiceConn = new VoiceConnection(
this,
guild_id,
user_id,
channel_id,
(conn) => {
resolve(conn);
},
);
this._voiceConnection = voiceConn;
this._gatewayEmitter.on(
"VOICE_STATE_UPDATE",
(d: { user_id: string; session_id: string }) => {
if (user_id !== d.user_id) return;
voiceConn.setSession(d.session_id);
},
);
this._gatewayEmitter.on(
"VOICE_SERVER_UPDATE",
(d: {
guild_id: string | null;
channel_id?: string;
endpoint: string;
token: string;
}) => {
if (guild_id !== d.guild_id) return;
// channel_id is not set for guild voice calls
if (d.channel_id && channel_id !== d.channel_id) return;
voiceConn.setTokens(d.endpoint, d.token);
},
);
this.signalVideo(false);
});
}
/** Create a GoLive stream (screen share) on top of the voice connection. */
createStream(): Promise<WebRtcConnWrapper> {
return new Promise((resolve, reject) => {
if (!this.client.user) {
reject(new Error("Client not logged in"));
return;
}
if (!this.voiceConnection) {
reject(
new Error("cannot start stream without first joining voice channel"),
);
return;
}
this.signalStream();
const {
guildId: clientGuildId,
channelId: clientChannelId,
session_id,
} = this.voiceConnection;
const clientUserId = this.client.user.id;
if (!session_id) throw new Error("Session doesn't exist yet");
const streamConn = new StreamConnection(
this,
clientGuildId,
clientUserId,
clientChannelId,
(conn) => {
resolve(conn);
},
);
this.voiceConnection.streamConnection = streamConn;
this._gatewayEmitter.on(
"STREAM_CREATE",
(d: { stream_key: string; rtc_server_id: string }) => {
const { channelId, guildId, userId } = parseStreamKey(d.stream_key);
if (
clientGuildId !== guildId ||
clientChannelId !== channelId ||
clientUserId !== userId
) {
return;
}
streamConn.serverId = d.rtc_server_id;
streamConn.streamKey = d.stream_key;
streamConn.setSession(session_id);
},
);
this._gatewayEmitter.on(
"STREAM_SERVER_UPDATE",
(d: { stream_key: string; endpoint: string; token: string }) => {
const { channelId, guildId, userId } = parseStreamKey(d.stream_key);
if (
clientGuildId !== guildId ||
clientChannelId !== channelId ||
clientUserId !== userId
) {
return;
}
streamConn.setTokens(d.endpoint, d.token);
},
);
});
}
async setStreamPreview(image: Buffer): Promise<void> {
if (!this.client.token) throw new Error("Please login :)");
if (!this.voiceConnection?.streamConnection?.guildId) return;
const data = `data:image/jpeg;base64,${image.toString("base64")}`;
const { guildId } = this.voiceConnection.streamConnection;
if (!this.client.guilds) return;
const server = await this.client.guilds.fetch(guildId);
// biome-ignore lint/suspicious/noExplicitAny: discord.js-selfbot dynamic
(server as any).members.me?.voice?.postPreview(data);
}
stopStream(): void {
const stream = this.voiceConnection?.streamConnection;
if (!stream) return;
stream.stop();
this.signalStopStream();
this.voiceConnection.streamConnection = null;
this._gatewayEmitter.removeAllListeners("STREAM_CREATE");
this._gatewayEmitter.removeAllListeners("STREAM_SERVER_UPDATE");
}
leaveVoice(): void {
this.voiceConnection?.stop();
this.signalLeaveVoice();
this._voiceConnection = null;
this._gatewayEmitter.removeAllListeners("VOICE_STATE_UPDATE");
this._gatewayEmitter.removeAllListeners("VOICE_SERVER_UPDATE");
}
signalVideo(video_enabled: boolean): void {
if (!this.voiceConnection) return;
const { guildId: guild_id, channelId: channel_id } = this.voiceConnection;
this.sendOpcode(GatewayOpCodes.VOICE_STATE_UPDATE, {
guild_id: guild_id,
channel_id,
self_mute: false,
self_deaf: true,
self_video: video_enabled,
});
}
signalStream(): void {
if (!this.voiceConnection) return;
const {
type,
guildId: guild_id,
channelId: channel_id,
botId: user_id,
} = this.voiceConnection;
this.sendOpcode(GatewayOpCodes.STREAM_CREATE, {
type,
guild_id,
channel_id,
preferred_region: null,
});
this.sendOpcode(GatewayOpCodes.STREAM_SET_PAUSED, {
stream_key: generateStreamKey(type, guild_id, channel_id, user_id),
paused: false,
});
}
signalStopStream(): void {
if (!this.voiceConnection) return;
const {
type,
guildId: guild_id,
channelId: channel_id,
botId: user_id,
} = this.voiceConnection;
this.sendOpcode(GatewayOpCodes.STREAM_DELETE, {
stream_key: generateStreamKey(type, guild_id, channel_id, user_id),
});
}
signalLeaveVoice(): void {
this.sendOpcode(GatewayOpCodes.VOICE_STATE_UPDATE, {
guild_id: null,
channel_id: null,
self_mute: true,
self_deaf: false,
self_video: false,
});
}
}
export type { NativePeerConnection };
@@ -1,20 +0,0 @@
/**
* VideoStream feeds encoded H264 frames into the WebRTC connection.
* Ported from @dank074/discord-video-stream VideoStream.js.
*/
import { BaseMediaStream } from "./BaseMediaStream.js";
import type { WebRtcConnWrapper } from "./WebRtcWrapper.js";
export class VideoStream extends BaseMediaStream {
_conn: WebRtcConnWrapper;
constructor(conn: WebRtcConnWrapper, noSleep = false) {
super("video", noSleep);
this._conn = conn;
}
async _sendFrame(frame: Buffer, frametime: number): Promise<void> {
this._conn.sendVideoFrame(frame, frametime);
}
}
@@ -1,25 +0,0 @@
/**
* VoiceConnection guild/DM voice channel GoLive connection.
* Ported from @dank074/discord-video-stream VoiceConnection.js.
*/
import { BaseMediaConnection } from "./BaseMediaConnection.js";
import type { StreamConnection } from "./StreamConnection.js";
export class VoiceConnection extends BaseMediaConnection {
streamConnection: StreamConnection | null = null;
get daveChannelId(): string {
return this.channelId;
}
get serverId(): string | null {
// for guild vc it is the guild id, for dm voice it is the channel id
return this.guildId ?? this.channelId;
}
stop(): void {
super.stop();
this.streamConnection?.stop();
}
}
@@ -1,38 +0,0 @@
/** Discord voice WebSocket opcodes — ported from @dank074/discord-video-stream. */
export enum VoiceOpCodes {
IDENTIFY = 0,
SELECT_PROTOCOL = 1,
READY = 2,
HEARTBEAT = 3,
SELECT_PROTOCOL_ACK = 4,
SPEAKING = 5,
HEARTBEAT_ACK = 6,
RESUME = 7,
HELLO = 8,
RESUMED = 9,
CLIENTS_CONNECT = 11,
VIDEO = 12,
CLIENT_DISCONNECT = 13,
SESSION_UPDATE = 14,
MEDIA_SINK_WANTS = 15,
VOICE_BACKEND_VERSION = 16,
CHANNEL_OPTIONS_UPDATE = 17,
FLAGS = 18,
SPEED_TEST = 19,
PLATFORM = 20,
DAVE_PREPARE_TRANSITION = 21,
DAVE_EXECUTE_TRANSITION = 22,
DAVE_TRANSITION_READY = 23,
DAVE_PREPARE_EPOCH = 24,
MLS_INVALID_COMMIT_WELCOME = 31,
}
/** Binary voice WebSocket opcodes (DAVE / MLS). */
export enum VoiceOpCodesBinary {
MLS_EXTERNAL_SENDER = 25,
MLS_KEY_PACKAGE = 26,
MLS_PROPOSALS = 27,
MLS_COMMIT_WELCOME = 28,
MLS_ANNOUNCE_COMMIT_TRANSITION = 29,
MLS_WELCOME = 30,
}
@@ -1,205 +0,0 @@
/**
* WebRTC connection wrapper for GoLive ported from
* @dank074/discord-video-stream WebRtcWrapper.js, with the media stack
* (packetizers, RTCP SR/NACK, pacing) provided by the libdatachannel-min
* binding instead of node-datachannel's JS-exposed media classes.
*/
import {
H264Helpers,
H264NalUnitTypes,
splitNalu,
startCode3,
} from "./AnnexBHelper.js";
import { CodecPayloadType } from "./CodecPayloadType.js";
import type { NativePeerConnection, NativeTrack } from "./native.js";
import { loadNative } from "./native.js";
import { rewriteSPSVUI } from "./SPSVUIRewriter.js";
import { normalizeVideoCodec } from "./utils.js";
export type WebRtcVideoCodec = "H264" | "H265" | "VP8" | "VP9" | "AV1";
export interface WebRtcParams {
address: string;
port: number;
audioSsrc: number;
videoSsrc: number;
rtxSsrc: number;
supportedEncryptionModes: string[];
}
/** Minimal surface of the media connection that WebRtcWrapper drives. */
export interface VideoAttribute {
fps: number;
width: number;
height: number;
}
export interface MediaConnectionLike {
daveReady: boolean;
daveSession: {
encryptOpus(frame: Buffer): Buffer;
encrypt(mediaType: number, codec: number, frame: Buffer): Buffer;
} | null;
webRtcParams: WebRtcParams | null;
setSpeaking(speaking: boolean): void;
setVideoAttributes(enabled: boolean, attr?: VideoAttribute): void;
}
/** Media types used by DAVE encryption (from @dank074). */
export enum DaveMediaType {
AUDIO = 0,
VIDEO = 1,
}
/** DAVE codec ids (from @dank074). */
export enum DaveCodec {
UNKNOWN = 0,
VP8 = 2,
VP9 = 3,
H264 = 4,
H265 = 5,
AV1 = 6,
}
export class WebRtcConnWrapper {
private _mediaConn: MediaConnectionLike;
private _webRtcConn: NativePeerConnection | null = null;
private _audioTrack: NativeTrack | null = null;
private _videoTrack: NativeTrack | null = null;
private _videoCodec: WebRtcVideoCodec | null = null;
/** Assigned by BaseMediaConnection to send the gathered SDP to Discord. */
onLocalDescription: ((sdp: string) => void) | null = null;
constructor(mediaConn: MediaConnectionLike) {
this._mediaConn = mediaConn;
}
initWebRtc(): NativePeerConnection {
const native = loadNative();
this._webRtcConn = new native.PeerConnection({
iceServers: ["stun:stun.l.google.com:19302"],
});
// Track mids must match @dank074: "0" audio, "1" video.
this._audioTrack = this._webRtcConn.addTrack("0", "audio");
this._videoTrack = this._webRtcConn.addTrack("1", "video");
return this._webRtcConn;
}
close(): void {
this._webRtcConn?.close();
this._webRtcConn = null;
}
get webRtcConn(): NativePeerConnection | null {
return this._webRtcConn;
}
get ready(): boolean {
return this._webRtcConn?.state() === "connected";
}
get mediaConnection(): MediaConnectionLike {
return this._mediaConn;
}
sendAudioFrame(frame: Buffer, frametime: number): void {
if (!this.ready || !this._audioTrack) return;
const clockRate = CodecPayloadType.opus.clockRate;
if (this.mediaConnection.daveReady && this.mediaConnection.daveSession) {
frame = this.mediaConnection.daveSession.encryptOpus(frame);
}
this._audioTrack.sendFrame(frame);
this._audioTrack.addTimestamp(Math.round((frametime * clockRate) / 1000));
}
sendVideoFrame(frame: Buffer, frametime: number): void {
if (!this.ready || !this._videoTrack) return;
const clockRate = CodecPayloadType[this._videoCodec ?? "H264"].clockRate;
if (this._videoCodec === "H264") {
let spsRewritten = false;
const nalus = splitNalu(frame).map((el) => {
if (H264Helpers.getUnitType(el) === H264NalUnitTypes.SPS) {
spsRewritten = true;
return rewriteSPSVUI(el);
}
return el;
});
if (spsRewritten)
frame = Buffer.concat(nalus.flatMap((el) => [startCode3, el]));
}
if (this.mediaConnection.daveReady && this.mediaConnection.daveSession) {
let daveCodec = DaveCodec.UNKNOWN;
switch (this._videoCodec) {
case "H264":
daveCodec = DaveCodec.H264;
break;
case "H265":
daveCodec = DaveCodec.H265;
break;
case "VP8":
daveCodec = DaveCodec.VP8;
break;
case "VP9":
daveCodec = DaveCodec.VP9;
break;
case "AV1":
daveCodec = DaveCodec.AV1;
break;
default:
break;
}
frame = this.mediaConnection.daveSession.encrypt(
DaveMediaType.VIDEO,
daveCodec,
frame,
);
}
this._videoTrack.sendFrame(frame);
this._videoTrack.addTimestamp(Math.round((frametime * clockRate) / 1000));
}
setPacketizer(videoCodec: string): void {
if (!this.mediaConnection.webRtcParams) {
throw new Error("WebRTC connection not ready");
}
const { audioSsrc, videoSsrc } = this.mediaConnection.webRtcParams;
this._videoCodec = normalizeVideoCodec(videoCodec);
// Audio packetizer: opus 120 @ 48kHz, playout delay ext id 5 (like @dank074)
this._audioTrack?.setPacketizer(
"audio",
audioSsrc,
CodecPayloadType.opus.payload_type,
CodecPayloadType.opus.clockRate,
5,
0,
1,
);
// Video packetizer: H264/H265/AV1 with their payload types
const codecEntry = CodecPayloadType[this._videoCodec];
if (!codecEntry) {
throw new Error(`Packetizer not implemented for ${this._videoCodec}`);
}
const nativeKind =
this._videoCodec === "H264"
? "h264"
: this._videoCodec === "H265"
? "h265"
: this._videoCodec === "AV1"
? "av1"
: (() => {
throw new Error(
`Packetizer not implemented for ${this._videoCodec}`,
);
})();
this._videoTrack?.setPacketizer(
nativeKind,
videoSsrc,
codecEntry.payload_type,
codecEntry.clockRate,
5,
0,
10,
);
}
}
@@ -1,19 +0,0 @@
/**
* goLive public API re-exports the ported @dank074 modules.
* Drop-in replacement for `@dank074/discord-video-stream` in
* screenShareController.ts.
*/
export { AudioStream } from "./AudioStream.js";
export { BaseMediaConnection } from "./BaseMediaConnection.js";
export { BaseMediaStream } from "./BaseMediaStream.js";
export { CodecPayloadType } from "./CodecPayloadType.js";
export { demux } from "./Demuxer.js";
export { Encoders } from "./Encoders.js";
export { playStream, prepareStream } from "./prepareStream.js";
export { StreamConnection } from "./StreamConnection.js";
export { Streamer } from "./Streamer.js";
export { normalizeVideoCodec } from "./utils.js";
export { VideoStream } from "./VideoStream.js";
export { VoiceConnection } from "./VoiceConnection.js";
export { WebRtcConnWrapper } from "./WebRtcWrapper.js";
@@ -1,116 +0,0 @@
/**
* Loader + typings for the minimal libdatachannel N-API binding
* (native/libdatachannel-min). The binding exposes ONLY what GoLive needs:
* PeerConnection, DataChannel, Track (raw RTP + media packetizer chain).
*
* The .node file is built by node-gyp against libdatachannel 0.24.0 (built
* from source nixpkgs 0.24.1 is glibc-incompatible with this host). It is
* NOT shipped via npm; the Nix flake builds it as part of the gateway.
*/
export interface NativeTrack {
/** Send a RAW RTP/RTCP packet (no media handler installed). */
send(buffer: Uint8Array): void;
/** Send an ENCODED frame; the packetizer chain turns it into RTP. */
sendFrame(buffer: Uint8Array): void;
/** Advance the packetizer RTP timestamp by delta (clock-rate units). */
addTimestamp(delta: number): void;
/** Install the media-handler chain (packetizer → RTCP SR → NACK → pacing). */
setPacketizer(
kind: "audio" | "h264" | "h265" | "av1",
ssrc: number,
payloadType: number,
clockRate: number,
playoutDelayId: number,
playoutDelayMin: number,
playoutDelayMax: number,
): void;
isOpen(): boolean;
close(): void;
}
export interface NativePeerConnection {
/** mid must be "0" (audio) or "1" (video) — matches @dank074's track defs. */
addTrack(mid: string, kind: "audio" | "video"): NativeTrack;
/** Resolves with the full SDP (incl. candidates) after gathering completes. */
createOffer(): Promise<string>;
/** Resolves with the auto-generated answer SDP. */
createAnswer(offerSdp: string): Promise<string>;
setRemoteDescription(sdp: string, type: "offer" | "answer"): void;
state(): string;
close(): void;
onStateChange(cb: (state: string) => void): void;
}
export interface NativeBinding {
PeerConnection: new (config: {
iceServers: string[];
}) => NativePeerConnection;
DataChannel: unknown;
Track: unknown;
}
let cached: NativeBinding | null = null;
/** Load the native binding. Throws only if the .node is truly missing
* callers (screen share) guard with `isNativeAvailable()`. */
export function loadNative(): NativeBinding {
if (cached) return cached;
// Resolve relative to this file: src/goLive/ → native/libdatachannel-min/
const candidates = [
new URL(
"../../native/libdatachannel-min/build/Release/datachannel_min.node",
import.meta.url,
),
new URL(
"../../../native/libdatachannel-min/build/Release/datachannel_min.node",
import.meta.url,
),
];
let lastErr: unknown;
for (const url of candidates) {
try {
// @ts-expect-error — .node modules are not typed; dynamic require via file URL
const mod = process.dlopen ? null : null;
void mod;
const nativePath = url.pathname;
// eslint-disable-next-line @typescript-eslint/no-require-imports
const req = createRequire(import.meta.url);
const binding = req(nativePath) as NativeBinding;
if (typeof binding.PeerConnection === "function") {
cached = binding;
return binding;
}
} catch (e) {
lastErr = e;
}
}
// Fallback: plain relative require (tsx / jest environments)
try {
const req = createRequire(import.meta.url);
const binding = req(
"../../native/libdatachannel-min/build/Release/datachannel_min.node",
) as NativeBinding;
if (typeof binding.PeerConnection === "function") {
cached = binding;
return binding;
}
} catch (e) {
lastErr = e;
}
throw new Error(
`libdatachannel-min native binding not built (${String(lastErr)}). Run: cd native/libdatachannel-min && npx node-gyp rebuild`,
);
}
import { createRequire } from "node:module";
/** True when the native binding is built — screen share stays disabled otherwise. */
export function isNativeAvailable(): boolean {
try {
loadNative();
return true;
} catch {
return false;
}
}
@@ -1,325 +0,0 @@
/**
* prepareStream & playStream ported from @dank074/discord-video-stream
* newApi.js (Encoders/prepareStream/playStream), but uses `child_process.spawn`
* + ffmpeg CLI args directly instead of fluent-ffmpeg + node-av.
*
* Replaces the @dank074 video pipeline entirely:
* input (URL or Readable) ffmpeg spawn H264 AnnexB frames
* Demuxer stream VideoStream/AudioStream WebRtcConnWrapper
*/
import { type ChildProcess, spawn } from "node:child_process";
import { existsSync, readdirSync } from "node:fs";
import { join } from "node:path";
import { PassThrough, type Readable } from "node:stream";
import { demux } from "./Demuxer.js";
import { type EncoderSettings, Encoders } from "./Encoders.js";
import { VideoStream } from "./VideoStream.js";
import type { WebRtcConnWrapper } from "./WebRtcWrapper.js";
export interface PrepareStreamResult {
command: ChildProcess;
output: PassThrough;
encoder: () => Record<string, EncoderSettings>;
options: Record<string, unknown>;
videoCodec: string;
width: number;
height: number;
frameRate?: number;
includeAudio: boolean;
}
function isFiniteNonZero(n: unknown): n is number {
return typeof n === "number" && !!n && Number.isFinite(n);
}
const DEFAULT_HEADERS = {
"User-Agent":
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36",
Connection: "keep-alive",
};
/** Resolve ffmpeg binary (env override → PATH → Nix store ffmpeg-headless). */
function resolveFfmpeg(): string {
if (process.env.FFMPEG_PATH && existsSync(process.env.FFMPEG_PATH)) {
return process.env.FFMPEG_PATH;
}
const store = "/nix/store";
if (existsSync(store)) {
const entries = readdirSync(store);
for (const entry of entries) {
if (!entry.includes("ffmpeg-headless-")) continue;
const candidate = join(store, entry, "bin", "ffmpeg");
if (existsSync(candidate)) return candidate;
}
}
return "ffmpeg";
}
const FFMPEG_BIN = resolveFfmpeg();
/**
* prepareStream build an ffmpeg command (as spawn args + PassThrough output)
* that transcodes the input into a pipe we can demux. Mirrors @dank074's
* prepareStream but produces a raw spawn instead of a fluent-ffmpeg command.
*/
export function prepareStream(
input: string | Readable,
options: Record<string, unknown> = {},
): PrepareStreamResult {
const mergedOptions = {
noTranscoding: false,
width: isFiniteNonZero(options.width)
? Math.round(options.width as number)
: -2,
height: isFiniteNonZero(options.height)
? Math.round(options.height as number)
: -2,
frameRate:
isFiniteNonZero(options.frameRate) && (options.frameRate as number) > 0
? options.frameRate
: undefined,
videoCodec: (options.videoCodec as string) ?? "H264",
bitrateVideo:
isFiniteNonZero(options.bitrateVideo) &&
(options.bitrateVideo as number) > 0
? Math.round(options.bitrateVideo as number)
: 5000,
bitrateVideoMax:
isFiniteNonZero(options.bitrateVideoMax) &&
(options.bitrateVideoMax as number) > 0
? Math.round(options.bitrateVideoMax as number)
: 7000,
bitrateAudio:
isFiniteNonZero(options.bitrateAudio) &&
(options.bitrateAudio as number) > 0
? Math.round(options.bitrateAudio as number)
: 128,
includeAudio: options.includeAudio ?? true,
encoder:
(options.encoder as () => Record<string, EncoderSettings>) ??
Encoders.software(),
customHeaders: {
...DEFAULT_HEADERS,
...(options.customHeaders as Record<string, string> | undefined),
},
customInputOptions: (options.customInputOptions as string[]) ?? [],
customFfmpegFlags: (options.customFfmpegFlags as string[]) ?? [],
minimizeLatency: options.minimizeLatency ?? false,
};
const output = new PassThrough();
const args: string[] = [
"-hide_banner",
"-loglevel",
"error",
...(typeof input === "string" ? ["-i", input] : ["-i", "pipe:0"]),
...mergedOptions.customInputOptions,
];
if (mergedOptions.minimizeLatency) {
args.push("-fflags", "nobuffer", "-analyzeduration", "0");
}
if (typeof input === "string" && input.startsWith("http")) {
const headerStr = Object.entries(mergedOptions.customHeaders)
.map(([k, v]) => `${k}: ${v}`)
.join("\r\n");
args.push(
"-headers",
headerStr,
"-reconnect",
"1",
"-reconnect_at_eof",
"1",
"-reconnect_streamed",
"1",
"-reconnect_delay_max",
"4294",
);
}
// Video
args.push("-map", "0:v:0");
if (mergedOptions.noTranscoding) {
args.push("-c:v", "copy");
} else {
args.push(`-vf`, `scale=${mergedOptions.width}:${mergedOptions.height}`);
if (mergedOptions.frameRate)
args.push("-r", String(mergedOptions.frameRate));
const enc = mergedOptions.encoder()[mergedOptions.videoCodec];
if (!enc)
throw new Error(
`Encoder settings not specified for ${mergedOptions.videoCodec}`,
);
// Encoder options are declared as single strings like "-forced-idr 1";
// spawn needs each flag and value as separate argv entries.
const encOptions = enc.options.flatMap((opt) =>
opt.split(/\s+/).filter(Boolean),
);
args.push(
"-b:v",
`${mergedOptions.bitrateVideo}k`,
"-maxrate:v",
`${mergedOptions.bitrateVideoMax}k`,
"-bufsize:v",
`${Math.round(mergedOptions.bitrateVideo / 2)}k`,
"-bf",
"0",
"-pix_fmt",
"yuv420p",
"-force_key_frames",
"expr:gte(t,n_forced*1)",
"-c:v",
enc.name,
...encOptions,
...(enc.globalOptions ?? []).flatMap((opt) =>
opt.split(/\s+/).filter(Boolean),
),
);
}
// Audio
if (mergedOptions.includeAudio) {
args.push("-map", "0:a:0?");
args.push(
"-c:a",
"libopus",
"-b:a",
`${mergedOptions.bitrateAudio}k`,
"-ar",
"48000",
"-ac",
"2",
);
} else {
args.push("-an");
}
args.push(...mergedOptions.customFfmpegFlags);
args.push("-f", "h264", "pipe:1");
const isUrl = typeof input === "string";
const proc: ChildProcess = isUrl
? spawn(FFMPEG_BIN, args, { stdio: ["ignore", "pipe", "pipe"] })
: spawn(FFMPEG_BIN, args, { stdio: ["pipe", "pipe", "pipe"] });
if (proc.stdin && !isUrl) {
input.on("data", (chunk: Buffer) => proc.stdin?.write(chunk));
input.on("end", () => proc.stdin?.end());
input.on("error", () => proc.stdin?.destroy());
}
proc.stdout?.pipe(output);
proc.stderr?.on("data", () => {
/* swallow ffmpeg stderr */
});
proc.on("error", (err) => {
// spawn failed (e.g. ffmpeg missing). If someone is consuming output
// (demux attaches an 'error' listener) propagate; otherwise just end.
if (output.listenerCount("error") > 0) {
output.destroy(err);
} else {
output.end();
}
});
proc.on("close", () => {
output.end();
});
return {
command: proc,
output,
encoder: mergedOptions.encoder,
options: mergedOptions,
videoCodec: mergedOptions.videoCodec,
width: mergedOptions.width,
height: mergedOptions.height,
frameRate: mergedOptions.frameRate,
includeAudio: !!mergedOptions.includeAudio,
};
}
export interface PlayStreamOptions {
type?: "go-live" | "video";
format?: string;
width?: number | ((v: unknown) => number);
height?: number | ((v: unknown) => number);
frameRate?: number | ((v: unknown) => number);
readrateInitialBurst?: number;
streamPreview?: boolean;
}
/**
* playStream demux the prepareStream output and pipe frames into the
* WebRTC connection's video/audio streams. Resolves when the video stream
* ends (natural EOF or the ffmpeg command is killed via cleanup/stop).
*/
export async function playStream(
prepared: PrepareStreamResult,
streamer: { createStream: () => Promise<WebRtcConnWrapper> },
options: PlayStreamOptions = {},
): Promise<void> {
const conn = await streamer.createStream();
const { video, close: demuxClose } = await demux(prepared.output, {
format: options.format ?? "nut",
});
if (!video) throw new Error("No video stream in media");
conn.setPacketizer(video.codecName);
conn.mediaConnection.setSpeaking(true);
const w =
typeof options.width === "function"
? options.width(video)
: (options.width ?? video.width);
const h =
typeof options.height === "function"
? options.height(video)
: (options.height ?? video.height);
const fr =
typeof options.frameRate === "function"
? options.frameRate(video)
: (options.frameRate ??
(video.framerate_num / video.framerate_den || 30));
conn.mediaConnection.setVideoAttributes(true, {
width: Math.round(w),
height: Math.round(h),
fps: Math.round(fr),
});
const vStream = new VideoStream(conn);
video.stream.pipe(vStream);
const cleanup = () => {
try {
prepared.command.kill("SIGTERM");
} catch {
/* already dead */
}
demuxClose();
try {
conn.mediaConnection.setSpeaking(false);
conn.mediaConnection.setVideoAttributes(false);
} catch {
/* connection already torn down */
}
};
return new Promise<void>((resolve) => {
vStream.once("finish", () => {
cleanup();
resolve();
});
vStream.once("error", () => {
cleanup();
resolve();
});
});
}
export { Encoders };
@@ -1,82 +0,0 @@
/** GoLive helpers — ported from @dank074/discord-video-stream/utils.js. */
export function normalizeVideoCodec(
codec: string,
): "H264" | "H265" | "VP8" | "VP9" | "AV1" {
if (/H\.?264|AVC/i.test(codec)) return "H264";
if (/H\.?265|HEVC/i.test(codec)) return "H265";
if (/VP(8|9)/i.test(codec)) return codec.toUpperCase() as "VP8" | "VP9";
if (/AV1/i.test(codec)) return "AV1";
throw new Error(`Unknown codec: ${codec}`);
}
/**
* The available video streams are sent by the client on connection to the
* voice gateway using OpCode Identify (0); the server replies with the ssrc
* and rtxssrc for each available stream using OpCode Ready (2). RID
* distinguishes simulcast streams of the same video source we only send one
* quality stream, so a single entry is hardcoded.
*/
export const STREAMS_SIMULCAST = [{ type: "screen", rid: "100", quality: 100 }];
export const max_int16bit = 2 ** 16;
export const max_int32bit = 2 ** 32;
export function isFiniteNonZero(n: unknown): n is number {
return typeof n === "number" && !!n && Number.isFinite(n);
}
export interface ParsedStreamKey {
type: "guild" | "call";
channelId: string;
guildId: string | null;
userId: string;
}
export function parseStreamKey(streamKey: string): ParsedStreamKey {
const streamKeyArray = streamKey.split(":");
const type = streamKeyArray.shift();
if (type !== "guild" && type !== "call") {
throw new Error(`Invalid stream key type: ${type}`);
}
if (
(type === "guild" && streamKeyArray.length < 3) ||
(type === "call" && streamKey.length < 2)
) {
throw new Error(`Invalid stream key: ${streamKey}`);
}
let guildId: string | null = null;
if (type === "guild") {
guildId = streamKeyArray.shift() ?? null;
}
const channelId = streamKeyArray.shift();
const userId = streamKeyArray.shift();
if (!channelId || !userId) {
throw new Error(`Invalid stream key: ${streamKey}`);
}
return { type, channelId, guildId, userId };
}
export function generateStreamKey(
type: "guild" | "call",
guildId: string | null,
channelId: string,
userId: string,
): string {
return `${type}${type === "guild" ? `:${guildId}` : ""}:${channelId}:${userId}`;
}
export interface VoiceChannelLike {
type: string;
id: string;
guildId?: string | null;
}
export function isVoiceChannel(channel: VoiceChannelLike): boolean {
return (
channel.type === "DM" ||
channel.type === "GROUP_DM" ||
channel.type === "GUILD_STAGE_VOICE" ||
channel.type === "GUILD_VOICE"
);
}
@@ -316,11 +316,13 @@ async function processBatch(job: {
// The orchestrator handles text/media split + caching + parallel paths
// internally, so a 20-message batch = 1 text LLM call (+1 media call
// when media is present), not N per-message calls.
const analysisStart = Date.now();
const moderationResult = await runModerationAnalysis({
targets: readyMessages,
contextBlock,
attachments,
});
const analysisDurationMs = Date.now() - analysisStart;
const results = moderationResult.results.map((r) =>
normalizeResult(
@@ -342,6 +344,7 @@ async function processBatch(job: {
confidence: result.confidence,
recommendedAction: result.recommendedAction,
analyzedAt: Date.now(),
analysisDurationMs,
error: result.status === "error" ? result.analysis : null,
},
}));
@@ -136,9 +136,11 @@ export function startPendingAIAnalysisWorker(
import("./cultureLearner.js")
.then((m) => m.startCultureLearnerWorker())
.catch(console.error);
import("./userProfileLearner.js")
.then((m) => m.startUserProfileLearnerWorker())
.catch(console.error);
if (config.AI_USER_PROFILE_LEARNING_ENABLED) {
import("./userProfileLearner.js")
.then((m) => m.startUserProfileLearnerWorker())
.catch(console.error);
}
setInterval(() => {
// [D] Periodic cache hygiene: purge expired moderation verdicts from
@@ -18,7 +18,20 @@ const log = createChildLogger("llm-client");
// Concurrency limiter for LLM API calls (inlined from concurrencyLimiter.ts)
// ---------------------------------------------------------------------------
const llmSemaphore = pLimit(config.AI_LLM_MAX_CONCURRENT ?? 5);
// The limiter is cached per configured concurrency value so it can be tuned
// (env / BWS) without a code change and always reflects the current config —
// a module-level `pLimit(config.X)` would freeze the cap at import time.
let llmSemaphore = pLimit(config.AI_LLM_MAX_CONCURRENT ?? 5);
let llmSemaphoreLimit = config.AI_LLM_MAX_CONCURRENT ?? 5;
function getLlmSemaphore() {
const wanted = config.AI_LLM_MAX_CONCURRENT ?? 5;
if (wanted !== llmSemaphoreLimit) {
llmSemaphore = pLimit(wanted);
llmSemaphoreLimit = wanted;
}
return llmSemaphore;
}
let activeCount = 0;
let pendingCount = 0;
@@ -30,7 +43,7 @@ export async function withLlmConcurrency<T>(fn: () => Promise<T>): Promise<T> {
"Queuing LLM request",
);
return llmSemaphore(async () => {
return getLlmSemaphore()(async () => {
pendingCount--;
activeCount++;
@@ -147,12 +160,78 @@ export interface LlmCallOpts {
top_p?: number;
/** Force JSON output via response_format: { type: "json_object" }. */
jsonResponse?: { type: "json_object" };
/**
* Disable LLM chain-of-thought (reasoning/thinking) for faster analysis.
* Defaults to config.AI_LLM_DISABLE_THINKING when omitted.
*/
disableThinking?: boolean;
/** Extra retries beyond DEFAULT_RETRIES (default 2). */
retries?: number;
/** Whether to use streaming (if true, will consume stream and return aggregated result) */
stream?: boolean;
/** Optional AbortSignal to cancel the API request */
signal?: AbortSignal;
/**
* Per-request timeout in ms. Falls back to the client-level default
* (60s) when omitted. Vision/image analysis passes a longer budget here
* so a single large-image call isn't killed early by the shared default.
*/
timeout?: number;
}
/**
* Build the request params for an LLM chat completion. Pulled out of `llmChat`
* so the thinking-disable injection can be unit-tested without network access.
*
* Optional params (temperature/top_p/max_tokens) are only attached when
* explicitly provided, to maximise compatibility with various providers/local
* APIs. When `disableThinking` is set, we inject the common provider params
* used to switch OFF chain-of-thought reasoning. OpenAI-compatible routers
* ignore the variants their backend does not understand, so sending the
* OpenAI (`reasoning_effort`), OpenRouter (`reasoning.enabled`) and
* vLLM/Qwen/litellm (`chat_template_kwargs.enable_thinking`) forms together
* covers the popular reasoning backends behind a proxy.
*/
export function buildLlmParams(
opts: LlmCallOpts,
disableThinking: boolean,
): OpenAI.Chat.Completions.ChatCompletionCreateParams {
const {
messages,
model = config.AI_LLM_MODEL,
max_tokens,
temperature,
top_p,
jsonResponse,
stream,
} = opts;
const params = {
model,
messages,
...(stream !== undefined ? { stream } : {}),
} as OpenAI.Chat.Completions.ChatCompletionCreateParams;
if (temperature !== undefined) params.temperature = temperature;
if (top_p !== undefined) params.top_p = top_p;
if (max_tokens !== undefined) params.max_tokens = max_tokens;
if (jsonResponse) params.response_format = jsonResponse;
if (disableThinking) {
Object.assign(params, {
// OpenAI o-series
reasoning_effort: "none",
// OpenRouter
reasoning: { enabled: false },
// vLLM / Qwen / litellm
chat_template_kwargs: { enable_thinking: false },
// Anthropic / Claude-format (9router exposes thinkingFormat
// "claude-adaptive" / "claude-budget" on its reasoning models)
thinking: { type: "disabled" },
} as Record<string, unknown>);
}
return params;
}
/**
@@ -167,33 +246,12 @@ export async function llmChat(
const client = getClient();
if (!client) return null;
const {
messages,
model = config.AI_LLM_MODEL,
max_tokens,
temperature,
top_p,
jsonResponse,
retries = DEFAULT_RETRIES,
stream,
signal,
} = opts;
const { retries = DEFAULT_RETRIES, signal } = opts;
const disableThinking =
opts.disableThinking ?? config.AI_LLM_DISABLE_THINKING;
const params = {
model,
messages,
...(stream !== undefined ? { stream } : {}),
} as OpenAI.Chat.Completions.ChatCompletionCreateParams;
// Attach optional parameters only if explicitly provided to maintain
// maximum compatibility with various LLM providers and local APIs.
if (temperature !== undefined) params.temperature = temperature;
if (top_p !== undefined) params.top_p = top_p;
if (max_tokens !== undefined) params.max_tokens = max_tokens;
if (jsonResponse) {
params.response_format = jsonResponse;
}
const params = buildLlmParams(opts, disableThinking);
const model = params.model;
return retryWithBackoff(
async () => {
@@ -203,6 +261,7 @@ export async function llmChat(
) => {
const response = await client.chat.completions.create(currentParams, {
signal,
...(opts.timeout ? { timeout: opts.timeout } : {}),
});
if (currentParams.stream) {
let content = "";
@@ -283,6 +342,12 @@ export async function llmChat(
* Convenience for vision (image/sticker/emoji) analysis.
* Returns the raw completion content (trimmed) or null.
*
* Vision routes through the SAME router/base URL as text moderation
* (AI_LLM_BASE_URL) the dedicated NVIDIA multimodal endpoint was removed.
* It uses AI_LLM_VISION_MODEL (a different model alias from the text combo)
* so image analysis stays on a vision-capable model. Thinking-disable from
* config.AI_LLM_DISABLE_THINKING applies automatically via buildLlmParams.
*
* NOTE: retries are disabled here on purpose visionAnalyzer.ts already
* wraps this call in its own 3-attempt loop with exponential backoff.
* A second retry layer would multiply worst-case API calls (3×3=9/image).
@@ -307,6 +372,7 @@ export async function llmVision(
top_p: 0.9,
retries: 0,
stream: true, // router always streams SSE; non-stream waits for full body and times out
timeout: config.AI_LLM_VISION_ANALYSIS_TIMEOUT_MS ?? 60_000,
});
if (!completion) return null;
@@ -6,7 +6,6 @@
*/
export {
acquireMediaAnalysisLock,
computeImagePhash,
deleteCachedMediaAnalysis,
getCachedMediaAnalysis,
setCachedMediaAnalysis,
@@ -16,10 +16,8 @@ import { getChannelCulture } from "./channelCultureStore.js";
import type { RetryState } from "./llmCaller.js";
import { callModerationLLM } from "./llmCaller.js";
import { prepareMediaMessage } from "./mediaAnalysisClient.js";
import { buildUserProfilesBlock } from "./moderationBuilders.js";
import { buildSystemPrompt as buildSystemPromptModular } from "./moderationPrompt.js";
import { buildCorrectedFewShotExamples } from "./textBatchProcessor.js";
import { getUserProfile } from "./userProfileStore.js";
const log = createChildLogger("mediaBatchProcessor");
@@ -65,32 +63,14 @@ export async function runMediaBatch(
channelCulture,
});
// Gather user profiles ONCE for the whole batch and emit a deduplicated
// <user_profiles> map (with last-generated timestamp); per-message blocks
// (from prepareMediaMessage) reference it via <user_profile_ref>.
const profileByUser = new Map<
string,
{
text: string;
asOf?: number | null;
}
>();
for (const t of targets) {
if (profileByUser.has(t.user_id)) continue;
const profile = await getUserProfile(t.user_id);
profileByUser.set(t.user_id, {
text: profile?.profile_summary ?? "",
asOf: profile?.last_analyzed_at ?? null,
});
}
const userProfilesBlock = buildUserProfilesBlock(profileByUser);
// Per-message blocks (from prepareMediaMessage) contain the message
// content + reference/reply context only — no per-user reputation or
// profile context is injected (kept minimal per user request).
const messagesBlock = prepared.map((p) => p.messageBlock).join("\n");
// Data/instruction separation: the system prompt is stable per mode — all
// per-batch context (profiles, conversation) lives in the USER payload,
// ordered oldest-first so targets come last.
// per-batch context (conversation) lives in the USER payload, ordered
// oldest-first so targets come last.
const userBlocks = [
userProfilesBlock?.trimEnd() ?? "",
contextBlock?.trimEnd() ?? "",
`<messages_to_analyze>\n${messagesBlock}\n</messages_to_analyze>`,
].filter((b) => b.trim().length > 0);
@@ -7,28 +7,22 @@
import { LRUCache } from "lru-cache";
import {
acquireMediaAnalysisLock,
computeImagePhash,
deleteCachedMediaAnalysis,
getCachedMediaAnalysis,
getCachedMediaByPhash,
makeCustomEmojiCacheKey,
makeImageCacheKey,
makeStickerCacheKey,
upsertCachedMediaAnalysis,
upsertCachedMediaByPhash,
} from "./textCacheStore.js";
export {
acquireMediaAnalysisLock,
computeImagePhash,
deleteCachedMediaAnalysis,
getCachedMediaAnalysis,
getCachedMediaByPhash,
makeCustomEmojiCacheKey,
makeImageCacheKey,
makeStickerCacheKey,
upsertCachedMediaAnalysis,
upsertCachedMediaByPhash,
};
/** Convenience alias for upsertCachedMediaAnalysis. */
@@ -450,7 +450,7 @@ export async function downloadMediaCandidate(
if (candidate.customEmojiId || candidate.stickerName) {
const vck = candidate.customEmojiId
? makeCustomEmojiCacheKey(candidate.customEmojiId)
: makeStickerCacheKey(candidate.stickerName!);
: makeStickerCacheKey(candidate.stickerName ?? "");
const cached = await getCachedMediaAnalysis(vck);
if (cached) {
const existing = mediaAnalysisMap.get(targetId) ?? [];
@@ -194,7 +194,7 @@ export async function runModerationAnalysis(
if (embeddings && embeddings.length === texts.length) {
// index-aligned with semanticCandidates
for (let i = 0; i < semanticCandidates.length; i++) {
const { target, cacheKey } = semanticCandidates[i];
const { cacheKey } = semanticCandidates[i];
embeddingsByKey.set(cacheKey, embeddings[i]);
}
@@ -47,7 +47,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
title: "Pesan bersih dengan slang",
input: "[target] id=12345 user=budi: anjay wkwk gaskeun santuy bro",
output:
'{"results":[{"message_id":"12345","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Slang Indonesia umum, tanpa pelanggaran."}]}',
'{"results":[{"message_id":"12345","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Slang Indonesia umum, tanpa pelanggaran."}]}',
modes: ["text", "mixed"],
},
{
@@ -56,7 +56,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=67890 user=anon: lu goblok banget sih kontol, mampus aja lo",
output:
'{"results":[{"message_id":"67890","status":"flagged","flags":["harassment","vulgar_language"],"score":0.85,"severity":"high","confidence":0.9,"recommended_action":"delete","evidence":["lu goblok banget sih kontol"],"analysis":"Insult langsung dengan kata kasar terarah ke individu."}]}',
'{"results":[{"message_id":"67890","status":"flagged","flags":["harassment","vulgar_language"],"severity":"high","evidence":["lu goblok banget sih kontol"],"analysis":"Insult langsung dengan kata kasar terarah ke individu."}]}',
modes: ["text", "mixed"],
},
{
@@ -64,7 +64,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
title: "Emoji Huruf (Evasion)",
input: "[target] id=16161 user=sneaky: gsap expo 🇬 🇦 🇾",
output:
'{"results":[{"message_id":"16161","status":"flagged","flags":["sexual_deviation"],"score":0.8,"severity":"medium","confidence":0.95,"recommended_action":"delete","evidence":["🇬 🇦 🇾"],"analysis":"Regional indicator mengeja kata terlarang — evasi untuk topik yang dibatasi server."}]}',
'{"results":[{"message_id":"16161","status":"flagged","flags":["sexual_deviation"],"severity":"medium","evidence":["🇬 🇦 🇾"],"analysis":"Regional indicator mengeja kata terlarang — evasi untuk topik yang dibatasi server."}]}',
modes: ["text", "mixed"],
},
{
@@ -72,7 +72,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
title: "Typo QWERTY Programming (False Positive Prevention)",
input: "[target] id=17171 user=dian432: Apakah bisa ngodonf disitu?",
output:
'{"results":[{"message_id":"17171","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Kata \'ngodonf\' typo QWERTY natural (f-g, o-i) dari \'ngoding\', bukan obfuscation. Diskusi teknis wajar."}]}',
'{"results":[{"message_id":"17171","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Kata \'ngodonf\' typo QWERTY natural (f-g, o-i) dari \'ngoding\', bukan obfuscation. Diskusi teknis wajar."}]}',
modes: ["text", "mixed"],
},
{
@@ -81,7 +81,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=17172 user=dev: TypeError: Cannot read properties of undefined (reading 'map'). Kalo ada yang tau solusinya tolong bantu",
output:
'{"results":[{"message_id":"17172","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Diskusi teknis normal tentang error programming, bukan pelanggaran."}]}',
'{"results":[{"message_id":"17172","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Diskusi teknis normal tentang error programming, bukan pelanggaran."}]}',
modes: ["text", "mixed"],
},
{
@@ -90,19 +90,9 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=17173 user=dev: Guys pake Cursor lebih enak daripada VSCode? Gw pake Claude buat bantuin debugging",
output:
'{"results":[{"message_id":"17173","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Diskusi tools programming biasa (Cursor, VSCode, Claude), tidak ada pelanggaran."}]}',
'{"results":[{"message_id":"17173","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Diskusi tools programming biasa (Cursor, VSCode, Claude), tidak ada pelanggaran."}]}',
modes: ["text", "mixed"],
},
{
id: "19",
title: "Diskusi orientasi seksual LGBT (dilarang — zero tolerance)",
input:
"[target] id=17174 user=alex: btw gw gay, semoga ga masalah ya. Gw cmn mau jujur aja",
output:
'{"results":[{"message_id":"17174","status":"flagged","flags":["sexual_deviation"],"score":0.85,"severity":"high","confidence":0.95,"recommended_action":"delete","evidence":["btw gw gay, semoga ga masalah ya"],"analysis":"Pengakuan orientasi LGBT — server zero tolerance terhadap segala diskusi/coming out. Dihapus."}]}',
modes: ["text", "mixed"],
},
// ── Mixed examples ──
{
id: "3",
@@ -110,7 +100,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
'[target] id=11111 user=citra: <:singa_injek:123456> [sticker: "Singa injek pejabat"]',
output:
'{"results":[{"message_id":"11111","status":"clean","flags":[],"score":0.1,"severity":"none","confidence":0.8,"recommended_action":"none","evidence":[],"analysis":"Sticker kartun satir dengan nama provokatif, bukan ancaman nyata."}]}',
'{"results":[{"message_id":"11111","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Sticker kartun satir dengan nama provokatif, bukan ancaman nyata."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -119,7 +109,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=88888 user=sexist: dasar perempuan ngerti apa sih, logika lo aja kagak bener",
output:
'{"results":[{"message_id":"88888","status":"flagged","flags":["hate_speech","harassment"],"score":0.82,"severity":"high","confidence":0.9,"recommended_action":"delete","evidence":["dasar perempuan ngerti apa sih"],"analysis":"Komentar seksis merendahkan yang menyasar gender — stereotip ofensif terarah."}]}',
'{"results":[{"message_id":"88888","status":"flagged","flags":["hate_speech","harassment"],"severity":"high","evidence":["dasar perempuan ngerti apa sih"],"analysis":"Komentar seksis merendahkan yang menyasar gender — stereotip ofensif terarah."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -128,7 +118,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=99999 user=drama: si budi kemarin ngomongin lo di belakang, masa tega banget dia, ayo kita konfrontasi di sini aja",
output:
'{"results":[{"message_id":"99999","status":"warn","flags":["conflict_instigation"],"score":0.65,"severity":"low","confidence":0.75,"recommended_action":"warn","evidence":["ayo kita konfrontasi di sini aja"],"analysis":"Mengajak konfrontasi masalah personal di channel publik, berpotensi memicu drama."}]}',
'{"results":[{"message_id":"99999","status":"warn","flags":["conflict_instigation"],"severity":"low","evidence":["ayo kita konfrontasi di sini aja"],"analysis":"Mengajak konfrontasi masalah personal di channel publik, berpotensi memicu drama."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -137,7 +127,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=10101 user=fox: mau liat foto pake kostum hewan? DM aja, khusus 18+",
output:
'{"results":[{"message_id":"10101","status":"flagged","flags":["sexual_deviation"],"score":0.85,"severity":"high","confidence":0.9,"recommended_action":"delete","evidence":["mau liat foto pake kostum hewan? DM aja, khusus 18+"],"analysis":"Ajakan aktivitas seksual eksplisit \'DM khusus 18+\'. Melanggar kebijakan."}]}',
'{"results":[{"message_id":"10101","status":"flagged","flags":["sexual_deviation"],"severity":"high","evidence":["mau liat foto pake kostum hewan? DM aja, khusus 18+"],"analysis":"Ajakan aktivitas seksual eksplisit \'DM khusus 18+\'. Melanggar kebijakan."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -146,7 +136,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=10505 user=dev: ERROR: Cannot read properties of undefined (reading 'data'). Stack trace: at Module._compile (node:internal/modules/cjs/loader:1256:14)",
output:
'{"results":[{"message_id":"10505","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Error log programming biasa antara developer, aman."}]}',
'{"results":[{"message_id":"10505","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Error log programming biasa antara developer, aman."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -155,7 +145,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=12121 user=pejabat_munafik_dajjal: Halo teman-teman, ada yang main game?",
output:
'{"results":[{"message_id":"12121","status":"flagged","flags":["offensive_username"],"score":0.3,"severity":"low","confidence":0.95,"recommended_action":"warn","evidence":["Username \'pejabat_munafik_dajjal\' mengandung unsur ofensif/SARA"],"analysis":"Username ofensif menyerang pejabat dengan label SARA, tapi isi pesan bersih — flag ringan."}]}',
'{"results":[{"message_id":"12121","status":"flagged","flags":["offensive_username"],"severity":"low","evidence":["Username \'pejabat_munafik_dajjal\' mengandung unsur ofensif/SARA"],"analysis":"Username ofensif menyerang pejabat dengan label SARA, tapi isi pesan bersih — flag ringan."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -164,7 +154,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=13131 user=nazi_babi_itu: bener tuh nih ras emang harus dibasmi",
output:
'{"results":[{"message_id":"13131","status":"flagged","flags":["offensive_username","hate_speech","sara"],"score":0.9,"severity":"high","confidence":0.95,"recommended_action":"delete","evidence":["bener tuh nih ras emang harus dibasmi"],"analysis":"Username SARA dan isi pesan memperkuat kebencian dengan ajakan kekerasan. Pelanggaran berat."}]}',
'{"results":[{"message_id":"13131","status":"flagged","flags":["offensive_username","hate_speech","sara"],"severity":"high","evidence":["bener tuh nih ras emang harus dibasmi"],"analysis":"Username SARA dan isi pesan memperkuat kebencian dengan ajakan kekerasan. Pelanggaran berat."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -173,7 +163,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
'[target] id=14141 user=hater: ++++++K1[[ your $€/F" "~`| \\\\0ve $ 1F ¥°U |}iE ®©',
output:
'{"results":[{"message_id":"14141","status":"flagged","flags":["harassment","hate_speech"],"score":0.95,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["++++++K1[[ your $€/F"],"analysis":"Obfuscation simbol menyembunyikan frasa ancaman \'Kill yourself\'. Ancaman berat yang disamarkan."}]}',
'{"results":[{"message_id":"14141","status":"flagged","flags":["harassment","hate_speech"],"severity":"critical","evidence":["++++++K1[[ your $€/F"],"analysis":"Obfuscation simbol menyembunyikan frasa ancaman \'Kill yourself\'. Ancaman berat yang disamarkan."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -181,7 +171,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
title: "Vulgaritas Bahasa Asing / All-Caps",
input: "[target] id=15151 user=troll: AKU RAJA TITTEN",
output:
'{"results":[{"message_id":"15151","status":"flagged","flags":["vulgar_language"],"score":0.85,"severity":"medium","confidence":0.9,"recommended_action":"delete","evidence":["AKU RAJA TITTEN"],"analysis":"Kata vulgar bahasa asing (\'titten\') dengan huruf kapital — melanggar meski format seperti candaan."}]}',
'{"results":[{"message_id":"15151","status":"flagged","flags":["vulgar_language"],"severity":"medium","evidence":["AKU RAJA TITTEN"],"analysis":"Kata vulgar bahasa asing (\'titten\') dengan huruf kapital — melanggar meski format seperti candaan."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -190,7 +180,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
'[target] id=20201 user=parodist: Kitabonia 11:17 — "Dan bersabdalah Sang Admin: barang siapa yang melakukan spam, niscaya akan kena mute tujuh hari tujuh malam"',
output:
'{"results":[{"message_id":"20201","status":"flagged","flags":["sara"],"score":0.92,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["Kitabonia 11:17"],"analysis":"Ayat palsu dengan format kitab suci memparodikan wahyu — penistaan agama, bukan humor."}]}',
'{"results":[{"message_id":"20201","status":"flagged","flags":["sara"],"severity":"critical","evidence":["Kitabonia 11:17"],"analysis":"Ayat palsu dengan format kitab suci memparodikan wahyu — penistaan agama, bukan humor."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -199,7 +189,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=22223 user=edgy: Shirkmaxxing grindset, nanti halalmaxxing juga",
output:
'{"results":[{"message_id":"22223","status":"flagged","flags":["sara"],"score":0.88,"severity":"high","confidence":0.95,"recommended_action":"delete","evidence":["Shirkmaxxing grindset"],"analysis":"Istilah suci agama (shirk, halal) sebagai bahan candaan meme — penistaan konsep teologis."}]}',
'{"results":[{"message_id":"22223","status":"flagged","flags":["sara"],"severity":"high","evidence":["Shirkmaxxing grindset"],"analysis":"Istilah suci agama (shirk, halal) sebagai bahan candaan meme — penistaan konsep teologis."}]}',
modes: ["text", "media", "mixed"],
},
{
@@ -207,7 +197,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
title: "Ekspresi keagamaan normal (AMAN, BUKAN SARA)",
input: "[target] id=27278 user=muslim_user: Astaghfirullah, sabar ya bro",
output:
'{"results":[{"message_id":"27278","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Istighfar untuk menenangkan teman — ekspresi keagamaan wajar Indonesia, bukan penistaan."}]}',
'{"results":[{"message_id":"27278","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Istighfar untuk menenangkan teman — ekspresi keagamaan wajar Indonesia, bukan penistaan."}]}',
modes: ["text", "media", "mixed"],
},
@@ -218,7 +208,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=22222 user=rina: Aku suka nasgor loh [Media analysis for message 22222] [gambar di atas adalah attachment foto.jpg dari pesan id=22222]: Gambar menampilkan tangkapan layar aplikasi chat dengan teks percakapan biasa. Tidak ada konten melanggar terlihat. Aman.",
output:
'{"results":[{"message_id":"22222","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Percakapan sehari-hari tentang makanan; gambar screenshot chat biasa tanpa pelanggaran."}]}',
'{"results":[{"message_id":"22222","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Percakapan sehari-hari tentang makanan; gambar screenshot chat biasa tanpa pelanggaran."}]}',
modes: ["media", "mixed"],
},
{
@@ -227,7 +217,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
'[target] id=33333 user=spammer: MAIN DI SINI GACOR PARAH https://judionline.xyz [Media analysis for message 33333] [gambar di atas adalah attachment slot.jpg dari pesan id=33333]: Gambar menampilkan antarmuka situs judi online dengan mesin slot, chip, dan tombol deposit. Terlihat logo "JudiOnline" dan odds taruhan.',
output:
'{"results":[{"message_id":"33333","status":"flagged","flags":["gambling"],"score":0.92,"severity":"high","confidence":0.92,"recommended_action":"delete","evidence":["MAIN DI SINI GACOR PARAH","https://judionline.xyz"],"analysis":"Promosi situs judi dengan link, teks promosi, dan gambar antarmuka judi yang jelas."}]}',
'{"results":[{"message_id":"33333","status":"flagged","flags":["gambling"],"severity":"high","evidence":["MAIN DI SINI GACOR PARAH","https://judionline.xyz"],"analysis":"Promosi situs judi dengan link, teks promosi, dan gambar antarmuka judi yang jelas."}]}',
modes: ["media", "mixed"],
},
{
@@ -236,7 +226,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=44444 user=dev: [Media analysis for message 44444] [gambar di atas adalah attachment screenshot.png dari pesan id=44444]: Screenshot terminal Linux dengan background hitam dan teks hijau. Terlihat output command 'ls -la' dan 'git status'. Tidak ada teks atau elemen mencurigakan.",
output:
'{"results":[{"message_id":"44444","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.95,"recommended_action":"none","evidence":[],"analysis":"Screenshot terminal Linux (ls -la, git status) — aktivitas coding biasa, tidak ada pelanggaran."}]}',
'{"results":[{"message_id":"44444","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Screenshot terminal Linux (ls -la, git status) — aktivitas coding biasa, tidak ada pelanggaran."}]}',
modes: ["media", "mixed"],
},
{
@@ -245,7 +235,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
'[target] id=55555 user=promotor: [Media analysis for message 55555] [gambar di atas adalah attachment promo.jpg dari pesan id=55555]: Screenshot website dengan background merah dan emas. Terlihat teks "DEPOSIT NOW", "BONUS 100%", "SLOT GACOR", chip poker, dan roda roulette. Ada tombol "DAFTAR" dan "LOGIN".',
output:
'{"results":[{"message_id":"55555","status":"flagged","flags":["gambling"],"score":0.94,"severity":"high","confidence":0.94,"recommended_action":"delete","evidence":["Gambar antarmuka situs judi dengan chip, roulette, tombol deposit"],"analysis":"Gambar promosi judi jelas: chip, roulette, teks deposit dan bonus. Melanggar gambling."}]}',
'{"results":[{"message_id":"55555","status":"flagged","flags":["gambling"],"severity":"high","evidence":["Gambar antarmuka situs judi dengan chip, roulette, tombol deposit"],"analysis":"Gambar promosi judi jelas: chip, roulette, teks deposit dan bonus. Melanggar gambling."}]}',
modes: ["media", "mixed"],
},
{
@@ -254,7 +244,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=21212 user=troll: ayo bikin agama, aku tuhannya, yg mau jadi malaikat DM aku",
output:
'{"results":[{"message_id":"21212","status":"flagged","flags":["sara"],"score":0.95,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["ayo bikin agama, aku tuhannya"],"analysis":"Mengajak membuat agama palsu dan mengaku Tuhan — penistaan agama serius, bukan candaan."}]}',
'{"results":[{"message_id":"21212","status":"flagged","flags":["sara"],"severity":"critical","evidence":["ayo bikin agama, aku tuhannya"],"analysis":"Mengajak membuat agama palsu dan mengaku Tuhan — penistaan agama serius, bukan candaan."}]}',
modes: ["media", "mixed"],
},
{
@@ -263,7 +253,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=23234 user=edgelord: Hashem is watching you jerk off lol",
output:
'{"results":[{"message_id":"23234","status":"flagged","flags":["sara"],"score":0.94,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["Hashem is watching you jerk off lol"],"analysis":"Nama suci (Hashem) dalam konteks vulgar merendahkan — blasphemy serius."}]}',
'{"results":[{"message_id":"23234","status":"flagged","flags":["sara"],"severity":"critical","evidence":["Hashem is watching you jerk off lol"],"analysis":"Nama suci (Hashem) dalam konteks vulgar merendahkan — blasphemy serius."}]}',
modes: ["media", "mixed"],
},
{
@@ -272,7 +262,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=24245 user=provokator: nabi palsu ba'al, sembah aja patung",
output:
'{"results":[{"message_id":"24245","status":"flagged","flags":["sara","hate_speech"],"score":0.9,"severity":"high","confidence":0.95,"recommended_action":"delete","evidence":["nabi palsu ba\'al"],"analysis":"Sindiran konsep nabi dengan referensi Ba\'al + ajakan menyembah patung — penistaan monoteis."}]}',
'{"results":[{"message_id":"24245","status":"flagged","flags":["sara","hate_speech"],"severity":"high","evidence":["nabi palsu ba\'al"],"analysis":"Sindiran konsep nabi dengan referensi Ba\'al + ajakan menyembah patung — penistaan monoteis."}]}',
modes: ["media", "mixed"],
},
{
@@ -281,7 +271,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=25256 user=sindiran: Bodoh admin-admin kita itu. Mereka tidak minta petunjuk dari Tuhan",
output:
'{"results":[{"message_id":"25256","status":"flagged","flags":["sara","hate_speech"],"score":0.82,"severity":"high","confidence":0.9,"recommended_action":"delete","evidence":["tidak minta petunjuk dari Tuhan"],"analysis":"Framework agama (\'petunjuk dari Tuhan\') dipakai untuk mengejek — parodi merendahkan."}]}',
'{"results":[{"message_id":"25256","status":"flagged","flags":["sara","hate_speech"],"severity":"high","evidence":["tidak minta petunjuk dari Tuhan"],"analysis":"Framework agama (\'petunjuk dari Tuhan\') dipakai untuk mengejek — parodi merendahkan."}]}',
modes: ["media", "mixed"],
},
{
@@ -290,7 +280,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=26267 user=provokator: Islam/Kristen itu agama sesat semua, cuma agama gue yg bener",
output:
'{"results":[{"message_id":"26267","status":"flagged","flags":["sara","hate_speech","conflict_instigation"],"score":0.95,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["Islam/Kristen itu agama sesat semua"],"analysis":"Menyerang agama lain sebagai \'sesat\' dengan truth claim provokatif — provokasi antar-agama serius."}]}',
'{"results":[{"message_id":"26267","status":"flagged","flags":["sara","hate_speech","conflict_instigation"],"severity":"critical","evidence":["Islam/Kristen itu agama sesat semua"],"analysis":"Menyerang agama lain sebagai \'sesat\' dengan truth claim provokatif — provokasi antar-agama serius."}]}',
modes: ["media", "mixed"],
},
{
@@ -299,7 +289,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=28289 user=thinker: Menurutku konsep trinitas dalam Kristen menarik dibandingkan tauhid dalam Islam. Apa pendapat kalian?",
output:
'{"results":[{"message_id":"28289","status":"clean","flags":[],"score":0.0,"severity":"none","confidence":0.9,"recommended_action":"none","evidence":[],"analysis":"Perbandingan teologis dengan bahasa sopan dan tone diskusi — bukan provokasi. Aman."}]}',
'{"results":[{"message_id":"28289","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Perbandingan teologis dengan bahasa sopan dan tone diskusi — bukan provokasi. Aman."}]}',
modes: ["media", "mixed"],
},
@@ -310,7 +300,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=29292 user=promotor: Join sini bro https://discord.gg/xyzk123 diskusi coding seru",
output:
'{"results":[{"message_id":"29292","status":"warn","flags":["spam"],"score":0.55,"severity":"low","confidence":0.7,"recommended_action":"warn","evidence":["https://discord.gg/xyzk123"],"analysis":"Promosi server lain via invite di channel publik tanpa izin — potensi spam."}]}',
'{"results":[{"message_id":"29292","status":"warn","flags":["spam"],"severity":"low","evidence":["https://discord.gg/xyzk123"],"analysis":"Promosi server lain via invite di channel publik tanpa izin — potensi spam."}]}',
modes: ["text", "media", "mixed"],
},
@@ -321,7 +311,7 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=30303 user=rainbow: jujur aja gua lesbian, terima apa adanya",
output:
'{"results":[{"message_id":"30303","status":"flagged","flags":["sexual_deviation"],"score":0.9,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["jujur aja gua lesbian"],"analysis":"Pengakuan orientasi lesbian — zero tolerance LGBT di server ini. Dihapus."}]}',
'{"results":[{"message_id":"30303","status":"flagged","flags":["sexual_deviation"],"severity":"critical","evidence":["jujur aja gua lesbian"],"analysis":"Pengakuan orientasi lesbian — zero tolerance LGBT di server ini. Dihapus."}]}',
modes: ["text", "media", "mixed"],
},
@@ -332,9 +322,21 @@ export const ALL_EXAMPLES: ExampleDef[] = [
input:
"[target] id=31313 user=hot_takes: gw sih dukung palestina, israel biadab banget",
output:
'{"results":[{"message_id":"31313","status":"flagged","flags":["conflict_instigation","sara"],"score":0.95,"severity":"critical","confidence":0.95,"recommended_action":"delete","evidence":["gw sih dukung palestina"],"analysis":"Segala diskusi Israel/Palestina/Yahudi dilarang total — tidak ada debat, dukungan, atau berita. Dihapus."}]}',
'{"results":[{"message_id":"31313","status":"flagged","flags":["conflict_instigation","sara"],"severity":"critical","evidence":["gw sih dukung palestina"],"analysis":"Segala diskusi Israel/Palestina/Yahudi dilarang total — tidak ada debat, dukungan, atau berita. Dihapus."}]}',
modes: ["text", "media", "mixed"],
},
// ── Physics / Technology Discussions (false positive prevention) ──
{
id: "32",
title: "Diskusi fisika/kinetik dalam konteks teknis (AMAN, bukan ancaman)",
input:
"[target] id=32323 user=physics_student: Cukup cuman tubuh manusia vs gravitasi. Konsep energy conservation di sini penting buat analisis statis.",
output:
'{"results":[{"message_id":"32323","status":"clean","flags":[],"severity":"none","evidence":[],"analysis":"Diskusi fisika teknis tentang kinetik dan gravitasi dalam konteks analisis statis tidak ada ancaman atau konten melanggar. Penggunaan istilah fisika untuk perhitungan teknis adalah hal wajar."}]}',
modes: ["text", "mixed"],
},
// ── Invite link / promosi server ──
];
// Derive per-mode strings from the single ALL_EXAMPLES array (zero duplication)
@@ -35,13 +35,11 @@ Instruksi per field:
- "message_id": WAJIB sama persis dengan id di input. Setiap <message> di <messages_to_analyze> menghasilkan SATU hasil. Jangan gabungkan beberapa pesan, jangan lewati, jangan karang id.
- "evidence": kutipan PERSIS frasa yang melanggar (maks 1 baris). Pelanggaran di gambar/sticker kutip deskripsi Media analysis. Pelanggaran lewat balasan/referensi sebut konteks pesan yang dibalas. Boleh tambah label sumber, mis. [media analysis] / [web_search] / [reply]. Kosong jika clean.
## PERSONALITY & MEMORI Profil Pengguna dan Kultur Channel
Data konteks tersedia: <user_profiles> (peta ringkasan kepribadian, di pesan USER), <user_reputation> (skor trust), dan <channel_culture> (topik/vibe channel). Setiap <message> dapat memuat <user_profile_ref user_id="..."/> yang menunjuk ke entri di peta <user_profiles>.
## KONTEKS Kultur Channel
Data konteks tersedia: <channel_culture> (topik/vibe channel). Tidak ada data profil/reputasi per-user nilai tiap pesan murni dari isinya.
Gunakan untuk personalisasi analysis, tapi:
- Profil adalah KONTEKS, bukan bukti. Profil mencurigakan flag; profil bersih loloskan pelanggaran.
- Perubahan perilaku mencolok (biasanya teknis tiba-tiba provokatif) layak dicatat di analysis.
- <user_history> (kutipan pesan yang pernah di-flag) = pola pelanggaran lama. Gunakan untuk mendeteksi PENGULANGAN (mis. spam link yang sama, provokasi berulang), tapi JANGAN memflag pesan bersih hanya karena riwayat.
- JANGAN paksa referensi profil jika tidak relevan analysis natural lebih baik.
- Konteks adalah KONTEKS, bukan bukti. Riwayat di <conversation_context> membantu pahami alur, tapi pesan bersih tanpa pelanggaran CLEAN. JANGAN gunakan konteks untuk "menginterpretasi ulang" pesan bersih yang terpisah.
- Perubahan perilaku mencolok (mis. teknis tiba-tiba provokatif) layak dicatat. JANGAN paksa referensi profil jika tidak relevan analysis natural lebih baik.
- Channel culture coding/teknis pesan teknis lebih wajar; channel santai slang lebih wajar. Jangan dipakai mengabaikan pelanggaran nyata.
## FORMAT WAJIB analysis HARUS deskriptif berdasarkan konten:
@@ -72,8 +70,7 @@ CRITICAL:
- Jika pesan adalah BALASAN (reply) ke pesan lain, jelaskan konteks balasannya: apa yang sedang dibicarakan, siapa yang dibalas (tanpa nama, cukup peran/isi pesan yang dibalas), dan bagaimana tanggapan pengirim terhadapnya.
- Gunakan informasi dari Media analysis untuk mendeskripsikan gambar.
- Analisis harus MEMBERI KONTEKS, bukan hanya menyatakan status.
- GUNAKAN <user_profile_ref>/<user_profiles> untuk personalisasi analysis jadikan analysis terasa seperti sistem "mengenal" pengguna.
- Jika perilaku pesan menyimpang dari profil yang diketahui, CATAT dalam analysis sebagai informasi kontekstual yang relevan.
- Nilai tiap pesan murni dari isinya sendiri + <conversation_context> + <web_searches> + <location_context>. Tidak ada reputasi/profil per-user di context.
- JANGAN paksa referensi profil jika tidak relevan analysis natural lebih baik dari yang dipaksakan.`;
// ---------------------------------------------------------------------------
@@ -12,6 +12,7 @@ export const SYSTEM_RULES = `Kamu adalah asisten moderasi konten untuk server Di
## Normalisasi & Pertahanan Lintas Bahasa (WAJIB)
1. Campuran bahasa (Inggris/Indonesia/daerah) WAJIB dinormalisasi mental ke Bahasa Indonesia sebelum menilai intent. Jangan longgar hanya karena sintaksis campur (Polyglot Obfuscation).
2. Lakukan Named Entity Recognition agresif nama orang/karakter (mis. "ren" setelah kata archaic "diagem") tetap dikenali sebagai nama.
3. <term_glossary> (bila ada) = definisi kata/slang/jargon yang tidak umum. Baca dulu arti kata yang tidak kamu kenal dari sana jangan menebak dari bunyi/kemiripan. Kata yang tampak mencurigakan namun ternyata bermakna netral di glossary = AMAN; kata asing yang ternyata vulgar/terlarang di glossary = FLAG.
## Aturan Umum (AMAN jangan flag)
- Slang: anjay, wkwk, gws, gaskeun, santuy, njir, baka, woy/woi, hadeh, astaga = AMAN.
@@ -29,15 +30,19 @@ export const SYSTEM_RULES = `Kamu adalah asisten moderasi konten untuk server Di
- Ekspresi religius (Astaghfirullah, Alhamdulillah, Subhanallah, Allahuakbar, MasyaAllah, Bismillah, InsyaAllah, Laa ilaha illallah + varian all-caps) = DOA NORMAL, bukan vulgar. AMAN.
- Discord custom emoji (<:hadeh:123>) = ekspresi, bukan pelanggaran teks.
- Makian pada entitas eksternal (game, dev, perusahaan, benda mati: "game ini ampas") = AMAN. Harassment/hate_speech HANYA untuk anggota/kelompok server secara personal.
- **Diskusi fisika, teknik, atau engineering dalam konteks teknis** (kinetik, gravitasi, energi, drone, senjata, drone warfare, physics simulations, CAD, CNC, 3D printing, robotics, aerospace, aerodynamika) = AMAN. Penggunaan istilah teknis untuk perhitungan atau analisis bukan ancaman. JANGAN flag hanya karena istilah "senjata" atau "drone" dalam konteks diskusi teori teknis. Flag HANYA jika ada ajuan aksi eksplisit atau ancaman nyata terarah.
- **Riwayat pengguna dengan pelanggaran sebelumnya** tidak boleh memengaruhi penilaian pesan bersih yang TERPISAH dan tidak mengandung pelanggaran aktual. Setiap pesan dinilai berdasarkan ISINYA SENDIRI.
## Zero Tolerance Vulgaritas Anatomi/Seksual
Kata alat kelamin/anatomi seksual (kontol, memek, titten, tit, dick) atau istilah seksual eksplisit WAJIB di-flag sebagai vulgar_language/sexual_content TANPA pengecualian bercanda, slang, atau "santai".
## Nilai Server Diskriminasi
- Seksisme ("dasar perempuan", "logika cewek") hate_speech (umum) / harassment (terarah).
- Ageisme ("dasar bocil", "tau aja lo tua") hate_speech / harassment.
- Diskriminasi fisik ("gendut", "iteman", "cungkring") harassment jika terarah.
- Serangan personal, penghinaan, merendahkan = tidak ditoleransi. Perbedaan pendapat wajar.
-Ketika sesuatu yang melanggar terjadi di channel, flag jika relevan. Setiap pesan dinilai BERDASARKAN ISINYA SENDIRI, bukan sekadar histori pengguna.
-Seksisme ("dasar perempuan", "logika cewek") hate_speech (umum) / harassment (terarah).
-Ageisme ("dasar bocil", "tau aja lo tua") hate_speech / harassment.
-Diskriminasi fisik ("gendut", "iteman", "cungkring") harassment jika terarah.
-Serangan personal, penghinaan, merendahkan = tidak ditoleransi. Perbedaan pendapat wajar.
+**PESAN DINILAI SECARA STANDALONE:** Setiap pesan baru dinilai BERDASARKAN ISINYA SENDIRI. <user_history> (jika ada) HANYA untuk mendeteksi POLA PENGULANGAN dengan JAMAK (spam link yang SAMA, provokasi berulang yang MENGANDALKAN KONTEN YANG SAMA). JANGAN gunakan history untuk "menginterpretasi ulang" pesan bersih yang TERPISAH DARI riwayat pelanggaran sebelumnya. Jika pesan tidak mengandung unsur yang BERPANDUAN PADA riwayat tetap CLEAN.
## LARANGAN BERAT (ZERO TOLERANCE)
- **LGBT:** Segala promosi, diskusi, pengakuan orientasi, coming out, atau curhat personal tentang LGBT WAJIB di-flag "sexual_deviation". Tidak ada pengecualian.
@@ -73,6 +78,7 @@ RENDAH: harassment, vulgar_language terarah, offensive_username (Scunthorpe: "Sa
## Web Sebagai Bukti Utama
- <web_searches> ADALAH BUKTI UTAMA. Jika ada, WAJIB pakai hasilnya (hentai/scam/narkoba flag; aman clean). JANGAN abaikan. Jika tidak ada gunakan pengetahuan internal.
- <term_glossary> = REFERENSI ARTI KATA, bukan bukti pelanggaran. Dipakai untuk memahami istilah yang tidak dikenal sebelum memutuskan.
- Prioritas bukti: <web_searches> > <web_content> > <media_analysis> > pengetahuan internal. <web_content> (URL fetch): gunakan isi, jangan flag hanya dari domain name.
## Pohon Keputusan
@@ -103,35 +103,21 @@ export function buildSystemPrompt(options: BuildSystemPromptOptions): string {
}
parts.push(
`## Blok Data di Pesan USER\n` +
`Semua data dinamis per-batch dikirim di pesan USER — system prompt ini TIDAK memuat data batch:\n` +
`- <location_context .../> = metadata channel/thread (channel_id, channel_name, thread_name, topic, nsfw, age_restricted). topic = deskripsi resmi channel — pakai untuk menilai kesesuaian pesan dengan tujuan channel.\n` +
`- <conversation_context> = obrolan SEBELUM pesan target. Baris "[context]" di dalamnya BUKAN yang dinilai.\n` +
`- <user_profiles> = peta ringkasan kepribadian per user_id (attr as_of = kapan profil terakhir dibuat — profil lama mungkin tidak mencerminkan perilaku terkini); setiap <message> merujuk lewat <user_profile_ref user_id="..."/>.\n` +
`- <web_searches> / <web_content> = bukti web (lihat "Web Sebagai Bukti Utama").\n` +
`- <messages_to_analyze> = pesan-pesan TARGET yang WAJIB dinilai. Atribut <message>: id, user (nama server), time (ISO — kapan pesan dikirim), repetitions (N = teks pendek sama muncul N kali di batch sinyal spam), bot (true jika dari bot), edited (true jika konten adalah hasil edit setelah posting).`,
`## Blok Data (di pesan USER)\n` +
`System prompt ini TIDAK berisi data batch — semua data per-batch ada di pesan USER:\n` +
`- <location_context .../>: metadata channel/thread (channel_name, thread_name, topic, nsfw, age_restricted). topic = tujuan resmi channel; gunakan menilai kesesuaian pesan.\n` +
`- <conversation_context>: obrolan SEBELUM target. Baris pertama "[conversation_flow] status=... context_msgs=... dropped=..." = metadata sistem (ongoing/sparse/cold_start), BUKAN pesan dinilai. Baris "[context] id=... time=... user=...: isi" = konteks, BUKAN target.\n` +
`- Tidak ada data profil/reputasi per-user di context — nilai tiap pesan murni dari isinya + <conversation_context> + <web_searches> + <location_context>.\n` +
`- <web_searches>/<web_content>: bukti web (prioritas tertinggi). <term_glossary>: definisi kata/slang/jargon (SearXNG) — pakai pahami kata asing, JANGAN tebak arti.\n` +
`- <messages_to_analyze>: pesan TARGET yang WAJIB dinilai. Atribut <message>: id, user, time (ISO), repetitions (N = teks sama muncul N× di batch sinyal spam), bot (true = bot), edited (true = hasil edit setelah posting → evasi potensial).`,
);
parts.push(
`## Konteks Pengguna (Referensi, Bukan Bukti)\n` +
`Konteks per pengguna hanya indikator **referensi** untuk personalisasi analisis, BUKAN bukti pelanggaran:\n` +
`- <user_reputation trust_score="..." total_infractions="..." clean_streak="..." last_offense_days_ago="..." repeat_offender="..."> = histori moderasi pengguna. Skor rendah BUKAN alasan memflag pesan bersih; skor tinggi BUKAN alasan mengabaikan pelanggaran nyata. repeat_offender="true" = ada pelanggaran dalam 7 hari terakhir.\n` +
`- <user_history> (di dalam <user_reputation>) = kutipan pesan-pesan pengguna yang PERNAH di-flag. Gunakan untuk mengenali POLA berulang (spam link sama, provokasi), tapi JANGAN memflag pesan bersih hanya karena riwayat.\n` +
`- <user_profiles> (di pesan USER) = peta ringkasan kepribadian per user_id. <user_profile_ref user_id="..."/> dalam sebuah pesan menunjuk ke peta itu. Tanpa ref = tidak ada profil untuk pengguna tersebut.\n` +
`- Profil berguna untuk mengenali penyimpangan perilaku mencolok (mis. pengguna teknis tiba-tiba provokatif), tapi JANGAN memflag atau meloloskan hanya karena profil.\n` +
`**Setiap pesan dinilai berdasarkan isinya sendiri.**`,
);
parts.push(
`## Framing: Konteks vs Target\n` +
`- Baris dalam <conversation_context> berformat "[context] id=... time=<ISO> user=<nama>: isi", diurutkan paling lama → paling baru. Baris pertama biasanya "[conversation_flow] status=... context_msgs=... dropped=..." — metadata sistem tentang status percakapan (ongoing/sparse/cold_start), BUKAN pesan yang dinilai.\n` +
`- <messages_to_analyze> berisi pesan-pesan TARGET yang WAJIB dinilai. Hasilkan SATU hasil per message_id — jangan menggabungkan beberapa pesan, jangan melewati, jangan mengarang id.\n` +
`- Setiap target dinilai berdasarkan isinya sendiri; konteks percakapan memengaruhi interpretasi, bukan menggantikan isi pesan.\n` +
`- Marker "…[pesan dipotong: terlalu panjang]" = konten TARGET sengaja dipotong; marker "…[konteks dipotong: terlalu panjang]" = konten pesan KONTEKS dipotong. Nilai dari bagian yang terlihat; pemotongan BUKAN pelanggaran dan BUKAN teknik evasi.\n` +
`- Atribut time= pada <message> target = kapan pesan dikirim (ISO). Pakai untuk menilai kerelevanan waktu (mis. pesan lama di-bump, spam beruntun dalam menit yang sama).\n` +
`- repetitions="N" pada <message> = teks pendek yang sama muncul N kali dalam batch — pertimbangkan sebagai sinyal spam, tapi nilai tetap dari isi pesan.\n` +
`- bot="true" = pengirim adalah bot (otomatisasi), bukan pengguna manusia — jangan perlakukan sebagai pelanggaran personal, tapi kontennya tetap dinilai.\n` +
`- edited="true" = konten yang ditampilkan adalah hasil edit setelah posting (sinyal potensi evasi), nilai konten saat ini apa adanya.`,
`## Framing & Aturan Konteks\n` +
`- Hasilkan SATU hasil per message_id — jangan gabung, lewati, atau karang id.\n` +
`- Setiap target dinilai BERDASARKAN ISINYA SENDIRI. Konteks memengaruhi interpretasi, tapi TIDAK menggantikan isi pesan. Profil/riwayat = REFERENSI personalisasi, BUKAN bukti pelanggaran (lihat "PERSONALITY & MEMORI").\n` +
`- Marker "[pesan dipotong: terlalu panjang]" = TARGET dipotong; "[konteks dipotong: ...]" = konteks dipotong. Nilai dari bagian terlihat; pemotongan BUKAN pelanggaran/evasi.\n` +
`- time= = kapan dikirim (rekonsiliasi spam beruntun / bump pesan lama). bot=true = otomatisasi, bukan pelanggaran personal.`,
);
parts.push(OUTPUT_INSTRUCTIONS);
@@ -17,6 +17,18 @@ import { config } from "../../shared/config/config.js";
const log = createChildLogger("qdrant");
// ensureQdrantCollection performs a network round-trip (GET, possibly
// DELETE+PUT). Running it on every upsert adds 1-3 HTTP calls per
// moderation verdict, which under Qdrant load pushes the upsert past the
// request timeout and aborts it ("This operation was aborted"). Memoise the
// result so the collection is only verified once per process lifetime.
let ensureCollectionPromise: Promise<boolean> | null = null;
/** Reset the memoised ensure result (used by tests / config reload). */
export function resetQdrantCollectionCache(): void {
ensureCollectionPromise = null;
}
export interface QdrantVerdictPayload {
text: string;
flags: string; // JSON string of the full moderation result
@@ -97,46 +109,53 @@ export function qdrantPointId(cacheKey: string): number {
export async function ensureQdrantCollection(
vectorSize: number,
): Promise<boolean> {
try {
// 404 = collection doesn't exist yet → create it.
let existing: {
result?: { config?: { params?: { vectors?: { size?: number } } } };
} | null = null;
if (ensureCollectionPromise) return ensureCollectionPromise;
ensureCollectionPromise = (async () => {
try {
existing = (await request("GET", `/collections/${collectionName()}`)) as {
// 404 = collection doesn't exist yet → create it.
let existing: {
result?: { config?: { params?: { vectors?: { size?: number } } } };
};
} catch (error) {
if (!(error instanceof Error) || !error.message.includes("-> 404")) {
throw error;
} | null = null;
try {
existing = (await request(
"GET",
`/collections/${collectionName()}`,
)) as {
result?: { config?: { params?: { vectors?: { size?: number } } } };
};
} catch (error) {
if (!(error instanceof Error) || !error.message.includes("-> 404")) {
throw error;
}
}
}
const size = existing?.result?.config?.params?.vectors?.size;
if (size === vectorSize) return true;
const size = existing?.result?.config?.params?.vectors?.size;
if (size === vectorSize) return true;
if (size !== undefined && size !== vectorSize) {
log.warn(
{ collection: collectionName(), oldSize: size, newSize: vectorSize },
"Qdrant collection vector size changed — recreating collection",
if (size !== undefined && size !== vectorSize) {
log.warn(
{ collection: collectionName(), oldSize: size, newSize: vectorSize },
"Qdrant collection vector size changed — recreating collection",
);
await request("DELETE", `/collections/${collectionName()}`);
}
await request("PUT", `/collections/${collectionName()}`, {
vectors: { size: vectorSize, distance: "Cosine" },
});
return true;
} catch (error) {
log.error(
{
error: error instanceof Error ? error.message : String(error),
collection: collectionName(),
},
"Failed to ensure Qdrant collection",
);
await request("DELETE", `/collections/${collectionName()}`);
return false;
}
await request("PUT", `/collections/${collectionName()}`, {
vectors: { size: vectorSize, distance: "Cosine" },
});
return true;
} catch (error) {
log.error(
{
error: error instanceof Error ? error.message : String(error),
collection: collectionName(),
},
"Failed to ensure Qdrant collection",
);
return false;
}
})();
return ensureCollectionPromise;
}
/** Upsert one embedding + verdict payload point. Returns false on failure. */
@@ -147,10 +166,15 @@ export async function upsertQdrantPoint(
): Promise<boolean> {
try {
if (!(await ensureQdrantCollection(vector.length))) return false;
await request("PUT", `/collections/${collectionName()}/points`, {
points: [{ id: qdrantPointId(cacheKey), vector, payload }],
wait: true,
});
await request(
"PUT",
`/collections/${collectionName()}/points`,
{
points: [{ id: qdrantPointId(cacheKey), vector, payload }],
wait: true,
},
30_000,
);
return true;
} catch (error) {
log.warn(
@@ -1,10 +1,11 @@
import Redis from "ioredis";
import { createChildLogger } from "@/shared/logger/index";
import { createAbortControllerWithTimeout } from "@/shared/utils/index";
import { config } from "../../shared/config/config.js";
const log = createChildLogger("searxng-search");
const SEARXNG_BASE_URL = "https://searxng.imrnes.team";
const SEARXNG_BASE_URL = config.SEARXNG_BASE_URL;
const MAX_RESULTS = 3;
const TIMEOUT_MS = 8000;
const CACHE_TTL = 86400; // 24 hours
@@ -12,6 +13,42 @@ const CACHE_PREFIX = "searxng:";
let redis: Redis | null = null;
/**
* Exposes the shared SearXNG Redis connection so other modules (e.g. the
* term glossary) reuse the same connection and cache prefix instead of
* opening their own. Returns null when Redis is unavailable.
*/
export function getSearxngRedis(): Redis | null {
return redis;
}
/** Builds a namespaced SearXNG cache key (shared across modules). */
export function makeSearxngCacheKey(namespace: string, key: string): string {
return `${CACHE_PREFIX}${namespace}:${key.toLowerCase().trim()}`;
}
/** Reads a value from the SearXNG Redis cache; null on miss/unavailable. */
export async function searxngCacheGet(key: string): Promise<string | null> {
if (!redis) return null;
try {
return await redis.get(key);
} catch {
return null;
}
}
/** Writes a value to the SearXNG Redis cache, fire-and-forget. */
export function searxngCacheSet(
key: string,
value: string,
ttlSeconds: number,
): void {
if (!redis) return;
redis.setex(key, ttlSeconds, value).catch(() => {
// Cache write failed silently
});
}
/**
* Initialize Redis connection for SearXNG cache.
* Safe to call multiple times only creates one connection.
@@ -51,19 +88,26 @@ export interface SearxngResult {
/**
* Search SearXNG for a query and return structured results.
* Uses Redis cache when available same query within 24h returns cached results.
*
* @param engines Optional comma-separated SearXNG engine list to constrain
* the search (e.g. "wikipedia"). When set, results are cached under a
* separate cache namespace so engine-specific results never collide.
*/
export async function searchSearxng(
query: string,
category: "general" | "news" | "science" = "general",
engines?: string,
timeoutMs: number = TIMEOUT_MS,
): Promise<SearxngResult[]> {
const cacheKey = `${CACHE_PREFIX}${category}:${query.toLowerCase().trim()}`;
const engineNs = engines ? `eng:${engines}` : "auto";
const cacheKey = makeSearxngCacheKey(`${category}:${engineNs}`, query);
// Try cache first
if (redis) {
try {
const cached = await redis.get(cacheKey);
if (cached) {
log.debug({ query, category }, "SearXNG cache HIT");
log.debug({ query, category, engines }, "SearXNG cache HIT");
return JSON.parse(cached) as SearxngResult[];
}
} catch {
@@ -73,8 +117,11 @@ export async function searchSearxng(
// Cache miss — hit SearXNG API
try {
const url = `${SEARXNG_BASE_URL}/search?q=${encodeURIComponent(query)}&format=json&language=id&categories=${category}`;
const { controller, clear } = createAbortControllerWithTimeout(TIMEOUT_MS);
const engineParam = engines
? `&engines=${encodeURIComponent(engines)}`
: "";
const url = `${SEARXNG_BASE_URL}/search?q=${encodeURIComponent(query)}&format=json&language=id&categories=${category}${engineParam}`;
const { controller, clear } = createAbortControllerWithTimeout(timeoutMs);
try {
const response = await fetch(url, {
@@ -0,0 +1,551 @@
/**
* termGlossary.ts
*
* Per-word "kamus" enrichment for LLM moderation.
*
* Problem: the moderation LLM often meets words it does not know regional
* slang (Jawa/Sunda), foreign terms, niche anime/game jargon, or obscure
* technical vocabulary. When it guesses, it either invents a wrong meaning
* (false positive on a safe word) or misses a violation hidden in unfamiliar
* wording (false negative on an unknown vulgar/slang term).
*
* Solution: extract candidate "unknown-looking" words from message content,
* look each one up on Wikipedia via SearXNG, and inject the definitions into
* the LLM prompt as a `<term_glossary>` block so verdicts are based on facts
* instead of guesses.
*
* Cost control & persistence:
* - successfully resolved definitions are PERSISTED PERMANENTLY in Postgres
* (`term_glossary_cache`) definitions rarely change, so a resolved term
* is never searched again; only misses stay ephemeral (Redis/LRU, 1h);
* - in-memory LRU + Redis (shared with the SearXNG cache) sit in front of
* the DB as fast read caches, so repeat lookups are effectively free;
* - lookups per batch are bounded (AI_GLOSSARY_MAX_TERMS);
* - live SearXNG calls are rate-limit aware: concurrency 2 + stagger, retry
* once on empty results, and misses cached for only 1h so a limiter/
* network blip is not treated as a permanent miss;
* - only results that read like actual definitions are accepted (Wikipedia
* preferred; disambiguation/ads/translate-homepages rejected);
* - everything degrades gracefully: no Redis, no SearXNG, no match
* the block is simply omitted and moderation proceeds as before.
*/
import { LRUCache } from "lru-cache";
import pLimit from "p-limit";
import { createChildLogger } from "@/shared/logger/index";
import { delay } from "@/shared/utils/index";
import { config } from "../../shared/config/config.js";
import { escapeXml } from "./moderationBuilders.js";
import {
makeSearxngCacheKey,
searchSearxng,
searxngCacheGet,
searxngCacheSet,
} from "./searxngSearch.js";
import {
getTermDefinitionFromDb,
setTermDefinitionInDb,
} from "./termGlossaryStore.js";
const log = createChildLogger("term-glossary");
// ---------------------------------------------------------------------------
// Constants
// ---------------------------------------------------------------------------
/** Redis TTL for a successfully resolved definition (definitions are stable). */
const DEF_TTL_SECONDS = 7 * 24 * 60 * 60;
/**
* Redis TTL for a lookup that found nothing. Kept SHORT (1h): SearXNG
* instances silently return empty result sets when rate-limited, so an empty
* response is often a transient failure, not a real miss. A short TTL lets
* the term be retried on a later batch instead of poisoning it for a day.
*/
const MISS_TTL_SECONDS = 60 * 60;
const MISS_TTL_MS = MISS_TTL_SECONDS * 1000;
/** Sentinel stored in caches for "term has no resolvable definition". */
const EMPTY_SENTINEL = "__not_found__";
/** Per-search timeout — keep glossary lookups snappy even on a slow SearXNG. */
const GLOSSARY_SEARCH_TIMEOUT_MS = 5000;
/** Delay before retrying a search that returned zero results. */
const RETRY_DELAY_MS = 350;
/** Max definition snippet length kept in the prompt. */
const MAX_DEFINITION_CHARS = 300;
/**
* SearXNG rate-limits aggressive parallel bursts (returns 200 with empty
* results). Never fire all terms at once cap live searches at 2 concurrent
* and stagger the start times slightly.
*/
const LIVE_SEARCH_CONCURRENCY = 2;
const LIVE_SEARCH_STAGGER_MS = 250;
/** In-memory cache: term (lowercase) → definition | NOT_FOUND sentinel. */
const NOT_FOUND: TermDefinition = {
term: "__not_found__",
definition: "",
sourceUrl: "",
};
const termLru = new LRUCache<string, TermDefinition>({
max: 2000,
ttl: 24 * 60 * 60 * 1000,
});
/** Serializes live SearXNG lookups (rate-limit aware) with a small stagger. */
const liveSearchLimit = pLimit(LIVE_SEARCH_CONCURRENCY);
let lastLiveSearchAt = 0;
async function acquireLiveSlot(): Promise<void> {
const now = Date.now();
const wait = lastLiveSearchAt + LIVE_SEARCH_STAGGER_MS - now;
if (wait > 0) await delay(wait);
lastLiveSearchAt = Date.now();
}
// ---------------------------------------------------------------------------
// Term extraction
// ---------------------------------------------------------------------------
/** Word tokenizer letters/digits plus internal -_'· (handles "well-known",
* "node_modules", diacritics). */
const WORD_RE = /[\p{L}\p{N}]+(?:[-_'·][\p{L}\p{N}]+)*/gu;
/** Removes URLs, Discord mentions/custom emoji, code fences, markdown noise. */
function cleanContent(raw: string): string {
return raw
.replace(/https?:\/\/\S+/gi, " ")
.replace(/<@!?\d+>/g, " ")
.replace(/<#\d+>/g, " ")
.replace(/<a?:\w+:\d+>/g, " ")
.replace(/[`*_~|>[\]]/g, " ")
.replace(/[\p{Emoji}\p{Extended_Pictographic}]/gu, " ")
.replace(/\s+/g, " ")
.trim();
}
/** Filters out tokens that are useless as glossary candidates (numbers,
* repeated-char noise, mega-tokens). */
function isNoiseWord(word: string): boolean {
if (word.length > 28) return true;
if (/^\d+$/.test(word)) return true;
const lower = word.toLowerCase();
// "aaaa…", "wwwwww" — single repeated character
if (/^(.)\1{2,}$/.test(lower)) return true;
// "wkwk", "hehe", "69" alternations — repeated 23 char base. "meme" is
// the one legit 4-letter word this matches; it is whitelisted below.
if (/^([a-z]{2,3})\1{1,}$/.test(lower)) return true;
return false;
}
/** Deterministic bonus for words that look like proper nouns or foreign. */
function scoreWord(word: string): number {
let score = 1;
// Capitalized first letter (proper noun / title) but not ALL-CAPS acronyms
if (/^[A-Z]/.test(word) && !/^[A-Z]{2,}$/.test(word)) score += 3;
// Contains a letter outside basic latin → regional/foreign spelling
if (/[\p{L}]/u.test(word.replace(/[A-Za-z]/g, ""))) score += 2;
// Contains an internal apostrophe or hyphen → likely a named entity
if (/[-_'’·]/.test(word)) score += 2;
return score;
}
const STOPWORDS = new Set(
// ── Bahasa Indonesia ────────────────────────────────────────────────
(
" yang dan di ke dari ini itu dengan untuk pada dalam adalah akan telah sudah bisa dapat harus tidak juga saya kamu kita kami mereka dia aku kau gua lu lo gw gue elu anda kalian nya kah lah pun ya yah kan sih dong deh kok loh toh aja saja gitu gini begitu begini tapi tetapi namun atau karena sebab jika kalau bila maka supaya agar meski meskipun walau walaupun ketika saat setelah sebelum selama antara terhadap tentang mengenai bagi oleh secara sebagai seperti daripada tanpa hingga sampai sejak menuju bahwa padahal sebenarnya sepertinya mungkin memang jadi lalu terus akhirnya misalnya contohnya banyak sedikit semua seluruh setiap tiap beberapa ada bukan jangan boleh mau ingin pengen nggak ngak gak ga kagak ngga ndak nanti kemarin besok hari ini sekarang waktu itu masih sedang belum pernah sering selalu kadang jarang cepat lambat awal akhir baru lama besar kecil tinggi rendah panjang pendek baik buruk benar salah sama beda penting biasanya selamat terima kasih makasih sangat sekali paling cuma cuman hanya lebih kurang sekitar hampir ternyata rupanya begitu gimana bagaimana kenapa mengapa siapa apa mana kapan darimana kemana bilang ngomong omong kata tadi dulu terus lagi tetap pasti seharusnya sebaiknya seakan seolah kayaknya keliatan kelihatan ketahuan disini disitu disana kesini kesana bener pake pakai kayak emang lagian mulu istilah istilahnya banget" +
// ── English ───────────────────────────────────────────────────────
" the a an and or but if then else for to in on at by with without from of is are was were be been being have has had do does did will would can could should may might must shall this that these those it its i you he she we they them their there here when where why how what which who whom whose only very just about above after before below under over into onto within upon against between among during through across along around behind beyond near off out up down now then so as not no yes ok okay" +
// ── Common net slang / acronyms the LLM already knows ──────────────
" lol omg wtf idk btw tbh imo aka fyi nsfw smh nvm asap afk brb gg wp ty np mb sry thx kk oke okk ygy frfr"
).split(/\s+/),
);
/**
* Words that are either already defined by the moderation rules, or are so
* common (brands, tech vocabulary, project names) that a Wikipedia lookup is
* a guaranteed miss/waste. Keeps the glossary focused on genuinely unknown
* terms.
*/
const KNOWN_SAFE_TERMS = new Set(
(
"discord youtube google facebook instagram twitter tiktok whatsapp telegram netflix spotify steam github gitlab bitbucket chatgpt openai anthropic claude deepseek gemini llama copilot cursor vscode vscodium jetbrains intellij pycharm webstorm sublime codeblocks" +
" docker kubernetes k8s linux ubuntu debian arch fedora manjaro kali windows macos android ios chrome firefox safari edge opera brave" +
" react nextjs next vue svelte angular node nodejs deno bun pnpm yarn npm javascript typescript python golang go rust java kotlin swift cplusplus cpp css html json xml yaml toml regex backend frontend database mysql postgres postgresql mongodb redis qdrant sqlite nosql graphql rest websocket webhook" +
" bug crash error debug fix issue pr merge commit push pull branch main master dev staging production server client app website web browser" +
" stream streaming video audio voice call camera screen share screenshare gameplay gaming game play steam epic xbox playstation nintendo switch console" +
" bot discordbot moderation moderator admin member user profile avatar channel server guild message chat dm reply forward embed sticker emoji role permission" +
" meme code coding ngoding programmer program developer engineer software hardware cpu gpu ram rom storage disk network internet wifi lan ip dns vpn proxy cloud aws azure gcp vercel netlify heroku railway render vps hosting domain ssl login logout register account password email username" +
" anime manga waifu husbando tsundere moe otaku wibu weeb otome isekai shonen seinen josei manga manhwa manhua doujin" +
" anjay wkwk wkwkwk gws gaskeun santuy njir baka woy woi hadeh astaga asu anjing bangsat ngehe asal alay lebay caper mabar" +
" asus bete imphnen impnhen ngab" +
" syahadat sholat shalat solat puasa zakat haji umrah doa tuhan nabi allah yesus muhammad hashem" +
" loli shota incest exhibition furry fursuit cosplay costume" +
" gaza palestine israel yahudi yahud israel palestina israeli" +
" hokkian mandarin arabic jawa sunda betawi minang bugis batak melayu inggris indonesia"
).split(/\s+/),
);
function isKnownTerm(word: string): boolean {
return STOPWORDS.has(word) || KNOWN_SAFE_TERMS.has(word);
}
/** True when a quoted phrase is mostly filler words (skip it). */
function isMostlyStopwords(phrase: string): boolean {
const words = phrase
.toLowerCase()
.split(/[^a-zà-öø-ÿ]+/i)
.filter(Boolean);
if (words.length === 0) return true;
const stopCount = words.filter((w) => STOPWORDS.has(w)).length;
return stopCount / words.length >= 0.6;
}
export interface ExtractGlossaryOptions {
maxTerms?: number;
minWordLength?: number;
}
/**
* Extracts candidate terms that the LLM might not know from message content.
* Returns at most `maxTerms` terms (default from config), scored by how
* "unknown-looking" they are (proper nouns, foreign spelling, quoted phrases).
*/
export function extractGlossaryTerms(
contents: string[],
options: ExtractGlossaryOptions = {},
): string[] {
const maxTerms = options.maxTerms ?? config.AI_GLOSSARY_MAX_TERMS;
const minWordLength =
options.minWordLength ?? config.AI_GLOSSARY_MIN_WORD_LENGTH;
const candidates = new Map<string, { word: string; score: number }>();
const push = (rawWord: string, score: number): void => {
const clean = rawWord
.trim()
.replace(/^[^\p{L}\p{N}]+|[^\p{L}\p{N}]+$/gu, "");
if (clean.length < minWordLength) return;
const key = clean.toLowerCase();
if (isKnownTerm(key) || isNoiseWord(clean)) return;
const existing = candidates.get(key);
if (existing) {
existing.score += score + 1;
} else {
candidates.set(key, { word: clean, score });
}
};
for (const content of contents) {
if (!content) continue;
const cleaned = cleanContent(content);
if (!cleaned) continue;
// Quoted phrases — explicit terms the user called out
for (const m of cleaned.matchAll(/"([^"]{2,80})"/g)) {
const phrase = m[1].trim();
const wordCount = phrase.split(/\s+/).length;
if (wordCount >= 2 && wordCount <= 6 && !isMostlyStopwords(phrase)) {
push(phrase, 10);
}
}
// Individual words
for (const m of cleaned.matchAll(WORD_RE)) {
const w = m[0];
if (w.length < minWordLength) continue;
if (isNoiseWord(w)) continue;
const key = w.toLowerCase();
if (isKnownTerm(key)) continue;
push(w, scoreWord(w));
}
}
return Array.from(candidates.values())
.sort((a, b) => b.score - a.score)
.slice(0, maxTerms)
.map((c) => c.word);
}
// ---------------------------------------------------------------------------
// Definition lookup (cached: LRU → Redis → SearXNG/Wikipedia)
// ---------------------------------------------------------------------------
export interface TermDefinition {
term: string;
definition: string;
sourceUrl: string;
}
/** Definition-like markers for accepting a non-Wikipedia search result. */
const DEF_MARKERS =
/adalah|merupakan|istilah (?:untuk|yang|yg)|artinya|sebutan|berarti|refers? to|known as|also called|short for|a term (?:for|used)|istilah dalam|kata (?:asing|serapan)? ?untuk/i;
/** True when the term appears in the result text (or a 4+ char word in the
* result is part of the term). Lenient "kafircel" matches a "Kafir"
* article via substring, while a Google-Translate homepage snippet does not. */
function hasTermOverlap(term: string, title: string, snippet: string): boolean {
const termLower = term.toLowerCase();
const text = `${title} ${snippet}`.toLowerCase();
if (text.includes(termLower)) return true;
const words = text.match(/[a-z0-9]{4,}/gi) ?? [];
return words.some((w) => termLower.includes(w));
}
/** Quality gate: is this result good enough to quote as a definition? */
function isUsableDefinition(
r: { title: string; url: string; snippet: string },
term: string,
isWiki: boolean,
): boolean {
const text = `${r.title} ${r.snippet}`;
// Wikipedia disambiguation pages are not definitions
if (/disambiguasi|disambiguation/i.test(text)) return false;
if ((r.snippet ?? "").trim().length < 25) return false;
if (!hasTermOverlap(term, r.title, r.snippet)) return false;
// Wikipedia articles are accepted with just the overlap+length gate;
// everything else must read like an actual definition, not an ad,
// a translate homepage, or a navigation blurb.
if (isWiki) return true;
return DEF_MARKERS.test(r.snippet);
}
/** Picks the best definition from search results, preferring a genuine
* Wikipedia article; otherwise the first result that reads like a
* definition. Returns null when nothing qualifies. */
function pickDefinition(
results: Array<{ title: string; url: string; snippet: string }>,
term: string,
): TermDefinition | null {
const wiki = results.find((r) => /wikipedia\.org/i.test(r.url));
const best = wiki && isUsableDefinition(wiki, term, true) ? wiki : null;
if (!best) {
for (const r of results) {
if (isUsableDefinition(r, term, false)) {
return buildDefinition(r, term);
}
}
return null;
}
return buildDefinition(best, term);
}
function buildDefinition(
best: { title: string; url: string; snippet: string },
term: string,
): TermDefinition {
const snippet = (best.snippet || best.title || "").trim();
const definition =
snippet.length > MAX_DEFINITION_CHARS
? `${snippet.slice(0, MAX_DEFINITION_CHARS - 1).trimEnd()}`
: snippet;
return { term, definition, sourceUrl: best.url };
}
/** Live (network) lookup — runs under the shared SearXNG rate-limit gate. */
async function fetchDefinitionLive(
term: string,
key: string,
cacheKey: string,
): Promise<TermDefinition | null> {
return liveSearchLimit(async () => {
await acquireLiveSlot();
try {
let results = await searchSearxng(
key,
"general",
undefined,
GLOSSARY_SEARCH_TIMEOUT_MS,
);
let def = pickDefinition(results, term);
// Zero results is usually the limiter kicking in, not a real miss —
// retry once. Results-but-unusable = genuine miss, no retry.
if (!def && results.length === 0) {
await delay(RETRY_DELAY_MS);
results = await searchSearxng(
key,
"general",
undefined,
GLOSSARY_SEARCH_TIMEOUT_MS,
);
def = pickDefinition(results, term);
}
if (def) {
// Persist permanently (definitions rarely change) — best-effort,
// then warm the fast caches.
void setTermDefinitionInDb(key, def.definition, def.sourceUrl);
searxngCacheSet(
cacheKey,
JSON.stringify({
definition: def.definition,
sourceUrl: def.sourceUrl,
}),
DEF_TTL_SECONDS,
);
termLru.set(key, def);
log.debug({ term: key }, "Term glossary resolved definition");
return def;
}
} catch (err) {
log.debug(
{ term: key, error: err instanceof Error ? err.message : String(err) },
"Term glossary lookup failed — skipping term",
);
}
// No definition — cache the miss with a SHORT TTL so a transient
// limiter/network failure is retried on a later batch.
searxngCacheSet(cacheKey, EMPTY_SENTINEL, MISS_TTL_SECONDS);
termLru.set(key, NOT_FOUND, { ttl: MISS_TTL_MS });
return null;
});
}
/** Resolve one term: LRU Redis Postgres (permanent) live SearXNG
* (rate-limited). The fast caches sit in front of the DB; the DB is the
* source of truth for successfully resolved definitions. */
async function resolveTerm(term: string): Promise<TermDefinition | null> {
const key = term.toLowerCase().trim();
// 1. In-memory LRU — same process, instant
const lruHit = termLru.get(key);
if (lruHit) return lruHit === NOT_FOUND ? null : lruHit;
// 2. Redis — shared across processes/workers. A miss sentinel here is NOT
// a definitive answer: it may predate a permanent DB entry written by
// another process, so we keep going and let the DB decide.
const cacheKey = makeSearxngCacheKey("def", key);
const cached = await searxngCacheGet(cacheKey);
let redisMiss = false;
if (cached !== null) {
if (cached === EMPTY_SENTINEL) {
redisMiss = true;
} else {
try {
const parsed = JSON.parse(cached) as {
definition?: string;
sourceUrl?: string;
};
if (parsed.definition) {
const def: TermDefinition = {
term,
definition: parsed.definition,
sourceUrl: parsed.sourceUrl ?? "",
};
termLru.set(key, def);
return def;
}
} catch {
// malformed cache entry — fall through to DB/live
}
}
}
// 3. Postgres — permanent store for resolved definitions. A hit re-warms
// the fast caches so the DB is not hit on every batch.
const dbDef = await getTermDefinitionFromDb(key);
if (dbDef) {
const def: TermDefinition = {
term,
definition: dbDef.definition,
sourceUrl: dbDef.sourceUrl,
};
termLru.set(key, def);
searxngCacheSet(
cacheKey,
JSON.stringify({ definition: def.definition, sourceUrl: def.sourceUrl }),
DEF_TTL_SECONDS,
);
log.debug({ term: key }, "Term glossary DB hit");
return def;
}
// 4. Redis already said "miss" recently and the DB has nothing — respect
// that instead of hammering SearXNG again within the miss window.
if (redisMiss) {
termLru.set(key, NOT_FOUND, { ttl: MISS_TTL_MS });
return null;
}
// 5. Live search (rate-limited + staggered)
return fetchDefinitionLive(term, key, cacheKey);
}
/**
* Looks up definitions for a batch of terms, in parallel. Returns a map of
* term definition for the terms that resolved. Errors/misses are skipped.
* Live SearXNG calls are throttled internally (concurrency 2 + stagger).
*/
export async function lookupTermDefinitions(
terms: string[],
): Promise<Map<string, TermDefinition>> {
const map = new Map<string, TermDefinition>();
if (terms.length === 0) return map;
const results = await Promise.allSettled(terms.map(resolveTerm));
for (let i = 0; i < terms.length; i++) {
const r = results[i];
if (r.status === "fulfilled" && r.value) {
map.set(r.value.term, r.value);
}
}
return map;
}
// ---------------------------------------------------------------------------
// Prompt formatting
// ---------------------------------------------------------------------------
/**
* Formats definitions as a `<term_glossary>` XML block for the LLM prompt:
*
* <term_glossary>
* <term word="ngab" source="https://…">definisi</term>
* </term_glossary>
*
* Returns "" when there are no definitions (the block is then omitted).
*/
export function formatTermGlossary(
defs: ReadonlyMap<string, TermDefinition>,
): string {
if (!defs || defs.size === 0) return "";
const lines = Array.from(defs.values()).map(
(d) =>
` <term word="${escapeXml(d.term)}" source="${escapeXml(d.sourceUrl)}">${escapeXml(d.definition)}</term>`,
);
return `<term_glossary>\n${lines.join("\n")}\n</term_glossary>`;
}
// ---------------------------------------------------------------------------
// Convenience: full pipeline
// ---------------------------------------------------------------------------
export interface GlossaryBlockOptions extends ExtractGlossaryOptions {
enabled?: boolean;
}
/**
* One-shot helper: extract terms from message contents, look up definitions,
* and return the formatted `<term_glossary>` block ("" when disabled or no
* definitions found). Safe to call on every batch cached lookups make it
* cheap.
*/
export async function buildTermGlossaryBlock(
contents: string[],
options: GlossaryBlockOptions = {},
): Promise<string> {
const enabled = options.enabled ?? config.AI_GLOSSARY_ENABLED;
if (!enabled) return "";
if (contents.length === 0) return "";
const terms = extractGlossaryTerms(contents, options);
if (terms.length === 0) return "";
const defs = await lookupTermDefinitions(terms);
if (defs.size === 0) return "";
const block = formatTermGlossary(defs);
log.debug(
{ terms: terms.length, definitions: defs.size },
"Term glossary block built",
);
return block;
}
@@ -0,0 +1,86 @@
/**
* termGlossaryStore.ts
*
* Permanent Postgres layer for the term glossary. Resolved definitions
* (which carry content) are persisted here because they rarely change
* Redis/LRU only act as fast read caches in front of this table. Terms with
* no definition (misses) are deliberately NOT persisted; they stay ephemeral
* in Redis with a short TTL so transient lookup failures get retried.
*
* All calls are best-effort: any DB error degrades to a cache miss (the
* glossary then falls through to Redis/live search as if the DB layer
* didn't exist).
*/
import { createChildLogger } from "@/shared/logger/index";
import { executeAll, executeGet } from "../../shared/database/drizzle.js";
const log = createChildLogger("term-glossary-store");
export interface StoredTermDefinition {
definition: string;
sourceUrl: string;
}
/**
* Read a permanently stored definition for a term (lowercase key).
* Returns null when missing or on any DB error (callers fall through).
* A successful read bumps hit_count for observability (fire-and-forget).
*/
export async function getTermDefinitionFromDb(
term: string,
): Promise<StoredTermDefinition | null> {
try {
const row = await executeGet(
`SELECT definition, source_url FROM term_glossary_cache WHERE term = $1`,
[term.toLowerCase().trim()],
);
if (!row) return null;
try {
await executeAll(
`UPDATE term_glossary_cache SET hit_count = hit_count + 1 WHERE term = $1`,
[term.toLowerCase().trim()],
);
} catch {
// hit_count is observability only — never fail a read for it
}
return {
definition: row.definition as string,
sourceUrl: (row.source_url as string | null) ?? "",
};
} catch (error) {
log.debug(
{ error: error instanceof Error ? error.message : String(error) },
"getTermDefinitionFromDb failed — falling back to live search",
);
return null;
}
}
/**
* Persist a resolved definition permanently (UPSERT by term).
* Only called for successful resolutions never for misses.
* Best-effort: a DB write failure does not affect the returned definition.
*/
export async function setTermDefinitionInDb(
term: string,
definition: string,
sourceUrl: string,
): Promise<void> {
try {
await executeAll(
`INSERT INTO term_glossary_cache (term, definition, source_url, resolved_at, hit_count)
VALUES ($1, $2, $3, $4, 0)
ON CONFLICT (term) DO UPDATE SET
definition = EXCLUDED.definition,
source_url = EXCLUDED.source_url,
resolved_at = EXCLUDED.resolved_at`,
[term.toLowerCase().trim(), definition, sourceUrl, Date.now()],
);
} catch (error) {
log.warn(
{ error: error instanceof Error ? error.message : String(error) },
"setTermDefinitionInDb failed — definition stays memory/Redis only",
);
}
}
@@ -19,11 +19,7 @@ import { callModerationLLM } from "./llmCaller.js";
import { analyzeSingleMediaImage } from "./mediaAnalysisClient.js";
import {
buildReferenceXml,
buildUserHistoryXml,
buildUserProfileRef,
buildUserProfilesBlock,
escapeXml,
formatReputationAttrs,
getAnalysisContent,
resolveDisplayName,
resolveIsBot,
@@ -37,13 +33,9 @@ import {
formatSearchResults,
searchSearxng,
} from "./searxngSearch.js";
import { buildTermGlossaryBlock } from "./termGlossary.js";
import { getRecentCorrectedModerations } from "./textCacheStore.js";
import { extractUrlsFromText, fetchUrlSafely } from "./urlFetcher.js";
import { getUserProfile } from "./userProfileStore.js";
import {
getUserRecentInfractions,
initializeUserReputation,
} from "./userReputationStore.js";
import type { MessageImagePart } from "./visionAnalyzer.js";
const log = createChildLogger("textBatchProcessor");
@@ -145,9 +137,17 @@ export async function runTextOnlyBatch(
return map;
})();
const [urlFetchMaps, searxngResults] = await Promise.all([
// Term glossary — per-word Wikipedia lookups for words the LLM may not
// know (slang, jargon, regional language). Cached in Redis + in-memory, so
// repeat terms resolve instantly and only genuinely new words hit SearXNG.
const glossaryPromise = buildTermGlossaryBlock(
targets.map((msg) => getAnalysisContent(msg)),
).catch(() => "");
const [urlFetchMaps, searxngResults, glossaryBlock] = await Promise.all([
urlFetchPromise,
searxngPromise,
glossaryPromise,
]);
const urlFetchMap = urlFetchMaps.text;
@@ -190,57 +190,19 @@ export async function runTextOnlyBatch(
? await getChannelCulture(channelId)
: null;
const channelCulture = channelCultureObj?.culture_summary;
// Corrected false-positive examples are static per batch — fetch ONCE
// here instead of inside the per-sub-batch retry closure (which would
// re-query the DB on every sub-batch and every parse-error retry).
const correctedExamples = await buildCorrectedFewShotExamples();
for (let i = 0; i < subBatches.length; i++) {
const batch = subBatches[i];
const targetIds = batch.map((t) => t.id);
// User reputation + profiles (raw summary text — deduplicated into a
// single <user_profiles> map per batch; messages only reference it).
const userContexts = new Map<string, string>();
const userProfiles = new Map<
string,
{
text: string;
asOf?: number | null;
}
>();
for (const msg of batch) {
if (!userContexts.has(msg.user_id)) {
const rep = await initializeUserReputation(msg.user_id, msg.guild_id);
const repAttrs = formatReputationAttrs(rep);
let repXml = `<user_reputation ${repAttrs}/>`;
// Repeat offenders get their last flagged messages as <user_history>
// so the LLM can recognize PATTERNS (same scam link, repeated
// provocation) — history is reference, never proof. Best-effort.
if (rep.total_infractions > 0) {
try {
const history = await getUserRecentInfractions(msg.user_id, 2);
const historyXml = buildUserHistoryXml(
history.map((h) => ({
content: h.content ?? "",
severity: h.severity,
created_at: h.created_at,
})),
);
if (historyXml) {
repXml = `<user_reputation ${repAttrs}>\n${historyXml}\n</user_reputation>`;
}
} catch {
// history is a bonus — fall back to attrs-only reputation
}
}
userContexts.set(msg.user_id, repXml);
}
if (!userProfiles.has(msg.user_id)) {
const profile = await getUserProfile(msg.user_id);
userProfiles.set(msg.user_id, {
text: profile?.profile_summary ?? "",
asOf: profile?.last_analyzed_at ?? null,
});
}
}
const userProfilesBlock = buildUserProfilesBlock(userProfiles);
// No per-user reputation/profile context is injected into the prompt —
// the user asked to keep the AI analysis context minimal (raw messages
// only). Trust/infraction state is still tracked in the DB for
// enforcement, just not shown to the LLM.
// ── URL images → multimodal vision evidence ─────────────────────────
// The text batch fetches inline URLs; whenever one resolved to an image
@@ -264,7 +226,8 @@ export async function runTextOnlyBatch(
if (pics.length === 0) return { id: msg.id, lines: [] as string[] };
const lines = await Promise.all(
pics.map(async (url) => {
const img = urlImages.get(url)!;
const img = urlImages.get(url);
if (!img) return null;
try {
const { data: resizedBuffer, mimeType: resizedMime } =
await resizeImageForVision(img.data, maxDim);
@@ -310,7 +273,6 @@ export async function runTextOnlyBatch(
preview: state.lastInvalidContent?.slice(0, 800) ?? "<empty>",
}
: undefined;
const correctedExamples = await buildCorrectedFewShotExamples();
const systemText = buildSystemPromptModular({
mode: batchHasImageEvidence ? "mixed" : "text",
correction,
@@ -337,17 +299,11 @@ export async function runTextOnlyBatch(
const mediaEvidenceCtx = (batchImageEvidence.get(msg.id) ?? [])
.map((line) => `\n${line}`)
.join("");
const userCtx = userContexts.get(msg.user_id) ?? "";
const userProfileRef = (
userProfiles.get(msg.user_id)?.text ?? ""
).trim()
? buildUserProfileRef(msg.user_id)
: "";
const refXml = await buildReferenceXml(msg);
const repetitionCount = groupMapping.get(msg.id)?.length ?? 1;
const isBot = resolveIsBot(msg);
const isEdited = resolveIsEdited(msg);
return `<message id="${escapeXml(msg.id)}" user="${escapeXml(resolveDisplayName(msg))}" time="${new Date(msg.created_at).toISOString()}"${repetitionCount > 1 ? ` repetitions="${repetitionCount}"` : ""}${isBot ? ` bot="true"` : ""}${isEdited ? ` edited="true"` : ""}>\n ${userCtx}${userProfileRef ? `\n ${userProfileRef}` : ""}${refXml ? `\n ${refXml}` : ""}\n <content>${escapeXml(content)}</content>${webContext}${mediaEvidenceCtx}\n</message>`;
return `<message id="${escapeXml(msg.id)}" user="${escapeXml(resolveDisplayName(msg))}" time="${new Date(msg.created_at).toISOString()}"${repetitionCount > 1 ? ` repetitions="${repetitionCount}"` : ""}${isBot ? ` bot="true"` : ""}${isEdited ? ` edited="true"` : ""}>\n ${refXml ? `\n ${refXml}` : ""}\n <content>${escapeXml(content)}</content>${webContext}${mediaEvidenceCtx}\n</message>`;
}),
)
).join("\n");
@@ -362,12 +318,13 @@ export async function runTextOnlyBatch(
.join("\n")}\n</web_searches>`
: "";
// Data/instruction separation: the system prompt is stable per mode —
// all per-batch context (profiles, conversation, web evidence) lives in
// the USER payload, ordered oldest-first so targets come last.
// all per-batch context (conversation, web evidence) lives in the USER
// payload, ordered oldest-first so targets come last. Personal user
// profile descriptions are intentionally omitted (see above).
const userBlocks = [
userProfilesBlock?.trimEnd() ?? "",
contextBlock?.trimEnd() ?? "",
searxngBlock,
glossaryBlock,
`<messages_to_analyze>\n${messagesBlock}\n</messages_to_analyze>`,
].filter((b) => b.trim().length > 0);
return {
@@ -3,7 +3,6 @@ import { createChildLogger } from "@/shared/logger/index";
import { executeAll, executeGet } from "../../shared/database/drizzle.js";
import { findBestEmbeddingMatch } from "./embeddingClient.js";
import {
deleteExpiredQdrantPoints,
deleteQdrantPoint,
deleteQdrantPointsByContentHash,
isQdrantConfigured,
@@ -62,13 +61,30 @@ export function makeCustomEmojiCacheKey(emojiId: string): string {
}
/**
* Generate a deterministic cache key for an image data URL.
* Hashes the first 128 chars of the data URL (enough to identify the image
* without storing the full base64 string as the key).
* Generate a deterministic cache key for an image from its source URL
* (Discord CDN / embed URL / inline URL).
*
* The CDN URL is the stable identity of an attachment: re-analysis of the
* same message (recovery worker, retries) always hits the cache regardless
* of resize/encoding output. Query params are stripped (Discord signed
* tokens `?ex=&is=&hm=` and render variants `?format=&width=`) so the same
* attachment resolves to the same key even when fetched with different
* signatures or sizes.
*
* No SHA/phash the CDN URL is the cache key itself. This makes
* re-analysis of the SAME attachment cache-hit, while different attachments
* (different URLs) never collide.
*/
export function makeImageCacheKey(dataUrl: string): string {
const prefix = dataUrl.slice(0, 128);
const hash = createHash("sha256").update(prefix).digest("hex").slice(0, 16);
export function makeImageCacheKey(imageUrl: string): string {
// Hash the URL to a fixed-length key. The raw Discord CDN URL is short,
// but callers sometimes pass base64 data URLs (can be multi-MB) or very
// long signed/external URLs. text_analysis_cache.text is the PK and lives
// in a B-tree index with an 8191-byte per-row limit — inserting a long URL
// as the key aborts the whole INSERT ("index row requires N bytes, maximum
// size is 8191"), which fails acquireMediaAnalysisLock and silently skips
// every media analysis. A 32-char sha256 keeps the key well under the limit
// and is still deterministic (same attachment → same key).
const hash = createHash("sha256").update(imageUrl).digest("hex").slice(0, 32);
return `image:${hash}`;
}
@@ -515,64 +531,6 @@ export async function setCachedTextModeration(
}
}
// ---------------------------------------------------------------------------
// Perceptual hash helpers for image deduplication
// ---------------------------------------------------------------------------
/**
* Generate a deterministic cache key for a perceptual hash.
* The phash value is a string like "a1b2c3d4e5f6..." from the imghash library.
*/
export function makePhashCacheKey(phash: string): string {
return `phash:${phash.slice(0, 16)}`;
}
/**
* Look up a cached media analysis by perceptual hash.
* Returns the cached analysis string or null if not found/expired.
*/
export async function getCachedMediaByPhash(
phash: string,
): Promise<string | null> {
const cacheKey = makePhashCacheKey(phash);
return getCachedMediaAnalysis(cacheKey);
}
/**
* Store a media analysis result keyed by perceptual hash.
*/
export async function upsertCachedMediaByPhash(
phash: string,
analysisResult: string,
source: "vision_llm",
expiresAt: number,
): Promise<void> {
const cacheKey = makePhashCacheKey(phash);
return upsertCachedMediaAnalysis(cacheKey, analysisResult, source, expiresAt);
}
/**
* Compute perceptual hash from image buffer using imghash.
* Returns a hexadecimal string representation of the hash.
* Returns null if hashing fails (e.g., invalid image data).
*/
export async function computeImagePhash(
buffer: Buffer,
): Promise<string | null> {
try {
// Dynamic import — imghash is ESM with a default export containing { hash, hashRaw, ... }
const imghashModule: {
default?: { hash?: (buf: Buffer) => Promise<string> };
} = await import("imghash");
const hashFn = imghashModule.default?.hash;
if (typeof hashFn !== "function") return null;
const hash = await hashFn(buffer);
return hash;
} catch {
return null;
}
}
// ---------------------------------------------------------------------------
// Corrected Moderation (false-positive) helpers for dynamic few-shot injection
// ---------------------------------------------------------------------------
@@ -16,17 +16,14 @@ import type {
import { llmVision } from "./llmClient.js";
import {
acquireMediaAnalysisLock,
computeImagePhash,
deleteCachedMediaAnalysis,
FAILED_ANALYSIS_PREFIX,
getCachedMediaAnalysis,
getCachedMediaByPhash,
inFlightVisionCalls,
makeCustomEmojiCacheKey,
makeImageCacheKey,
makeStickerCacheKey,
upsertCachedMediaAnalysis,
upsertCachedMediaByPhash,
visionLruCache,
} from "./mediaCache.js";
@@ -66,10 +63,7 @@ import {
} from "./mediaDownloader.js";
import {
buildReferenceXml,
buildUserHistoryXml,
buildUserProfileRef,
escapeXml,
formatReputationAttrs,
getAnalysisContent,
resolveDisplayName,
resolveIsBot,
@@ -87,12 +81,8 @@ import {
formatSearchResults,
searchSearxng,
} from "./searxngSearch.js";
import { buildTermGlossaryBlock } from "./termGlossary.js";
import { extractUrlsFromText } from "./urlFetcher.js";
import { getUserProfile } from "./userProfileStore.js";
import {
getUserRecentInfractions,
initializeUserReputation,
} from "./userReputationStore.js";
// ---------------------------------------------------------------------------
// Types
@@ -162,7 +152,10 @@ export const analyzeSingleMediaImage = async (
const cached = await getCachedMediaAnalysis(cacheKey);
if (cached && !isNoImageSeenText(cached)) {
visionLruCache.set(cacheKey, cached);
log.debug({ cacheKey }, "Media analysis cache HIT (DB → LRU)");
log.debug(
{ cacheKey, messageId, cachedLen: cached.length },
"Media analysis cache HIT (DB → LRU)",
);
return `[Media analysis for message ${messageId}] ${image.sourceLabel}: ${cached}`;
}
if (cached) {
@@ -206,45 +199,19 @@ export const analyzeSingleMediaImage = async (
return FAILED_ANALYSIS_PREFIX;
}
// phash check
let phash: string | null = null;
if (image.image_url.url.startsWith("data:")) {
try {
const base64Data = image.image_url.url.split(",")[1];
if (base64Data) {
const imgBuffer = Buffer.from(base64Data, "base64");
phash = await computeImagePhash(imgBuffer);
if (phash) {
const phashCached = await getCachedMediaByPhash(phash);
if (phashCached && !isNoImageSeenText(phashCached)) {
visionLruCache.set(cacheKey, phashCached);
await upsertCachedMediaAnalysis(
cacheKey,
phashCached,
"vision_llm",
Date.now() + 24 * 60 * 60 * 1000,
).catch(() => {});
return phashCached;
}
if (phashCached) {
log.warn(
{ phash, cacheKey },
"phash cache HIT was no-image-seen — ignoring",
);
}
}
}
} catch {
phash = null;
}
}
// Vision API call
let lastError: Error | null = null;
for (let attempt = 0; attempt < 3; attempt++) {
try {
const content = await llmVision(promptText, image.image_url);
if (content && !isNoImageSeenText(content)) {
// Defensive: log when a vision analysis is cached so we can trace
// if the SAME analysis text is being stored for DIFFERENT cache keys
// (which would indicate the vision model is returning duplicates).
log.debug(
{ cacheKey, messageId, contentLen: content.length },
"Vision analysis cached (new entry)",
);
await upsertCachedMediaAnalysis(
cacheKey,
content,
@@ -252,20 +219,11 @@ export const analyzeSingleMediaImage = async (
Date.now() + 24 * 60 * 60 * 1000,
);
visionLruCache.set(cacheKey, content);
if (phash) {
upsertCachedMediaByPhash(
phash,
content,
"vision_llm",
Date.now() + 7 * 24 * 60 * 60 * 1000,
).catch(() => {});
}
return content;
}
if (content) {
// Model claims it saw no image — same as a null response: NOT a
// valid analysis, and caching it would poison the key for every
// re-analysis of the same image (phash TTL is 7 days).
// valid analysis, and caching it would poison the cache key.
log.warn(
{ messageId, cacheKey },
"Vision returned no-image-seen text — not caching",
@@ -413,6 +371,12 @@ export async function prepareMediaMessage(
searxngXml = `\n<web_searches>\n${parts.join("\n")}\n</web_searches>`;
}
// Term glossary — cached per-word Wikipedia definitions for words the LLM
// may not know. Bounded and cached (in-memory + Redis), so this adds no
// meaningful latency to the media path either.
const glossaryXml = await buildTermGlossaryBlock([content]).catch(() => "");
const glossaryCtx = glossaryXml ? `\n${glossaryXml}` : "";
// Build XML block
const webTexts = webTextMap.get(targetId) ?? [];
const mediaAnalyses = mediaAnalysisMap.get(targetId) ?? [];
@@ -432,40 +396,13 @@ export async function prepareMediaMessage(
.filter(Boolean)
.join(" ");
const rep = await initializeUserReputation(target.user_id, target.guild_id);
const profile = await getUserProfile(target.user_id);
const refXml = await buildReferenceXml(target);
// Profile is emitted ONCE per batch in a <user_profiles> map (see
// mediaBatchProcessor); here we only reference it to avoid repeating the
// full summary on every message of the same user.
const profileRef = profile?.profile_summary?.trim()
? buildUserProfileRef(target.user_id)
: "";
// Rich reputation — same shape as the text path: attrs + optional
// <user_history> with the last flagged messages for repeat offenders.
const repAttrs = formatReputationAttrs(rep);
let repXml = `<user_reputation ${repAttrs}/>`;
if (rep.total_infractions > 0) {
try {
const history = await getUserRecentInfractions(target.user_id, 2);
const historyXml = buildUserHistoryXml(
history.map((h) => ({
content: h.content ?? "",
severity: h.severity,
created_at: h.created_at,
})),
);
if (historyXml) {
repXml = `<user_reputation ${repAttrs}>\n${historyXml}\n</user_reputation>`;
}
} catch {
// history is a bonus — fall back to attrs-only reputation
}
}
// No per-user reputation/profile context is injected into the prompt —
// keep the AI analysis context minimal (raw messages only). Trust state is
// still tracked in the DB for enforcement, just not shown to the LLM.
const isBot = resolveIsBot(target);
const isEdited = resolveIsEdited(target);
const messageBlock = `<message id="${escapeXml(target.id)}" user="${escapeXml(resolveDisplayName(target))}" time="${new Date(target.created_at).toISOString()}"${isBot ? ` bot="true"` : ""}${isEdited ? ` edited="true"` : ""}>\n ${repXml}${profileRef ? `\n ${profileRef}` : ""}${refXml ? `\n ${refXml}` : ""}\n <content>${escapeXml(truncateForAi(content))}</content>${mediaContext ? ` ${escapeXml(mediaContext)}` : ""}${webContext}${mediaAnalysisContext}${searxngXml}\n</message>`;
const messageBlock = `<message id="${escapeXml(target.id)}" user="${escapeXml(resolveDisplayName(target))}" time="${new Date(target.created_at).toISOString()}"${isBot ? ` bot="true"` : ""}${isEdited ? ` edited="true"` : ""}>\n ${refXml ? `\n ${refXml}` : ""}\n <content>${escapeXml(truncateForAi(content))}</content>${mediaContext ? ` ${escapeXml(mediaContext)}` : ""}${webContext}${mediaAnalysisContext}${searxngXml}${glossaryCtx}\n</message>`;
return { targetId, messageBlock };
}

Some files were not shown because too many files have changed in this diff Show More