Add tests/llmE2e.test.ts — 7 end-to-end tests driving the REAL
moderation prompt pipeline (buildSystemPrompt → XML payload → llmChat
→ parseModerationResponse) against a live model via omniroute.
Covers: clean technical content (no false positives), harassment
(flagged), username-only offenses including 'Pecinta Pria' +
sexual/provocative usernames + SARA-in-username (always warn/low,
NEVER delete — the nickname-reset path), and spam bursts.
Gated behind AI_LLM_BASE_URL + AI_LLM_API_KEY: CI (no creds) skips
the file → 216 unit tests stay green, zero LLM cost. Run locally via
pnpm test:e2e:live (scripts/run-llm-e2e.sh injects creds from bws).
Verified: 223/223 tests pass with live LLM, stability across 4 runs,
typecheck + biome clean. docs: TESTING.md. ignore .hermes/ plans.
- new tinyFishSearch module: GET api.search.tinyfish.ai with X-API-Key,
maps top-3 to SearchResult shape, never throws (all failure modes -> [])
- wikipediaSearch: on wiki miss, one tinyfish attempt; hits cached 6h
under the same key so fallback latency is paid once
- termGlossary: on summary miss, top tinyfish hit becomes the definition
(persisted permanently like wiki defs); miss keeps 1h sentinel
- config: TINYFISH_API_KEY (empty = fallback disabled), ENABLED,
BASE_URL, TIMEOUT_MS, LOCATION, LANGUAGE knobs
- tests: 6 coverage for disabled/mapping/non-OK/network/bad-json
API key NOT committed — set TINYFISH_API_KEY in BWS gmw secrets.
Verified: typecheck + lint clean, 216/216 tests pass, live probe
'gubernur jawa barat' returned 3 mapped results
Even with the prompt firewall (username vs content), the LLM can still
occasionally mis-apply a content-level zero-tolerance flag (sara /
conflict_instigation) to a message whose ONLY violation is the username
(e.g. 'matikanetanyahu'). The auto-delete eligibility check only recognized
exact offensive_username flags, so such false positives still deleted the
message.
Add a belt-and-suspenders guard in isNicknameOnlyViolation: if the flag set
is entirely username-attributable (offensive_username/sara/conflict_instigation)
AND the analysis text corroborates that the violation is username-only with
clean message content, route to nickname-reset instead of message deletion.
Adds 6 test cases covering the real matikanetanyahu scenario and the
false-positive/negative boundaries.
result.score was required by zod; the LLM (gemini-3.5-flash-lite via 9router)
occasionally omits it for media batches, hard-failing the whole batch parse
('Zod validation failed: expected number, received undefined' at
results[0].score). Callers already null-coalesce (result.score ?? 0) and the
parser clampScore()s it, so requiring it only caused parse failures.
Adds regression tests: media-batch without score parses (score->0), and
score-present responses still parse with the value.
A selfbot (user token) cannot auto-detect other members' camera/share
(no VOICE_STATE_UPDATE for others, 403 on member fetch). The only
selfbot-viable path to capture another member's SCREEN SHARE is an
operator-initiated STREAM_WATCH (gateway op 20, not gated on bot-vs-user).
Add video:watch / video:unwatch Redis commands routed via the existing
command handler to startStreamWatch/stopStreamWatch, which then does the
DAVE handshake + per-burst MP4 segmentation + DB insert + Tele upload
(already implemented in streamWatchReceiver).
- new VideoHandler (command-handler/video.handler.ts)
- register video:watch / video:unwatch in handler-registry + CommandHandler
- command constants COMMAND_VIDEO_WATCH / COMMAND_VIDEO_UNWATCH
- resolve active voice channel from voice controller + client cache
- 8 unit tests (videoHandler.test.ts)
- biome fixes for pre-existing test import ordering
All green: typecheck, build, lint (174 files), 200 tests.
Real root cause of 'tidak mendeteksi user yg sudah ada di voice':
The selfbot's discord.js-selfbot-v13 GUILD_CREATE handler only sends
GUILD_SUBSCRIPTIONS_BULK — it DROPS d.voice_states from the payload.
So when the gateway (re)starts, every user who was ALREADY in the voice
channel (with camera or screen-share on) is invisible: channel.members is
empty, and REST fallbacks don't work for user tokens (verified live:
GET /channels/{id}/voice-states -> 404, GET /guilds/{id}/members -> 403).
Fix: register our own raw listener for GUILD_CREATE, capture
d.voice_states per guild (buffer it), and consume it in
scanExistingStreamers when the channel gets tracked (voice join happens
after GUILD_CREATE). This is the ONLY user-token-compatible source of
'who is in voice with video right now'.
- videoRecorder: +pendingGuildCreateVoiceStates buffer, +raw GUILD_CREATE
listener, Path C consumes buffered states (self_video/self_stream,
matching channel, skip self) before the REST fallback
- scanExistingStreamers: Path A cache -> Path C GUILD_CREATE -> Path D
REST members.fetch() best-effort
- +1 test: GUILD_CREATE voice_states buffer -> watch camera+screen-share,
skip self/no-video/other-channel (192/192 pass)
- typecheck + build + biome clean (1 pre-existing warning)
Root cause: scanExistingStreamers only read channel.members, which is
EMPTY after a gateway restart because the selfbot's guild member cache
hasn't been populated yet. So any user who was ALREADY on camera /
screen-sharing when the bot (re)joined was never detected → no video.
Fix: when the member cache is empty, fall back to guild.members.fetch()
(REST GET /guilds/{id}/members — user-token compatible; the bot-only
GET /channels/{id}/voice-states returns 404 for selfbots, verified live)
which populates member.voice states, then re-scan channel.members for
streaming/selfVideo.
- scanExistingStreamers is now async; trackChannel fire-and-forgets it
- logs source=cache vs source=rest-members for observability
- +1 test: cold-start channel.members empty → REST fetch → watch camera user
- 191/191 tests pass, typecheck + build + biome clean
Fork @discordjs/voice 0.19.2 into vendor/discord-voice-fork and patch the
voice gateway handshake so Discord sends camera/screen-share RTP video:
- Identify payload now declares video:true + streams:[] (derived from
Discord-RE/Discord-video-stream) - this is what makes Discord deliver
H264 (PT 96-127) to the voice socket. Previously the client never
declared video capability, so Discord omitted all video SSRCs/packets.
- op-12 Speaking handler maps streams[].ssrc -> videoSSRC alongside the
audio SSRC, so ssrcMap carries camera/screen-share stream IDs.
- SSRCMap.get() now also resolves video SSRCs (video RTP arrives on a
different SSRC than audio), enabling attribution for videoReceiver.ts.
- Both dist/index.js (CJS) and dist/index.mjs (ESM) patched; 3 isolated
hunks vs upstream, verified by diff.
- package.json points @discordjs/voice -> file:vendor/discord-voice-fork
(pnpm lockfile updated, CI --frozen-lockfile compatible).
- .gitignore: replace global dist/ with per-service explicit patterns so
the vendored fork's dist/ is committed while build outputs stay ignored.
- tests: +3 SSRCMap videoSSRC cases (190 total, all pass).
The receive-side pipeline (H264Depacketizer -> .h264 -> muxToMp4 -> MP4)
already exists in videoReceiver.ts; this unblocks it by making Discord
actually deliver video packets.
CRITICAL: djs/voice stores the DAVESession wrapper at net.state.dave
(createDaveSession assigns to state.dave on op4 SessionDescription), NOT
inside connectionData. decryptVideoPacket looked up connectionData.dave which
is ALWAYS undefined -> every video packet hit '!dave?.session' guard and was
silently dropped, so no .h264/.mp4 ever got written despite the handshake
reaching Ready.
Fix: pass net.state.dave as a separate arg (the wrapper has .session ->
Davey.DAVESession) so the DAVE-layer decrypt (MediaType.VIDEO) actually runs.
Typecheck + build pass, lint clean (src/), 179/179 tests.
If someone is ALREADY sharing screen / camera on when the bot joins the
channel, no voiceStateUpdate with streaming:true fires for them, so the
bot never sent STREAM_WATCH and missed their video entirely. trackChannel
now scans channel.members and starts a watch for anyone already streaming
(ignoring the bot itself and non-streamers). Idempotent: startStreamWatch
no-ops if a watch already exists. +2 tests (9/9 in videoRecorder).
Replace the dead selfbot-v13 video path (WS 4017 DAVE). videoRecorder
now delegates to a new streamWatchReceiver that:
- sends STREAM_WATCH (op 20) on voiceState.streaming
- opens a @discordjs/voice Networking to the watch RTC (STREAM_CREATE +
STREAM_SERVER_UPDATE) with DAVE enabled
- decrypts H264 via Davey MediaType.VIDEO, depacketizes + muxes to mp4
- tears down on streaming-stop / leave / untrack
Remove ensureSelfbotVoice/createVideoStream/joinStreamConnection (dead).
recorder.ts no longer fires the futile eager selfbot join.
Video capture (camera + screen share) recorded ZERO frames because the selfbot
ClientVoiceManager.connection was created LAZILY — only when a user started
streaming. At that point the bot is already connected via @discordjs/voice, so
the selfbot re-join never gets a fresh VOICE_SERVER_UPDATE and times out with
VOICE_CONNECTION_TIMEOUT after 15s. joinStreamConnection (STREAM_WATCH) +
receiver.createVideoStream both need that selfbot VoiceConnection CONNECTED.
Fix: establish the selfbot VoiceConnection eagerly in recorder.startRecording,
BEFORE joinVoiceChannel, so it rides the bot's fresh join (Discord emits
VOICE_SERVER_UPDATE → selfbot authenticates). videoRecorder reuses the cached
connection per guild, tears it down on voice stop/destroy. Best-effort — never
blocks audio recording.
Verified: typecheck + build + biome (src/) green; 9/9 videoRecorder tests.
Root cause of why Phase A/B captured zero video: @discordjs/voice is audio-only
and never sends the gateway STREAM_WATCH signal, so Discord never forwards a
member's video RTP to the bot. Live diagnostic confirmed: while members were
sharing, audio .ogg files flowed for many users but no non-opus RTP ever
arrived.
Correct path: discord.js-selfbot-v13 ships a complete native watch/record stack.
New src/modules/voice-recording/videoRecorder.ts:
- Detects streamers via voiceState.streaming on a single idempotent
voiceStateUpdate listener.
- client.voice.joinChannel() (reuses the single session alongside
@discordjs/voice) + joinStreamConnection(userId) -> STREAM_WATCH (op 20).
- receiver.createVideoStream(userId, path) -> PacketHandler -> Recorder
(ffmpeg over UDP loopback) -> Matroska .mkv, decryption handled internally.
Wired in recorder.ts (trackChannel/untrackChannel) + bootstrap.ts
(setVideoRecorderClient/setVideoRecordingsDir). All best-effort; failures log
and never break existing voice/audio. Unit tests 6/6 (single listener, watch
handshake + mkv path, skip own video, idempotence, teardown). Full suite
20 files / 176 tests green; typecheck + build + biome clean.
UNVERIFIED live yet: needs deploy + a streamer to confirm Recorder ready +
playable .mkv.
Phase A captured raw .h264 streams but left them as non-playable elementary
streams. Phase B adds automatic muxing: when a video burst closes, the raw
.h264 is remuxed to a self-contained MP4 via `ffmpeg -c copy` (no re-encode,
fast) with `+faststart`, waits for the write stream to fully flush first so the
mux never reads a truncated tail, and deletes the raw .h264 on success (keeping
it on failure). Output: <RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.mp4.
muxToMp4 is exported + covered by a real-ffmpeg vitest (tests/videoReceiver.test.ts):
generates a tiny baseline h264, remuxes, asserts mp4 exists/non-empty & raw deleted
(also the 5 depacketizer tests). Full gateway suite 170/170 green, tsc + biome clean.
@discordjs/voice only decrypts/forwards AUDIO (opus) — its onUdpMessage drops
every non-opus RTP packet (dist/index.mjs:2068 `!== RTP_OPUS_PAYLOAD_TYPE`, 120).
Video RTP (H264 camera + screen share, plus VP8/VP9/AV1) arrives on the same UDP
socket but was silently discarded.
New videoReceiver.ts wraps receiver.onUdpMessage (like screenShareAudio.ts):
- detects video payload types (96/98/101/102/106/116/126/127),
- decrypts them with the connection secret key/encryptionMode via the SAME
receiver.parsePacket path @discordjs/voice uses for audio (so DAVE + voice
encryption are handled identically),
- depacketizes H264 to AnnexB (single NAL, STAP-A, and FU-A fragmentation),
waiting for a keyframe (SPS/PPS/IDR) before writing,
- writes a raw .h264 elementary stream per user per burst under
<RECORDINGS_DIR>/<uid>/video-<ssrc>-<ts>.h264.
Attribution: a videoSSRC→user index is built from ssrcMap updates; a proximity
fallback mirrors screenShareAudio's inferScreenShareOwner. Bot's own video is
skipped.
Unit tests: tests/videoReceiver.test.ts (AnnexB start code, keyframe gating,
FU-A reassembly, orphan-fragment tolerance) — 5/5 green.
Phase A only (capture raw h264). Phase B (ffmpeg decode+mux to MP4/WebM +
persist) and Phase C (frontend playback) are follow-ups.
- mediaDownloader: flip URL candidate order so discord_url is tried
before uploaded_url (uploaded_url is archive-only fallback)
- ai-analysis-worker: remove upload-pending race guard that blocked
analysis until Tele upload completed; analysis now runs immediately
on the Discord CDN URL
- batchProcessor: remove upload-pending defer/poll-backoff logic
- individualFallbackProcessor: remove upload_pending requeue loop
- batchOutcomeClassifier/fallbackResultClassifier: drop upload_pending
classification (no longer needed)
- tests: update batchOutcomeClassifier + fallbackResultClassifier tests
to reflect removed upload_pending signal
Switch GMW's AI LLM base URL from 9router (https://9router.asepharyana.my.id/v1)
to omniroute on imrnes (http://100.121.180.82:20128/api/v1).
- Update default AI_LLM_BASE_URL in discord-gateway + backend config schemas
- Update .env.example documentation
- Update all 9router references in comments/docs/tests to omniroute
- Production BWS secret gmw_ai_llm_base_url already updated
Omniroute uses /api/v1 prefix (not /v1 like 9router), so the base URL
now correctly points at the right API path for the OpenAI SDK.
Batch race guard balikin {ok:true, rows:[]} tanpa sinyal saat semua target
masih upload-pending -> processor klasifikasi semua incomplete -> fanout ke
individual queue -> di situ requeue + reschedule 250ms -> balik ke batch:
hot loop ~300ms sepanjang upload (10 siklus/3 dtk di log prod 08:13).
Fix: worker batch kini return uploadPendingIds eksplisit; classifier pure
baru (partitionBatchOutcome) partisi completed/upload_pending/incomplete/
parse_failed/api_failed; target upload-pending DEFERRED dengan poll backoff
linear (AI_ANALYSIS_UPLOAD_POLL_MS 1500 base, cap AI_ANALYSIS_MAX_UPLOAD_POLL_MS
8000), tidak pernah masuk fanout; tail shouldScheduleNext tak menimpa defer.
Test: tests/batchOutcomeClassifier.test.ts (8 kasus, pure tanpa DB/Piscina).
Analisis pertama tetap berkonteks (chat history) demi akurasi, tapi
verdict clean non-actionable (conf>=0.85) juga ditulis di bare key
tanpa konteks. Repeat teks sama di channel lain -> exact cache HIT,
bukan LLM call baru. Guard sama dgn read path; dedupe LRU per proses;
bare row tanpa embedding (tier semantic sudah global).
- Fase-1 exact-cache lookup: N query serial -> SATU query ANY($1::text[])
- Global reuse utk bare key legacy, HANYA verdict non-actionable
(clean/flagless/action=none, conf>=0.85, umur<=72h) — flagged/warn
tetap context-scoped
- Semantic cache dua-band: clean band 0.92 default, actionable tetap
0.97; di antara band -> LLM (fail-open ke akurasi)
- hit_count kini di-increment (bulk UPDATE per batch) -> hit-rate terukur
- Cache hasil wikipediaSearch di Redis (6h, hanya hasil non-kosong)
- Memoize fetchUrlSafely utk type=text (LRU 30m + in-flight dedupe)
- makeImageCacheKey strip query CDN Discord (?ex/is/hm, format/width)
-> attachment sama = satu key vision, skip re-download+re-vision
Spec: .hermes/plans/2026-08-24-ai-analysis-cache-optimization.md
Tests: +33 (cacheGuards, discordImageKeyNormalize, cacheBatchLookup)
- pickBatchWithinBudget: stop di overflow pertama (break), bukan skip —
batch tetap prefix kronologis tanpa gap analisis di tengah timeline.
Diekstrak ke batchBudget.ts (pure, estimator di-inject) + regression test.
- callModerationLLM: param opsional maxTokens; text/media caller menghitung
ceiling dari estimasi prompt (floor 2048, cap 16384) — batch kecil tak
lagi reserve window completion 16k.
- getPending/IncompleteMessagesByConversation: sort hasil UPDATE..RETURNING
by created_at ASC — Postgres tak menjamin urutan, konsumen (anchor konteks
messages[0], prefix batch) bergantung pada urutan kronologis.
- normalizeStoredStatus(): exact-hash & semantic (Qdrant/PG) cache reader
sebelumnya menipiskan 'warn' jadi 'flagged'/'clean' (type narrowing
legacy clean|flagged) — merusak gating auto-delete & label dashboard.
Kini status tersimpan dipertahankan penuh (clean/warn/flagged).
- prompts: hapus referensi <user_history> yang tak pernah di-inject,
SearXNG -> Wikipedia (sudah migrasi), referensi section yang tak ada,
typo 'secifik', dan baris list rusak '|-'.
- moderationBuilders: buang dead code buildUserProfilesBlock/
buildUserProfileRef/UserProfileEntry/buildUserHistoryXml (tanpa caller
produksi sejak context minimization) + test-nya.
- test baru: tests/storedStatusNormalization.test.ts (regresi warn).
- gateway-metrics: collectors now run per scrape so Prometheus sees real
data (process memory/uptime + live AI-analysis pipeline gauges) instead
of an always-empty stub. bootstrap registers the pipeline collectors.
- systemd: MemoryMax 512M -> 1G (live RSS ~500MiB, peak 508MiB; 512M left
~2% headroom and risked an OOM-kill restart; host has 8GB free).
- config: POSTGRES_POOL_MIN 2 -> 0 so main + 4 Piscina worker threads don't
hold ~10 permanently-open idle pg connections against PgBouncer.
- docs: rewrite stale ARCHITECTURE.md / MODULE_STRUCTURE.md (winston ->
pino, removed mock-crc/indonesianTextNormalizer, renamed
aiAnalysisWorker/llmModerationClient).
Verified: tsc clean, 129 vitest pass, biome clean on changed files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two root causes behind 'all image analysis failing':
1. imageResizer still emitted lossless PNG for vision input. A 1024px
Facebook photo balloons to multi-MB PNG base64 that the vision model
silently rejects ('Vision API null response'). Switch to JPEG q85
(no upscaling) — same photo drops to ~100-400KB, model processes fine.
Re-encodes even already-small images so raw originals never bloat the
data URL. Added tests/imageResizer.test.ts covering both cases.
2. acquireMediaAnalysisLock INSERT aborted with 'index row requires N
bytes, maximum size is 8191'. text_analysis_cache.text is the PK in a
B-tree index (8191-byte/row cap); callers pass the raw image URL as the
key, and base64 data URLs / very long URLs blow past the limit, so the
lock INSERT fails and every media analysis is skipped. Hash the URL in
makeImageCacheKey (image:<sha256[:32]>) — fixed-length, deterministic,
well under the limit. All store/get/lock/delete callers already route
through this function so lookup stays consistent.
yt-dlp 2026.07.04 rewrites the --cookies file on close. Handing it the
root-owned /etc/.../ytcookies.txt (not writable by the gmw service user)
caused PermissionError -> exit 1 on every screen-share download attempt.
- buildCookieArgs on-disk branch now copies the system cookie file into a
per-run temp file (like the env branch) so write-back lands somewhere we
own; unreadable -> anonymous.
- resolveInputWithRetry Invidious fallback regex now also matches
permission|EACCES|cookie, so a cookie failure triggers the link-alternative
(no-auth Invidious mirror) path instead of failing all retries.
- adds regression test asserting the original cookie path is never passed to yt-dlp
The live pipe (yt-dlp -o - -> ffmpeg) delivers data at network speed with
unreliable PTS, which defeats ffmpeg -re and made x264 -r 30 force-duplicate
held frames -> ~1fps video (the patah-patah symptom). Per user suggestion,
download the FULL clip to a temp file first (downloadScreenInput), then feed
that FILE PATH to prepareStream. String inputs already get -re, so the
encoder now paces cleanly at 1x against a monotonic-PTS file — proven
reliable in local tests (vs the live pipe which always bursted). Temp file
is removed on stream end / stop.
- getDirectScreenInput -> downloadScreenInput (returns file path)
- resolveInputWithRetry now awaits a completed file + retries on failure
- screenShareController.stops/cleanup removes the per-run tmpdir
- screenShareInput.test.ts updated to the file-download contract
Verifies that two data URLs sharing the first 128 chars (same MIME prefix
+ identical base64 header — the real-world scenario that caused ALL images
to reuse the same cached vision analysis) produce DIFFERENT cache keys
under the fixed full-dataURL hashing, whereas the old 128-char-prefix
approach would collide. Also includes consistency + prefix tests.
Screen share showed a single frozen frame: the GoLive pipeline sent video
only (-"-an", h264 muxer cannot carry audio) so the audio SSRC never
transmitted and Discord kept the stream in thumbnail state.
- prepareStream: mux NUT when includeAudio (h264 muxer drops audio) and
return the actual container format
- Demuxer: support NUT input with a second output pipe (fd3) carrying
Ogg Opus; parse OGG pages into opus frames (20ms, 48kHz) emitted as
GoLiveFrames; fix metadata parsing that dropped the audio stream line
when it arrived in a later stderr chunk (early parsedMeta return)
- playStream: pipe audio.stream into AudioStream → RTP on the audio SSRC
- Encoders: -x264-params repeat-headers=1 → SPS/PPS inline before EVERY
IDR (NUT remux drops container extradata; also enables PLI recovery)
- screenShareController: includeAudio true
- tests: demuxerNut.test.ts — OGG parser unit test + real ffmpeg NUT
integration (video access units + parsed opus frames)
Second root cause (2026-08-12): even with yt-dlp http_headers forwarded,
YouTube still returns 403 when a signed DASH URL from --dump-single-json is
fetched raw by ffmpeg/curl on some videos (verified on fONoh7Pc6VU: curl
with the EXACT headers got 403; yt-dlp's own downloader succeeded). The
signature is tied to the extracting client context (po_token/visitor), not
just UA/IP.
Fix: getDirectScreenInput now spawns 'yt-dlp -o -' and returns its stdout
as a Readable — the same mechanism resolveMediaUrl already uses for music.
yt-dlp handles auth, cookies and transient retries internally. Merge
fragments go to /tmp/gmw-ytdlp-tmp (Nix store CWD is read-only → EACCES).
Removed resolveScreenInput + mergeScreenStreams (dead code).
Controller resolveInputWithRetry unchanged: tees the stream, waits for the
first byte (12s), retries with a fresh yt-dlp run up to 3x on error/EOF/
timeout, and destroys stuck inputs (EPIPE) so no process leaks.
Tests: rewritten for streaming (yt-dlp emits bytes; fail mode = exit 8
without stdout → stream must terminate with zero bytes).
Root cause (2026-08-12 11:50 test): merge ffmpeg hit a transient YouTube
403 and exited code 8 BEFORE prepareStream attached its input listeners
(voice release+join takes ~10s). The input's end/error events fired into
the void, the encoder stdin never received EOF, demux resolved with
fallback 0x0 metadata, setSpeaking fired anyway → stream 'started' with
zero frames for 8+ minutes (black tile, both ffmpeg processes hung).
Fixes:
- mediaSource: pass yt-dlp http_headers (UA/referer) to the merge ffmpeg
via -headers to suppress transient 403s; destroy the returned stream
with an error when the merge exits non-zero before producing bytes.
- screenShareController: resolveInputWithRetry — tee the merge stream and
wait for the first readable byte (12s timeout) before proceeding; on
error/EOF/timeout retry the whole resolution with a FRESH yt-dlp run
(signed DASH URLs expire fast) up to 3 attempts. Stuck merges get
EPIPE via input.destroy() so no process leaks per attempt.
- prepareStream: race guard — if the input already ended/destroyed before
listeners attach, EOF the encoder stdin immediately; first-frame
watchdog in playStream rejects 'started but nothing flowing' after 10s
instead of resolving with a silent black stream.
Tests: +2 (merge-fail zero-byte terminal state, -headers forwarding).
SDP offer advertises profile-level-id=42e01f (constrained baseline) but
x264 encoded the default High profile — Discord's receiver configures its
decoder from the negotiated profile, so the High-profile bitstream failed
to decode → black GoLive tile despite valid access units + correct RTP
timestamps (fixed in 42a503c).
- Add -profile:v baseline to H264 encoder options (matches @dank074's
proven config; SPS now 6742c01e → profile_idc=66 baseline, aligns with
the 42e01f fmtp).
- Default x264 tune film → zerolatency (no lookahead — correct for live
GoLive; @dank074 uses it).
- Update goLive-port test to assert baseline + zerolatency.
Demuxer emitted each AnnexB NAL as its own WebRTC frame (SPS/PPS/SEI
separate from slices) with a near-zero timestamp delta (duration=1 in a
1/90000 timebase → RTP +1/frame instead of +3000 @30fps). Discord's H264
receiver never receives a complete decodable access unit → black GoLive
tile despite frames flowing.
- Group NALs into access units: buffer param-set/SEI NALs, flush one
frame per slice with preceding parameter sets (AnnexB start codes kept
so the H264RtpPacketizer finds NAL boundaries).
- Timestamp each frame at the video frame rate: duration=1, timeBase
1/fps → BaseMediaStream frametime=1000/fps ms → RTP +clockRate/fps
(3000 @ 30fps/90kHz) and correct pacing.
- Thread explicit frameRate from playStream options (raw H264 has no
timing info; ffmpeg guesses 25fps on stderr).
- Strengthen golive-demux-live-e2e: validates every frame has a slice,
no bare param-set frames, keyframes carry SPS/PPS, timeBase 1/30.
Root cause of 'tile appears but content empty': demux() spooled the live
NUT/H264 input to a temp file and awaited stream 'finish' — but the merge
ffmpeg output never ends during playback, so demux deadlocked, no probe,
no transcode, 0 frames sent.
- Demuxer: pipe input straight into ffmpeg stdin (-i pipe:0), parse NAL
frames live from stdout; parse video metadata from ffmpeg stderr with a
1.5s race (fall back to H264 defaults). No spool, no await-end.
- screenShareController: pass width/height/frameRate (1280x720@30) to
playStream — matches the prepareStream encode settings, so setVideoAttributes
gets real dimensions even when ffmpeg can't report metadata on an open pipe.
- Add tests/golive-demux-live-e2e.ts: proves frames flow while input is
still open (regression test for the deadlock).
Root cause (3rd layer after 50371bd + 4f4c435): a vision model run
(2026-08-10) returned 'Maaf, saya tidak melihat gambar apapun yang terlampir...'
and that text was cached as a VALID vision_llm result (image + phash keys,
24h/7d TTL). Every subsequent analysis of the same image (same hash/phash)
hit the poisoned cache, so image analysis looked broken forever even though
9router responded fine — the moderation LLM wrote 'lampiran yang gagal
terbaca' from a cache hit.
Also: mimo via 9router streams reasoning in delta.reasoning +
delta.reasoning_details[].text (content:"") — extractChunkText only read
delta.reasoning_content, so those runs aggregated empty → 'Vision API null
response' (observed 08:54/09:07/09:38).
Fixes:
- llmClient.extractChunkText: fall back to delta.reasoning and
reasoning_details[].text (mimo), on top of reasoning_content (gemma).
- visionAnalyzer: isNoImageSeenText() detects 'no image' style outputs;
such results are NEVER cached, and poisoned entries are purged when hit
(LRU/DB/phash) so re-analysis actually re-runs vision.
- Tests: reasoning/reasoning_details extraction + isNoImageSeenText
(Indonesian + English, no false positives on real descriptions).
Root cause: 9router combo 'multimodal' routes to cloudflare-ai/@cf/google/
gemma-4-26b-a4b-it which streams ALL output in delta.reasoning_content
(content:"") and finishes with 'length' at max_tokens. llmClient only read
delta.content, so llmVision returned empty → every image moderation fell back
to text-only analysis ('Meskipun analisis gambar gagal' in every ai_analysis).
Fix: extractChunkText() prefers delta.content then falls back to
delta.reasoning_content (also handles message/text/response fields), with
unit tests for the exact 9router chunk shape. Verified live against a real
DB image: oc/mimo-v2.5-free (new first model in the multimodal combo) returns
a proper description in delta.content.
- <message> targets now carry time (ISO), repetitions (N identical short texts = spam signal), bot and edited flags; escape id/user XML
- rich <user_reputation>: total_infractions, clean_streak, last_offense_days_ago, repeat_offender (7-day window)
- <user_history> with last flagged messages for repeat offenders (wires dead getUserRecentInfractions)
- <user_profile as_of> staleness signal; <location_context topic> from captured channel topic
- prompt framing + output instructions teach the LLM to use the new signals without treating history as proof
- tests: contextEnrichment.test.ts (13) + topic cases in conversationContext.test.ts
When the ONLY violation is offensive_username (message content clean):
- Message is NOT deleted (nickname-only violation bypasses auto-delete)
- Member's server nickname is reset to default username via
setNickname(null) (Discord shows the global username again)
- Action 'reset_nickname' logged to moderation_actions; cooldown
10min per guild:user (LRU) so repeated messages by same member
don't hammer the Discord PATCH
- Config: AUTO_NICKNAME_RESET_ENABLED / AUTO_NICKNAME_RESET_COOLDOWN_MS
- resolveDisplayName(): member.displayName from captured metadata,
falls back to global username
- Applied to context lines, target message blocks, and media message
blocks — LLM sees the name the channel actually sees (nickname can
carry moderation signal itself)