Implements a context-aware moderation system by tracking user behavior
and channel-specific norms to improve AI decision-making accuracy.
- Adds `user_reputations` table to track trust scores, clean streaks,
and infraction history.
- Adds `channel_cultures` table to store AI-generated summaries of
channel-specific norms and slang.
- Implements `userReputationStore` to autonomously update user scores
based on moderation outcomes (clean vs. flagged).
- Implements `cultureLearner` and `channelCultureStore` to manage
evolving channel contexts.
- Enhances LLM prompts to inject user reputation (trust scores,
history) and channel culture summaries, enabling "wisdom-based"
moderation (e.g., giving benefit of the doubt to high-trust users).
- Integrates reputation and culture updates into the existing
`aiAnalyzer` pipeline.
Refactors the AI moderation pipeline to improve concurrency control and
cache efficiency by moving from user-centric to content-centric caching.
- Implements a distributed locking mechanism for media analysis using
`acquireMediaAnalysisLock` to prevent redundant LLM vision calls across
multiple pods.
- Transitions text moderation caching from `user_mod:userId:hash` to a
purely content-based `text_mod:hash` approach to increase hit rates.
- Enhances `getPendingMessagesByConversation` with atomic transactions
and `FOR UPDATE SKIP LOCKED` to safely transition messages from
`pending` to `processing` state.
- Adds `processing` status to the `AIStatus` type and database schema to
track active analysis lifecycles.
- Implements polling logic in `llmModerationClient.ts` to wait for
in-progress media analyses.
- Instructs the AI to decode combinations of regional indicator emojis (e.g., 🇬 🇦 🇾) and custom letters spelling out words, rather than dismissing them as 'just a series of emojis'
- Added a specific few-shot example (Contoh 15) to demonstrate flagging this technique when used to spell banned words
- Added explicit zero-tolerance rule for anatomical/sexual vulgarity (e.g. titten, kontol), explicitly forbidding the AI from passing them off as 'casual conversation' or 'jokes'
- Expanded sexual_deviation rule to explicitly cover brief mentions of BL (Boys Love), yaoi, yuri, and LGBT topics, instructing the AI to flag them regardless of casual context
Sebelumnya: 'Pesan hanya berisi attachment tanpa teks yang melanggar'
Sekarang: harus deskriptif berdasarkan tipe konten:
- Text only: '[user] membahas tentang <topik>. <konteks>.'
- Image only: 'Gambar berupa <jenis>. Terlihat <isi>.'
- Text+Image: '[user] mengirim <gambar> sambil membahas <topik>.'
Tambahkan contoh baik vs buruk di OUTPUT_INSTRUCTIONS sebagai format wajib.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sebelumnya aturan 'percaya teks terlebih dahulu' membuat model abaikan
deskripsi gambar saat teks kosong. Semua image-only message di-clean.
Fix:
- SYSTEM_RULES: pisah Mode 1 (teks+gambar) dan Mode 2 (hanya gambar)
- Mode 2: deskripsi gambar jadi bukti utama, WAJIB dibaca
- Gambar terminal/chat/editor kode/casual → clean
- Gambar dengan elemen judi NYATA (chip, roulette, odds) → flag
- 2 contoh few-shot baru: terminal clean, situs judi flag
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Root cause: vision model diminta untuk 'flag' dan 'menilai' gambar,
sehingga screenshot terminal/chat biasa diklaim sebagai 'situs perjudian'.
Fix:
- vision prompt: HANYA deskripsi objektif (objek, teks, layout, jenis gambar)
- larang tegas kata 'gambling', 'judi', 'pelanggaran', 'harus dihapus'
- tambah buildGeneralImageVisionPrompt di discord-gateway stickerPrompt
- MEDIA_INSTRUCTIONS: tegaskan batch LLM adalah hakim, vision hanya saksi mata
- deskripsi netral (terminal, chat, editor kode) tidak boleh jadi dasar flag
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- Extract services into services/{frontend,backend,discord-gateway}
- Create packages/shared/ for shared logger, errors, utils, types
- Setup Modular MVC pattern in backend (controller→service→repository)
- Setup event-driven architecture in discord-gateway with Redis pub/sub
- Move Docker files to infra/docker/ with per-service Dockerfiles
- Update docker-compose.yml to use Traefik-only routing (no port exposes)
- Update GitHub Actions deploy workflow for multi-service matrix build
- Fix all import paths and resolve type errors across all services
- All 3 services pass tsc --noEmit clean
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>