fix(ai-moderation): audit prompts to reduce false positives

- ENTROPY rule: replace 'PILIH WARN' default with 'PILIH CLEAN'
- Add safe-list for code/log/stack traces/project names
- Narrow sexual_deviation: only flag explicit sexual solicitation
- Remove overbroad ontological graph (kostum hewan → furry) detection
- Add clean few-shot examples: error logs, project names, orientation disclosure
- Update simple fallback prompt with FP prevention rules
- Remove LGBT/furry from offensive username criteria

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
MythEclipse
2026-06-13 00:01:09 +07:00
co-authored by Claude
parent 12f9b55000
commit 0ac056dc8a
2 changed files with 64 additions and 27 deletions
@@ -1465,7 +1465,11 @@ Aturan:
- warn: spam ringan, promosi tidak jelas, atau pelanggaran ringan
- flagged: harassment, SARA, NSFW, judi, ancaman, atau pelanggaran serius
PENTING: Slang Indonesia ("anjay", "wkwk", "njir", "gws", dll) dan makian umum ("asu", "anjing", "bangsat") yang TIDAK ditujukan ke orang lain = clean.
PENTING (False Positive Prevention):
- Slang Indonesia ("anjay", "wkwk", "njir", "gws", dll) dan makian umum ("asu", "anjing", "bangsat") yang TIDAK ditujukan ke orang lain = clean.
- Konten coding/programming (kode, log error, SQL, command line, error message, stack trace, nama library) = clean. JANGAN flag hanya karena ada kata "error" atau "crash" dalam konteks teknis.
- Nama proyek, tools, framework (IMPHNEN, Bete, Cursor, Claude, React, Discord) = clean.
- Percakapan multilingual (campuran Indonesia-Inggris) = clean.
Pesan: "${truncatedContent}"