refactor(ai): remove semantic embedding cache + Qdrant vector store
Hapus seluruh fitur embedding/Qdrant (tidak dipakai lagi): - gateway: drop embeddingClient.ts, qdrantClient.ts, archiveEmbedder.ts dan tes qdrantEnsure.test.ts; moderationOrchestrator kembali ke exact-hash cache -> LLM (tanpa phase-2 semantic lookup); textCacheStore kehilangan findSimilarTextModeration / parseQdrantVerdict / isSemanticBandAccepted / upsertBareKeyToQdrant; cache-prune hanya menyapu Postgres. - backend: drop embed.ts + qdrant.ts, endpoint messages.semanticSearch dan schema/type terkait; kolom embedding dilepas dari schema text_analysis_cache. - frontend: hapus toggle EXACT/SEMANTIC, hook useSemanticSearch, API client + tipe SemanticSearchResult. - config: buang AI_LLM_EMBEDDING_* dan QDRANT_* (env + .env.example). - docs: ARCHITECTURE.md / AGENTS.md / README.md / diagram arsitektur disesuaikan (LLM caller - vision, cache = exact-hash saja). Verifikasi: tsc 0 (backend, gateway, frontend); bun test 135 pass + 37 pass, 0 fail; biome 0 error.
This commit is contained in:
@@ -64,11 +64,6 @@ AI_LLM_API_KEY= # REQUIRED if AI_ANALYSIS_ENABLED=true. L
|
||||
AI_LLM_BASE_URL=http://100.121.180.82:20128/api/v1 # LLM API base URL (omniroute — OpenAI-compatible router on imrnes, Tailscale 100.121.180.82)
|
||||
AI_LLM_MODEL=text # LLM text model name (default: text)
|
||||
# AI_LLM_VISION_MODEL= # Vision model for image analysis (falls back to AI_LLM_MODEL)
|
||||
# AI_LLM_EMBEDDING_MODEL= # Embedding model for semantic moderation cache (optional; enables near-duplicate text reuse to save LLM calls)
|
||||
# AI_LLM_EMBEDDING_MIN_SIMILARITY=0.97 # Min cosine similarity to reuse a cached verdict (default: 0.97)
|
||||
QDRANT_URL=http://100.121.180.82:6333 # Qdrant vector store for embeddings (semantic cache); when set, vectors are stored/searched in Qdrant instead of Postgres
|
||||
# QDRANT_COLLECTION=gmw_text_moderation # Qdrant collection name (default: gmw_text_moderation)
|
||||
# QDRANT_API_KEY= # Qdrant API key (optional)
|
||||
AI_LLM_MAX_CONCURRENT=5 # Max concurrent LLM API calls (default: 5)
|
||||
AI_LLM_IMAGE_MAX_DIMENSION=1024 # Max image dimension in pixels before resize (default: 1024)
|
||||
AI_LLM_TEXT_BATCH_SIZE=20 # Max messages per text-only moderation batch (default: 20)
|
||||
|
||||
Reference in New Issue
Block a user