Commit Graph
11 Commits
Author SHA1 Message Date
mytheclipsebotreview 2b6ec59286 refactor(ai): remove semantic embedding cache + Qdrant vector store
Hapus seluruh fitur embedding/Qdrant (tidak dipakai lagi):

- gateway: drop embeddingClient.ts, qdrantClient.ts, archiveEmbedder.ts
  dan tes qdrantEnsure.test.ts; moderationOrchestrator kembali ke
  exact-hash cache -> LLM (tanpa phase-2 semantic lookup); textCacheStore
  kehilangan findSimilarTextModeration / parseQdrantVerdict /
  isSemanticBandAccepted / upsertBareKeyToQdrant; cache-prune hanya
  menyapu Postgres.
- backend: drop embed.ts + qdrant.ts, endpoint messages.semanticSearch
  dan schema/type terkait; kolom embedding dilepas dari schema
  text_analysis_cache.
- frontend: hapus toggle EXACT/SEMANTIC, hook useSemanticSearch,
  API client + tipe SemanticSearchResult.
- config: buang AI_LLM_EMBEDDING_* dan QDRANT_* (env + .env.example).
- docs: ARCHITECTURE.md / AGENTS.md / README.md / diagram arsitektur
  disesuaikan (LLM caller - vision, cache = exact-hash saja).

Verifikasi: tsc 0 (backend, gateway, frontend); bun test 135 pass +
37 pass, 0 fail; biome 0 error.
2026-09-25 01:38:51 +07:00
asepharyana ef7708bf7d feat(gmw): route all LLM traffic through 9router
GMW moves off omniroute (100.121.180.82:20128) and off the direct NVIDIA
vision endpoint onto 9router, which runs on the same host as both services
(127.0.0.1:4014) — loopback avoids the TLS/proxy hop and localhost calls
bypass 9router's remote-key guard.

- gateway + backend: AI_LLM_BASE_URL default -> http://127.0.0.1:4014/v1
- drop stale 'omniroute' router references from comments/docs now that the
  active router is 9router (llmClient, llmCaller, ARCHITECTURE, AGENTS)

Verified against 9router before wiring: model 'text' -> gemini-3.5-flash-lite
(SSE, as the pipeline expects), 'multimodal' -> nemotron-3-nano-omni answers
image input, and gemini/gemini-embedding-001 returns 3072 dims — matching the
existing Qdrant collections (no reindex needed). The GMW key is already
registered in 9router's apiKeys table.

typecheck + lint + tests green (gateway 138, backend 37 excluding e2e).
2026-09-24 16:05:36 +07:00
asepharyana 750f3aa598 docs(gateway): correct metrics port — 4016 was wrong
The metrics server binds METRICS_PORT (code default 9090; this host runs it
on 4018 — 4016 is occupied by another process). Docs said 4016 in three
places, which is not what the service does.
2026-09-24 15:50:15 +07:00
asepharyana c57ee12da1 docs(gateway): consolidate README/ARCHITECTURE, drop stale MODULE_STRUCTURE
README.md was the extraction-era document (referenced winston, mock-crc.ts,
llmModerationClient.ts, indonesianTextNormalizer.ts — all long gone) and
duplicated ARCHITECTURE.md. Rewritten as a short run-the-service guide;
layout/design lives only in ARCHITECTURE.md.

MODULE_STRUCTURE.md deleted: it was a stale duplicate of ARCHITECTURE.md,
referenced by nothing but itself.

ARCHITECTURE.md updated to the post-refactor reality: app/ lifecycle split
(bootstrap/lifecycle/process-guards/metrics-collector), ai-moderation
recovery-worker + cache-prune, per-module index.ts facades, one-way
dependency rule, corrected init/shutdown/observability sections.
2026-09-24 15:09:13 +07:00
asepharyana e34dcd6bc6 fix(gateway): separate text and media lanes in AI analysis queue (#86)
Image messages previously blocked the whole analysis pipeline:
- conversationProcessing was a single lock per conversation; processBatch
  awaited BOTH text and media worker jobs before releasing it, so a fast
  text verdict sat unused until the slow vision/media batch finished
- one global LLM semaphore (AI_LLM_MAX_CONCURRENT) was shared by text and
  media, so a vision backlog could starve text inference
- recovery worker gated on conversationProcessing.size

Now the queue is split into independent text/media lanes:
- conversationProcessing maps key -> Partial<Record<lane, startedAt>>;
  each lane holds its own lock and frees it the moment ITS worker job
  resolves (ownership-guarded clear prevents stale timers clearing newer
  slots)
- two LLM semaphores: AI_LLM_MAX_CONCURRENT (text, default 8) and
  AI_LLM_MEDIA_MAX_CONCURRENT (media, default 4) via
  withLlmConcurrency(fn, { lane })
- batchScheduler schedules per conversation+lane (timer keys
  '<key>::<lane>'); splitMessagesByLane/laneOfMessage moved to pure
  analysisLanes.ts (unit-testable without Piscina)
- ai-analysis-worker batch jobs carry a lane field; per-lane active
  request gauges (active_text_requests / active_media_requests)
- added tests/analysisLaneLock.test.ts (7 tests: independent lane locks,
  preserving other-lane lock, clear-all, ownership guard, lane split)

Docs: ARCHITECTURE.md + AGENTS.md concurrency model updated.
typecheck/lint/test(138)/build all green.
2026-09-24 14:09:53 +07:00
asepharyana a13884d80f docs: remove voice/recording/media references from AGENTS/ARCHITECTURE/README 2026-09-23 21:25:03 +07:00
asepharyana 12cc956329 feat(gateway): implement separate Piscina pools for text and media analysis to optimize processing 2026-08-31 22:59:27 +07:00
asepharyana ffbe9959ab chore: migrate AI LLM router from 9router to omniroute
Switch GMW's AI LLM base URL from 9router (https://9router.asepharyana.my.id/v1)
to omniroute on imrnes (http://100.121.180.82:20128/api/v1).

- Update default AI_LLM_BASE_URL in discord-gateway + backend config schemas
- Update .env.example documentation
- Update all 9router references in comments/docs/tests to omniroute
- Production BWS secret gmw_ai_llm_base_url already updated

Omniroute uses /api/v1 prefix (not /v1 like 9router), so the base URL
now correctly points at the right API path for the OpenAI SDK.
2026-08-28 20:18:48 +07:00
asepharyana 2a8f6d9062 refactor(gateway): remove user reputation feature entirely
Drop trust-score/infraction system: delete userReputationStore, remove call sites in fallback/batch processors, drop formatReputationAttrs, drop user_reputations table (migration 0016), delete trust-model test, update docs.
2026-08-18 18:27:15 +07:00
asepharyanaandClaude Opus 5 d2e97ae11d audit(gateway): fix dead /metrics endpoint, raise OOM-prone MemoryMax, trim DB pool
- gateway-metrics: collectors now run per scrape so Prometheus sees real
  data (process memory/uptime + live AI-analysis pipeline gauges) instead
  of an always-empty stub. bootstrap registers the pipeline collectors.
- systemd: MemoryMax 512M -> 1G (live RSS ~500MiB, peak 508MiB; 512M left
  ~2% headroom and risked an OOM-kill restart; host has 8GB free).
- config: POSTGRES_POOL_MIN 2 -> 0 so main + 4 Piscina worker threads don't
  hold ~10 permanently-open idle pg connections against PgBouncer.
- docs: rewrite stale ARCHITECTURE.md / MODULE_STRUCTURE.md (winston ->
  pino, removed mock-crc/indonesianTextNormalizer, renamed
  aiAnalysisWorker/llmModerationClient).

Verified: tsc clean, 129 vitest pass, biome clean on changed files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 08:41:54 +07:00
MythEclipseandClaude Opus 4.8 c48a0c5e3b refactor: split monolith into 3 microservices (frontend, backend, discord-gateway)
- Extract services into services/{frontend,backend,discord-gateway}
- Create packages/shared/ for shared logger, errors, utils, types
- Setup Modular MVC pattern in backend (controller→service→repository)
- Setup event-driven architecture in discord-gateway with Redis pub/sub
- Move Docker files to infra/docker/ with per-service Dockerfiles
- Update docker-compose.yml to use Traefik-only routing (no port exposes)
- Update GitHub Actions deploy workflow for multi-service matrix build
- Fix all import paths and resolve type errors across all services
- All 3 services pass tsc --noEmit clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 21:44:29 +07:00