Files
GMW/services
asepharyana 7ef86c81ca feat(ai): audit + harden embedding pipeline
- Normalize text before embedding (strip mentions/URLs/emoji/markdown/control chars, lowercase, truncate) on both write and query sides so vectors aren't diluted and tokens aren't wasted
- embeddingClient: retry embeddings (maxRetries 2), validate batch dimension consistency, preserve index alignment for empty-normalized texts
- archiveEmbedder: store normalized text in archive payload, skip empty-normalized content
- backend: normalize search queries, make archive search similarity threshold configurable (AI_LLM_EMBEDDING_ARCHIVE_MIN_SIMILARITY, default 0.6)
2026-09-01 18:01:47 +07:00
..