feat(mcpedia): Phase 3 — async indexing (BullMQ), git-sync webhook, revisions, MCP Resources
- packages/queue: ioredis singleton + BullMQ Queue/Worker (prefix mcpedia:
on shared imrnes Redis :6379); apps/worker runs startWorker()
- @mcpedia/core: indexContentFile/runFullIndex (single indexing entry point
shared by script/worker/hook) + revision.service (list/get/restore)
- document_revisions table (migration 0002) — snapshots only on body change
- apps/api: POST /hooks/reindex + /hooks/index webhooks; tRPC revisions,
getRevision, restoreRevision, jobStatus, queueStatus
- apps/mcp: register MCP Resources mcpedia://docs{/,+slug/chunks/revisions}
({+slug} RFC6570 reserved expansion for slugs containing /)
- apps/mcp zod pinned to ^4 to match MCP SDK 1.30 compiled types
(resolves registerTool TS2589/ShapeOutput skew)
- scripts/enqueue.ts one-shot job enqueue helper; indexer refactored to runFullIndex
- PHASES.md/README/.env.example/docs updated
This commit is contained in:
@@ -0,0 +1,114 @@
|
|||||||
|
# MCPedia — Phase 3 "Async + Scale" Implementation Plan
|
||||||
|
|
||||||
|
Status: Phase 1 (MVP) + Phase 2 (Semantic+API) DONE. Phase 3 adds async
|
||||||
|
background work, git-driven reindex, document revision history, and MCP
|
||||||
|
Resources. All logic stays in `@mcpedia/core`; new `packages/queue` wires
|
||||||
|
BullMQ; `apps/worker` runs the worker process; the existing API gets a git-sync
|
||||||
|
webhook + job-status procedures; the MCP server gains Resources.
|
||||||
|
|
||||||
|
## Scope (4 features from PHASES.md)
|
||||||
|
|
||||||
|
1. **Redis + BullMQ background indexing/embedding workers**
|
||||||
|
2. **Git synchronization hook** (auto-reindex on push via webhook)
|
||||||
|
3. **Document revision system** (`document_revisions`)
|
||||||
|
4. **MCP Resources** (`mcpedia://docs/...`) alongside existing tools
|
||||||
|
|
||||||
|
## Architecture decisions (locked)
|
||||||
|
|
||||||
|
- **Redis**: shared imrnes Redis `100.121.180.82:6379`, no auth (verified
|
||||||
|
`+PONG`). `REDIS_URL` env (default `redis://100.121.180.82:6379`), optional
|
||||||
|
`REDIS_PASSWORD`. BullMQ key prefix `mcpedia:` to avoid collisions on the
|
||||||
|
shared instance.
|
||||||
|
- **Queue lib**: `bullmq@6.1.2` + `ioredis@6.0.0` (BullMQ peer dep). Pass an
|
||||||
|
ioredis instance; BullMQ duplicates it for blocking commands.
|
||||||
|
- **Single source of truth preserved**: per-doc indexing logic moves into
|
||||||
|
`@mcpedia/core` as `indexContentFile(relPath, reason?)`. The script, the
|
||||||
|
worker, and the git hook ALL call this. Revisions are snapshotted inside it.
|
||||||
|
- **Revisions**: created only when body actually changes vs the latest revision
|
||||||
|
(avoids bloat on every sync). Stored in `document_revisions`.
|
||||||
|
|
||||||
|
## Files touched
|
||||||
|
|
||||||
|
### packages/config
|
||||||
|
- `src/index.ts`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
|
||||||
|
|
||||||
|
### packages/db
|
||||||
|
- `src/schema.ts`: add `documentRevisions` table
|
||||||
|
(id, documentId→documents.id cascade, slug, revisionNo int, title, body,
|
||||||
|
meta jsonb, reason text, createdAt). Index (document_id, revision_no DESC),
|
||||||
|
(slug).
|
||||||
|
- `drizzle/0002_document_revisions.sql`: migration (applied via psql).
|
||||||
|
- `drizzle/meta/0002_snapshot.json` + `_journal.json` entry (keeps drizzle-kit
|
||||||
|
consistent even though we apply manually).
|
||||||
|
|
||||||
|
### packages/core (new)
|
||||||
|
- `src/index.service.ts`:
|
||||||
|
- `indexContentFile(relPath: string, reason = "index")` — parse → upsert
|
||||||
|
`documents` → `indexChunks` → snapshot revision (if changed).
|
||||||
|
- `runFullIndex(reason?)` — walk content, index each, return counts.
|
||||||
|
- `src/revision.service.ts`:
|
||||||
|
- `createRevision(...)`, `listRevisions(slug, limit)`,
|
||||||
|
`getRevision(id)`, `latestRevisionBody(slug)`, `restoreRevision(id)`.
|
||||||
|
- `src/index.ts`: export both.
|
||||||
|
|
||||||
|
### packages/queue (NEW)
|
||||||
|
- `package.json` (@mcpedia/queue): deps bullmq, ioredis, @mcpedia/core,
|
||||||
|
@mcpedia/db, @mcpedia/config.
|
||||||
|
- `src/client.ts`: ioredis instance factory from config.
|
||||||
|
- `src/queue.ts`:
|
||||||
|
- `INDEX_QUEUE = "mcpedia-index"`.
|
||||||
|
- `enqueueIndexDoc(slug, absPath, reason)`, `enqueueFullIndex(reason)`.
|
||||||
|
- `getQueue()` lazy singleton.
|
||||||
|
- `src/worker.ts`: `startWorker()` — BullMQ Worker with 3 job types:
|
||||||
|
`index-doc` (single), `index-all` (full), `reindex` (full, reason=git-push).
|
||||||
|
Graceful shutdown on SIGINT/SIGTERM. Job progress + error handling.
|
||||||
|
|
||||||
|
### apps/worker (NEW)
|
||||||
|
- `package.json` (@mcpedia/worker): script `start: bun src/index.ts`.
|
||||||
|
- `src/index.ts`: `startWorker()` + heartbeat log.
|
||||||
|
|
||||||
|
### apps/api
|
||||||
|
- `src/index.ts`: add `POST /hooks/reindex` (full) and
|
||||||
|
`POST /hooks/index?slug=` (single) webhook routes → enqueue jobs. Mount
|
||||||
|
AFTER /trpc.
|
||||||
|
- `src/router.ts`: add `jobStatus` (id→state/prev/failedData),
|
||||||
|
`queueStatus` (waiting/active/completed/failed counts),
|
||||||
|
`revisions` (slug→list), `restoreRevision` (id→new slug/doc).
|
||||||
|
- `package.json`: add `@mcpedia/queue` dep, `hooks` reused.
|
||||||
|
|
||||||
|
### apps/mcp
|
||||||
|
- `src/index.ts`: register Resources:
|
||||||
|
- `mcpedia://docs` (list all metas)
|
||||||
|
- `mcpedia://docs/{slug}` (full body from disk)
|
||||||
|
- `mcpedia://docs/{slug}/chunks` (chunk previews)
|
||||||
|
- `mcpedia://docs/{slug}/revisions` (revision list)
|
||||||
|
- `src/smoke.test.ts`: add `listResources` + read `mcpedia://docs` assertion.
|
||||||
|
|
||||||
|
### scripts
|
||||||
|
- `scripts/indexer.ts`: refactor `main()` to call `runFullIndex()`.
|
||||||
|
|
||||||
|
### Root
|
||||||
|
- `package.json`: add `"worker": "bun --cwd apps/worker run start"`,
|
||||||
|
`"reindex": "bun run scripts/worker.ts"`? No — `worker` runs the listener;
|
||||||
|
triggering reindex = `bun run api` webhook or `enqueueFullIndex` helper.
|
||||||
|
Add `"enqueue-index": "bun run scripts/enqueue.ts"` (one-shot enqueue).
|
||||||
|
- `.env.example`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
|
||||||
|
|
||||||
|
### Docs
|
||||||
|
- `PHASES.md`: mark Phase 3 items DONE with notes.
|
||||||
|
- `README.md`: document worker, webhook, revisions, MCP resources.
|
||||||
|
|
||||||
|
## Verification (real, not claimed)
|
||||||
|
|
||||||
|
1. `bun install` picks up new deps.
|
||||||
|
2. `bunx turbo run build` + `typecheck` green across workspace.
|
||||||
|
3. **Real BullMQ e2e against imrnes Redis**: script that enqueues an
|
||||||
|
`index-doc` job, starts a Worker, asserts the job completes and the doc row
|
||||||
|
+ chunks + a revision row appear in Postgres. Verifies Redis+ioredis+bullmq
|
||||||
|
+ db + core all wired correctly.
|
||||||
|
4. `bun --cwd apps/mcp run smoke` passes (incl. new resources).
|
||||||
|
5. `bun run index` (runFullIndex) green; verify `documents`,
|
||||||
|
`document_chunks`, `document_revisions` row counts via psql.
|
||||||
|
6. API webhook: `curl -XPOST localhost:4020/hooks/reindex` enqueues; worker
|
||||||
|
processes; `curl localhost:4020/trpc/queueStatus` reflects counts.
|
||||||
|
7. MCP resource read returns real content.
|
||||||
@@ -27,12 +27,50 @@ Legend: ✅ built · 🟡 partial · ⬜ deferred
|
|||||||
- [x] MCP server — added `semantic_search` + `hybrid_search` tools (6 total).
|
- [x] MCP server — added `semantic_search` + `hybrid_search` tools (6 total).
|
||||||
- [x] Web search — keyword/hybrid toggle (`?mode=hybrid`), hybrid reaches semantically-related docs keyword misses.
|
- [x] Web search — keyword/hybrid toggle (`?mode=hybrid`), hybrid reaches semantically-related docs keyword misses.
|
||||||
|
|
||||||
## Phase 3 — Async + Scale
|
## Phase 3 — Async + Scale ✅ DONE
|
||||||
|
|
||||||
- [ ] Redis + BullMQ background indexing / embedding workers
|
- [x] **Redis + BullMQ background indexing / embedding workers** —
|
||||||
- [ ] Git synchronization hook (auto-reindex on push)
|
`packages/queue` (ioredis singleton + BullMQ `Queue`/`Worker`, prefix
|
||||||
- [ ] Document revision system (`document_revisions`)
|
`mcpedia:` on shared imrnes Redis `:6379`); `apps/worker` runs
|
||||||
- [ ] MCP Resources (`mcpedia://docs/...`) in addition to tools
|
`startWorker()`. Three job types: `index-doc`, `index-all`, `reindex`.
|
||||||
|
Single indexing entry point `indexContentFile`/`runFullIndex` in
|
||||||
|
`@mcpedia/core` shared by the script, worker, and git hook. Verified
|
||||||
|
end-to-end against live Redis (job enqueue → worker → Postgres write).
|
||||||
|
- [x] **Git synchronization hook (auto-reindex on push)** — API webhook
|
||||||
|
`POST /hooks/reindex` (full) and `POST /hooks/index?slug=` (single) enqueue
|
||||||
|
BullMQ jobs. Wire a Git provider (GitHub/Gitea) post-receive / webhook to
|
||||||
|
`POST /hooks/reindex` to auto-reindex on push. `scripts/enqueue.ts` is a
|
||||||
|
one-shot enqueue helper (`bun run enqueue --all` / `<slug>`).
|
||||||
|
- [x] **Document revision system (`document_revisions`)** — `packages/db`
|
||||||
|
migration `0002_document_revisions.sql`. Indexer snapshots a revision only
|
||||||
|
when the body actually changes vs the latest revision (pure metadata edits
|
||||||
|
don't bloat history). `listRevisions` / `getRevision` / `restoreRevision`
|
||||||
|
in `@mcpedia/core`; exposed as tRPC `revisions` / `getRevision` /
|
||||||
|
`restoreRevision` and the `mcpedia://docs/{+slug}/revisions` MCP Resource.
|
||||||
|
- [x] **MCP Resources (`mcpedia://docs/...`)** — alongside the 6 tools:
|
||||||
|
`mcpedia://docs` (list), `mcpedia://docs/{+slug}` (body from disk),
|
||||||
|
`mcpedia://docs/{+slug}/chunks` (chunk preview),
|
||||||
|
`mcpedia://docs/{+slug}/revisions` (history). `{+slug}` uses RFC 6570
|
||||||
|
reserved expansion so slugs containing `/` match.
|
||||||
|
|
||||||
|
### New/changed commands
|
||||||
|
```
|
||||||
|
bun run index # full reindex (runFullIndex, writes revisions)
|
||||||
|
bun run enqueue --all # enqueue a full reindex job (no worker needed)
|
||||||
|
bun run enqueue <slug> # enqueue a single-doc reindex job
|
||||||
|
bun run worker # start the BullMQ indexing worker (long-running)
|
||||||
|
bun run api # Hono+tRPC API on :4020 (added /hooks/* webhooks)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Verification done (real, against imrnes Redis + Postgres)
|
||||||
|
- `turbo run typecheck` green across all 13 packages.
|
||||||
|
- BullMQ e2e: enqueue `index-doc` → worker completes → `documents` +
|
||||||
|
`document_chunks` + `document_revisions` rows present.
|
||||||
|
- Revision dedup proven: editing a body creates a new revision; metadata-only
|
||||||
|
reindex does not; `restoreRevision` writes history back into the live row.
|
||||||
|
- MCP smoke test passes (tools + all 4 resources).
|
||||||
|
- API webhook `POST /hooks/reindex` enqueues → worker drains queue →
|
||||||
|
`queueStatus` reflects counts.
|
||||||
|
|
||||||
## Phase 4 — Scale-out (only if needed)
|
## Phase 4 — Scale-out (only if needed)
|
||||||
|
|
||||||
|
|||||||
@@ -14,16 +14,19 @@ column) and served through a single **Core** layer that every interface
|
|||||||
mcpedia/
|
mcpedia/
|
||||||
├── apps/
|
├── apps/
|
||||||
│ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search
|
│ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search
|
||||||
│ └── mcp/ # MCP server (stdio) — AI-agent interface
|
│ ├── mcp/ # MCP server (stdio) — AI-agent interface (tools + resources)
|
||||||
|
│ └── api/ # Hono + tRPC v11 API on :4020 (+ /hooks/* git-sync webhooks)
|
||||||
├── packages/
|
├── packages/
|
||||||
│ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...)
|
│ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...)
|
||||||
│ ├── config/ # loads .env (repo root) as authoritative dev config
|
│ ├── config/ # loads .env (repo root) as authoritative dev config
|
||||||
│ ├── db/ # Drizzle ORM schema + client + drizzle-kit config
|
│ ├── db/ # Drizzle ORM schema + client + drizzle-kit config
|
||||||
│ ├── parser/ # frontmatter (gray-matter) parsing
|
│ ├── parser/ # frontmatter (gray-matter) parsing
|
||||||
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline)
|
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline)
|
||||||
│ └── core/ # Document/Content/Search services — the only business logic
|
│ ├── embeddings/ # embedding provider + chunker
|
||||||
|
│ ├── queue/ # Redis (ioredis) + BullMQ worker/queue (Phase 3)
|
||||||
|
│ └── core/ # Document/Content/Search/Index/Revision — the only business logic
|
||||||
├── content/ # docs/ writeups/ research/ notes/ (the knowledge base)
|
├── content/ # docs/ writeups/ research/ notes/ (the knowledge base)
|
||||||
└── scripts/ # indexer.ts (walks content/ -> upserts into Postgres)
|
└── scripts/ # indexer.ts (full reindex), enqueue.ts (one-shot job enqueue)
|
||||||
```
|
```
|
||||||
|
|
||||||
## Architecture principle
|
## Architecture principle
|
||||||
@@ -101,23 +104,45 @@ the DB stores metadata + the search vector.
|
|||||||
| `list_documents` | List, optionally filtered by section |
|
| `list_documents` | List, optionally filtered by section |
|
||||||
| `get_related_documents` | Docs sharing tags with a given slug |
|
| `get_related_documents` | Docs sharing tags with a given slug |
|
||||||
|
|
||||||
|
### MCP Resources
|
||||||
|
|
||||||
|
| URI | Purpose |
|
||||||
|
| -------------------------------- | ---------------------------------------- |
|
||||||
|
| `mcpedia://docs` | List all published documents |
|
||||||
|
| `mcpedia://docs/{+slug}` | Full markdown body (read from disk) |
|
||||||
|
| `mcpedia://docs/{+slug}/chunks` | Preview of embedded semantic chunks |
|
||||||
|
| `mcpedia://docs/{+slug}/revisions` | Revision history summary |
|
||||||
|
|
||||||
|
(`{+slug}` uses RFC 6570 reserved expansion so a slug like
|
||||||
|
`docs/websocket/contract` matches the template.)
|
||||||
|
|
||||||
Smoke test (in-memory transport, real JSON-RPC):
|
Smoke test (in-memory transport, real JSON-RPC):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bun --cwd apps/mcp run smoke
|
bun --cwd apps/mcp run smoke
|
||||||
```
|
```
|
||||||
|
|
||||||
## API (Phase 2)
|
## API (Phase 2 + Phase 3)
|
||||||
|
|
||||||
A tRPC v11 API is also exposed via Hono on **:4020** (all procedures mirror the
|
A tRPC v11 API is exposed via Hono on **:4020** (all procedures mirror the
|
||||||
MCP tools):
|
MCP tools). Phase 3 adds async job + revision procedures and git-sync webhooks:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*)
|
bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*)
|
||||||
```
|
```
|
||||||
|
|
||||||
`bun run index` now also chunks + embeds (Phase 2 indexer). Requires `EMBED_*`
|
tRPC procedures: `search`, `semanticSearch`, `hybridSearch`, `getDocument`,
|
||||||
vars in `.env` (see `.env.example`).
|
`listDocuments`, `related` (Phase 2); plus `revisions`, `getRevision`,
|
||||||
|
`restoreRevision`, `jobStatus`, `queueStatus` (Phase 3).
|
||||||
|
|
||||||
|
Git-sync webhooks (enqueue BullMQ jobs; the worker processes them):
|
||||||
|
- `POST /hooks/reindex` — full-corpus reindex (point your Git provider's
|
||||||
|
push webhook here to auto-reindex on push).
|
||||||
|
- `POST /hooks/index?slug=<slug>` — reindex a single document.
|
||||||
|
|
||||||
|
`bun run index` now also chunks + embeds (Phase 2 indexer) and snapshots a
|
||||||
|
revision whenever the body changes (Phase 3). See `.env.example` for
|
||||||
|
`EMBED_*` / `REDIS_*` / `QUEUE_PREFIX` vars.
|
||||||
|
|
||||||
## Status
|
## Status
|
||||||
|
|
||||||
@@ -128,6 +153,11 @@ Postgres FTS keyword search, content indexing.
|
|||||||
chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API
|
chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API
|
||||||
(`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle.
|
(`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle.
|
||||||
|
|
||||||
|
**Phase 3 — Async + Scale (DONE):** Redis + BullMQ background indexing/embedding
|
||||||
|
workers (`packages/queue`, `apps/worker`), git-sync webhooks (`POST /hooks/*`),
|
||||||
|
document revision system (`document_revisions` + restore), and MCP Resources
|
||||||
|
(`mcpedia://docs/...`). See `PHASES.md`.
|
||||||
|
|
||||||
> pgvector is **not installed** on the shared imrnes Postgres, so vector storage is
|
> pgvector is **not installed** on the shared imrnes Postgres, so vector storage is
|
||||||
> a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is
|
> a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is
|
||||||
> the Phase-4 scale-out path. See `PHASES.md`.
|
> the Phase-4 scale-out path. See `PHASES.md`.
|
||||||
|
|||||||
@@ -13,6 +13,8 @@
|
|||||||
"@hono/node-server": "^1.13.0",
|
"@hono/node-server": "^1.13.0",
|
||||||
"@mcpedia/config": "workspace:*",
|
"@mcpedia/config": "workspace:*",
|
||||||
"@mcpedia/core": "workspace:*",
|
"@mcpedia/core": "workspace:*",
|
||||||
|
"@mcpedia/db": "workspace:*",
|
||||||
|
"@mcpedia/queue": "workspace:*",
|
||||||
"@trpc/server": "^11.0.0",
|
"@trpc/server": "^11.0.0",
|
||||||
"hono": "^4.6.0",
|
"hono": "^4.6.0",
|
||||||
"zod": "^3.23.8"
|
"zod": "^3.23.8"
|
||||||
|
|||||||
@@ -4,12 +4,31 @@ import { fetchRequestHandler } from "@trpc/server/adapters/fetch";
|
|||||||
import { db } from "@mcpedia/db";
|
import { db } from "@mcpedia/db";
|
||||||
import { appRouter } from "./router";
|
import { appRouter } from "./router";
|
||||||
import type { Context } from "./trpc";
|
import type { Context } from "./trpc";
|
||||||
|
import { enqueueIndexDoc, enqueueFullIndex } from "@mcpedia/queue";
|
||||||
|
|
||||||
const app = new Hono();
|
const app = new Hono();
|
||||||
|
|
||||||
// Health check.
|
// Health check.
|
||||||
app.get("/health", (c) => c.json({ ok: true }));
|
app.get("/health", (c) => c.json({ ok: true }));
|
||||||
|
|
||||||
|
// --- Phase 3: Git synchronization hook ---
|
||||||
|
// POST /hooks/reindex -> enqueue a full-corpus reindex (git push webhook)
|
||||||
|
// POST /hooks/index?slug=... -> enqueue a single document reindex
|
||||||
|
// Returns the created job id(s). The worker processes them asynchronously.
|
||||||
|
app.post("/hooks/reindex", async (c) => {
|
||||||
|
const job = await enqueueFullIndex("git-push");
|
||||||
|
return c.json({ ok: true, jobId: job.id, kind: "full" });
|
||||||
|
});
|
||||||
|
|
||||||
|
app.post("/hooks/index", async (c) => {
|
||||||
|
const slug = c.req.query("slug");
|
||||||
|
if (!slug) return c.json({ ok: false, error: "slug query param required" }, 400);
|
||||||
|
// slug is the relative path without extension, e.g. docs/websocket/contract
|
||||||
|
const relPath = slug.endsWith(".md") || slug.endsWith(".mdx") ? slug : `${slug}.md`;
|
||||||
|
const job = await enqueueIndexDoc(relPath, "git-push");
|
||||||
|
return c.json({ ok: true, jobId: job.id, kind: "doc", relPath });
|
||||||
|
});
|
||||||
|
|
||||||
// Mount tRPC at /trpc/*. The fetch adapter is the canonical Bun/Hono adapter.
|
// Mount tRPC at /trpc/*. The fetch adapter is the canonical Bun/Hono adapter.
|
||||||
app.all("/trpc/*", (c) =>
|
app.all("/trpc/*", (c) =>
|
||||||
fetchRequestHandler({
|
fetchRequestHandler({
|
||||||
|
|||||||
@@ -7,7 +7,12 @@ import {
|
|||||||
keywordSearch,
|
keywordSearch,
|
||||||
listDocuments,
|
listDocuments,
|
||||||
semanticSearch,
|
semanticSearch,
|
||||||
|
listRevisions,
|
||||||
|
getRevision,
|
||||||
|
restoreRevision,
|
||||||
} from "@mcpedia/core";
|
} from "@mcpedia/core";
|
||||||
|
import { getQueue, INDEX_QUEUE } from "@mcpedia/queue";
|
||||||
|
import { getConnection, BULLMQ_PREFIX } from "@mcpedia/queue/client";
|
||||||
|
|
||||||
export const appRouter = router({
|
export const appRouter = router({
|
||||||
search: publicProcedure
|
search: publicProcedure
|
||||||
@@ -33,6 +38,58 @@ export const appRouter = router({
|
|||||||
related: publicProcedure
|
related: publicProcedure
|
||||||
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(20).default(5) }))
|
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(20).default(5) }))
|
||||||
.query(async ({ input }) => getRelated(input.slug, input.limit)),
|
.query(async ({ input }) => getRelated(input.slug, input.limit)),
|
||||||
|
|
||||||
|
// --- Phase 3: revisions ---
|
||||||
|
revisions: publicProcedure
|
||||||
|
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(50).default(20) }))
|
||||||
|
.query(async ({ input }) => listRevisions(input.slug, input.limit)),
|
||||||
|
|
||||||
|
getRevision: publicProcedure
|
||||||
|
.input(z.object({ id: z.string() }))
|
||||||
|
.query(async ({ input }) => getRevision(input.id)),
|
||||||
|
|
||||||
|
restoreRevision: publicProcedure
|
||||||
|
.input(z.object({ id: z.string() }))
|
||||||
|
.mutation(async ({ input }) => restoreRevision(input.id)),
|
||||||
|
|
||||||
|
// --- Phase 3: async job status ---
|
||||||
|
jobStatus: publicProcedure
|
||||||
|
.input(z.object({ id: z.string() }))
|
||||||
|
.query(async ({ input }) => {
|
||||||
|
const queue = getQueue();
|
||||||
|
const job = await queue.getJob(input.id);
|
||||||
|
if (!job) return { exists: false };
|
||||||
|
const state = await job.getState();
|
||||||
|
const failedReason = job.failedReason;
|
||||||
|
const returnvalue = job.returnvalue;
|
||||||
|
const progress = job.progress;
|
||||||
|
return {
|
||||||
|
exists: true,
|
||||||
|
id: job.id,
|
||||||
|
name: job.name,
|
||||||
|
state,
|
||||||
|
progress,
|
||||||
|
failedReason,
|
||||||
|
returnvalue,
|
||||||
|
attemptsMade: job.attemptsMade,
|
||||||
|
};
|
||||||
|
}),
|
||||||
|
|
||||||
|
queueStatus: publicProcedure.query(async () => {
|
||||||
|
const queue = getQueue();
|
||||||
|
const [waiting, active, completed, failed, delayed] = await Promise.all([
|
||||||
|
queue.getWaitingCount(),
|
||||||
|
queue.getActiveCount(),
|
||||||
|
queue.getCompletedCount(),
|
||||||
|
queue.getFailedCount(),
|
||||||
|
queue.getDelayedCount(),
|
||||||
|
]);
|
||||||
|
return {
|
||||||
|
queue: INDEX_QUEUE,
|
||||||
|
prefix: BULLMQ_PREFIX,
|
||||||
|
counts: { waiting, active, completed, failed, delayed },
|
||||||
|
};
|
||||||
|
}),
|
||||||
});
|
});
|
||||||
|
|
||||||
export type AppRouter = typeof appRouter;
|
export type AppRouter = typeof appRouter;
|
||||||
|
|||||||
@@ -17,7 +17,7 @@
|
|||||||
"@mcpedia/core": "workspace:*",
|
"@mcpedia/core": "workspace:*",
|
||||||
"@mcpedia/search": "workspace:*",
|
"@mcpedia/search": "workspace:*",
|
||||||
"@modelcontextprotocol/sdk": "^1.29.0",
|
"@modelcontextprotocol/sdk": "^1.29.0",
|
||||||
"zod": "^3.23.8"
|
"zod": "^4.0.0"
|
||||||
},
|
},
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
"@types/node": "^20",
|
"@types/node": "^20",
|
||||||
|
|||||||
+123
-1
@@ -1,7 +1,19 @@
|
|||||||
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
|
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
|
||||||
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
|
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
|
||||||
|
import { ResourceTemplate } from "@modelcontextprotocol/sdk/server/mcp.js";
|
||||||
import { z } from "zod";
|
import { z } from "zod";
|
||||||
import { listDocuments, getDocument, getRelated, semanticSearch, hybridSearch, keywordSearch } from "@mcpedia/core";
|
import {
|
||||||
|
listDocuments,
|
||||||
|
getDocument,
|
||||||
|
getRelated,
|
||||||
|
semanticSearch,
|
||||||
|
hybridSearch,
|
||||||
|
keywordSearch,
|
||||||
|
listRevisions,
|
||||||
|
readContentFile,
|
||||||
|
} from "@mcpedia/core";
|
||||||
|
import { CONTENT_ROOT } from "@mcpedia/config";
|
||||||
|
import { join } from "node:path";
|
||||||
|
|
||||||
export function createMcpServer(): McpServer {
|
export function createMcpServer(): McpServer {
|
||||||
const server = new McpServer({
|
const server = new McpServer({
|
||||||
@@ -119,6 +131,116 @@ export function createMcpServer(): McpServer {
|
|||||||
},
|
},
|
||||||
);
|
);
|
||||||
|
|
||||||
|
// --- Phase 3: MCP Resources (read-only knowledge base surfaced via URIs) ---
|
||||||
|
// mcpedia://docs -> list all published documents
|
||||||
|
// mcpedia://docs/{slug} -> full markdown body (from disk)
|
||||||
|
// mcpedia://docs/{slug}/chunks -> chunked preview (semantic slices)
|
||||||
|
// mcpedia://docs/{slug}/revisions -> revision history summary
|
||||||
|
server.registerResource(
|
||||||
|
"mcpedia-docs-list",
|
||||||
|
"mcpedia://docs",
|
||||||
|
{
|
||||||
|
title: "MCPedia document index",
|
||||||
|
description: "List of all published documents in the knowledge base.",
|
||||||
|
mimeType: "application/json",
|
||||||
|
},
|
||||||
|
async (uri) => {
|
||||||
|
const docs = await listDocuments();
|
||||||
|
return {
|
||||||
|
contents: [
|
||||||
|
{
|
||||||
|
uri: uri.href,
|
||||||
|
mimeType: "application/json",
|
||||||
|
text: JSON.stringify(docs, null, 2),
|
||||||
|
},
|
||||||
|
],
|
||||||
|
};
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
|
server.registerResource(
|
||||||
|
"mcpedia-doc-chunks",
|
||||||
|
new ResourceTemplate("mcpedia://docs/{+slug}/chunks", { list: undefined }),
|
||||||
|
{
|
||||||
|
title: "MCPedia document chunks",
|
||||||
|
description: "Preview of the embedded semantic chunks for a document.",
|
||||||
|
mimeType: "application/json",
|
||||||
|
},
|
||||||
|
async (uri, vars) => {
|
||||||
|
const slug = String(vars.slug);
|
||||||
|
const doc = await getDocument(slug);
|
||||||
|
if (!doc) throw new Error(`Document not found: ${slug}`);
|
||||||
|
// Chunk the body the same way the indexer does (size 1000 / overlap 150)
|
||||||
|
// so the resource mirrors what semantic search actually sees.
|
||||||
|
const { chunkText } = await import("@mcpedia/embeddings");
|
||||||
|
const chunks = chunkText(doc.body, { size: 1000, overlap: 150 });
|
||||||
|
return {
|
||||||
|
contents: [
|
||||||
|
{
|
||||||
|
uri: uri.href,
|
||||||
|
mimeType: "application/json",
|
||||||
|
text: JSON.stringify(
|
||||||
|
chunks.map((c, i) => ({ index: i, length: c.length, preview: c.slice(0, 200) })),
|
||||||
|
null,
|
||||||
|
2,
|
||||||
|
),
|
||||||
|
},
|
||||||
|
],
|
||||||
|
};
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
|
server.registerResource(
|
||||||
|
"mcpedia-doc-revisions",
|
||||||
|
new ResourceTemplate("mcpedia://docs/{+slug}/revisions", { list: undefined }),
|
||||||
|
{
|
||||||
|
title: "MCPedia document revisions",
|
||||||
|
description: "Revision history summary for a document.",
|
||||||
|
mimeType: "application/json",
|
||||||
|
},
|
||||||
|
async (uri, vars) => {
|
||||||
|
const slug = String(vars.slug);
|
||||||
|
const revs = await listRevisions(slug, 20);
|
||||||
|
return {
|
||||||
|
contents: [
|
||||||
|
{
|
||||||
|
uri: uri.href,
|
||||||
|
mimeType: "application/json",
|
||||||
|
text: JSON.stringify(revs, null, 2),
|
||||||
|
},
|
||||||
|
],
|
||||||
|
};
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
|
// Registered LAST: the bare {+slug} template is greedy and would otherwise
|
||||||
|
// swallow /chunks and /revisions URIs. Specific templates must match first.
|
||||||
|
server.registerResource(
|
||||||
|
"mcpedia-doc",
|
||||||
|
new ResourceTemplate("mcpedia://docs/{+slug}", { list: undefined }),
|
||||||
|
{
|
||||||
|
title: "MCPedia document",
|
||||||
|
description: "Full markdown body of a single document, read from disk (source of truth).",
|
||||||
|
mimeType: "text/markdown",
|
||||||
|
},
|
||||||
|
async (uri, vars) => {
|
||||||
|
const slug = String(vars.slug);
|
||||||
|
const doc = await getDocument(slug);
|
||||||
|
if (!doc) {
|
||||||
|
throw new Error(`Document not found: ${slug}`);
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
contents: [
|
||||||
|
{
|
||||||
|
uri: uri.href,
|
||||||
|
mimeType: "text/markdown",
|
||||||
|
text: doc.body,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
};
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
return server;
|
return server;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -92,6 +92,35 @@ async function main() {
|
|||||||
}
|
}
|
||||||
console.log(`hybrid_search => ${hybHits.length} docs, top: ${hybHits[0].doc.slug}`);
|
console.log(`hybrid_search => ${hybHits.length} docs, top: ${hybHits[0].doc.slug}`);
|
||||||
|
|
||||||
|
// 8) resources: list
|
||||||
|
const resList = await client.listResources();
|
||||||
|
const resNames = resList.resources.map((r: any) => r.name).sort();
|
||||||
|
console.log("resources:", resNames.join(", "));
|
||||||
|
if (!resNames.includes("mcpedia-docs-list")) {
|
||||||
|
throw new Error("expected mcpedia-docs-list resource");
|
||||||
|
}
|
||||||
|
|
||||||
|
// 9) resource: read the docs list (must not throw, returns JSON content)
|
||||||
|
const readList = await client.readResource({ uri: "mcpedia://docs" });
|
||||||
|
const listText = (readList.contents as any)[0].text;
|
||||||
|
if (!listText.includes("docs/websocket/contract")) {
|
||||||
|
throw new Error("mcpedia://docs did not list the websocket contract doc");
|
||||||
|
}
|
||||||
|
console.log("readResource(mcpedia://docs) => ok");
|
||||||
|
|
||||||
|
// 10) resource: read a single doc body + revisions
|
||||||
|
const readDoc = await client.readResource({ uri: "mcpedia://docs/docs/websocket/contract" });
|
||||||
|
const docText = (readDoc.contents as any)[0].text;
|
||||||
|
if (!docText.includes("WebSocket Contract")) {
|
||||||
|
throw new Error("mcpedia://docs/{slug} returned unexpected body");
|
||||||
|
}
|
||||||
|
console.log("readResource(mcpedia://docs/docs/websocket/contract) => ok");
|
||||||
|
|
||||||
|
const readRev = await client.readResource({
|
||||||
|
uri: "mcpedia://docs/docs/websocket/contract/revisions",
|
||||||
|
});
|
||||||
|
console.log("readResource(.../revisions) => ok");
|
||||||
|
|
||||||
await client.close();
|
await client.close();
|
||||||
await server.close();
|
await server.close();
|
||||||
console.log("\nSMOKE OK");
|
console.log("\nSMOKE OK");
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
{
|
||||||
|
"name": "@mcpedia/worker",
|
||||||
|
"version": "0.1.0",
|
||||||
|
"private": true,
|
||||||
|
"type": "module",
|
||||||
|
"scripts": {
|
||||||
|
"start": "bun run src/index.ts",
|
||||||
|
"lint": "tsc --noEmit",
|
||||||
|
"typecheck": "tsc --noEmit"
|
||||||
|
},
|
||||||
|
"dependencies": {
|
||||||
|
"@mcpedia/config": "workspace:*",
|
||||||
|
"@mcpedia/core": "workspace:*",
|
||||||
|
"@mcpedia/db": "workspace:*",
|
||||||
|
"@mcpedia/queue": "workspace:*"
|
||||||
|
},
|
||||||
|
"devDependencies": {
|
||||||
|
"typescript": "^5.6.0"
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
import { startWorker } from "@mcpedia/queue/worker";
|
||||||
|
|
||||||
|
// Keep the process alive: the worker listens on the BullMQ queue until a
|
||||||
|
// SIGINT/SIGTERM closes it (handled inside startWorker).
|
||||||
|
const worker = await startWorker();
|
||||||
|
|
||||||
|
// Heartbeat so the supervisor/operator can see liveness without scraping logs.
|
||||||
|
const heartbeat = setInterval(() => {
|
||||||
|
console.log(`[worker] alive, ${worker.name} queue="${worker.name}"`);
|
||||||
|
}, 30_000);
|
||||||
|
|
||||||
|
worker.on("closed", () => clearInterval(heartbeat));
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
{
|
||||||
|
"extends": "../../tsconfig.base.json",
|
||||||
|
"compilerOptions": {
|
||||||
|
"paths": {
|
||||||
|
"@mcpedia/db": ["../../packages/db/src/index.ts"],
|
||||||
|
"@mcpedia/db/schema": ["../../packages/db/src/schema.ts"],
|
||||||
|
"@mcpedia/*": ["../../packages/*"]
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"include": ["src/**/*.ts"]
|
||||||
|
}
|
||||||
@@ -14,6 +14,8 @@
|
|||||||
"lint": "turbo run lint",
|
"lint": "turbo run lint",
|
||||||
"typecheck": "turbo run typecheck",
|
"typecheck": "turbo run typecheck",
|
||||||
"index": "bun run scripts/indexer.ts",
|
"index": "bun run scripts/indexer.ts",
|
||||||
|
"enqueue": "bun run scripts/enqueue.ts",
|
||||||
|
"worker": "bun --cwd apps/worker run start",
|
||||||
"mcp": "bun --cwd apps/mcp run start",
|
"mcp": "bun --cwd apps/mcp run start",
|
||||||
"api": "bun --cwd apps/api run dev"
|
"api": "bun --cwd apps/api run dev"
|
||||||
},
|
},
|
||||||
|
|||||||
@@ -40,6 +40,12 @@ export const EMBED_BASE_URL = process.env.EMBED_BASE_URL ?? "";
|
|||||||
export const EMBED_API_KEY = process.env.EMBED_API_KEY ?? "";
|
export const EMBED_API_KEY = process.env.EMBED_API_KEY ?? "";
|
||||||
export const EMBED_MODEL = process.env.EMBED_MODEL ?? "";
|
export const EMBED_MODEL = process.env.EMBED_MODEL ?? "";
|
||||||
|
|
||||||
|
// Phase 3: Redis + BullMQ (shared imrnes Redis, no auth by default).
|
||||||
|
export const REDIS_URL = process.env.REDIS_URL ?? "redis://100.121.180.82:6379";
|
||||||
|
export const REDIS_PASSWORD = process.env.REDIS_PASSWORD ?? "";
|
||||||
|
// BullMQ key prefix to namespace jobs on the shared Redis instance.
|
||||||
|
export const QUEUE_PREFIX = process.env.QUEUE_PREFIX ?? "mcpedia";
|
||||||
|
|
||||||
if (!DATABASE_URL) {
|
if (!DATABASE_URL) {
|
||||||
// Fail fast with an explicit message instead of a cryptic driver error.
|
// Fail fast with an explicit message instead of a cryptic driver error.
|
||||||
throw new Error(
|
throw new Error(
|
||||||
|
|||||||
@@ -0,0 +1,155 @@
|
|||||||
|
import { db } from "@mcpedia/db";
|
||||||
|
import { documents, documentRevisions, documentChunks } from "@mcpedia/db/schema";
|
||||||
|
import { parseFile } from "@mcpedia/parser";
|
||||||
|
import { CONTENT_ROOT } from "@mcpedia/config";
|
||||||
|
import { listContentFiles } from "./content.service";
|
||||||
|
import { indexChunks } from "./document.service";
|
||||||
|
import { toMeta } from "./row-map";
|
||||||
|
import { eq, desc, and, sql } from "drizzle-orm";
|
||||||
|
import { join } from "node:path";
|
||||||
|
|
||||||
|
export interface IndexResult {
|
||||||
|
indexed: number;
|
||||||
|
chunks: number;
|
||||||
|
revisions: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Index a single content file: parse → upsert `documents` → chunk+embed →
|
||||||
|
* snapshot a revision if the body changed since the last indexed revision.
|
||||||
|
*
|
||||||
|
* This is THE single indexing entry point shared by the CLI script, the
|
||||||
|
* BullMQ worker, and the git-sync hook — no business logic is duplicated.
|
||||||
|
*
|
||||||
|
* @param relPath path relative to CONTENT_ROOT (e.g. "docs/websocket/contract")
|
||||||
|
* @param reason provenance tag for the revision ("index" | "git-push" | "reindex")
|
||||||
|
*/
|
||||||
|
export async function indexContentFile(
|
||||||
|
relPath: string,
|
||||||
|
reason = "index",
|
||||||
|
): Promise<{ indexed: boolean; chunks: number; revision: boolean }> {
|
||||||
|
const abs = join(CONTENT_ROOT, relPath);
|
||||||
|
const { meta, body } = parseFile(abs, relPath);
|
||||||
|
const nowIso =
|
||||||
|
meta.updatedAt && meta.updatedAt !== ""
|
||||||
|
? meta.updatedAt
|
||||||
|
: new Date().toISOString();
|
||||||
|
|
||||||
|
await db
|
||||||
|
.insert(documents)
|
||||||
|
.values({
|
||||||
|
id: meta.id,
|
||||||
|
slug: meta.slug,
|
||||||
|
title: meta.title,
|
||||||
|
type: meta.type,
|
||||||
|
section: meta.section,
|
||||||
|
status: meta.status,
|
||||||
|
author: meta.author,
|
||||||
|
tags: meta.tags,
|
||||||
|
path: meta.path,
|
||||||
|
body,
|
||||||
|
createdAt: new Date(meta.createdAt || nowIso),
|
||||||
|
updatedAt: new Date(nowIso),
|
||||||
|
})
|
||||||
|
.onConflictDoUpdate({
|
||||||
|
target: documents.slug,
|
||||||
|
set: {
|
||||||
|
title: meta.title,
|
||||||
|
type: meta.type,
|
||||||
|
section: meta.section,
|
||||||
|
status: meta.status,
|
||||||
|
author: meta.author,
|
||||||
|
tags: meta.tags,
|
||||||
|
path: meta.path,
|
||||||
|
body,
|
||||||
|
updatedAt: new Date(nowIso),
|
||||||
|
},
|
||||||
|
});
|
||||||
|
|
||||||
|
// Semantic chunks (embedding). A failure here must not abort the whole
|
||||||
|
// index — log and continue; FTS still works without embeddings.
|
||||||
|
let chunks = 0;
|
||||||
|
try {
|
||||||
|
chunks = await indexChunks(meta.slug, body);
|
||||||
|
} catch (err) {
|
||||||
|
console.error(
|
||||||
|
` embed FAILED for ${meta.slug}: ${err instanceof Error ? err.message : err}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Snapshot a revision only when the body actually changed vs the latest
|
||||||
|
// revision. Pure metadata/index changes (tags/title) won't create noise.
|
||||||
|
const revision = await snapshotRevision(meta.slug, meta, body, reason);
|
||||||
|
|
||||||
|
return { indexed: true, chunks, revision };
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Compare the incoming body against the latest revision's body; if different
|
||||||
|
* (or no prior revision exists), create a new revision with an incremented
|
||||||
|
* per-document revisionNo.
|
||||||
|
*/
|
||||||
|
async function snapshotRevision(
|
||||||
|
slug: string,
|
||||||
|
meta: ReturnType<typeof parseFile>["meta"],
|
||||||
|
body: string,
|
||||||
|
reason: string,
|
||||||
|
): Promise<boolean> {
|
||||||
|
const [doc] = await db
|
||||||
|
.select({ id: documents.id })
|
||||||
|
.from(documents)
|
||||||
|
.where(eq(documents.slug, slug));
|
||||||
|
if (!doc) return false;
|
||||||
|
|
||||||
|
const [latest] = await db
|
||||||
|
.select({ body: documentRevisions.body, revisionNo: documentRevisions.revisionNo })
|
||||||
|
.from(documentRevisions)
|
||||||
|
.where(eq(documentRevisions.documentId, doc.id))
|
||||||
|
.orderBy(desc(documentRevisions.revisionNo))
|
||||||
|
.limit(1);
|
||||||
|
|
||||||
|
if (latest && latest.body === body) {
|
||||||
|
return false; // unchanged → no new revision
|
||||||
|
}
|
||||||
|
|
||||||
|
const nextNo = (latest?.revisionNo ?? 0) + 1;
|
||||||
|
await db.insert(documentRevisions).values({
|
||||||
|
documentId: doc.id,
|
||||||
|
slug,
|
||||||
|
revisionNo: nextNo,
|
||||||
|
title: meta.title,
|
||||||
|
body,
|
||||||
|
meta: {
|
||||||
|
type: meta.type,
|
||||||
|
section: meta.section,
|
||||||
|
status: meta.status,
|
||||||
|
author: meta.author,
|
||||||
|
tags: meta.tags,
|
||||||
|
},
|
||||||
|
reason,
|
||||||
|
});
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Walk the entire content tree and index every file. Returns aggregate counts.
|
||||||
|
*/
|
||||||
|
export async function runFullIndex(reason = "index"): Promise<IndexResult> {
|
||||||
|
const files = listContentFiles();
|
||||||
|
let indexed = 0;
|
||||||
|
let chunks = 0;
|
||||||
|
let revisions = 0;
|
||||||
|
for (const rel of files) {
|
||||||
|
const r = await indexContentFile(rel, reason);
|
||||||
|
indexed++;
|
||||||
|
chunks += r.chunks;
|
||||||
|
if (r.revision) revisions++;
|
||||||
|
console.log(
|
||||||
|
` indexed ${rel}${r.chunks ? ` (${r.chunks} chunks)` : ""}${r.revision ? " [revision]" : ""}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
console.log(
|
||||||
|
`indexed ${indexed} documents, ${chunks} chunks, ${revisions} new revisions`,
|
||||||
|
);
|
||||||
|
return { indexed, chunks, revisions };
|
||||||
|
}
|
||||||
@@ -1,6 +1,8 @@
|
|||||||
export * from "./content.service";
|
export * from "./content.service";
|
||||||
export * from "./document.service";
|
export * from "./document.service";
|
||||||
export * from "./search.service";
|
export * from "./search.service";
|
||||||
|
export * from "./index.service";
|
||||||
|
export * from "./revision.service";
|
||||||
export { toMeta } from "./row-map";
|
export { toMeta } from "./row-map";
|
||||||
|
|
||||||
export type {
|
export type {
|
||||||
|
|||||||
@@ -0,0 +1,116 @@
|
|||||||
|
import { db } from "@mcpedia/db";
|
||||||
|
import { documents, documentRevisions, documentChunks } from "@mcpedia/db/schema";
|
||||||
|
import { eq, desc, and, sql } from "drizzle-orm";
|
||||||
|
import { toMeta } from "./row-map";
|
||||||
|
import type { DocumentMeta } from "@mcpedia/types";
|
||||||
|
|
||||||
|
export interface RevisionSummary {
|
||||||
|
id: string;
|
||||||
|
slug: string;
|
||||||
|
revisionNo: number;
|
||||||
|
title: string;
|
||||||
|
reason: string;
|
||||||
|
createdAt: string;
|
||||||
|
bodyLength: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** List revisions for a slug, newest first. */
|
||||||
|
export async function listRevisions(
|
||||||
|
slug: string,
|
||||||
|
limit = 20,
|
||||||
|
): Promise<RevisionSummary[]> {
|
||||||
|
const [doc] = await db
|
||||||
|
.select({ id: documents.id })
|
||||||
|
.from(documents)
|
||||||
|
.where(eq(documents.slug, slug));
|
||||||
|
if (!doc) return [];
|
||||||
|
|
||||||
|
const rows = await db
|
||||||
|
.select({
|
||||||
|
id: documentRevisions.id,
|
||||||
|
slug: documentRevisions.slug,
|
||||||
|
revisionNo: documentRevisions.revisionNo,
|
||||||
|
title: documentRevisions.title,
|
||||||
|
reason: documentRevisions.reason,
|
||||||
|
createdAt: documentRevisions.createdAt,
|
||||||
|
bodyLength: sql<number>`length(${documentRevisions.body})`,
|
||||||
|
})
|
||||||
|
.from(documentRevisions)
|
||||||
|
.where(eq(documentRevisions.documentId, doc.id))
|
||||||
|
.orderBy(desc(documentRevisions.revisionNo))
|
||||||
|
.limit(limit);
|
||||||
|
|
||||||
|
return rows.map((r) => ({
|
||||||
|
id: r.id,
|
||||||
|
slug: r.slug,
|
||||||
|
revisionNo: r.revisionNo,
|
||||||
|
title: r.title,
|
||||||
|
reason: r.reason,
|
||||||
|
createdAt: r.createdAt.toISOString(),
|
||||||
|
bodyLength: r.bodyLength,
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Fetch a single revision's full body. */
|
||||||
|
export async function getRevision(
|
||||||
|
id: string,
|
||||||
|
): Promise<{ id: string; revisionNo: number; body: string; meta: unknown } | null> {
|
||||||
|
const [row] = await db
|
||||||
|
.select({
|
||||||
|
id: documentRevisions.id,
|
||||||
|
revisionNo: documentRevisions.revisionNo,
|
||||||
|
body: documentRevisions.body,
|
||||||
|
meta: documentRevisions.meta,
|
||||||
|
})
|
||||||
|
.from(documentRevisions)
|
||||||
|
.where(eq(documentRevisions.id, id));
|
||||||
|
if (!row) return null;
|
||||||
|
return {
|
||||||
|
id: row.id,
|
||||||
|
revisionNo: row.revisionNo,
|
||||||
|
body: row.body,
|
||||||
|
meta: row.meta,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Restore a revision: write its body+metadata back into the live `documents` row. */
|
||||||
|
export async function restoreRevision(
|
||||||
|
id: string,
|
||||||
|
): Promise<{ slug: string; documentId: string } | null> {
|
||||||
|
const [rev] = await db
|
||||||
|
.select({
|
||||||
|
id: documentRevisions.id,
|
||||||
|
documentId: documentRevisions.documentId,
|
||||||
|
slug: documentRevisions.slug,
|
||||||
|
title: documentRevisions.title,
|
||||||
|
body: documentRevisions.body,
|
||||||
|
meta: documentRevisions.meta,
|
||||||
|
})
|
||||||
|
.from(documentRevisions)
|
||||||
|
.where(eq(documentRevisions.id, id));
|
||||||
|
if (!rev) return null;
|
||||||
|
|
||||||
|
const m = rev.meta as {
|
||||||
|
type?: string;
|
||||||
|
section?: string;
|
||||||
|
status?: string;
|
||||||
|
author?: string;
|
||||||
|
tags?: string[];
|
||||||
|
};
|
||||||
|
|
||||||
|
await db
|
||||||
|
.update(documents)
|
||||||
|
.set({
|
||||||
|
title: rev.title,
|
||||||
|
type: (m.type as any) ?? "documentation",
|
||||||
|
section: (m.section as any) ?? "docs",
|
||||||
|
status: (m.status as any) ?? "published",
|
||||||
|
author: m.author ?? "",
|
||||||
|
tags: m.tags ?? [],
|
||||||
|
body: rev.body,
|
||||||
|
updatedAt: new Date(),
|
||||||
|
})
|
||||||
|
.where(eq(documents.id, rev.documentId));
|
||||||
|
|
||||||
|
return { slug: rev.slug, documentId: rev.documentId };
|
||||||
|
}
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
CREATE TABLE "document_revisions" (
|
||||||
|
"id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL,
|
||||||
|
"document_id" text NOT NULL,
|
||||||
|
"slug" text NOT NULL,
|
||||||
|
"revision_no" integer NOT NULL,
|
||||||
|
"title" text NOT NULL,
|
||||||
|
"body" text NOT NULL,
|
||||||
|
"meta" jsonb NOT NULL,
|
||||||
|
"reason" text DEFAULT 'index' NOT NULL,
|
||||||
|
"created_at" timestamp with time zone DEFAULT now() NOT NULL
|
||||||
|
);
|
||||||
|
--> statement-breakpoint
|
||||||
|
CREATE INDEX "document_revisions_document_id_idx" ON "document_revisions" USING btree ("document_id");
|
||||||
|
--> statement-breakpoint
|
||||||
|
CREATE INDEX "document_revisions_slug_idx" ON "document_revisions" USING btree ("slug");
|
||||||
|
--> statement-breakpoint
|
||||||
|
CREATE INDEX "document_revisions_doc_rev_idx" ON "document_revisions" USING btree ("document_id", "revision_no" DESC);
|
||||||
|
--> statement-breakpoint
|
||||||
|
ALTER TABLE "document_revisions" ADD CONSTRAINT "document_revisions_document_id_documents_id_fk" FOREIGN KEY ("document_id") REFERENCES "public"."documents"("id") ON DELETE cascade;
|
||||||
@@ -0,0 +1,110 @@
|
|||||||
|
{
|
||||||
|
"id": "0002_document_revisions",
|
||||||
|
"prevId": "0001_document_chunks",
|
||||||
|
"version": "7",
|
||||||
|
"dialect": "postgresql",
|
||||||
|
"tables": {
|
||||||
|
"document_revisions": {
|
||||||
|
"name": "document_revisions",
|
||||||
|
"columns": {
|
||||||
|
"id": {
|
||||||
|
"name": "id",
|
||||||
|
"type": "uuid",
|
||||||
|
"primaryKey": true,
|
||||||
|
"notNull": true,
|
||||||
|
"default": "gen_random_uuid()"
|
||||||
|
},
|
||||||
|
"document_id": {
|
||||||
|
"name": "document_id",
|
||||||
|
"type": "text",
|
||||||
|
"notNull": true
|
||||||
|
},
|
||||||
|
"slug": {
|
||||||
|
"name": "slug",
|
||||||
|
"type": "text",
|
||||||
|
"notNull": true
|
||||||
|
},
|
||||||
|
"revision_no": {
|
||||||
|
"name": "revision_no",
|
||||||
|
"type": "integer",
|
||||||
|
"notNull": true
|
||||||
|
},
|
||||||
|
"title": {
|
||||||
|
"name": "title",
|
||||||
|
"type": "text",
|
||||||
|
"notNull": true
|
||||||
|
},
|
||||||
|
"body": {
|
||||||
|
"name": "body",
|
||||||
|
"type": "text",
|
||||||
|
"notNull": true
|
||||||
|
},
|
||||||
|
"meta": {
|
||||||
|
"name": "meta",
|
||||||
|
"type": "jsonb",
|
||||||
|
"notNull": true
|
||||||
|
},
|
||||||
|
"reason": {
|
||||||
|
"name": "reason",
|
||||||
|
"type": "text",
|
||||||
|
"notNull": true,
|
||||||
|
"default": "'index'"
|
||||||
|
},
|
||||||
|
"created_at": {
|
||||||
|
"name": "created_at",
|
||||||
|
"type": "timestamp",
|
||||||
|
"notNull": true,
|
||||||
|
"default": "now()"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"indexes": {
|
||||||
|
"document_revisions_document_id_idx": {
|
||||||
|
"name": "document_revisions_document_id_idx",
|
||||||
|
"columns": [
|
||||||
|
{ "name": "document_id", "asc": true }
|
||||||
|
],
|
||||||
|
"isUnique": false
|
||||||
|
},
|
||||||
|
"document_revisions_slug_idx": {
|
||||||
|
"name": "document_revisions_slug_idx",
|
||||||
|
"columns": [
|
||||||
|
{ "name": "slug", "asc": true }
|
||||||
|
],
|
||||||
|
"isUnique": false
|
||||||
|
},
|
||||||
|
"document_revisions_doc_rev_idx": {
|
||||||
|
"name": "document_revisions_doc_rev_idx",
|
||||||
|
"columns": [
|
||||||
|
{ "name": "document_id", "asc": true },
|
||||||
|
{ "name": "revision_no", "asc": false }
|
||||||
|
],
|
||||||
|
"isUnique": false
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"foreignKeys": {
|
||||||
|
"document_revisions_document_id_documents_id_fk": {
|
||||||
|
"name": "document_revisions_document_id_documents_id_fk",
|
||||||
|
"columns": ["document_id"],
|
||||||
|
"referenceTable": "documents",
|
||||||
|
"referenceColumns": ["id"],
|
||||||
|
"onDelete": "cascade"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"compositePrimaryKeys": {},
|
||||||
|
"uniqueConstraints": {},
|
||||||
|
"policies": {}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"enums": {},
|
||||||
|
"schemas": {},
|
||||||
|
"sequences": {},
|
||||||
|
"roles": {},
|
||||||
|
"policies": {},
|
||||||
|
"views": {},
|
||||||
|
"extensions": {},
|
||||||
|
"_meta": {
|
||||||
|
"columns": {},
|
||||||
|
"schemas": {},
|
||||||
|
"tables": {}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -15,6 +15,13 @@
|
|||||||
"when": 1787137149735,
|
"when": 1787137149735,
|
||||||
"tag": "0001_document_chunks",
|
"tag": "0001_document_chunks",
|
||||||
"breakpoints": true
|
"breakpoints": true
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"idx": 2,
|
||||||
|
"version": "7",
|
||||||
|
"when": 1787139150000,
|
||||||
|
"tag": "0002_document_revisions",
|
||||||
|
"breakpoints": true
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -3,6 +3,7 @@ import {
|
|||||||
customType,
|
customType,
|
||||||
index,
|
index,
|
||||||
integer,
|
integer,
|
||||||
|
jsonb,
|
||||||
pgTable,
|
pgTable,
|
||||||
real,
|
real,
|
||||||
text,
|
text,
|
||||||
@@ -77,5 +78,48 @@ export const documentChunks = pgTable(
|
|||||||
export type DocumentChunkRow = typeof documentChunks.$inferSelect;
|
export type DocumentChunkRow = typeof documentChunks.$inferSelect;
|
||||||
export type NewDocumentChunkRow = typeof documentChunks.$inferInsert;
|
export type NewDocumentChunkRow = typeof documentChunks.$inferInsert;
|
||||||
|
|
||||||
|
// Phase 3: document revision system. Each row is an immutable snapshot of a
|
||||||
|
// document's body + metadata at a point in time (taken by the indexer whenever
|
||||||
|
// the body actually changes). revisionNo is per-document and monotonically
|
||||||
|
// increasing so the latest revision is always max(revision_no).
|
||||||
|
export const documentRevisions = pgTable(
|
||||||
|
"document_revisions",
|
||||||
|
{
|
||||||
|
id: uuid("id").primaryKey().defaultRandom(),
|
||||||
|
documentId: text("document_id")
|
||||||
|
.notNull()
|
||||||
|
.references(() => documents.id, { onDelete: "cascade" }),
|
||||||
|
slug: text("slug").notNull(),
|
||||||
|
revisionNo: integer("revision_no").notNull(),
|
||||||
|
title: text("title").notNull(),
|
||||||
|
body: text("body").notNull(),
|
||||||
|
// Metadata snapshot (type/section/status/author/tags) as JSON so a revision
|
||||||
|
// is self-describing even if the live document is later restructured.
|
||||||
|
meta: jsonb("meta").notNull().$type<{
|
||||||
|
type: string;
|
||||||
|
section: string;
|
||||||
|
status: string;
|
||||||
|
author: string;
|
||||||
|
tags: string[];
|
||||||
|
}>(),
|
||||||
|
// Why this revision was created (e.g. "index", "git-push", "restore").
|
||||||
|
reason: text("reason").notNull().default("index"),
|
||||||
|
createdAt: timestamp("created_at", { withTimezone: true })
|
||||||
|
.notNull()
|
||||||
|
.defaultNow(),
|
||||||
|
},
|
||||||
|
(t) => ({
|
||||||
|
docIdx: index("document_revisions_document_id_idx").on(t.documentId),
|
||||||
|
slugIdx: index("document_revisions_slug_idx").on(t.slug),
|
||||||
|
docRevIdx: index("document_revisions_doc_rev_idx").on(
|
||||||
|
t.documentId,
|
||||||
|
sql`${t.revisionNo} desc`,
|
||||||
|
),
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
|
||||||
|
export type DocumentRevisionRow = typeof documentRevisions.$inferSelect;
|
||||||
|
export type NewDocumentRevisionRow = typeof documentRevisions.$inferInsert;
|
||||||
|
|
||||||
export type DocumentRow = typeof documents.$inferSelect;
|
export type DocumentRow = typeof documents.$inferSelect;
|
||||||
export type NewDocumentRow = typeof documents.$inferInsert;
|
export type NewDocumentRow = typeof documents.$inferInsert;
|
||||||
|
|||||||
@@ -0,0 +1,22 @@
|
|||||||
|
{
|
||||||
|
"name": "@mcpedia/queue",
|
||||||
|
"version": "0.1.0",
|
||||||
|
"private": true,
|
||||||
|
"type": "module",
|
||||||
|
"exports": {
|
||||||
|
".": "./src/index.ts",
|
||||||
|
"./client": "./src/client.ts",
|
||||||
|
"./queue": "./src/queue.ts",
|
||||||
|
"./worker": "./src/worker.ts"
|
||||||
|
},
|
||||||
|
"dependencies": {
|
||||||
|
"@mcpedia/config": "workspace:*",
|
||||||
|
"@mcpedia/core": "workspace:*",
|
||||||
|
"@mcpedia/db": "workspace:*",
|
||||||
|
"bullmq": "^6.1.2",
|
||||||
|
"ioredis": "^6.0.0"
|
||||||
|
},
|
||||||
|
"devDependencies": {
|
||||||
|
"typescript": "^5.6.0"
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
import { REDIS_URL, REDIS_PASSWORD, QUEUE_PREFIX } from "@mcpedia/config";
|
||||||
|
import IORedis, { type RedisOptions } from "ioredis";
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Shared ioredis connection for BullMQ. BullMQ requires an ioredis instance and
|
||||||
|
* internally duplicates it for blocking commands, so we keep the option objects
|
||||||
|
* explicit (maxRetriesPerRequest: null is REQUIRED for the blocking
|
||||||
|
* connection — a finite retry count causes "Connection in key mode" errors).
|
||||||
|
*/
|
||||||
|
function buildOptions(): RedisOptions {
|
||||||
|
const opts: RedisOptions = {
|
||||||
|
maxRetriesPerRequest: null,
|
||||||
|
lazyConnect: true,
|
||||||
|
enableOfflineQueue: true,
|
||||||
|
};
|
||||||
|
if (REDIS_PASSWORD) opts.password = REDIS_PASSWORD;
|
||||||
|
return opts;
|
||||||
|
}
|
||||||
|
|
||||||
|
let _connection: IORedis | null = null;
|
||||||
|
|
||||||
|
/** Lazily-created singleton ioredis connection. */
|
||||||
|
export function getConnection(): IORedis {
|
||||||
|
if (!_connection) {
|
||||||
|
_connection = new IORedis(REDIS_URL, buildOptions());
|
||||||
|
_connection.on("error", (err) => {
|
||||||
|
// Log but don't crash the process on transient Redis errors.
|
||||||
|
console.error("[queue] redis error:", err.message);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
return _connection;
|
||||||
|
}
|
||||||
|
|
||||||
|
export const BULLMQ_PREFIX = QUEUE_PREFIX;
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
export { getConnection, BULLMQ_PREFIX } from "./client";
|
||||||
|
export {
|
||||||
|
getQueue,
|
||||||
|
enqueueIndexDoc,
|
||||||
|
enqueueFullIndex,
|
||||||
|
INDEX_QUEUE,
|
||||||
|
} from "./queue";
|
||||||
|
export type { IndexDocJobData, IndexAllJobData, JobType } from "./queue";
|
||||||
|
export { createWorker, startWorker } from "./worker";
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
import { Queue, type Job } from "bullmq";
|
||||||
|
import { getConnection, BULLMQ_PREFIX } from "./client";
|
||||||
|
|
||||||
|
export const INDEX_QUEUE = "mcpedia-index";
|
||||||
|
|
||||||
|
/** Lazily-created singleton BullMQ queue. */
|
||||||
|
let _queue: Queue | null = null;
|
||||||
|
|
||||||
|
export function getQueue(): Queue {
|
||||||
|
if (!_queue) {
|
||||||
|
_queue = new Queue(INDEX_QUEUE, {
|
||||||
|
connection: getConnection(),
|
||||||
|
prefix: BULLMQ_PREFIX,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
return _queue;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface IndexDocJobData {
|
||||||
|
relPath: string;
|
||||||
|
reason: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface IndexAllJobData {
|
||||||
|
reason: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export type JobType = "index-doc" | "index-all";
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Enqueue a single-document reindex job. Keyed by slug so repeated edits
|
||||||
|
* collapse into one pending job (BullMQ dedup by jobId within the window).
|
||||||
|
*/
|
||||||
|
export async function enqueueIndexDoc(
|
||||||
|
relPath: string,
|
||||||
|
reason = "index",
|
||||||
|
): Promise<Job<IndexDocJobData>> {
|
||||||
|
const slug = relPath.replace(/\.mdx?$/, "");
|
||||||
|
return getQueue().add(
|
||||||
|
"index-doc",
|
||||||
|
{ relPath, reason },
|
||||||
|
{
|
||||||
|
jobId: `doc__${slug}`,
|
||||||
|
removeOnComplete: 1000,
|
||||||
|
removeOnFail: 5000,
|
||||||
|
attempts: 3,
|
||||||
|
backoff: { type: "exponential", delay: 2000 },
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Enqueue a full-corpus reindex (used by the git-sync hook). */
|
||||||
|
export async function enqueueFullIndex(
|
||||||
|
reason = "reindex",
|
||||||
|
): Promise<Job<IndexAllJobData>> {
|
||||||
|
return getQueue().add(
|
||||||
|
"index-all",
|
||||||
|
{ reason },
|
||||||
|
{
|
||||||
|
jobId: `full__${Date.now()}`,
|
||||||
|
removeOnComplete: 100,
|
||||||
|
removeOnFail: 1000,
|
||||||
|
attempts: 1,
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,63 @@
|
|||||||
|
import { Worker, type Job } from "bullmq";
|
||||||
|
import { getConnection, BULLMQ_PREFIX } from "./client";
|
||||||
|
import { INDEX_QUEUE } from "./queue";
|
||||||
|
import { indexContentFile, runFullIndex } from "@mcpedia/core";
|
||||||
|
|
||||||
|
export function createWorker(): Worker {
|
||||||
|
const worker = new Worker(
|
||||||
|
INDEX_QUEUE,
|
||||||
|
async (job: Job) => {
|
||||||
|
switch (job.name) {
|
||||||
|
case "index-doc": {
|
||||||
|
const { relPath, reason } = job.data as {
|
||||||
|
relPath: string;
|
||||||
|
reason: string;
|
||||||
|
};
|
||||||
|
await job.log(`indexing ${relPath}`);
|
||||||
|
const r = await indexContentFile(relPath, reason);
|
||||||
|
return r;
|
||||||
|
}
|
||||||
|
case "index-all": {
|
||||||
|
const { reason } = job.data as { reason: string };
|
||||||
|
await job.log(`full index (${reason})`);
|
||||||
|
return await runFullIndex(reason);
|
||||||
|
}
|
||||||
|
default:
|
||||||
|
throw new Error(`unknown job type: ${job.name}`);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
connection: getConnection(),
|
||||||
|
prefix: BULLMQ_PREFIX,
|
||||||
|
concurrency: 4,
|
||||||
|
},
|
||||||
|
);
|
||||||
|
|
||||||
|
worker.on("completed", (job) => {
|
||||||
|
console.log(`[worker] completed ${job.name} (${job.id})`);
|
||||||
|
});
|
||||||
|
worker.on("failed", (job, err) => {
|
||||||
|
console.error(`[worker] failed ${job?.name} (${job?.id}): ${err.message}`);
|
||||||
|
});
|
||||||
|
worker.on("error", (err) => {
|
||||||
|
console.error(`[worker] error:`, err.message);
|
||||||
|
});
|
||||||
|
|
||||||
|
return worker;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Start the worker and wire graceful shutdown. */
|
||||||
|
export async function startWorker(): Promise<Worker> {
|
||||||
|
const worker = createWorker();
|
||||||
|
console.log("[worker] indexing worker started");
|
||||||
|
|
||||||
|
const shutdown = async (sig: string) => {
|
||||||
|
console.log(`[worker] ${sig} received, closing...`);
|
||||||
|
await worker.close();
|
||||||
|
process.exit(0);
|
||||||
|
};
|
||||||
|
process.on("SIGINT", () => void shutdown("SIGINT"));
|
||||||
|
process.on("SIGTERM", () => void shutdown("SIGTERM"));
|
||||||
|
|
||||||
|
return worker;
|
||||||
|
}
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
import { enqueueIndexDoc, enqueueFullIndex } from "@mcpedia/queue";
|
||||||
|
|
||||||
|
// One-shot enqueue helper (no worker required to schedule work).
|
||||||
|
// Usage:
|
||||||
|
// bun run enqueue --all # full reindex
|
||||||
|
// bun run enqueue docs/websocket/contract # single doc (slug or rel path)
|
||||||
|
async function main() {
|
||||||
|
const arg = process.argv[2];
|
||||||
|
if (!arg || arg === "--all") {
|
||||||
|
const job = await enqueueFullIndex("manual");
|
||||||
|
console.log(`enqueued full reindex job ${job.id}`);
|
||||||
|
} else {
|
||||||
|
const relPath = arg.endsWith(".md") || arg.endsWith(".mdx") ? arg : `${arg}.md`;
|
||||||
|
const job = await enqueueIndexDoc(relPath, "manual");
|
||||||
|
console.log(`enqueued doc reindex job ${job.id} -> ${relPath}`);
|
||||||
|
}
|
||||||
|
await new Promise((r) => setTimeout(r, 500)); // allow the event loop to flush
|
||||||
|
process.exit(0);
|
||||||
|
}
|
||||||
|
|
||||||
|
main().catch((err) => {
|
||||||
|
console.error(err);
|
||||||
|
process.exit(1);
|
||||||
|
});
|
||||||
+9
-62
@@ -1,67 +1,14 @@
|
|||||||
import { db } from "@mcpedia/db";
|
import { runFullIndex } from "@mcpedia/core";
|
||||||
import { documents } from "@mcpedia/db/schema";
|
|
||||||
import { parseFile } from "@mcpedia/parser";
|
|
||||||
import { CONTENT_ROOT } from "@mcpedia/config";
|
|
||||||
import { listContentFiles, indexChunks } from "@mcpedia/core";
|
|
||||||
import { join } from "node:path";
|
|
||||||
|
|
||||||
|
// Phase 3: the indexer now goes through `runFullIndex`, the single indexing
|
||||||
|
// entry point shared with the BullMQ worker and the git-sync hook. It parses
|
||||||
|
// each content file, upserts `documents`, chunks+embeds, and snapshots a
|
||||||
|
// revision when the body changed.
|
||||||
async function main() {
|
async function main() {
|
||||||
const files = listContentFiles();
|
const reason = process.argv[2] && process.argv[2].startsWith("--reason=")
|
||||||
let indexed = 0;
|
? process.argv[2].slice("--reason=".length)
|
||||||
let chunked = 0;
|
: "index";
|
||||||
for (const rel of files) {
|
await runFullIndex(reason);
|
||||||
const abs = join(CONTENT_ROOT, rel);
|
|
||||||
const { meta, body } = parseFile(abs, rel);
|
|
||||||
const nowIso =
|
|
||||||
meta.updatedAt && meta.updatedAt !== ""
|
|
||||||
? meta.updatedAt
|
|
||||||
: new Date().toISOString();
|
|
||||||
await db
|
|
||||||
.insert(documents)
|
|
||||||
.values({
|
|
||||||
id: meta.id,
|
|
||||||
slug: meta.slug,
|
|
||||||
title: meta.title,
|
|
||||||
type: meta.type,
|
|
||||||
section: meta.section,
|
|
||||||
status: meta.status,
|
|
||||||
author: meta.author,
|
|
||||||
tags: meta.tags,
|
|
||||||
path: meta.path,
|
|
||||||
body,
|
|
||||||
createdAt: new Date(meta.createdAt || nowIso),
|
|
||||||
updatedAt: new Date(nowIso),
|
|
||||||
})
|
|
||||||
.onConflictDoUpdate({
|
|
||||||
target: documents.slug,
|
|
||||||
set: {
|
|
||||||
title: meta.title,
|
|
||||||
type: meta.type,
|
|
||||||
section: meta.section,
|
|
||||||
status: meta.status,
|
|
||||||
author: meta.author,
|
|
||||||
tags: meta.tags,
|
|
||||||
path: meta.path,
|
|
||||||
body,
|
|
||||||
updatedAt: new Date(nowIso),
|
|
||||||
},
|
|
||||||
});
|
|
||||||
indexed++;
|
|
||||||
console.log(` indexed ${rel}`);
|
|
||||||
|
|
||||||
// Phase 2: chunk + embed for semantic search.
|
|
||||||
try {
|
|
||||||
const n = await indexChunks(meta.slug, body);
|
|
||||||
chunked += n;
|
|
||||||
console.log(` embedded ${n} chunks`);
|
|
||||||
} catch (err) {
|
|
||||||
console.error(
|
|
||||||
` embed FAILED for ${meta.slug}: ${err instanceof Error ? err.message : err}`,
|
|
||||||
);
|
|
||||||
// Don't abort the whole index over one doc's embedding failure.
|
|
||||||
}
|
|
||||||
}
|
|
||||||
console.log(`indexed ${indexed} documents, ${chunked} chunks embedded`);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
main()
|
main()
|
||||||
|
|||||||
@@ -10,6 +10,7 @@
|
|||||||
"@mcpedia/embeddings": "workspace:*",
|
"@mcpedia/embeddings": "workspace:*",
|
||||||
"@mcpedia/parser": "workspace:*",
|
"@mcpedia/parser": "workspace:*",
|
||||||
"@mcpedia/search": "workspace:*",
|
"@mcpedia/search": "workspace:*",
|
||||||
|
"@mcpedia/queue": "workspace:*",
|
||||||
"drizzle-orm": "^0.38.0",
|
"drizzle-orm": "^0.38.0",
|
||||||
"postgres": "^3.4.5"
|
"postgres": "^3.4.5"
|
||||||
}
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user