feat(mcpedia): Phase 3 — async indexing (BullMQ), git-sync webhook, revisions, MCP Resources

- packages/queue: ioredis singleton + BullMQ Queue/Worker (prefix mcpedia:
  on shared imrnes Redis :6379); apps/worker runs startWorker()
- @mcpedia/core: indexContentFile/runFullIndex (single indexing entry point
  shared by script/worker/hook) + revision.service (list/get/restore)
- document_revisions table (migration 0002) — snapshots only on body change
- apps/api: POST /hooks/reindex + /hooks/index webhooks; tRPC revisions,
  getRevision, restoreRevision, jobStatus, queueStatus
- apps/mcp: register MCP Resources mcpedia://docs{/,+slug/chunks/revisions}
  ({+slug} RFC6570 reserved expansion for slugs containing /)
- apps/mcp zod pinned to ^4 to match MCP SDK 1.30 compiled types
  (resolves registerTool TS2589/ShapeOutput skew)
- scripts/enqueue.ts one-shot job enqueue helper; indexer refactored to runFullIndex
- PHASES.md/README/.env.example/docs updated
This commit is contained in:
asepharyana
2026-08-19 20:18:28 +07:00
parent 9397303f01
commit 8f2229d447
30 changed files with 1233 additions and 78 deletions
+114
View File
@@ -0,0 +1,114 @@
# MCPedia — Phase 3 "Async + Scale" Implementation Plan
Status: Phase 1 (MVP) + Phase 2 (Semantic+API) DONE. Phase 3 adds async
background work, git-driven reindex, document revision history, and MCP
Resources. All logic stays in `@mcpedia/core`; new `packages/queue` wires
BullMQ; `apps/worker` runs the worker process; the existing API gets a git-sync
webhook + job-status procedures; the MCP server gains Resources.
## Scope (4 features from PHASES.md)
1. **Redis + BullMQ background indexing/embedding workers**
2. **Git synchronization hook** (auto-reindex on push via webhook)
3. **Document revision system** (`document_revisions`)
4. **MCP Resources** (`mcpedia://docs/...`) alongside existing tools
## Architecture decisions (locked)
- **Redis**: shared imrnes Redis `100.121.180.82:6379`, no auth (verified
`+PONG`). `REDIS_URL` env (default `redis://100.121.180.82:6379`), optional
`REDIS_PASSWORD`. BullMQ key prefix `mcpedia:` to avoid collisions on the
shared instance.
- **Queue lib**: `bullmq@6.1.2` + `ioredis@6.0.0` (BullMQ peer dep). Pass an
ioredis instance; BullMQ duplicates it for blocking commands.
- **Single source of truth preserved**: per-doc indexing logic moves into
`@mcpedia/core` as `indexContentFile(relPath, reason?)`. The script, the
worker, and the git hook ALL call this. Revisions are snapshotted inside it.
- **Revisions**: created only when body actually changes vs the latest revision
(avoids bloat on every sync). Stored in `document_revisions`.
## Files touched
### packages/config
- `src/index.ts`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
### packages/db
- `src/schema.ts`: add `documentRevisions` table
(id, documentId→documents.id cascade, slug, revisionNo int, title, body,
meta jsonb, reason text, createdAt). Index (document_id, revision_no DESC),
(slug).
- `drizzle/0002_document_revisions.sql`: migration (applied via psql).
- `drizzle/meta/0002_snapshot.json` + `_journal.json` entry (keeps drizzle-kit
consistent even though we apply manually).
### packages/core (new)
- `src/index.service.ts`:
- `indexContentFile(relPath: string, reason = "index")` — parse → upsert
`documents` → `indexChunks` → snapshot revision (if changed).
- `runFullIndex(reason?)` — walk content, index each, return counts.
- `src/revision.service.ts`:
- `createRevision(...)`, `listRevisions(slug, limit)`,
`getRevision(id)`, `latestRevisionBody(slug)`, `restoreRevision(id)`.
- `src/index.ts`: export both.
### packages/queue (NEW)
- `package.json` (@mcpedia/queue): deps bullmq, ioredis, @mcpedia/core,
@mcpedia/db, @mcpedia/config.
- `src/client.ts`: ioredis instance factory from config.
- `src/queue.ts`:
- `INDEX_QUEUE = "mcpedia-index"`.
- `enqueueIndexDoc(slug, absPath, reason)`, `enqueueFullIndex(reason)`.
- `getQueue()` lazy singleton.
- `src/worker.ts`: `startWorker()` — BullMQ Worker with 3 job types:
`index-doc` (single), `index-all` (full), `reindex` (full, reason=git-push).
Graceful shutdown on SIGINT/SIGTERM. Job progress + error handling.
### apps/worker (NEW)
- `package.json` (@mcpedia/worker): script `start: bun src/index.ts`.
- `src/index.ts`: `startWorker()` + heartbeat log.
### apps/api
- `src/index.ts`: add `POST /hooks/reindex` (full) and
`POST /hooks/index?slug=` (single) webhook routes → enqueue jobs. Mount
AFTER /trpc.
- `src/router.ts`: add `jobStatus` (id→state/prev/failedData),
`queueStatus` (waiting/active/completed/failed counts),
`revisions` (slug→list), `restoreRevision` (id→new slug/doc).
- `package.json`: add `@mcpedia/queue` dep, `hooks` reused.
### apps/mcp
- `src/index.ts`: register Resources:
- `mcpedia://docs` (list all metas)
- `mcpedia://docs/{slug}` (full body from disk)
- `mcpedia://docs/{slug}/chunks` (chunk previews)
- `mcpedia://docs/{slug}/revisions` (revision list)
- `src/smoke.test.ts`: add `listResources` + read `mcpedia://docs` assertion.
### scripts
- `scripts/indexer.ts`: refactor `main()` to call `runFullIndex()`.
### Root
- `package.json`: add `"worker": "bun --cwd apps/worker run start"`,
`"reindex": "bun run scripts/worker.ts"`? No — `worker` runs the listener;
triggering reindex = `bun run api` webhook or `enqueueFullIndex` helper.
Add `"enqueue-index": "bun run scripts/enqueue.ts"` (one-shot enqueue).
- `.env.example`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
### Docs
- `PHASES.md`: mark Phase 3 items DONE with notes.
- `README.md`: document worker, webhook, revisions, MCP resources.
## Verification (real, not claimed)
1. `bun install` picks up new deps.
2. `bunx turbo run build` + `typecheck` green across workspace.
3. **Real BullMQ e2e against imrnes Redis**: script that enqueues an
`index-doc` job, starts a Worker, asserts the job completes and the doc row
+ chunks + a revision row appear in Postgres. Verifies Redis+ioredis+bullmq
+ db + core all wired correctly.
4. `bun --cwd apps/mcp run smoke` passes (incl. new resources).
5. `bun run index` (runFullIndex) green; verify `documents`,
`document_chunks`, `document_revisions` row counts via psql.
6. API webhook: `curl -XPOST localhost:4020/hooks/reindex` enqueues; worker
processes; `curl localhost:4020/trpc/queueStatus` reflects counts.
7. MCP resource read returns real content.
+43 -5
View File
@@ -27,12 +27,50 @@ Legend: ✅ built · 🟡 partial · ⬜ deferred
- [x] MCP server — added `semantic_search` + `hybrid_search` tools (6 total). - [x] MCP server — added `semantic_search` + `hybrid_search` tools (6 total).
- [x] Web search — keyword/hybrid toggle (`?mode=hybrid`), hybrid reaches semantically-related docs keyword misses. - [x] Web search — keyword/hybrid toggle (`?mode=hybrid`), hybrid reaches semantically-related docs keyword misses.
## Phase 3 — Async + Scale ## Phase 3 — Async + Scale ✅ DONE
- [ ] Redis + BullMQ background indexing / embedding workers - [x] **Redis + BullMQ background indexing / embedding workers** —
- [ ] Git synchronization hook (auto-reindex on push) `packages/queue` (ioredis singleton + BullMQ `Queue`/`Worker`, prefix
- [ ] Document revision system (`document_revisions`) `mcpedia:` on shared imrnes Redis `:6379`); `apps/worker` runs
- [ ] MCP Resources (`mcpedia://docs/...`) in addition to tools `startWorker()`. Three job types: `index-doc`, `index-all`, `reindex`.
Single indexing entry point `indexContentFile`/`runFullIndex` in
`@mcpedia/core` shared by the script, worker, and git hook. Verified
end-to-end against live Redis (job enqueue → worker → Postgres write).
- [x] **Git synchronization hook (auto-reindex on push)** — API webhook
`POST /hooks/reindex` (full) and `POST /hooks/index?slug=` (single) enqueue
BullMQ jobs. Wire a Git provider (GitHub/Gitea) post-receive / webhook to
`POST /hooks/reindex` to auto-reindex on push. `scripts/enqueue.ts` is a
one-shot enqueue helper (`bun run enqueue --all` / `<slug>`).
- [x] **Document revision system (`document_revisions`)** — `packages/db`
migration `0002_document_revisions.sql`. Indexer snapshots a revision only
when the body actually changes vs the latest revision (pure metadata edits
don't bloat history). `listRevisions` / `getRevision` / `restoreRevision`
in `@mcpedia/core`; exposed as tRPC `revisions` / `getRevision` /
`restoreRevision` and the `mcpedia://docs/{+slug}/revisions` MCP Resource.
- [x] **MCP Resources (`mcpedia://docs/...`)** — alongside the 6 tools:
`mcpedia://docs` (list), `mcpedia://docs/{+slug}` (body from disk),
`mcpedia://docs/{+slug}/chunks` (chunk preview),
`mcpedia://docs/{+slug}/revisions` (history). `{+slug}` uses RFC 6570
reserved expansion so slugs containing `/` match.
### New/changed commands
```
bun run index # full reindex (runFullIndex, writes revisions)
bun run enqueue --all # enqueue a full reindex job (no worker needed)
bun run enqueue <slug> # enqueue a single-doc reindex job
bun run worker # start the BullMQ indexing worker (long-running)
bun run api # Hono+tRPC API on :4020 (added /hooks/* webhooks)
```
### Verification done (real, against imrnes Redis + Postgres)
- `turbo run typecheck` green across all 13 packages.
- BullMQ e2e: enqueue `index-doc` → worker completes → `documents` +
`document_chunks` + `document_revisions` rows present.
- Revision dedup proven: editing a body creates a new revision; metadata-only
reindex does not; `restoreRevision` writes history back into the live row.
- MCP smoke test passes (tools + all 4 resources).
- API webhook `POST /hooks/reindex` enqueues → worker drains queue →
`queueStatus` reflects counts.
## Phase 4 — Scale-out (only if needed) ## Phase 4 — Scale-out (only if needed)
+38 -8
View File
@@ -14,16 +14,19 @@ column) and served through a single **Core** layer that every interface
mcpedia/ mcpedia/
├── apps/ ├── apps/
│ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search │ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search
│ └── mcp/ # MCP server (stdio) — AI-agent interface │ ├── mcp/ # MCP server (stdio) — AI-agent interface (tools + resources)
│ └── api/ # Hono + tRPC v11 API on :4020 (+ /hooks/* git-sync webhooks)
├── packages/ ├── packages/
│ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...) │ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...)
│ ├── config/ # loads .env (repo root) as authoritative dev config │ ├── config/ # loads .env (repo root) as authoritative dev config
│ ├── db/ # Drizzle ORM schema + client + drizzle-kit config │ ├── db/ # Drizzle ORM schema + client + drizzle-kit config
│ ├── parser/ # frontmatter (gray-matter) parsing │ ├── parser/ # frontmatter (gray-matter) parsing
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline) │ ├── search/ # Postgres FTS query (ts_rank + ts_headline)
│ └── core/ # Document/Content/Search services — the only business logic │ ├── embeddings/ # embedding provider + chunker
│ ├── queue/ # Redis (ioredis) + BullMQ worker/queue (Phase 3)
│ └── core/ # Document/Content/Search/Index/Revision — the only business logic
├── content/ # docs/ writeups/ research/ notes/ (the knowledge base) ├── content/ # docs/ writeups/ research/ notes/ (the knowledge base)
└── scripts/ # indexer.ts (walks content/ -> upserts into Postgres) └── scripts/ # indexer.ts (full reindex), enqueue.ts (one-shot job enqueue)
``` ```
## Architecture principle ## Architecture principle
@@ -101,23 +104,45 @@ the DB stores metadata + the search vector.
| `list_documents` | List, optionally filtered by section | | `list_documents` | List, optionally filtered by section |
| `get_related_documents` | Docs sharing tags with a given slug | | `get_related_documents` | Docs sharing tags with a given slug |
### MCP Resources
| URI | Purpose |
| -------------------------------- | ---------------------------------------- |
| `mcpedia://docs` | List all published documents |
| `mcpedia://docs/{+slug}` | Full markdown body (read from disk) |
| `mcpedia://docs/{+slug}/chunks` | Preview of embedded semantic chunks |
| `mcpedia://docs/{+slug}/revisions` | Revision history summary |
(`{+slug}` uses RFC 6570 reserved expansion so a slug like
`docs/websocket/contract` matches the template.)
Smoke test (in-memory transport, real JSON-RPC): Smoke test (in-memory transport, real JSON-RPC):
```bash ```bash
bun --cwd apps/mcp run smoke bun --cwd apps/mcp run smoke
``` ```
## API (Phase 2) ## API (Phase 2 + Phase 3)
A tRPC v11 API is also exposed via Hono on **:4020** (all procedures mirror the A tRPC v11 API is exposed via Hono on **:4020** (all procedures mirror the
MCP tools): MCP tools). Phase 3 adds async job + revision procedures and git-sync webhooks:
```bash ```bash
bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*) bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*)
``` ```
`bun run index` now also chunks + embeds (Phase 2 indexer). Requires `EMBED_*` tRPC procedures: `search`, `semanticSearch`, `hybridSearch`, `getDocument`,
vars in `.env` (see `.env.example`). `listDocuments`, `related` (Phase 2); plus `revisions`, `getRevision`,
`restoreRevision`, `jobStatus`, `queueStatus` (Phase 3).
Git-sync webhooks (enqueue BullMQ jobs; the worker processes them):
- `POST /hooks/reindex` — full-corpus reindex (point your Git provider's
push webhook here to auto-reindex on push).
- `POST /hooks/index?slug=<slug>` — reindex a single document.
`bun run index` now also chunks + embeds (Phase 2 indexer) and snapshots a
revision whenever the body changes (Phase 3). See `.env.example` for
`EMBED_*` / `REDIS_*` / `QUEUE_PREFIX` vars.
## Status ## Status
@@ -128,6 +153,11 @@ Postgres FTS keyword search, content indexing.
chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API
(`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle. (`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle.
**Phase 3 — Async + Scale (DONE):** Redis + BullMQ background indexing/embedding
workers (`packages/queue`, `apps/worker`), git-sync webhooks (`POST /hooks/*`),
document revision system (`document_revisions` + restore), and MCP Resources
(`mcpedia://docs/...`). See `PHASES.md`.
> pgvector is **not installed** on the shared imrnes Postgres, so vector storage is > pgvector is **not installed** on the shared imrnes Postgres, so vector storage is
> a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is > a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is
> the Phase-4 scale-out path. See `PHASES.md`. > the Phase-4 scale-out path. See `PHASES.md`.
+2
View File
@@ -13,6 +13,8 @@
"@hono/node-server": "^1.13.0", "@hono/node-server": "^1.13.0",
"@mcpedia/config": "workspace:*", "@mcpedia/config": "workspace:*",
"@mcpedia/core": "workspace:*", "@mcpedia/core": "workspace:*",
"@mcpedia/db": "workspace:*",
"@mcpedia/queue": "workspace:*",
"@trpc/server": "^11.0.0", "@trpc/server": "^11.0.0",
"hono": "^4.6.0", "hono": "^4.6.0",
"zod": "^3.23.8" "zod": "^3.23.8"
+19
View File
@@ -4,12 +4,31 @@ import { fetchRequestHandler } from "@trpc/server/adapters/fetch";
import { db } from "@mcpedia/db"; import { db } from "@mcpedia/db";
import { appRouter } from "./router"; import { appRouter } from "./router";
import type { Context } from "./trpc"; import type { Context } from "./trpc";
import { enqueueIndexDoc, enqueueFullIndex } from "@mcpedia/queue";
const app = new Hono(); const app = new Hono();
// Health check. // Health check.
app.get("/health", (c) => c.json({ ok: true })); app.get("/health", (c) => c.json({ ok: true }));
// --- Phase 3: Git synchronization hook ---
// POST /hooks/reindex -> enqueue a full-corpus reindex (git push webhook)
// POST /hooks/index?slug=... -> enqueue a single document reindex
// Returns the created job id(s). The worker processes them asynchronously.
app.post("/hooks/reindex", async (c) => {
const job = await enqueueFullIndex("git-push");
return c.json({ ok: true, jobId: job.id, kind: "full" });
});
app.post("/hooks/index", async (c) => {
const slug = c.req.query("slug");
if (!slug) return c.json({ ok: false, error: "slug query param required" }, 400);
// slug is the relative path without extension, e.g. docs/websocket/contract
const relPath = slug.endsWith(".md") || slug.endsWith(".mdx") ? slug : `${slug}.md`;
const job = await enqueueIndexDoc(relPath, "git-push");
return c.json({ ok: true, jobId: job.id, kind: "doc", relPath });
});
// Mount tRPC at /trpc/*. The fetch adapter is the canonical Bun/Hono adapter. // Mount tRPC at /trpc/*. The fetch adapter is the canonical Bun/Hono adapter.
app.all("/trpc/*", (c) => app.all("/trpc/*", (c) =>
fetchRequestHandler({ fetchRequestHandler({
+57
View File
@@ -7,7 +7,12 @@ import {
keywordSearch, keywordSearch,
listDocuments, listDocuments,
semanticSearch, semanticSearch,
listRevisions,
getRevision,
restoreRevision,
} from "@mcpedia/core"; } from "@mcpedia/core";
import { getQueue, INDEX_QUEUE } from "@mcpedia/queue";
import { getConnection, BULLMQ_PREFIX } from "@mcpedia/queue/client";
export const appRouter = router({ export const appRouter = router({
search: publicProcedure search: publicProcedure
@@ -33,6 +38,58 @@ export const appRouter = router({
related: publicProcedure related: publicProcedure
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(20).default(5) })) .input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(20).default(5) }))
.query(async ({ input }) => getRelated(input.slug, input.limit)), .query(async ({ input }) => getRelated(input.slug, input.limit)),
// --- Phase 3: revisions ---
revisions: publicProcedure
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(50).default(20) }))
.query(async ({ input }) => listRevisions(input.slug, input.limit)),
getRevision: publicProcedure
.input(z.object({ id: z.string() }))
.query(async ({ input }) => getRevision(input.id)),
restoreRevision: publicProcedure
.input(z.object({ id: z.string() }))
.mutation(async ({ input }) => restoreRevision(input.id)),
// --- Phase 3: async job status ---
jobStatus: publicProcedure
.input(z.object({ id: z.string() }))
.query(async ({ input }) => {
const queue = getQueue();
const job = await queue.getJob(input.id);
if (!job) return { exists: false };
const state = await job.getState();
const failedReason = job.failedReason;
const returnvalue = job.returnvalue;
const progress = job.progress;
return {
exists: true,
id: job.id,
name: job.name,
state,
progress,
failedReason,
returnvalue,
attemptsMade: job.attemptsMade,
};
}),
queueStatus: publicProcedure.query(async () => {
const queue = getQueue();
const [waiting, active, completed, failed, delayed] = await Promise.all([
queue.getWaitingCount(),
queue.getActiveCount(),
queue.getCompletedCount(),
queue.getFailedCount(),
queue.getDelayedCount(),
]);
return {
queue: INDEX_QUEUE,
prefix: BULLMQ_PREFIX,
counts: { waiting, active, completed, failed, delayed },
};
}),
}); });
export type AppRouter = typeof appRouter; export type AppRouter = typeof appRouter;
+1 -1
View File
@@ -17,7 +17,7 @@
"@mcpedia/core": "workspace:*", "@mcpedia/core": "workspace:*",
"@mcpedia/search": "workspace:*", "@mcpedia/search": "workspace:*",
"@modelcontextprotocol/sdk": "^1.29.0", "@modelcontextprotocol/sdk": "^1.29.0",
"zod": "^3.23.8" "zod": "^4.0.0"
}, },
"devDependencies": { "devDependencies": {
"@types/node": "^20", "@types/node": "^20",
+123 -1
View File
@@ -1,7 +1,19 @@
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js"; import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"; import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { ResourceTemplate } from "@modelcontextprotocol/sdk/server/mcp.js";
import { z } from "zod"; import { z } from "zod";
import { listDocuments, getDocument, getRelated, semanticSearch, hybridSearch, keywordSearch } from "@mcpedia/core"; import {
listDocuments,
getDocument,
getRelated,
semanticSearch,
hybridSearch,
keywordSearch,
listRevisions,
readContentFile,
} from "@mcpedia/core";
import { CONTENT_ROOT } from "@mcpedia/config";
import { join } from "node:path";
export function createMcpServer(): McpServer { export function createMcpServer(): McpServer {
const server = new McpServer({ const server = new McpServer({
@@ -119,6 +131,116 @@ export function createMcpServer(): McpServer {
}, },
); );
// --- Phase 3: MCP Resources (read-only knowledge base surfaced via URIs) ---
// mcpedia://docs -> list all published documents
// mcpedia://docs/{slug} -> full markdown body (from disk)
// mcpedia://docs/{slug}/chunks -> chunked preview (semantic slices)
// mcpedia://docs/{slug}/revisions -> revision history summary
server.registerResource(
"mcpedia-docs-list",
"mcpedia://docs",
{
title: "MCPedia document index",
description: "List of all published documents in the knowledge base.",
mimeType: "application/json",
},
async (uri) => {
const docs = await listDocuments();
return {
contents: [
{
uri: uri.href,
mimeType: "application/json",
text: JSON.stringify(docs, null, 2),
},
],
};
},
);
server.registerResource(
"mcpedia-doc-chunks",
new ResourceTemplate("mcpedia://docs/{+slug}/chunks", { list: undefined }),
{
title: "MCPedia document chunks",
description: "Preview of the embedded semantic chunks for a document.",
mimeType: "application/json",
},
async (uri, vars) => {
const slug = String(vars.slug);
const doc = await getDocument(slug);
if (!doc) throw new Error(`Document not found: ${slug}`);
// Chunk the body the same way the indexer does (size 1000 / overlap 150)
// so the resource mirrors what semantic search actually sees.
const { chunkText } = await import("@mcpedia/embeddings");
const chunks = chunkText(doc.body, { size: 1000, overlap: 150 });
return {
contents: [
{
uri: uri.href,
mimeType: "application/json",
text: JSON.stringify(
chunks.map((c, i) => ({ index: i, length: c.length, preview: c.slice(0, 200) })),
null,
2,
),
},
],
};
},
);
server.registerResource(
"mcpedia-doc-revisions",
new ResourceTemplate("mcpedia://docs/{+slug}/revisions", { list: undefined }),
{
title: "MCPedia document revisions",
description: "Revision history summary for a document.",
mimeType: "application/json",
},
async (uri, vars) => {
const slug = String(vars.slug);
const revs = await listRevisions(slug, 20);
return {
contents: [
{
uri: uri.href,
mimeType: "application/json",
text: JSON.stringify(revs, null, 2),
},
],
};
},
);
// Registered LAST: the bare {+slug} template is greedy and would otherwise
// swallow /chunks and /revisions URIs. Specific templates must match first.
server.registerResource(
"mcpedia-doc",
new ResourceTemplate("mcpedia://docs/{+slug}", { list: undefined }),
{
title: "MCPedia document",
description: "Full markdown body of a single document, read from disk (source of truth).",
mimeType: "text/markdown",
},
async (uri, vars) => {
const slug = String(vars.slug);
const doc = await getDocument(slug);
if (!doc) {
throw new Error(`Document not found: ${slug}`);
}
return {
contents: [
{
uri: uri.href,
mimeType: "text/markdown",
text: doc.body,
},
],
};
},
);
return server; return server;
} }
+29
View File
@@ -92,6 +92,35 @@ async function main() {
} }
console.log(`hybrid_search => ${hybHits.length} docs, top: ${hybHits[0].doc.slug}`); console.log(`hybrid_search => ${hybHits.length} docs, top: ${hybHits[0].doc.slug}`);
// 8) resources: list
const resList = await client.listResources();
const resNames = resList.resources.map((r: any) => r.name).sort();
console.log("resources:", resNames.join(", "));
if (!resNames.includes("mcpedia-docs-list")) {
throw new Error("expected mcpedia-docs-list resource");
}
// 9) resource: read the docs list (must not throw, returns JSON content)
const readList = await client.readResource({ uri: "mcpedia://docs" });
const listText = (readList.contents as any)[0].text;
if (!listText.includes("docs/websocket/contract")) {
throw new Error("mcpedia://docs did not list the websocket contract doc");
}
console.log("readResource(mcpedia://docs) => ok");
// 10) resource: read a single doc body + revisions
const readDoc = await client.readResource({ uri: "mcpedia://docs/docs/websocket/contract" });
const docText = (readDoc.contents as any)[0].text;
if (!docText.includes("WebSocket Contract")) {
throw new Error("mcpedia://docs/{slug} returned unexpected body");
}
console.log("readResource(mcpedia://docs/docs/websocket/contract) => ok");
const readRev = await client.readResource({
uri: "mcpedia://docs/docs/websocket/contract/revisions",
});
console.log("readResource(.../revisions) => ok");
await client.close(); await client.close();
await server.close(); await server.close();
console.log("\nSMOKE OK"); console.log("\nSMOKE OK");
+20
View File
@@ -0,0 +1,20 @@
{
"name": "@mcpedia/worker",
"version": "0.1.0",
"private": true,
"type": "module",
"scripts": {
"start": "bun run src/index.ts",
"lint": "tsc --noEmit",
"typecheck": "tsc --noEmit"
},
"dependencies": {
"@mcpedia/config": "workspace:*",
"@mcpedia/core": "workspace:*",
"@mcpedia/db": "workspace:*",
"@mcpedia/queue": "workspace:*"
},
"devDependencies": {
"typescript": "^5.6.0"
}
}
+12
View File
@@ -0,0 +1,12 @@
import { startWorker } from "@mcpedia/queue/worker";
// Keep the process alive: the worker listens on the BullMQ queue until a
// SIGINT/SIGTERM closes it (handled inside startWorker).
const worker = await startWorker();
// Heartbeat so the supervisor/operator can see liveness without scraping logs.
const heartbeat = setInterval(() => {
console.log(`[worker] alive, ${worker.name} queue="${worker.name}"`);
}, 30_000);
worker.on("closed", () => clearInterval(heartbeat));
+11
View File
@@ -0,0 +1,11 @@
{
"extends": "../../tsconfig.base.json",
"compilerOptions": {
"paths": {
"@mcpedia/db": ["../../packages/db/src/index.ts"],
"@mcpedia/db/schema": ["../../packages/db/src/schema.ts"],
"@mcpedia/*": ["../../packages/*"]
}
},
"include": ["src/**/*.ts"]
}
+75 -1
View File
File diff suppressed because one or more lines are too long
+2
View File
@@ -14,6 +14,8 @@
"lint": "turbo run lint", "lint": "turbo run lint",
"typecheck": "turbo run typecheck", "typecheck": "turbo run typecheck",
"index": "bun run scripts/indexer.ts", "index": "bun run scripts/indexer.ts",
"enqueue": "bun run scripts/enqueue.ts",
"worker": "bun --cwd apps/worker run start",
"mcp": "bun --cwd apps/mcp run start", "mcp": "bun --cwd apps/mcp run start",
"api": "bun --cwd apps/api run dev" "api": "bun --cwd apps/api run dev"
}, },
+6
View File
@@ -40,6 +40,12 @@ export const EMBED_BASE_URL = process.env.EMBED_BASE_URL ?? "";
export const EMBED_API_KEY = process.env.EMBED_API_KEY ?? ""; export const EMBED_API_KEY = process.env.EMBED_API_KEY ?? "";
export const EMBED_MODEL = process.env.EMBED_MODEL ?? ""; export const EMBED_MODEL = process.env.EMBED_MODEL ?? "";
// Phase 3: Redis + BullMQ (shared imrnes Redis, no auth by default).
export const REDIS_URL = process.env.REDIS_URL ?? "redis://100.121.180.82:6379";
export const REDIS_PASSWORD = process.env.REDIS_PASSWORD ?? "";
// BullMQ key prefix to namespace jobs on the shared Redis instance.
export const QUEUE_PREFIX = process.env.QUEUE_PREFIX ?? "mcpedia";
if (!DATABASE_URL) { if (!DATABASE_URL) {
// Fail fast with an explicit message instead of a cryptic driver error. // Fail fast with an explicit message instead of a cryptic driver error.
throw new Error( throw new Error(
+155
View File
@@ -0,0 +1,155 @@
import { db } from "@mcpedia/db";
import { documents, documentRevisions, documentChunks } from "@mcpedia/db/schema";
import { parseFile } from "@mcpedia/parser";
import { CONTENT_ROOT } from "@mcpedia/config";
import { listContentFiles } from "./content.service";
import { indexChunks } from "./document.service";
import { toMeta } from "./row-map";
import { eq, desc, and, sql } from "drizzle-orm";
import { join } from "node:path";
export interface IndexResult {
indexed: number;
chunks: number;
revisions: number;
}
/**
* Index a single content file: parse → upsert `documents` → chunk+embed →
* snapshot a revision if the body changed since the last indexed revision.
*
* This is THE single indexing entry point shared by the CLI script, the
* BullMQ worker, and the git-sync hook — no business logic is duplicated.
*
* @param relPath path relative to CONTENT_ROOT (e.g. "docs/websocket/contract")
* @param reason provenance tag for the revision ("index" | "git-push" | "reindex")
*/
export async function indexContentFile(
relPath: string,
reason = "index",
): Promise<{ indexed: boolean; chunks: number; revision: boolean }> {
const abs = join(CONTENT_ROOT, relPath);
const { meta, body } = parseFile(abs, relPath);
const nowIso =
meta.updatedAt && meta.updatedAt !== ""
? meta.updatedAt
: new Date().toISOString();
await db
.insert(documents)
.values({
id: meta.id,
slug: meta.slug,
title: meta.title,
type: meta.type,
section: meta.section,
status: meta.status,
author: meta.author,
tags: meta.tags,
path: meta.path,
body,
createdAt: new Date(meta.createdAt || nowIso),
updatedAt: new Date(nowIso),
})
.onConflictDoUpdate({
target: documents.slug,
set: {
title: meta.title,
type: meta.type,
section: meta.section,
status: meta.status,
author: meta.author,
tags: meta.tags,
path: meta.path,
body,
updatedAt: new Date(nowIso),
},
});
// Semantic chunks (embedding). A failure here must not abort the whole
// index — log and continue; FTS still works without embeddings.
let chunks = 0;
try {
chunks = await indexChunks(meta.slug, body);
} catch (err) {
console.error(
` embed FAILED for ${meta.slug}: ${err instanceof Error ? err.message : err}`,
);
}
// Snapshot a revision only when the body actually changed vs the latest
// revision. Pure metadata/index changes (tags/title) won't create noise.
const revision = await snapshotRevision(meta.slug, meta, body, reason);
return { indexed: true, chunks, revision };
}
/**
* Compare the incoming body against the latest revision's body; if different
* (or no prior revision exists), create a new revision with an incremented
* per-document revisionNo.
*/
async function snapshotRevision(
slug: string,
meta: ReturnType<typeof parseFile>["meta"],
body: string,
reason: string,
): Promise<boolean> {
const [doc] = await db
.select({ id: documents.id })
.from(documents)
.where(eq(documents.slug, slug));
if (!doc) return false;
const [latest] = await db
.select({ body: documentRevisions.body, revisionNo: documentRevisions.revisionNo })
.from(documentRevisions)
.where(eq(documentRevisions.documentId, doc.id))
.orderBy(desc(documentRevisions.revisionNo))
.limit(1);
if (latest && latest.body === body) {
return false; // unchanged → no new revision
}
const nextNo = (latest?.revisionNo ?? 0) + 1;
await db.insert(documentRevisions).values({
documentId: doc.id,
slug,
revisionNo: nextNo,
title: meta.title,
body,
meta: {
type: meta.type,
section: meta.section,
status: meta.status,
author: meta.author,
tags: meta.tags,
},
reason,
});
return true;
}
/**
* Walk the entire content tree and index every file. Returns aggregate counts.
*/
export async function runFullIndex(reason = "index"): Promise<IndexResult> {
const files = listContentFiles();
let indexed = 0;
let chunks = 0;
let revisions = 0;
for (const rel of files) {
const r = await indexContentFile(rel, reason);
indexed++;
chunks += r.chunks;
if (r.revision) revisions++;
console.log(
` indexed ${rel}${r.chunks ? ` (${r.chunks} chunks)` : ""}${r.revision ? " [revision]" : ""}`,
);
}
console.log(
`indexed ${indexed} documents, ${chunks} chunks, ${revisions} new revisions`,
);
return { indexed, chunks, revisions };
}
+2
View File
@@ -1,6 +1,8 @@
export * from "./content.service"; export * from "./content.service";
export * from "./document.service"; export * from "./document.service";
export * from "./search.service"; export * from "./search.service";
export * from "./index.service";
export * from "./revision.service";
export { toMeta } from "./row-map"; export { toMeta } from "./row-map";
export type { export type {
+116
View File
@@ -0,0 +1,116 @@
import { db } from "@mcpedia/db";
import { documents, documentRevisions, documentChunks } from "@mcpedia/db/schema";
import { eq, desc, and, sql } from "drizzle-orm";
import { toMeta } from "./row-map";
import type { DocumentMeta } from "@mcpedia/types";
export interface RevisionSummary {
id: string;
slug: string;
revisionNo: number;
title: string;
reason: string;
createdAt: string;
bodyLength: number;
}
/** List revisions for a slug, newest first. */
export async function listRevisions(
slug: string,
limit = 20,
): Promise<RevisionSummary[]> {
const [doc] = await db
.select({ id: documents.id })
.from(documents)
.where(eq(documents.slug, slug));
if (!doc) return [];
const rows = await db
.select({
id: documentRevisions.id,
slug: documentRevisions.slug,
revisionNo: documentRevisions.revisionNo,
title: documentRevisions.title,
reason: documentRevisions.reason,
createdAt: documentRevisions.createdAt,
bodyLength: sql<number>`length(${documentRevisions.body})`,
})
.from(documentRevisions)
.where(eq(documentRevisions.documentId, doc.id))
.orderBy(desc(documentRevisions.revisionNo))
.limit(limit);
return rows.map((r) => ({
id: r.id,
slug: r.slug,
revisionNo: r.revisionNo,
title: r.title,
reason: r.reason,
createdAt: r.createdAt.toISOString(),
bodyLength: r.bodyLength,
}));
}
/** Fetch a single revision's full body. */
export async function getRevision(
id: string,
): Promise<{ id: string; revisionNo: number; body: string; meta: unknown } | null> {
const [row] = await db
.select({
id: documentRevisions.id,
revisionNo: documentRevisions.revisionNo,
body: documentRevisions.body,
meta: documentRevisions.meta,
})
.from(documentRevisions)
.where(eq(documentRevisions.id, id));
if (!row) return null;
return {
id: row.id,
revisionNo: row.revisionNo,
body: row.body,
meta: row.meta,
};
}
/** Restore a revision: write its body+metadata back into the live `documents` row. */
export async function restoreRevision(
id: string,
): Promise<{ slug: string; documentId: string } | null> {
const [rev] = await db
.select({
id: documentRevisions.id,
documentId: documentRevisions.documentId,
slug: documentRevisions.slug,
title: documentRevisions.title,
body: documentRevisions.body,
meta: documentRevisions.meta,
})
.from(documentRevisions)
.where(eq(documentRevisions.id, id));
if (!rev) return null;
const m = rev.meta as {
type?: string;
section?: string;
status?: string;
author?: string;
tags?: string[];
};
await db
.update(documents)
.set({
title: rev.title,
type: (m.type as any) ?? "documentation",
section: (m.section as any) ?? "docs",
status: (m.status as any) ?? "published",
author: m.author ?? "",
tags: m.tags ?? [],
body: rev.body,
updatedAt: new Date(),
})
.where(eq(documents.id, rev.documentId));
return { slug: rev.slug, documentId: rev.documentId };
}
@@ -0,0 +1,19 @@
CREATE TABLE "document_revisions" (
"id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL,
"document_id" text NOT NULL,
"slug" text NOT NULL,
"revision_no" integer NOT NULL,
"title" text NOT NULL,
"body" text NOT NULL,
"meta" jsonb NOT NULL,
"reason" text DEFAULT 'index' NOT NULL,
"created_at" timestamp with time zone DEFAULT now() NOT NULL
);
--> statement-breakpoint
CREATE INDEX "document_revisions_document_id_idx" ON "document_revisions" USING btree ("document_id");
--> statement-breakpoint
CREATE INDEX "document_revisions_slug_idx" ON "document_revisions" USING btree ("slug");
--> statement-breakpoint
CREATE INDEX "document_revisions_doc_rev_idx" ON "document_revisions" USING btree ("document_id", "revision_no" DESC);
--> statement-breakpoint
ALTER TABLE "document_revisions" ADD CONSTRAINT "document_revisions_document_id_documents_id_fk" FOREIGN KEY ("document_id") REFERENCES "public"."documents"("id") ON DELETE cascade;
+110
View File
@@ -0,0 +1,110 @@
{
"id": "0002_document_revisions",
"prevId": "0001_document_chunks",
"version": "7",
"dialect": "postgresql",
"tables": {
"document_revisions": {
"name": "document_revisions",
"columns": {
"id": {
"name": "id",
"type": "uuid",
"primaryKey": true,
"notNull": true,
"default": "gen_random_uuid()"
},
"document_id": {
"name": "document_id",
"type": "text",
"notNull": true
},
"slug": {
"name": "slug",
"type": "text",
"notNull": true
},
"revision_no": {
"name": "revision_no",
"type": "integer",
"notNull": true
},
"title": {
"name": "title",
"type": "text",
"notNull": true
},
"body": {
"name": "body",
"type": "text",
"notNull": true
},
"meta": {
"name": "meta",
"type": "jsonb",
"notNull": true
},
"reason": {
"name": "reason",
"type": "text",
"notNull": true,
"default": "'index'"
},
"created_at": {
"name": "created_at",
"type": "timestamp",
"notNull": true,
"default": "now()"
}
},
"indexes": {
"document_revisions_document_id_idx": {
"name": "document_revisions_document_id_idx",
"columns": [
{ "name": "document_id", "asc": true }
],
"isUnique": false
},
"document_revisions_slug_idx": {
"name": "document_revisions_slug_idx",
"columns": [
{ "name": "slug", "asc": true }
],
"isUnique": false
},
"document_revisions_doc_rev_idx": {
"name": "document_revisions_doc_rev_idx",
"columns": [
{ "name": "document_id", "asc": true },
{ "name": "revision_no", "asc": false }
],
"isUnique": false
}
},
"foreignKeys": {
"document_revisions_document_id_documents_id_fk": {
"name": "document_revisions_document_id_documents_id_fk",
"columns": ["document_id"],
"referenceTable": "documents",
"referenceColumns": ["id"],
"onDelete": "cascade"
}
},
"compositePrimaryKeys": {},
"uniqueConstraints": {},
"policies": {}
}
},
"enums": {},
"schemas": {},
"sequences": {},
"roles": {},
"policies": {},
"views": {},
"extensions": {},
"_meta": {
"columns": {},
"schemas": {},
"tables": {}
}
}
+7
View File
@@ -15,6 +15,13 @@
"when": 1787137149735, "when": 1787137149735,
"tag": "0001_document_chunks", "tag": "0001_document_chunks",
"breakpoints": true "breakpoints": true
},
{
"idx": 2,
"version": "7",
"when": 1787139150000,
"tag": "0002_document_revisions",
"breakpoints": true
} }
] ]
} }
+44
View File
@@ -3,6 +3,7 @@ import {
customType, customType,
index, index,
integer, integer,
jsonb,
pgTable, pgTable,
real, real,
text, text,
@@ -77,5 +78,48 @@ export const documentChunks = pgTable(
export type DocumentChunkRow = typeof documentChunks.$inferSelect; export type DocumentChunkRow = typeof documentChunks.$inferSelect;
export type NewDocumentChunkRow = typeof documentChunks.$inferInsert; export type NewDocumentChunkRow = typeof documentChunks.$inferInsert;
// Phase 3: document revision system. Each row is an immutable snapshot of a
// document's body + metadata at a point in time (taken by the indexer whenever
// the body actually changes). revisionNo is per-document and monotonically
// increasing so the latest revision is always max(revision_no).
export const documentRevisions = pgTable(
"document_revisions",
{
id: uuid("id").primaryKey().defaultRandom(),
documentId: text("document_id")
.notNull()
.references(() => documents.id, { onDelete: "cascade" }),
slug: text("slug").notNull(),
revisionNo: integer("revision_no").notNull(),
title: text("title").notNull(),
body: text("body").notNull(),
// Metadata snapshot (type/section/status/author/tags) as JSON so a revision
// is self-describing even if the live document is later restructured.
meta: jsonb("meta").notNull().$type<{
type: string;
section: string;
status: string;
author: string;
tags: string[];
}>(),
// Why this revision was created (e.g. "index", "git-push", "restore").
reason: text("reason").notNull().default("index"),
createdAt: timestamp("created_at", { withTimezone: true })
.notNull()
.defaultNow(),
},
(t) => ({
docIdx: index("document_revisions_document_id_idx").on(t.documentId),
slugIdx: index("document_revisions_slug_idx").on(t.slug),
docRevIdx: index("document_revisions_doc_rev_idx").on(
t.documentId,
sql`${t.revisionNo} desc`,
),
}),
);
export type DocumentRevisionRow = typeof documentRevisions.$inferSelect;
export type NewDocumentRevisionRow = typeof documentRevisions.$inferInsert;
export type DocumentRow = typeof documents.$inferSelect; export type DocumentRow = typeof documents.$inferSelect;
export type NewDocumentRow = typeof documents.$inferInsert; export type NewDocumentRow = typeof documents.$inferInsert;
+22
View File
@@ -0,0 +1,22 @@
{
"name": "@mcpedia/queue",
"version": "0.1.0",
"private": true,
"type": "module",
"exports": {
".": "./src/index.ts",
"./client": "./src/client.ts",
"./queue": "./src/queue.ts",
"./worker": "./src/worker.ts"
},
"dependencies": {
"@mcpedia/config": "workspace:*",
"@mcpedia/core": "workspace:*",
"@mcpedia/db": "workspace:*",
"bullmq": "^6.1.2",
"ioredis": "^6.0.0"
},
"devDependencies": {
"typescript": "^5.6.0"
}
}
+34
View File
@@ -0,0 +1,34 @@
import { REDIS_URL, REDIS_PASSWORD, QUEUE_PREFIX } from "@mcpedia/config";
import IORedis, { type RedisOptions } from "ioredis";
/**
* Shared ioredis connection for BullMQ. BullMQ requires an ioredis instance and
* internally duplicates it for blocking commands, so we keep the option objects
* explicit (maxRetriesPerRequest: null is REQUIRED for the blocking
* connection — a finite retry count causes "Connection in key mode" errors).
*/
function buildOptions(): RedisOptions {
const opts: RedisOptions = {
maxRetriesPerRequest: null,
lazyConnect: true,
enableOfflineQueue: true,
};
if (REDIS_PASSWORD) opts.password = REDIS_PASSWORD;
return opts;
}
let _connection: IORedis | null = null;
/** Lazily-created singleton ioredis connection. */
export function getConnection(): IORedis {
if (!_connection) {
_connection = new IORedis(REDIS_URL, buildOptions());
_connection.on("error", (err) => {
// Log but don't crash the process on transient Redis errors.
console.error("[queue] redis error:", err.message);
});
}
return _connection;
}
export const BULLMQ_PREFIX = QUEUE_PREFIX;
+9
View File
@@ -0,0 +1,9 @@
export { getConnection, BULLMQ_PREFIX } from "./client";
export {
getQueue,
enqueueIndexDoc,
enqueueFullIndex,
INDEX_QUEUE,
} from "./queue";
export type { IndexDocJobData, IndexAllJobData, JobType } from "./queue";
export { createWorker, startWorker } from "./worker";
+66
View File
@@ -0,0 +1,66 @@
import { Queue, type Job } from "bullmq";
import { getConnection, BULLMQ_PREFIX } from "./client";
export const INDEX_QUEUE = "mcpedia-index";
/** Lazily-created singleton BullMQ queue. */
let _queue: Queue | null = null;
export function getQueue(): Queue {
if (!_queue) {
_queue = new Queue(INDEX_QUEUE, {
connection: getConnection(),
prefix: BULLMQ_PREFIX,
});
}
return _queue;
}
export interface IndexDocJobData {
relPath: string;
reason: string;
}
export interface IndexAllJobData {
reason: string;
}
export type JobType = "index-doc" | "index-all";
/**
* Enqueue a single-document reindex job. Keyed by slug so repeated edits
* collapse into one pending job (BullMQ dedup by jobId within the window).
*/
export async function enqueueIndexDoc(
relPath: string,
reason = "index",
): Promise<Job<IndexDocJobData>> {
const slug = relPath.replace(/\.mdx?$/, "");
return getQueue().add(
"index-doc",
{ relPath, reason },
{
jobId: `doc__${slug}`,
removeOnComplete: 1000,
removeOnFail: 5000,
attempts: 3,
backoff: { type: "exponential", delay: 2000 },
},
);
}
/** Enqueue a full-corpus reindex (used by the git-sync hook). */
export async function enqueueFullIndex(
reason = "reindex",
): Promise<Job<IndexAllJobData>> {
return getQueue().add(
"index-all",
{ reason },
{
jobId: `full__${Date.now()}`,
removeOnComplete: 100,
removeOnFail: 1000,
attempts: 1,
},
);
}
+63
View File
@@ -0,0 +1,63 @@
import { Worker, type Job } from "bullmq";
import { getConnection, BULLMQ_PREFIX } from "./client";
import { INDEX_QUEUE } from "./queue";
import { indexContentFile, runFullIndex } from "@mcpedia/core";
export function createWorker(): Worker {
const worker = new Worker(
INDEX_QUEUE,
async (job: Job) => {
switch (job.name) {
case "index-doc": {
const { relPath, reason } = job.data as {
relPath: string;
reason: string;
};
await job.log(`indexing ${relPath}`);
const r = await indexContentFile(relPath, reason);
return r;
}
case "index-all": {
const { reason } = job.data as { reason: string };
await job.log(`full index (${reason})`);
return await runFullIndex(reason);
}
default:
throw new Error(`unknown job type: ${job.name}`);
}
},
{
connection: getConnection(),
prefix: BULLMQ_PREFIX,
concurrency: 4,
},
);
worker.on("completed", (job) => {
console.log(`[worker] completed ${job.name} (${job.id})`);
});
worker.on("failed", (job, err) => {
console.error(`[worker] failed ${job?.name} (${job?.id}): ${err.message}`);
});
worker.on("error", (err) => {
console.error(`[worker] error:`, err.message);
});
return worker;
}
/** Start the worker and wire graceful shutdown. */
export async function startWorker(): Promise<Worker> {
const worker = createWorker();
console.log("[worker] indexing worker started");
const shutdown = async (sig: string) => {
console.log(`[worker] ${sig} received, closing...`);
await worker.close();
process.exit(0);
};
process.on("SIGINT", () => void shutdown("SIGINT"));
process.on("SIGTERM", () => void shutdown("SIGTERM"));
return worker;
}
+24
View File
@@ -0,0 +1,24 @@
import { enqueueIndexDoc, enqueueFullIndex } from "@mcpedia/queue";
// One-shot enqueue helper (no worker required to schedule work).
// Usage:
// bun run enqueue --all # full reindex
// bun run enqueue docs/websocket/contract # single doc (slug or rel path)
async function main() {
const arg = process.argv[2];
if (!arg || arg === "--all") {
const job = await enqueueFullIndex("manual");
console.log(`enqueued full reindex job ${job.id}`);
} else {
const relPath = arg.endsWith(".md") || arg.endsWith(".mdx") ? arg : `${arg}.md`;
const job = await enqueueIndexDoc(relPath, "manual");
console.log(`enqueued doc reindex job ${job.id} -> ${relPath}`);
}
await new Promise((r) => setTimeout(r, 500)); // allow the event loop to flush
process.exit(0);
}
main().catch((err) => {
console.error(err);
process.exit(1);
});
+9 -62
View File
@@ -1,67 +1,14 @@
import { db } from "@mcpedia/db"; import { runFullIndex } from "@mcpedia/core";
import { documents } from "@mcpedia/db/schema";
import { parseFile } from "@mcpedia/parser";
import { CONTENT_ROOT } from "@mcpedia/config";
import { listContentFiles, indexChunks } from "@mcpedia/core";
import { join } from "node:path";
// Phase 3: the indexer now goes through `runFullIndex`, the single indexing
// entry point shared with the BullMQ worker and the git-sync hook. It parses
// each content file, upserts `documents`, chunks+embeds, and snapshots a
// revision when the body changed.
async function main() { async function main() {
const files = listContentFiles(); const reason = process.argv[2] && process.argv[2].startsWith("--reason=")
let indexed = 0; ? process.argv[2].slice("--reason=".length)
let chunked = 0; : "index";
for (const rel of files) { await runFullIndex(reason);
const abs = join(CONTENT_ROOT, rel);
const { meta, body } = parseFile(abs, rel);
const nowIso =
meta.updatedAt && meta.updatedAt !== ""
? meta.updatedAt
: new Date().toISOString();
await db
.insert(documents)
.values({
id: meta.id,
slug: meta.slug,
title: meta.title,
type: meta.type,
section: meta.section,
status: meta.status,
author: meta.author,
tags: meta.tags,
path: meta.path,
body,
createdAt: new Date(meta.createdAt || nowIso),
updatedAt: new Date(nowIso),
})
.onConflictDoUpdate({
target: documents.slug,
set: {
title: meta.title,
type: meta.type,
section: meta.section,
status: meta.status,
author: meta.author,
tags: meta.tags,
path: meta.path,
body,
updatedAt: new Date(nowIso),
},
});
indexed++;
console.log(` indexed ${rel}`);
// Phase 2: chunk + embed for semantic search.
try {
const n = await indexChunks(meta.slug, body);
chunked += n;
console.log(` embedded ${n} chunks`);
} catch (err) {
console.error(
` embed FAILED for ${meta.slug}: ${err instanceof Error ? err.message : err}`,
);
// Don't abort the whole index over one doc's embedding failure.
}
}
console.log(`indexed ${indexed} documents, ${chunked} chunks embedded`);
} }
main() main()
+1
View File
@@ -10,6 +10,7 @@
"@mcpedia/embeddings": "workspace:*", "@mcpedia/embeddings": "workspace:*",
"@mcpedia/parser": "workspace:*", "@mcpedia/parser": "workspace:*",
"@mcpedia/search": "workspace:*", "@mcpedia/search": "workspace:*",
"@mcpedia/queue": "workspace:*",
"drizzle-orm": "^0.38.0", "drizzle-orm": "^0.38.0",
"postgres": "^3.4.5" "postgres": "^3.4.5"
} }