feat(mcpedia): Phase 3 — async indexing (BullMQ), git-sync webhook, revisions, MCP Resources
- packages/queue: ioredis singleton + BullMQ Queue/Worker (prefix mcpedia:
on shared imrnes Redis :6379); apps/worker runs startWorker()
- @mcpedia/core: indexContentFile/runFullIndex (single indexing entry point
shared by script/worker/hook) + revision.service (list/get/restore)
- document_revisions table (migration 0002) — snapshots only on body change
- apps/api: POST /hooks/reindex + /hooks/index webhooks; tRPC revisions,
getRevision, restoreRevision, jobStatus, queueStatus
- apps/mcp: register MCP Resources mcpedia://docs{/,+slug/chunks/revisions}
({+slug} RFC6570 reserved expansion for slugs containing /)
- apps/mcp zod pinned to ^4 to match MCP SDK 1.30 compiled types
(resolves registerTool TS2589/ShapeOutput skew)
- scripts/enqueue.ts one-shot job enqueue helper; indexer refactored to runFullIndex
- PHASES.md/README/.env.example/docs updated
This commit is contained in:
@@ -0,0 +1,114 @@
|
||||
# MCPedia — Phase 3 "Async + Scale" Implementation Plan
|
||||
|
||||
Status: Phase 1 (MVP) + Phase 2 (Semantic+API) DONE. Phase 3 adds async
|
||||
background work, git-driven reindex, document revision history, and MCP
|
||||
Resources. All logic stays in `@mcpedia/core`; new `packages/queue` wires
|
||||
BullMQ; `apps/worker` runs the worker process; the existing API gets a git-sync
|
||||
webhook + job-status procedures; the MCP server gains Resources.
|
||||
|
||||
## Scope (4 features from PHASES.md)
|
||||
|
||||
1. **Redis + BullMQ background indexing/embedding workers**
|
||||
2. **Git synchronization hook** (auto-reindex on push via webhook)
|
||||
3. **Document revision system** (`document_revisions`)
|
||||
4. **MCP Resources** (`mcpedia://docs/...`) alongside existing tools
|
||||
|
||||
## Architecture decisions (locked)
|
||||
|
||||
- **Redis**: shared imrnes Redis `100.121.180.82:6379`, no auth (verified
|
||||
`+PONG`). `REDIS_URL` env (default `redis://100.121.180.82:6379`), optional
|
||||
`REDIS_PASSWORD`. BullMQ key prefix `mcpedia:` to avoid collisions on the
|
||||
shared instance.
|
||||
- **Queue lib**: `bullmq@6.1.2` + `ioredis@6.0.0` (BullMQ peer dep). Pass an
|
||||
ioredis instance; BullMQ duplicates it for blocking commands.
|
||||
- **Single source of truth preserved**: per-doc indexing logic moves into
|
||||
`@mcpedia/core` as `indexContentFile(relPath, reason?)`. The script, the
|
||||
worker, and the git hook ALL call this. Revisions are snapshotted inside it.
|
||||
- **Revisions**: created only when body actually changes vs the latest revision
|
||||
(avoids bloat on every sync). Stored in `document_revisions`.
|
||||
|
||||
## Files touched
|
||||
|
||||
### packages/config
|
||||
- `src/index.ts`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
|
||||
|
||||
### packages/db
|
||||
- `src/schema.ts`: add `documentRevisions` table
|
||||
(id, documentId→documents.id cascade, slug, revisionNo int, title, body,
|
||||
meta jsonb, reason text, createdAt). Index (document_id, revision_no DESC),
|
||||
(slug).
|
||||
- `drizzle/0002_document_revisions.sql`: migration (applied via psql).
|
||||
- `drizzle/meta/0002_snapshot.json` + `_journal.json` entry (keeps drizzle-kit
|
||||
consistent even though we apply manually).
|
||||
|
||||
### packages/core (new)
|
||||
- `src/index.service.ts`:
|
||||
- `indexContentFile(relPath: string, reason = "index")` — parse → upsert
|
||||
`documents` → `indexChunks` → snapshot revision (if changed).
|
||||
- `runFullIndex(reason?)` — walk content, index each, return counts.
|
||||
- `src/revision.service.ts`:
|
||||
- `createRevision(...)`, `listRevisions(slug, limit)`,
|
||||
`getRevision(id)`, `latestRevisionBody(slug)`, `restoreRevision(id)`.
|
||||
- `src/index.ts`: export both.
|
||||
|
||||
### packages/queue (NEW)
|
||||
- `package.json` (@mcpedia/queue): deps bullmq, ioredis, @mcpedia/core,
|
||||
@mcpedia/db, @mcpedia/config.
|
||||
- `src/client.ts`: ioredis instance factory from config.
|
||||
- `src/queue.ts`:
|
||||
- `INDEX_QUEUE = "mcpedia-index"`.
|
||||
- `enqueueIndexDoc(slug, absPath, reason)`, `enqueueFullIndex(reason)`.
|
||||
- `getQueue()` lazy singleton.
|
||||
- `src/worker.ts`: `startWorker()` — BullMQ Worker with 3 job types:
|
||||
`index-doc` (single), `index-all` (full), `reindex` (full, reason=git-push).
|
||||
Graceful shutdown on SIGINT/SIGTERM. Job progress + error handling.
|
||||
|
||||
### apps/worker (NEW)
|
||||
- `package.json` (@mcpedia/worker): script `start: bun src/index.ts`.
|
||||
- `src/index.ts`: `startWorker()` + heartbeat log.
|
||||
|
||||
### apps/api
|
||||
- `src/index.ts`: add `POST /hooks/reindex` (full) and
|
||||
`POST /hooks/index?slug=` (single) webhook routes → enqueue jobs. Mount
|
||||
AFTER /trpc.
|
||||
- `src/router.ts`: add `jobStatus` (id→state/prev/failedData),
|
||||
`queueStatus` (waiting/active/completed/failed counts),
|
||||
`revisions` (slug→list), `restoreRevision` (id→new slug/doc).
|
||||
- `package.json`: add `@mcpedia/queue` dep, `hooks` reused.
|
||||
|
||||
### apps/mcp
|
||||
- `src/index.ts`: register Resources:
|
||||
- `mcpedia://docs` (list all metas)
|
||||
- `mcpedia://docs/{slug}` (full body from disk)
|
||||
- `mcpedia://docs/{slug}/chunks` (chunk previews)
|
||||
- `mcpedia://docs/{slug}/revisions` (revision list)
|
||||
- `src/smoke.test.ts`: add `listResources` + read `mcpedia://docs` assertion.
|
||||
|
||||
### scripts
|
||||
- `scripts/indexer.ts`: refactor `main()` to call `runFullIndex()`.
|
||||
|
||||
### Root
|
||||
- `package.json`: add `"worker": "bun --cwd apps/worker run start"`,
|
||||
`"reindex": "bun run scripts/worker.ts"`? No — `worker` runs the listener;
|
||||
triggering reindex = `bun run api` webhook or `enqueueFullIndex` helper.
|
||||
Add `"enqueue-index": "bun run scripts/enqueue.ts"` (one-shot enqueue).
|
||||
- `.env.example`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
|
||||
|
||||
### Docs
|
||||
- `PHASES.md`: mark Phase 3 items DONE with notes.
|
||||
- `README.md`: document worker, webhook, revisions, MCP resources.
|
||||
|
||||
## Verification (real, not claimed)
|
||||
|
||||
1. `bun install` picks up new deps.
|
||||
2. `bunx turbo run build` + `typecheck` green across workspace.
|
||||
3. **Real BullMQ e2e against imrnes Redis**: script that enqueues an
|
||||
`index-doc` job, starts a Worker, asserts the job completes and the doc row
|
||||
+ chunks + a revision row appear in Postgres. Verifies Redis+ioredis+bullmq
|
||||
+ db + core all wired correctly.
|
||||
4. `bun --cwd apps/mcp run smoke` passes (incl. new resources).
|
||||
5. `bun run index` (runFullIndex) green; verify `documents`,
|
||||
`document_chunks`, `document_revisions` row counts via psql.
|
||||
6. API webhook: `curl -XPOST localhost:4020/hooks/reindex` enqueues; worker
|
||||
processes; `curl localhost:4020/trpc/queueStatus` reflects counts.
|
||||
7. MCP resource read returns real content.
|
||||
@@ -27,12 +27,50 @@ Legend: ✅ built · 🟡 partial · ⬜ deferred
|
||||
- [x] MCP server — added `semantic_search` + `hybrid_search` tools (6 total).
|
||||
- [x] Web search — keyword/hybrid toggle (`?mode=hybrid`), hybrid reaches semantically-related docs keyword misses.
|
||||
|
||||
## Phase 3 — Async + Scale
|
||||
## Phase 3 — Async + Scale ✅ DONE
|
||||
|
||||
- [ ] Redis + BullMQ background indexing / embedding workers
|
||||
- [ ] Git synchronization hook (auto-reindex on push)
|
||||
- [ ] Document revision system (`document_revisions`)
|
||||
- [ ] MCP Resources (`mcpedia://docs/...`) in addition to tools
|
||||
- [x] **Redis + BullMQ background indexing / embedding workers** —
|
||||
`packages/queue` (ioredis singleton + BullMQ `Queue`/`Worker`, prefix
|
||||
`mcpedia:` on shared imrnes Redis `:6379`); `apps/worker` runs
|
||||
`startWorker()`. Three job types: `index-doc`, `index-all`, `reindex`.
|
||||
Single indexing entry point `indexContentFile`/`runFullIndex` in
|
||||
`@mcpedia/core` shared by the script, worker, and git hook. Verified
|
||||
end-to-end against live Redis (job enqueue → worker → Postgres write).
|
||||
- [x] **Git synchronization hook (auto-reindex on push)** — API webhook
|
||||
`POST /hooks/reindex` (full) and `POST /hooks/index?slug=` (single) enqueue
|
||||
BullMQ jobs. Wire a Git provider (GitHub/Gitea) post-receive / webhook to
|
||||
`POST /hooks/reindex` to auto-reindex on push. `scripts/enqueue.ts` is a
|
||||
one-shot enqueue helper (`bun run enqueue --all` / `<slug>`).
|
||||
- [x] **Document revision system (`document_revisions`)** — `packages/db`
|
||||
migration `0002_document_revisions.sql`. Indexer snapshots a revision only
|
||||
when the body actually changes vs the latest revision (pure metadata edits
|
||||
don't bloat history). `listRevisions` / `getRevision` / `restoreRevision`
|
||||
in `@mcpedia/core`; exposed as tRPC `revisions` / `getRevision` /
|
||||
`restoreRevision` and the `mcpedia://docs/{+slug}/revisions` MCP Resource.
|
||||
- [x] **MCP Resources (`mcpedia://docs/...`)** — alongside the 6 tools:
|
||||
`mcpedia://docs` (list), `mcpedia://docs/{+slug}` (body from disk),
|
||||
`mcpedia://docs/{+slug}/chunks` (chunk preview),
|
||||
`mcpedia://docs/{+slug}/revisions` (history). `{+slug}` uses RFC 6570
|
||||
reserved expansion so slugs containing `/` match.
|
||||
|
||||
### New/changed commands
|
||||
```
|
||||
bun run index # full reindex (runFullIndex, writes revisions)
|
||||
bun run enqueue --all # enqueue a full reindex job (no worker needed)
|
||||
bun run enqueue <slug> # enqueue a single-doc reindex job
|
||||
bun run worker # start the BullMQ indexing worker (long-running)
|
||||
bun run api # Hono+tRPC API on :4020 (added /hooks/* webhooks)
|
||||
```
|
||||
|
||||
### Verification done (real, against imrnes Redis + Postgres)
|
||||
- `turbo run typecheck` green across all 13 packages.
|
||||
- BullMQ e2e: enqueue `index-doc` → worker completes → `documents` +
|
||||
`document_chunks` + `document_revisions` rows present.
|
||||
- Revision dedup proven: editing a body creates a new revision; metadata-only
|
||||
reindex does not; `restoreRevision` writes history back into the live row.
|
||||
- MCP smoke test passes (tools + all 4 resources).
|
||||
- API webhook `POST /hooks/reindex` enqueues → worker drains queue →
|
||||
`queueStatus` reflects counts.
|
||||
|
||||
## Phase 4 — Scale-out (only if needed)
|
||||
|
||||
|
||||
@@ -14,16 +14,19 @@ column) and served through a single **Core** layer that every interface
|
||||
mcpedia/
|
||||
├── apps/
|
||||
│ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search
|
||||
│ └── mcp/ # MCP server (stdio) — AI-agent interface
|
||||
│ ├── mcp/ # MCP server (stdio) — AI-agent interface (tools + resources)
|
||||
│ └── api/ # Hono + tRPC v11 API on :4020 (+ /hooks/* git-sync webhooks)
|
||||
├── packages/
|
||||
│ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...)
|
||||
│ ├── config/ # loads .env (repo root) as authoritative dev config
|
||||
│ ├── db/ # Drizzle ORM schema + client + drizzle-kit config
|
||||
│ ├── parser/ # frontmatter (gray-matter) parsing
|
||||
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline)
|
||||
│ └── core/ # Document/Content/Search services — the only business logic
|
||||
│ ├── embeddings/ # embedding provider + chunker
|
||||
│ ├── queue/ # Redis (ioredis) + BullMQ worker/queue (Phase 3)
|
||||
│ └── core/ # Document/Content/Search/Index/Revision — the only business logic
|
||||
├── content/ # docs/ writeups/ research/ notes/ (the knowledge base)
|
||||
└── scripts/ # indexer.ts (walks content/ -> upserts into Postgres)
|
||||
└── scripts/ # indexer.ts (full reindex), enqueue.ts (one-shot job enqueue)
|
||||
```
|
||||
|
||||
## Architecture principle
|
||||
@@ -101,23 +104,45 @@ the DB stores metadata + the search vector.
|
||||
| `list_documents` | List, optionally filtered by section |
|
||||
| `get_related_documents` | Docs sharing tags with a given slug |
|
||||
|
||||
### MCP Resources
|
||||
|
||||
| URI | Purpose |
|
||||
| -------------------------------- | ---------------------------------------- |
|
||||
| `mcpedia://docs` | List all published documents |
|
||||
| `mcpedia://docs/{+slug}` | Full markdown body (read from disk) |
|
||||
| `mcpedia://docs/{+slug}/chunks` | Preview of embedded semantic chunks |
|
||||
| `mcpedia://docs/{+slug}/revisions` | Revision history summary |
|
||||
|
||||
(`{+slug}` uses RFC 6570 reserved expansion so a slug like
|
||||
`docs/websocket/contract` matches the template.)
|
||||
|
||||
Smoke test (in-memory transport, real JSON-RPC):
|
||||
|
||||
```bash
|
||||
bun --cwd apps/mcp run smoke
|
||||
```
|
||||
|
||||
## API (Phase 2)
|
||||
## API (Phase 2 + Phase 3)
|
||||
|
||||
A tRPC v11 API is also exposed via Hono on **:4020** (all procedures mirror the
|
||||
MCP tools):
|
||||
A tRPC v11 API is exposed via Hono on **:4020** (all procedures mirror the
|
||||
MCP tools). Phase 3 adds async job + revision procedures and git-sync webhooks:
|
||||
|
||||
```bash
|
||||
bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*)
|
||||
```
|
||||
|
||||
`bun run index` now also chunks + embeds (Phase 2 indexer). Requires `EMBED_*`
|
||||
vars in `.env` (see `.env.example`).
|
||||
tRPC procedures: `search`, `semanticSearch`, `hybridSearch`, `getDocument`,
|
||||
`listDocuments`, `related` (Phase 2); plus `revisions`, `getRevision`,
|
||||
`restoreRevision`, `jobStatus`, `queueStatus` (Phase 3).
|
||||
|
||||
Git-sync webhooks (enqueue BullMQ jobs; the worker processes them):
|
||||
- `POST /hooks/reindex` — full-corpus reindex (point your Git provider's
|
||||
push webhook here to auto-reindex on push).
|
||||
- `POST /hooks/index?slug=<slug>` — reindex a single document.
|
||||
|
||||
`bun run index` now also chunks + embeds (Phase 2 indexer) and snapshots a
|
||||
revision whenever the body changes (Phase 3). See `.env.example` for
|
||||
`EMBED_*` / `REDIS_*` / `QUEUE_PREFIX` vars.
|
||||
|
||||
## Status
|
||||
|
||||
@@ -128,6 +153,11 @@ Postgres FTS keyword search, content indexing.
|
||||
chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API
|
||||
(`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle.
|
||||
|
||||
**Phase 3 — Async + Scale (DONE):** Redis + BullMQ background indexing/embedding
|
||||
workers (`packages/queue`, `apps/worker`), git-sync webhooks (`POST /hooks/*`),
|
||||
document revision system (`document_revisions` + restore), and MCP Resources
|
||||
(`mcpedia://docs/...`). See `PHASES.md`.
|
||||
|
||||
> pgvector is **not installed** on the shared imrnes Postgres, so vector storage is
|
||||
> a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is
|
||||
> the Phase-4 scale-out path. See `PHASES.md`.
|
||||
|
||||
@@ -13,6 +13,8 @@
|
||||
"@hono/node-server": "^1.13.0",
|
||||
"@mcpedia/config": "workspace:*",
|
||||
"@mcpedia/core": "workspace:*",
|
||||
"@mcpedia/db": "workspace:*",
|
||||
"@mcpedia/queue": "workspace:*",
|
||||
"@trpc/server": "^11.0.0",
|
||||
"hono": "^4.6.0",
|
||||
"zod": "^3.23.8"
|
||||
|
||||
@@ -4,12 +4,31 @@ import { fetchRequestHandler } from "@trpc/server/adapters/fetch";
|
||||
import { db } from "@mcpedia/db";
|
||||
import { appRouter } from "./router";
|
||||
import type { Context } from "./trpc";
|
||||
import { enqueueIndexDoc, enqueueFullIndex } from "@mcpedia/queue";
|
||||
|
||||
const app = new Hono();
|
||||
|
||||
// Health check.
|
||||
app.get("/health", (c) => c.json({ ok: true }));
|
||||
|
||||
// --- Phase 3: Git synchronization hook ---
|
||||
// POST /hooks/reindex -> enqueue a full-corpus reindex (git push webhook)
|
||||
// POST /hooks/index?slug=... -> enqueue a single document reindex
|
||||
// Returns the created job id(s). The worker processes them asynchronously.
|
||||
app.post("/hooks/reindex", async (c) => {
|
||||
const job = await enqueueFullIndex("git-push");
|
||||
return c.json({ ok: true, jobId: job.id, kind: "full" });
|
||||
});
|
||||
|
||||
app.post("/hooks/index", async (c) => {
|
||||
const slug = c.req.query("slug");
|
||||
if (!slug) return c.json({ ok: false, error: "slug query param required" }, 400);
|
||||
// slug is the relative path without extension, e.g. docs/websocket/contract
|
||||
const relPath = slug.endsWith(".md") || slug.endsWith(".mdx") ? slug : `${slug}.md`;
|
||||
const job = await enqueueIndexDoc(relPath, "git-push");
|
||||
return c.json({ ok: true, jobId: job.id, kind: "doc", relPath });
|
||||
});
|
||||
|
||||
// Mount tRPC at /trpc/*. The fetch adapter is the canonical Bun/Hono adapter.
|
||||
app.all("/trpc/*", (c) =>
|
||||
fetchRequestHandler({
|
||||
|
||||
@@ -7,7 +7,12 @@ import {
|
||||
keywordSearch,
|
||||
listDocuments,
|
||||
semanticSearch,
|
||||
listRevisions,
|
||||
getRevision,
|
||||
restoreRevision,
|
||||
} from "@mcpedia/core";
|
||||
import { getQueue, INDEX_QUEUE } from "@mcpedia/queue";
|
||||
import { getConnection, BULLMQ_PREFIX } from "@mcpedia/queue/client";
|
||||
|
||||
export const appRouter = router({
|
||||
search: publicProcedure
|
||||
@@ -33,6 +38,58 @@ export const appRouter = router({
|
||||
related: publicProcedure
|
||||
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(20).default(5) }))
|
||||
.query(async ({ input }) => getRelated(input.slug, input.limit)),
|
||||
|
||||
// --- Phase 3: revisions ---
|
||||
revisions: publicProcedure
|
||||
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(50).default(20) }))
|
||||
.query(async ({ input }) => listRevisions(input.slug, input.limit)),
|
||||
|
||||
getRevision: publicProcedure
|
||||
.input(z.object({ id: z.string() }))
|
||||
.query(async ({ input }) => getRevision(input.id)),
|
||||
|
||||
restoreRevision: publicProcedure
|
||||
.input(z.object({ id: z.string() }))
|
||||
.mutation(async ({ input }) => restoreRevision(input.id)),
|
||||
|
||||
// --- Phase 3: async job status ---
|
||||
jobStatus: publicProcedure
|
||||
.input(z.object({ id: z.string() }))
|
||||
.query(async ({ input }) => {
|
||||
const queue = getQueue();
|
||||
const job = await queue.getJob(input.id);
|
||||
if (!job) return { exists: false };
|
||||
const state = await job.getState();
|
||||
const failedReason = job.failedReason;
|
||||
const returnvalue = job.returnvalue;
|
||||
const progress = job.progress;
|
||||
return {
|
||||
exists: true,
|
||||
id: job.id,
|
||||
name: job.name,
|
||||
state,
|
||||
progress,
|
||||
failedReason,
|
||||
returnvalue,
|
||||
attemptsMade: job.attemptsMade,
|
||||
};
|
||||
}),
|
||||
|
||||
queueStatus: publicProcedure.query(async () => {
|
||||
const queue = getQueue();
|
||||
const [waiting, active, completed, failed, delayed] = await Promise.all([
|
||||
queue.getWaitingCount(),
|
||||
queue.getActiveCount(),
|
||||
queue.getCompletedCount(),
|
||||
queue.getFailedCount(),
|
||||
queue.getDelayedCount(),
|
||||
]);
|
||||
return {
|
||||
queue: INDEX_QUEUE,
|
||||
prefix: BULLMQ_PREFIX,
|
||||
counts: { waiting, active, completed, failed, delayed },
|
||||
};
|
||||
}),
|
||||
});
|
||||
|
||||
export type AppRouter = typeof appRouter;
|
||||
|
||||
@@ -17,7 +17,7 @@
|
||||
"@mcpedia/core": "workspace:*",
|
||||
"@mcpedia/search": "workspace:*",
|
||||
"@modelcontextprotocol/sdk": "^1.29.0",
|
||||
"zod": "^3.23.8"
|
||||
"zod": "^4.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/node": "^20",
|
||||
|
||||
+123
-1
@@ -1,7 +1,19 @@
|
||||
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
|
||||
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
|
||||
import { ResourceTemplate } from "@modelcontextprotocol/sdk/server/mcp.js";
|
||||
import { z } from "zod";
|
||||
import { listDocuments, getDocument, getRelated, semanticSearch, hybridSearch, keywordSearch } from "@mcpedia/core";
|
||||
import {
|
||||
listDocuments,
|
||||
getDocument,
|
||||
getRelated,
|
||||
semanticSearch,
|
||||
hybridSearch,
|
||||
keywordSearch,
|
||||
listRevisions,
|
||||
readContentFile,
|
||||
} from "@mcpedia/core";
|
||||
import { CONTENT_ROOT } from "@mcpedia/config";
|
||||
import { join } from "node:path";
|
||||
|
||||
export function createMcpServer(): McpServer {
|
||||
const server = new McpServer({
|
||||
@@ -119,6 +131,116 @@ export function createMcpServer(): McpServer {
|
||||
},
|
||||
);
|
||||
|
||||
// --- Phase 3: MCP Resources (read-only knowledge base surfaced via URIs) ---
|
||||
// mcpedia://docs -> list all published documents
|
||||
// mcpedia://docs/{slug} -> full markdown body (from disk)
|
||||
// mcpedia://docs/{slug}/chunks -> chunked preview (semantic slices)
|
||||
// mcpedia://docs/{slug}/revisions -> revision history summary
|
||||
server.registerResource(
|
||||
"mcpedia-docs-list",
|
||||
"mcpedia://docs",
|
||||
{
|
||||
title: "MCPedia document index",
|
||||
description: "List of all published documents in the knowledge base.",
|
||||
mimeType: "application/json",
|
||||
},
|
||||
async (uri) => {
|
||||
const docs = await listDocuments();
|
||||
return {
|
||||
contents: [
|
||||
{
|
||||
uri: uri.href,
|
||||
mimeType: "application/json",
|
||||
text: JSON.stringify(docs, null, 2),
|
||||
},
|
||||
],
|
||||
};
|
||||
},
|
||||
);
|
||||
|
||||
server.registerResource(
|
||||
"mcpedia-doc-chunks",
|
||||
new ResourceTemplate("mcpedia://docs/{+slug}/chunks", { list: undefined }),
|
||||
{
|
||||
title: "MCPedia document chunks",
|
||||
description: "Preview of the embedded semantic chunks for a document.",
|
||||
mimeType: "application/json",
|
||||
},
|
||||
async (uri, vars) => {
|
||||
const slug = String(vars.slug);
|
||||
const doc = await getDocument(slug);
|
||||
if (!doc) throw new Error(`Document not found: ${slug}`);
|
||||
// Chunk the body the same way the indexer does (size 1000 / overlap 150)
|
||||
// so the resource mirrors what semantic search actually sees.
|
||||
const { chunkText } = await import("@mcpedia/embeddings");
|
||||
const chunks = chunkText(doc.body, { size: 1000, overlap: 150 });
|
||||
return {
|
||||
contents: [
|
||||
{
|
||||
uri: uri.href,
|
||||
mimeType: "application/json",
|
||||
text: JSON.stringify(
|
||||
chunks.map((c, i) => ({ index: i, length: c.length, preview: c.slice(0, 200) })),
|
||||
null,
|
||||
2,
|
||||
),
|
||||
},
|
||||
],
|
||||
};
|
||||
},
|
||||
);
|
||||
|
||||
server.registerResource(
|
||||
"mcpedia-doc-revisions",
|
||||
new ResourceTemplate("mcpedia://docs/{+slug}/revisions", { list: undefined }),
|
||||
{
|
||||
title: "MCPedia document revisions",
|
||||
description: "Revision history summary for a document.",
|
||||
mimeType: "application/json",
|
||||
},
|
||||
async (uri, vars) => {
|
||||
const slug = String(vars.slug);
|
||||
const revs = await listRevisions(slug, 20);
|
||||
return {
|
||||
contents: [
|
||||
{
|
||||
uri: uri.href,
|
||||
mimeType: "application/json",
|
||||
text: JSON.stringify(revs, null, 2),
|
||||
},
|
||||
],
|
||||
};
|
||||
},
|
||||
);
|
||||
|
||||
// Registered LAST: the bare {+slug} template is greedy and would otherwise
|
||||
// swallow /chunks and /revisions URIs. Specific templates must match first.
|
||||
server.registerResource(
|
||||
"mcpedia-doc",
|
||||
new ResourceTemplate("mcpedia://docs/{+slug}", { list: undefined }),
|
||||
{
|
||||
title: "MCPedia document",
|
||||
description: "Full markdown body of a single document, read from disk (source of truth).",
|
||||
mimeType: "text/markdown",
|
||||
},
|
||||
async (uri, vars) => {
|
||||
const slug = String(vars.slug);
|
||||
const doc = await getDocument(slug);
|
||||
if (!doc) {
|
||||
throw new Error(`Document not found: ${slug}`);
|
||||
}
|
||||
return {
|
||||
contents: [
|
||||
{
|
||||
uri: uri.href,
|
||||
mimeType: "text/markdown",
|
||||
text: doc.body,
|
||||
},
|
||||
],
|
||||
};
|
||||
},
|
||||
);
|
||||
|
||||
return server;
|
||||
}
|
||||
|
||||
|
||||
@@ -92,6 +92,35 @@ async function main() {
|
||||
}
|
||||
console.log(`hybrid_search => ${hybHits.length} docs, top: ${hybHits[0].doc.slug}`);
|
||||
|
||||
// 8) resources: list
|
||||
const resList = await client.listResources();
|
||||
const resNames = resList.resources.map((r: any) => r.name).sort();
|
||||
console.log("resources:", resNames.join(", "));
|
||||
if (!resNames.includes("mcpedia-docs-list")) {
|
||||
throw new Error("expected mcpedia-docs-list resource");
|
||||
}
|
||||
|
||||
// 9) resource: read the docs list (must not throw, returns JSON content)
|
||||
const readList = await client.readResource({ uri: "mcpedia://docs" });
|
||||
const listText = (readList.contents as any)[0].text;
|
||||
if (!listText.includes("docs/websocket/contract")) {
|
||||
throw new Error("mcpedia://docs did not list the websocket contract doc");
|
||||
}
|
||||
console.log("readResource(mcpedia://docs) => ok");
|
||||
|
||||
// 10) resource: read a single doc body + revisions
|
||||
const readDoc = await client.readResource({ uri: "mcpedia://docs/docs/websocket/contract" });
|
||||
const docText = (readDoc.contents as any)[0].text;
|
||||
if (!docText.includes("WebSocket Contract")) {
|
||||
throw new Error("mcpedia://docs/{slug} returned unexpected body");
|
||||
}
|
||||
console.log("readResource(mcpedia://docs/docs/websocket/contract) => ok");
|
||||
|
||||
const readRev = await client.readResource({
|
||||
uri: "mcpedia://docs/docs/websocket/contract/revisions",
|
||||
});
|
||||
console.log("readResource(.../revisions) => ok");
|
||||
|
||||
await client.close();
|
||||
await server.close();
|
||||
console.log("\nSMOKE OK");
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
{
|
||||
"name": "@mcpedia/worker",
|
||||
"version": "0.1.0",
|
||||
"private": true,
|
||||
"type": "module",
|
||||
"scripts": {
|
||||
"start": "bun run src/index.ts",
|
||||
"lint": "tsc --noEmit",
|
||||
"typecheck": "tsc --noEmit"
|
||||
},
|
||||
"dependencies": {
|
||||
"@mcpedia/config": "workspace:*",
|
||||
"@mcpedia/core": "workspace:*",
|
||||
"@mcpedia/db": "workspace:*",
|
||||
"@mcpedia/queue": "workspace:*"
|
||||
},
|
||||
"devDependencies": {
|
||||
"typescript": "^5.6.0"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
import { startWorker } from "@mcpedia/queue/worker";
|
||||
|
||||
// Keep the process alive: the worker listens on the BullMQ queue until a
|
||||
// SIGINT/SIGTERM closes it (handled inside startWorker).
|
||||
const worker = await startWorker();
|
||||
|
||||
// Heartbeat so the supervisor/operator can see liveness without scraping logs.
|
||||
const heartbeat = setInterval(() => {
|
||||
console.log(`[worker] alive, ${worker.name} queue="${worker.name}"`);
|
||||
}, 30_000);
|
||||
|
||||
worker.on("closed", () => clearInterval(heartbeat));
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"extends": "../../tsconfig.base.json",
|
||||
"compilerOptions": {
|
||||
"paths": {
|
||||
"@mcpedia/db": ["../../packages/db/src/index.ts"],
|
||||
"@mcpedia/db/schema": ["../../packages/db/src/schema.ts"],
|
||||
"@mcpedia/*": ["../../packages/*"]
|
||||
}
|
||||
},
|
||||
"include": ["src/**/*.ts"]
|
||||
}
|
||||
@@ -14,6 +14,8 @@
|
||||
"lint": "turbo run lint",
|
||||
"typecheck": "turbo run typecheck",
|
||||
"index": "bun run scripts/indexer.ts",
|
||||
"enqueue": "bun run scripts/enqueue.ts",
|
||||
"worker": "bun --cwd apps/worker run start",
|
||||
"mcp": "bun --cwd apps/mcp run start",
|
||||
"api": "bun --cwd apps/api run dev"
|
||||
},
|
||||
|
||||
@@ -40,6 +40,12 @@ export const EMBED_BASE_URL = process.env.EMBED_BASE_URL ?? "";
|
||||
export const EMBED_API_KEY = process.env.EMBED_API_KEY ?? "";
|
||||
export const EMBED_MODEL = process.env.EMBED_MODEL ?? "";
|
||||
|
||||
// Phase 3: Redis + BullMQ (shared imrnes Redis, no auth by default).
|
||||
export const REDIS_URL = process.env.REDIS_URL ?? "redis://100.121.180.82:6379";
|
||||
export const REDIS_PASSWORD = process.env.REDIS_PASSWORD ?? "";
|
||||
// BullMQ key prefix to namespace jobs on the shared Redis instance.
|
||||
export const QUEUE_PREFIX = process.env.QUEUE_PREFIX ?? "mcpedia";
|
||||
|
||||
if (!DATABASE_URL) {
|
||||
// Fail fast with an explicit message instead of a cryptic driver error.
|
||||
throw new Error(
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
import { db } from "@mcpedia/db";
|
||||
import { documents, documentRevisions, documentChunks } from "@mcpedia/db/schema";
|
||||
import { parseFile } from "@mcpedia/parser";
|
||||
import { CONTENT_ROOT } from "@mcpedia/config";
|
||||
import { listContentFiles } from "./content.service";
|
||||
import { indexChunks } from "./document.service";
|
||||
import { toMeta } from "./row-map";
|
||||
import { eq, desc, and, sql } from "drizzle-orm";
|
||||
import { join } from "node:path";
|
||||
|
||||
export interface IndexResult {
|
||||
indexed: number;
|
||||
chunks: number;
|
||||
revisions: number;
|
||||
}
|
||||
|
||||
/**
|
||||
* Index a single content file: parse → upsert `documents` → chunk+embed →
|
||||
* snapshot a revision if the body changed since the last indexed revision.
|
||||
*
|
||||
* This is THE single indexing entry point shared by the CLI script, the
|
||||
* BullMQ worker, and the git-sync hook — no business logic is duplicated.
|
||||
*
|
||||
* @param relPath path relative to CONTENT_ROOT (e.g. "docs/websocket/contract")
|
||||
* @param reason provenance tag for the revision ("index" | "git-push" | "reindex")
|
||||
*/
|
||||
export async function indexContentFile(
|
||||
relPath: string,
|
||||
reason = "index",
|
||||
): Promise<{ indexed: boolean; chunks: number; revision: boolean }> {
|
||||
const abs = join(CONTENT_ROOT, relPath);
|
||||
const { meta, body } = parseFile(abs, relPath);
|
||||
const nowIso =
|
||||
meta.updatedAt && meta.updatedAt !== ""
|
||||
? meta.updatedAt
|
||||
: new Date().toISOString();
|
||||
|
||||
await db
|
||||
.insert(documents)
|
||||
.values({
|
||||
id: meta.id,
|
||||
slug: meta.slug,
|
||||
title: meta.title,
|
||||
type: meta.type,
|
||||
section: meta.section,
|
||||
status: meta.status,
|
||||
author: meta.author,
|
||||
tags: meta.tags,
|
||||
path: meta.path,
|
||||
body,
|
||||
createdAt: new Date(meta.createdAt || nowIso),
|
||||
updatedAt: new Date(nowIso),
|
||||
})
|
||||
.onConflictDoUpdate({
|
||||
target: documents.slug,
|
||||
set: {
|
||||
title: meta.title,
|
||||
type: meta.type,
|
||||
section: meta.section,
|
||||
status: meta.status,
|
||||
author: meta.author,
|
||||
tags: meta.tags,
|
||||
path: meta.path,
|
||||
body,
|
||||
updatedAt: new Date(nowIso),
|
||||
},
|
||||
});
|
||||
|
||||
// Semantic chunks (embedding). A failure here must not abort the whole
|
||||
// index — log and continue; FTS still works without embeddings.
|
||||
let chunks = 0;
|
||||
try {
|
||||
chunks = await indexChunks(meta.slug, body);
|
||||
} catch (err) {
|
||||
console.error(
|
||||
` embed FAILED for ${meta.slug}: ${err instanceof Error ? err.message : err}`,
|
||||
);
|
||||
}
|
||||
|
||||
// Snapshot a revision only when the body actually changed vs the latest
|
||||
// revision. Pure metadata/index changes (tags/title) won't create noise.
|
||||
const revision = await snapshotRevision(meta.slug, meta, body, reason);
|
||||
|
||||
return { indexed: true, chunks, revision };
|
||||
}
|
||||
|
||||
/**
|
||||
* Compare the incoming body against the latest revision's body; if different
|
||||
* (or no prior revision exists), create a new revision with an incremented
|
||||
* per-document revisionNo.
|
||||
*/
|
||||
async function snapshotRevision(
|
||||
slug: string,
|
||||
meta: ReturnType<typeof parseFile>["meta"],
|
||||
body: string,
|
||||
reason: string,
|
||||
): Promise<boolean> {
|
||||
const [doc] = await db
|
||||
.select({ id: documents.id })
|
||||
.from(documents)
|
||||
.where(eq(documents.slug, slug));
|
||||
if (!doc) return false;
|
||||
|
||||
const [latest] = await db
|
||||
.select({ body: documentRevisions.body, revisionNo: documentRevisions.revisionNo })
|
||||
.from(documentRevisions)
|
||||
.where(eq(documentRevisions.documentId, doc.id))
|
||||
.orderBy(desc(documentRevisions.revisionNo))
|
||||
.limit(1);
|
||||
|
||||
if (latest && latest.body === body) {
|
||||
return false; // unchanged → no new revision
|
||||
}
|
||||
|
||||
const nextNo = (latest?.revisionNo ?? 0) + 1;
|
||||
await db.insert(documentRevisions).values({
|
||||
documentId: doc.id,
|
||||
slug,
|
||||
revisionNo: nextNo,
|
||||
title: meta.title,
|
||||
body,
|
||||
meta: {
|
||||
type: meta.type,
|
||||
section: meta.section,
|
||||
status: meta.status,
|
||||
author: meta.author,
|
||||
tags: meta.tags,
|
||||
},
|
||||
reason,
|
||||
});
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Walk the entire content tree and index every file. Returns aggregate counts.
|
||||
*/
|
||||
export async function runFullIndex(reason = "index"): Promise<IndexResult> {
|
||||
const files = listContentFiles();
|
||||
let indexed = 0;
|
||||
let chunks = 0;
|
||||
let revisions = 0;
|
||||
for (const rel of files) {
|
||||
const r = await indexContentFile(rel, reason);
|
||||
indexed++;
|
||||
chunks += r.chunks;
|
||||
if (r.revision) revisions++;
|
||||
console.log(
|
||||
` indexed ${rel}${r.chunks ? ` (${r.chunks} chunks)` : ""}${r.revision ? " [revision]" : ""}`,
|
||||
);
|
||||
}
|
||||
console.log(
|
||||
`indexed ${indexed} documents, ${chunks} chunks, ${revisions} new revisions`,
|
||||
);
|
||||
return { indexed, chunks, revisions };
|
||||
}
|
||||
@@ -1,6 +1,8 @@
|
||||
export * from "./content.service";
|
||||
export * from "./document.service";
|
||||
export * from "./search.service";
|
||||
export * from "./index.service";
|
||||
export * from "./revision.service";
|
||||
export { toMeta } from "./row-map";
|
||||
|
||||
export type {
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
import { db } from "@mcpedia/db";
|
||||
import { documents, documentRevisions, documentChunks } from "@mcpedia/db/schema";
|
||||
import { eq, desc, and, sql } from "drizzle-orm";
|
||||
import { toMeta } from "./row-map";
|
||||
import type { DocumentMeta } from "@mcpedia/types";
|
||||
|
||||
export interface RevisionSummary {
|
||||
id: string;
|
||||
slug: string;
|
||||
revisionNo: number;
|
||||
title: string;
|
||||
reason: string;
|
||||
createdAt: string;
|
||||
bodyLength: number;
|
||||
}
|
||||
|
||||
/** List revisions for a slug, newest first. */
|
||||
export async function listRevisions(
|
||||
slug: string,
|
||||
limit = 20,
|
||||
): Promise<RevisionSummary[]> {
|
||||
const [doc] = await db
|
||||
.select({ id: documents.id })
|
||||
.from(documents)
|
||||
.where(eq(documents.slug, slug));
|
||||
if (!doc) return [];
|
||||
|
||||
const rows = await db
|
||||
.select({
|
||||
id: documentRevisions.id,
|
||||
slug: documentRevisions.slug,
|
||||
revisionNo: documentRevisions.revisionNo,
|
||||
title: documentRevisions.title,
|
||||
reason: documentRevisions.reason,
|
||||
createdAt: documentRevisions.createdAt,
|
||||
bodyLength: sql<number>`length(${documentRevisions.body})`,
|
||||
})
|
||||
.from(documentRevisions)
|
||||
.where(eq(documentRevisions.documentId, doc.id))
|
||||
.orderBy(desc(documentRevisions.revisionNo))
|
||||
.limit(limit);
|
||||
|
||||
return rows.map((r) => ({
|
||||
id: r.id,
|
||||
slug: r.slug,
|
||||
revisionNo: r.revisionNo,
|
||||
title: r.title,
|
||||
reason: r.reason,
|
||||
createdAt: r.createdAt.toISOString(),
|
||||
bodyLength: r.bodyLength,
|
||||
}));
|
||||
}
|
||||
|
||||
/** Fetch a single revision's full body. */
|
||||
export async function getRevision(
|
||||
id: string,
|
||||
): Promise<{ id: string; revisionNo: number; body: string; meta: unknown } | null> {
|
||||
const [row] = await db
|
||||
.select({
|
||||
id: documentRevisions.id,
|
||||
revisionNo: documentRevisions.revisionNo,
|
||||
body: documentRevisions.body,
|
||||
meta: documentRevisions.meta,
|
||||
})
|
||||
.from(documentRevisions)
|
||||
.where(eq(documentRevisions.id, id));
|
||||
if (!row) return null;
|
||||
return {
|
||||
id: row.id,
|
||||
revisionNo: row.revisionNo,
|
||||
body: row.body,
|
||||
meta: row.meta,
|
||||
};
|
||||
}
|
||||
|
||||
/** Restore a revision: write its body+metadata back into the live `documents` row. */
|
||||
export async function restoreRevision(
|
||||
id: string,
|
||||
): Promise<{ slug: string; documentId: string } | null> {
|
||||
const [rev] = await db
|
||||
.select({
|
||||
id: documentRevisions.id,
|
||||
documentId: documentRevisions.documentId,
|
||||
slug: documentRevisions.slug,
|
||||
title: documentRevisions.title,
|
||||
body: documentRevisions.body,
|
||||
meta: documentRevisions.meta,
|
||||
})
|
||||
.from(documentRevisions)
|
||||
.where(eq(documentRevisions.id, id));
|
||||
if (!rev) return null;
|
||||
|
||||
const m = rev.meta as {
|
||||
type?: string;
|
||||
section?: string;
|
||||
status?: string;
|
||||
author?: string;
|
||||
tags?: string[];
|
||||
};
|
||||
|
||||
await db
|
||||
.update(documents)
|
||||
.set({
|
||||
title: rev.title,
|
||||
type: (m.type as any) ?? "documentation",
|
||||
section: (m.section as any) ?? "docs",
|
||||
status: (m.status as any) ?? "published",
|
||||
author: m.author ?? "",
|
||||
tags: m.tags ?? [],
|
||||
body: rev.body,
|
||||
updatedAt: new Date(),
|
||||
})
|
||||
.where(eq(documents.id, rev.documentId));
|
||||
|
||||
return { slug: rev.slug, documentId: rev.documentId };
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
CREATE TABLE "document_revisions" (
|
||||
"id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL,
|
||||
"document_id" text NOT NULL,
|
||||
"slug" text NOT NULL,
|
||||
"revision_no" integer NOT NULL,
|
||||
"title" text NOT NULL,
|
||||
"body" text NOT NULL,
|
||||
"meta" jsonb NOT NULL,
|
||||
"reason" text DEFAULT 'index' NOT NULL,
|
||||
"created_at" timestamp with time zone DEFAULT now() NOT NULL
|
||||
);
|
||||
--> statement-breakpoint
|
||||
CREATE INDEX "document_revisions_document_id_idx" ON "document_revisions" USING btree ("document_id");
|
||||
--> statement-breakpoint
|
||||
CREATE INDEX "document_revisions_slug_idx" ON "document_revisions" USING btree ("slug");
|
||||
--> statement-breakpoint
|
||||
CREATE INDEX "document_revisions_doc_rev_idx" ON "document_revisions" USING btree ("document_id", "revision_no" DESC);
|
||||
--> statement-breakpoint
|
||||
ALTER TABLE "document_revisions" ADD CONSTRAINT "document_revisions_document_id_documents_id_fk" FOREIGN KEY ("document_id") REFERENCES "public"."documents"("id") ON DELETE cascade;
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"id": "0002_document_revisions",
|
||||
"prevId": "0001_document_chunks",
|
||||
"version": "7",
|
||||
"dialect": "postgresql",
|
||||
"tables": {
|
||||
"document_revisions": {
|
||||
"name": "document_revisions",
|
||||
"columns": {
|
||||
"id": {
|
||||
"name": "id",
|
||||
"type": "uuid",
|
||||
"primaryKey": true,
|
||||
"notNull": true,
|
||||
"default": "gen_random_uuid()"
|
||||
},
|
||||
"document_id": {
|
||||
"name": "document_id",
|
||||
"type": "text",
|
||||
"notNull": true
|
||||
},
|
||||
"slug": {
|
||||
"name": "slug",
|
||||
"type": "text",
|
||||
"notNull": true
|
||||
},
|
||||
"revision_no": {
|
||||
"name": "revision_no",
|
||||
"type": "integer",
|
||||
"notNull": true
|
||||
},
|
||||
"title": {
|
||||
"name": "title",
|
||||
"type": "text",
|
||||
"notNull": true
|
||||
},
|
||||
"body": {
|
||||
"name": "body",
|
||||
"type": "text",
|
||||
"notNull": true
|
||||
},
|
||||
"meta": {
|
||||
"name": "meta",
|
||||
"type": "jsonb",
|
||||
"notNull": true
|
||||
},
|
||||
"reason": {
|
||||
"name": "reason",
|
||||
"type": "text",
|
||||
"notNull": true,
|
||||
"default": "'index'"
|
||||
},
|
||||
"created_at": {
|
||||
"name": "created_at",
|
||||
"type": "timestamp",
|
||||
"notNull": true,
|
||||
"default": "now()"
|
||||
}
|
||||
},
|
||||
"indexes": {
|
||||
"document_revisions_document_id_idx": {
|
||||
"name": "document_revisions_document_id_idx",
|
||||
"columns": [
|
||||
{ "name": "document_id", "asc": true }
|
||||
],
|
||||
"isUnique": false
|
||||
},
|
||||
"document_revisions_slug_idx": {
|
||||
"name": "document_revisions_slug_idx",
|
||||
"columns": [
|
||||
{ "name": "slug", "asc": true }
|
||||
],
|
||||
"isUnique": false
|
||||
},
|
||||
"document_revisions_doc_rev_idx": {
|
||||
"name": "document_revisions_doc_rev_idx",
|
||||
"columns": [
|
||||
{ "name": "document_id", "asc": true },
|
||||
{ "name": "revision_no", "asc": false }
|
||||
],
|
||||
"isUnique": false
|
||||
}
|
||||
},
|
||||
"foreignKeys": {
|
||||
"document_revisions_document_id_documents_id_fk": {
|
||||
"name": "document_revisions_document_id_documents_id_fk",
|
||||
"columns": ["document_id"],
|
||||
"referenceTable": "documents",
|
||||
"referenceColumns": ["id"],
|
||||
"onDelete": "cascade"
|
||||
}
|
||||
},
|
||||
"compositePrimaryKeys": {},
|
||||
"uniqueConstraints": {},
|
||||
"policies": {}
|
||||
}
|
||||
},
|
||||
"enums": {},
|
||||
"schemas": {},
|
||||
"sequences": {},
|
||||
"roles": {},
|
||||
"policies": {},
|
||||
"views": {},
|
||||
"extensions": {},
|
||||
"_meta": {
|
||||
"columns": {},
|
||||
"schemas": {},
|
||||
"tables": {}
|
||||
}
|
||||
}
|
||||
@@ -15,6 +15,13 @@
|
||||
"when": 1787137149735,
|
||||
"tag": "0001_document_chunks",
|
||||
"breakpoints": true
|
||||
},
|
||||
{
|
||||
"idx": 2,
|
||||
"version": "7",
|
||||
"when": 1787139150000,
|
||||
"tag": "0002_document_revisions",
|
||||
"breakpoints": true
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -3,6 +3,7 @@ import {
|
||||
customType,
|
||||
index,
|
||||
integer,
|
||||
jsonb,
|
||||
pgTable,
|
||||
real,
|
||||
text,
|
||||
@@ -77,5 +78,48 @@ export const documentChunks = pgTable(
|
||||
export type DocumentChunkRow = typeof documentChunks.$inferSelect;
|
||||
export type NewDocumentChunkRow = typeof documentChunks.$inferInsert;
|
||||
|
||||
// Phase 3: document revision system. Each row is an immutable snapshot of a
|
||||
// document's body + metadata at a point in time (taken by the indexer whenever
|
||||
// the body actually changes). revisionNo is per-document and monotonically
|
||||
// increasing so the latest revision is always max(revision_no).
|
||||
export const documentRevisions = pgTable(
|
||||
"document_revisions",
|
||||
{
|
||||
id: uuid("id").primaryKey().defaultRandom(),
|
||||
documentId: text("document_id")
|
||||
.notNull()
|
||||
.references(() => documents.id, { onDelete: "cascade" }),
|
||||
slug: text("slug").notNull(),
|
||||
revisionNo: integer("revision_no").notNull(),
|
||||
title: text("title").notNull(),
|
||||
body: text("body").notNull(),
|
||||
// Metadata snapshot (type/section/status/author/tags) as JSON so a revision
|
||||
// is self-describing even if the live document is later restructured.
|
||||
meta: jsonb("meta").notNull().$type<{
|
||||
type: string;
|
||||
section: string;
|
||||
status: string;
|
||||
author: string;
|
||||
tags: string[];
|
||||
}>(),
|
||||
// Why this revision was created (e.g. "index", "git-push", "restore").
|
||||
reason: text("reason").notNull().default("index"),
|
||||
createdAt: timestamp("created_at", { withTimezone: true })
|
||||
.notNull()
|
||||
.defaultNow(),
|
||||
},
|
||||
(t) => ({
|
||||
docIdx: index("document_revisions_document_id_idx").on(t.documentId),
|
||||
slugIdx: index("document_revisions_slug_idx").on(t.slug),
|
||||
docRevIdx: index("document_revisions_doc_rev_idx").on(
|
||||
t.documentId,
|
||||
sql`${t.revisionNo} desc`,
|
||||
),
|
||||
}),
|
||||
);
|
||||
|
||||
export type DocumentRevisionRow = typeof documentRevisions.$inferSelect;
|
||||
export type NewDocumentRevisionRow = typeof documentRevisions.$inferInsert;
|
||||
|
||||
export type DocumentRow = typeof documents.$inferSelect;
|
||||
export type NewDocumentRow = typeof documents.$inferInsert;
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
{
|
||||
"name": "@mcpedia/queue",
|
||||
"version": "0.1.0",
|
||||
"private": true,
|
||||
"type": "module",
|
||||
"exports": {
|
||||
".": "./src/index.ts",
|
||||
"./client": "./src/client.ts",
|
||||
"./queue": "./src/queue.ts",
|
||||
"./worker": "./src/worker.ts"
|
||||
},
|
||||
"dependencies": {
|
||||
"@mcpedia/config": "workspace:*",
|
||||
"@mcpedia/core": "workspace:*",
|
||||
"@mcpedia/db": "workspace:*",
|
||||
"bullmq": "^6.1.2",
|
||||
"ioredis": "^6.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"typescript": "^5.6.0"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,34 @@
|
||||
import { REDIS_URL, REDIS_PASSWORD, QUEUE_PREFIX } from "@mcpedia/config";
|
||||
import IORedis, { type RedisOptions } from "ioredis";
|
||||
|
||||
/**
|
||||
* Shared ioredis connection for BullMQ. BullMQ requires an ioredis instance and
|
||||
* internally duplicates it for blocking commands, so we keep the option objects
|
||||
* explicit (maxRetriesPerRequest: null is REQUIRED for the blocking
|
||||
* connection — a finite retry count causes "Connection in key mode" errors).
|
||||
*/
|
||||
function buildOptions(): RedisOptions {
|
||||
const opts: RedisOptions = {
|
||||
maxRetriesPerRequest: null,
|
||||
lazyConnect: true,
|
||||
enableOfflineQueue: true,
|
||||
};
|
||||
if (REDIS_PASSWORD) opts.password = REDIS_PASSWORD;
|
||||
return opts;
|
||||
}
|
||||
|
||||
let _connection: IORedis | null = null;
|
||||
|
||||
/** Lazily-created singleton ioredis connection. */
|
||||
export function getConnection(): IORedis {
|
||||
if (!_connection) {
|
||||
_connection = new IORedis(REDIS_URL, buildOptions());
|
||||
_connection.on("error", (err) => {
|
||||
// Log but don't crash the process on transient Redis errors.
|
||||
console.error("[queue] redis error:", err.message);
|
||||
});
|
||||
}
|
||||
return _connection;
|
||||
}
|
||||
|
||||
export const BULLMQ_PREFIX = QUEUE_PREFIX;
|
||||
@@ -0,0 +1,9 @@
|
||||
export { getConnection, BULLMQ_PREFIX } from "./client";
|
||||
export {
|
||||
getQueue,
|
||||
enqueueIndexDoc,
|
||||
enqueueFullIndex,
|
||||
INDEX_QUEUE,
|
||||
} from "./queue";
|
||||
export type { IndexDocJobData, IndexAllJobData, JobType } from "./queue";
|
||||
export { createWorker, startWorker } from "./worker";
|
||||
@@ -0,0 +1,66 @@
|
||||
import { Queue, type Job } from "bullmq";
|
||||
import { getConnection, BULLMQ_PREFIX } from "./client";
|
||||
|
||||
export const INDEX_QUEUE = "mcpedia-index";
|
||||
|
||||
/** Lazily-created singleton BullMQ queue. */
|
||||
let _queue: Queue | null = null;
|
||||
|
||||
export function getQueue(): Queue {
|
||||
if (!_queue) {
|
||||
_queue = new Queue(INDEX_QUEUE, {
|
||||
connection: getConnection(),
|
||||
prefix: BULLMQ_PREFIX,
|
||||
});
|
||||
}
|
||||
return _queue;
|
||||
}
|
||||
|
||||
export interface IndexDocJobData {
|
||||
relPath: string;
|
||||
reason: string;
|
||||
}
|
||||
|
||||
export interface IndexAllJobData {
|
||||
reason: string;
|
||||
}
|
||||
|
||||
export type JobType = "index-doc" | "index-all";
|
||||
|
||||
/**
|
||||
* Enqueue a single-document reindex job. Keyed by slug so repeated edits
|
||||
* collapse into one pending job (BullMQ dedup by jobId within the window).
|
||||
*/
|
||||
export async function enqueueIndexDoc(
|
||||
relPath: string,
|
||||
reason = "index",
|
||||
): Promise<Job<IndexDocJobData>> {
|
||||
const slug = relPath.replace(/\.mdx?$/, "");
|
||||
return getQueue().add(
|
||||
"index-doc",
|
||||
{ relPath, reason },
|
||||
{
|
||||
jobId: `doc__${slug}`,
|
||||
removeOnComplete: 1000,
|
||||
removeOnFail: 5000,
|
||||
attempts: 3,
|
||||
backoff: { type: "exponential", delay: 2000 },
|
||||
},
|
||||
);
|
||||
}
|
||||
|
||||
/** Enqueue a full-corpus reindex (used by the git-sync hook). */
|
||||
export async function enqueueFullIndex(
|
||||
reason = "reindex",
|
||||
): Promise<Job<IndexAllJobData>> {
|
||||
return getQueue().add(
|
||||
"index-all",
|
||||
{ reason },
|
||||
{
|
||||
jobId: `full__${Date.now()}`,
|
||||
removeOnComplete: 100,
|
||||
removeOnFail: 1000,
|
||||
attempts: 1,
|
||||
},
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
import { Worker, type Job } from "bullmq";
|
||||
import { getConnection, BULLMQ_PREFIX } from "./client";
|
||||
import { INDEX_QUEUE } from "./queue";
|
||||
import { indexContentFile, runFullIndex } from "@mcpedia/core";
|
||||
|
||||
export function createWorker(): Worker {
|
||||
const worker = new Worker(
|
||||
INDEX_QUEUE,
|
||||
async (job: Job) => {
|
||||
switch (job.name) {
|
||||
case "index-doc": {
|
||||
const { relPath, reason } = job.data as {
|
||||
relPath: string;
|
||||
reason: string;
|
||||
};
|
||||
await job.log(`indexing ${relPath}`);
|
||||
const r = await indexContentFile(relPath, reason);
|
||||
return r;
|
||||
}
|
||||
case "index-all": {
|
||||
const { reason } = job.data as { reason: string };
|
||||
await job.log(`full index (${reason})`);
|
||||
return await runFullIndex(reason);
|
||||
}
|
||||
default:
|
||||
throw new Error(`unknown job type: ${job.name}`);
|
||||
}
|
||||
},
|
||||
{
|
||||
connection: getConnection(),
|
||||
prefix: BULLMQ_PREFIX,
|
||||
concurrency: 4,
|
||||
},
|
||||
);
|
||||
|
||||
worker.on("completed", (job) => {
|
||||
console.log(`[worker] completed ${job.name} (${job.id})`);
|
||||
});
|
||||
worker.on("failed", (job, err) => {
|
||||
console.error(`[worker] failed ${job?.name} (${job?.id}): ${err.message}`);
|
||||
});
|
||||
worker.on("error", (err) => {
|
||||
console.error(`[worker] error:`, err.message);
|
||||
});
|
||||
|
||||
return worker;
|
||||
}
|
||||
|
||||
/** Start the worker and wire graceful shutdown. */
|
||||
export async function startWorker(): Promise<Worker> {
|
||||
const worker = createWorker();
|
||||
console.log("[worker] indexing worker started");
|
||||
|
||||
const shutdown = async (sig: string) => {
|
||||
console.log(`[worker] ${sig} received, closing...`);
|
||||
await worker.close();
|
||||
process.exit(0);
|
||||
};
|
||||
process.on("SIGINT", () => void shutdown("SIGINT"));
|
||||
process.on("SIGTERM", () => void shutdown("SIGTERM"));
|
||||
|
||||
return worker;
|
||||
}
|
||||
@@ -0,0 +1,24 @@
|
||||
import { enqueueIndexDoc, enqueueFullIndex } from "@mcpedia/queue";
|
||||
|
||||
// One-shot enqueue helper (no worker required to schedule work).
|
||||
// Usage:
|
||||
// bun run enqueue --all # full reindex
|
||||
// bun run enqueue docs/websocket/contract # single doc (slug or rel path)
|
||||
async function main() {
|
||||
const arg = process.argv[2];
|
||||
if (!arg || arg === "--all") {
|
||||
const job = await enqueueFullIndex("manual");
|
||||
console.log(`enqueued full reindex job ${job.id}`);
|
||||
} else {
|
||||
const relPath = arg.endsWith(".md") || arg.endsWith(".mdx") ? arg : `${arg}.md`;
|
||||
const job = await enqueueIndexDoc(relPath, "manual");
|
||||
console.log(`enqueued doc reindex job ${job.id} -> ${relPath}`);
|
||||
}
|
||||
await new Promise((r) => setTimeout(r, 500)); // allow the event loop to flush
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
main().catch((err) => {
|
||||
console.error(err);
|
||||
process.exit(1);
|
||||
});
|
||||
+9
-62
@@ -1,67 +1,14 @@
|
||||
import { db } from "@mcpedia/db";
|
||||
import { documents } from "@mcpedia/db/schema";
|
||||
import { parseFile } from "@mcpedia/parser";
|
||||
import { CONTENT_ROOT } from "@mcpedia/config";
|
||||
import { listContentFiles, indexChunks } from "@mcpedia/core";
|
||||
import { join } from "node:path";
|
||||
import { runFullIndex } from "@mcpedia/core";
|
||||
|
||||
// Phase 3: the indexer now goes through `runFullIndex`, the single indexing
|
||||
// entry point shared with the BullMQ worker and the git-sync hook. It parses
|
||||
// each content file, upserts `documents`, chunks+embeds, and snapshots a
|
||||
// revision when the body changed.
|
||||
async function main() {
|
||||
const files = listContentFiles();
|
||||
let indexed = 0;
|
||||
let chunked = 0;
|
||||
for (const rel of files) {
|
||||
const abs = join(CONTENT_ROOT, rel);
|
||||
const { meta, body } = parseFile(abs, rel);
|
||||
const nowIso =
|
||||
meta.updatedAt && meta.updatedAt !== ""
|
||||
? meta.updatedAt
|
||||
: new Date().toISOString();
|
||||
await db
|
||||
.insert(documents)
|
||||
.values({
|
||||
id: meta.id,
|
||||
slug: meta.slug,
|
||||
title: meta.title,
|
||||
type: meta.type,
|
||||
section: meta.section,
|
||||
status: meta.status,
|
||||
author: meta.author,
|
||||
tags: meta.tags,
|
||||
path: meta.path,
|
||||
body,
|
||||
createdAt: new Date(meta.createdAt || nowIso),
|
||||
updatedAt: new Date(nowIso),
|
||||
})
|
||||
.onConflictDoUpdate({
|
||||
target: documents.slug,
|
||||
set: {
|
||||
title: meta.title,
|
||||
type: meta.type,
|
||||
section: meta.section,
|
||||
status: meta.status,
|
||||
author: meta.author,
|
||||
tags: meta.tags,
|
||||
path: meta.path,
|
||||
body,
|
||||
updatedAt: new Date(nowIso),
|
||||
},
|
||||
});
|
||||
indexed++;
|
||||
console.log(` indexed ${rel}`);
|
||||
|
||||
// Phase 2: chunk + embed for semantic search.
|
||||
try {
|
||||
const n = await indexChunks(meta.slug, body);
|
||||
chunked += n;
|
||||
console.log(` embedded ${n} chunks`);
|
||||
} catch (err) {
|
||||
console.error(
|
||||
` embed FAILED for ${meta.slug}: ${err instanceof Error ? err.message : err}`,
|
||||
);
|
||||
// Don't abort the whole index over one doc's embedding failure.
|
||||
}
|
||||
}
|
||||
console.log(`indexed ${indexed} documents, ${chunked} chunks embedded`);
|
||||
const reason = process.argv[2] && process.argv[2].startsWith("--reason=")
|
||||
? process.argv[2].slice("--reason=".length)
|
||||
: "index";
|
||||
await runFullIndex(reason);
|
||||
}
|
||||
|
||||
main()
|
||||
|
||||
@@ -10,6 +10,7 @@
|
||||
"@mcpedia/embeddings": "workspace:*",
|
||||
"@mcpedia/parser": "workspace:*",
|
||||
"@mcpedia/search": "workspace:*",
|
||||
"@mcpedia/queue": "workspace:*",
|
||||
"drizzle-orm": "^0.38.0",
|
||||
"postgres": "^3.4.5"
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user