Reinterpreted from the plan's YAGNI 'Scale-out' (OpenSearch/object-storage /multi-tenant deferred at KB scale). Phase 4 = make the Phase 3 async + revision system correct, secure, observable, deployable. - T1 (correctness bug): restoreRevision now rebuilds semantic chunks via new @mcpedia/core reindexChunks(slug) so semantic/hybrid search stay consistent after a restore (previously document_chunks held the NEW body while documents.body held the restored OLD body -> stale search). - T2 (security): /hooks/* git-sync webhooks now require x-webhook-secret header matching WEBHOOK_SECRET (401 otherwise); API fails fast at startup if unset. Added WEBHOOK_SECRET to @mcpedia/config + .env.example; set real secret in .env. - T3 (UX): web doc page shows a History panel (revision no/reason/date/length) with per-revision Restore; app/api/revisions/restore/route.ts calls restoreRevision + revalidatePath (server-component only, no client JS). - T4: listRevisions gains offset paging; summary never includes body. - T5 (ops): deploy/mcpedia-api.service + deploy/mcpedia-worker.service systemd units (Restart=on-failure, EnvironmentFile=.env). Not auto-enabled on host. Verified against live imrnes Redis + Postgres: turbo typecheck+build green; restore-rebuilds-chunks (marker present -> gone after restore); webhook 401/200; web restore route redirects to doc + reverts body; revisions API returns summary (no body); systemd-analyze verify passes.
187 lines
7.4 KiB
Markdown
187 lines
7.4 KiB
Markdown
# MCPedia
|
||
|
||
> A content-first knowledge base — readable as Markdown/MDX in Git, queryable by
|
||
> humans via a Web UI and by AI agents via the Model Context Protocol (MCP).
|
||
|
||
MCPedia keeps content as plain Markdown files under `content/`. A Git-tracked
|
||
source of truth, indexed into PostgreSQL (metadata + a `tsvector` full-text
|
||
column) and served through a single **Core** layer that every interface
|
||
(Web, MCP) shares — no business logic duplicated per surface.
|
||
|
||
## Monorepo layout
|
||
|
||
```
|
||
mcpedia/
|
||
├── apps/
|
||
│ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search
|
||
│ ├── mcp/ # MCP server (stdio) — AI-agent interface (tools + resources)
|
||
│ └── api/ # Hono + tRPC v11 API on :4020 (+ /hooks/* git-sync webhooks)
|
||
├── packages/
|
||
│ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...)
|
||
│ ├── config/ # loads .env (repo root) as authoritative dev config
|
||
│ ├── db/ # Drizzle ORM schema + client + drizzle-kit config
|
||
│ ├── parser/ # frontmatter (gray-matter) parsing
|
||
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline)
|
||
│ ├── embeddings/ # embedding provider + chunker
|
||
│ ├── queue/ # Redis (ioredis) + BullMQ worker/queue (Phase 3)
|
||
│ └── core/ # Document/Content/Search/Index/Revision — the only business logic
|
||
├── content/ # docs/ writeups/ research/ notes/ (the knowledge base)
|
||
└── scripts/ # indexer.ts (full reindex), enqueue.ts (one-shot job enqueue)
|
||
```
|
||
|
||
## Architecture principle
|
||
|
||
```
|
||
Web ─┐
|
||
├──► Core ──► Repository (@mcpedia/db) ──► PostgreSQL
|
||
MCP ─┘
|
||
```
|
||
|
||
All interfaces go through `@mcpedia/core`. Nothing outside `packages/db` and
|
||
`packages/core` touches the database directly.
|
||
|
||
## Quick start
|
||
|
||
```bash
|
||
bun install # install workspace deps
|
||
cp .env.example .env # set DATABASE_URL (dev uses imrnes Postgres :6432)
|
||
bunx turbo run build # typecheck + build every package
|
||
|
||
bun run index # walk content/ -> upsert into Postgres
|
||
bun --cwd apps/web run dev # Web UI on :3000
|
||
bun run mcp # MCP server on stdio (pipe to an MCP client)
|
||
```
|
||
|
||
### Database
|
||
|
||
Schema is defined in `packages/db/src/schema.ts` (`documents` with a weighted
|
||
`search_vector` tsvector + GIN index, and `document_chunks` with an `embedding real[]`).
|
||
The `pgvector` extension is **not available** on the shared imrnes Postgres, so
|
||
semantic search stores vectors as `real[]` and ranks by in-app cosine similarity.
|
||
|
||
Migrations live in `packages/db/drizzle/`. They were applied manually via `psql`
|
||
(`drizzle-kit push` is unreliable under PgBouncer transaction pooling); to
|
||
re-apply on a fresh DB:
|
||
|
||
```bash
|
||
psql $DATABASE_URL -f packages/db/drizzle/0000_grey_toro.sql
|
||
psql $DATABASE_URL -f packages/db/drizzle/0001_document_chunks.sql
|
||
```
|
||
|
||
> Note: on imrnes (PgBouncer `:6432`) a leaked `DATABASE_URL` shell var can
|
||
> shadow `.env`. `@mcpedia/config` loads `.env` **last** so the repo config
|
||
> always wins for local/dev.
|
||
|
||
## Content
|
||
|
||
Each Markdown file carries YAML frontmatter:
|
||
|
||
```yaml
|
||
---
|
||
id: websocket-contract
|
||
title: WebSocket Contract
|
||
type: documentation
|
||
tags: [typescript, websocket, rpc]
|
||
status: published
|
||
author: asep
|
||
created_at: 2026-08-19
|
||
updated_at: 2026-08-19
|
||
---
|
||
```
|
||
|
||
`slug` = relative path under `content/` (e.g. `docs/websocket/contract`). The
|
||
`body` shown in the UI is always read from the on-disk file (source of truth);
|
||
the DB stores metadata + the search vector.
|
||
|
||
## MCP tools
|
||
|
||
| Tool | Purpose |
|
||
| --------------------- | ------------------------------------------------ |
|
||
| `search_documents` | Postgres FTS over the corpus (ranked + snippet) |
|
||
| `semantic_search` | Embedding/cosine search over chunked content |
|
||
| `hybrid_search` | FTS + semantic fused via RRF |
|
||
| `get_document` | Full markdown body by slug |
|
||
| `list_documents` | List, optionally filtered by section |
|
||
| `get_related_documents` | Docs sharing tags with a given slug |
|
||
|
||
### MCP Resources
|
||
|
||
| URI | Purpose |
|
||
| -------------------------------- | ---------------------------------------- |
|
||
| `mcpedia://docs` | List all published documents |
|
||
| `mcpedia://docs/{+slug}` | Full markdown body (read from disk) |
|
||
| `mcpedia://docs/{+slug}/chunks` | Preview of embedded semantic chunks |
|
||
| `mcpedia://docs/{+slug}/revisions` | Revision history summary |
|
||
|
||
(`{+slug}` uses RFC 6570 reserved expansion so a slug like
|
||
`docs/websocket/contract` matches the template.)
|
||
|
||
Smoke test (in-memory transport, real JSON-RPC):
|
||
|
||
```bash
|
||
bun --cwd apps/mcp run smoke
|
||
```
|
||
|
||
## API (Phase 2 + Phase 3)
|
||
|
||
A tRPC v11 API is exposed via Hono on **:4020** (all procedures mirror the
|
||
MCP tools). Phase 3 adds async job + revision procedures and git-sync webhooks:
|
||
|
||
```bash
|
||
bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*)
|
||
```
|
||
|
||
tRPC procedures: `search`, `semanticSearch`, `hybridSearch`, `getDocument`,
|
||
`listDocuments`, `related` (Phase 2); plus `revisions`, `getRevision`,
|
||
`restoreRevision`, `jobStatus`, `queueStatus` (Phase 3).
|
||
|
||
Git-sync webhooks (enqueue BullMQ jobs; the worker processes them):
|
||
- `POST /hooks/reindex` — full-corpus reindex (point your Git provider's
|
||
push webhook here to auto-reindex on push).
|
||
- `POST /hooks/index?slug=<slug>` — reindex a single document.
|
||
|
||
> **Security:** both webhooks require an `x-webhook-secret` header that matches
|
||
> `WEBHOOK_SECRET` (set in `.env`). The API refuses to start if `WEBHOOK_SECRET`
|
||
> is unset, so the hooks are never left open.
|
||
|
||
`bun run index` now also chunks + embeds (Phase 2 indexer) and snapshots a
|
||
revision whenever the body changes (Phase 3). See `.env.example` for
|
||
`EMBED_*` / `REDIS_*` / `QUEUE_PREFIX` / `WEBHOOK_SECRET` vars.
|
||
|
||
### Run as a supervised service (Phase 4)
|
||
|
||
`deploy/mcpedia-api.service` + `deploy/mcpedia-worker.service` are systemd units
|
||
(`Restart=on-failure`, `EnvironmentFile=.env`, `WorkingDirectory=/home/code/mcpedia`).
|
||
Enable them with:
|
||
|
||
```bash
|
||
sudo cp deploy/*.service /etc/systemd/system/
|
||
sudo systemctl daemon-reload
|
||
sudo systemctl enable --now mcpedia-api mcpedia-worker
|
||
# tail logs
|
||
journalctl -u mcpedia-api -u mcpedia-worker -f
|
||
```
|
||
|
||
The API should sit behind Caddy (or your reverse proxy) for TLS; expose only
|
||
`:4020` internally and the web app publicly.
|
||
|
||
## Status
|
||
|
||
**Phase 1 — MVP (DONE):** monorepo, Core, Web UI (home/doc/search), MCP server,
|
||
Postgres FTS keyword search, content indexing.
|
||
|
||
**Phase 2 — Semantic + API (DONE):** embeddings provider (OpenRouter via 9router),
|
||
chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API
|
||
(`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle.
|
||
|
||
**Phase 3 — Async + Scale (DONE):** Redis + BullMQ background indexing/embedding
|
||
workers (`packages/queue`, `apps/worker`), git-sync webhooks (`POST /hooks/*`),
|
||
document revision system (`document_revisions` + restore), and MCP Resources
|
||
(`mcpedia://docs/...`). See `PHASES.md`.
|
||
|
||
> pgvector is **not installed** on the shared imrnes Postgres, so vector storage is
|
||
> a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is
|
||
> the Phase-4 scale-out path. See `PHASES.md`.
|
||
|
||
See `PHASES.md` for Phase 3–4 (Redis/BullMQ, auth, revisions, scale-out).
|