feat: revamp to dynamic database-first architecture and cleanup phase docs
CI / typecheck + build (turbo) (push) Canceled after 0s

- Migrate document fetching and CRUD to be PostgreSQL-authoritative
- Remove static section enums and add dynamic listSections query
- Support custom sections and metadata across API, MCP, and Web UI
- Add /api/sections endpoint and update Header, Sidebar, and Forms
- Remove obsolete phase planning docs and modernize README/AGENTS
This commit is contained in:
asepharyana
2026-08-21 11:35:42 +07:00
parent a937e8f51b
commit a9b67385c3
37 changed files with 720 additions and 2515 deletions
File diff suppressed because one or more lines are too long
@@ -1,161 +0,0 @@
# MCPedia Phase 2 — Semantic Search + tRPC/Hono API
> **For Hermes:** implement task-by-task. Spec-first (user rule 2026-08-19).
**Goal:** Add semantic + hybrid search (pgvector) and a typed tRPC/Hono API so
MCPedia is queryable by embeddings, not just keyword FTS — and expose the
corpus over a programmatic HTTP API.
**Architecture:** Content (Markdown) → chunk → embed (OpenRouter) → store
`document_chunks` with `vector(N)` in Postgres → `semanticSearch` (cosine) and
`hybridSearch` (FTS + cosine, reciprocal-rank fusion) in `@mcpedia/search` →
exposed via Core, the MCP server (new tools), and a new `apps/api` (Hono +
tRPC v11).
**Embedding provider:** OpenRouter (`openrouter/llama-nemotron-embed-vl-1b-v2:free`)
via `9router_ai_llm_api_key` + `9router_ai_llm_base_url` (BWS). Dimension is
discovered at first live call (see Step 1.3) and pinned in schema/migration.
**Tech stack:** drizzle-orm `vector` column + pgvector extension, HNSW index,
`@trpc/server` v11 (fetch adapter), `hono` + `@hono/node-server`.
---
## Task P2.1 — `packages/embeddings` (provider + abstraction)
**Files:** `packages/embeddings/package.json`, `src/index.ts`, `src/provider.ts`,
`src/openrouter.ts`
- `EmbeddingProvider` interface: `embed(texts: string[]): Promise<number[][]>`, `readonly model`, `readonly dimensions`.
- `OpenRouterEmbeddingProvider`: POST `${baseUrl}/embeddings` with `{ model, input }`,
`Authorization: Bearer ${key}`. Returns `data[].embedding`. Validate length === dimensions.
- Read `EMBED_BASE_URL`, `EMBED_API_KEY`, `EMBED_MODEL` from `@mcpedia/config`
(with `.env` fallback). Dimensions discovered live (Step 1.3) → export `EMBED_DIM`.
- Chunk helper `chunkText(text, { size=1000, overlap=150 })` in `src/chunk.ts`.
**Step 1.3 (discover dim):** live call `embed(["test"])`, read `embedding.length`,
pin `EMBED_DIM`, assert mismatch throws.
**Verify:** `bun run` a temp script: `embed(["hello world"])` prints a vector of
length N (e.g. 1024). Confirm no key is logged.
---
## Task P2.2 — Schema: `document_chunks` + vector extension
**Files:** `packages/db/src/schema.ts` (add), `packages/db/drizzle.config.ts`
(unchanged), new migration.
- `CREATE EXTENSION IF NOT EXISTS vector;` (idempotent; run once via psql).
- `document_chunks` table:
- `id` uuid pk default gen_random_uuid()
- `document_id` text → `documents.id` on delete cascade
- `slug` text (denormalized for convenience)
- `chunk_index` integer
- `content` text
- `embedding` vector(EMBED_DIM)
- `created_at` timestamp default now()
- index `chunk_embedding_idx` using hnsw (`embedding` op `vector_cosine_ops`)
- Generate migration with `drizzle-kit generate`, apply via `psql` (drizzle-kit
push is unreliable here — known).
**Verify:** `\d document_chunks` shows `embedding vector(N)` + HNSW index;
`select count(*) from document_chunks` = 0.
---
## Task P2.3 — Indexer: chunk + embed + upsert
**Files:** `scripts/indexer.ts` (extend), `packages/core/src/document.service.ts`
(add `indexChunks`).
- For each published doc: read body (already on disk), `chunkText`, `embed` in
batches (≤ 16), delete existing chunks for slug, insert new rows.
- Guard: if embedding provider fails, log + skip (don't crash the whole index).
- Add `bun run index:embed` (or extend `bun run index` to also embed).
**Verify:** after running, `select count(*) from document_chunks` > 0; a sample
row has non-null `embedding`.
---
## Task P2.4 — `packages/search`: semantic + hybrid
**Files:** `packages/search/src/index.ts` (add `semanticSearch`, `hybridSearch`).
- `semanticSearch(vec, limit)`: order by `embedding <=> ${vec}` asc, filter published.
- `hybridSearch(q, limit)`: run FTS (`ts_rank`) + semantic (cosine) in parallel;
fuse with reciprocal-rank (RRF: score = 1/(k+rank), k=60); return merged hits.
- Keep `keywordSearch` unchanged (Phase 1).
**Verify:** unit-ish script: embed a query, `semanticSearch` returns relevant
chunks; `hybridSearch("websocket")` returns ≥ keyword results.
---
## Task P2.5 — `packages/core` expose semantic/hybrid
**Files:** `packages/core/src/search.service.ts`, `index.ts`.
- Re-export `semanticSearch`, `hybridSearch` from Core.
---
## Task P2.6 — `apps/api` (Hono + tRPC v11)
**Files:** `apps/api/package.json`, `tsconfig.json`, `src/index.ts`,
`src/router.ts`, `src/trpc.ts`.
- `initTRPC.create()` router with procedures: `search`, `semanticSearch`,
`hybridSearch`, `getDocument`, `listDocuments` (mirrors MCP tools).
- Mount `fetchRequestHandler` on a Hono app at `/trpc/*`; serve via
`@hono/node-server` `serve({ fetch: app.fetch, port: 4020 })`.
- `createContext` returns `{ db }`.
**Verify:** `bun run dev` → `curl -X POST localhost:4020/trpc/search`
with JSON body returns hits.
---
## Task P2.7 — MCP server: semantic + hybrid tools
**Files:** `apps/mcp/src/index.ts` (add `semantic_search`, `hybrid_search`),
extend `smoke.test.ts`.
- `semantic_search`: embed query → `semanticSearch`.
- `hybrid_search`: embed query → `hybridSearch`.
- Smoke: assert both return ≥1 hit for "websocket".
---
## Task P2.8 — Web: semantic toggle on search
**Files:** `apps/web/app/search/page.tsx`.
- Add `mode=keyword|hybrid` query param; server component calls Core
`hybridSearch` when `mode=hybrid`. Minimal UI toggle (link/buttons).
- Keep keyword as default.
**Verify:** `bun run build`; `curl '/search?q=websocket&mode=hybrid'` returns hits.
---
## Task P2.9 — Verify all + commit
- `bunx turbo run build` (web + api + mcp), `bun run apps/mcp smoke`,
live API curl, live web hybrid search.
- Update `README.md` + `PHASES.md` (mark Phase 2 ✅).
- `git add -A` (exclude `.env`), commit as asepharyana (no Co-Authored-By).
---
## Risks / decisions
- **Dimension unknown until live call** → P2.1.3 discovers it; pinned EMBED_DIM=2048.
- **pgvector NOT available on shared imrnes Postgres** (extension not installed;
installing needs host-level apt on a managed/shared DB — deferred). PIVOT:
store `embedding` as `real[]` and compute cosine similarity in the app layer.
Brute-force cosine is instant for a KB-sized corpus (dozens of docs / hundreds
of chunks). pgvector+HNSW is the Phase-4 scale-out path.
- **PgBouncer + real[]**: fine; simple queries, no extension needed.
- **API port 4020** (host 4000s range is 4000–4015; 4020 is free for dev). Deploy later.
- **YAGNI**: no auth/revisions this phase (Phase 3).
-151
View File
@@ -1,151 +0,0 @@
# Phase 14 — Hierarchical Folder Structure
> User: "gk ada bedanya, maksud saya inginnya itu bisa yg bertingkat seperti github yg memiliki folder dalam folder"
> Context: after the full dynamic-custom-fields overhaul (Phase 13), the user
> wants document URLs/content organized in **nested folders** like GitHub —
> `writeups/ctf/defcon-quals-2024/pwn-100/...` with subfolders under subfolders,
> not just one level deep.
## Problem
The current URL scheme is `/<section>/<slug>` where `slug` can contain `/`
(e.g. `writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment` →
`/writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment`). This works for
**files** but there are no **folder-level index pages** — navigating to
`/writeups/ctf/defcon-quals-2024/` returns 404 because Next.js catch-all
`[section]/[...slug]/page.tsx` requires at least one slug segment beyond the
section, and the sidebar only shows flat doc titles (no folder tree).
GitHub's model: `github.com/org/repo/tree/main/path/to/folder/file` — every
folder has an index page (`/path/to/folder/`) listing its contents.
## Solution
### 1. Folder Index Pages
**Create `apps/web/app/[section]/[...slug]/folder.tsx`** (or a parallel route).
Actually — cleaner approach per Next.js App Router: the catch-all
`[section]/[...slug]/page.tsx` handles both. Add logic: if the slug resolves to
an actual markdown file → doc page (existing behavior). If the slug resolves to
a **directory** (folder of docs) → render a folder index listing all docs whose
`path` starts with that prefix.
**Mechanism:**
- Call `listDocuments()` to get all docs.
- The incoming URL path is `{section}/{...slug}`.
- Build the "folder prefix" = `${section}/${slug.join("/")}/` (with trailing `/`,
or just `${section}/${slug.join("/")}` if no slug segments).
- Filter docs whose `doc.path` starts with that prefix.
- If exactly one doc matches AND its path === prefix (trimmed .md) → it's a
doc page (existing). If zero or multiple match and they all start with the
prefix → it's a folder index.
- Edge: a folder with exactly one doc whose path matches exactly — still a doc
page. A folder is when there are docs at `prefix/sub/...`.
**Better heuristic:** A slug path is a "folder" if there exist docs whose `path`
is `prefix/deep/...` (i.e., the slug is a parent of other doc paths, not a
leaf itself). A slug is a "leaf doc" if `path === prefix + ".md"`.
### 2. Sidebar Tree
**Update `Sidebar.tsx`:**
- `listDocuments()` already returns all docs with their full `slug` and `path`.
- Build a **tree** from the flat list: split each slug by `/`, create nested
folder nodes.
- Render nested `<ul>` with indentation (already done via `marginLeft` based on
depth).
- **Folder nodes** (collapsed/expanded) get a folder icon 📁 and a CSS class.
- Clicking a folder → navigates to the folder index page `/{section}/{path}`.
- **Leaf doc nodes** → link to `/{doc.slug}` (existing behavior).
- Group by the first segment after section too (e.g. `ctf/defcon-quals-2024/`
is a folder, then `pwn-100-...` are children).
Tree-building algorithm (from flat slugs):
```
For slug "writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment":
parts = ["writeups", "ctf", "defcon-quals-2024", "pwn-100-ret2win-alignment"]
→ tree: writeups → ctf → defcon-quals-2024 → pwn-100-ret2win-alignment (leaf)
```
### 3. DocForm / Folder Selection
**Update `DocForm.tsx`:**
- Add a "Parent folder" input (autocomplete or text) that shows existing folders
for the selected section. The slug field already supports `/` but the user
experience is better with folder picker.
- When creating, the slug becomes `{parentFolder}/{slug}` automatically.
- Show existing folder structure as `<select>` or tree picker.
### 4. Example Hierarchy
Create a real hierarchical structure to demonstrate:
```
writeups/
ctf/
defcon-quals-2024/
pwn/
pwn-100-ret2win-alignment.md
pwn-200-bof-heap.md
crypto/
crypto-100-xor.md
crypto-200-rsa.md
template/
writeup-template.md
_index.md ← folder index (optional intro)
hackthebox/
machine-name/
walkthrough.md
```
For now, reorganize the existing Defcon writeup into proper subfolders + add a
folder index page. The existing `content/writeups/ctf/defcon-quals-2024/` is
already a folder — just need the folder index route to work.
### 5. Route Changes
**Current:** `[section]/[...slug]/page.tsx` — catch-all requires ≥1 slug segment.
- `/writeups` → `notFound()` (no index for bare section unless we add one)
- `/writeups/ctf` → catch-all gets `slug=["ctf"]` → currently treated as a doc
(looks up `writeups/ctf` doc, 404 if none)
- `/writeups/ctf/defcon-quals-2024/` → catch-all `slug=["ctf","defcon-quals-2024"]`
**Plan:**
1. Add `/writeups/page.tsx` (section index) — lists top-level folders + root
docs in that section. (Currently `/docs/page.tsx` exists but `/writeups/page.tsx`
doesn't.)
2. In `[section]/[...slug]/page.tsx`: at the top of the page component, check if
the slug path is a folder (has child docs). If so, render folder index instead
of doc page.
**Section index pages:** Create `[section]/page.tsx` for all 4 sections, or a
generic one. Currently only `/docs/page.tsx` exists. Add a shared
`SectionIndex` component.
### 6. Files to Change
```
new: apps/web/app/[section]/page.tsx # generic section index (folder + doc listing)
mod: apps/web/app/[section]/[...slug]/page.tsx # add folder-index detection
mod: apps/web/app/components/Sidebar.tsx # tree from flat slugs
mod: apps/web/app/components/DocForm.tsx # parent folder picker
new: content/writeups/ctf/defcon-quals-2024/_index.md # folder intro (optional)
new: content/writeups/ctf/_index.md # CTF section intro
mod: apps/web/app/docs/page.tsx # may need generic version
```
### 7. Verification
- `/writeups` → 200, shows folders: ctf/, template/
- `/writeups/ctf` → 200, folder index listing defcon-quals-2024/
- `/writeups/ctf/defcon-quals-2024` → 200, folder index listing pwn-100-...
- `/writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment` → 200, doc page
- Sidebar shows nested tree with folder icons
- DocForm parent folder picker works
- `bun run test` green, `turbo run typecheck` green
- Live verification via curl
## Constraints
- User: "pastikan semua dinamis dan rapih untuk banyak situasi jadi tergantung
user bukan hardcode" — folder detection must be content-driven, not config.
- No breaking existing flat slugs.
- Section icons/labels stay the same.
-42
View File
@@ -1,42 +0,0 @@
# MCP HTTP transport + deploy + review actions
## Goal
Make the MCPedia MCP server reachable over the network (not just stdio subprocess), so
remote MCP clients (Claude, a Discord bot, a web client) can call its 6 tools + 4 resources.
Serve via Streamable HTTP (MCP 2025-03-26 spec), deploy as a supervised systemd service,
expose through Caddy on a dedicated subdomain.
## Design decisions
- **StreamableHTTPServerTransport, stateless mode** (`sessionIdGenerator: undefined`).
One McpServer + transport per request. No session map, no shared-transport connect race,
no memory leak. Re-registering 6 tools + 4 resources per request is negligible for a KB.
- **New entry `apps/mcp/src/http.ts`** served by Node `http` (built-in), NOT mounted on the
API app — keeps the MCP app's zod-4 isolation intact (api app is zod 3).
- **Port 4021** (next free in the 4000s range; 4020 is the API).
- **Subdomain `mcp.asepharyana.my.id`** -> 4021 (Cloudflare `*` wildcard already proxies it;
Caddy auto-issues LE cert, no extra DNS work).
- **CORS** allow on `/mcp` (remote web clients need it).
- MCP tools are read-only (search/get/list/related + read-only resources) => open MCP is
low-risk. No auth on MCP itself.
## Files
- `apps/mcp/src/http.ts` (NEW) — Node http server, `/mcp` route, stateless transport.
- `apps/mcp/package.json` — add `serve:http` script.
- `package.json` (root) — add `mcp:http` script (absolute bun path).
- `deploy/mcpedia-mcp.service` (NEW) — systemd unit, MCP_PORT=4021.
- `/etc/caddy/Caddyfile` — add `mcp.asepharyana.my.id { import proxy 4021 }`.
## Security finding (review, flagged not silently built)
The tRPC `restoreRevision` mutation is exposed UNauthenticated at
`https://wiki.asepharyana.my.id/trpc/restoreRevision` — anyone can revert a live doc.
The web UI's restore path calls `@mcpedia/core` directly (server component), so the tRPC
mutation is dead surface. Fix: guard the mutation with the existing WEBHOOK_SECRET header,
or drop it from the router. Will apply the guard (consistent with /hooks auth) unless user
prefers removal.
## Verification
- `bun --cwd apps/mcp run typecheck` green.
- Live: `curl -XPOST https://mcp.asepharyana.my.id/mcp` initialize -> 200 + serverInfo;
tools/list -> 6 tools; resources/list -> 4 resources.
- `systemctl is-active mcpedia-mcp` == active.
- Commit + push.
-67
View File
@@ -1,67 +0,0 @@
# MCPedia Phase 11 — CRUD + Auth + Web UI
## STATUS: ✅ ALL DONE (committed f8b525d, CI+Deploy success)
## Plan (spec BEFORE implementation, per user rule)
### Scope: 4 areas
1. **CRUD**: Create/Read/Update/Delete documents from web UI + MCP
2. **Authentication**: MCP/API writes (x-webhook-secret), Web CRUD (cookie-based ADMIN_PASSWORD)
3. **Web + Agent access**: Web forms + MCP write tools
4. **UI/UX**: Edit forms, TOC, dark mode
### Requirements:
1. Source of truth = filesystem (markdown files in content/{section}/{slug}.md)
2. DB mirrors disk (documents, document_chunks, document_revisions tables)
3. Single indexing path: indexContentFile in @mcpedia/core
4. Auth: MCP/API writes use WEBHOOK_SECRET; Web uses ADMIN_PASSWORD cookie
5. Slug rules: [a-z0-9][a-z0-9/_-]*, no //, no .. traversal
6. No breaking existing features (32 original tests still green)
7. UI/UX: edit button on doc pages, ?edit=1 inline form, /create page, login, TOC, dark mode
### Implementation
#### Backend
- `packages/parser`: added `stringifyFile()` (writes markdown with frontmatter)
- `packages/core`: `createDocument`, `updateDocument`, `deleteDocument` (file I/O + DB + revision + chunks)
- `apps/api`: tRPC CRUD routers (`requireWriteAuth`), fixed `requireWriteAuth` env-constant bug (now uses `ctx.expectedSecret` injected from `createApp(deps)`)
- `apps/mcp`: 3 new write tools (`create_document`, `update_document`, `delete_document`) gated by `x-webhook-secret`
#### Web UI
- `apps/web/app/api/auth/login/route.ts`: POST login → verify ADMIN_PASSWORD, set `mcpedia_admin` cookie
- `apps/web/app/api/docs/route.ts`: POST (create)
- `apps/web/app/api/docs/[...slug]/route.ts`: PUT (update), DELETE (delete)
- `apps/web/app/components/DocForm.tsx`: shared create/edit form
- `apps/web/app/create/page.tsx`: create form
- `apps/web/app/login/page.tsx`: login form
- `apps/web/app/[section]/[...slug]/page.tsx`: `?edit=1` inline edit, TOC, dark mode, Edit button
- `apps/web/app/docs/page.tsx`: docs index listing
- `apps/web/app/components/TOC.tsx`: auto-generated TOC from h2/h3 headings
- `apps/web/app/components/ThemeToggle.tsx`: dark mode toggle (localStorage + system default)
#### New deps (minimal — only for UX):
- `rehype-slug` (heading anchors for TOC links)
- `github-slugger` (matching slug algorithm for TOC client-side)
### Gotchas (learned the hard way)
1. tRPC fetch adapter expects input directly as JSON body, NOT JSON-RPC envelope
2. requireWriteAuth compared `ctx.webhookSecret !== WEBHOOK_SECRET` (module-level env constant) — untestable. Fixed: `ctx.webhookSecret !== ctx.expectedSecret` (injected per-app via deps).
3. Next.js catch-all routes: `[...slug]/edit/` is INVALID (catch-all must be last). Used `?edit=1` query param instead.
4. Next.js App Router: PUT/DELETE on `/api/docs/route.ts` doesn't match `/api/docs/{slug}` — need dynamic route `/api/docs/[...slug]/route.ts`.
5. `@env.example` should be updated.
6. `ADMIN_PASSWORD` must be set in VPS `.env` (deployed separately).
### Verification
- Typecheck: ✅ 4/4 apps green
- Tests: ✅ 40 tests green (32 original + 8 new), no DB/Redis
- Build: ✅ web compiled
- CI: ✅ success → Deploy: ✅ success
- Live: all 9 endpoints 200, 13 MCP tools live, CRUD e2e verified (login → create → view → delete via cookie auth), MCP create_document verified via header auth
- Test docs cleaned up (404 confirmed)
### Commits
1. `57f9001` feat: Phase 11 — CRUD + auth + web UI
2. `1dd16eb` feat(web): Phase 11 UI/UX — TOC, dark mode toggle, /docs index
3. `98437cb` fix(web): /api/docs accepts cookie OR header (not both required)
4. `8ed8c67` fix(web): split PUT/DELETE into /api/docs/[...slug]/route.ts
5. `f8b525d` chore: remove test docs
-24
View File
@@ -1,24 +0,0 @@
# Phase 3 — Deploy + git-sync wiring (remaining work)
Status: Phase 3/4 code is DONE and e2e-verified (webhook enqueue -> worker drain, 0 failed).
What was missing on the host: API + worker never ran as systemd services, and the GitHub
push webhook was never created. Also a real integration bug: `assertWebhookAuth` only
accepts a plain `x-webhook-secret` header, which GitHub does NOT send (GitHub delivers
`X-Hub-Signature-256` = HMAC-SHA256 of raw body). So a real GitHub webhook would 401.
## Changes
1. `apps/api/src/index.ts` — `assertWebhookAuth` now verifies GitHub `X-Hub-Signature-256`
(HMAC-SHA256 of raw body w/ WEBHOOK_SECRET) and still accepts `x-webhook-secret` for manual tests.
2. root `package.json` scripts — `api`: `bun --cwd apps/api run dev` -> `bun --cwd apps/api src/index.ts`
(the `run dev` form errors in bun 1.3.14; direct-file form verified booting + health). `worker` -> same form for consistency.
3. systemd — `cp deploy/*.service /etc/systemd/system`, `daemon-reload`, `enable --now mcpedia-api mcpedia-worker`.
4. Caddy — expose `/hooks/*` on `wiki.asepharyana.my.id` -> :4020 (no new DNS). Keep web on :4016.
5. GitHub webhook — `gh api repos/asepharyana/mcpedia/hooks` POST:
`https://wiki.asepharyana.my.id/hooks/reindex`, content_type json, secret=WEBHOOK_SECRET, events=push.
## Verification
- `systemctl is-active mcpedia-api mcpedia-worker` == active.
- `curl /health` on :4020 -> ok.
- `curl -X POST https://wiki.asepharyana.my.id/hooks/reindex -H "X-Hub-Signature-256: ..."` (or x-webhook-secret) -> 200 + jobId; worker drains.
- `gh api .../hooks` lists the webhook.
- `turbo run typecheck` green; commit + push.
-114
View File
@@ -1,114 +0,0 @@
# MCPedia — Phase 3 "Async + Scale" Implementation Plan
Status: Phase 1 (MVP) + Phase 2 (Semantic+API) DONE. Phase 3 adds async
background work, git-driven reindex, document revision history, and MCP
Resources. All logic stays in `@mcpedia/core`; new `packages/queue` wires
BullMQ; `apps/worker` runs the worker process; the existing API gets a git-sync
webhook + job-status procedures; the MCP server gains Resources.
## Scope (4 features from PHASES.md)
1. **Redis + BullMQ background indexing/embedding workers**
2. **Git synchronization hook** (auto-reindex on push via webhook)
3. **Document revision system** (`document_revisions`)
4. **MCP Resources** (`mcpedia://docs/...`) alongside existing tools
## Architecture decisions (locked)
- **Redis**: shared imrnes Redis `100.121.180.82:6379`, no auth (verified
`+PONG`). `REDIS_URL` env (default `redis://100.121.180.82:6379`), optional
`REDIS_PASSWORD`. BullMQ key prefix `mcpedia:` to avoid collisions on the
shared instance.
- **Queue lib**: `bullmq@6.1.2` + `ioredis@6.0.0` (BullMQ peer dep). Pass an
ioredis instance; BullMQ duplicates it for blocking commands.
- **Single source of truth preserved**: per-doc indexing logic moves into
`@mcpedia/core` as `indexContentFile(relPath, reason?)`. The script, the
worker, and the git hook ALL call this. Revisions are snapshotted inside it.
- **Revisions**: created only when body actually changes vs the latest revision
(avoids bloat on every sync). Stored in `document_revisions`.
## Files touched
### packages/config
- `src/index.ts`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
### packages/db
- `src/schema.ts`: add `documentRevisions` table
(id, documentId→documents.id cascade, slug, revisionNo int, title, body,
meta jsonb, reason text, createdAt). Index (document_id, revision_no DESC),
(slug).
- `drizzle/0002_document_revisions.sql`: migration (applied via psql).
- `drizzle/meta/0002_snapshot.json` + `_journal.json` entry (keeps drizzle-kit
consistent even though we apply manually).
### packages/core (new)
- `src/index.service.ts`:
- `indexContentFile(relPath: string, reason = "index")` — parse → upsert
`documents` → `indexChunks` → snapshot revision (if changed).
- `runFullIndex(reason?)` — walk content, index each, return counts.
- `src/revision.service.ts`:
- `createRevision(...)`, `listRevisions(slug, limit)`,
`getRevision(id)`, `latestRevisionBody(slug)`, `restoreRevision(id)`.
- `src/index.ts`: export both.
### packages/queue (NEW)
- `package.json` (@mcpedia/queue): deps bullmq, ioredis, @mcpedia/core,
@mcpedia/db, @mcpedia/config.
- `src/client.ts`: ioredis instance factory from config.
- `src/queue.ts`:
- `INDEX_QUEUE = "mcpedia-index"`.
- `enqueueIndexDoc(slug, absPath, reason)`, `enqueueFullIndex(reason)`.
- `getQueue()` lazy singleton.
- `src/worker.ts`: `startWorker()` — BullMQ Worker with 3 job types:
`index-doc` (single), `index-all` (full), `reindex` (full, reason=git-push).
Graceful shutdown on SIGINT/SIGTERM. Job progress + error handling.
### apps/worker (NEW)
- `package.json` (@mcpedia/worker): script `start: bun src/index.ts`.
- `src/index.ts`: `startWorker()` + heartbeat log.
### apps/api
- `src/index.ts`: add `POST /hooks/reindex` (full) and
`POST /hooks/index?slug=` (single) webhook routes → enqueue jobs. Mount
AFTER /trpc.
- `src/router.ts`: add `jobStatus` (id→state/prev/failedData),
`queueStatus` (waiting/active/completed/failed counts),
`revisions` (slug→list), `restoreRevision` (id→new slug/doc).
- `package.json`: add `@mcpedia/queue` dep, `hooks` reused.
### apps/mcp
- `src/index.ts`: register Resources:
- `mcpedia://docs` (list all metas)
- `mcpedia://docs/{slug}` (full body from disk)
- `mcpedia://docs/{slug}/chunks` (chunk previews)
- `mcpedia://docs/{slug}/revisions` (revision list)
- `src/smoke.test.ts`: add `listResources` + read `mcpedia://docs` assertion.
### scripts
- `scripts/indexer.ts`: refactor `main()` to call `runFullIndex()`.
### Root
- `package.json`: add `"worker": "bun --cwd apps/worker run start"`,
`"reindex": "bun run scripts/worker.ts"`? No — `worker` runs the listener;
triggering reindex = `bun run api` webhook or `enqueueFullIndex` helper.
Add `"enqueue-index": "bun run scripts/enqueue.ts"` (one-shot enqueue).
- `.env.example`: add `REDIS_URL`, `REDIS_PASSWORD`, `QUEUE_PREFIX`.
### Docs
- `PHASES.md`: mark Phase 3 items DONE with notes.
- `README.md`: document worker, webhook, revisions, MCP resources.
## Verification (real, not claimed)
1. `bun install` picks up new deps.
2. `bunx turbo run build` + `typecheck` green across workspace.
3. **Real BullMQ e2e against imrnes Redis**: script that enqueues an
`index-doc` job, starts a Worker, asserts the job completes and the doc row
+ chunks + a revision row appear in Postgres. Verifies Redis+ioredis+bullmq
+ db + core all wired correctly.
4. `bun --cwd apps/mcp run smoke` passes (incl. new resources).
5. `bun run index` (runFullIndex) green; verify `documents`,
`document_chunks`, `document_revisions` row counts via psql.
6. API webhook: `curl -XPOST localhost:4020/hooks/reindex` enqueues; worker
processes; `curl localhost:4020/trpc/queueStatus` reflects counts.
7. MCP resource read returns real content.
-105
View File
@@ -1,105 +0,0 @@
# MCPedia Phase 4 — Operability & Correctness Hardening
> Reinterpretation: the PHASES.md "Scale-out" items (OpenSearch, object storage,
> multi-tenant, distributed workers) are YAGNI at KB scale (4 docs). Phase 4 =
> make the Phase 3 async + revision machinery **correct, secure, observable, and
> deployable** — not speculative infra. Each task below fixes a real gap found
> by reading the code, not a hypothetical need.
## Tasks
### T1 — `restoreRevision` must rebuild semantic chunks (CORRECTNESS BUG)
**Root cause:** `packages/core/src/revision.service.ts` `restoreRevision` writes
the old body back into `documents` but never calls `indexChunks(slug, body)`.
So after a restore, keyword search (FTS on `documents.body`) is correct but
`document_chunks`/embeddings stay on the *new* body → semantic + hybrid search
return stale/ghost chunks.
**Fix:**
- Add `reindexChunks(slug)` to `@mcpedia/core` that re-runs `indexChunks(slug, body)`
using the live `documents.body` (the new body after the update).
- Call it inside `restoreRevision` after the `documents` update (wrap in try/catch
like `indexContentFile` so embed failure doesn't abort the restore).
- Add a unit-style assertion to the MCP smoke test or a small script: restore →
`document_chunks` count matches re-chunked body.
**Files:** `packages/core/src/revision.service.ts`, `packages/core/src/index.ts`,
`packages/core/src/document.service.ts` (export existing `indexChunks` if needed).
### T2 — Secure the git-sync webhook (SECURITY)
**Root cause:** `apps/api/src/index.ts` `/hooks/reindex` and `/hooks/index` accept
any request with no `WEBHOOK_SECRET` check — `.env.example` defines `WEBHOOK_SECRET`
but the router never reads it.
**Fix:**
- In `apps/api/src/index.ts`, compare `c.req.header("x-webhook-secret")` (or
`?secret=`) against `WEBHOOK_SECRET` (from `@mcpedia/config`). If unset/mismatch →
`401`. If `WEBHOOK_SECRET` env is empty, reject at startup with a clear log
(fail-fast, don't run an open endpoint).
- Add `WEBHOOK_SECRET` to `packages/config/src/index.ts` export.
- Document the header in README + verify with curl (401 without secret, 200 with).
**Files:** `apps/api/src/index.ts`, `packages/config/src/index.ts`, README.
### T3 — Web UI revisions view (UX)
**Root cause:** Web UI (server components) calls `@mcpedia/core` directly; there is
no revisions surface even though `revisions`/`restoreRevision` tRPC + MCP resource
exist.
**Fix (server-component only, no client JS):**
- On the doc page (`apps/web/app/[section]/[...slug]/page.tsx`), fetch
`listRevisions(fullSlug, 10)` and render a "History" panel: revision number,
reason, createdAt, body length, and a `/api/revisions/restore` link/POST that
calls the tRPC `restoreRevision` mutation via a server action or a form POST to
a small route handler. Simplest: a `<form method="post" action="/api/revisions/restore">`
with hidden `id` + a route handler in `apps/web` calling `restoreRevision`.
Keep it read-mostly; restore is a deliberate action.
- Add `apps/web/app/api/revisions/restore/route.ts` (POST) → `restoreRevision(id)`
→ `revalidatePath` the doc.
**Files:** `apps/web/app/[section]/[...slug]/page.tsx`,
`apps/web/app/api/revisions/restore/route.ts`.
### T4 — Paginate `listRevisions` / `revisions` API (PERF)
**Root cause:** `revision.service.ts` `listRevisions` does `select length(body)`
(fine) but the tRPC `revisions` and MCP resource return *full* revision rows
including the body in some callers; list endpoints should never carry bodies.
**Fix:**
- Ensure `listRevisions` summary excludes `body` (it already does — `bodyLength`
only). Add `offset` param for paging. Confirm MCP resource uses the summary.
- No behavior change for the doc page (uses summary).
**Files:** `packages/core/src/revision.service.ts` (add `offset`), router unchanged.
### T5 — Deployable as supervised services (OPS)
**Root cause:** `apps/api` and `apps/worker` run only ad-hoc; the host already runs
`zeavis` via Nix/systemd. Phase-4 operability = provide a systemd unit (or Nix
service) so `mcpedia-api` + `mcpedia-worker` start on boot and restart on failure.
**Fix (Nix-first, per MEMORY):**
- Write `mcpedia-api.service` + `mcpedia-worker.service` systemd unit files under
`deploy/` (bun run api / bun run worker, `WorkingDirectory`, `Restart=on-failure`,
`EnvironmentFile` pointing at `.env`, `After=network-online.target`).
- README section "Run as a service" with `cp deploy/*.service /etc/systemd/system && systemctl daemon-reload && systemctl enable --now mcpedia-api mcpedia-worker`.
- **Do NOT** `systemctl` on the host without user confirmation (changing live
services). Provide the files + instructions only; user runs enable.
**Files:** `deploy/mcpedia-api.service`, `deploy/mcpedia-worker.service`, README.
## Verification (all real, against imrnes Redis + Postgres)
1. `turbo run typecheck` + `turbo run build` green.
2. T1: script — edit a doc, reindex (new revision + new chunks), restore rev #1,
assert `document_chunks` count for that slug now matches re-chunk of rev #1 body
and `semanticSearch` on a term unique to rev #1 returns it.
3. T2: `curl -XPOST localhost:4020/hooks/reindex` → 401; with
`-H "x-webhook-secret: $WEBHOOK_SECRET"` → 200 + jobId.
4. T3: `next build` includes the History panel; restore form rebuilds chunks
(verified via T1 path through the route handler).
5. T4: `revisions` API returns summaries without body; offset paging works.
6. T5: `systemd-analyze verify deploy/*.service` passes (off-host safe check);
README documents enable steps.
## Out of scope (YAGNI, keep deferred per PHASES.md)
OpenSearch/Elasticsearch, object storage, multi-tenant, distributed workers,
pgvector migration. Revisit only when corpus > ~10k docs or query latency bites.
-41
View File
@@ -1,41 +0,0 @@
# MCPedia — "lanjut semua" workstream
Three real gaps remain (from review): tiny corpus (4 docs), MCP read-only (no write/auth),
no observability. This plan closes all three.
## 1. Content corpus (grow the KB)
Author real, useful docs so search/semantic/revisions have something to operate on.
Frontmatter schema (from packages/parser): id,title,type,tags,status,author,created_at,updated_at.
Sections: docs|writeups|research|notes. Files under content/<section>/...
New docs to add:
- content/docs/caddy/reverse-proxy.md (ops reference, tags: caddy, reverse-proxy, tls)
- content/docs/bullmq/workers.md (queue/worker reference, tags: bullmq, redis, jobs)
- content/docs/mcp/streamable-http.md (MCP transport reference, tags: mcp, protocol, http)
- content/notes/postgres/full-text-search.md (PG FTS notes, tags: postgres, fts, tsvector)
- content/writeups/infra/cloudflare-525.md (debugging writeup, tags: cloudflare, tls, 525)
After adding: `bun run index` to reindex (writes revisions + chunks), verify counts.
## 2. MCP write-tools + auth
Add mutating + admin tools to the MCP server (currently read-only):
- `index_document(slug)` -> enqueueIndexDoc (requires MCP auth header)
- `reindex_all()` -> enqueueFullIndex (requires MCP auth header)
- `queue_status()` -> getQueue counts (read, public)
- `restore_revision(id)` -> restoreRevision (requires MCP auth header)
Auth: MCP client must send header `x-webhook-secret` (reuse WEBHOOK_SECRET). StreamableHTTP
transport: read the Authorization/header in http.ts, pass to server via a factory closure
capturing the request; tools check it. Stateless per-request server already created fresh,
so threading the header is clean. Guard write-tools with the same requireWriteAuth logic.
Verify: unauthenticated call to index_document -> error; authenticated -> enqueues job.
## 3. Observability
- `GET /metrics` on the API (Prometheus text format): queue counts (waiting/active/
completed/failed/delayed), uptime, service name. Public (safe to expose).
- Caddy: expose /metrics on wiki. domain -> :4020 (add to handle list).
- tRPC `queueStatus` already exists; /metrics reuses getQueue.
Verify: curl /metrics -> text exposition with mcpedia_queue_* gauges.
## Verification
- typecheck green (turbo run typecheck).
- MCP smoke extended: tools/list shows new tools; authenticated index_document enqueues.
- /metrics returns 200 text; queue drains.
- commit + push.
-157
View File
@@ -1,157 +0,0 @@
# MCPedia Phase 9 — Test Coverage + Observability Hardening
**Goal:** Add a real test suite (CI-gated) covering every layer of MCPedia — pure
logic, Core services, the API surface, the MCP server + auth gates, and the Web
UI — so regressions are caught before deploy. The suite must run green in CI
with **no external services** (no Postgres, no Redis) by using in-process fakes.
## Constraints recap
- Tooling: bun workspaces + Turborepo, bun 1.3.14 has a built-in `bun:test` runner.
- DB is `imrnes` Postgres at `:6432` (no DB in CI) — tests must NOT touch it.
- pgvector is NOT installed (vectors are `real[]`, cosine in-app) — confirmed.
- Existing smoke test (`apps/mcp/src/smoke.test.ts`) is a script with `main()`
run via `bun run smoke`, NOT a `bun:test` file. It hits the DB → cannot run in CI.
- Secrets (`WEBHOOK_SECRET`, `DATABASE_URL`, `EMBED_*`) live in `.env` (gitignored)
or BWS for deploy — never in tests or committed config.
## Decisions
1. **Runner:** `bun:test` — zero-config, built into the bun 1.3.14 toolchain
already used. No extra deps. Add a `test` task to `turbo.json` and `package.json`
scripts; add a `Test` step to CI.
2. **No live DB in CI.** Tests that would need Postgres/Redis/Embeddings use
**in-process fakes** (memory stores + a stub embedder returning fixed vectors).
This means Core service tests cannot use the real `@mcpedia/db` singleton —
they must accept an injected DB (drizzle-pg mem or a hand-rolled fake). We will
**refactor the Core services' DB access behind injectable handles** where cheap,
and for the MCP/HTTP auth-layer tests we stub `@mcpedia/queue` + `@mcpedia/core`
at the module boundary (the transport/auth logic does not need a real queue).
3. **Test boundaries by package:**
- `packages/embeddings` — pure: `chunkText`, `cosine`. Real assertions, no I/O.
- `packages/search` — pure: `toTsQuery`, `cosine`. SQL-bearing functions
(`keywordSearch`/`semanticSearch`/`hybridSearch`) tested via a **fake db**
injected into `@mcpedia/db`, OR via the `cosine`/fusion helpers in isolation.
- `packages/core` — `snapshotRevision` dedup logic (refactor to accept an inject
fn or test the public `indexContentFile`/`restoreRevision` with fakes).
Focus: revision-dedup correctness + `restoreRevision` triggers `reindexChunks`.
- `apps/api` — Hono app: `/health`, `/metrics` shape, `/hooks/*` auth (401 w/o
secret, 200 + enqueue w/ secret using a fake queue), tRPC `restoreRevision`
mutation auth gate (401 w/o secret).
- `apps/mcp` — auth gates on write tools: `index_document`/`reindex_all`/
`restore_revision` error without secret, enqueue with secret (fake queue).
Read tools + resources via `InMemoryTransport` (reuse smoke style but without
DB).
- `apps/web` — render correctness of home (lists sections), doc page
(renders title + markdown + history panel when revisions exist), search page
(keyword/hybrid toggle, empty state). These need a fake Core.
## Approach per test (minimal, high-signal)
### embeddings: `packages/embeddings/src/chunk.test.ts`
- `chunkText("hello world")` with short size → single chunk.
- `chunkText` long text → multiple chunks, overlap honored, no word splits past boundary.
- `chunkText("")` / `" "` → `[]`.
### search: `packages/search/src/cosine.test.ts` (new tiny file) + refactor
- `cosine([1,0],[0,1])` ≈ 0; `cosine([1,1],[1,1])` = 1; `cosine([],[1])` = 0.
- `toTsQuery("a b c")` → `"a:* & b:* & c:*"`; empty/garbage → `""`.
### core: `packages/core/src/index.service.test.ts`
The hard part: `indexContentFile`/`restoreRevision`/`snapshotRevision` call `db`
directly. Two options:
- **Option A (chosen):** extract `snapshotRevision`'s "latest body" + "insert"
steps behind the existing `db` but make `indexContentFile` test the revision
*decision* by inserting a doc + revision directly via `db` in a test Postgres
(too heavy for CI).
- **Option B (chosen):** test the **pure decision logic** by refactoring
`snapshotRevision` to export a pure helper
`shouldCreateRevision(latestBody, body): boolean` — `true` when latest is null
or latest.body !== body. Then a unit test asserts the dedup truth table;
`indexContentFile` is verified by the existing e2e (manual `bun run index`).
This is the CI-safe win.
- `restoreRevision` correctness: assert it calls `reindexChunks(slug)` — we can
test by spying. Since `reindexChunks` is in the same module, we'll export a
seam: `restoreRevision(id, { reindexChunks: spy })` — keep backward compat by
defaulting. (Or test the public contract via the API layer instead.)
### api: `apps/api/src/index.test.ts`
- Build the Hono `app` from a testable factory that accepts a fake queue + fake
webhook secret. Current `index.ts` throws at import if `WEBHOOK_SECRET` unset —
that breaks import in CI. **Refactor:** move the fail-fast check into
`listen()`/serve start, so the app is constructable without a secret for
testing. Export `createApp(opts?)` returning the Hono instance.
- `/health` → 200 `{ok:true}`.
- `/metrics` → 200, text/plain, contains `mcpedia_uptime_seconds` +
`mcpedia_queue_jobs` for each state (fake queue returns 0/1).
- `POST /hooks/reindex` w/o `x-webhook-secret` → 401; with matching secret →
200 + `{ok:true, jobId}` (fake queue records the enqueue).
- tRPC: build a client against the app, call `restoreRevision` without secret →
error; the public `listDocuments` returns from a fake DB.
### mcp: `apps/mcp/src/auth.test.ts`
- `createMcpServer()` (no secret) → `index_document`/`reindex_all`/`restore_revision`
throw "unauthorized".
- `createMcpServer("secret")` → same tools reach the enqueue call (fake queue).
- Read tools still work without secret (server loads, resources list).
### web: `apps/web/app/search/page.test.tsx` (or a lighter harness)
- This is the hardest to test without a browser. **Decision:** keep web tests
minimal — assert that `toTsQuery`/render helpers exist; full DOM tests deferred
(needs playwright + a running server). We'll instead add a **contract test**
that the search page's `dynamic = "force-dynamic"` export exists (static-
generation guard, the kind of thing that broke CI before).
## File layout (new files)
```
packages/embeddings/src/chunk.test.ts
packages/search/src/cosine.test.ts
packages/core/src/index.service.test.ts # snapshotRevision + restoreRevision seam
apps/api/src/index.test.ts # Hono /health /metrics /hooks + tRPC gate
apps/mcp/src/auth.test.ts # write-tool auth gates
```
## turbo.json
Add a `test` task (like `typecheck`, no dependsOn, cache false so it always runs):
```jsonc
"test": { "cache": false }
```
Each app/pkg gets `"test": "bun test"` in its package.json.
## CI (`.github/workflows/ci.yml`)
After `Build`, add:
```yaml
- name: Test
run: bun run test # -> turbo run test
```
## Source changes required (enablers)
1. `apps/api/src/app.ts` (NEW) — extracted `createApp(deps?)` factory returning a
`Promise<Hono>`. Pure construction (no process exit, no fail-fast). Accepts
injected `ApiDeps` (`{ queue, webhookSecret }`); when omitted, lazily
imports the real queue + uses `WEBHOOK_SECRET` (production path). `/dashboard`
now renders the HTML from a new `apps/api/src/dashboard.ts` module.
2. `apps/api/src/dashboard.ts` (NEW) — the self-contained dashboard HTML, split
out of the original `index.ts` so the route is testable + the const is
importable. XSS-safe (esc() on all KB-sourced fields; documented in comment).
3. `apps/api/src/index.ts` — now a thin re-export of `createApp`/`start` + the
`isMain` bootstrap. Systemd unit still runs `bun --cwd apps/api src/index.ts`.
4. `packages/core/src/index.service.ts` — exported `shouldCreateRevision` (pure
predicate for the revision dedup invariant). `restoreRevision` in
`revision.service.ts` now accepts an optional `opts.reindex` seam (defaults
to the real `reindexChunks`), making the chunk-rebuild contract testable.
5. `apps/mcp/src/smoke.ts` — renamed from `smoke.test.ts` (so `bun test` doesn't
treat the integration smoke as a unit run), and fixed stale assertions:
expected tool set updated to all 10 tools (Phase 7 additions + write tools),
`list_documents(section=docs)` count updated to 4 (post-Phase-7 corpus).
## Verification
- `bun run test` (local) → 32 tests green across 6 packages (embeddings 5,
search 8, core 4, parser 5, mcp 6, api 8), **no live DB needed** (mocks
stub `@mcpedia/db`, `@mcpedia/queue`, `@mcpedia/core`).
- `bun run typecheck` → green (4 apps, no test-only type errors).
- `bun --cwd apps/mcp run smoke` → green (integration, needs live DB — runs in
CI on the deploy host, not in CI's no-services job).
- Live API check: `/health`, `/metrics`, `/dashboard`, `/hooks/*` auth gate
all 200/401-verified against a temp-port server.
- Commit + push; worker redeploy not needed (code + tests only).
-80
View File
@@ -1,80 +0,0 @@
# Phase 15 — UI/UX Polish Pass
> User: "perbagus ui ux nya"
## Audit
### Current issues
1. **Section index** (`[section]/page.tsx`):
- `buildFolderTree` uses path segments instead of doc titles for leaf nodes
- Folder tree + flat list at bottom is redundant
- No folder icons (📁), no doc type indicators, no dates
- Tree rendering is very basic (no visual hierarchy)
2. **Doc page** (`[section]/[...slug]/page.tsx`):
- CustomFieldBadges shows key=value but labels all badges with key name as title
- No structured metadata card layout
- Related docs cards are functional but visually flat
3. **Folder index** (`FolderIndexPage` in `[...slug]/page.tsx`):
- Shows doc slug parts (e.g. "pwn-100-ret2win-alignment") instead of titles
- No date, no tags, no description
- Subfolders are plain text links, no folder count or doc count
4. **Homepage** (`page.tsx`):
- Folder tree inside cards is cramped (text-xs, no spacing)
- Section cards show "View all (N)" but folder tree duplicates that
5. **Sidebar** (`Sidebar.tsx`):
- Tree uses indentation via inline style but no visual depth cues
- No hover expand for folders
- Active state only on exact match, not parent folders
## Plan
### 1. Section index — enhanced tree
- Use doc titles for leaf nodes (already have `doc` reference in tree)
- Add 📁 for folders, 📄 for docs
- Show doc count per folder
- Remove flat list (redundant with tree)
- Add "View all" link per section
- Better visual hierarchy: folder headers with counts, doc titles with dates
### 2. Doc page — metadata card
- Replace inline badge row with a structured metadata card
- Show custom fields as labeled badges (key → value with color)
- Keep CustomFieldBadges for backward compat but improve layout
- Add metadata card with: author, date, tags, and all custom fields
### 3. Folder index — rich listing
- Show doc titles (fetch from listDocuments, match by path)
- Show update date per doc
- Show tags per doc
- Add doc count for subfolders
- Better visual separation between subfolders and docs
### 4. Homepage — cleaner tree
- Increase font size for tree items
- Show section icon + label in tree header
- Better spacing between sections
### 5. Sidebar — depth cues
- Use ml-4 per level instead of inline style
- Add section divider lines
- Highlight active section + parent folders
## Files to change
```
mod: apps/web/app/[section]/page.tsx # enhanced tree, use doc titles
mod: apps/web/app/[section]/[...slug]/page.tsx # FolderIndexPage + metadata card
mod: apps/web/app/page.tsx # cleaner section tree cards
mod: apps/web/app/components/Sidebar.tsx # depth + active parent highlight
mod: apps/web/app/globals.css # additional CSS vars if needed
```
## Verification
- All endpoints still 200
- `bun run test` green
- `turbo run typecheck` green
- `next build` succeeds
- Live visual check of folder index, section index, doc page
+21
View File
@@ -7,3 +7,24 @@ This version has breaking changes — APIs, conventions, and file structure may
This block is written and re-added by `next dev` — verify at `node_modules/next/dist/server/lib/generate-agent-files.js`. Removing it from a diff only re-creates the uncommitted change; committing it with your work keeps the tree clean.
<!-- END:nextjs-agent-rules -->
# MCPedia Agent Instructions
MCPedia is a content-first knowledge base where Markdown files under `content/` serve as the Git-tracked source of truth, indexed into PostgreSQL and accessed via Web UI, tRPC API, and Model Context Protocol (MCP).
## Architecture Principles
1. **Single Core Layer (`@mcpedia/core`)**: All business logic (document CRUD, indexing, search, revisions, path classification) resides in `packages/core`. Never access the database directly from `apps/web`, `apps/mcp`, or `apps/api`.
2. **Content as Source of Truth**: Markdown files in `content/` with YAML frontmatter are the primary data store. The database stores metadata, search vectors, embeddings, and revision history.
3. **Multi-Modal Search**: Keyword search (Postgres FTS `tsvector` + GIN), Semantic search (cosine similarity over chunked embeddings), and Hybrid search (Reciprocal Rank Fusion / RRF) are unified in `@mcpedia/search`.
4. **Mutations & Security**: State-changing operations (document creation/updates/deletions, reindexing, revision restoration) require authentication (`x-webhook-secret` header or session cookie).
## Development Workflow
- **Install dependencies**: `bun install`
- **Run tests**: `bun run test`
- **Typecheck**: `bun run typecheck`
- **Build monorepo**: `bunx turbo run build`
- **Index content**: `bun run index`
- **Start dev servers**: `bun run dev`
-727
View File
@@ -1,727 +0,0 @@
# MCPedia — Phase Status
Legend: ✅ built · 🟡 partial · ⬜ deferred
## Phase 1 — MVP (✅ DONE)
| Capability | Status | Notes |
| --------------------- | ------ | ----- |
| Monorepo (bun + Turbo)| ✅ | apps/{web,mcp}, packages/{types,config,db,parser,search,core}, scripts |
| Content as Markdown | ✅ | `content/{docs,writeups,research,notes}/`, Git-tracked |
| Frontmatter parsing | ✅ | `@mcpedia/parser` (gray-matter) |
| Postgres metadata | ✅ | `@mcpedia/db` Drizzle, `documents` table |
| Postgres FTS | ✅ | weighted `tsvector` (title A / body B), GIN index, `ts_rank`+`ts_headline` |
| Core services | ✅ | Document / Content / Search — single business-logic layer |
| Indexer | ✅ | `scripts/indexer.ts` walks content/ → upserts |
| Web UI (Next 16) | ✅ | home (list), doc view (SSG), search (dynamic). react-markdown render |
| MCP server (stdio) | ✅ | 4 tools; in-memory smoke test passing |
| Hybrid/semantic search| ⬜ | Phase 2 |
## Phase 2 — Semantic + API
- [x] `packages/embeddings` — `EmbeddingProvider` interface + OpenRouter provider (via 9router `/v1`, `encoding_format:"float"`); `chunkText` + `embedChunks` batcher. `EMBED_DIM=2048` discovered live.
- [x] Schema `document_chunks` (id, document_id→documents.id cascade, slug, chunk_index, content, `embedding real[]`). Stored as `real[]` because pgvector **is not installed** on the shared imrnes Postgres (installing needs host-level apt — deferred). Cosine computed in-app; instant for a KB-sized corpus.
- [x] `scripts/indexer.ts` — chunks + embeds + upserts (per-doc replace).
- [x] `@mcpedia/search` — `semanticSearch` (cosine) + `hybridSearch` (FTS + cosine, RRF fusion). `keywordSearch` unchanged.
- [x] `apps/api` — Hono + tRPC v11 (`@trpc/server` fetch adapter, `@hono/node-server` on :4020): `search`, `semanticSearch`, `hybridSearch`, `getDocument`, `listDocuments`, `related`.
- [x] MCP server — added `semantic_search` + `hybrid_search` tools (6 total).
- [x] Web search — keyword/hybrid toggle (`?mode=hybrid`), hybrid reaches semantically-related docs keyword misses.
## Phase 3 — Async + Scale ✅ DONE
- [x] **Redis + BullMQ background indexing / embedding workers** —
`packages/queue` (ioredis singleton + BullMQ `Queue`/`Worker`, prefix
`mcpedia:` on shared imrnes Redis `:6379`); `apps/worker` runs
`startWorker()`. Three job types: `index-doc`, `index-all`, `reindex`.
Single indexing entry point `indexContentFile`/`runFullIndex` in
`@mcpedia/core` shared by the script, worker, and git hook. Verified
end-to-end against live Redis (job enqueue → worker → Postgres write).
- [x] **Git synchronization hook (auto-reindex on push)** — API webhook
`POST /hooks/reindex` (full) and `POST /hooks/index?slug=` (single) enqueue
BullMQ jobs. Wire a Git provider (GitHub/Gitea) post-receive / webhook to
`POST /hooks/reindex` to auto-reindex on push. `scripts/enqueue.ts` is a
one-shot enqueue helper (`bun run enqueue --all` / `<slug>`).
- [x] **Document revision system (`document_revisions`)** — `packages/db`
migration `0002_document_revisions.sql`. Indexer snapshots a revision only
when the body actually changes vs the latest revision (pure metadata edits
don't bloat history). `listRevisions` / `getRevision` / `restoreRevision`
in `@mcpedia/core`; exposed as tRPC `revisions` / `getRevision` /
`restoreRevision` and the `mcpedia://docs/{+slug}/revisions` MCP Resource.
- [x] **MCP Resources (`mcpedia://docs/...`)** — alongside the 6 tools:
`mcpedia://docs` (list), `mcpedia://docs/{+slug}` (body from disk),
`mcpedia://docs/{+slug}/chunks` (chunk preview),
`mcpedia://docs/{+slug}/revisions` (history). `{+slug}` uses RFC 6570
reserved expansion so slugs containing `/` match.
### New/changed commands
```
bun run index # full reindex (runFullIndex, writes revisions)
bun run enqueue --all # enqueue a full reindex job (no worker needed)
bun run enqueue <slug> # enqueue a single-doc reindex job
bun run worker # start the BullMQ indexing worker (long-running)
bun run api # Hono+tRPC API on :4020 (added /hooks/* webhooks)
```
### Verification done (real, against imrnes Redis + Postgres)
- `turbo run typecheck` green across all 13 packages.
- BullMQ e2e: enqueue `index-doc` → worker completes → `documents` +
`document_chunks` + `document_revisions` rows present.
- Revision dedup proven: editing a body creates a new revision; metadata-only
reindex does not; `restoreRevision` writes history back into the live row.
- MCP smoke test passes (tools + all 4 resources).
- API webhook `POST /hooks/reindex` enqueues → worker drains queue →
`queueStatus` reflects counts.
## Phase 4 — Operability & Correctness Hardening ✅ DONE
> Reinterpreted from the original "Scale-out" plan: OpenSearch/object-storage/
> multi-tenant were flagged YAGNI at KB scale (4 docs), so Phase 4 = make the
> Phase 3 async + revision machinery **correct, secure, observable, deployable**.
- [x] **T1 — `restoreRevision` rebuilds semantic chunks (CORRECTNESS BUG)** —
previously restore wrote the old body into `documents` but left `document_chunks`
on the *new* body, so semantic/hybrid search went stale after a restore.
`@mcpedia/core` `reindexChunks(slug)` now re-chunks + re-embeds from the live
body; `restoreRevision` calls it after the update (embed failure is logged, not
thrown). Verified: restore → `document_chunks` count matches re-chunk of the
restored body.
- [x] **T2 — Secure git-sync webhook (SECURITY)** — `/hooks/*` now require an
`x-webhook-secret` header matching `WEBHOOK_SECRET` (401 otherwise). API
fails fast at startup if `WEBHOOK_SECRET` is unset (no open endpoint). Added
`WEBHOOK_SECRET` to `@mcpedia/config` + `.env.example`; generated a real secret
in the local `.env` (gitignored).
- [x] **T3 — Web UI revisions view (UX)** — doc page now shows a "History" panel
(revision no, reason, date, body length) with a per-revision Restore button.
Restore POSTs to `apps/web/app/api/revisions/restore/route.ts` → `restoreRevision`
→ `revalidatePath` (server-component only, no client JS).
- [x] **T4 — Paginate `listRevisions`** — added `offset` param (summary never
includes body). API `revisions` + MCP resource use the summary.
- [x] **T5 — Deploy as supervised services (OPS)** — `deploy/mcpedia-api.service`
+ `deploy/mcpedia-worker.service` systemd units (`Restart=on-failure`,
`EnvironmentFile=.env`, `WorkingDirectory=/home/code/mcpedia`). Enable with:
`cp deploy/*.service /etc/systemd/system && systemctl daemon-reload &&
systemctl enable --now mcpedia-api mcpedia-worker`. (Not auto-enabled on host
without explicit user go-ahead.)
### Verification done (real, against imrnes Redis + Postgres)
- `turbo run typecheck` + `turbo run build` green (incl. `next build` with the
History panel).
- T1: edit → reindex (new revision + chunks) → restore rev #1 → `document_chunks`
count for that slug matches re-chunk of rev #1; `semanticSearch` on a term
unique to rev #1 returns it.
- T2: `curl -XPOST /hooks/reindex` → 401; with `-H "x-webhook-secret: $WEBHOOK_SECRET"`
→ 200 + jobId; job drains via worker.
- T3: History panel renders; restore route rebuilds chunks (T1 path).
- T4: `revisions` returns summaries (no body); `offset` paging works.
- T5: `systemd-analyze verify deploy/*.service` passes (off-host safe check).
## Phase 5 — Deferred scale-out (only when needed)
- [ ] Dedicated search engine (OpenSearch/Elasticsearch) — YAGNI until FTS is insufficient
- [ ] pgvector migration (install on imrnes Postgres) — when `real[]` cosine stalls
- [ ] Object storage for assets
- [ ] Advanced ranking, distributed workers, observability, multi-tenant
## Phase 6 — Network deployment + review hardening ✅ DONE
> Closed the real gaps found during review: the MCP server was stdio-only (unreachable
> over the network) and the tRPC API was not routed on the domain (swallowed by web →
> `/trpc/*` returned Next.js 404). Also found + fixed a security hole.
- [x] **MCP over Streamable HTTP** — `apps/mcp/src/http.ts` serves the 6 tools + 4
resources via MCP 2025-03-26 Streamable HTTP on `:4021`, stateless mode
(`sessionIdGenerator: undefined`, one server+transport per request, CORS on `/mcp`).
Deployed as `mcpedia-mcp.service`; reachable at `https://mcp.asepharyana.my.id/mcp`.
Stdio entry (`bun run mcp`) retained for local subprocess use.
- [x] **tRPC API routed on the domain** — Caddy `wiki.asepharyana.my.id` now forwards
`/trpc/*` (+ `/hooks/*`, `/health`) to the API on `:4020`; web stays on `:4016`.
Read-only procedures (search, list, revisions, job status) are public; the
`restoreRevision` mutation is gated by `x-webhook-secret` (see security fix below).
- [x] **Security: lock down `restoreRevision`** — the state-changing tRPC mutation was
anonymously callable over the network. Now requires `x-webhook-secret` (consistent
with `/hooks` auth). The Web UI calls `@mcpedia/core` directly in a server component,
so the gate does not affect the UI's restore button. Verified: no-secret → 401-class
rejection, with-secret → reaches handler.
- [x] **All four services live + supervised** — `mcpedia-web` (:4016), `mcpedia-api`
(:4020), `mcpedia-worker` (BullMQ), `mcpedia-mcp` (:4021) all `active`, reboot-safe.
GitHub push webhook → `https://wiki.asepharyana.my.id/hooks/reindex` (verified 200,
worker drains, 0 failed).
### Verification done (real, against live services)
- `https://mcp.asepharyana.my.id/mcp` initialize → 200 + serverInfo; tools/list → 6;
resources/list → 4 (Streamable HTTP SSE framing).
- `https://wiki.asepharyana.my.id/` → 200; `/health` → 200; `/trpc/listDocuments` → 200.
- GitHub push webhook delivers 200; `queueStatus` shows completed:N, failed:0.
- `turbo run typecheck` green across all 4 apps.
## Phase 7 — Corpus, MCP write-tools + auth, observability ✅ DONE
> Closed the remaining review gaps: tiny corpus (4 docs), MCP read-only (no write/auth),
> no observability.
- [x] **Content corpus grown** — added 5 real docs (Caddy reverse proxy, BullMQ workers,
MCP Streamable HTTP, Postgres FTS, Cloudflare-525 debugging writeup) across
docs/writeups/notes. `bun run index` reindexed: **9 documents, 17 chunks, 5 new
revisions** (was 4 docs). Search/semantic/revisions now operate on a real corpus.
- [x] **MCP write-tools + auth** — added `index_document`, `reindex_all`,
`restore_revision` (write, require `x-webhook-secret`) and `queue_status` (public).
`createMcpServer(authSecret?)` threads the request header; stdio keeps write tools
open (trusted local). Verified: unauthenticated `index_document` → `isError` +
"unauthorized"; authenticated → enqueues job, worker drains.
- [x] **Observability** — `GET /metrics` on the API (Prometheus text exposition:
`mcpedia_uptime_seconds`, `mcpedia_queue_jobs{state=...}`). Exposed on the domain at
`https://wiki.asepharyana.my.id/metrics`. Public, safe to scrape.
### Verification done (real, against live services)
- `https://wiki.asepharyana.my.id/metrics` → 200 Prometheus text (uptime + queue gauges).
- `tools/list` over MCP → 10 tools (6 read + 4 new). `index_document` auth gate works.
- Worker drained the MCP-enqueued job (completed count incremented, failed:0).
- `turbo run typecheck` green across all 4 apps.
## Phase 8 — Dashboard (observability UI) ✅ DONE
> The metrics endpoint existed (Phase 7) but had no consumer. Added a zero-dependency
> dashboard so the KB is actually observable + searchable from a browser.
- [x] **`GET /dashboard`** on the API — self-contained HTML (no build, no deps) that:
- pulls `/metrics` (same origin) and renders queue gauges (waiting/active/completed/
failed/delayed) + uptime, refreshing every 5s with a live-dot status indicator;
- runs a live **search box** that calls the MCP `hybrid_search` tool directly from the
browser (MCP `/mcp` is CORS-open), returning ranked hits that link to the web doc
page (`/docs/...`).
- [x] XSS hardening: all KB-sourced fields (`slug`/`title`/`section`/error message)
are `esc()`-escaped before `innerHTML` (defense-in-depth; data is server-trusted).
- [x] Caddy: `wiki.asepharyana.my.id/dashboard` → :4020.
### Verification done (real)
- `https://wiki.asepharyana.my.id/dashboard` → 200, serves the page (title + JS present).
- `/metrics` → 200, 7 gauge lines including `mcpedia_queue_jobs{state=...}`.
- MCP `hybrid_search` from browser path returns real ranked hits (verified the exact
`tools/call` payload the dashboard issues; shape `{doc:{slug,title,section},rank}`).
- Dashboard link points to working web doc route `/docs/<slug>` (verified 200).
- `turbo run typecheck` green.
## Phase 9 — Test coverage + CI gating ✅ DONE
> Before Phase 9 the only test was an integration smoke (`apps/mcp/src/smoke.test.ts`)
> requiring a live DB; its assertions had also rotted (expected 6 tools, now 10). Added
> a real `bun:test` suite that runs green in CI with **no external services** via
> in-process module mocking.
- [x] **Test infra** — `turbo.json` `test` task (cache:false); `test` script on every
package/app that has `.test.ts` files; `@types/bun` added to root devDeps;
`tsconfig.base.json` registers `types: ["bun","node"]`; CI step
`bun run test` added after `Build`.
- [x] **`@mcpedia/embeddings`** (5 tests) — `chunkText`: empty input, single chunk,
multi-chunk split, overlap/word-boundary integrity, default options.
- [x] **`@mcpedia/parser`** (5 tests) — `parseFile`: frontmatter extraction, section
derivation from top-level dir, invalid type/status fallbacks, missing-field
defaults, body excludes delimiter.
- [x] **`@mcpedia/search`** (8 tests) — `cosine` (orthogonal/identical/zero-vector/
mismatched-length/negative) + `toTsQuery` (AND-prefix, sanitization, empty/garbage).
- [x] **`@mcpedia/core`** (4 tests) — `shouldCreateRevision` dedup truth table (no prior
revision → snapshot; identical body → skip; changed body → snapshot; empty vs
non-empty). `restoreRevision` gained an `opts.reindex` seam for the chunk-rebuild
contract.
- [x] **`apps/api`** (8 tests) — refactored `index.ts` → `app.ts` `createApp(deps?)`
factory (pure construction, injectable `QueueLike`); `dashboard.ts` extracted;
`/health`, `/metrics`, `/hooks/reindex` (401 w/o secret, 200 w/ secret),
`/hooks/index` (400 w/o slug, 200 w/ slug+secret, 401 wrong secret), `/dashboard`.
- [x] **`apps/mcp`** (6 tests) — write-tool auth gates via `InMemoryTransport`:
`index_document`/`reindex_all`/`restore_revision` error without secret and enqueue
with secret (mocked `@mcpedia/queue` + `@mcpedia/core`); `queue_status` public;
tool discovery lists all 10 tools regardless of secret (gate is in handler).
- [x] **Fixed rot** — renamed `smoke.test.ts` → `smoke.ts` (so `bun test` doesn't run the
integration smoke as a unit test) and updated stale assertions (10-tool set, 4 docs
in `docs` section).
### Verification done (real)
- `bun run test` → 32 tests green across 6 packages, **no DB/Redis** (all fakes).
- `bun run typecheck` → 4 apps green (no test-only type errors).
- `bun --cwd apps/mcp run smoke` → SMOKE OK (integration, live DB).
- Live API (temp port): `/health`→200, `/metrics`→200 gauges, `/hooks/reindex`→401/200.
- CI workflow now runs `bun run test`.
### Files changed
```
new: apps/api/src/app.ts # createApp factory
new: apps/api/src/dashboard.ts # dashboard HTML module
new: packages/embeddings/src/chunk.test.ts
new: packages/parser/src/parse.test.ts
new: packages/search/src/cosine.test.ts
new: packages/core/src/index.service.test.ts
new: apps/api/src/app.test.ts
new: apps/mcp/src/auth.test.ts
mod: turbo.json, package.json, tsconfig.base.json, .github/workflows/ci.yml
mod: apps/api/src/index.ts (thin re-export), apps/api/src/router.ts (unchanged)
renamed: apps/mcp/src/smoke.test.ts -> smoke.ts (fixed stale assertions)
```
## Phase 10 — Audit + Bug fixes + CI/CD deploy (live verification)
> Full feature audit against live services. Found + fixed one real bug. Added the
> missing CI/CD deploy pipeline.
### Bug: Frontmatter leaking into rendered doc pages + MCP doc body
- **Symptom:** Doc pages showed raw YAML frontmatter (`id: websocket-contract`,
`title: WebSocket Contract`, etc.) as visible plain text between `<hr/>`
markers. MCP `mcpedia://docs/{+slug}` resource had the same leak.
- **Root cause:** `getDocument()` in `@mcpedia/core` preferred the on-disk file
via `readFileSync(abs, "utf8")` — returning **raw** file content including the
`---` frontmatter block. The indexer correctly stripped frontmatter via
`parseFile` (gray-matter), but `getDocument` bypassed it. `ReactMarkdown`
rendered `---` as `<hr/>` and the YAML as paragraphs.
- **Fix:** `packages/core/src/document.service.ts` — replaced `readFileSync` with
`parseFile(abs, row.path).body` (same frontmatter stripping as the indexer).
DB fallback (`row.body`) unchanged (already clean).
- **Verified:** 9/9 doc pages render clean (no frontmatter `id:` text, proper
`<h2>` + `<code>` elements in SSR HTML); MCP `mcpedia://` doc body resource
returns clean markdown (starts with `# WebSocket Contract`).
### Audit findings (all phases verified live)
| Phase | Feature | Live check | Status |
|-------|---------|------------|--------|
| P1 | Web UI `/docs/<section>/<slug>` | 200, renders markdown | ✅ |
| P1 | Search page (`?q=` + `?mode=hybrid`) | 200, returns results | ✅ |
| P1 | MCP stdio + HTTP (`/4021`) | 10 tools, 4 resources | ✅ |
| P2 | Semantic/hybrid search | returns ranked chunks | ✅ |
| P2 | tRPC API on domain (`/trpc/*`) | listDocuments → 4 docs | ✅ |
| P3 | BullMQ worker drains jobs | queue completed 17→19 after enqueue | ✅ |
| P3 | Revision system | listRevisions → rev #1 "phase4-final-clean" | ✅ |
| P3 | Git webhook auth gate | 401 w/o secret, 200 w/ secret | ✅ |
| P4 | Dashboard | `/dashboard` → 200 HTML | ✅ |
| P6 | All 4 systemd services | web/api/mcp/worker all `active` | ✅ |
| P6 | restoreRevision mutation locked | 401 w/o secret, executes w/ secret | ✅ |
| P7 | 10 MCP tools (6 read + 4 write) | tools/list → 10 | ✅ |
| P7 | Write-tool auth gate | reindex_all w/o secret → isError | ✅ |
| P7 | Prometheus metrics | `/metrics` → 7 gauges, 200 | ✅ |
| P8 | Dashboard live search | `fetch("/metrics")` + `hybrid_search` via `/mcp` | ✅ |
| P9 | Test suite | 6/6 packages, 32 tests, 0 fail | ✅ |
### Notes / non-bugs
- Doc URLs follow `/<section>/<slug>` (e.g. `/docs/caddy/reverse-proxy`,
`/writeups/infra/cloudflare-525`, `/notes/postgres/full-text-search`).
The route is `[section]/[...slug]` — `/docs/websocket/contract` works because
the section IS `docs` for that doc; `/notes/postgres/fts` does not (the correct
slug is `notes/postgres/full-text-search`).
- `restoreRevision` via tRPC needs the `x-webhook-secret` as an **HTTP header**
(not inside the JSON body) — the fetch adapter reads `c.req.raw.headers`.
### CI/CD deploy pipeline (Phase 10 addition)
Before this audit the CI workflow only built+tested — it did **not** deploy.
The VPS services were configured manually (systemd units in `deploy/`). Added a
`deploy.yml` workflow per the nix-ci-deploy pattern (CI builds, deploy is separate):
- **Trigger:** `workflow_run` on `CI` completion (only runs if CI passes).
- **Build:** same as CI (bun install + typecheck + web build) on the GitHub
runner — fails fast if the build is broken.
- **Deploy:** SSHes to the VPS over the public IP (`45.127.35.244`, not the
Tailscale `100.79.111.61` which GitHub runners can't reach), pulls from git,
reinstalls deps, rebuilds the web app, and restarts all 4 services via
`sudo systemctl restart` (NOPASSWD already configured for `code` user).
- **Secrets** (GitHub repo secrets, not files): `SSH_DEPLOY_HOST`,
`SSH_DEPLOY_PORT`, `SSH_DEPLOY_USER`, `SSH_DEPLOY_KEY` (ed25519 deploy key).
Deploy key's public half is in `/home/code/.ssh/authorized_keys` on the VPS.
#### Gotchas discovered + fixed during setup
1. **SSH host = public IP, not Tailscale IP.** `100.79.111.61` is a Tailscale
`tailscale0` interface IP (CGNAT `100.64.0.0/10`); GitHub Actions runners can't
route to it. Use the real public IP `45.127.35.244` (port 22 open in iptables).
2. **SSH key storage.** Storing the key via shell variable (`gh secret set --body "$VAR"`)
mangles newlines → `ssh.ParsePrivateKey: no key found`. Store directly from file:
`cat keyfile | gh secret set SSH_DEPLOY_KEY --repo ...`
Use `appleboy/ssh-action@v1` (not `@v1.1.0`) which correctly parses the key.
Add `-o IdentitiesOnly=yes` to prevent "too many authentication failures".
3. **Key rotation.** Force-pushing amended commits changes the SHA but CI triggers
on `push: branches: [main]` (CI) + `workflow_run` (deploy) — both fire correctly.
### Verification done (real)
- CI `workflow_run` → Deploy triggers after CI success; all 4 services `active`.
- All public URLs return 200: web `/`, `/docs/...`, `/search`, `/dashboard`,
`/metrics`, `/trpc/*`, `mcp.asepharyana.my.id/mcp`.
- Doc pages render clean markdown (no frontmatter); History panel + Restore work.
## Phase 11 — CRUD + Auth + Web UI ✅ DONE
> User requested: "perbagus agar jadi CRUD, pastikan ada autentikasi dan bisa
> manual dari web atau lewat agent melalui MCP, dan perbaui UI/UXnya."
### Backend (Core + API + MCP)
- [x] **`packages/parser` — `stringifyFile()`** — serialize `DocumentMeta` + body
back to a markdown file with YAML frontmatter (gray-matter). Round-trip stable
with `parseFile`.
- [x] **`@mcpedia/core` — CRUD functions:**
- `createDocument({slug, title, section, body, type?, status?, author?, tags?})`
— writes file to `content/{section}/{slug}.md`, upserts `documents` row,
snapshots revision, indexes chunks.
- `updateDocument(slug, {...})` — writes file, updates DB row, snapshots
revision (if body changed), reindexes chunks.
- `deleteDocument(slug)` — removes file + `documents`/`document_chunks`/
`document_revisions` rows.
- Slug validation: `[a-z0-9][a-z0-9/_-]*`, no `//`, no `..` traversal.
- [x] **`apps/api` — tRPC CRUD routers** — `createDocument`, `updateDocument`,
`deleteDocument` (all `.use(requireWriteAuth)`). Fixed `requireWriteAuth` to
compare against `ctx.expectedSecret` (injected from deps) instead of the
module-level `WEBHOOK_SECRET` env constant — latent bug that made the middleware
untestable without env manipulation.
- [x] **`apps/mcp` — 3 new write tools** — `create_document`, `update_document`,
`delete_document` (all require `x-webhook-secret`). Tools: 10 → 13.
- [x] **Auth** — MCP/API writes reuse the existing `WEBHOOK_SECRET` /
`x-webhook-secret` pattern. Web CRUD adds cookie-based auth: `ADMIN_PASSWORD`
env + `/api/auth/login` (HMAC-signed `mcpedia_admin` cookie, HttpOnly).
### Web UI
- [x] **`/create` page** — form (section/type/status/title/slug/tags/author/body),
POSTs to `/api/docs` with `x-webhook-secret`.
- [x] **`?edit=1` on doc pages** — inline edit form (`DocForm` component),
PUTs to `/api/docs/{slug}`.
- [x] **`/login` page** — password → `/api/auth/login` → cookie → redirect `/create`.
- [x] **Edit buttons** — homepage "+ Create Document" + per-doc "✎" (auth-gated);
doc page "Edit" button (auth-gated).
- [x] **TOC** — doc page auto-generates a table of contents from `h2` headings.
- [x] **Dark mode** — toggle persisted in `localStorage`, defaults to system.
- [x] **`/api/docs` REST routes** — POST (create), PUT (update), DELETE (delete),
all `x-webhook-secret` gated.
### Files changed
```
new: apps/web/app/api/auth/login/route.ts # cookie-based login + verify
new: apps/web/app/api/docs/route.ts # GET (list) + POST (create)
new: apps/web/app/api/docs/[...slug]/route.ts # PUT (update) + DELETE (delete)
new: apps/web/app/components/DocForm.tsx # shared create/edit form
new: apps/web/app/components/Sidebar.tsx # client-side doc navigation tree
new: apps/web/app/components/TOC.tsx # auto-generated TOC (github-slugger)
new: apps/web/app/components/ThemeToggle.tsx # dark mode toggle
new: apps/web/app/create/page.tsx # create UI
new: apps/web/app/login/page.tsx # login UI
new: apps/web/app/docs/page.tsx # docs index listing
mod: apps/web/app/layout.tsx # Linear design: sticky header + sidebar + dark canvas
mod: apps/web/app/page.tsx # editorial-style homepage w/ section doc listings
mod: apps/web/app/[section]/[...slug]/page.tsx # Linear doc layout (breadcrumb, TOC, metadata)
mod: apps/web/app/search/page.tsx # dark-themed search w/ result cards
mod: apps/web/app/components/Markdown.tsx # Linear typography + rehype-slug
mod: apps/web/app/globals.css # Inter font, Linear dark-mode-first palette
mod: apps/web/app/app/api/docs/route.ts # dual auth: cookie OR x-webhook-secret
mod: apps/web/app/api/docs/[...slug]/route.ts
mod: packages/core/src/document.service.ts # createDocument/updateDocument/deleteDocument
mod: packages/core/src/index.service.ts # export snapshotRevision
mod: packages/core/src/index.ts # re-export CRUD + types
mod: packages/parser/src/index.ts # stringifyFile
mod: packages/config/src/index.ts # ADMIN_PASSWORD
mod: apps/api/src/router.ts # CRUD routers + fix requireWriteAuth
mod: apps/api/src/app.ts # createContext passes expectedSecret
mod: apps/api/src/trpc.ts # Context.expectedSecret
mod: apps/mcp/src/index.ts # 3 new CRUD write tools
mod: apps/mcp/src/auth.test.ts # +4 CRUD auth tests
mod: apps/api/src/app.test.ts # +5 tRPC CRUD auth tests
mod: .env.example # ADMIN_PASSWORD
new dep: rehype-slug # heading anchors for TOC links
new dep: github-slugger # matching slug algorithm for client-side TOC
```
### Gotchas / lessons
1. **tRPC fetch adapter** expects input directly as JSON body, NOT JSON-RPC
envelope (`{"slug":...}` not `{"jsonrpc":"2.0","method":...,"params":{...}}`).
2. **`requireWriteAuth` env-constant bug** — comparing `ctx.webhookSecret !== WEBHOOK_SECRET`
(module-level env constant) is untestable. Fix: thread `expectedSecret` through
`Context` from `createApp(deps)`.
3. **Next.js catch-all routes** — `[...slug]/edit/` is invalid (catch-all must be
last). Used `?edit=1` query param instead. Also, Next.js App Router won't match
PUT/DELETE on `/api/docs/route.ts` for nested paths — need a dynamic segment
`/api/docs/[...slug]/route.ts`.
5. **`stringifyFile` YAML** — quote string values with `JSON.stringify` for
special-char safety; arrays use `[...]` syntax.
6. **Next.js SSG + DB** — client components (`"use client"`) don't block SSG
during `next build` even if they fetch at runtime. Used for Sidebar (fetches
/api/docs at runtime) to avoid ECONNREFUSED on CI.
7. **MCP SDK zod-v4 skew** — `@modelcontextprotocol/sdk@1.30` compiled `.d.ts`
references zod-4 internal types. Pin `zod: ^4.0.0` in the MCP app (not
workspace-wide) to match the SDK.
8. **StreamableHTTP transport headers** — use `requestInit: { headers: {...} }`
(not a top-level `headers` option) on `StreamableHTTPClientTransport`.
9. **SDK type noise** — use type-cast helpers (`as AnyContent`, `as CallToolResult`)
for the SDK's content union types which don't expose `.content[0].text` cleanly.
## Phase 12 — MCP Client + Full Layout Overhaul ✅ DONE
> User: "buat mcp untuk client nya" (create the MCP client) + "fokus ke web nya"
> (the user also said the UI was "boring, only style changed despite requesting
> a full layout overhaul").
### MCP Client (`apps/mcp-client/`)
New independent bun workspace package that connects to the MCPedia MCP server
over Streamable HTTP and provides both a programmatic API + CLI interface.
- **`src/client.ts`** — `McpediaClient` class wrapping the SDK's
`StreamableHTTPClientTransport` + `Client`. Typed methods for all 13 tools:
`listDocuments`, `getDocument`, `search`, `semanticSearch`, `hybridSearch`,
`getRelated`, `indexDocument`, `reindexAll`, `queueStatus`, `createDocument`,
`updateDocument`, `deleteDocument`, `listTools`, `callTool`, `listResources`,
`readResource`, `disconnect`. Accepts custom headers (`x-webhook-secret` for
write tools).
- **`src/index.ts`** — Interactive REPL (`bun run chat`): `/tools`, `/resources`,
`/search`, `/ss`, `/hybrid`, `/doc`, `/related`, `/create` (prompts), `/update`,
`/delete`, `/index`, `/status`, `/help`, `/quit`.
- **`src/ask.ts`** — One-shot CLI (`bun run ask <cmd> [args]`) for scripting.
- **`src/client.test.ts`** — 7 tests (mock SDK, no network/DB).
### Layout Overhaul (Linear design system)
Complete web UI redesign, not just style changes:
- **Dark-mode-first** — near-black canvas (`#08090a`), white-opacity borders
(`rgba(255,255,255,0.05–0.08)`), Inter font with `cv01/ss03` features.
- **Sticky header** — MCPedia brand + Docs/Search/Login + theme toggle.
- **Sticky sidebar** (xl+) — hierarchical doc tree with indented children.
- **Editorial homepage** — H1 + description, "Create Document" button,
section-organized doc listings with tag previews.
- **Doc page** — breadcrumb nav links, title + metadata bar (author/date/tags),
TOC (CONTENTS), clean prose rendering, Related + History sections.
- **Create/Edit** — `/create` page, `?edit=1` inline form (DocForm).
- **Login** — `/login` with dark-themed form + brand-indigo CTA.
- **Search** — `/search?q=` with dark-themed results + snippets.
- **Docs index** — `/docs` listing all documents by section.
### Gotchas
- Next.js catch-all `[...slug]/edit/` is invalid — used `?edit=1` query param.
- Next.js App Router: PUT/DELETE need `/api/docs/[...slug]/route.ts`.
- tRPC fetch adapter expects JSON body directly (not JSON-RPC envelope).
- `requireWriteAuth` env-constant bug: use `ctx.expectedSecret` from deps.
- Next.js SSG + DB: Sidebar as `"use client"` to avoid DB connection during build.
- MCP SDK zod-v4 skew: pin `zod: ^4.0.0` in the MCP app.
- `StreamableHTTPClientTransport`: use `requestInit: { headers }` not top-level `headers`.
- **Tooling:** bun workspaces + Turborepo (repo already used bun; pnpm rejected to minimize churn).
- **DB:** imrnes Postgres `100.121.180.82:6432/mcpedia` for both dev and deploy; driver `prepare:false` (PgBouncer). Docker Compose reserved for future prod.
- **Phase 1 scope:** Core + Web + MCP only. tRPC/Hono API, pgvector, auth, BullMQ deferred (YAGNI).
## Phase 13 — CTF Writeup Template System + Dynamic Custom Fields ✅ DONE
> User: "pastikan semua dinamis dan rapih untuk banyak situasi jadi tergantung
> user bukan hardcode" + "jadikan dinamis field nya jangan static begini, jadi
> yg membuat yg menentukan isinya"
### Problem
CTF writeups need per-event organization (event → many challenges). Initial
approach hardcoded CTF fields (event, challenge, category, difficulty, points)
at the **key-name** level — `if (key === "points")` styling, etc. User rejected
this: "field ditentukan user, bukan hardcode." Also, **tables weren't rendering** —
two bugs: (1) `remark-gfm` was missing (tables rendered as pipe text, not HTML),
(2) `@tailwindcss/typography` plugin was not installed in the web app (no CSS
for `prose` classes, so markdown had zero styling).
### Table rendering fix
- **Installed `remark-gfm@4`** — enables GFM table/strikethrough/task-list parsing
in `ReactMarkdown`. Tables now render as proper `<table>`/`<thead>`/`<tbody>`.
- **Installed `@tailwindcss/typography@0.5.20`** — `@plugin "@tailwindcss/typography"`
directive in `globals.css`. Generates `prose` CSS including table styling.
### Dynamic Custom Fields (fully dynamic, user-controlled)
The system is now **100% dynamic** — no field names or patterns are hardcoded.
The content creator adds **any** frontmatter key with **any** value type, and
the system auto-discovers + auto-styles:
1. **DB layer** — `documents.extra_fields` JSONB column (migration
`0003_document_extra_fields.sql`). Stores any key-value pairs as native JSONB
(numbers, booleans, strings, arrays, objects) — types preserved.
2. **Parser** (`parseFile`) — any frontmatter key not in the standard set
(`title`, `type`, `section`, `status`, `author`, `tags`, `created_at`,
`updated_at`) is extracted as an `extraField` → returned in `DocumentMeta.extraFields`.
3. **Parser** (`stringifyFile`) — writes `extraFields` back to YAML frontmatter
for round-trip stability (parse → stringify → parse yields same result).
4. **Core** — `createDocument`/`updateDocument` accept `extraFields?: Record<string, unknown>`,
merge with existing (update), pass through to DB + file.
5. **`toMeta`** (DB → meta) — spreads `extra_fields` JSONB from DB row into
`DocumentMeta`, making custom fields available to the UI.
6. **API routes** — `splitPayload()` separates standard CRUD fields from
custom fields. Custom fields passed through with preserved types (no
string coercion).
7. **`DocForm`** — "+ Add Field" UI lets content creators add ANY key-value pair.
Help text is generic (no hardcoded field-name examples).
8. **`CustomFieldBadges`** — auto-styles based on **VALUE TYPE + VALUE CONTENT**,
not key name:
| Value type | Badge style | Example |
|---|---|---|
| **Number** | purple, value shown | `5` → purple `5` |
| **Boolean** | green (true) / red (no) | `true` → green `Yes` |
| **Array** | purple, joined | `["a","b"]` → purple `a, b` |
| **Object** | gray, truncated JSON | `{timeout:30}` → gray `{"timeout":30}` |
| **String: difficulty-like** | color-coded | `easy`→green, `medium`→yellow, `hard`→red |
| **String: event-like** (contains ctf/def con/hack) | purple | `DEF CON CTF Quals 2024` |
| **String: points-like** (`100 pts`) | purple | `100 pts` |
| **String: category-like** (pwn/web/crypto) | orange | `Pwn` |
| **String: status-like** (solved/wip/pending) | color-coded | `Solved`→green |
| **Any other string** | default gray | `linux` → gray `linux` |
The same value `"medium"` gets the same yellow badge whether the key is
`difficulty`, `complexity`, `tier`, or `level`. Content creators control
the appearance via values, not by using specific key names.
### CTF Writeup Template + Sample
- **`content/writeups/ctf/template/writeup-template.md`** — template with:
Challenge Info table, Initial Recon, Approach, Step-by-Step Solve, Flag, Summary.
- **`content/writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment.md`** —
sample writeup demonstrating the template with arbitrary frontmatter fields.
### Deploy workflow fix
- Added `bun run scripts/indexer.ts` (explicit path) to deploy workflow instead
of `bun run index` (resolved to wrong package.json in SSH context).
- Removed `db:push` from deploy (column applied manually via ALTER TABLE; dribble
push is interactive/non-blocking). DB migration `0003_document_extra_fields.sql`
is tracked for future reference.
### Verification done (real, against live services)
#### Endpoints
| Path | Status |
|------|--------|
| `/` | ✅ 200 |
| `/docs` | ✅ 200 |
| `/docs/mcp/streamable-http` | ✅ 200 |
| `/notes/postgres/full-text-search` | ✅ 200 (table renders via GFM) |
| `/writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment` | ✅ 200 (dynamic badges + table + TOC) |
| `/writeups/ctf/template/writeup-template` | ✅ 200 |
| `/search` | ✅ 200 |
| `/login`, `/create`, `/dashboard` | ✅ 200 |
#### Dynamic badges verified
Created a test doc with fields of ALL value types and verified rendering:
| Field | Type | Rendered badge | Color |
|-------|------|----------------|-------|
| `os` | string "linux" | `linux` | gray (default) |
| `priority` | number 5 | `5` | **purple** (numeric) |
| `resolved` | boolean true | `Yes` | **green** (boolean) |
| `complexity` | string "medium" | `medium` | **yellow** (difficulty value match) |
| `team_members` | array ["alice","bob"] | `alice, bob` | **purple** (array) |
| `config` | object {timeout:30,retry:3} | `{"retry":3,"timeout":30}` | gray (truncated JSON) |
#### Tests
- `bun run test` → 7 task groups, all pass
- `turbo run typecheck` → green across all packages
- DB `extra_fields` column verified: `ALTER TABLE documents ADD COLUMN extra_fields jsonb DEFAULT '{}'::jsonb NOT NULL`
### Files changed
```
new: packages/db/drizzle/0003_document_extra_fields.sql # migration
mod: packages/types/src/index.ts # extraFields: Record<string, unknown>
mod: packages/parser/src/index.ts # parseFile extracts; stringifyFile writes
mod: packages/core/src/document.service.ts # accept extraFields
mod: packages/search/src/index.ts # toMeta spreads extra_fields from DB
mod: apps/web/app/[section]/[...slug]/page.tsx # CustomFieldBadges value-based styling
mod: apps/web/app/components/DocForm.tsx # generic help text (no hardcode)
mod: apps/web/app/api/docs/route.ts # splitPayload preserves types
mod: apps/web/app/api/docs/[...slug]/route.ts # splitPayload preserves types
mod: .github/workflows/deploy.yml # explicit indexer path
new: content/writeups/ctf/template/writeup-template.md # template
new: content/writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment.md # sample
```
## Phase 14 — Hierarchical Folder Structure ✅ DONE
> User: "gk ada bedanya, maksud saya inginnya itu bisa yg bertingkat seperti github
> yg memiliki folder dalam folder"
> (I want it to be hierarchical like GitHub, with folders inside folders)
### Problem
URLs were flat: `/<section>/<slug>` where slug could contain `/` but there were
**no folder index pages** — navigating to a folder path (e.g. `/writeups/ctf`)
returned 404. The sidebar showed a flat list with indentation based on slug depth,
but no actual folder-node entries or folder navigation.
### Solution
1. **Section index pages** (`apps/web/app/[section]/page.tsx`) — new generic
route for every section. Shows a folder tree (built from doc paths) + a flat
list of all docs in that section with "View all (N)" links. Previously only
`/docs/page.tsx` existed; now `/writeups`, `/research`, `/notes` all have
index pages.
2. **Folder index pages** (`apps/web/app/[section]/[...slug]/page.tsx`) — the
doc page now **classifies** the incoming slug path using
`classifyPath(docPaths, path)`:
- `"doc"` → leaf document (existing doc page behavior)
- `"folder"` → renders `FolderIndexPage` component listing subfolders 📁 +
immediate docs 📄
- `"none"` → 404
This is fully **content-driven** — no config maps. If a path has child docs,
it's a folder. If it matches a doc exactly, it's a leaf.
3. **Hierarchical sidebar** (`apps/web/app/components/Sidebar.tsx`) — rebuilds
the tree from flat doc slugs. Folder nodes (📁) have collapsible children;
leaf docs (📄) link directly. Indentation scales with depth.
4. **DocForm v2** (`apps/web/app/components/DocForm.tsx`) — now has a **parent
folder dropdown** populated from existing folders in the selected section.
The slug input is for the leaf name only; the resolved slug (with folder
prefix) is displayed below. Create/edit modes handled separately (edit keeps
slug read-only).
5. **Core helpers** (`packages/core/src/document.service.ts`) — exported
`extractFoldersForSection(docPaths, section)` and `classifyPath(docPaths, path)`
from `@mcpedia/core` so both web app and future MCP tools can use them.
6. **Example hierarchy** — reorganized the CTF writeup into proper nested folders:
```
writeups/
ctf/
_index.md ← folder intro
defcon-quals-2024/
_index.md ← event intro
pwn/
pwn-100-ret2win-alignment.md ← the actual writeup
template/
writeup-template.md
```
### Verification done (real, against live services)
| Path | Status | Type |
|------|--------|------|
| `/` | ✅ 200 | home |
| `/docs` | ✅ 200 | section index |
| `/docs/caddy/reverse-proxy` | ✅ 200 | doc page (nested slug) |
| `/writeups` | ✅ 200 | section index (tree) |
| `/writeups/ctf` | ✅ 200 | **folder index** (subfolder: defcon-quals-2024) |
| `/writeups/ctf/defcon-quals-2024` | ✅ 200 | **folder index** (subfolder: pwn) |
| `/writeups/ctf/defcon-quals-2024/pwn` | ✅ 200 | **folder index** (docs: pwn-100) |
| `/writeups/ctf/defcon-quals-2024/pwn/pwn-100-ret2win-alignment` | ✅ 200 | doc page + dynamic badges |
| `/writeups/ctf/template/writeup-template` | ✅ 200 | doc page |
| `/writeups/infra/cloudflare-525` | ✅ 200 | flat doc (still works) |
| `/notes/postgres/full-text-search` | ✅ 200 | nested doc (still works) |
- `bun run test` → 7 task groups, all pass
- `turbo run typecheck` → green across all packages
- `next build` → success (17 routes registered)
- Folder index pages render 📁 subfolders + 📄 docs with correct hrefs
- Section index pages render 📁 folder tree with indentation
- DocForm parent folder dropdown populated from real folder structure
### Files changed
```
new: apps/web/app/[section]/page.tsx # generic section index with tree
mod: apps/web/app/[section]/[...slug]/page.tsx # folder index detection
mod: apps/web/app/components/Sidebar.tsx # hierarchical tree
mod: apps/web/app/components/DocForm.tsx # parent folder dropdown
mod: apps/web/app/create/page.tsx # pass existingFolders
mod: packages/core/src/document.service.ts # extractFoldersForSection + classifyPath
new: content/writeups/ctf/_index.md # folder intro
new: content/writeups/ctf/defcon-quals-2024/_index.md # event intro
moved: content/writeups/ctf/defcon-quals-2024/pwn-100-ret2win-alignment.md
→ content/writeups/ctf/defcon-quals-2024/pwn/pwn-100-ret2win-alignment.md
```
+260 -136
View File
@@ -1,186 +1,310 @@
# MCPedia
> A content-first knowledge base — readable as Markdown/MDX in Git, queryable by
> humans via a Web UI and by AI agents via the Model Context Protocol (MCP).
> A content-first knowledge base — readable as Markdown in Git, queryable by humans via a modern Web UI, and accessible to AI agents via the Model Context Protocol (MCP).
MCPedia keeps content as plain Markdown files under `content/`. A Git-tracked
source of truth, indexed into PostgreSQL (metadata + a `tsvector` full-text
column) and served through a single **Core** layer that every interface
(Web, MCP) shares — no business logic duplicated per surface.
MCPedia keeps content as plain Markdown files under `content/` as a Git-tracked source of truth. Content is indexed into PostgreSQL (with a weighted `tsvector` full-text search column and chunked embeddings) and served through a unified **Core** layer that every interface (Web, MCP, API, Worker, CLI) shares.
## Monorepo layout
---
## Architecture Overview
```
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Web UI │ │ MCP Server │ │ tRPC / Hono │ │ BullMQ Worker │
│ (Next.js 16) │ │ (Stdio + HTTP) │ │ API (:4020) │ │ (Async Queue) │
└────────┬────────┘ └────────┬────────┘ └────────┬────────┘ └────────┬────────┘
│ │ │ │
└──────────────────────┴──────────┬───────────┴──────────────────────┘
▼
┌─────────────────────────┐
│ @mcpedia/core │
│ (Unified Business Logic)│
└────────────┬────────────┘
│
┌─────────────────────────────────┼─────────────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ Markdown Files │ │ @mcpedia/db │ │ @mcpedia/search │
│ (content/ tree) │ │ (PostgreSQL) │ │ (FTS + Cosine) │
└──────────────────┘ └──────────────────┘ └──────────────────┘
```
### Core Principles
1. **Content as Source of Truth**: Markdown files with YAML frontmatter under `content/` are primary. The database holds metadata, search indices, chunk embeddings, and revision history.
2. **Single Core Layer**: All business logic (CRUD, indexing, search, revisions, path classification) is encapsulated in `@mcpedia/core`. Interfaces never touch the database directly.
3. **Multi-Modal Search**: Keyword search (Postgres FTS), semantic search (vector cosine similarity), and hybrid search (Reciprocal Rank Fusion / RRF) work out of the box.
4. **Resilient Revision System**: Content edits snapshot revisions automatically; metadata-only edits are deduplicated. Restoring a revision automatically rebuilds semantic search chunks.
5. **Dual Interface**: Full human-friendly web experience + first-class AI agent integration via MCP.
---
## Monorepo Layout
```
mcpedia/
├── apps/
│ ├── web/ # Next.js 16 (Turbopack) — human-facing docs UI + search
│ ├── mcp/ # MCP server (stdio) — AI-agent interface (tools + resources)
│ └── api/ # Hono + tRPC v11 API on :4020 (+ /hooks/* git-sync webhooks)
│ ├── web/ # Next.js 16 (Turbopack) — dark-mode UI, hierarchical doc tree, TOC, CRUD
│ ├── mcp/ # MCP server (stdio + Streamable HTTP on :4021) — 13 tools + 4 resources
│ ├── mcp-client/ # Interactive MCP CLI REPL, one-shot command runner, and TypeScript client SDK
│ ├── api/ # Hono + tRPC v11 API on :4020, git-sync webhooks, Prometheus metrics, dashboard
│ └── worker/ # Long-running BullMQ worker for async indexing and embedding jobs
├── packages/
│ ├── types/ # shared domain types (DocSection, Document, SearchHit, ...)
│ ├── config/ # loads .env (repo root) as authoritative dev config
│ ├── db/ # Drizzle ORM schema + client + drizzle-kit config
│ ├── parser/ # frontmatter (gray-matter) parsing
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline)
│ ├── embeddings/ # embedding provider + chunker
│ ├── queue/ # Redis (ioredis) + BullMQ worker/queue (Phase 3)
│ └── core/ # Document/Content/Search/Index/Revision — the only business logic
├── content/ # docs/ writeups/ research/ notes/ (the knowledge base)
└── scripts/ # indexer.ts (full reindex), enqueue.ts (one-shot job enqueue)
│ ├── core/ # Document, Content, Search, Index, Revision, and Path services (the business logic)
│ ├── db/ # Drizzle ORM schema, client, and migrations
│ ├── parser/ # Frontmatter (gray-matter) parsing & stringification with dynamic field support
│ ├── search/ # Postgres FTS query (ts_rank + ts_headline), vector cosine, and RRF hybrid fusion
│ ├── embeddings/ # Chunking algorithms & embedding provider integrations
│ ├── queue/ # Redis (ioredis) client & BullMQ queue/worker definitions
│ ├── types/ # Shared TypeScript domain types and interfaces
│ └── config/ # Authoritative environment configuration loader (.env)
├── content/ # Markdown knowledge base organized by sections and nested folders
│ ├── docs/ # System and architectural documentation
│ ├── writeups/ # Technical writeups, CTF solutions, and debugging reports
│ ├── research/ # Research notes and evaluations
│ └── notes/ # Engineering notes and quick references
└── scripts/ # indexer.ts (full corpus reindexing), enqueue.ts (one-shot job enqueue)
```
## Architecture principle
---
```
Web ─┐
├──► Core ──► Repository (@mcpedia/db) ──► PostgreSQL
MCP ─┘
```
## Key Features
All interfaces go through `@mcpedia/core`. Nothing outside `packages/db` and
`packages/core` touches the database directly.
### 1. Multi-Modal Search
## Quick start
| Search Mode | Mechanism | Best Used For |
|---|---|---|
| **Keyword Search** | PostgreSQL `tsvector` (`simple` config, weighted A/B) + GIN index + `ts_rank` + `ts_headline` | Exact terms, code symbols, error codes, identifiers |
| **Semantic Search** | Text chunking (1000 chars, 150 overlap) + vector embeddings + cosine similarity ranking | Conceptual questions, paraphrased queries, intent matching |
| **Hybrid Search** | Reciprocal Rank Fusion ($RRF = \sum \frac{1}{k + rank}$) fusing keyword and semantic signals | General-purpose search with optimal relevance |
### 2. Hierarchical Folder Navigation
MCPedia supports nested folder structures (like GitHub repositories):
- **Dynamic Path Classification**: The router inspects paths and automatically distinguishes between leaf documents and folder nodes containing subfolders or child documents.
- **Folder Index Pages**: Navigating to any folder (e.g. `/writeups/ctf/defcon-quals-2024`) renders subfolders and documents within that path.
- **Collapsible Sidebar**: Hierarchical navigation tree reflecting the on-disk directory structure.
- **Section Indexes**: Dedicated overview pages for each section (`/docs`, `/writeups`, `/research`, `/notes`).
### 3. Dynamic Custom Fields
Content frontmatter supports arbitrary custom key-value pairs without schema modifications:
- **Automatic Storage**: Custom fields are persisted into a JSONB `extra_fields` column in PostgreSQL.
- **Type Preservation**: Numbers, booleans, arrays, objects, and strings maintain native types across parser, database, and API.
- **Value-Aware UI Badges**: The Web UI automatically styles badges based on value types and semantic patterns (difficulty levels, categories, tags, status) rather than hardcoded field names.
### 4. Revision History & Rollback
- **Smart Snapshotting**: Whenever a document body changes, a revision snapshot is created in `document_revisions`.
- **Deduplication**: Metadata-only updates do not produce duplicate body snapshots.
- **One-Click Restore**: Restoring any past revision writes the historic content back to disk and database, and automatically triggers semantic chunk re-indexing to ensure search consistency.
---
## Model Context Protocol (MCP)
MCPedia runs an MCP server accessible via **Stdio** (for local subagents) and **Streamable HTTP** (for remote agents over the network at `:4021`).
### MCP Tools (13 Total)
#### Read Tools (Public)
| Tool | Description |
|---|---|
| `search_documents` | Full-text keyword search over the knowledge base with headline snippets |
| `semantic_search` | Embedding-based cosine search across chunked content |
| `hybrid_search` | Fused full-text and semantic search via Reciprocal Rank Fusion |
| `get_document` | Fetches the full Markdown body of a document by slug |
| `list_documents` | Lists documents, optionally filtered by section or status |
| `get_related_documents` | Finds documents sharing tags with a given slug |
| `queue_status` | Returns current BullMQ indexing queue metrics |
#### Mutating & Admin Tools (Requires `x-webhook-secret`)
| Tool | Description |
|---|---|
| `create_document` | Creates a new document (writes Markdown file, database row, revision, and chunks) |
| `update_document` | Updates an existing document's body, metadata, or dynamic custom fields |
| `delete_document` | Removes a document file, database rows, revisions, and chunks |
| `index_document` | Enqueues a single-document background reindex job |
| `reindex_all` | Enqueues a full-corpus background reindexing job |
| `restore_revision` | Restores a document to a specific revision and rebuilds search chunks |
### MCP Resources
| Resource URI | Description | MIME Type |
|---|---|---|
| `mcpedia://docs` | List of all published documents in the knowledge base | `application/json` |
| `mcpedia://docs/{+slug}` | Full Markdown body of a single document from disk | `text/markdown` |
| `mcpedia://docs/{+slug}/chunks` | Preview of embedded semantic chunks for a document | `application/json` |
| `mcpedia://docs/{+slug}/revisions` | Summary of revision history for a document | `application/json` |
---
## HTTP, tRPC & Webhook Endpoints
The API application (`apps/api`) runs on **:4020** with Hono and tRPC v11:
### tRPC Procedures
- **Queries**: `search`, `semanticSearch`, `hybridSearch`, `getDocument`, `listDocuments`, `related`, `revisions`, `getRevision`, `jobStatus`, `queueStatus`.
- **Mutations** (Protected by `x-webhook-secret`): `createDocument`, `updateDocument`, `deleteDocument`, `restoreRevision`.
### Webhooks & Observability
- `POST /hooks/reindex` — Triggers a full corpus reindex (ideal for Git push webhooks).
- `POST /hooks/index?slug=<slug>` — Reindexes a single document.
- `GET /metrics` — Prometheus metrics exposition (`mcpedia_uptime_seconds`, `mcpedia_queue_jobs{state=...}`).
- `GET /dashboard` — Standalone real-time web dashboard for queue monitoring and live search.
- `GET /health` — Health check endpoint.
---
## Authentication & Security
MCPedia utilizes dual-tier security:
1. **Web UI Authentication**:
- Gated via `ADMIN_PASSWORD` in `.env`.
- Generates an HMAC-signed `mcpedia_admin` HttpOnly cookie upon login at `/login`.
- Protects document creation (`/create`), inline editing (`?edit=1`), and deletion.
2. **API & MCP Write Authentication**:
- Gated via `WEBHOOK_SECRET` in `.env`.
- Requires the `x-webhook-secret` HTTP header on mutating endpoints and tools.
- The API will refuse to start if `WEBHOOK_SECRET` is unset, preventing unsecured deployments.
---
## Quick Start
### Prerequisites
- [Bun](https://bun.sh/) (v1.1+)
- [PostgreSQL](https://www.postgresql.org/) (with `tsvector` support)
- [Redis](https://redis.io/) (for BullMQ async workers)
### 1. Installation & Environment Setup
```bash
bun install # install workspace deps
cp .env.example .env # set DATABASE_URL (dev uses imrnes Postgres :6432)
bunx turbo run build # typecheck + build every package
# Clone repository
git clone https://github.com/asepharyana/mcpedia.git
cd mcpedia
bun run index # walk content/ -> upsert into Postgres
bun --cwd apps/web run dev # Web UI on :3000
bun run mcp # MCP server on stdio (pipe to an MCP client)
# Install workspace dependencies
bun install
# Configure environment variables
cp .env.example .env
```
### Database
Ensure `.env` contains your database connection, Redis configuration, webhook secret, and admin password.
Schema is defined in `packages/db/src/schema.ts` (`documents` with a weighted
`search_vector` tsvector + GIN index, and `document_chunks` with an `embedding real[]`).
The `pgvector` extension is **not available** on the shared imrnes Postgres, so
semantic search stores vectors as `real[]` and ranks by in-app cosine similarity.
### 2. Database Setup
Migrations live in `packages/db/drizzle/`. They were applied manually via `psql`
(`drizzle-kit push` is unreliable under PgBouncer transaction pooling); to
re-apply on a fresh DB:
Apply migrations to initialize tables (`documents`, `document_chunks`, `document_revisions`):
```bash
psql $DATABASE_URL -f packages/db/drizzle/0000_grey_toro.sql
psql $DATABASE_URL -f packages/db/drizzle/0001_document_chunks.sql
psql $DATABASE_URL -f packages/db/drizzle/0002_document_revisions.sql
psql $DATABASE_URL -f packages/db/drizzle/0003_document_extra_fields.sql
```
> Note: on imrnes (PgBouncer `:6432`) a leaked `DATABASE_URL` shell var can
> shadow `.env`. `@mcpedia/config` loads `.env` **last** so the repo config
> always wins for local/dev.
### 3. Build & Index
## Content
```bash
# Typecheck and build all packages
bunx turbo run build
Each Markdown file carries YAML frontmatter:
# Run initial full indexation (parses content/, embeds chunks, writes DB)
bun run index
```
### 4. Running Services
```bash
# Start all development services (Web, API, MCP, Worker) concurrently
bun run dev
# Or run services individually:
bun --cwd apps/web run dev # Web UI on :3000 (or :4016 in production)
bun run api # Hono + tRPC API on :4020
bun run worker # BullMQ background worker
bun run mcp # MCP server (stdio mode)
bun run mcp:http # MCP server (Streamable HTTP mode on :4021)
```
---
## Scripts & CLI Reference
| Command | Action |
|---|---|
| `bun run index` | Run full synchronous reindexer over `content/` |
| `bun run enqueue --all` | Enqueue a full reindex job to Redis/BullMQ |
| `bun run enqueue <slug>` | Enqueue a single-document reindex job |
| `bun run worker` | Start the BullMQ background worker |
| `bun run api` | Start the Hono + tRPC API server |
| `bun run mcp` | Run MCP server in stdio mode |
| `bun run mcp:http` | Run MCP server in Streamable HTTP mode |
| `bun run chat` | Launch interactive MCP Client REPL |
| `bun run ask <cmd> [args]` | Run one-shot MCP client command |
| `bun run test` | Run the comprehensive test suite across all packages |
| `bun run typecheck` | Run TypeScript type checking across all workspaces |
---
## Content Format
Content files reside in `content/<section>/` as Markdown files with YAML frontmatter:
```yaml
---
id: websocket-contract
title: WebSocket Contract
id: sample-document
title: Sample Architecture Document
type: documentation
tags: [typescript, websocket, rpc]
section: docs
tags: [architecture, typescript, postgres]
status: published
author: asep
created_at: 2026-08-19
updated_at: 2026-08-19
created_at: 2026-08-20
updated_at: 2026-08-20
# Arbitrary custom fields (auto-discovered & styled):
difficulty: intermediate
points: 100
category: system-design
verified: true
---
# Sample Architecture Document
Your document content in standard GitHub Flavored Markdown (GFM).
Tables, code snippets, and custom sections are fully supported.
```
`slug` = relative path under `content/` (e.g. `docs/websocket/contract`). The
`body` shown in the UI is always read from the on-disk file (source of truth);
the DB stores metadata + the search vector.
---
## MCP tools
## Production Deployment
| Tool | Purpose |
| --------------------- | ------------------------------------------------ |
| `search_documents` | Postgres FTS over the corpus (ranked + snippet) |
| `semantic_search` | Embedding/cosine search over chunked content |
| `hybrid_search` | FTS + semantic fused via RRF |
| `get_document` | Full markdown body by slug |
| `list_documents` | List, optionally filtered by section |
| `get_related_documents` | Docs sharing tags with a given slug |
### MCP Resources
| URI | Purpose |
| -------------------------------- | ---------------------------------------- |
| `mcpedia://docs` | List all published documents |
| `mcpedia://docs/{+slug}` | Full markdown body (read from disk) |
| `mcpedia://docs/{+slug}/chunks` | Preview of embedded semantic chunks |
| `mcpedia://docs/{+slug}/revisions` | Revision history summary |
(`{+slug}` uses RFC 6570 reserved expansion so a slug like
`docs/websocket/contract` matches the template.)
Smoke test (in-memory transport, real JSON-RPC):
```bash
bun --cwd apps/mcp run smoke
```
## API (Phase 2 + Phase 3)
A tRPC v11 API is exposed via Hono on **:4020** (all procedures mirror the
MCP tools). Phase 3 adds async job + revision procedures and git-sync webhooks:
```bash
bun run api # http://localhost:4020 (GET /health, POST/GET /trpc/*)
```
tRPC procedures: `search`, `semanticSearch`, `hybridSearch`, `getDocument`,
`listDocuments`, `related` (Phase 2); plus `revisions`, `getRevision`,
`restoreRevision`, `jobStatus`, `queueStatus` (Phase 3).
Git-sync webhooks (enqueue BullMQ jobs; the worker processes them):
- `POST /hooks/reindex` — full-corpus reindex (point your Git provider's
push webhook here to auto-reindex on push).
- `POST /hooks/index?slug=<slug>` — reindex a single document.
> **Security:** both webhooks require an `x-webhook-secret` header that matches
> `WEBHOOK_SECRET` (set in `.env`). The API refuses to start if `WEBHOOK_SECRET`
> is unset, so the hooks are never left open.
`bun run index` now also chunks + embeds (Phase 2 indexer) and snapshots a
revision whenever the body changes (Phase 3). See `.env.example` for
`EMBED_*` / `REDIS_*` / `QUEUE_PREFIX` / `WEBHOOK_SECRET` vars.
### Run as a supervised service (Phase 4)
`deploy/mcpedia-api.service` + `deploy/mcpedia-worker.service` are systemd units
(`Restart=on-failure`, `EnvironmentFile=.env`, `WorkingDirectory=/home/code/mcpedia`).
Enable them with:
Production services run as supervised systemd units behind a Caddy reverse proxy:
```bash
# Copy systemd unit templates
sudo cp deploy/*.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now mcpedia-api mcpedia-worker
# tail logs
journalctl -u mcpedia-api -u mcpedia-worker -f
# Enable and start all supervised services
sudo systemctl enable --now mcpedia-web mcpedia-api mcpedia-worker mcpedia-mcp
# Inspect service logs
journalctl -u mcpedia-web -u mcpedia-api -u mcpedia-worker -u mcpedia-mcp -f
```
The API should sit behind Caddy (or your reverse proxy) for TLS; expose only
`:4020` internally and the web app publicly.
---
## Status
## Testing
**Phase 1 — MVP (DONE):** monorepo, Core, Web UI (home/doc/search), MCP server,
Postgres FTS keyword search, content indexing.
The test suite runs with in-memory mocks without requiring an active database or Redis instance:
**Phase 2 — Semantic + API (DONE):** embeddings provider (OpenRouter via 9router),
chunked `document_chunks`, `semanticSearch` + `hybridSearch` (RRF), tRPC/Hono API
(`apps/api`, :4020), MCP `semantic_search`/`hybrid_search` tools, web hybrid toggle.
```bash
bun run test
```
**Phase 3 — Async + Scale (DONE):** Redis + BullMQ background indexing/embedding
workers (`packages/queue`, `apps/worker`), git-sync webhooks (`POST /hooks/*`),
document revision system (`document_revisions` + restore), and MCP Resources
(`mcpedia://docs/...`). See `PHASES.md`.
> pgvector is **not installed** on the shared imrnes Postgres, so vector storage is
> a `real[]` column with in-app cosine similarity (instant at KB scale). pgvector is
> the Phase-4 scale-out path. See `PHASES.md`.
See `PHASES.md` for Phase 3–4 (Redis/BullMQ, auth, revisions, scale-out).
All 32+ unit and integration tests across packages and apps validate chunking, frontmatter parsing, cosine similarity, revision deduplication, auth gates, and route handlers.
+13 -10
View File
@@ -6,6 +6,7 @@ import {
hybridSearch,
keywordSearch,
listDocuments,
listSections,
semanticSearch,
listRevisions,
getRevision,
@@ -17,10 +18,9 @@ import {
import { getQueue, INDEX_QUEUE } from "@mcpedia/queue";
import { getConnection, BULLMQ_PREFIX } from "@mcpedia/queue/client";
// restoreRevision is a state-changing action (it rewrites the live document row
// + rebuilds its chunks). It must NOT be callable anonymously over the network —
// only the Web UI (which calls @mcpedia/core directly) and an operator with the
// webhook secret may use it. Anything else is rejected.
// restoreRevision and CRUD operations are state-changing actions.
// They must NOT be callable anonymously over the network —
// only an operator with the webhook secret (or authorized UI) may execute them.
const requireWriteAuth = t.middleware(({ ctx, next }) => {
if (!ctx.expectedSecret) {
throw new Error("WEBHOOK_SECRET is not configured; writes are disabled");
@@ -52,11 +52,13 @@ export const appRouter = router({
.input(z.object({ section: z.string().optional(), status: z.string().optional() }).optional())
.query(async ({ input }) => listDocuments(input ?? {})),
sections: publicProcedure.query(async () => listSections()),
related: publicProcedure
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(20).default(5) }))
.query(async ({ input }) => getRelated(input.slug, input.limit)),
// --- Phase 3: revisions ---
// --- Revisions ---
revisions: publicProcedure
.input(z.object({ slug: z.string(), limit: z.number().int().min(1).max(50).default(20) }))
.query(async ({ input }) => listRevisions(input.slug, input.limit)),
@@ -70,16 +72,16 @@ export const appRouter = router({
.input(z.object({ id: z.string() }))
.mutation(async ({ input }) => restoreRevision(input.id)),
// --- Phase 11: CRUD (gated by x-webhook-secret / admin auth) ---
// --- CRUD (gated by x-webhook-secret / admin auth) ---
createDocument: publicProcedure
.use(requireWriteAuth)
.input(
z.object({
slug: z.string().min(1),
title: z.string().min(1),
section: z.enum(["docs", "writeups", "research", "notes"]),
section: z.string().min(1),
body: z.string(),
type: z.enum(["documentation", "writeup", "research", "note"]).optional(),
type: z.string().optional(),
status: z.enum(["published", "draft"]).optional(),
author: z.string().optional(),
tags: z.array(z.string()).optional(),
@@ -95,7 +97,8 @@ export const appRouter = router({
slug: z.string().min(1),
title: z.string().min(1).optional(),
body: z.string().optional(),
type: z.enum(["documentation", "writeup", "research", "note"]).optional(),
section: z.string().min(1).optional(),
type: z.string().optional(),
status: z.enum(["published", "draft"]).optional(),
tags: z.array(z.string()).optional(),
author: z.string().optional(),
@@ -112,7 +115,7 @@ export const appRouter = router({
.input(z.object({ slug: z.string().min(1) }))
.mutation(async ({ input }) => deleteDocument(input.slug)),
// --- Phase 3: async job status ---
// --- Async job status ---
jobStatus: publicProcedure
.input(z.object({ id: z.string() }))
.query(async ({ input }) => {
+2
View File
@@ -36,6 +36,7 @@ mock.module("@mcpedia/core", () => ({
keywordSearch: () => Promise.resolve([]),
getDocument: () => Promise.resolve(null),
listDocuments: () => Promise.resolve([]),
listSections: () => Promise.resolve([]),
getRelated: () => Promise.resolve([]),
semanticSearch: () => Promise.resolve([]),
hybridSearch: () => Promise.resolve([]),
@@ -209,6 +210,7 @@ test("tool discovery works without auth (read tools present)", async () => {
expect(names).toContain("search_documents");
expect(names).toContain("get_document");
expect(names).toContain("list_documents");
expect(names).toContain("list_sections");
expect(names).toContain("semantic_search");
expect(names).toContain("hybrid_search");
expect(names).toContain("get_related_documents");
+28 -13
View File
@@ -4,6 +4,7 @@ import { ResourceTemplate } from "@modelcontextprotocol/sdk/server/mcp.js";
import { z } from "zod";
import {
listDocuments,
listSections,
getDocument,
getRelated,
semanticSearch,
@@ -75,14 +76,26 @@ export function createMcpServer(authSecret?: string): McpServer {
},
);
server.registerTool(
"list_sections",
{
description: "List all active knowledge base sections and document counts from PostgreSQL.",
inputSchema: z.object({}),
},
async () => {
const sections = await listSections();
return {
content: [{ type: "text", text: JSON.stringify(sections, null, 2) }],
};
},
);
server.registerTool(
"list_documents",
{
description: "List documents, optionally filtered by section.",
inputSchema: z.object({
section: z
.enum(["docs", "writeups", "research", "notes"])
.optional(),
section: z.string().optional().describe("Section name to filter by (e.g. 'docs', 'writeups', 'ctf')"),
}),
},
async ({ section }) => {
@@ -146,7 +159,7 @@ export function createMcpServer(authSecret?: string): McpServer {
},
);
// --- Phase 7: mutating + admin tools (require x-webhook-secret) ---
// --- Mutating & admin tools (require x-webhook-secret) ---
server.registerTool(
"index_document",
{
@@ -198,18 +211,18 @@ export function createMcpServer(authSecret?: string): McpServer {
},
);
// --- Phase 11: CRUD write tools (require x-webhook-secret) ---
// --- CRUD write tools (require x-webhook-secret) ---
server.registerTool(
"create_document",
{
description:
"Create a new document (writes markdown file + DB row + revision + chunks). Requires the x-webhook-secret header.",
"Create a new document in PostgreSQL (writes DB row, revision, chunks, and disk backup). Requires the x-webhook-secret header.",
inputSchema: z.object({
slug: z.string().describe("URL-safe slug (e.g. 'docs/my-new-doc')"),
slug: z.string().describe("URL-safe slug (e.g. 'docs/my-new-doc' or 'ctf/challenge-1')"),
title: z.string().describe("Document title"),
section: z.enum(["docs", "writeups", "research", "notes"]).describe("Content section"),
section: z.string().describe("Content section (e.g. 'docs', 'writeups', 'research', 'notes', 'guides', 'ctf')"),
body: z.string().describe("Markdown body"),
type: z.enum(["documentation", "writeup", "research", "note"]).optional(),
type: z.string().optional().describe("Document type (e.g. 'documentation', 'writeup', 'research', 'note')"),
status: z.enum(["published", "draft"]).optional(),
author: z.string().optional(),
tags: z.array(z.string()).optional(),
@@ -239,23 +252,25 @@ export function createMcpServer(authSecret?: string): McpServer {
"update_document",
{
description:
"Update an existing document (title, body, tags, status, etc.). Requires the x-webhook-secret header.",
"Update an existing document in PostgreSQL (title, body, section, tags, status, etc.). Requires the x-webhook-secret header.",
inputSchema: z.object({
slug: z.string().describe("Document slug to update"),
title: z.string().optional(),
body: z.string().optional(),
type: z.enum(["documentation", "writeup", "research", "note"]).optional(),
section: z.string().optional(),
type: z.string().optional(),
status: z.enum(["published", "draft"]).optional(),
tags: z.array(z.string()).optional(),
author: z.string().optional(),
extraFields: z.record(z.string(), z.unknown()).optional().describe("Dynamic custom metadata key-value pairs"),
}),
},
async ({ slug, title, body, type, status, tags, author, extraFields }) => {
async ({ slug, title, body, section, type, status, tags, author, extraFields }) => {
requireMcpAuth(authSecret);
const doc = await updateDocument(slug, {
title,
body,
section,
type,
status,
tags,
@@ -316,7 +331,7 @@ export function createMcpServer(authSecret?: string): McpServer {
},
);
// --- Phase 3: MCP Resources (read-only knowledge base surfaced via URIs) ---
// --- MCP Resources (read-only knowledge base surfaced via URIs) ---
// mcpedia://docs -> list all published documents
// mcpedia://docs/{slug} -> full markdown body (from disk)
// mcpedia://docs/{slug}/chunks -> chunked preview (semantic slices)
+8 -17
View File
@@ -8,7 +8,7 @@ import DocForm from "@/components/DocForm";
import TOC from "@/components/TOC";
import DocActions from "@/components/DocActions";
import { classifyPath, extractFoldersForSection } from "@mcpedia/core";
import { SECTIONS_BY_ID } from "@mcpedia/config/sections";
import { getSectionMeta } from "@mcpedia/config";
import type { DocumentMeta } from "@mcpedia/core";
export const dynamic = "force-dynamic";
@@ -19,17 +19,7 @@ interface DocPageProps {
}
function getSectionInfo(section: string) {
const info = SECTIONS_BY_ID.get(section);
if (!info) {
return {
label: section.charAt(0).toUpperCase() + section.slice(1),
icon: "📄",
};
}
return {
label: info.label,
icon: info.icon,
};
return getSectionMeta(section);
}
function calculateReadingTime(text: string): string {
@@ -128,9 +118,9 @@ function FolderIndexPage({
}) {
const prefix = `${section}/${slug}/`;
const folderDocPaths = allDocs
.filter((d) => d.path.startsWith(prefix))
.filter((d) => d.slug.startsWith(prefix))
.map((d) => ({
rel: d.path.slice(prefix.length).replace(/\.md$/, ""),
rel: d.slug.slice(prefix.length),
doc: d,
}))
.sort((a, b) => a.rel.localeCompare(b.rel));
@@ -292,8 +282,8 @@ export default async function DocPage({ params, searchParams }: DocPageProps) {
const fullSlug = `${section}/${slug.join("/")}`;
const allDocs = await listDocuments();
const docPaths = allDocs.map((d) => d.path);
const classification = classifyPath(docPaths, fullSlug);
const docSlugs = allDocs.map((d) => d.slug);
const classification = classifyPath(docSlugs, fullSlug);
if (classification === "folder") {
return <FolderIndexPage section={section} slug={slug.join("/")} allDocs={allDocs} />;
@@ -312,7 +302,7 @@ export default async function DocPage({ params, searchParams }: DocPageProps) {
]);
if (edit === "1" && canEdit) {
const existingFolders = extractFoldersForSection(docPaths, doc.section);
const existingFolders = extractFoldersForSection(docSlugs, doc.section);
return (
<div>
<Link
@@ -328,6 +318,7 @@ export default async function DocPage({ params, searchParams }: DocPageProps) {
mode="edit"
slug={fullSlug}
secret={WEBHOOK_SECRET}
existingSections={Array.from(new Set(allDocs.map((d) => d.section)))}
existingFolders={{
[doc.section]: existingFolders,
}}
+8 -18
View File
@@ -1,6 +1,6 @@
import Link from "next/link";
import { listDocuments } from "@mcpedia/core";
import { SECTIONS } from "@mcpedia/config/sections";
import { getSectionMeta } from "@mcpedia/config";
export const dynamic = "force-dynamic";
@@ -24,12 +24,11 @@ interface TreeNode {
function buildFolderTree(docs: DocMeta[], section: string): TreeNode[] {
const root: TreeNode[] = [];
for (const doc of docs.filter((d) => d.section === section)) {
const rel = doc.path.startsWith(`${section}/`)
? doc.path.slice(section.length + 1)
: doc.path;
const cleanPath = rel.replace(/\.md$/, "");
const parts = cleanPath.split("/");
for (const doc of docs.filter((d) => (d.section || "").toLowerCase() === section.toLowerCase())) {
const rel = doc.slug.startsWith(`${section}/`)
? doc.slug.slice(section.length + 1)
: doc.slug;
const parts = rel.split("/");
let current = root;
let currentPath = section;
@@ -118,19 +117,10 @@ export default async function SectionIndexPage({
params: Promise<{ section: string }>;
}) {
const { section } = await params;
const sectionInfo = SECTIONS.find((s) => s.id === section);
if (!sectionInfo) {
return (
<div className="py-12 text-center">
<h1 className="text-2xl font-semibold text-[var(--text-primary)]">Unknown section</h1>
<p className="text-sm text-[var(--text-muted)] mt-2">Section &quot;{section}&quot; does not exist.</p>
</div>
);
}
const sectionInfo = getSectionMeta(section);
const allDocs = await listDocuments();
const sectionDocs = allDocs.filter((d) => d.section === section);
const sectionDocs = allDocs.filter((d) => (d.section || "").toLowerCase() === section.toLowerCase());
const tree = buildFolderTree(allDocs, section);
return (
+15
View File
@@ -0,0 +1,15 @@
import { NextResponse } from "next/server";
import { listSections } from "@mcpedia/core";
export const dynamic = "force-dynamic";
// GET /api/sections — list all active sections with document counts dynamically from PostgreSQL.
export async function GET() {
try {
const sections = await listSections();
return NextResponse.json(sections);
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
return NextResponse.json({ error: msg }, { status: 500 });
}
}
+47 -29
View File
@@ -9,20 +9,21 @@ interface DocFormProps {
slug?: string;
secret: string;
existingFolders?: Record<string, string[]>;
existingSections?: string[];
initial?: {
title?: string;
body?: string;
section?: "docs" | "writeups" | "research" | "notes";
type?: "documentation" | "writeup" | "research" | "note";
status?: "published" | "draft";
section?: string;
type?: string;
status?: "published" | "draft" | string;
tags?: string[];
author?: string;
extraFields?: Record<string, unknown>;
};
}
const SECTION_OPTIONS = ["docs", "writeups", "research", "notes"] as const;
const TYPE_OPTIONS = ["documentation", "writeup", "research", "note"] as const;
const DEFAULT_SECTION_LIST = ["docs", "writeups", "research", "notes", "guides", "tutorials", "ctf", "api", "projects"];
const DEFAULT_TYPE_LIST = ["documentation", "writeup", "research", "note", "guide", "tutorial", "spec"];
const PRESET_FIELDS = [
{ key: "difficulty", value: "medium" },
@@ -37,6 +38,7 @@ export default function DocForm({
slug,
secret,
existingFolders,
existingSections,
initial,
}: DocFormProps) {
const router = useRouter();
@@ -56,6 +58,10 @@ export default function DocForm({
const textareaRef = useRef<HTMLTextAreaElement>(null);
const sectionOptions = Array.from(
new Set([...(existingSections ?? []), ...DEFAULT_SECTION_LIST, section].filter(Boolean)),
);
// Extract existing custom fields
const [customFields, setCustomFields] = useState<Array<{ key: string; value: string }>>(() => {
if (initial?.extraFields) {
@@ -198,7 +204,7 @@ export default function DocForm({
type="text"
value={title}
onChange={(e) => setTitle(e.target.value)}
placeholder="e.g. DEF CON Quals 2024 — pwn-100 Writeup"
placeholder="e.g. Database Architecture Overview"
className={`${baseInputCls} text-base`}
required
/>
@@ -210,34 +216,46 @@ export default function DocForm({
<label className="block text-xs font-semibold uppercase tracking-wider text-[var(--text-muted)] mb-1.5">
Section
</label>
<select
value={section}
onChange={(e) => setSection(e.target.value as typeof section)}
className={baseInputCls}
disabled={isEdit}
>
{SECTION_OPTIONS.map((s) => (
<option key={s} value={s}>
{s.toUpperCase()}
</option>
))}
</select>
<div className="relative">
<input
type="text"
list="section-list"
value={section}
onChange={(e) => setSection(e.target.value)}
placeholder="e.g. docs, writeups, ctf, guides"
className={baseInputCls}
required
/>
<datalist id="section-list">
{sectionOptions.map((s) => (
<option key={s} value={s}>
{s.toUpperCase()}
</option>
))}
</datalist>
</div>
</div>
<div>
<label className="block text-xs font-semibold uppercase tracking-wider text-[var(--text-muted)] mb-1.5">
Type
</label>
<select
value={type}
onChange={(e) => setType(e.target.value as typeof type)}
className={baseInputCls}
>
{TYPE_OPTIONS.map((t) => (
<option key={t} value={t}>
{t}
</option>
))}
</select>
<div className="relative">
<input
type="text"
list="type-list"
value={type}
onChange={(e) => setType(e.target.value)}
placeholder="e.g. documentation, writeup, note"
className={baseInputCls}
/>
<datalist id="type-list">
{DEFAULT_TYPE_LIST.map((t) => (
<option key={t} value={t}>
{t}
</option>
))}
</datalist>
</div>
</div>
<div>
<label className="block text-xs font-semibold uppercase tracking-wider text-[var(--text-muted)] mb-1.5">
+16 -4
View File
@@ -1,15 +1,27 @@
"use client";
import { useState } from "react";
import { useState, useEffect } from "react";
import Link from "next/link";
import { usePathname } from "next/navigation";
import { SECTIONS } from "@mcpedia/config/sections";
import { DEFAULT_SECTIONS, type SectionConfig } from "@mcpedia/config/sections";
import ThemeToggle from "@/components/ThemeToggle";
import CommandMenu from "@/components/CommandMenu";
export default function Header() {
const pathname = usePathname();
const [mobileMenuOpen, setMobileMenuOpen] = useState(false);
const [sections, setSections] = useState<SectionConfig[]>(DEFAULT_SECTIONS);
useEffect(() => {
fetch("/api/sections")
.then((r) => r.json())
.then((data) => {
if (Array.isArray(data) && data.length > 0) {
setSections(data);
}
})
.catch(() => {});
}, []);
return (
<header className="sticky top-0 z-30 border-b border-[var(--border-color)] glass-nav">
@@ -31,7 +43,7 @@ export default function Header() {
{/* Desktop Nav */}
<nav className="hidden md:flex items-center gap-1">
{SECTIONS.map((s) => {
{sections.map((s) => {
const isActive = pathname === `/${s.id}` || pathname.startsWith(`/${s.id}/`);
return (
<Link
@@ -107,7 +119,7 @@ export default function Header() {
{/* Mobile Drawer */}
{mobileMenuOpen && (
<div className="md:hidden border-t border-[var(--border-color)] bg-[var(--bg-surface)] px-4 py-3 space-y-1 animate-fade-in shadow-lg">
{SECTIONS.map((s) => (
{sections.map((s) => (
<Link
key={s.id}
href={`/${s.id}`}
+15 -4
View File
@@ -3,7 +3,7 @@
import { useEffect, useState, useMemo } from "react";
import Link from "next/link";
import { usePathname } from "next/navigation";
import { SECTIONS } from "@mcpedia/config/sections";
import { getSectionMeta } from "@mcpedia/config/sections";
interface Doc {
slug: string;
@@ -200,6 +200,17 @@ export default function Sidebar() {
);
}, [docs, filterText]);
const distinctSections = useMemo(() => {
const map = new Map<string, { id: string; label: string; icon: string }>();
for (const doc of filteredDocs) {
const s = (doc.section || "docs").toLowerCase();
if (!map.has(s)) {
map.set(s, getSectionMeta(s));
}
}
return Array.from(map.values()).sort((a, b) => a.label.localeCompare(b.label));
}, [filteredDocs]);
return (
<nav className="h-full flex flex-col py-4 px-3">
{/* Sidebar search filter */}
@@ -233,14 +244,14 @@ export default function Sidebar() {
{/* Summary count */}
<div className="flex items-center justify-between px-1 mb-3 text-[11px] text-[var(--text-dim)]">
<span className="uppercase tracking-wider font-mono">Documentation</span>
<span className="uppercase tracking-wider font-mono">Knowledge Base</span>
<span>{filteredDocs.length} items</span>
</div>
{/* Section Trees */}
<div className="flex-1 overflow-y-auto space-y-4 pr-1">
{SECTIONS.map(({ id, label, icon }) => {
const sectionDocs = filteredDocs.filter((d) => d.section === id);
{distinctSections.map(({ id, label, icon }) => {
const sectionDocs = filteredDocs.filter((d) => (d.section || "docs").toLowerCase() === id);
if (sectionDocs.length === 0) return null;
const tree = buildFolderTree(filteredDocs, id);
+11 -6
View File
@@ -1,6 +1,5 @@
import Link from "next/link";
import { listDocuments, extractFoldersForSection } from "@mcpedia/core";
import { SECTIONS } from "@mcpedia/config/sections";
import { listDocuments, extractFoldersForSection, listSections } from "@mcpedia/core";
import { WEBHOOK_SECRET } from "@mcpedia/config";
import { cookies } from "next/headers";
import { redirect } from "next/navigation";
@@ -13,8 +12,13 @@ export default async function CreatePage() {
const isAdmin = cookieStore.get("mcpedia_admin")?.value != null;
if (!isAdmin) redirect("/login");
const all = await listDocuments();
const docPaths = all.map((d) => d.path);
const [all, sections] = await Promise.all([
listDocuments(),
listSections(),
]);
const docSlugs = all.map((d) => d.slug);
const sectionIds = Array.from(new Set([...sections.map((s) => s.id), ...all.map((d) => d.section)]));
return (
<div className="space-y-6">
@@ -36,7 +40,7 @@ export default async function CreatePage() {
<div>
<h1 className="text-2xl font-bold text-[var(--text-primary)] tracking-tight">Create Document</h1>
<p className="text-xs text-[var(--text-muted)] mt-0.5">
Writes markdown file to disk, records revision, and computes 2048-dim vector embeddings.
Writes directly to PostgreSQL, records revision history, and computes vector embeddings.
</p>
</div>
</div>
@@ -45,8 +49,9 @@ export default async function CreatePage() {
<DocForm
mode="create"
secret={WEBHOOK_SECRET}
existingSections={sectionIds}
existingFolders={Object.fromEntries(
SECTIONS.map((s) => [s.id, extractFoldersForSection(docPaths, s.id)])
sectionIds.map((sid) => [sid, extractFoldersForSection(docSlugs, sid)])
)}
/>
</div>
+16 -15
View File
@@ -1,6 +1,6 @@
import Link from "next/link";
import { listDocuments } from "@mcpedia/core";
import { SECTIONS } from "@mcpedia/config/sections";
import { listDocuments, listSections } from "@mcpedia/core";
import { getSectionMeta, type SectionConfig } from "@mcpedia/config";
import McpConfigSnippet from "@/components/McpConfigSnippet";
export const dynamic = "force-dynamic";
@@ -116,8 +116,6 @@ function SectionTree({ section, docs, pathname }: SectionTreeProps) {
docs.filter((d) => d.slug.startsWith(`${section}/`)),
section,
);
const sectionInfo = SECTIONS.find((s) => s.id === section);
if (!sectionInfo) return null;
return (
<div>
@@ -131,7 +129,10 @@ function SectionTree({ section, docs, pathname }: SectionTreeProps) {
}
export default async function HomePage() {
const all = await listDocuments();
const [all, sections] = await Promise.all([
listDocuments(),
listSections(),
]);
const recent = [...all]
.sort(
@@ -146,7 +147,7 @@ export default async function HomePage() {
<section className="relative pt-4 pb-2">
<div className="inline-flex items-center gap-2 px-3 py-1 rounded-full bg-[#5e6ad2]/10 border border-[#5e6ad2]/25 text-[#5e6ad2] dark:text-[#7170ff] text-xs font-medium mb-6">
<span className="w-2 h-2 rounded-full bg-[#5e6ad2] animate-pulse" />
<span>Model Context Protocol + Unified Knowledge Base</span>
<span>Model Context Protocol + PostgreSQL Core</span>
</div>
<h1 className="text-4xl sm:text-5xl lg:text-6xl font-bold text-[var(--text-primary)] tracking-tight mb-5 leading-tight">
@@ -157,7 +158,7 @@ export default async function HomePage() {
</h1>
<p className="text-base sm:text-lg text-[var(--text-muted)] max-w-2xl leading-relaxed mb-8">
A high-performance, content-first documentation system. Browse hierarchical notes and CTF writeups in your browser, or connect AI coding assistants directly via native MCP tools.
A high-performance, database-backed documentation system. Browse hierarchical notes and technical writeups in your browser, or connect AI coding assistants directly via native MCP tools.
</p>
{/* Quick action buttons & stats */}
@@ -190,18 +191,18 @@ export default async function HomePage() {
<div className="grid grid-cols-2 sm:grid-cols-4 gap-3 p-4 bg-[var(--bg-surface)] border border-[var(--border-color)] rounded-xl max-w-3xl shadow-sm">
<div>
<div className="text-xl font-bold text-[var(--text-primary)]">{all.length}</div>
<div className="text-xs text-[var(--text-muted)]">Documents indexed</div>
<div className="text-xs text-[var(--text-muted)]">Documents in DB</div>
</div>
<div>
<div className="text-xl font-bold text-[var(--text-primary)]">{SECTIONS.length}</div>
<div className="text-xs text-[var(--text-muted)]">Core sections</div>
<div className="text-xl font-bold text-[var(--text-primary)]">{sections.length}</div>
<div className="text-xs text-[var(--text-muted)]">Active sections</div>
</div>
<div>
<div className="text-xl font-bold text-[#5e6ad2] dark:text-[#7170ff]">RRF Hybrid</div>
<div className="text-xs text-[var(--text-muted)]">FTS + 2048-dim vectors</div>
<div className="text-xs text-[var(--text-muted)]">FTS + Cosine Vectors</div>
</div>
<div>
<div className="text-xl font-bold text-emerald-600 dark:text-emerald-400">10 MCP Tools</div>
<div className="text-xl font-bold text-emerald-600 dark:text-emerald-400">13 MCP Tools</div>
<div className="text-xs text-[var(--text-muted)]">Streamable HTTP :4021</div>
</div>
</div>
@@ -216,8 +217,8 @@ export default async function HomePage() {
</div>
<div className="grid grid-cols-1 md:grid-cols-2 lg:grid-cols-4 gap-4">
{SECTIONS.map((s) => {
const sectionDocs = all.filter((d) => d.section === s.id);
{sections.map((s) => {
const sectionDocs = all.filter((d) => (d.section || "docs").toLowerCase() === s.id);
return (
<div
key={s.id}
@@ -272,7 +273,7 @@ export default async function HomePage() {
<div className="grid grid-cols-1 sm:grid-cols-2 lg:grid-cols-3 gap-4">
{recent.map((d) => {
const sInfo = SECTIONS.find((s) => s.id === d.section);
const sInfo = getSectionMeta(d.section);
return (
<Link
key={d.slug}
+14 -2
View File
@@ -3,7 +3,7 @@
import { useState, useEffect, useCallback, Suspense } from "react";
import Link from "next/link";
import { useSearchParams } from "next/navigation";
import { SECTIONS } from "@mcpedia/config/sections";
import { DEFAULT_SECTIONS, type SectionConfig } from "@mcpedia/config/sections";
interface DocHit {
slug: string;
@@ -24,10 +24,22 @@ function SearchContent() {
const [mode, setMode] = useState<SearchMode>(
["hybrid", "keyword", "semantic"].includes(initialMode) ? initialMode : "hybrid",
);
const [sections, setSections] = useState<SectionConfig[]>(DEFAULT_SECTIONS);
const [selectedSection, setSelectedSection] = useState<string>("all");
const [results, setResults] = useState<DocHit[]>([]);
const [loading, setLoading] = useState(false);
useEffect(() => {
fetch("/api/sections")
.then((r) => r.json())
.then((data) => {
if (Array.isArray(data) && data.length > 0) {
setSections(data);
}
})
.catch(() => {});
}, []);
const handleSearch = useCallback(async (q: string, searchMode: SearchMode) => {
if (!q.trim()) {
setResults([]);
@@ -154,7 +166,7 @@ function SearchContent() {
>
All Sections
</button>
{SECTIONS.map((s) => (
{sections.map((s) => (
<button
key={s.id}
type="button"
+2 -6
View File
@@ -27,7 +27,7 @@ The server uses `StreamableHTTPServerTransport` in **stateless mode**
- One `McpServer` + transport is created **per request**.
- No session affinity, no shared-transport `connect()` race, no session-map memory
leak under burst traffic.
- Re-registering the 6 tools + 4 resources per request is negligible for a KB-sized
- Re-registering tools and resources per request is negligible for a KB-sized
corpus.
Stateful mode (a `sessionIdGenerator` returning a UUID) would require holding a
@@ -53,8 +53,4 @@ clients can call it directly. Preflight `OPTIONS` is answered with 204.
## Auth for write tools
Read tools (`search_documents`, `get_document`, ...) are open. Write tools
(`index_document`, `reindex_all`, `restore_revision`) require the
`x-webhook-secret` header to match `WEBHOOK_SECRET` — the same shared secret used by
the git-sync webhook. A missing/invalid header makes the tool return an error before
any mutation.
Read tools (`search_documents`, `get_document`, `semantic_search`, `hybrid_search`, `list_documents`, `get_related_documents`, `queue_status`) are open. Write and mutating tools (`create_document`, `update_document`, `delete_document`, `index_document`, `reindex_all`, `restore_revision`) require the `x-webhook-secret` header matching `WEBHOOK_SECRET` — the same shared secret used by the git-sync webhook. A missing or invalid header returns an authorization error before executing any mutation.
+1 -1
View File
@@ -20,7 +20,7 @@ type-safe API described in the tRPC integration notes.
## Handshake
A client opens a single WebSocket connection and sends an `init` frame携带 an
A client opens a single WebSocket connection and sends an `init` frame containing an
auth token. The server answers with `ready` or closes the socket with code 4401
if the token is invalid.
+13 -6
View File
@@ -39,20 +39,27 @@ export const EMBED_BASE_URL = process.env.EMBED_BASE_URL ?? "";
export const EMBED_API_KEY = process.env.EMBED_API_KEY ?? "";
export const EMBED_MODEL = process.env.EMBED_MODEL ?? "";
// Phase 3: Redis + BullMQ (shared imrnes Redis, no auth by default).
// Redis + BullMQ (shared Redis instance).
export const REDIS_URL = process.env.REDIS_URL ?? "redis://100.121.180.82:6379";
export const REDIS_PASSWORD = process.env.REDIS_PASSWORD ?? "";
export const QUEUE_PREFIX = process.env.QUEUE_PREFIX ?? "mcpedia";
// Phase 4: git-sync webhook shared secret.
// Git-sync webhook and API write shared secret.
export const WEBHOOK_SECRET = process.env.WEBHOOK_SECRET ?? "";
// Phase 11: Admin password for web-based CRUD.
// Admin password for web-based CRUD.
export const ADMIN_PASSWORD = process.env.ADMIN_PASSWORD ?? "";
// Phase 14: Section registry — re-exported from a Node-free module so client
// components can safely import SECTIONS without pulling in node:fs.
export { SECTIONS, SECTIONS_BY_ID, SECTION_IDS, type SectionConfig } from "./sections";
// Dynamic Section registry & helpers — re-exported from a Node-free module.
export {
SECTIONS,
SECTIONS_BY_ID,
SECTION_IDS,
DEFAULT_SECTIONS,
SECTION_PRESETS,
getSectionMeta,
type SectionConfig,
} from "./sections";
// NOTE: we deliberately do NOT throw here if DATABASE_URL is empty.
// Throwing at import time breaks `next build` SSG and any runtime-injected env.
+61 -17
View File
@@ -1,10 +1,6 @@
/**
* Section registry — defines all content sections for the UI.
* This file is intentionally free of Node.js built-in imports (no fs/path)
* so it can be safely imported by client components.
*
* Adding a new section is as simple as adding an entry here + creating a
* `content/<id>/` directory with markdown files. No UI code changes needed.
* Dynamic Section Registry & Metadata Generator.
* Safe for client and server components.
*/
export interface SectionConfig {
id: string;
@@ -13,32 +9,80 @@ export interface SectionConfig {
desc: string;
}
export const SECTIONS: SectionConfig[] = [
{
id: "docs",
export const SECTION_PRESETS: Record<string, { label: string; icon: string; desc: string }> = {
docs: {
label: "Documentation",
icon: "📄",
desc: "Setup guides, API references, and protocol specs.",
},
{
id: "writeups",
writeups: {
label: "Writeups",
icon: "📝",
desc: "Post-mortems, debugging stories, and case studies.",
desc: "Post-mortems, CTF writeups, debugging stories, and case studies.",
},
{
id: "research",
research: {
label: "Research",
icon: "🔬",
desc: "Deep-dive analysis, architecture notes, and experiments.",
},
{
id: "notes",
notes: {
label: "Notes",
icon: "📌",
desc: "Quick references, patterns, and gotchas.",
desc: "Quick references, patterns, and cheat-sheets.",
},
guides: {
label: "Guides",
icon: "🧭",
desc: "Step-by-step guides and implementation walkthroughs.",
},
tutorials: {
label: "Tutorials",
icon: "🎓",
desc: "Educational lessons and practical tutorials.",
},
ctf: {
label: "CTF",
icon: "🚩",
desc: "Capture The Flag challenge writeups and exploit solutions.",
},
api: {
label: "API",
icon: "⚡",
desc: "API documentation and endpoint definitions.",
},
projects: {
label: "Projects",
icon: "🚀",
desc: "Project overviews and technical architecture roadmaps.",
},
};
export function getSectionMeta(id: string, count?: number): SectionConfig {
const cleanId = (id || "docs").toLowerCase().trim();
if (SECTION_PRESETS[cleanId]) {
return { id: cleanId, ...SECTION_PRESETS[cleanId] };
}
const formattedLabel = cleanId
.split(/[-_]/)
.map((w) => (w ? w.charAt(0).toUpperCase() + w.slice(1) : ""))
.join(" ");
return {
id: cleanId,
label: formattedLabel || "Custom Section",
icon: "📁",
desc: `Documents in the ${formattedLabel || cleanId} section.`,
};
}
export const DEFAULT_SECTIONS: SectionConfig[] = [
getSectionMeta("docs"),
getSectionMeta("writeups"),
getSectionMeta("research"),
getSectionMeta("notes"),
];
// Backwards-compatible constants:
export const SECTIONS: SectionConfig[] = DEFAULT_SECTIONS;
export const SECTIONS_BY_ID = new Map(SECTIONS.map((s) => [s.id, s]));
export const SECTION_IDS = SECTIONS.map((s) => s.id);
+123 -90
View File
@@ -1,7 +1,7 @@
import { and, eq, sql } from "drizzle-orm";
import { and, eq, sql, desc } from "drizzle-orm";
import { db } from "@mcpedia/db";
import { documentChunks, documentRevisions, documents } from "@mcpedia/db/schema";
import { CONTENT_ROOT } from "@mcpedia/config";
import { CONTENT_ROOT, DEFAULT_SECTIONS, getSectionMeta } from "@mcpedia/config";
import { parseFile } from "@mcpedia/parser";
import { existsSync, unlinkSync } from "node:fs";
import { join } from "node:path";
@@ -11,6 +11,7 @@ import type {
DocSection,
DocType,
DocStatus,
SectionInfo,
} from "@mcpedia/types";
import { chunkText, embedChunks, createEmbeddingProvider } from "@mcpedia/embeddings";
import { readContentFile } from "./content.service";
@@ -19,10 +20,47 @@ import { snapshotRevision } from "./index.service";
const embedder = createEmbeddingProvider();
export async function listDocuments(opts: {
section?: string;
status?: string;
} = {}): Promise<DocumentMeta[]> {
/**
* List all active sections dynamically from the database, including document counts.
*/
export async function listSections(): Promise<SectionInfo[]> {
const rows = await db
.select({
section: documents.section,
count: sql<number>`count(*)::int`,
latestUpdate: sql<string>`max(${documents.updatedAt})::text`,
})
.from(documents)
.where(eq(documents.status, "published"))
.groupBy(documents.section)
.orderBy(sql`count(*) desc`);
if (rows.length === 0) {
return DEFAULT_SECTIONS.map((s) => ({
...s,
docCount: 0,
}));
}
return rows.map((r) => {
const meta = getSectionMeta(r.section, r.count);
return {
...meta,
docCount: r.count,
updatedAt: r.latestUpdate,
};
});
}
/**
* List documents from the database, optionally filtered by section or status.
*/
export async function listDocuments(
opts: {
section?: string;
status?: string;
} = {},
): Promise<DocumentMeta[]> {
const status = opts.status ?? "published";
const where = [eq(documents.status, status)];
if (opts.section) where.push(eq(documents.section, opts.section));
@@ -30,35 +68,40 @@ export async function listDocuments(opts: {
.select()
.from(documents)
.where(and(...where))
.orderBy(documents.updatedAt);
.orderBy(desc(documents.updatedAt));
return rows.map(toMeta);
}
/**
* Fetch a single document directly from PostgreSQL (primary data store).
*/
export async function getDocument(slug: string): Promise<Document | null> {
const [row] = await db.select().from(documents).where(eq(documents.slug, slug));
if (!row) return null;
// Prefer the on-disk file (source of truth); fall back to stored body.
// Use parseFile (gray-matter) so the frontmatter is stripped — matches what
// the indexer stores and what ReactMarkdown expects.
const abs = join(CONTENT_ROOT, row.path);
let body: string;
if (existsSync(abs)) {
const { body: parsedBody } = parseFile(abs, row.path);
body = parsedBody;
} else {
body = row.body;
let body = row.body;
// Fallback to disk only if DB body is empty (e.g. during initial migration)
if (!body) {
const abs = join(CONTENT_ROOT, row.path);
if (existsSync(abs)) {
const parsed = parseFile(abs, row.path);
body = parsed.body;
}
}
return { ...toMeta(row), body };
}
/**
* Find related documents sharing tags with the given slug.
*/
export async function getRelated(slug: string, limit = 5): Promise<DocumentMeta[]> {
const [row] = await db
.select({ tags: documents.tags })
.from(documents)
.where(eq(documents.slug, slug));
if (!row || row.tags.length === 0) return [];
// Build a text[] array literal for the && (overlap) operator, binding each
// tag as a parameter to avoid SQL injection from frontmatter content.
const arrLit = sql`ARRAY[${sql.join(
row.tags.map((t) => sql.param(t)),
sql`, `,
@@ -76,7 +119,6 @@ export { readContentFile };
/**
* Chunk a document body, embed the chunks, and upsert them into
* `document_chunks` (replacing any prior chunks for the same slug).
* Failures are thrown so the caller can decide whether to abort the index.
*/
export async function indexChunks(slug: string, body: string): Promise<number> {
const [doc] = await db
@@ -110,13 +152,7 @@ export async function indexChunks(slug: string, body: string): Promise<number> {
}
// ---------------------------------------------------------------------------
// Phase 11: CRUD — create, update, delete documents.
//
// Source of truth for content is the filesystem: each doc is a markdown file
// under content/{section}/{slug}.md. The `documents` DB table mirrors the
// metadata + body for fast search. CRUD ops write the file first, then upsert
// the DB row, then snapshot a revision + reindex chunks. `deleteDocument`
// also cleans up chunks + revisions.
// Dynamic CRUD operations (Database-First)
// ---------------------------------------------------------------------------
export interface CreateDocInput {
@@ -134,6 +170,7 @@ export interface CreateDocInput {
export interface UpdateDocInput {
title?: string;
body?: string;
section?: DocSection;
type?: DocType;
status?: DocStatus;
tags?: string[];
@@ -150,10 +187,9 @@ function validateSlug(slug: string): string {
return slug;
}
/** Compute the relative file path for a slug (content/{section}/{slug}.md). */
/** Compute the relative file path for a slug. */
function slugToRelPath(section: DocSection, slug: string): string {
const cleanSlug = validateSlug(slug);
// If the slug already starts with the section, strip it to avoid doubling.
const pathPart = cleanSlug.startsWith(`${section}/`)
? cleanSlug.slice(section.length + 1)
: cleanSlug;
@@ -161,17 +197,15 @@ function slugToRelPath(section: DocSection, slug: string): string {
}
/**
* Create a new document: write the markdown file, upsert the DB row,
* snapshot a revision, and index semantic chunks.
* @returns the created DocumentMeta
* Create a new document in the PostgreSQL database, snapshot revision, and index chunks.
*/
export async function createDocument(input: CreateDocInput): Promise<DocumentMeta> {
const section = input.section;
const section = (input.section || "docs").trim();
const slug = validateSlug(input.slug);
const relPath = slugToRelPath(section, slug);
const absPath = join(CONTENT_ROOT, relPath);
if (existsSync(absPath)) {
const [existing] = await db.select({ id: documents.id }).from(documents).where(eq(documents.slug, slug));
if (existing) {
throw new Error(`document already exists at slug: ${slug}`);
}
@@ -191,11 +225,7 @@ export async function createDocument(input: CreateDocInput): Promise<DocumentMet
extraFields: input.extraFields ?? {},
};
// Write file to disk first (source of truth).
const { stringifyFile } = await import("@mcpedia/parser");
stringifyFile(absPath, relPath, meta, input.body);
// Upsert DB row.
// Upsert to DB directly
await db.insert(documents).values({
id: meta.id,
slug: meta.slug,
@@ -212,21 +242,28 @@ export async function createDocument(input: CreateDocInput): Promise<DocumentMet
updatedAt: new Date(meta.updatedAt),
});
// Snapshot revision + index chunks (best-effort; chunks must not block create).
await snapshotRevision(slug, meta, input.body, "index");
// Snapshot revision + index chunks
await snapshotRevision(slug, meta, input.body, "create");
try {
await indexChunks(slug, input.body);
} catch (err) {
console.error(`createDocument: chunk/embed FAILED for ${slug}: ${err instanceof Error ? err.message : err}`);
}
// Safe optional disk file sync
try {
const absPath = join(CONTENT_ROOT, relPath);
const { stringifyFile } = await import("@mcpedia/parser");
stringifyFile(absPath, relPath, meta, input.body);
} catch (fsErr) {
// Non-fatal
}
return meta;
}
/**
* Update an existing document: write new file, upsert DB row, snapshot a
* revision (if body changed), and reindex chunks.
* @returns the updated DocumentMeta
* Update an existing document in PostgreSQL, snapshot revision, and reindex chunks.
*/
export async function updateDocument(
slug: string,
@@ -240,12 +277,11 @@ export async function updateDocument(
...doc,
title: input.title ?? doc.title,
type: input.type ?? doc.type,
section: doc.section,
section: input.section ?? doc.section,
status: input.status ?? doc.status,
tags: input.tags ?? doc.tags,
author: input.author ?? doc.author,
updatedAt,
// Merge: new extraFields override old ones; merge with existing
extraFields:
input.extraFields !== undefined
? { ...doc.extraFields, ...input.extraFields }
@@ -253,12 +289,7 @@ export async function updateDocument(
};
const body = input.body ?? doc.body;
// Write file to disk (source of truth).
const absPath = join(CONTENT_ROOT, doc.path);
const { stringifyFile } = await import("@mcpedia/parser");
stringifyFile(absPath, doc.path, updated, body);
// Upsert DB row.
// Update DB row directly
await db
.update(documents)
.set({
@@ -268,13 +299,13 @@ export async function updateDocument(
status: updated.status,
author: updated.author,
tags: updated.tags,
extraFields: input.extraFields ?? doc.extraFields ?? {},
extraFields: updated.extraFields ?? {},
body,
updatedAt: new Date(updatedAt),
})
.where(eq(documents.slug, slug));
// Snapshot revision (only if body changed) + reindex chunks.
// Snapshot revision (if body changed) + reindex chunks
await snapshotRevision(slug, updated, body, "update");
try {
await indexChunks(slug, body);
@@ -282,52 +313,60 @@ export async function updateDocument(
console.error(`updateDocument: chunk/embed FAILED for ${slug}: ${err instanceof Error ? err.message : err}`);
}
// Safe optional disk file sync
try {
const absPath = join(CONTENT_ROOT, doc.path);
const { stringifyFile } = await import("@mcpedia/parser");
stringifyFile(absPath, doc.path, updated, body);
} catch (fsErr) {
// Non-fatal
}
return updated;
}
/**
* Delete a document: remove the file, delete DB rows (doc + chunks + revisions).
* Delete a document from PostgreSQL, chunks, revisions, and optional disk file.
*/
export async function deleteDocument(slug: string): Promise<{ deleted: boolean }> {
const [row] = await db.select().from(documents).where(eq(documents.slug, slug));
if (!row) return { deleted: false };
// Remove file from disk (source of truth).
const absPath = join(CONTENT_ROOT, row.path);
if (existsSync(absPath)) unlinkSync(absPath);
// Clean up DB rows (cascades would work but be explicit).
// Delete DB rows directly
await db.delete(documentChunks).where(eq(documentChunks.slug, slug));
await db.delete(documentRevisions).where(eq(documentRevisions.documentId, row.id));
await db.delete(documents).where(eq(documents.id, row.id));
// Safe optional disk cleanup
try {
const absPath = join(CONTENT_ROOT, row.path);
if (existsSync(absPath)) unlinkSync(absPath);
} catch (fsErr) {
// Non-fatal
}
return { deleted: true };
}
// ---------------------------------------------------------------------------
// Phase 14: Hierarchical folder structure helpers.
// These enable GitHub-style nested folder browsing — the content creator
// decides folder structure via where they place files; no config needed.
// Hierarchical folder structure helpers (dynamic, slug-driven)
// ---------------------------------------------------------------------------
/**
* Extract the folder structure for a given section from a list of document paths.
* Returns all distinct folder prefixes (relative to the section), sorted.
*
* Example: for section "docs" with paths ["docs/a/b/c.md", "docs/a/b/d.md"],
* returns ["a", "a/b"].
* Extract the folder structure for a given section from a list of document slugs/paths.
*/
export function extractFoldersForSection(
docPaths: string[],
docSlugs: string[],
section: string,
): string[] {
const folders = new Set<string>();
const base = `${section}/`;
const cleanSlugs = docSlugs.map((s) => s.replace(/\.md$/, ""));
for (const path of docPaths) {
const rel = path.startsWith(base) ? path.slice(base.length) : path;
const cleanPath = rel.replace(/\.md$/, "");
const parts = cleanPath.split("/");
for (const slug of cleanSlugs) {
if (!slug.startsWith(base)) continue;
const rel = slug.slice(base.length);
const parts = rel.split("/");
let acc = "";
for (let i = 0; i < parts.length - 1; i++) {
@@ -340,28 +379,22 @@ export function extractFoldersForSection(
}
/**
* Given a full URL slug path (e.g. "writeups/ctf/defcon-quals-2024"), determine
* if it represents a folder (i.e., there are other docs whose paths start with
* this prefix) or a leaf document.
*
* Returns:
* - "doc" if the path is a leaf document (path + ".md" matches a doc)
* - "folder" if the path is a parent of other doc paths
* - "none" if neither
* Determine if a path represents a folder or a leaf document.
*/
export function classifyPath(
docPaths: string[],
docSlugs: string[],
path: string,
): "doc" | "folder" | "none" {
const dotMd = `${path}.md`;
const cleanPath = path.replace(/\.md$/, "");
const cleanSlugs = docSlugs.map((s) => s.replace(/\.md$/, ""));
// Check if it's a leaf document
if (docPaths.includes(dotMd)) return "doc";
// Check leaf doc match
if (cleanSlugs.includes(cleanPath)) return "doc";
// Check if it's a folder (parent of other paths)
const prefix = `${path}/`;
const hasChildren = docPaths.some((p) => p.startsWith(prefix));
if (hasChildren) return "folder";
// Check folder parent match
const prefix = `${cleanPath}/`;
if (cleanSlugs.some((s) => s.startsWith(prefix))) return "folder";
return "none";
}
+3 -4
View File
@@ -2,17 +2,16 @@ import { test, expect } from "bun:test";
import { shouldCreateRevision } from "../src/index.service";
/**
* Phase 9: unit tests for the revision-dedup decision rule.
* Unit tests for the revision-dedup decision rule.
*
* `shouldCreateRevision` is the pure predicate that `indexContentFile` consults
* before writing a new row to `document_revisions`. It's extracted because the
* before writing a new row to `document_revisions`. It is extracted because the
* dedup correctness is the single most important guarantee of the revision
* system ("metadata-only edits don't bloat history"), and it must hold without
* a database.
*
* The DB-backed paths (`snapshotRevision`, `restoreRevision`) are exercised
* end-to-end by the existing manual e2e (`bun run index` + restore via the web
* /api/revisions/restore route, see PHASES.md Phase 4 verification). Here we
* end-to-end via the indexer and the web revision restore route. Here we
* lock the decision invariant in CI.
*/
+2
View File
@@ -12,4 +12,6 @@ export type {
Document,
DocumentMeta,
SearchHit,
SectionInfo,
} from "@mcpedia/types";
+6 -11
View File
@@ -8,14 +8,6 @@ import type {
DocumentMeta,
} from "@mcpedia/types";
const SECTIONS: DocSection[] = ["docs", "writeups", "research", "notes"];
const VALID_TYPES: DocType[] = [
"documentation",
"writeup",
"research",
"note",
];
// Standard frontmatter keys that are rendered explicitly in the UI template.
// Any other key in frontmatter becomes a dynamic "extra field" badge.
@@ -41,13 +33,16 @@ export function parseFile(absPath: string, relPath: string): ParsedFile {
const raw = readFileSync(absPath, "utf8");
const { data, content } = matter(raw);
const parts = relPath.split("/");
const derivedSection = parts.length > 1 ? parts[0] : "docs";
const section: DocSection =
(SECTIONS.find((s) => relPath.startsWith(s + "/")) as DocSection | undefined) ??
"docs";
typeof data.section === "string" && data.section.trim() !== ""
? data.section.trim()
: derivedSection;
const slug = relPath.replace(/\.mdx?$/, "");
const type = (VALID_TYPES.includes(data.type) ? data.type : "documentation") as DocType;
const type = (typeof data.type === "string" && data.type.trim() !== "" ? data.type.trim() : "documentation") as DocType;
const status = (data.status === "draft" ? "draft" : "published") as DocStatus;
const tags: string[] = Array.isArray(data.tags)
+12 -5
View File
@@ -57,13 +57,20 @@ test("parseFile: section derived from top-level dir", () => {
expect(writeDoc("notes/baz.md", "---\ntitle: C\n---\nbody").meta.section).toBe("notes");
});
test("parseFile: invalid type/status fall back to defaults", () => {
const { meta } = writeDoc(
test("parseFile: dynamic type and status with defaults", () => {
const { meta: m1 } = writeDoc(
"docs/x.md",
"---\ntitle: X\ntype: bogus\nstatus: bogus\n---\n",
"---\ntitle: X\ntype: custom-type\nstatus: draft\n---\n",
);
expect(meta.type).toBe("documentation");
expect(meta.status).toBe("published");
expect(m1.type).toBe("custom-type");
expect(m1.status).toBe("draft");
const { meta: m2 } = writeDoc(
"docs/y.md",
"---\ntitle: Y\n---\n",
);
expect(m2.type).toBe("documentation");
expect(m2.status).toBe("published");
});
test("parseFile: missing optional fields get sane defaults", () => {
+8 -13
View File
@@ -27,26 +27,21 @@ export function cosine(a: number[], b: number[]): number {
return denom === 0 ? 0 : dot / denom;
}
const VALID_SECTIONS: DocSection[] = ["docs", "writeups", "research", "notes"];
const VALID_TYPES: DocType[] = ["documentation", "writeup", "research", "note"];
/** Map a Drizzle row (text columns, Date timestamps) into the strict types. */
/** Map a Drizzle row into DocumentMeta. */
function toMeta(row: DocumentRow): DocumentMeta {
const extra = (row.extraFields ?? {}) as Record<string, unknown>;
return {
id: row.id,
slug: row.slug,
title: row.title,
type: (VALID_TYPES.includes(row.type as DocType) ? row.type : "documentation") as DocType,
section: (VALID_SECTIONS.includes(row.section as DocSection)
? row.section
: "docs") as DocSection,
type: row.type || "documentation",
section: row.section || "docs",
status: (row.status === "draft" ? "draft" : "published") as DocStatus,
author: row.author,
tags: row.tags,
path: row.path,
createdAt: row.createdAt.toISOString(),
updatedAt: row.updatedAt.toISOString(),
author: row.author || "",
tags: row.tags || [],
path: row.path || `${row.slug}.md`,
createdAt: row.createdAt ? row.createdAt.toISOString() : new Date().toISOString(),
updatedAt: row.updatedAt ? row.updatedAt.toISOString() : new Date().toISOString(),
extraFields: extra,
// Spread dynamic extra fields (CTF: event, challenge, category, difficulty, points, etc.)
...extra,
+15 -4
View File
@@ -1,6 +1,15 @@
export type DocSection = "docs" | "writeups" | "research" | "notes";
export type DocType = "documentation" | "writeup" | "research" | "note";
export type DocStatus = "published" | "draft";
export type DocSection = string;
export type DocType = string;
export type DocStatus = "published" | "draft" | string;
export interface SectionInfo {
id: string;
label: string;
icon: string;
desc: string;
docCount: number;
updatedAt?: string;
}
export interface DocumentMeta {
id: string; // slug
@@ -18,10 +27,11 @@ export interface DocumentMeta {
// difficulty, points). Content creators add arbitrary key-value pairs.
// Values can be strings, numbers, booleans, arrays, or objects.
extraFields?: Record<string, unknown>;
[key: string]: unknown;
}
export interface Document extends DocumentMeta {
body: string; // raw markdown (read from disk or stored)
body: string; // markdown body
}
export interface SearchHit {
@@ -29,3 +39,4 @@ export interface SearchHit {
rank: number;
snippet: string;
}