Commit Graph
34 Commits
Author SHA1 Message Date
asepharyana 99d47c8ff8 feat(worker): drive Hermes gateway API server instead of claude -p
Claude Code could not complete AI fixes / conflict resolution against this
host's provider setup: it hung 900s spawning an MCP server, then failed with
'body is JSON but not a Message' (Anthropic-Messages transport mismatch), then
exited 1 with empty stderr. The Hermes gateway already runs continuously with
the working 9router provider config and a full toolset, so drive it directly:

- POST http://127.0.0.1:8642/v1/chat/completions (OpenAI-compatible API server)
- bearer auth from API_SERVER_KEY (env or ~/.hermes/.env), overridable via
  API_SERVER_URL
- model_options.max_turns caps a runaway run; 900s timeout for sync, 600s for
  PR fixes
- every failure maps to an [INFRA] string so the existing skip-once logic works
- the agent commits locally; the WORKER pushes (agents must never push)

run_ai_fix no longer shells out to claude; it calls the API server, then pushes
the agent's commit itself and reports push failures explicitly. Sync call sites
keep their contract via _run_claude_sync -> _run_hermes_sync alias, with labels
renamed hermes_sync_conflicts / hermes_sync_quality.

Tests: 59/59 (8 new assertions exercise a real local HTTP round-trip: path,
bearer auth, OpenAI message shape, max_turns cap, HTTP-error/missing-key/
unreachable -> [INFRA]).
2026-09-21 18:24:47 +07:00
asepharyana 0653f6c508 fix(worker): pure --dry runs + Claude Code MCP-server hang
- --dry now touches nothing: no state file writes (run_upstream_sync uses a
  throwaway state dict; sync_fork_repo refuses to mutate in dry), no Discord,
  no push, no PR (protected-branch dry prints would-open instead). The earlier
  dry run polluted /tmp/pr-queue-sync-state.json and posted a false 'skip'
  notification — both are gone.
- Claude Code conflict/quality runs now use --strict-mcp-config (with
  --mcp-config ''): --mcp-config '' alone still lets claude -p spawn MCP
  servers from settings.json/managed/plugins (observed ouroboros mcp serve
  hanging 15+ min with zero output until the 900s timeout). These calls only
  read/edit a throwaway clone and run git — no MCP server is ever needed.
- tests: 51/51 (added dry-purity cases: clean-merge, conflict-failure).
2026-09-21 17:31:16 +07:00
asepharyana c48adea6c3 feat(worker): upstream fork auto-sync — pull+merge fork repos hourly
Merge new upstream (parent) commits into every fork in the App installation,
gated by a per-repo interval (default 1h), inside the existing 5-minute tick
(STEP 0, max 2 forks/tick, oldest-first).

- Conflicted merges are resolved by Claude Code (merge-reconciler rules:
  never wholesale --ours/--theirs, verify with the repo's own
  typecheck+tests, commit --no-edit; Claude never pushes — harness does).
- Clean merges get a single Claude Code quality pass commit.
- Push path: owner PAT (gh CLI) first — the App lacks workflows:write and a
  workflows-touching merge is rejected for the App token; App token fallback.
- Protected default branch: detected from the push result (GH006 /
  required-status-check) → upstream-sync-<ts> branch + PR through the normal
  pipeline; duplicate open sync PRs are skipped.
- CI safety: after a direct push, ticks verify the fork CI at our merge sha;
  red CI at OUR merge (still the tip) → sha-guarded force-revert to
  pre-merge sha + Discord notify; never reverts foreign commits.
- Discord: synced / PR opened / reverted / skipped-once on pr-agent-ops.
- merge_pr gains the same PAT fallback (a PR merge touching workflows is a
  workflow-file push).
- CLI: --sync-status, --sync-only <repo> [--dry].
- Tests: scripts/test_pr_queue_sync.py (46 assertions, monkeypatched, no
  network); py_compile clean.
- Plan: .hermes/plans/2026-09-21-upstream-auto-sync.md
2026-09-21 17:04:39 +07:00
asepharyana 1dc4d098fa feat(worker): skip-on-error no loop + conflict auto-fix + deduped notif
1. Skip ONCE on infra errors (Claude Code CLI missing/timeout/unreachable):
   - run_ai_fix returns '[INFRA] ...' reasons; worker marks the PR permanently
     skipped at that head SHA in fix-state (no more retry every 5 min)
   - skip is recorded per {repo,pr,sha}; cleared when head SHA changes
2. Merge conflict auto-fix via Claude Code:
   - mergeable=False or merge HTTP 409 now trigger run_ai_fix (prompt already
     merges base + resolves conflicts) once per head SHA
   - success → next tick re-checks mergeable and merges; failure → skip once
3. Discord skip notification dedupe:
   - notify_skip_once(): posts '⏭️ Skipped: <reason>' exactly once per
     PR+head_sha (state.notified flag); no repeated spam every cron tick
   - skip reason + which PR is visible in the notification
4. State migration: legacy {repo:{pr:'sha'}} → dict form handled in _pr_entry
Verified: py_compile clean, state-helper unit tests pass (skip/fixed/notify
dedupe/legacy migration). Cron wrapper execs repo copy — no manual sync.
2026-09-21 15:58:52 +07:00
asepharyana 296f825347 fix(discord): strip HTML table markup from review notifications
Discord embeds render markdown, not HTML — the PR Reviewer Guide table
(<table><tr><td>…) was appearing as literal HTML in the webhook message.
Added htmlToDiscordPlain(): collapses the table into readable lines
(score/effort/security/key-issues), keeps emoji + **bold** markdown, and
decodes entities. Also fixed review score suffix '/10' → '/100' (the LLM
score scale is 0-100, matching the table's 'Score: 72').
2026-09-21 15:38:21 +07:00
asepharyana 391743780d fix(discord): review score scale is 0-100 not /10
The LLM review score comes from the PR Reviewer Guide table which uses a
0-100 scale (72, 78, 100...). The Discord notification appended '/10',
making it read 'score 72/10'. Fixed to '/100' to match the actual scale.
2026-09-21 15:35:30 +07:00
asepharyana 4fa7d8b6d0 fix(cli): fallback to user-local private key + record analytics
- cli.ts: key resolution now falls back to ~/.hermes/keys/pr-agent-key.pem
  when /opt/pr-agent-server/private-key.pem is EACCES/ENOENT (CLI as
  non-root user works out of the box)
- cli.ts+index.ts: export logReviewEvent; CLI now records an analytics
  event (describe/improve/review) in the legacy pr-agent.*.log format,
  best-effort (never fails the CLI on a log write), honoring
  PR_AGENT_ANALYTICS_DIR
- Verified: 16/16 tests, tsc clean, describe publishes, review publishes
  (comment 5757282509), webhook 403/ping/ignored paths correct
2026-09-21 15:11:36 +07:00
asepharyana 570b5b8707 chore(cleanup): remove retired Python/Nix pr_agent server entirely
- delete src/*.py (run_server, start_server, callback_server, health-check,
  auto_merge_bot, trivial_merge, sync-key) — Python pr_agent retired
- delete scripts/{setup_app,setup_all,generate_manifest,generate_manifest_domain}.py
- delete flake.nix + flake.lock + flakehub-publish-rolling.yaml — Nix build retired
- deploy.yml: Nix/Python CI → Bun CI (setup-bun, typecheck, tests,
  bun build --compile → scp binary → swap /opt/.../pr-agent-bun →
  restart pr-agent-bun.service → health check :4023)
- README: document Bun era; legacy Python/Nix section
- pr-agent-bun.service is the sole production server (port 4023)
2026-09-21 14:49:45 +07:00
asepharyana 608b52a0e5 feat(server): port /describe and /improve tools + real analytics
Phase 8 (extras) complete:
- describe.ts — full port of pr_description: type/title/description/diagram/
  per-file walkthrough table, PATCH /pulls to update description, optional
  labels. Verified live: PR #19 body updated with AI description (Type,
  Description bullets, mermaid diagram, file walkthrough table).
- improve.ts — full port of pr_code_suggestions summarize path:
  getPrMultiDiffs chunking, parallel LLM calls, score filtering, category
  table with unified diff snippets + persistent comment with history.
  Verified live: persistent update of existing '## PR Code Suggestions ✨'
  comment (updated_at 07:26:53) with new table.
- prompts.ts — verbatim pr_description_prompts.toml + code_suggestions prompts
- github.ts — getLineLink (SHA-256 diff anchor like pr_agent), getLabels,
  updateDescription (pulls.update); getPrMultiDiffs in diff.ts
- index.ts — real analytics: readLegacy pr-agent.*.log files (mtime-sorted,
  fixes pid-filename ordering bug), logReviewEvent writes legacy-format
  events, /api/metrics + /api/analytics now serve real data (427 events,
  per-command breakdown)
- cli.ts — --tool review|describe|improve
- tests 16/16, tsc clean
2026-09-21 14:36:02 +07:00
asepharyana bc69998d8e feat(server): production cut-over to Bun — notify_review, setup/callback, key resolution
- index.ts: add POST /api/v1/notify_review (queue-worker Discord bridge),
  /setup/callback (GitHub App manifest conversion), error logging on review
  failure, PR_AGENT_APP_DIR-based private key resolution
- config.ts: API key falls back to on-disk omniroute_key (same source as
  run_server.py) when no env key present — fixes 401 in systemd context
- markdown.ts: don't hyperlink non-URL ticket values (N/A)
- secrets.ts: fallback private key from ~/.hermes/keys + omni key ~/.hermes
- pr-queue-worker.py / auto_merge_bot.py: notify_review default port 4002→4023

Deploy: pr-agent-bun.service (Bun binary, port 4023) replaces
pr-agent-server.service (Python, port 4002, disabled). Caddy route updated in
asepharyana/infra (proxy 4023). Verified live: webhook → review → claude-opus-5
→ persistent GitHub comment published after Python shutdown.
2026-09-21 13:47:47 +07:00
asepharyana 055501aed1 feat(server): Bun/TypeScript re-implementation of PR review engine
Reimplements the pr_agent review pipeline (previously Python package) in
Bun/TypeScript, verified end-to-end against a real GitHub App + 9router:

- src/diff.ts: git patch processing (extend_patch, line-numbered hunks,
  token-budget generate_full_patch, generated/invalid file filters)
- src/review.ts: orchestrator (fetch PR/diff -> render prompts -> LLM ->
  YAML parse -> markdown render -> publish), fallback models
- src/github.ts: GitHub App auth via @octokit/auth-app (JWT -> installation
  token, custom authStrategy for Bun), PR/diff/languages/comments REST
- src/prompts.ts: verbatim pr_reviewer_prompts.toml port (nunjucks templates)
- src/yaml.ts: resilient load_yaml with 9 fallback fixers
- src/markdown.ts: convert_to_markdown_v2 port ('PR Reviewer Guide' comment)
- src/llm.ts: 9router OpenAI-compatible chat completions (non-stream,
  temp omitted for claude-opus-5), timeout 600s
- src/token.ts: js-tiktoken o200k counting
- src/index.ts: webhook server (HMAC verify, segment analytics, /health,
  /api/metrics) + cli.ts one-shot review
- e2e.ts: real-world harness (bun e2e --repo o/r --pr N [--publish])

Verified: bun test 13/13, tsc clean, E2E published a real review comment
(## PR Reviewer Guide) on asepharyana/nextjs-template#19 via GitHub App
MythEclipseBotReview + claude-opus-5 through 9router.

Python run_server.py + pr_agent remain for the live queue worker; server/
is the replacement path.
2026-09-21 12:50:10 +07:00
asepharyana 835552c5ca feat(ops): add PR queue worker with toolchain pin guard
- pr-queue-worker.py: cron orchestrator (review trigger → AI fix → safety →
  CI gate → approve/merge) now versioned in-repo
- TOOLCHAIN_PINS: close dependabot PRs bumping pinned majors
  (typescript/eslint/@tsparticles/eslint-config-next/eslint-plugin-react)
- STALE_CI_CLOSE_DAYS=2: close dependabot PRs stuck failing CI
- Secrets externalized to env (PR_AGENT_*), hydrated from ~/.hermes/.env —
  file is safe for the public repo; no inline secrets
2026-09-21 11:36:14 +07:00
asepharyana d516a0b563 fix(health-check): handle PermissionError on on-disk key file
File is pr-agent:pr-agent 0600 — code user can't read directly.
Added sudo -n cat fallback for the on-disk key file, matching the
existing pattern used for /etc/bws-token. Key resolution order:
1. Direct read (works when gateway has bws group)
2. sudo -n cat (works with NOPASSWD sudo)
3. BWS CLI fallback
2026-09-09 21:09:02 +07:00
asepharyana ae5356a4ea fix(health-check): read key from on-disk file first, no BWS dependency
Root cause: health-check binary runs inside a Nix venv that doesn't have
/usr/local/bin/bws on PATH. When the shell wrapper's BWS_ACCESS_TOKEN
export fails (e.g. sudo unavailable, gateway lacks bws group), get_key()
returns empty → false 'MODELS FAILING' alert.

Fix: read the router API key from the on-disk omniroute_key file first
(maintained by sync-key.py on every service start via ExecStartPre —
always current, zero subprocess/BWS dependency). Fall back to BWS CLI
only if the file is missing/stale.

Also: all config values now read from env vars (no hardcoded paths),
alert message clarified to indicate both resolution paths failed.
2026-09-09 21:05:10 +07:00
asepharyana bc8e1739e9 refactor: restructure into proper project layout + improve docs
Project layout:
- src/: application modules (run_server, auto_merge_bot, health-check, sync-key, trivial_merge, callback_server, start_server)
- scripts/: setup/deployment helpers (setup_all, setup_app, generate_manifest)
- templates/: manifest.json (GitHub App manifest template)
- docs/  + CONTRIBUTING.md: documentation

Improvements:
- flake.nix: added pr-agent-auto-merge wrapper binary, updated installPhase paths
- deploy.yml: syntax check covers all modules including health-check.py and sync-key.py
- README.md: comprehensive with architecture, layout, dev, ops, deployment
- CONTRIBUTING.md: standards and testing checklist
- .gitignore: added *.log, *.pid, .env.*
- Cleanup: removed duplicate manifest_current.json / manifest_final.json
- Fix: health-check.py docstring updated to claude-opus-5

Verification:
- ✅ python3 -m py_compile: all 11 modules pass
- ✅ nix flake check: passes
2026-08-20 11:38:24 +07:00
asepharyana 017656d97b feat: update models to claude-opus-5/sonnet-5/haiku-4-5-20251001 (tested live on 9router)
- PRIMARY: openai/claude-opus-5 (was openai/claude-opus-4-8)
- FALLBACKS: added openai/claude-sonnet-5, openai/claude-haiku-4-5-20251001
- Verified all three return valid responses via raw HTTP to 9router
- Updated run_server.py, health-check.py, setup_all.py
2026-08-20 11:30:10 +07:00
asepharyana 54cec6bf0c improve: add README, .editorconfig, fix fallback models in setup_all.py, update .gitignore 2026-08-20 11:26:03 +07:00
asepharyana 2293428421 fix: health watchdog false 504s — Caddy header timeout + per-model logic
Root cause: 9router combo models (deepseek-v4-flash-free on fallback) have
TTFT up to 30-40s. Caddy 9router route inherited the default
response_header_timeout 30s / read_timeout 60s → 504 'timeout awaiting
response headers' even though 9router was still processing. Cloudflare/log
showed repeated 504s; health watchdog (correctly) flagged the outage.

Fixes:
1. Caddy: dedicated 9router route with response_header_timeout 120s +
   read/write 300s (was default 30/60). Removed invalid top-level
   flush_interval on upload block that broke caddy reload (2.11 rejects it as
   transport subdirective).
2. health-check: HTTP timeout 60→150s (mirror Caddy), and alert ONLY when
   EVERY model fails — any working model means the server's fallback chain
   succeeds. Early-exit on first success to bound runtime (~3s healthy).
Verified: 3 runs green, ~3.6s each, silent exit 0.
2026-08-04 21:06:54 +07:00
asepharyana bd3d739250 ci: add Nix GC cleanup job on VPS after deploy 2026-08-04 13:57:47 +07:00
asepharyana 875ddfcca9 fix: health-check BWS token read as non-root cron user
Cron runs as user code (not root). /etc/bws-token is root:bws 640, so direct
read fails with Permission denied → watchdog exited 1 every run. Fix:
- wrapper uses sudo -n cat (code is in sudo group, NOPASSWD)
- health-check.py get_key() falls back to sudo -n cat too
Verified as code user: silent exit 0 when healthy.
2026-08-04 12:12:42 +07:00
asepharyana 407819b647 fix: parse PR-Agent analytics record-wrapped JSON format
Real PR-Agent analytics logs wrap fields under 'record': {...}. The parser
now unwraps that before extracting command/pr_url/message/level, so
/api/analytics and /api/metrics show real data (verified with actual format
from production logs).
2026-08-04 11:43:13 +07:00
asepharyana efb99b73b7 chore: gitignore nix result symlink 2026-08-04 11:39:34 +07:00
asepharyana 690f96aacb feat: add health watchdog, key auto-sync, analytics, Discord notifications, trivial-PR merge
- health-check.py: model health watchdog (silent when healthy, alert on 2+
  consecutive failures) — catches stale-key/model-breakage like 2026-08-04
- sync-key.py: systemd ExecStartPre syncs 9router key from BWS to disk,
  prevents silent 401s after key rotation
- run_server.py: CONFIG__ANALYTICS_FOLDER + /api/metrics (Prometheus) +
  /api/analytics (JSON) + /api/v1/notify_review (Discord webhook)
- trivial_merge.py: trivial-PR fast-path (docs/deps/tiny diffs) approve+merge
- auto_merge_bot.py: Discord notifications on merge, trivial integration
- flake.nix: ship pr-agent-sync-key + pr-agent-health-check binaries
2026-08-04 11:39:29 +07:00
asepharyana 06f3314819 fix: replace broken openai/auto fallback models with working 9router aliases
openai/auto/best-coding and openai/auto/claude-sonnet return
'No active credentials for provider: auto' on 9router (broken upstream
key). Primary openai/claude-opus-4-8 + fallbacks now all verified
working via litellm against 9router.asepharyana.my.id.
2026-08-04 10:17:29 +07:00
aseph ca67f0328b ci: use free GHA Nix cache (disable FlakeHub cache, not subscribed) 2026-08-03 16:44:18 +07:00
asepharyana fa0e031ba9 ci: enable FlakeHub Cache (id-token: write + use-flakehub) 2026-08-03 16:22:48 +07:00
asepharyana e0e591a63c fix: unterminated triple-quote in setup_all.py secrets template 2026-08-03 13:39:36 +07:00
asepharyana eb97e020dd ci: add python syntax-check gate before Nix deploy 2026-08-03 13:37:29 +07:00
asepharyana 86ff584b15 chore: webhook URLs to domain, manifest.json sync 2026-08-02 16:22:32 +07:00
asepharyana b9ba5ad9fa chore: use domain for webhook URLs 2026-08-02 16:21:47 +07:00
asepharyana 637b44a9e0 chore: port 3000 to 4002, use domain instead of IP 2026-08-02 16:15:43 +07:00
asepharyana 081b2ad4b5 fix(nix): restrict flake to x86_64-linux (nixpkgs 26.11 dropped darwin) 2026-08-01 18:03:43 +07:00
asepharyana eadc674565 ci: publish flake to FlakeHub (rolling) 2026-08-01 17:58:21 +07:00
asepharyana 18cf0d26aa ci: migrate CI to GitHub Actions (deploy nix + mirror ke Gitea backup)
Build & Deploy (Nix) / build-and-deploy (push) Failing after 32m50s
Mirror to Gitea / mirror (push) Successful in 9s
2026-08-01 16:41:46 +07:00