1. Skip ONCE on infra errors (Claude Code CLI missing/timeout/unreachable):
- run_ai_fix returns '[INFRA] ...' reasons; worker marks the PR permanently
skipped at that head SHA in fix-state (no more retry every 5 min)
- skip is recorded per {repo,pr,sha}; cleared when head SHA changes
2. Merge conflict auto-fix via Claude Code:
- mergeable=False or merge HTTP 409 now trigger run_ai_fix (prompt already
merges base + resolves conflicts) once per head SHA
- success → next tick re-checks mergeable and merges; failure → skip once
3. Discord skip notification dedupe:
- notify_skip_once(): posts '⏭️ Skipped: <reason>' exactly once per
PR+head_sha (state.notified flag); no repeated spam every cron tick
- skip reason + which PR is visible in the notification
4. State migration: legacy {repo:{pr:'sha'}} → dict form handled in _pr_entry
Verified: py_compile clean, state-helper unit tests pass (skip/fixed/notify
dedupe/legacy migration). Cron wrapper execs repo copy — no manual sync.
Discord embeds render markdown, not HTML — the PR Reviewer Guide table
(<table><tr><td>…) was appearing as literal HTML in the webhook message.
Added htmlToDiscordPlain(): collapses the table into readable lines
(score/effort/security/key-issues), keeps emoji + **bold** markdown, and
decodes entities. Also fixed review score suffix '/10' → '/100' (the LLM
score scale is 0-100, matching the table's 'Score: 72').
The LLM review score comes from the PR Reviewer Guide table which uses a
0-100 scale (72, 78, 100...). The Discord notification appended '/10',
making it read 'score 72/10'. Fixed to '/100' to match the actual scale.
- cli.ts: key resolution now falls back to ~/.hermes/keys/pr-agent-key.pem
when /opt/pr-agent-server/private-key.pem is EACCES/ENOENT (CLI as
non-root user works out of the box)
- cli.ts+index.ts: export logReviewEvent; CLI now records an analytics
event (describe/improve/review) in the legacy pr-agent.*.log format,
best-effort (never fails the CLI on a log write), honoring
PR_AGENT_ANALYTICS_DIR
- Verified: 16/16 tests, tsc clean, describe publishes, review publishes
(comment 5757282509), webhook 403/ping/ignored paths correct
- pr-queue-worker.py: cron orchestrator (review trigger → AI fix → safety →
CI gate → approve/merge) now versioned in-repo
- TOOLCHAIN_PINS: close dependabot PRs bumping pinned majors
(typescript/eslint/@tsparticles/eslint-config-next/eslint-plugin-react)
- STALE_CI_CLOSE_DAYS=2: close dependabot PRs stuck failing CI
- Secrets externalized to env (PR_AGENT_*), hydrated from ~/.hermes/.env —
file is safe for the public repo; no inline secrets
File is pr-agent:pr-agent 0600 — code user can't read directly.
Added sudo -n cat fallback for the on-disk key file, matching the
existing pattern used for /etc/bws-token. Key resolution order:
1. Direct read (works when gateway has bws group)
2. sudo -n cat (works with NOPASSWD sudo)
3. BWS CLI fallback
Root cause: health-check binary runs inside a Nix venv that doesn't have
/usr/local/bin/bws on PATH. When the shell wrapper's BWS_ACCESS_TOKEN
export fails (e.g. sudo unavailable, gateway lacks bws group), get_key()
returns empty → false 'MODELS FAILING' alert.
Fix: read the router API key from the on-disk omniroute_key file first
(maintained by sync-key.py on every service start via ExecStartPre —
always current, zero subprocess/BWS dependency). Fall back to BWS CLI
only if the file is missing/stale.
Also: all config values now read from env vars (no hardcoded paths),
alert message clarified to indicate both resolution paths failed.
Root cause: 9router combo models (deepseek-v4-flash-free on fallback) have
TTFT up to 30-40s. Caddy 9router route inherited the default
response_header_timeout 30s / read_timeout 60s → 504 'timeout awaiting
response headers' even though 9router was still processing. Cloudflare/log
showed repeated 504s; health watchdog (correctly) flagged the outage.
Fixes:
1. Caddy: dedicated 9router route with response_header_timeout 120s +
read/write 300s (was default 30/60). Removed invalid top-level
flush_interval on upload block that broke caddy reload (2.11 rejects it as
transport subdirective).
2. health-check: HTTP timeout 60→150s (mirror Caddy), and alert ONLY when
EVERY model fails — any working model means the server's fallback chain
succeeds. Early-exit on first success to bound runtime (~3s healthy).
Verified: 3 runs green, ~3.6s each, silent exit 0.
Cron runs as user code (not root). /etc/bws-token is root:bws 640, so direct
read fails with Permission denied → watchdog exited 1 every run. Fix:
- wrapper uses sudo -n cat (code is in sudo group, NOPASSWD)
- health-check.py get_key() falls back to sudo -n cat too
Verified as code user: silent exit 0 when healthy.
Real PR-Agent analytics logs wrap fields under 'record': {...}. The parser
now unwraps that before extracting command/pr_url/message/level, so
/api/analytics and /api/metrics show real data (verified with actual format
from production logs).
openai/auto/best-coding and openai/auto/claude-sonnet return
'No active credentials for provider: auto' on 9router (broken upstream
key). Primary openai/claude-opus-4-8 + fallbacks now all verified
working via litellm against 9router.asepharyana.my.id.