- pr-queue-worker.py: cron orchestrator (review trigger → AI fix → safety →
CI gate → approve/merge) now versioned in-repo
- TOOLCHAIN_PINS: close dependabot PRs bumping pinned majors
(typescript/eslint/@tsparticles/eslint-config-next/eslint-plugin-react)
- STALE_CI_CLOSE_DAYS=2: close dependabot PRs stuck failing CI
- Secrets externalized to env (PR_AGENT_*), hydrated from ~/.hermes/.env —
file is safe for the public repo; no inline secrets
File is pr-agent:pr-agent 0600 — code user can't read directly.
Added sudo -n cat fallback for the on-disk key file, matching the
existing pattern used for /etc/bws-token. Key resolution order:
1. Direct read (works when gateway has bws group)
2. sudo -n cat (works with NOPASSWD sudo)
3. BWS CLI fallback
Root cause: health-check binary runs inside a Nix venv that doesn't have
/usr/local/bin/bws on PATH. When the shell wrapper's BWS_ACCESS_TOKEN
export fails (e.g. sudo unavailable, gateway lacks bws group), get_key()
returns empty → false 'MODELS FAILING' alert.
Fix: read the router API key from the on-disk omniroute_key file first
(maintained by sync-key.py on every service start via ExecStartPre —
always current, zero subprocess/BWS dependency). Fall back to BWS CLI
only if the file is missing/stale.
Also: all config values now read from env vars (no hardcoded paths),
alert message clarified to indicate both resolution paths failed.
Patched run_server.py (ANTHROPIC_API_* routing for 9router, bare claude-opus-5)
lives in /opt. Nix venv (h4bkq...) reused but execs /opt/run_server.py.
LD_LIBRARY_PATH needed for libstdc++ (litellm tokenizers crate).
9router/omniroute reject all provider prefixes (openai/claude-* -> 404).
Model must be bare (claude-opus-5). litellm routes bare claude-* to the
Anthropic native provider, so we set ANTHROPIC_API_BASE/ANTHROPIC_API_KEY
env vars pointing at 9router instead of the OpenAI-shaped OPENAI__* env.
Bug: PR-Agent auto-review gagal karena model 'openai/claude-opus-4-8'
tidak valid di 9router (provider openai/ tidak ada).
Root cause: BWS secret pr_agent_pr_agent_model berisi prefix openai/
yang hanya valid untuk omniroute, bukan 9router. 9router pakai model
tanpa provider prefix (e.g. claude-opus-5).
Fix:
- Default CONFIG__MODEL: openai/claude-opus-5 -> claude-opus-5
- Fallback claude-sonnet-5 -> claude-sonnet-5 (drop openai/ prefix)
BWS secret juga sudah diupdate di production.
Root cause: 9router combo models (deepseek-v4-flash-free on fallback) have
TTFT up to 30-40s. Caddy 9router route inherited the default
response_header_timeout 30s / read_timeout 60s → 504 'timeout awaiting
response headers' even though 9router was still processing. Cloudflare/log
showed repeated 504s; health watchdog (correctly) flagged the outage.
Fixes:
1. Caddy: dedicated 9router route with response_header_timeout 120s +
read/write 300s (was default 30/60). Removed invalid top-level
flush_interval on upload block that broke caddy reload (2.11 rejects it as
transport subdirective).
2. health-check: HTTP timeout 60→150s (mirror Caddy), and alert ONLY when
EVERY model fails — any working model means the server's fallback chain
succeeds. Early-exit on first success to bound runtime (~3s healthy).
Verified: 3 runs green, ~3.6s each, silent exit 0.
Cron runs as user code (not root). /etc/bws-token is root:bws 640, so direct
read fails with Permission denied → watchdog exited 1 every run. Fix:
- wrapper uses sudo -n cat (code is in sudo group, NOPASSWD)
- health-check.py get_key() falls back to sudo -n cat too
Verified as code user: silent exit 0 when healthy.
Real PR-Agent analytics logs wrap fields under 'record': {...}. The parser
now unwraps that before extracting command/pr_url/message/level, so
/api/analytics and /api/metrics show real data (verified with actual format
from production logs).
openai/auto/best-coding and openai/auto/claude-sonnet return
'No active credentials for provider: auto' on 9router (broken upstream
key). Primary openai/claude-opus-4-8 + fallbacks now all verified
working via litellm against 9router.asepharyana.my.id.