- _hermes_api_post sends X-Hermes-Session-Id so each fork sync uses a
fresh gateway transcript (was: derived session reused 200k-token history
across upstream tips, making later runs pay growing context)
- SYNC_CLAUDE_TIMEOUT 2700→3600: a fresh session needs ~25min/120 calls
to resolve 17 files + verify + commit; 2700s cut the HTTP call while the
agent was still committing
- salvage-on-timeout: when the agent call ends early (timeout/transport)
but the merge is already committed (MERGE_HEAD gone), finish+push it
instead of aborting — verified live: 17 conflicts resolved by agent,
push ok, fork 27 ahead / 0 behind
- _sync_finish_merge blocks committing files that still carry conflict
markers (guard against git add of unresolved files)
- tests: 68/68 (session header, salvage completed merge, abort unfinished,
marker guard; fake-leak hardening in _install_git_fake)
Claude Code could not complete AI fixes / conflict resolution against this
host's provider setup: it hung 900s spawning an MCP server, then failed with
'body is JSON but not a Message' (Anthropic-Messages transport mismatch), then
exited 1 with empty stderr. The Hermes gateway already runs continuously with
the working 9router provider config and a full toolset, so drive it directly:
- POST http://127.0.0.1:8642/v1/chat/completions (OpenAI-compatible API server)
- bearer auth from API_SERVER_KEY (env or ~/.hermes/.env), overridable via
API_SERVER_URL
- model_options.max_turns caps a runaway run; 900s timeout for sync, 600s for
PR fixes
- every failure maps to an [INFRA] string so the existing skip-once logic works
- the agent commits locally; the WORKER pushes (agents must never push)
run_ai_fix no longer shells out to claude; it calls the API server, then pushes
the agent's commit itself and reports push failures explicitly. Sync call sites
keep their contract via _run_claude_sync -> _run_hermes_sync alias, with labels
renamed hermes_sync_conflicts / hermes_sync_quality.
Tests: 59/59 (8 new assertions exercise a real local HTTP round-trip: path,
bearer auth, OpenAI message shape, max_turns cap, HTTP-error/missing-key/
unreachable -> [INFRA]).
- --dry now touches nothing: no state file writes (run_upstream_sync uses a
throwaway state dict; sync_fork_repo refuses to mutate in dry), no Discord,
no push, no PR (protected-branch dry prints would-open instead). The earlier
dry run polluted /tmp/pr-queue-sync-state.json and posted a false 'skip'
notification — both are gone.
- Claude Code conflict/quality runs now use --strict-mcp-config (with
--mcp-config ''): --mcp-config '' alone still lets claude -p spawn MCP
servers from settings.json/managed/plugins (observed ouroboros mcp serve
hanging 15+ min with zero output until the 900s timeout). These calls only
read/edit a throwaway clone and run git — no MCP server is ever needed.
- tests: 51/51 (added dry-purity cases: clean-merge, conflict-failure).
Merge new upstream (parent) commits into every fork in the App installation,
gated by a per-repo interval (default 1h), inside the existing 5-minute tick
(STEP 0, max 2 forks/tick, oldest-first).
- Conflicted merges are resolved by Claude Code (merge-reconciler rules:
never wholesale --ours/--theirs, verify with the repo's own
typecheck+tests, commit --no-edit; Claude never pushes — harness does).
- Clean merges get a single Claude Code quality pass commit.
- Push path: owner PAT (gh CLI) first — the App lacks workflows:write and a
workflows-touching merge is rejected for the App token; App token fallback.
- Protected default branch: detected from the push result (GH006 /
required-status-check) → upstream-sync-<ts> branch + PR through the normal
pipeline; duplicate open sync PRs are skipped.
- CI safety: after a direct push, ticks verify the fork CI at our merge sha;
red CI at OUR merge (still the tip) → sha-guarded force-revert to
pre-merge sha + Discord notify; never reverts foreign commits.
- Discord: synced / PR opened / reverted / skipped-once on pr-agent-ops.
- merge_pr gains the same PAT fallback (a PR merge touching workflows is a
workflow-file push).
- CLI: --sync-status, --sync-only <repo> [--dry].
- Tests: scripts/test_pr_queue_sync.py (46 assertions, monkeypatched, no
network); py_compile clean.
- Plan: .hermes/plans/2026-09-21-upstream-auto-sync.md
1. Skip ONCE on infra errors (Claude Code CLI missing/timeout/unreachable):
- run_ai_fix returns '[INFRA] ...' reasons; worker marks the PR permanently
skipped at that head SHA in fix-state (no more retry every 5 min)
- skip is recorded per {repo,pr,sha}; cleared when head SHA changes
2. Merge conflict auto-fix via Claude Code:
- mergeable=False or merge HTTP 409 now trigger run_ai_fix (prompt already
merges base + resolves conflicts) once per head SHA
- success → next tick re-checks mergeable and merges; failure → skip once
3. Discord skip notification dedupe:
- notify_skip_once(): posts '⏭️ Skipped: <reason>' exactly once per
PR+head_sha (state.notified flag); no repeated spam every cron tick
- skip reason + which PR is visible in the notification
4. State migration: legacy {repo:{pr:'sha'}} → dict form handled in _pr_entry
Verified: py_compile clean, state-helper unit tests pass (skip/fixed/notify
dedupe/legacy migration). Cron wrapper execs repo copy — no manual sync.
Discord embeds render markdown, not HTML — the PR Reviewer Guide table
(<table><tr><td>…) was appearing as literal HTML in the webhook message.
Added htmlToDiscordPlain(): collapses the table into readable lines
(score/effort/security/key-issues), keeps emoji + **bold** markdown, and
decodes entities. Also fixed review score suffix '/10' → '/100' (the LLM
score scale is 0-100, matching the table's 'Score: 72').
The LLM review score comes from the PR Reviewer Guide table which uses a
0-100 scale (72, 78, 100...). The Discord notification appended '/10',
making it read 'score 72/10'. Fixed to '/100' to match the actual scale.
- cli.ts: key resolution now falls back to ~/.hermes/keys/pr-agent-key.pem
when /opt/pr-agent-server/private-key.pem is EACCES/ENOENT (CLI as
non-root user works out of the box)
- cli.ts+index.ts: export logReviewEvent; CLI now records an analytics
event (describe/improve/review) in the legacy pr-agent.*.log format,
best-effort (never fails the CLI on a log write), honoring
PR_AGENT_ANALYTICS_DIR
- Verified: 16/16 tests, tsc clean, describe publishes, review publishes
(comment 5757282509), webhook 403/ping/ignored paths correct
- pr-queue-worker.py: cron orchestrator (review trigger → AI fix → safety →
CI gate → approve/merge) now versioned in-repo
- TOOLCHAIN_PINS: close dependabot PRs bumping pinned majors
(typescript/eslint/@tsparticles/eslint-config-next/eslint-plugin-react)
- STALE_CI_CLOSE_DAYS=2: close dependabot PRs stuck failing CI
- Secrets externalized to env (PR_AGENT_*), hydrated from ~/.hermes/.env —
file is safe for the public repo; no inline secrets
File is pr-agent:pr-agent 0600 — code user can't read directly.
Added sudo -n cat fallback for the on-disk key file, matching the
existing pattern used for /etc/bws-token. Key resolution order:
1. Direct read (works when gateway has bws group)
2. sudo -n cat (works with NOPASSWD sudo)
3. BWS CLI fallback
Root cause: health-check binary runs inside a Nix venv that doesn't have
/usr/local/bin/bws on PATH. When the shell wrapper's BWS_ACCESS_TOKEN
export fails (e.g. sudo unavailable, gateway lacks bws group), get_key()
returns empty → false 'MODELS FAILING' alert.
Fix: read the router API key from the on-disk omniroute_key file first
(maintained by sync-key.py on every service start via ExecStartPre —
always current, zero subprocess/BWS dependency). Fall back to BWS CLI
only if the file is missing/stale.
Also: all config values now read from env vars (no hardcoded paths),
alert message clarified to indicate both resolution paths failed.
Patched run_server.py (ANTHROPIC_API_* routing for 9router, bare claude-opus-5)
lives in /opt. Nix venv (h4bkq...) reused but execs /opt/run_server.py.
LD_LIBRARY_PATH needed for libstdc++ (litellm tokenizers crate).
9router/omniroute reject all provider prefixes (openai/claude-* -> 404).
Model must be bare (claude-opus-5). litellm routes bare claude-* to the
Anthropic native provider, so we set ANTHROPIC_API_BASE/ANTHROPIC_API_KEY
env vars pointing at 9router instead of the OpenAI-shaped OPENAI__* env.
Bug: PR-Agent auto-review gagal karena model 'openai/claude-opus-4-8'
tidak valid di 9router (provider openai/ tidak ada).
Root cause: BWS secret pr_agent_pr_agent_model berisi prefix openai/
yang hanya valid untuk omniroute, bukan 9router. 9router pakai model
tanpa provider prefix (e.g. claude-opus-5).
Fix:
- Default CONFIG__MODEL: openai/claude-opus-5 -> claude-opus-5
- Fallback claude-sonnet-5 -> claude-sonnet-5 (drop openai/ prefix)
BWS secret juga sudah diupdate di production.
Root cause: 9router combo models (deepseek-v4-flash-free on fallback) have
TTFT up to 30-40s. Caddy 9router route inherited the default
response_header_timeout 30s / read_timeout 60s → 504 'timeout awaiting
response headers' even though 9router was still processing. Cloudflare/log
showed repeated 504s; health watchdog (correctly) flagged the outage.
Fixes:
1. Caddy: dedicated 9router route with response_header_timeout 120s +
read/write 300s (was default 30/60). Removed invalid top-level
flush_interval on upload block that broke caddy reload (2.11 rejects it as
transport subdirective).
2. health-check: HTTP timeout 60→150s (mirror Caddy), and alert ONLY when
EVERY model fails — any working model means the server's fallback chain
succeeds. Early-exit on first success to bound runtime (~3s healthy).
Verified: 3 runs green, ~3.6s each, silent exit 0.
Cron runs as user code (not root). /etc/bws-token is root:bws 640, so direct
read fails with Permission denied → watchdog exited 1 every run. Fix:
- wrapper uses sudo -n cat (code is in sudo group, NOPASSWD)
- health-check.py get_key() falls back to sudo -n cat too
Verified as code user: silent exit 0 when healthy.
Real PR-Agent analytics logs wrap fields under 'record': {...}. The parser
now unwraps that before extracting command/pr_url/message/level, so
/api/analytics and /api/metrics show real data (verified with actual format
from production logs).
openai/auto/best-coding and openai/auto/claude-sonnet return
'No active credentials for provider: auto' on 9router (broken upstream
key). Primary openai/claude-opus-4-8 + fallbacks now all verified
working via litellm against 9router.asepharyana.my.id.