fix(sla): 100% fleet SLA (64/64) - per-challenge SSH login, phew checker, sidecar detection

Three independent root causes, all found by measuring instead of assuming:

1. SSH failed on 10/16 challenges while state.json looked perfect.
   Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
   XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
   set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
   used -> 'Permission denied' on every team. Registry gains a per-challenge
   'ssh_user'; chpasswd targets the real login and reports failures loudly.

2. phew SLA timed out on a healthy service, four bugs stacked:
   - chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
     set) and does a fresh Pailier keygen (~12 s) before printing its menu;
   - _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
     TextIOWrapper buffer and select() then blocked on data already in memory;
   - its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
     the '> ' the next call waited for -> a race that failed intermittently;
   - reaping killed chall.py it did not own: a blanket pkill -f, a
     snapshot-diff (concurrent sessions diff against the same pre-spawn set),
     and a class-level _children shared across uvicorn's thread pool. The child
     now prints its own pid so exactly one session is reaped.
   Also: ONE interactive session per check instead of five spawns (Paillier is
   randomized per ciphertext, not per process) - 5 keygens were the CPU load
   that starved the checks. And the 6 orphan single-node containers from the
   original deploy were removed; one held 58 leaked chall.py and drove load
   average 76 on 2 CPUs.

3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
   every service sets an explicit container_name, so it reported all 16 running
   challenges as missing and hid the one real gap (anti-alchemy-db, which has
   no container_name). Now reads container_name when present and falls back to
   the compose default otherwise.

Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.

Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
This commit is contained in:
Cyrene
2026-09-26 15:37:31 +08:00
parent ae50acfe40
commit aa0bb45633
11 changed files with 589 additions and 145 deletions
+12
View File
@@ -0,0 +1,12 @@
#!/usr/bin/env bash
# Watch the host recover after the orphan single-node containers were removed.
# 2 CPUs + ~92 containers means the 6 leftovers (one with 58 leaked chall.py
# processes) were the dominant load source; SLA timeouts on a healthy service
# were a symptom of that, not of the service.
for i in 1 2 3 4 5 6; do
LOAD=$(cut -d' ' -f1-3 /proc/loadavg)
IDLE=$(vmstat 1 2 | tail -1 | awk '{print $15}')
CHALL=$(ps -eo args --no-headers | grep -c '[c]hall.py')
echo "t+$((i * 20))s load=$LOAD idle=${IDLE}% chall.py=$CHALL"
sleep 20
done