Files
Cyrene aa0bb45633 fix(sla): 100% fleet SLA (64/64) - per-challenge SSH login, phew checker, sidecar detection
Three independent root causes, all found by measuring instead of assuming:

1. SSH failed on 10/16 challenges while state.json looked perfect.
   Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
   XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
   set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
   used -> 'Permission denied' on every team. Registry gains a per-challenge
   'ssh_user'; chpasswd targets the real login and reports failures loudly.

2. phew SLA timed out on a healthy service, four bugs stacked:
   - chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
     set) and does a fresh Pailier keygen (~12 s) before printing its menu;
   - _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
     TextIOWrapper buffer and select() then blocked on data already in memory;
   - its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
     the '> ' the next call waited for -> a race that failed intermittently;
   - reaping killed chall.py it did not own: a blanket pkill -f, a
     snapshot-diff (concurrent sessions diff against the same pre-spawn set),
     and a class-level _children shared across uvicorn's thread pool. The child
     now prints its own pid so exactly one session is reaped.
   Also: ONE interactive session per check instead of five spawns (Paillier is
   randomized per ciphertext, not per process) - 5 keygens were the CPU load
   that starved the checks. And the 6 orphan single-node containers from the
   original deploy were removed; one held 58 leaked chall.py and drove load
   average 76 on 2 CPUs.

3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
   every service sets an explicit container_name, so it reported all 16 running
   challenges as missing and hid the one real gap (anti-alchemy-db, which has
   no container_name). Now reads container_name when present and falls back to
   the compose default otherwise.

Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.

Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
2026-09-26 15:37:31 +08:00

96 lines
3.3 KiB
Python

#!/usr/bin/env python3
"""Replay the Phew checker's exact interaction, timing every step.
Purpose: find WHICH read exceeds its budget. The receiver log only says
"Timeout waiting for '> '. Got so far: <empty>", which cannot distinguish a
slow keygen from a hang. This prints per-step wall time and the buffer state
at the moment of the timeout.
"""
import os
import select
import subprocess
import sys
import time
CONT = os.environ.get("PHEW_CONT", "phew_container_team1")
CMD = ["docker", "exec", "-i", "-e", "PYTHONUNBUFFERED=1", CONT,
"python3", "/home/ctfuser/chall/src/chall.py"]
def read_until(proc, needle, timeout):
"""Mirror of the checker's _read_until, but reporting the buffer."""
buf = ""
end = time.time() + timeout
raw = proc.stdout.buffer if hasattr(proc.stdout, "buffer") else proc.stdout
while needle not in buf:
left = end - time.time()
if left <= 0:
raise TimeoutError(f"timeout after {timeout}s, buffer={buf!r}")
r, _, _ = select.select([raw], [], [], min(left, 1.0))
if not r:
continue
chunk = raw.read1(4096) if hasattr(raw, "read1") else raw.read(4096)
if not chunk:
raise TimeoutError(f"EOF, buffer={buf!r}")
buf += chunk.decode(errors="replace")
return buf
def step(label, fn):
t0 = time.time()
try:
out = fn()
print(f" {label:<34} {time.time()-t0:6.2f}s ok")
return out
except TimeoutError as e:
print(f" {label:<34} {time.time()-t0:6.2f}s TIMEOUT {e}")
raise
def main():
proc = subprocess.Popen(CMD, stdin=subprocess.PIPE, stdout=subprocess.PIPE,
stderr=subprocess.STDOUT, text=True, bufsize=0)
try:
print(f"container={CONT}")
step("boot -> first menu (budget 45s)", lambda: read_until(proc, "> ", 45))
for name, send, budget in (("encrypt", "1", 30), ("decrypt", "3", 30),
("key?", "4", 30)):
proc.stdin.write(send + "\n")
proc.stdin.flush()
step(f"send '{send}' -> prompt (budget 30s)",
lambda: read_until(proc, "> ", 30))
if name == "encrypt":
step(" send plaintext -> prompt",
lambda: read_until(proc, "> ", 30)) if False else None
proc.stdin.write("414243\n")
proc.stdin.flush()
step(" send plaintext -> prompt",
lambda: read_until(proc, "> ", 30))
elif name == "key?":
proc.stdin.write("\n")
proc.stdin.flush()
step(" send blank -> prompt",
lambda: read_until(proc, "> ", 30))
else:
proc.stdin.write("0\n")
proc.stdin.flush()
step(" send ct -> prompt",
lambda: read_until(proc, "> ", 30))
print("ALL STEPS WITHIN BUDGET")
except TimeoutError:
print("=> a step needs a bigger budget (or a different prompt)")
raise SystemExit(1)
finally:
try:
subprocess.run(["docker", "exec", CONT, "pkill", "-f", "chall.py"],
capture_output=True, timeout=20)
except Exception:
pass
if proc.poll() is None:
proc.kill()
if __name__ == "__main__":
main()