Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
96 lines
3.3 KiB
Python
96 lines
3.3 KiB
Python
#!/usr/bin/env python3
|
|
"""Replay the Phew checker's exact interaction, timing every step.
|
|
|
|
Purpose: find WHICH read exceeds its budget. The receiver log only says
|
|
"Timeout waiting for '> '. Got so far: <empty>", which cannot distinguish a
|
|
slow keygen from a hang. This prints per-step wall time and the buffer state
|
|
at the moment of the timeout.
|
|
"""
|
|
import os
|
|
import select
|
|
import subprocess
|
|
import sys
|
|
import time
|
|
|
|
CONT = os.environ.get("PHEW_CONT", "phew_container_team1")
|
|
CMD = ["docker", "exec", "-i", "-e", "PYTHONUNBUFFERED=1", CONT,
|
|
"python3", "/home/ctfuser/chall/src/chall.py"]
|
|
|
|
|
|
def read_until(proc, needle, timeout):
|
|
"""Mirror of the checker's _read_until, but reporting the buffer."""
|
|
buf = ""
|
|
end = time.time() + timeout
|
|
raw = proc.stdout.buffer if hasattr(proc.stdout, "buffer") else proc.stdout
|
|
while needle not in buf:
|
|
left = end - time.time()
|
|
if left <= 0:
|
|
raise TimeoutError(f"timeout after {timeout}s, buffer={buf!r}")
|
|
r, _, _ = select.select([raw], [], [], min(left, 1.0))
|
|
if not r:
|
|
continue
|
|
chunk = raw.read1(4096) if hasattr(raw, "read1") else raw.read(4096)
|
|
if not chunk:
|
|
raise TimeoutError(f"EOF, buffer={buf!r}")
|
|
buf += chunk.decode(errors="replace")
|
|
return buf
|
|
|
|
|
|
def step(label, fn):
|
|
t0 = time.time()
|
|
try:
|
|
out = fn()
|
|
print(f" {label:<34} {time.time()-t0:6.2f}s ok")
|
|
return out
|
|
except TimeoutError as e:
|
|
print(f" {label:<34} {time.time()-t0:6.2f}s TIMEOUT {e}")
|
|
raise
|
|
|
|
|
|
def main():
|
|
proc = subprocess.Popen(CMD, stdin=subprocess.PIPE, stdout=subprocess.PIPE,
|
|
stderr=subprocess.STDOUT, text=True, bufsize=0)
|
|
try:
|
|
print(f"container={CONT}")
|
|
step("boot -> first menu (budget 45s)", lambda: read_until(proc, "> ", 45))
|
|
|
|
for name, send, budget in (("encrypt", "1", 30), ("decrypt", "3", 30),
|
|
("key?", "4", 30)):
|
|
proc.stdin.write(send + "\n")
|
|
proc.stdin.flush()
|
|
step(f"send '{send}' -> prompt (budget 30s)",
|
|
lambda: read_until(proc, "> ", 30))
|
|
if name == "encrypt":
|
|
step(" send plaintext -> prompt",
|
|
lambda: read_until(proc, "> ", 30)) if False else None
|
|
proc.stdin.write("414243\n")
|
|
proc.stdin.flush()
|
|
step(" send plaintext -> prompt",
|
|
lambda: read_until(proc, "> ", 30))
|
|
elif name == "key?":
|
|
proc.stdin.write("\n")
|
|
proc.stdin.flush()
|
|
step(" send blank -> prompt",
|
|
lambda: read_until(proc, "> ", 30))
|
|
else:
|
|
proc.stdin.write("0\n")
|
|
proc.stdin.flush()
|
|
step(" send ct -> prompt",
|
|
lambda: read_until(proc, "> ", 30))
|
|
print("ALL STEPS WITHIN BUDGET")
|
|
except TimeoutError:
|
|
print("=> a step needs a bigger budget (or a different prompt)")
|
|
raise SystemExit(1)
|
|
finally:
|
|
try:
|
|
subprocess.run(["docker", "exec", CONT, "pkill", "-f", "chall.py"],
|
|
capture_output=True, timeout=20)
|
|
except Exception:
|
|
pass
|
|
if proc.poll() is None:
|
|
proc.kill()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|