Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
70 lines
2.8 KiB
Python
70 lines
2.8 KiB
Python
#!/usr/bin/env python3
|
|
"""Verify each team's SSH passwords actually work in the live containers.
|
|
|
|
state.json can look perfect while the container holds a different password —
|
|
set_ssh_passwords() races container boot and its failures are easy to miss.
|
|
This proves the binding from the INSIDE (per the skill rule: never trust
|
|
config, prove it with a real login).
|
|
|
|
python3 panel/verify_ssh_creds.py [teamIdx ...]
|
|
"""
|
|
import json
|
|
import subprocess
|
|
import sys
|
|
from concurrent.futures import ThreadPoolExecutor
|
|
from pathlib import Path
|
|
|
|
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
|
import teams as orch
|
|
|
|
def probe(user, pw, port, host="127.0.0.1"):
|
|
r = subprocess.run(
|
|
["sshpass", "-p", pw, "ssh",
|
|
"-o", "StrictHostKeyChecking=no", "-o", "UserKnownHostsFile=/dev/null",
|
|
"-o", "ConnectTimeout=8", "-o", "LogLevel=ERROR",
|
|
"-p", str(port), f"{user}@{host}", "whoami; hostname"],
|
|
capture_output=True, text=True, timeout=30)
|
|
out = (r.stdout or "").strip().splitlines()
|
|
return (r.returncode == 0 and len(out) >= 2, out, (r.stderr or "").strip()[:80])
|
|
|
|
def main():
|
|
want = [int(a) for a in sys.argv[1:]]
|
|
teams = [t for t in orch.list_teams() if not want or t["index"] in want]
|
|
for t in teams:
|
|
idx = t["index"]
|
|
names = [c["name"] for c in orch.enabled_challenges()]
|
|
jobs = []
|
|
for n in names:
|
|
if n not in (t.get("ports") or {}):
|
|
continue
|
|
jobs.append((n, t["ports"][n]["ssh"], t.get("chall_passwords", {}).get(n)))
|
|
ok = bad = 0
|
|
details = []
|
|
users = orch.challenge_ssh_users()
|
|
with ThreadPoolExecutor(max_workers=6) as ex:
|
|
futs = {ex.submit(probe, users.get(n, "root"), pw, port): n
|
|
for n, port, pw in jobs if pw}
|
|
for fut, n in futs.items():
|
|
good, out, err = fut.result()
|
|
if good:
|
|
ok += 1
|
|
# hostname must be <challenge>_teamN — proves the binding
|
|
details.append((n, out[1] if len(out) > 1 else "?"))
|
|
else:
|
|
bad += 1
|
|
details.append((n, f"FAIL {err}"))
|
|
print(f"team{idx} ({t.get('label')}): ssh {ok} ok / {bad} fail")
|
|
for n, info in details:
|
|
if info == "FAIL" or info.startswith("FAIL"):
|
|
print(f" {n}: {info}")
|
|
hosts = [i for n, i in details if not i.startswith("FAIL")]
|
|
mism = [(n, i) for n, i in details
|
|
if not i.startswith("FAIL") and not i.endswith(f"_team{idx}")]
|
|
if mism:
|
|
print(f" !! hostname mismatch (not _team{idx}): {mism}")
|
|
else:
|
|
print(f" all hostnames correct (e.g. {hosts[0] if hosts else '-'})")
|
|
|
|
if __name__ == "__main__":
|
|
main()
|