Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
41 lines
1.4 KiB
Bash
41 lines
1.4 KiB
Bash
#!/usr/bin/env bash
|
|
# Full SLA sweep across every team, SEQUENTIALLY.
|
|
#
|
|
# Sequential is not optional: the receivers are sync Flask apps, so hitting
|
|
# several at once makes them contend for the same 2 CPUs and report false
|
|
# timeouts (a pitfall already documented, and re-violated once here).
|
|
set -uo pipefail
|
|
BASE=/opt/gemastik18-final
|
|
declare -A RPORT=( [1]=31080 [2]=32080 [3]=33080 [4]=34080 )
|
|
|
|
echo "load: $(cut -d' ' -f1-3 /proc/loadavg)"
|
|
for t in 1 2 3 4; do
|
|
ENVF=$BASE/teams/team$t/receiver/.env
|
|
[ -f "$ENVF" ] || { echo "=== team$t: no receiver env ==="; continue; }
|
|
U=$(grep -oP '^ADMIN_USERNAME=\K.*' "$ENVF")
|
|
P=$(grep -oP '^ADMIN_PASSWORD=\K.*' "$ENVF")
|
|
port=${RPORT[$t]}
|
|
S=$(date +%s)
|
|
code=$(timeout 900 curl -sS -u "$U:$P" \
|
|
"http://127.0.0.1:$port/check/all" \
|
|
-o "/tmp/sla_t$t.json" -w '%{http_code}' 2>/dev/null)
|
|
took=$(( $(date +%s) - S ))
|
|
echo "=== team$t (receiver :$port, HTTP $code, ${took}s) ==="
|
|
python3 - "$t" <<'PY'
|
|
import json, sys
|
|
t = sys.argv[1]
|
|
try:
|
|
d = json.load(open(f"/tmp/sla_t{t}.json"))
|
|
except Exception as e:
|
|
print(" no/invalid result:", e); raise SystemExit
|
|
r = d.get("results", d)
|
|
if not isinstance(r, dict):
|
|
print(" ", json.dumps(d)[:300]); raise SystemExit
|
|
ok = sorted(k for k, v in r.items() if v is True)
|
|
bad = sorted(k for k, v in r.items() if v is not True)
|
|
print(f" PASS {len(ok)}/{len(r)}")
|
|
if ok: print(" ok :", ", ".join(ok))
|
|
if bad: print(" BAD:", ", ".join(bad))
|
|
PY
|
|
done
|