Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
36 lines
1.2 KiB
Bash
36 lines
1.2 KiB
Bash
#!/usr/bin/env bash
|
|
# Concurrency regression test for the Phew checker.
|
|
#
|
|
# The bug this guards against: the checker reaped chall.py processes it did not
|
|
# own, so two overlapping checks killed each other's session and the victim
|
|
# reported "Process ended while waiting for '> '" on a healthy service. Firing
|
|
# several checks at once is the only way to reproduce it — sequential runs pass
|
|
# even with the bug present.
|
|
set -uo pipefail
|
|
BASE=/opt/gemastik18-final
|
|
TEAM=${1:-1}
|
|
N=${2:-5}
|
|
case "$TEAM" in 1) PORT=31080;; 2) PORT=32080;; 3) PORT=33080;; 4) PORT=34080;; esac
|
|
U=$(grep -oP '^ADMIN_USERNAME=\K.*' "$BASE/teams/team$TEAM/receiver/.env")
|
|
P=$(grep -oP '^ADMIN_PASSWORD=\K.*' "$BASE/teams/team$TEAM/receiver/.env")
|
|
|
|
count() { docker exec "phew_container_team$TEAM" sh -c 'ps ax | grep -c "[c]hall.py"'; }
|
|
echo "before: $(count) chall.py (1 = socat service only)"
|
|
|
|
pids=()
|
|
for i in $(seq 1 "$N"); do
|
|
( out=$(timeout 300 curl -sS -u "$U:$P" "http://127.0.0.1:$PORT/check/phew")
|
|
echo " req$i: $out" ) &
|
|
pids+=($!)
|
|
done
|
|
for p in "${pids[@]}"; do wait "$p"; done
|
|
|
|
sleep 5
|
|
after=$(count)
|
|
echo "after: $after chall.py"
|
|
if [ "$after" -le 1 ]; then
|
|
echo "RESULT: PASS (no leak)"
|
|
else
|
|
echo "RESULT: FAIL (leaked $((after - 1)))"
|
|
fi
|