Replaces the 3-line upstream stub with a manual that documents the platform
as it actually runs. Every claim is derived from the live code and registry
rather than from memory.
Challenge spec:
- 28-challenge tables (6 XVIII / 10 XVI / 12 XVII, 16 active) generated from
teams/challenge_registry.json, with per-challenge org_port, chall/ssh
offsets, and the real team-1 runtime ports read from state.json.
- Port formula corrected to the real one:
port = 30000 + idx*1000 + chall_offset. org_port is the native graveyard
port and is NOT used for runtime allocation, so two challenges sharing an
org_port (carbeat offset 1 vs anti-alchemy offset 30) never collide.
- Per-challenge ssh_user documented: only the 6 native XVIII images provision
ctfuser; all imported XVI/XVII images chpasswd root, so hardcoding ctfuser
breaks 10 of the 16 active challenges.
- Scoring: 100 per flag awarded to the ATTACKER (first solve only), +50 SLA
bonus at most once per 5-minute window, runtime threshold documented as
len(enabled_challenges()) rather than the hardcoded constant 6.
Setup and operations:
- Setup from clone: required /opt path, Docker, venv, panel credentials, both
systemd units verbatim, team creation, verification step.
- Full HTTP API split into public / admin / team, including why challenge
toggle and bulk team delete are async jobs.
- Troubleshooting and operational traps as declarative rules: the bare domain
is the receiver and not the panel, EOL base images, UFW default-deny
silently blackholing ports, the mandatory compose -p teamN project name,
and why docker image prune -af destroys services-* images that are in use.
- Image sizes measured from the host (189MB-903MB, ~8GB for 16 active)
instead of the incorrect "~3GB per challenge" figure.
- Topology section: PixiJS v8, on-demand rendering, and the parent-to-child
drag hierarchy derived from the edge list.
The flag example is redacted to a placeholder. No live credential, token, or
flag is committed. README.md is the only file touched.
Created a scratch team 5, booted all 16 of its containers, snapshotted its
footprint, deleted it via the API, and re-snapshotted:
before: 16 containers, 1 sidecar, team5_default network, teams/team5 dir
after : 0 containers, 0 sidecars, network gone, dir GONE, 34 UFW ports
closed, Traefik domain removed, 0 score rows
took 102s; the other 4 teams (64 containers, receivers active) were
untouched and still at 64/64 SLA.
Adds panel/team_footprint.sh (before/after proof of a delete),
team_health.sh, watch_load.sh.
Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
account nobody uses -> 'Permission denied' everywhere.
Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
(PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
timeout (26 orphans, container saturated) -> reaps the whole exec process
group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
_CRYPTO_TIMEOUT, not the 5 s prompt default.
Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.