Replaces the 3-line upstream stub with a manual that documents the platform
as it actually runs. Every claim is derived from the live code and registry
rather than from memory.
Challenge spec:
- 28-challenge tables (6 XVIII / 10 XVI / 12 XVII, 16 active) generated from
teams/challenge_registry.json, with per-challenge org_port, chall/ssh
offsets, and the real team-1 runtime ports read from state.json.
- Port formula corrected to the real one:
port = 30000 + idx*1000 + chall_offset. org_port is the native graveyard
port and is NOT used for runtime allocation, so two challenges sharing an
org_port (carbeat offset 1 vs anti-alchemy offset 30) never collide.
- Per-challenge ssh_user documented: only the 6 native XVIII images provision
ctfuser; all imported XVI/XVII images chpasswd root, so hardcoding ctfuser
breaks 10 of the 16 active challenges.
- Scoring: 100 per flag awarded to the ATTACKER (first solve only), +50 SLA
bonus at most once per 5-minute window, runtime threshold documented as
len(enabled_challenges()) rather than the hardcoded constant 6.
Setup and operations:
- Setup from clone: required /opt path, Docker, venv, panel credentials, both
systemd units verbatim, team creation, verification step.
- Full HTTP API split into public / admin / team, including why challenge
toggle and bulk team delete are async jobs.
- Troubleshooting and operational traps as declarative rules: the bare domain
is the receiver and not the panel, EOL base images, UFW default-deny
silently blackholing ports, the mandatory compose -p teamN project name,
and why docker image prune -af destroys services-* images that are in use.
- Image sizes measured from the host (189MB-903MB, ~8GB for 16 active)
instead of the incorrect "~3GB per challenge" figure.
- Topology section: PixiJS v8, on-demand rendering, and the parent-to-child
drag hierarchy derived from the edge list.
The flag example is redacted to a placeholder. No live credential, token, or
flag is committed. README.md is the only file touched.
User report: the graph "disappeared". Reproduced, and the cause was not the
graph at all — the panel's main thread was blocked hard enough that the browser
stopped responding to clicks.
Root cause, found by measuring rather than guessing:
rAF 2 FPS, setTimeout(0) lag 1353ms, with the graph rendering correctly the
whole time. The scene was repainting continuously at 60fps even though it is
static between the 10s data refreshes. Each repaint cost ~305ms on this
host's software GL, so the main thread never got a free slot.
Fixes:
- Render on demand. The ticker is registered but not started; it runs only while
an attack pulse is in flight or the user is dragging, and stops once the scene
is quiet. applyView() (the single funnel for pan/zoom/drag/refit) repaints
synchronously. Result: 2 FPS -> 73 FPS, 1353ms -> 2ms lag. That is faster than
a blank page in the same harness (29 FPS), which confirms the ticker was the
cost, not the scene complexity.
- Stop rebuilding Pixi objects every refresh. _redrawEdges destroyed and
recreated ~70 Graphics + ~36 Text each cycle; every new Text allocates a
canvas, rasterises glyphs and uploads a texture. Edges are now kept per
from->to key and only their geometry is redrawn; a label's text is re-set only
when the string actually changes. Same for node titles/subtitles.
- Bind tooltip handlers once, at node creation. Re-binding inside the refresh
loop added a pointerover listener every 10s, so a single hover after an hour
fired thousands of handlers.
Two correctness bugs fixed while in there:
- init() race. The Topology tab button calls showView('topo') AND loadTopo(), and
showView() itself calls loadTopo(), so two ran concurrently. The re-entry guard
tested `topoGraph && topoGraph.app`, but topoGraph was assigned BEFORE the
awaited init(), so app was still null and a second renderer was built. Both
cleared host.innerHTML and appended their own canvas, so the last init to
finish won the DOM while the bridge still pointed at the other — a live canvas
that was no longer on the page. Now guarded by a single-flight promise, and
the instance is published only after init() resolves.
- Dropped the dead topoLoaded flag left over from the SVG renderer.
Hierarchical drag: dragging a team now carries its 16 challenge nodes with it.
The parent->children index is built from the edge list, the same source the
lines are drawn from, so the drag hierarchy cannot disagree with the picture;
challenges missing from the edge list still attach by id prefix. Children
translate rigidly (verified: 0.000px deviation across all 16).
Verified from a cold panel restart: all 5 topology tests pass, 73 FPS, 2ms
lag, 37 nodes drawn, graph framed at 4 viewport widths, no page errors, and
platform SLA unregressed at 32/32.
The off-screen-layout bug (nodes centred on the scrollable width instead of the
viewport) only reproduced at the default window size, so a single-width test
would have shipped it. Verifies the graph stays framed at 1920/1400/1100/820,
and records which assumption in the pixel census is safe: the measurement uses
the canvas buffer size, not the screen size, because the CSS min-width makes
narrow viewports scroll.
Replaces the hand-rolled inline-SVG topology with a PixiJS 8 scene graph.
Why: the old renderer rebuilt all 37 nodes / 36 edges as one innerHTML string
every 10s, which tore down and recreated every DOM node. That restarted CSS
animations mid-flight and made dragging fight the browser's own hit-testing.
The scene graph gives per-node transforms, so pan/zoom is a single container
transform instead of getScreenCTM() matrix math.
Changes:
- static/topo_pixi.js: new self-contained renderer. Owns its Application and
tears it down on tab exit so a second WebGL context cannot leak.
- static/index.html: the <svg id=topoSvg> host becomes a <div id=topoHost>;
the 188-line SVG renderer is replaced by a bridge to the module.
- static/vendor/pixi.mjs: PixiJS 8.21.0 self-hosted (MIT). The .mjs build is
required; the .js build exports no global. See vendor/README.md.
- main.py: mount /static. Pages were served as inline HTMLResponse, so the
directory was never mounted and the module had no URL to load from.
Two real bugs found by measuring pixels rather than trusting init():
- preserveDrawingBuffer: without it WebGL clears the back buffer after
compositing, so any readback or screenshot of the canvas is a coin flip
depending on which frame it lands on. The graph rendered intermittently
blank. Now enabled: cheap for a 2D scene, and it makes the view capturable.
- Layout was centred on the SCROLLABLE width (nodes.length * 130), not the
viewport, so with 37 nodes every team and challenge node landed at
x=2230-2650 on a 1310px canvas: entirely off-screen. Layout now centres on
the visible width and reset() frames the whole graph to fit.
test_topo_pixels.js documents three wrong test designs it replaces, all of
which reported false failures against a working graph: counting scene-graph
children (passes on a blank canvas), diffing against the background colour
(the theme is dark by design, so a perfect render measures ~0%), and diffing
two Playwright screenshots (both can be captured after the scene was mutated).
The check now reads the GL back buffer via readPixels in one evaluate.
Verified: 37 nodes / 36 edges drawn (7.94% of frame, max channel delta 225),
graph bbox [437,46,881,476] inside the 1310x520 canvas, glGetError=0, no page
errors, zoom and frame-to-fit reset working. Platform unregressed: SLA 32/32.
The web SSH terminal and the credential API reported `ctfuser` for all 16
challenges, but only the 6 native GEMASTIK XVIII images provision ctfuser.
Every imported XVI/XVII image does `RUN echo root:${PASSWORD} | chpasswd`,
so 10 of 16 participant logins were refused with "Permission denied".
Root causes (all the same class of bug - login hardcoded in the wrong layer):
- main.py websocket ssh handler read st["ssh_user"], a single team-wide value
defaulting to ctfuser, instead of the per-challenge registry field
- /api/credential proxied the global receiver on :18080, which only knows the
6 native challenges, so the other 10 returned "Invalid challenge"
- team.html hardcoded the challenge picker to those same 6 challenges, making
the other 10 unreachable from the terminal entirely
- index.html rendered `<b>ctfuser</b>` and a stale hardcoded SSH port table
Fixes:
- orch.challenge_credential()/all_teams() read the TEAM's state.json, which
holds the same per-challenge password the panel chpasswds
- gen_receiver_services.py injects SSH_USER_<port> from the registry so the
receiver's /credential endpoint agrees with the panel
- receiver Challenge.credentials() honours SSH_USER_<port> (ctfuser fallback)
- new /api/team/{idx}/own-challenges feeds the picker; targets now carry
challenge + ssh_user
- UI takes user and port from the server instead of hardcoding them
Verified: 32/32 credential payloads correct across teams 1-2, and 32/32 real
paramiko SSH logins succeed with whoami confirming the expected account.
Also adds bulk team delete: POST /api/teams/bulk-delete runs one background
thread and is polled via GET /api/teams/bulk-delete/{job_id}, plus per-team
checkboxes with select-all/clear in the UI. Deletion must stay sequential
because delete_team() regenerates shared artifacts at the end.
Created a scratch team 5, booted all 16 of its containers, snapshotted its
footprint, deleted it via the API, and re-snapshotted:
before: 16 containers, 1 sidecar, team5_default network, teams/team5 dir
after : 0 containers, 0 sidecars, network gone, dir GONE, 34 UFW ports
closed, Traefik domain removed, 0 score rows
took 102s; the other 4 teams (64 containers, receivers active) were
untouched and still at 64/64 SLA.
Adds panel/team_footprint.sh (before/after proof of a delete),
team_health.sh, watch_load.sh.
Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
account nobody uses -> 'Permission denied' everywhere.
Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
(PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
timeout (26 orphans, container saturated) -> reaps the whole exec process
group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
_CRYPTO_TIMEOUT, not the 5 s prompt default.
Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.
Root causes found by prebuilding every challenge image in parallel:
- fjb: ghcr.io base is not anonymously pullable here -> official httpd:2.4.
pnpm 12 (via corepack on node:20) fails the install with
ERR_PNPM_IGNORED_BUILDS unless build scripts are approved; neither
onlyBuiltDependencies in pnpm-workspace.yaml nor --no-ignore-scripts
suppresses it. The working sequence is:
pnpm install --ignore-scripts && pnpm approve-builds --all && pnpm rebuild
- xl + kode-viewer: node:20-slim-bookworm is not a real tag -> node:20-bookworm-slim.
- burvesigner: python-dev no longer exists in bookworm -> dropped (python3-dev
was already there and the source has no py2 syntax).
- burvesigner/hirnfick/s3: apt update and install were separate RUN layers;
with the bundled apt-insecure.conf the second invocation re-resolved against
the EOL bullseye-security mirror and 404'd every package. Merged into one
'update && install' layer (fix_apt_layers.py, idempotent).
- consolidate_images.sh: teams used to build a private image per team
(team1-x ... team4-x) because no shared image existed. Since the password is
applied at runtime via chpasswd, one shared services-<name> build is enough;
this reclaims ~1.5 GB, which matters on a 79 GB disk.
- reconcile_team_state(): a challenge enabled while a team was down left
state.json without ports/flag/password, so the next compose render died with
KeyError. Now both the API and the CLI tools reconcile first.
- fix_dup_volumes.py: 4 canonical templates had TWO volumes: keys inside one
service (invalid YAML -> 'mapping key volumes already defined'), which broke
every enable for anti-alchemy/burvesigner/gemas-notes/kode-viewer.
- fjb: ghcr.io base is not anonymously pullable on this host; swapped to the
official httpd:2.4 (its httpd.conf only uses stock modules). Added
onlyBuiltDependencies to package.json (pnpm >=10 blocks esbuild's postinstall).
- xl + kode-viewer: node:20-slim-bookworm is not a real tag; use
node:20-bookworm-slim. gift-voucher: buster -> bookworm.
- prebuild_images.py: build each challenge's shared services-<name> image once
in parallel (passes a placeholder PASSWORD build-arg, since several Dockerfiles
run chpasswd and fail on an empty arg).
- set_enabled.py / sync_all_challenges.py: batch registry flip + runtime apply
that survives panel restarts and reports per-team results.
- Challenge toggle is now async: PATCH returns a job id, the client polls
/api/challenges/jobs/<id> so a multi-minute build no longer blocks the panel.
Added _SYNC_LOCK to serialize concurrent compose rewrites.
Found by testing a real enable/disable cycle (art, fjb, gift-card):
1. compose_gen always swapped build->image, so a never-built challenge
produced 'pull access denied for services-<name>'. Now it only reuses
the image when it exists locally, otherwise keeps build: so
'docker compose up --build' builds it.
2. Canonical templates use 'build: context: .' (written for the shared
services/ tree). In the per-team compose that resolves to the team dir
which has no Dockerfile -> 'failed to read dockerfile'. The renderer
now rewrites the main service's context to ./<name>.
3. Teams created before the XVI/XVII import had no xvi/xvii subpackages
under their local challenges/ dir, so the regenerated receiver main.py
crash-looped on import. gen_receiver_main now mirrors ALL shared
checkers (native + xvi + xvii) into every team receiver on each sync.
4. systemd Environment= keys can't contain hyphens, so
CHALLENGE_PORT_GIFT-CARD was silently dropped. Keys are now
normalized to underscores on both the writer and reader side.
5. Several checkers called 'docker exec' with no timeout; against a
container with accumulated chall.py zombies that blocks forever and
stalls the whole SLA loop. Added mandatory timeouts (Phew, Sheesh,
Carbeat, Poke, Warmup).
Also: enabling a challenge now copies its source tree into each team's
services/ dir (team dirs only held challenges enabled at create_team
time), and the XVII checkers were rewritten to be protocol-aware
(gift-card/gift-voucher are socat TCP, not HTTP) with strict timeouts.
- all 6 Dockerfiles: vim curl wget netcat git python3-pip now installed
- apt-insecure.conf (AllowInsecureRepositories) copied into images so
participants can apt-get install despite expired Ubuntu/Debian GPG keys
- warmup base ubuntu:20.04 (EOL, GPG expired) -> ubuntu:24.04
- installed vim+git live into all 18 running team containers
- team portal target dropdown reloads after login (was empty pre-auth)
- attack log endpoint + A/D submit (attacker vs target) verified e2e
Receivers were child Popen processes of gemastik-panel systemd cgroup;
restarting the panel killed all team receivers (SLA -> 0/6, 401 on
/api/team/N/status proxy). Now each team receiver is a systemd service
(gemastik-receiver-teamN.service) generated by gen_receiver_services.py with
per-team env (ports, containers, COMPOSE_LOCATION, .env). Verified: panel
restart no longer kills receivers; 18/18 SLA stays UP.
- guide link now server-side replaced to /team/<idx>/guide (no /team/0 403)
- _check_team_host() applied to ALL team endpoints (login, info, targets,
status, guide, portal, ssh-ws): host must match team domain; panel/gemastik
host only with admin session. Cross-domain session reuse -> 403.
- host check BEFORE auth on info/targets (no team-existence oracle)
- loading overlay (spinner + text) on start/stop all-team/set; JS util
showLoading/hideLoading
- PORT_BASE 20000->30000: team1=31xxx team2=32xxx; syncthing owns 22000
- create_team replaces build: with image: services-<name> so teams reuse base images (was rebuilding 6 images per team, disk 100%)
- recovery: receiver/main.py was corrupted by bad patch (write_file with read_file format); restored from team1 copy + original GitHub
- docker compose -p teamN: project isolation so team compose doesn't overlap (was showing team1 containers for team2)
- FastAPI app at panel/ proxying receiver API server-side (admin creds stay server-side)
- Login-protected dashboard: SLA status, rotate flag, restart/rollback/activate/deactivate, SSH creds, command history
- Runs as systemd service gemastik-panel.service on :18081
- Published at https://panel.gemastik.imrnes.team via Traefik
- blogpost: python:3.11-slim-bullseye is EOL (apt 404s), switch to bookworm
- utils/bashrc: host port 80 is taken by Traefik/Coolify, preexec posts to :18080
- ignore receiver .venv