The off-screen-layout bug (nodes centred on the scrollable width instead of the
viewport) only reproduced at the default window size, so a single-width test
would have shipped it. Verifies the graph stays framed at 1920/1400/1100/820,
and records which assumption in the pixel census is safe: the measurement uses
the canvas buffer size, not the screen size, because the CSS min-width makes
narrow viewports scroll.
Replaces the hand-rolled inline-SVG topology with a PixiJS 8 scene graph.
Why: the old renderer rebuilt all 37 nodes / 36 edges as one innerHTML string
every 10s, which tore down and recreated every DOM node. That restarted CSS
animations mid-flight and made dragging fight the browser's own hit-testing.
The scene graph gives per-node transforms, so pan/zoom is a single container
transform instead of getScreenCTM() matrix math.
Changes:
- static/topo_pixi.js: new self-contained renderer. Owns its Application and
tears it down on tab exit so a second WebGL context cannot leak.
- static/index.html: the <svg id=topoSvg> host becomes a <div id=topoHost>;
the 188-line SVG renderer is replaced by a bridge to the module.
- static/vendor/pixi.mjs: PixiJS 8.21.0 self-hosted (MIT). The .mjs build is
required; the .js build exports no global. See vendor/README.md.
- main.py: mount /static. Pages were served as inline HTMLResponse, so the
directory was never mounted and the module had no URL to load from.
Two real bugs found by measuring pixels rather than trusting init():
- preserveDrawingBuffer: without it WebGL clears the back buffer after
compositing, so any readback or screenshot of the canvas is a coin flip
depending on which frame it lands on. The graph rendered intermittently
blank. Now enabled: cheap for a 2D scene, and it makes the view capturable.
- Layout was centred on the SCROLLABLE width (nodes.length * 130), not the
viewport, so with 37 nodes every team and challenge node landed at
x=2230-2650 on a 1310px canvas: entirely off-screen. Layout now centres on
the visible width and reset() frames the whole graph to fit.
test_topo_pixels.js documents three wrong test designs it replaces, all of
which reported false failures against a working graph: counting scene-graph
children (passes on a blank canvas), diffing against the background colour
(the theme is dark by design, so a perfect render measures ~0%), and diffing
two Playwright screenshots (both can be captured after the scene was mutated).
The check now reads the GL back buffer via readPixels in one evaluate.
Verified: 37 nodes / 36 edges drawn (7.94% of frame, max channel delta 225),
graph bbox [437,46,881,476] inside the 1310x520 canvas, glGetError=0, no page
errors, zoom and frame-to-fit reset working. Platform unregressed: SLA 32/32.
The web SSH terminal and the credential API reported `ctfuser` for all 16
challenges, but only the 6 native GEMASTIK XVIII images provision ctfuser.
Every imported XVI/XVII image does `RUN echo root:${PASSWORD} | chpasswd`,
so 10 of 16 participant logins were refused with "Permission denied".
Root causes (all the same class of bug - login hardcoded in the wrong layer):
- main.py websocket ssh handler read st["ssh_user"], a single team-wide value
defaulting to ctfuser, instead of the per-challenge registry field
- /api/credential proxied the global receiver on :18080, which only knows the
6 native challenges, so the other 10 returned "Invalid challenge"
- team.html hardcoded the challenge picker to those same 6 challenges, making
the other 10 unreachable from the terminal entirely
- index.html rendered `<b>ctfuser</b>` and a stale hardcoded SSH port table
Fixes:
- orch.challenge_credential()/all_teams() read the TEAM's state.json, which
holds the same per-challenge password the panel chpasswds
- gen_receiver_services.py injects SSH_USER_<port> from the registry so the
receiver's /credential endpoint agrees with the panel
- receiver Challenge.credentials() honours SSH_USER_<port> (ctfuser fallback)
- new /api/team/{idx}/own-challenges feeds the picker; targets now carry
challenge + ssh_user
- UI takes user and port from the server instead of hardcoding them
Verified: 32/32 credential payloads correct across teams 1-2, and 32/32 real
paramiko SSH logins succeed with whoami confirming the expected account.
Also adds bulk team delete: POST /api/teams/bulk-delete runs one background
thread and is polled via GET /api/teams/bulk-delete/{job_id}, plus per-team
checkboxes with select-all/clear in the UI. Deletion must stay sequential
because delete_team() regenerates shared artifacts at the end.
Created a scratch team 5, booted all 16 of its containers, snapshotted its
footprint, deleted it via the API, and re-snapshotted:
before: 16 containers, 1 sidecar, team5_default network, teams/team5 dir
after : 0 containers, 0 sidecars, network gone, dir GONE, 34 UFW ports
closed, Traefik domain removed, 0 score rows
took 102s; the other 4 teams (64 containers, receivers active) were
untouched and still at 64/64 SLA.
Adds panel/team_footprint.sh (before/after proof of a delete),
team_health.sh, watch_load.sh.
Three independent root causes, all found by measuring instead of assuming:
1. SSH failed on 10/16 challenges while state.json looked perfect.
Only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password nobody
used -> 'Permission denied' on every team. Registry gains a per-challenge
'ssh_user'; chpasswd targets the real login and reports failures loudly.
2. phew SLA timed out on a healthy service, four bugs stacked:
- chall.py block-buffers stdout through the exec pipe (PYTHONUNBUFFERED now
set) and does a fresh Pailier keygen (~12 s) before printing its menu;
- _read_until read a TEXT pipe, so read(1) pulled 8 KB into Python's
TextIOWrapper buffer and select() then blocked on data already in memory;
- its buffer was per-call, so the read satisfying 'pt (hex)' also swallowed
the '> ' the next call waited for -> a race that failed intermittently;
- reaping killed chall.py it did not own: a blanket pkill -f, a
snapshot-diff (concurrent sessions diff against the same pre-spawn set),
and a class-level _children shared across uvicorn's thread pool. The child
now prints its own pid so exactly one session is reaped.
Also: ONE interactive session per check instead of five spawns (Paillier is
randomized per ciphertext, not per process) - 5 keygens were the CPU load
that starved the checks. And the 6 orphan single-node containers from the
original deploy were removed; one held 58 leaked chall.py and drove load
average 76 on 2 CPUs.
3. missing_sidecars() matched compose-generated names (teamN-<svc>-1) while
every service sets an explicit container_name, so it reported all 16 running
challenges as missing and hid the one real gap (anti-alchemy-db, which has
no container_name). Now reads container_name when present and falls back to
the compose default otherwise.
Verified: 64/64 SLA across 4 teams; 64/64 real SSH logins succeed with
correct <chall>_teamN hostnames; phew 3/3 sequential with no process leak.
Adds panel/verify_ssh_creds.py, audit_ssh_users.sh, reset_runtime.sh,
sla_sweep.sh, fix_sidecars.sh, phew_concurrency_test.sh, exec_probe_i.py.
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
account nobody uses -> 'Permission denied' everywhere.
Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
(PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
timeout (26 orphans, container saturated) -> reaps the whole exec process
group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
_CRYPTO_TIMEOUT, not the 5 s prompt default.
Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.
Root causes found by prebuilding every challenge image in parallel:
- fjb: ghcr.io base is not anonymously pullable here -> official httpd:2.4.
pnpm 12 (via corepack on node:20) fails the install with
ERR_PNPM_IGNORED_BUILDS unless build scripts are approved; neither
onlyBuiltDependencies in pnpm-workspace.yaml nor --no-ignore-scripts
suppresses it. The working sequence is:
pnpm install --ignore-scripts && pnpm approve-builds --all && pnpm rebuild
- xl + kode-viewer: node:20-slim-bookworm is not a real tag -> node:20-bookworm-slim.
- burvesigner: python-dev no longer exists in bookworm -> dropped (python3-dev
was already there and the source has no py2 syntax).
- burvesigner/hirnfick/s3: apt update and install were separate RUN layers;
with the bundled apt-insecure.conf the second invocation re-resolved against
the EOL bullseye-security mirror and 404'd every package. Merged into one
'update && install' layer (fix_apt_layers.py, idempotent).
- consolidate_images.sh: teams used to build a private image per team
(team1-x ... team4-x) because no shared image existed. Since the password is
applied at runtime via chpasswd, one shared services-<name> build is enough;
this reclaims ~1.5 GB, which matters on a 79 GB disk.
- reconcile_team_state(): a challenge enabled while a team was down left
state.json without ports/flag/password, so the next compose render died with
KeyError. Now both the API and the CLI tools reconcile first.
- fix_dup_volumes.py: 4 canonical templates had TWO volumes: keys inside one
service (invalid YAML -> 'mapping key volumes already defined'), which broke
every enable for anti-alchemy/burvesigner/gemas-notes/kode-viewer.
- fjb: ghcr.io base is not anonymously pullable on this host; swapped to the
official httpd:2.4 (its httpd.conf only uses stock modules). Added
onlyBuiltDependencies to package.json (pnpm >=10 blocks esbuild's postinstall).
- xl + kode-viewer: node:20-slim-bookworm is not a real tag; use
node:20-bookworm-slim. gift-voucher: buster -> bookworm.
- prebuild_images.py: build each challenge's shared services-<name> image once
in parallel (passes a placeholder PASSWORD build-arg, since several Dockerfiles
run chpasswd and fail on an empty arg).
- set_enabled.py / sync_all_challenges.py: batch registry flip + runtime apply
that survives panel restarts and reports per-team results.
- Challenge toggle is now async: PATCH returns a job id, the client polls
/api/challenges/jobs/<id> so a multi-minute build no longer blocks the panel.
Added _SYNC_LOCK to serialize concurrent compose rewrites.
Found by testing a real enable/disable cycle (art, fjb, gift-card):
1. compose_gen always swapped build->image, so a never-built challenge
produced 'pull access denied for services-<name>'. Now it only reuses
the image when it exists locally, otherwise keeps build: so
'docker compose up --build' builds it.
2. Canonical templates use 'build: context: .' (written for the shared
services/ tree). In the per-team compose that resolves to the team dir
which has no Dockerfile -> 'failed to read dockerfile'. The renderer
now rewrites the main service's context to ./<name>.
3. Teams created before the XVI/XVII import had no xvi/xvii subpackages
under their local challenges/ dir, so the regenerated receiver main.py
crash-looped on import. gen_receiver_main now mirrors ALL shared
checkers (native + xvi + xvii) into every team receiver on each sync.
4. systemd Environment= keys can't contain hyphens, so
CHALLENGE_PORT_GIFT-CARD was silently dropped. Keys are now
normalized to underscores on both the writer and reader side.
5. Several checkers called 'docker exec' with no timeout; against a
container with accumulated chall.py zombies that blocks forever and
stalls the whole SLA loop. Added mandatory timeouts (Phew, Sheesh,
Carbeat, Poke, Warmup).
Also: enabling a challenge now copies its source tree into each team's
services/ dir (team dirs only held challenges enabled at create_team
time), and the XVII checkers were rewritten to be protocol-aware
(gift-card/gift-voucher are socat TCP, not HTTP) with strict timeouts.
- all 6 Dockerfiles: vim curl wget netcat git python3-pip now installed
- apt-insecure.conf (AllowInsecureRepositories) copied into images so
participants can apt-get install despite expired Ubuntu/Debian GPG keys
- warmup base ubuntu:20.04 (EOL, GPG expired) -> ubuntu:24.04
- installed vim+git live into all 18 running team containers
- team portal target dropdown reloads after login (was empty pre-auth)
- attack log endpoint + A/D submit (attacker vs target) verified e2e
Receivers were child Popen processes of gemastik-panel systemd cgroup;
restarting the panel killed all team receivers (SLA -> 0/6, 401 on
/api/team/N/status proxy). Now each team receiver is a systemd service
(gemastik-receiver-teamN.service) generated by gen_receiver_services.py with
per-team env (ports, containers, COMPOSE_LOCATION, .env). Verified: panel
restart no longer kills receivers; 18/18 SLA stays UP.
- guide link now server-side replaced to /team/<idx>/guide (no /team/0 403)
- _check_team_host() applied to ALL team endpoints (login, info, targets,
status, guide, portal, ssh-ws): host must match team domain; panel/gemastik
host only with admin session. Cross-domain session reuse -> 403.
- host check BEFORE auth on info/targets (no team-existence oracle)
- loading overlay (spinner + text) on start/stop all-team/set; JS util
showLoading/hideLoading
- PORT_BASE 20000->30000: team1=31xxx team2=32xxx; syncthing owns 22000
- create_team replaces build: with image: services-<name> so teams reuse base images (was rebuilding 6 images per team, disk 100%)
- recovery: receiver/main.py was corrupted by bad patch (write_file with read_file format); restored from team1 copy + original GitHub
- docker compose -p teamN: project isolation so team compose doesn't overlap (was showing team1 containers for team2)
- FastAPI app at panel/ proxying receiver API server-side (admin creds stay server-side)
- Login-protected dashboard: SLA status, rotate flag, restart/rollback/activate/deactivate, SSH creds, command history
- Runs as systemd service gemastik-panel.service on :18081
- Published at https://panel.gemastik.imrnes.team via Traefik
- blogpost: python:3.11-slim-bullseye is EOL (apt 404s), switch to bookworm
- utils/bashrc: host port 80 is taken by Traefik/Coolify, preexec posts to :18080
- ignore receiver .venv