Replaces the hand-rolled inline-SVG topology with a PixiJS 8 scene graph.
Why: the old renderer rebuilt all 37 nodes / 36 edges as one innerHTML string
every 10s, which tore down and recreated every DOM node. That restarted CSS
animations mid-flight and made dragging fight the browser's own hit-testing.
The scene graph gives per-node transforms, so pan/zoom is a single container
transform instead of getScreenCTM() matrix math.
Changes:
- static/topo_pixi.js: new self-contained renderer. Owns its Application and
tears it down on tab exit so a second WebGL context cannot leak.
- static/index.html: the <svg id=topoSvg> host becomes a <div id=topoHost>;
the 188-line SVG renderer is replaced by a bridge to the module.
- static/vendor/pixi.mjs: PixiJS 8.21.0 self-hosted (MIT). The .mjs build is
required; the .js build exports no global. See vendor/README.md.
- main.py: mount /static. Pages were served as inline HTMLResponse, so the
directory was never mounted and the module had no URL to load from.
Two real bugs found by measuring pixels rather than trusting init():
- preserveDrawingBuffer: without it WebGL clears the back buffer after
compositing, so any readback or screenshot of the canvas is a coin flip
depending on which frame it lands on. The graph rendered intermittently
blank. Now enabled: cheap for a 2D scene, and it makes the view capturable.
- Layout was centred on the SCROLLABLE width (nodes.length * 130), not the
viewport, so with 37 nodes every team and challenge node landed at
x=2230-2650 on a 1310px canvas: entirely off-screen. Layout now centres on
the visible width and reset() frames the whole graph to fit.
test_topo_pixels.js documents three wrong test designs it replaces, all of
which reported false failures against a working graph: counting scene-graph
children (passes on a blank canvas), diffing against the background colour
(the theme is dark by design, so a perfect render measures ~0%), and diffing
two Playwright screenshots (both can be captured after the scene was mutated).
The check now reads the GL back buffer via readPixels in one evaluate.
Verified: 37 nodes / 36 edges drawn (7.94% of frame, max channel delta 225),
graph bbox [437,46,881,476] inside the 1310x520 canvas, glGetError=0, no page
errors, zoom and frame-to-fit reset working. Platform unregressed: SLA 32/32.
The web SSH terminal and the credential API reported `ctfuser` for all 16
challenges, but only the 6 native GEMASTIK XVIII images provision ctfuser.
Every imported XVI/XVII image does `RUN echo root:${PASSWORD} | chpasswd`,
so 10 of 16 participant logins were refused with "Permission denied".
Root causes (all the same class of bug - login hardcoded in the wrong layer):
- main.py websocket ssh handler read st["ssh_user"], a single team-wide value
defaulting to ctfuser, instead of the per-challenge registry field
- /api/credential proxied the global receiver on :18080, which only knows the
6 native challenges, so the other 10 returned "Invalid challenge"
- team.html hardcoded the challenge picker to those same 6 challenges, making
the other 10 unreachable from the terminal entirely
- index.html rendered `<b>ctfuser</b>` and a stale hardcoded SSH port table
Fixes:
- orch.challenge_credential()/all_teams() read the TEAM's state.json, which
holds the same per-challenge password the panel chpasswds
- gen_receiver_services.py injects SSH_USER_<port> from the registry so the
receiver's /credential endpoint agrees with the panel
- receiver Challenge.credentials() honours SSH_USER_<port> (ctfuser fallback)
- new /api/team/{idx}/own-challenges feeds the picker; targets now carry
challenge + ssh_user
- UI takes user and port from the server instead of hardcoding them
Verified: 32/32 credential payloads correct across teams 1-2, and 32/32 real
paramiko SSH logins succeed with whoami confirming the expected account.
Also adds bulk team delete: POST /api/teams/bulk-delete runs one background
thread and is polled via GET /api/teams/bulk-delete/{job_id}, plus per-team
checkboxes with select-all/clear in the UI. Deletion must stay sequential
because delete_team() regenerates shared artifacts at the end.
Passwords failed on 10/16 challenges while state.json looked correct:
- only the 6 native GEMASTIK XVIII images provision 'ctfuser'; every imported
XVI/XVII image does 'echo root:${PASSWORD} | chpasswd' and logs in as root.
set_ssh_passwords() hardcoded ctfuser, so chpasswd set a password on an
account nobody uses -> 'Permission denied' everywhere.
Registry gains a per-challenge 'ssh_user'; chpasswd now targets the real
login (and ctfuser/ctf when present) and reports failures loudly.
- phew checker: chall.py block-buffers stdout through the docker exec pipe
(PYTHONUNBUFFERED now set) and leaks chall.py inside the container on
timeout (26 orphans, container saturated) -> reaps the whole exec process
group. Startup does a fresh Pailier keygen (~12 s) so crypto reads need
_CRYPTO_TIMEOUT, not the 5 s prompt default.
Adds panel/verify_ssh_creds.py (proves the state->container binding from
inside via a real login), audit_ssh_users.sh, reset_runtime.sh.
- fix_dup_volumes.py: 4 canonical templates had TWO volumes: keys inside one
service (invalid YAML -> 'mapping key volumes already defined'), which broke
every enable for anti-alchemy/burvesigner/gemas-notes/kode-viewer.
- fjb: ghcr.io base is not anonymously pullable on this host; swapped to the
official httpd:2.4 (its httpd.conf only uses stock modules). Added
onlyBuiltDependencies to package.json (pnpm >=10 blocks esbuild's postinstall).
- xl + kode-viewer: node:20-slim-bookworm is not a real tag; use
node:20-bookworm-slim. gift-voucher: buster -> bookworm.
- prebuild_images.py: build each challenge's shared services-<name> image once
in parallel (passes a placeholder PASSWORD build-arg, since several Dockerfiles
run chpasswd and fail on an empty arg).
- set_enabled.py / sync_all_challenges.py: batch registry flip + runtime apply
that survives panel restarts and reports per-team results.
- Challenge toggle is now async: PATCH returns a job id, the client polls
/api/challenges/jobs/<id> so a multi-minute build no longer blocks the panel.
Added _SYNC_LOCK to serialize concurrent compose rewrites.
- guide link now server-side replaced to /team/<idx>/guide (no /team/0 403)
- _check_team_host() applied to ALL team endpoints (login, info, targets,
status, guide, portal, ssh-ws): host must match team domain; panel/gemastik
host only with admin session. Cross-domain session reuse -> 403.
- host check BEFORE auth on info/targets (no team-existence oracle)
- loading overlay (spinner + text) on start/stop all-team/set; JS util
showLoading/hideLoading
- FastAPI app at panel/ proxying receiver API server-side (admin creds stay server-side)
- Login-protected dashboard: SLA status, rotate flag, restart/rollback/activate/deactivate, SSH creds, command history
- Runs as systemd service gemastik-panel.service on :18081
- Published at https://panel.gemastik.imrnes.team via Traefik