- Nix flake now installs scrape_media.py to $out/bin so CI deploys it with the binary
- run_playwright_scraper locates the script alongside the running binary (Nix store),
the Cargo manifest dir (dev), /home/code/scraper, or $SCRAPER_SCRIPT_DIR
- Fix var_os() Option match
- Add scrape_media.py using Playwright+Chromium to bypass Cloudflare/anti-bot
- Add run_playwright_scraper() Rust helper (spawns Python, probes venv)
- Add playwright_to_download_result() JSON->DownloadResult converter
- Instagram/Facebook: try downr.org, fall back to Playwright (returns real cdninstagram/fbcdn URLs)
- Twitter: scope scraper::Html/Selector parsing in a block so the future stays Send, then Playwright fallback
- Revert --impersonate chrome (unsupported on Linux) to --user-agent
- Install Chromium browsers to /usr/local/share/ms-playwright for all users
The current otakudesu site changed paths and card layout, breaking two APIs:
1. genre pages: singular /genre/{slug}/page/{n}/ 301-redirects to a dead
otakudesu.io placeholder → 0 items. Fix: plural /genres/{slug}/page/{n}/.
Also re-target parse_genre_anime_document from the retired .venz ul li /
.thumbz h2.jdlflm layout to the new .col-anime card layout
(.col-anime-title a, .col-anime-eps, .col-anime-rating, .col-anime-cover img).
2. latest: /latest-anime/ was removed (301 → otakudesu.io). The homepage IS
the latest-episodes feed (.venz ul li with .thumbz h2.jdlflm/.epz), so
fetch_latest_anime_page now uses base_url() for page 1 and /page/N/ for
later pages instead of the dead path.
Blanket-switching the whole komik module to komiku.org broke the manga
list: komiku.org/manga/?tipe=manga 301-redirects to /pustaka/ (the new
library layout, 0 .bge items), while api.komiku.org/manga/?tipe=manga
serves the .bge grid the parser expects unredirected.
Correct split:
- api_url() = api.komiku.org -> manga/manhua/manhwa/genre/search lists (.bge layout)
- genre_list fetches base_url() = komiku.org root, which is POPULATED
(#Genre .ls3 .ls3p h4 + /genre/<slug>/ links) — api.komiku.org root is empty
Reverts the api_url() half of a46adf1; keeps the underlying insight that
the genre-list needs the populated komiku.org root.
api.komiku.org returns an empty body for its root path, which broke the
komik genre-list endpoint (it fetches the api root and parsed 0 genres).
komiku.org serves identical sub-path content (manga lists, genre, search)
plus a populated root. Consolidate the komik module onto the canonical
working domain. KOMIK2_API_URL env override still honoured.
The search parser targeted the index layout (.venz ul li, .thumbz h2.jdlflm,
img, .epz, genre-tag/status/rating classes) but otakudesu search results live
in <ul class=chivsrc><li> with <h2><a>Title</a></h2> plus .set label/value
rows. Search silently returned 0 items.
Rewrote to parse .chivsrc li, extracting title + url from h2 a and
Status/Rating/Genre from .set rows. No poster on the search page (empty).
The scraper cached a single RedisCache (one multiplexed connection) in a
OnceCell forever. When that one connection broke (Redis restart, idle
timeout, network blip), every cache op failed with 'cache io: broken pipe',
and because Cache::get_or_set propagated the post-compute write error,
EVERY API (anime, anime2, komik) returned 500 until a process restart.
Fix:
- Build a fresh RedisCache from a freshly checked-out deadpool connection
per call, so a broken connection self-heals without a process restart
(deadpool recycles/drops dead connections and reconnects on checkout).
- Make Cache::get_or_set treat a cache-write failure as non-fatal: return
the freshly computed value (cache is best-effort), so a transient Redis
outage degrades to cache-less instead of 500.
Pre-existing clippy warnings (repositories/parsers) untouched — out of scope.
mytheclipse-queue and mytheclipse-tracing are now published on crates.io
(v1.21.2). Switch both from local path deps to version deps so the whole
mytheclipse family resolves from crates.io — no more path deps in the
dependency graph.
The user published all mytheclipse crates to crates.io at 1.21.2
(including the mytheclipse-cache redis-0.32 fix). Switch the scraper to
version deps for the released crates (mytheclipse, -cache, -config,
-event) so it tracks the published releases instead of local path deps.
mytheclipse-queue + mytheclipse-tracing are not yet on crates.io, so they
remain path deps to the local workspace.