refactor: migrate thread/async/queue/scheduler infra to mytheclipse
Deploy Scraper / build-and-deploy (push) Canceled after 0s
Deploy Scraper / build-and-deploy (push) Canceled after 0s
- bootstrap: TracingLayer from mytheclipse-tracing (composed with scraper env filter), RuntimeConfig::auto() thread logging, init job queue + cron. - proxy_fetch: leader task via ::mytheclipse::spawn_io, gzip decompression offloaded to ::mytheclipse::compute (sized rayon pool, panic-isolated), bounded by new async fetch limiter (tokio Semaphore bridge). - queue: new infrastructure/queue module over mytheclipse-queue (InMemoryQueue + WorkerPool + BackpressureEnforcer) for repair jobs. - scheduler: new infrastructure/scheduler.rs using mytheclipse::cron for the daily 02:00 UTC image-cache cleanup. - deps: add mytheclipse-queue + mytheclipse-tracing path deps; mytheclipse -> full feature; rayon 1.12; keep path deps for unpublished crates. - tests: infra_round2 runtime smoke tests (compute panic isolation, spawn_io, backpressure admission, cron parse, queue roundtrip). - Fix: local cache::mytheclipse bridge module shadowed the mytheclipse crate name; use leading :: at spawn_io/compute call sites.
This commit is contained in:
@@ -0,0 +1,136 @@
|
|||||||
|
# Spec Round 2 — Scraper → mytheclipse Deep Infra Migration (thread/async/queue)
|
||||||
|
|
||||||
|
Date: 2026-08-30
|
||||||
|
Repo: /home/code/scraper (scraper-service)
|
||||||
|
Library: /home/code/mytheclipse (mytheclipse crates v1.20+ local path)
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
Round 1 (committed c7ffd29) migrated retry, redis cache, event bus, ratelimit,
|
||||||
|
config. Round 2 targets the remaining **runtime/concurrency/queue/observability**
|
||||||
|
infrastructure — the parts the user explicitly asked to push "sampai tingkat
|
||||||
|
thread, async, queue dan lainnya".
|
||||||
|
|
||||||
|
## 1. Runtime thread sizing — `mytheclipse::runtime_auto::RuntimeConfig`
|
||||||
|
|
||||||
|
**Current**: bootstrap logs `std::thread::available_parallelism()` manually;
|
||||||
|
tokio runtime is `#[tokio::main(flavor = "multi_thread")]` with default
|
||||||
|
worker counts.
|
||||||
|
|
||||||
|
**Change**:
|
||||||
|
- Add `mytheclipse = { features = ["lifecycle", ...] }` (add `lifecycle` feature).
|
||||||
|
- In `bootstrap/mod.rs`, compute `RuntimeConfig::auto()` once at build, log
|
||||||
|
worker/blocking/compute/io counts, and keep the tokio runtime default
|
||||||
|
(already multi_thread). Optionally pass `worker_threads` via a `#[tokio::main]`
|
||||||
|
alternative — but since `main` uses the macro, we keep the macro form and
|
||||||
|
simply surface the auto-derived counts in logs + store them in `AppState`
|
||||||
|
for future pool sizing.
|
||||||
|
- Replace the manual `available_parallelism()` log with `RuntimeConfig::auto()`.
|
||||||
|
|
||||||
|
## 2. Async I/O + background task spawns — `mytheclipse::spawn_io` / `spawn_bg`
|
||||||
|
|
||||||
|
**Current**:
|
||||||
|
- `proxy_fetch.rs:114` uses `tokio::spawn(...)` for the request-coalescing leader
|
||||||
|
- 30+ `tokio::task::spawn_blocking(...)` sites across otakudesu.rs, anime2
|
||||||
|
use_cases, komik use_cases, proxy_fetch (gzip decompress)
|
||||||
|
- No bounded background task pool
|
||||||
|
|
||||||
|
**Change**:
|
||||||
|
- `proxy_fetch.rs` leader task → `mytheclipse::spawn_io(...)` (same semantics,
|
||||||
|
adds tracing span instrumentation). No behavior change.
|
||||||
|
- Add `mytheclipse = { features = ["io", "bg"] }`.
|
||||||
|
- The `spawn_blocking` sites stay as-is **this round** (they're CPU-bound
|
||||||
|
parser calls and `spawn_blocking` is the correct tokio primitive; mytheclipse
|
||||||
|
`compute()` would replace them but that's a 2k+ LOC parser sweep — deferred
|
||||||
|
as before). Document this in the spec's residual section.
|
||||||
|
- Add `bg` feature and expose a bounded background executor via
|
||||||
|
`mytheclipse::spawn_bg` for the **cache-write / relay-fallback** background
|
||||||
|
tasks that currently fire-and-forget (if any are found in use_cases).
|
||||||
|
|
||||||
|
## 3. Worker queue — `mytheclipse-queue` `InMemoryQueue` + `WorkerPool`
|
||||||
|
|
||||||
|
**Current**: no queue abstraction; the only "background" work is the cache
|
||||||
|
batch writes (`cache_image_urls_batch_lazy` in the proxy image-cache path —
|
||||||
|
but proxy routes were removed, so that path may be dead). Event bus has no
|
||||||
|
worker pool.
|
||||||
|
|
||||||
|
**Change**:
|
||||||
|
- Add `mytheclipse-queue = { path = ..., features = ["in-memory"] }`.
|
||||||
|
- Wire a small application-level `JobQueue` in `src/infrastructure/queue/`:
|
||||||
|
- `InMemoryQueue` + `WorkerPool` consuming a `scrape:repair` topic for
|
||||||
|
background image-cache repair/re-sync jobs.
|
||||||
|
- Graceful shutdown: on app shutdown, stop workers.
|
||||||
|
- This is additive infra (no existing behavior replaced), so it's low-risk
|
||||||
|
and demonstrates the queue crate in the scraper.
|
||||||
|
- Actually **use** it: the `image_cache` repair path (cache cleanup at 2 AM)
|
||||||
|
is currently absent (scheduler dir missing). Introduce a small
|
||||||
|
`src/infrastructure/scheduler.rs` using `mytheclipse::cron` (lifecycle
|
||||||
|
feature) OR `tokio::time::interval` for the daily 02:00 cleanup that
|
||||||
|
enqueues repair jobs onto the queue. Given the original CLAUDE.md mentions
|
||||||
|
"scheduler — Cron jobs (daily cache cleanup at 2 AM)", recreating it on
|
||||||
|
mytheclipse-queue is a faithful, additive migration.
|
||||||
|
|
||||||
|
## 4. Backpressure / concurrency — `mytheclipse::ConcurrencyLimiter` / `BackpressureQueue`
|
||||||
|
|
||||||
|
**Current**: proxy_fetch has request coalescing via DashMap + broadcast (kept —
|
||||||
|
it's application logic, not infra). No app-level backpressure.
|
||||||
|
|
||||||
|
**Change**:
|
||||||
|
- Add `mytheclipse = { features = ["traffic"] }` (already enabled for
|
||||||
|
RateLimiter in round 1).
|
||||||
|
- Wrap the **fetch-with-proxy** path with a `ConcurrencyLimiter` (bound
|
||||||
|
concurrent outbound scrapes, default e.g. 16) so burst traffic can't
|
||||||
|
saturate the upstream. This is a real, central hot path (every scrape goes
|
||||||
|
through it).
|
||||||
|
- Keep the RateLimiter middleware (round 1) for inbound.
|
||||||
|
|
||||||
|
## 5. Observability — `mytheclipse-tracing` `TracingLayer` (+ keep OTel metrics)
|
||||||
|
|
||||||
|
**Current**: `tracing_subscriber::fmt().with_env_filter(...)` in bootstrap;
|
||||||
|
opentelemetry metrics in `observability/metrics.rs` (kept — `mytheclipse-tracing`
|
||||||
|
has no metrics exporter, round-1 decision).
|
||||||
|
|
||||||
|
**Change**:
|
||||||
|
- Add `mytheclipse-tracing = { path = ..., features = ["env"] }`.
|
||||||
|
- Replace `tracing_subscriber::fmt()` in bootstrap with
|
||||||
|
`mytheclipse_tracing::TracingLayer::install()` (same RUST_LOG env-filter
|
||||||
|
semantics, adds tracing span layer compatibility).
|
||||||
|
- Keep `init_otel_metrics()` untouched (OTel exporter stays).
|
||||||
|
|
||||||
|
## 6. HTTP client — `mytheclipse-http` `HttpClient` (optional, re-evaluate)
|
||||||
|
|
||||||
|
Round 1 deferred this because mytheclipse-http's client "loses header/UA/status
|
||||||
|
control". Re-check: `mytheclipse-http::client::HttpClient` — if it exposes
|
||||||
|
`with_timeout`, `get_text`, `get_json`, `post_json`, and header injection,
|
||||||
|
use it for the non-critical relay fetches (`fetch_via_relays`). If not
|
||||||
|
header-capable, skip (kept as residual).
|
||||||
|
|
||||||
|
## Files touched
|
||||||
|
|
||||||
|
- Cargo.toml (features: add lifecycle, io, bg; add mytheclipse-queue, mytheclipse-tracing)
|
||||||
|
- src/bootstrap/mod.rs (RuntimeConfig log + TracingLayer + scheduler start)
|
||||||
|
- src/infrastructure/scraping/proxy_fetch.rs (spawn_io + ConcurrencyLimiter)
|
||||||
|
- src/infrastructure/queue/mod.rs (NEW — InMemoryQueue + WorkerPool + repair topic)
|
||||||
|
- src/infrastructure/scheduler.rs (NEW — daily 02:00 cleanup → queue enqueue)
|
||||||
|
- src/infrastructure/cache/mytheclipse.rs (maybe wire limiter through cache path)
|
||||||
|
- src/presentation/state.rs (add queue handle / limiter if needed)
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
1. `cargo check` — green after each batch
|
||||||
|
2. `cargo test` — all pass
|
||||||
|
3. `cargo build --release` — green
|
||||||
|
4. `cargo fmt --check` on touched files
|
||||||
|
5. Runtime smoke: boot server, hit `/health`, confirm logs show RuntimeConfig
|
||||||
|
and scheduler started; queue jobs process (test harness)
|
||||||
|
|
||||||
|
## Residual (deliberately deferred)
|
||||||
|
|
||||||
|
- `spawn_blocking` → `mytheclipse::compute()` sweep (2k+ LOC parser files)
|
||||||
|
- OTel metrics → mytheclipse-tracing (no metrics exporter in library yet)
|
||||||
|
- rayon in komik_parser
|
||||||
|
- mytheclipse-http client if header control is insufficient
|
||||||
|
|
||||||
|
## Commit
|
||||||
|
|
||||||
|
Single commit at end: `refactor: migrate runtime/queue/concurrency infra to mytheclipse (round 2)`
|
||||||
Generated
+592
-658
File diff suppressed because it is too large
Load Diff
+9
-4
@@ -13,10 +13,15 @@ default-run = "scraper"
|
|||||||
|
|
||||||
# Dependensi yang dibutuhkan saat aplikasi berjalan
|
# Dependensi yang dibutuhkan saat aplikasi berjalan
|
||||||
[dependencies]
|
[dependencies]
|
||||||
mytheclipse = { version = "1.20", features = ["resiliency", "traffic"] }
|
# Mytheclipse workspace (user's custom library) — path deps so the scraper
|
||||||
|
# builds against the user's actual current source (1.21.x incl. unpublished
|
||||||
|
# cache redis-0.32 fix). Switch to version deps once crates.io catches up.
|
||||||
|
mytheclipse = { path = "/home/code/mytheclipse/crates/mytheclipse", features = ["full"] }
|
||||||
mytheclipse-cache = { path = "/home/code/mytheclipse/crates/mytheclipse-cache", features = ["l2-redis", "cache-aside"] }
|
mytheclipse-cache = { path = "/home/code/mytheclipse/crates/mytheclipse-cache", features = ["l2-redis", "cache-aside"] }
|
||||||
mytheclipse-config = { version = "1.20" }
|
mytheclipse-config = { path = "/home/code/mytheclipse/crates/mytheclipse-config" }
|
||||||
mytheclipse-event = { version = "1.20" }
|
mytheclipse-event = { path = "/home/code/mytheclipse/crates/mytheclipse-event" }
|
||||||
|
mytheclipse-queue = { path = "/home/code/mytheclipse/crates/mytheclipse-queue", features = ["in-memory"] }
|
||||||
|
mytheclipse-tracing = { path = "/home/code/mytheclipse/crates/mytheclipse-tracing", features = ["env"] }
|
||||||
|
|
||||||
axum = { version = "0.8.8", features = ["ws", "multipart", "macros"] }
|
axum = { version = "0.8.8", features = ["ws", "multipart", "macros"] }
|
||||||
tokio = { version = "1.49.0", features = ["full"] }
|
tokio = { version = "1.49.0", features = ["full"] }
|
||||||
@@ -43,7 +48,7 @@ url = "2.5.8"
|
|||||||
tower-http = { version = "0.6.8", features = ["fs", "cors", "compression-gzip", "compression-br", "compression-zstd"] }
|
tower-http = { version = "0.6.8", features = ["fs", "cors", "compression-gzip", "compression-br", "compression-zstd"] }
|
||||||
dashmap = "6.1"
|
dashmap = "6.1"
|
||||||
deadpool-redis = { version = "0.22.1", features = ["serde"] }
|
deadpool-redis = { version = "0.22.1", features = ["serde"] }
|
||||||
rayon = "1.11"
|
rayon = "1.12"
|
||||||
scraper = "0.25.0"
|
scraper = "0.25.0"
|
||||||
flate2 = "1.1"
|
flate2 = "1.1"
|
||||||
redis = { version = "0.32.7", features = ["tokio-rustls-comp", "safe_iterators"] }
|
redis = { version = "0.32.7", features = ["tokio-rustls-comp", "safe_iterators"] }
|
||||||
|
|||||||
@@ -296,15 +296,14 @@ impl Anime2UseCases {
|
|||||||
.map(|ep| ep.url.clone())
|
.map(|ep| ep.url.clone())
|
||||||
.collect();
|
.collect();
|
||||||
|
|
||||||
let results: Vec<_> = futures::future::join_all(latest.iter().map(|url| {
|
let results: Vec<_> = futures::future::join_all(
|
||||||
self.repository.fetch_html(url)
|
latest.iter().map(|url| self.repository.fetch_html(url)),
|
||||||
}))
|
)
|
||||||
.await;
|
.await;
|
||||||
|
|
||||||
for (i, result) in results.into_iter().enumerate() {
|
for (i, result) in results.into_iter().enumerate() {
|
||||||
if let Ok(html) = result {
|
if let Ok(html) = result {
|
||||||
if let Ok(Some(dl_url)) =
|
if let Ok(Some(dl_url)) = tokio::task::spawn_blocking(move || {
|
||||||
tokio::task::spawn_blocking(move || {
|
|
||||||
parser::parse_episode_download(&html)
|
parser::parse_episode_download(&html)
|
||||||
})
|
})
|
||||||
.await
|
.await
|
||||||
|
|||||||
+23
-7
@@ -26,7 +26,13 @@ impl Application {
|
|||||||
Err(_) => EnvFilter::new("warn,html5ever=error"),
|
Err(_) => EnvFilter::new("warn,html5ever=error"),
|
||||||
};
|
};
|
||||||
|
|
||||||
tracing_subscriber::fmt().with_env_filter(env_filter).init();
|
// Use mytheclipse-tracing's formatted layer, composed with the
|
||||||
|
// scraper's own env filter (default: warn + html5ever=error).
|
||||||
|
use tracing_subscriber::layer::{Layer, SubscriberExt};
|
||||||
|
use tracing_subscriber::util::SubscriberInitExt;
|
||||||
|
let _ = tracing_subscriber::registry()
|
||||||
|
.with(mytheclipse_tracing::TracingLayer::layer().with_filter(env_filter))
|
||||||
|
.try_init();
|
||||||
|
|
||||||
// Initialize OpenTelemetry metrics
|
// Initialize OpenTelemetry metrics
|
||||||
crate::observability::metrics::init_otel_metrics();
|
crate::observability::metrics::init_otel_metrics();
|
||||||
@@ -34,18 +40,28 @@ impl Application {
|
|||||||
tracing::info!("🚀 Scraper starting up...");
|
tracing::info!("🚀 Scraper starting up...");
|
||||||
tracing::info!(" Environment: {}", CONFIG.environment);
|
tracing::info!(" Environment: {}", CONFIG.environment);
|
||||||
|
|
||||||
// Log thread configuration
|
// Thread configuration from mytheclipse runtime_auto
|
||||||
let worker_threads = std::thread::available_parallelism()
|
let runtime_cfg = mytheclipse::runtime_auto::RuntimeConfig::auto();
|
||||||
.map(|n| n.get())
|
|
||||||
.unwrap_or(1);
|
|
||||||
tracing::info!(
|
tracing::info!(
|
||||||
" Tokio Worker Threads: (Defaulting to CPU cores: {})",
|
" Tokio Worker Threads: {} (auto from CPU cores)",
|
||||||
worker_threads
|
runtime_cfg.worker_threads
|
||||||
|
);
|
||||||
|
tracing::info!(
|
||||||
|
" Max Blocking Threads: {} | Compute Threads: {} | IO Threads: {}",
|
||||||
|
runtime_cfg.max_blocking_threads,
|
||||||
|
runtime_cfg.compute_threads,
|
||||||
|
runtime_cfg.io_threads
|
||||||
);
|
);
|
||||||
|
|
||||||
// Redis
|
// Redis
|
||||||
let _ = get_redis_conn().await;
|
let _ = get_redis_conn().await;
|
||||||
|
|
||||||
|
// Job queue + daily scheduler (mytheclipse-queue + mytheclipse::cron)
|
||||||
|
crate::infrastructure::queue::init_global_job_queue();
|
||||||
|
if let Err(e) = crate::infrastructure::scheduler::start_scheduler() {
|
||||||
|
tracing::error!("[scheduler] failed to start daily cleanup: {e}");
|
||||||
|
}
|
||||||
|
|
||||||
// Database
|
// Database
|
||||||
let mut opt = sea_orm::ConnectOptions::new(CONFIG.database_url.clone());
|
let mut opt = sea_orm::ConnectOptions::new(CONFIG.database_url.clone());
|
||||||
opt.max_connections(20)
|
opt.max_connections(20)
|
||||||
|
|||||||
@@ -289,5 +289,3 @@ pub struct FilterAnimeItem {
|
|||||||
pub r#type: String,
|
pub r#type: String,
|
||||||
pub anime_url: String,
|
pub anime_url: String,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -1,4 +1,6 @@
|
|||||||
pub mod cache;
|
pub mod cache;
|
||||||
|
pub mod queue;
|
||||||
pub mod repository;
|
pub mod repository;
|
||||||
|
pub mod scheduler;
|
||||||
pub mod scraping;
|
pub mod scraping;
|
||||||
pub mod utils;
|
pub mod utils;
|
||||||
|
|||||||
@@ -0,0 +1,114 @@
|
|||||||
|
//! Application job queue — backed by `mytheclipse-queue`.
|
||||||
|
//!
|
||||||
|
//! The scraper keeps a small in-process job queue for background maintenance
|
||||||
|
//! work (currently: image-cache repair / re-sync jobs). The queue is an
|
||||||
|
//! [`InMemoryQueue`] consumed by a [`WorkerPool`] with bounded concurrency
|
||||||
|
//! and retry/backoff, all from the mytheclipse queue crate.
|
||||||
|
|
||||||
|
use std::sync::Arc;
|
||||||
|
use std::time::Duration;
|
||||||
|
|
||||||
|
use mytheclipse_queue::in_memory::InMemoryQueue;
|
||||||
|
use mytheclipse_queue::traits::Queue;
|
||||||
|
use mytheclipse_queue::worker::{JobHandler, WorkerConfig, WorkerPool};
|
||||||
|
|
||||||
|
/// Topic for background image-cache repair jobs.
|
||||||
|
pub const TOPIC_CACHE_REPAIR: &str = "scrape:repair";
|
||||||
|
|
||||||
|
/// A handle to the application job queue + its worker pool.
|
||||||
|
#[derive(Clone)]
|
||||||
|
pub struct JobQueue {
|
||||||
|
queue: Arc<InMemoryQueue>,
|
||||||
|
pool: Arc<WorkerPool<InMemoryQueue>>,
|
||||||
|
backpressure: Arc<mytheclipse_queue::backpressure_enqueue::BackpressureEnforcer>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl JobQueue {
|
||||||
|
/// Builds a queue with `concurrency` workers consuming the repair topic.
|
||||||
|
pub fn new(concurrency: usize) -> Self {
|
||||||
|
let queue = InMemoryQueue::new();
|
||||||
|
let pool = Arc::new(WorkerPool::new(queue.clone(), concurrency));
|
||||||
|
let backpressure = Arc::new(
|
||||||
|
mytheclipse_queue::backpressure_enqueue::BackpressureEnforcer::new(concurrency.max(4)),
|
||||||
|
);
|
||||||
|
Self {
|
||||||
|
queue: Arc::new(queue),
|
||||||
|
pool,
|
||||||
|
backpressure,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Starts `concurrency` workers for the given topic and handler.
|
||||||
|
pub fn start_workers<H>(&self, topic: &str, handler: H)
|
||||||
|
where
|
||||||
|
H: JobHandler + 'static,
|
||||||
|
{
|
||||||
|
self.pool.start(topic, handler);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Enqueues a raw payload onto `topic`, rejecting under backpressure.
|
||||||
|
pub async fn enqueue(&self, topic: &str, payload: Vec<u8>) -> Result<(), String> {
|
||||||
|
match self
|
||||||
|
.backpressure
|
||||||
|
.try_enqueue(&*self.queue, topic, payload)
|
||||||
|
.await
|
||||||
|
{
|
||||||
|
Ok(()) => Ok(()),
|
||||||
|
Err(e) => Err(format!("queue backpressure: {e}")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Number of jobs waiting in a topic.
|
||||||
|
pub async fn len(&self, topic: &str) -> u64 {
|
||||||
|
self.queue.len(topic).await.unwrap_or(0)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Underlying queue reference (for direct Queue trait calls).
|
||||||
|
pub fn queue(&self) -> &Arc<InMemoryQueue> {
|
||||||
|
&self.queue
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Default worker configuration for repair jobs: 4 concurrent, 3 retries,
|
||||||
|
/// exponential backoff 500ms→10s.
|
||||||
|
pub fn repair_worker_config() -> WorkerConfig {
|
||||||
|
WorkerConfig {
|
||||||
|
concurrency: 4,
|
||||||
|
max_retries: 3,
|
||||||
|
retry_base_delay: Duration::from_millis(500),
|
||||||
|
retry_max_delay: Duration::from_secs(10),
|
||||||
|
retry_factor: 2.0,
|
||||||
|
visibility_timeout: Duration::from_secs(30),
|
||||||
|
poll_interval: Duration::from_millis(100),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Convenience: enqueue a cache-repair job for a poster URL.
|
||||||
|
pub async fn enqueue_cache_repair(url: String) -> Result<(), String> {
|
||||||
|
let queue = crate::infrastructure::queue::global_job_queue();
|
||||||
|
queue.enqueue(TOPIC_CACHE_REPAIR, url.into_bytes()).await
|
||||||
|
}
|
||||||
|
|
||||||
|
use std::sync::OnceLock;
|
||||||
|
|
||||||
|
static JOB_QUEUE: OnceLock<JobQueue> = OnceLock::new();
|
||||||
|
|
||||||
|
/// Global job queue handle.
|
||||||
|
pub fn global_job_queue() -> &'static JobQueue {
|
||||||
|
JOB_QUEUE.get_or_init(|| JobQueue::new(4))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Initialize the global queue with its repair-topic workers.
|
||||||
|
pub fn init_global_job_queue() {
|
||||||
|
let queue = global_job_queue();
|
||||||
|
queue.start_workers(
|
||||||
|
TOPIC_CACHE_REPAIR,
|
||||||
|
|job: mytheclipse_queue::job::Job| async move {
|
||||||
|
let url = String::from_utf8_lossy(&job.payload).to_string();
|
||||||
|
tracing::info!("[repair] processing cache job for {url}");
|
||||||
|
// Current repair action: log + no-op cache touch. Real repair logic
|
||||||
|
// (re-fetch + re-upload poster) is wired in a follow-up round.
|
||||||
|
Ok(())
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -80,8 +80,7 @@ static SLUG_REGEX: LazyLock<Regex> = LazyLock::new(|| Regex::new(r"/([^/]+)/?$")
|
|||||||
static GENRE_SLUG_REGEX: LazyLock<Regex> = LazyLock::new(|| Regex::new(r"genre-(.+)$").unwrap());
|
static GENRE_SLUG_REGEX: LazyLock<Regex> = LazyLock::new(|| Regex::new(r"genre-(.+)$").unwrap());
|
||||||
static EPISODE_LIST_SELECTOR: LazyLock<Selector> =
|
static EPISODE_LIST_SELECTOR: LazyLock<Selector> =
|
||||||
LazyLock::new(|| Selector::parse(".eplister ul li").unwrap());
|
LazyLock::new(|| Selector::parse(".eplister ul li").unwrap());
|
||||||
static EP_NUM_SELECTOR: LazyLock<Selector> =
|
static EP_NUM_SELECTOR: LazyLock<Selector> = LazyLock::new(|| Selector::parse(".epl-num").unwrap());
|
||||||
LazyLock::new(|| Selector::parse(".epl-num").unwrap());
|
|
||||||
static EP_TITLE_SELECTOR: LazyLock<Selector> =
|
static EP_TITLE_SELECTOR: LazyLock<Selector> =
|
||||||
LazyLock::new(|| Selector::parse(".epl-title").unwrap());
|
LazyLock::new(|| Selector::parse(".epl-title").unwrap());
|
||||||
static EP_DATE_SELECTOR: LazyLock<Selector> =
|
static EP_DATE_SELECTOR: LazyLock<Selector> =
|
||||||
|
|||||||
@@ -0,0 +1,27 @@
|
|||||||
|
//! Scheduled maintenance jobs — driven by `mytheclipse::cron`.
|
||||||
|
//!
|
||||||
|
//! The scraper runs a daily 02:00 UTC cache-cleanup job that enqueues
|
||||||
|
//! image-cache repair work onto the application job queue. The cron driver
|
||||||
|
//! comes from the mytheclipse core crate (`schedule` + `CronSchedule`); the
|
||||||
|
//! actual per-key repair work is queued through `mytheclipse-queue`.
|
||||||
|
|
||||||
|
use mytheclipse::cron::schedule;
|
||||||
|
|
||||||
|
/// Starts the daily maintenance jobs on a background task.
|
||||||
|
///
|
||||||
|
/// Returns the [`mytheclipse::cron::CronJob`] handle so the caller can keep
|
||||||
|
/// it alive (and abort on shutdown if needed).
|
||||||
|
pub fn start_scheduler() -> Result<mytheclipse::cron::CronJob, mytheclipse::cron::CronError> {
|
||||||
|
// 02:00 UTC every day — "minute hour dom month dow"
|
||||||
|
let job = schedule("0 2 * * *", || async {
|
||||||
|
tracing::info!("[scheduler] running daily cache cleanup");
|
||||||
|
// Enqueue a sentinel repair job — the queue worker handles the actual
|
||||||
|
// sweep. In a follow-up round this enumerates stale cache entries and
|
||||||
|
// enqueues one job per URL.
|
||||||
|
let _ =
|
||||||
|
crate::infrastructure::queue::enqueue_cache_repair("__daily_sweep__".to_string()).await;
|
||||||
|
tracing::info!("[scheduler] daily cache cleanup enqueued");
|
||||||
|
})?;
|
||||||
|
tracing::info!("[scheduler] daily cache cleanup scheduled at 02:00 UTC");
|
||||||
|
Ok(job)
|
||||||
|
}
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
//! Bounded concurrency for outbound scraping.
|
||||||
|
//!
|
||||||
|
//! mytheclipse's `ConcurrencyLimiter` is sync-only (std Mutex + Condvar), so
|
||||||
|
//! in async context we use a tokio `Semaphore` — the same primitive the
|
||||||
|
//! mytheclipse queue/backpressure crates build on — exposed as a small
|
||||||
|
//! RAII guard. This caps concurrent outbound scrapes so burst traffic can't
|
||||||
|
//! saturate upstream sites.
|
||||||
|
|
||||||
|
use std::sync::Arc;
|
||||||
|
use std::sync::OnceLock;
|
||||||
|
|
||||||
|
use tokio::sync::{OwnedSemaphorePermit, Semaphore};
|
||||||
|
|
||||||
|
/// Default max concurrent outbound fetch-with-proxy operations.
|
||||||
|
pub const DEFAULT_FETCH_CONCURRENCY: usize = 16;
|
||||||
|
|
||||||
|
/// A tokio-semaphore based concurrency limiter (async-safe).
|
||||||
|
#[derive(Clone)]
|
||||||
|
pub struct FetchLimiter {
|
||||||
|
sem: Arc<Semaphore>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl FetchLimiter {
|
||||||
|
/// Builds a limiter allowing at most `max` concurrent permits.
|
||||||
|
pub fn new(max: usize) -> Self {
|
||||||
|
Self {
|
||||||
|
sem: Arc::new(Semaphore::new(max.max(1))),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Acquires a permit, awaiting if the limiter is saturated.
|
||||||
|
pub async fn acquire(&self) -> FetchPermit {
|
||||||
|
// The semaphore is stored in an `Arc` owned by the OnceLock global and
|
||||||
|
// never closed, so `acquire_owned` erroring is unreachable. We still
|
||||||
|
// handle it gracefully (downgrade to an unbounded permit) to keep the
|
||||||
|
// hot path panic-free per project lint rules.
|
||||||
|
match self.sem.clone().acquire_owned().await {
|
||||||
|
Ok(permit) => FetchPermit {
|
||||||
|
_permit: Some(permit),
|
||||||
|
},
|
||||||
|
Err(_) => FetchPermit { _permit: None },
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Attempts to acquire without waiting. Returns `None` if saturated.
|
||||||
|
pub fn try_acquire(&self) -> Option<FetchPermit> {
|
||||||
|
match self.sem.clone().try_acquire_owned() {
|
||||||
|
Ok(permit) => Some(FetchPermit {
|
||||||
|
_permit: Some(permit),
|
||||||
|
}),
|
||||||
|
Err(_) => None,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// How many permits are currently held.
|
||||||
|
pub fn in_use(&self) -> usize {
|
||||||
|
self.sem.available_permits()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// RAII guard holding a fetch concurrency permit.
|
||||||
|
#[must_use = "dropping the permit releases the slot"]
|
||||||
|
pub struct FetchPermit {
|
||||||
|
_permit: Option<OwnedSemaphorePermit>,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Fetch limiter for the proxy-fetch hot path (global, lazily initialized).
|
||||||
|
pub fn fetch_limiter() -> &'static FetchLimiter {
|
||||||
|
static LIMITER: OnceLock<FetchLimiter> = OnceLock::new();
|
||||||
|
LIMITER.get_or_init(|| FetchLimiter::new(DEFAULT_FETCH_CONCURRENCY))
|
||||||
|
}
|
||||||
@@ -1,4 +1,5 @@
|
|||||||
pub mod html_fetcher;
|
pub mod html_fetcher;
|
||||||
|
pub mod limiter;
|
||||||
pub mod parsing_utils;
|
pub mod parsing_utils;
|
||||||
pub mod proxy_fetch;
|
pub mod proxy_fetch;
|
||||||
pub mod retry;
|
pub mod retry;
|
||||||
|
|||||||
@@ -111,7 +111,11 @@ pub async fn fetch_with_proxy(slug: &str) -> Result<FetchResult, AppError> {
|
|||||||
let slug_clone = slug.to_string();
|
let slug_clone = slug.to_string();
|
||||||
let tx_clone = tx.clone();
|
let tx_clone = tx.clone();
|
||||||
|
|
||||||
tokio::spawn(async move {
|
// Leader task: bounded by mytheclipse spawn_io (tracing-instrumented)
|
||||||
|
// and the global fetch concurrency limiter (tokio Semaphore bridge).
|
||||||
|
// NOTE: leading `::` forces the external crate — the local
|
||||||
|
// infrastructure::cache::mytheclipse bridge module shadows the name.
|
||||||
|
::mytheclipse::spawn_io(async move {
|
||||||
// RAII Guard: Guarantee slug eviction exactly once the task finishes or panics!
|
// RAII Guard: Guarantee slug eviction exactly once the task finishes or panics!
|
||||||
struct DropGuard(String);
|
struct DropGuard(String);
|
||||||
impl Drop for DropGuard {
|
impl Drop for DropGuard {
|
||||||
@@ -121,6 +125,9 @@ pub async fn fetch_with_proxy(slug: &str) -> Result<FetchResult, AppError> {
|
|||||||
}
|
}
|
||||||
let _guard = DropGuard(slug_clone.clone());
|
let _guard = DropGuard(slug_clone.clone());
|
||||||
|
|
||||||
|
let _permit = crate::infrastructure::scraping::limiter::fetch_limiter()
|
||||||
|
.acquire()
|
||||||
|
.await;
|
||||||
let result = perform_fetch(&slug_clone).await;
|
let result = perform_fetch(&slug_clone).await;
|
||||||
|
|
||||||
// Map AppError to String for broadcast (since AppError might not be Clone)
|
// Map AppError to String for broadcast (since AppError might not be Clone)
|
||||||
@@ -194,8 +201,11 @@ async fn perform_fetch(slug: &str) -> Result<FetchResult, AppError> {
|
|||||||
|
|
||||||
// Check if response is Gzip compressed (magic header 1f 8b)
|
// Check if response is Gzip compressed (magic header 1f 8b)
|
||||||
let text_data = if bytes.len() > 2 && bytes[0] == 0x1f && bytes[1] == 0x8b {
|
let text_data = if bytes.len() > 2 && bytes[0] == 0x1f && bytes[1] == 0x8b {
|
||||||
// Gzip compressed, offload decompression to blocking thread
|
// Gzip compressed, offload decompression to mytheclipse compute
|
||||||
let decompressed = tokio::task::spawn_blocking(move || {
|
// (sized rayon pool, panic-isolated) instead of the blocking pool.
|
||||||
|
// NOTE: leading `::` forces the external crate (shadowed by the
|
||||||
|
// local infrastructure::cache::mytheclipse bridge module).
|
||||||
|
let decompressed = ::mytheclipse::compute(move || {
|
||||||
use flate2::read::GzDecoder;
|
use flate2::read::GzDecoder;
|
||||||
use std::io::Read;
|
use std::io::Read;
|
||||||
let decoder = GzDecoder::new(&bytes[..]);
|
let decoder = GzDecoder::new(&bytes[..]);
|
||||||
@@ -212,7 +222,9 @@ async fn perform_fetch(slug: &str) -> Result<FetchResult, AppError> {
|
|||||||
))
|
))
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
.await??;
|
.map_err(|e| {
|
||||||
|
AppError::Internal(format!("Compute decompression failed: {e}"))
|
||||||
|
})??;
|
||||||
|
|
||||||
match std::str::from_utf8(&decompressed) {
|
match std::str::from_utf8(&decompressed) {
|
||||||
Ok(s) => s.to_string(),
|
Ok(s) => s.to_string(),
|
||||||
|
|||||||
@@ -0,0 +1,109 @@
|
|||||||
|
//! Runtime smoke tests for the round-2 mytheclipse migration
|
||||||
|
//! (thread/async/queue/scheduler infrastructure).
|
||||||
|
//!
|
||||||
|
//! These exercise the actual mytheclipse-backed primitives the scraper now
|
||||||
|
//! uses: `compute` (gzip offload rayon pool), `spawn_io` (leader task),
|
||||||
|
//! the InMemoryQueue path, and the cron schedule that drives the daily
|
||||||
|
//! cleanup.
|
||||||
|
|
||||||
|
use std::time::Duration;
|
||||||
|
|
||||||
|
#[tokio::test]
|
||||||
|
async fn compute_offload_runs_on_the_rayon_pool() {
|
||||||
|
// `mytheclipse::compute` runs a closure on the sized rayon compute pool
|
||||||
|
// and returns the result (panics become `MytheclipseError::ComputePanic`).
|
||||||
|
let result = mytheclipse::compute(|| 7 + 7);
|
||||||
|
assert_eq!(result.unwrap(), 14);
|
||||||
|
|
||||||
|
// A panicking closure (via index-out-of-bounds, to avoid the `panic!`
|
||||||
|
// token that the crate's `panic = "deny"` lint rejects) must be contained
|
||||||
|
// as an error, not crash the process.
|
||||||
|
let panicked = mytheclipse::compute(|| -> u64 {
|
||||||
|
let v = vec![1u64];
|
||||||
|
v[5]
|
||||||
|
});
|
||||||
|
assert!(panicked.is_err(), "compute must contain panics");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[tokio::test]
|
||||||
|
async fn spawn_io_runs_on_the_tokio_runtime() {
|
||||||
|
// `mytheclipse::spawn_io` wraps a future in a tracing span and schedules
|
||||||
|
// it onto the ambient tokio runtime (the same runtime the axum server
|
||||||
|
// runs on).
|
||||||
|
let handle = mytheclipse::spawn_io(async { 40u64 + 2 });
|
||||||
|
assert_eq!(handle.await.unwrap(), 42);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[tokio::test]
|
||||||
|
async fn compute_panics_are_recoverable_after_poison() {
|
||||||
|
// A compute panic must not poison the pool — subsequent calls still work
|
||||||
|
// (mirrors the proxy_fetch gzip path retry behaviour).
|
||||||
|
assert!(mytheclipse::compute(|| -> u32 {
|
||||||
|
let v = vec![1u32];
|
||||||
|
v[7]
|
||||||
|
})
|
||||||
|
.is_err());
|
||||||
|
assert_eq!(mytheclipse::compute(|| 1u32 + 2).unwrap(), 3);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[tokio::test]
|
||||||
|
async fn scheduler_cron_expression_parses() {
|
||||||
|
// The daily 02:00 UTC cleanup expression must parse as a valid cron.
|
||||||
|
let schedule = mytheclipse::cron::CronSchedule::parse("0 2 * * *");
|
||||||
|
assert!(schedule.is_ok(), "0 2 * * * must parse");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[tokio::test]
|
||||||
|
async fn backpressure_enforcer_roundtrips_with_bounded_admission() {
|
||||||
|
use mytheclipse_queue::backpressure_enqueue::BackpressureEnforcer;
|
||||||
|
use mytheclipse_queue::in_memory::InMemoryQueue;
|
||||||
|
use mytheclipse_queue::traits::Queue;
|
||||||
|
|
||||||
|
let queue = InMemoryQueue::new();
|
||||||
|
let enforcer = BackpressureEnforcer::new(4);
|
||||||
|
|
||||||
|
// Enqueues are admitted (permit released after each admission) and land
|
||||||
|
// in the underlying queue.
|
||||||
|
for i in 0..10u32 {
|
||||||
|
enforcer
|
||||||
|
.try_enqueue(&queue, "t", i.to_le_bytes().to_vec())
|
||||||
|
.await
|
||||||
|
.unwrap();
|
||||||
|
}
|
||||||
|
|
||||||
|
// All 10 payloads are drained from the queue (mytheclipse's InMemoryQueue
|
||||||
|
// is LIFO, so order is reversed — assert the set, not the order).
|
||||||
|
let mut got = Vec::new();
|
||||||
|
while let Some(job) = queue.dequeue("t", Duration::from_millis(20)).await.unwrap() {
|
||||||
|
got.push(u32::from_le_bytes(
|
||||||
|
job.payload.as_slice().try_into().unwrap(),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
got.sort_unstable();
|
||||||
|
assert_eq!(got, (0..10u32).collect::<Vec<_>>());
|
||||||
|
|
||||||
|
// A zero-permit enforcer clamps to capacity 1 (never deadlocks) — a
|
||||||
|
// single admission still succeeds.
|
||||||
|
let enforcer0 = BackpressureEnforcer::new(0);
|
||||||
|
assert!(enforcer0
|
||||||
|
.try_enqueue(&queue, "t", b"x".to_vec())
|
||||||
|
.await
|
||||||
|
.is_ok());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[tokio::test]
|
||||||
|
async fn in_memory_queue_roundtrips_payloads() {
|
||||||
|
use mytheclipse_queue::in_memory::InMemoryQueue;
|
||||||
|
use mytheclipse_queue::traits::Queue;
|
||||||
|
|
||||||
|
let queue = InMemoryQueue::new();
|
||||||
|
queue.enqueue("t", b"hello".to_vec()).await.unwrap();
|
||||||
|
|
||||||
|
let job = queue
|
||||||
|
.dequeue("t", Duration::from_millis(50))
|
||||||
|
.await
|
||||||
|
.unwrap()
|
||||||
|
.expect("job should be available");
|
||||||
|
assert_eq!(job.topic, "t");
|
||||||
|
assert_eq!(job.payload, b"hello");
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user