The interleaved metadata pass misses mangas to list drift (a title slips a pagination slot during the slow detail walk). Add a reconcile pass: a cheap, full, list-only walk (refs only, no detail visit, no early stop) that set-diffs the walked keys against manga_sources and enqueues the strictly- missing ones as SyncManga jobs. A set-diff is immune to drift — a manga only has to appear somewhere in the list. - Build the previously-dead SyncManga worker by extending RealChapterDispatcher to run the shared pipeline::process_manga_ref (fetch → upsert → cover → chapters), refactored out of run_metadata_pass so both paths stay in lockstep. The crawl worker now leases both sync_chapter_content and sync_manga (jobs::lease_kinds); both serialize on the single exclusive browser. - SyncManga payload carries url + title so the worker can rebuild the ref and the dead-jobs/history UI can label a missing manga that has no manga row yet. - "Missing" = strict NOT EXISTS (dropped rows count as present). Re-enqueue is skipped when a pending/running/dead SyncManga job already exists, so a gone manga (detail 404 → retries → dead) is left dead and not retried. - New POST /v1/admin/crawler/reconcile (fire-and-forget, shares manual_pass_lock) + "Reconcile missing" admin button + Reconciling status phase streamed over SSE. - dead-jobs/history queries surface payload title/url/key; tables fall back to them for sync_manga rows; both searches match the payload title. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
247 lines
8.3 KiB
Rust
247 lines
8.3 KiB
Rust
//! POST /admin/crawler/run — trigger an out-of-cycle metadata pass
|
|
//! POST /admin/crawler/browser/restart — coordinated restart
|
|
//! POST /admin/crawler/session — refresh PHPSESSID
|
|
//! POST /admin/crawler/session/clear-expired — clear sticky expired flag
|
|
//!
|
|
//! All four mutate live in-process state on `CrawlerControl` (browser
|
|
//! manager, session controller, manual-pass mutex). They each emit an
|
|
//! `admin_audit` row and `status.poke()` so SSE subscribers see the
|
|
//! change instantly without polling.
|
|
|
|
use axum::extract::State;
|
|
use axum::routing::post;
|
|
use axum::{Json, Router};
|
|
use serde::{Deserialize, Serialize};
|
|
use serde_json::json;
|
|
|
|
use crate::app::AppState;
|
|
use crate::auth::extractor::RequireAdmin;
|
|
use crate::error::{AppError, AppResult};
|
|
use crate::repo;
|
|
|
|
use super::require_crawler;
|
|
|
|
pub(super) fn routes() -> Router<AppState> {
|
|
Router::new()
|
|
.route("/admin/crawler/run", post(run_now))
|
|
.route("/admin/crawler/reconcile", post(reconcile_now))
|
|
.route("/admin/crawler/browser/restart", post(restart_browser))
|
|
.route("/admin/crawler/session", post(update_session))
|
|
.route(
|
|
"/admin/crawler/session/clear-expired",
|
|
post(clear_session_expired),
|
|
)
|
|
}
|
|
|
|
#[derive(Debug, Serialize)]
|
|
struct RunResponse {
|
|
started: bool,
|
|
}
|
|
|
|
async fn run_now(
|
|
State(state): State<AppState>,
|
|
admin: RequireAdmin,
|
|
) -> AppResult<Json<RunResponse>> {
|
|
let c = require_crawler(&state)?;
|
|
let mp = c.metadata_pass.as_ref().ok_or_else(|| {
|
|
AppError::ServiceUnavailable("no source configured (CRAWLER_START_URL unset)".into())
|
|
})?;
|
|
// Operator-click dedup. A pass holds the lock for its entire run
|
|
// (minutes); a second click while it's in flight returns 409 instead
|
|
// of fanning a second pass onto the single browser lease. The daily
|
|
// cron does NOT take this lock — cron is single-fire by definition,
|
|
// and its own contention with a manual pass already serialises
|
|
// through the browser lease + advisory lock at a lower layer.
|
|
let pass_guard = c
|
|
.manual_pass_lock
|
|
.clone()
|
|
.try_lock_owned()
|
|
.map_err(|_| AppError::Conflict("manual metadata pass already running".into()))?;
|
|
let mp = std::sync::Arc::clone(mp);
|
|
// Fire-and-forget: the pass can run for minutes; the dashboard
|
|
// streams progress over SSE. The guard moves into the task so the
|
|
// lock is released only when the pass finishes (or the task panics).
|
|
tokio::spawn(async move {
|
|
let _pass_guard = pass_guard;
|
|
if let Err(e) = mp.run().await {
|
|
tracing::warn!(error = ?e, "manual metadata pass failed");
|
|
}
|
|
});
|
|
repo::admin_audit::insert(&state.db, admin.0.id, "crawler_run", "crawler", None, json!({}))
|
|
.await?;
|
|
Ok(Json(RunResponse { started: true }))
|
|
}
|
|
|
|
async fn reconcile_now(
|
|
State(state): State<AppState>,
|
|
admin: RequireAdmin,
|
|
) -> AppResult<Json<RunResponse>> {
|
|
let c = require_crawler(&state)?;
|
|
let rp = c.reconcile_pass.as_ref().ok_or_else(|| {
|
|
AppError::ServiceUnavailable("no source configured (CRAWLER_START_URL unset)".into())
|
|
})?;
|
|
// Share `manual_pass_lock` with the metadata pass: reconcile's list walk
|
|
// and a metadata pass both contend for the single browser lease, so they
|
|
// must not run concurrently. A click while either is in flight gets 409.
|
|
let pass_guard = c
|
|
.manual_pass_lock
|
|
.clone()
|
|
.try_lock_owned()
|
|
.map_err(|_| AppError::Conflict("a metadata or reconcile pass is already running".into()))?;
|
|
let rp = std::sync::Arc::clone(rp);
|
|
// Fire-and-forget: the full list walk runs for minutes; progress streams
|
|
// over SSE (the Reconciling phase). The guard moves into the task so the
|
|
// lock releases only when the walk + enqueue finish.
|
|
tokio::spawn(async move {
|
|
let _pass_guard = pass_guard;
|
|
match rp.run().await {
|
|
Ok(stats) => tracing::info!(?stats, "manual reconcile pass complete"),
|
|
Err(e) => tracing::warn!(error = ?e, "manual reconcile pass failed"),
|
|
}
|
|
});
|
|
repo::admin_audit::insert(
|
|
&state.db,
|
|
admin.0.id,
|
|
"crawler_reconcile",
|
|
"crawler",
|
|
None,
|
|
json!({}),
|
|
)
|
|
.await?;
|
|
Ok(Json(RunResponse { started: true }))
|
|
}
|
|
|
|
#[derive(Debug, Serialize)]
|
|
struct RestartResponse {
|
|
ok: bool,
|
|
error: Option<String>,
|
|
}
|
|
|
|
async fn restart_browser(
|
|
State(state): State<AppState>,
|
|
admin: RequireAdmin,
|
|
) -> AppResult<Json<RestartResponse>> {
|
|
let c = require_crawler(&state)?;
|
|
let result = c.browser_manager.coordinated_restart(c.drain_deadline).await;
|
|
// A successful coordinated_restart re-runs on_launch, which re-injects
|
|
// PHPSESSID and re-probes — i.e. the session is live. Drop the sticky
|
|
// `session_expired` flag so chapter workers stop idling without
|
|
// requiring a second click on "Clear expired".
|
|
if result.is_ok() {
|
|
c.session.clear_expired();
|
|
}
|
|
// Push the post-restart browser phase to live subscribers immediately.
|
|
c.status.poke();
|
|
repo::admin_audit::insert(
|
|
&state.db,
|
|
admin.0.id,
|
|
"crawler_browser_restart",
|
|
"crawler",
|
|
None,
|
|
json!({ "ok": result.is_ok() }),
|
|
)
|
|
.await?;
|
|
Ok(Json(match result {
|
|
Ok(()) => RestartResponse {
|
|
ok: true,
|
|
error: None,
|
|
},
|
|
Err(e) => RestartResponse {
|
|
ok: false,
|
|
error: Some(format!("{e:#}")),
|
|
},
|
|
}))
|
|
}
|
|
|
|
#[derive(Debug, Deserialize)]
|
|
struct UpdateSessionRequest {
|
|
phpsessid: String,
|
|
}
|
|
|
|
#[derive(Debug, Serialize)]
|
|
struct UpdateSessionResponse {
|
|
/// Whether the post-update browser relaunch + session probe succeeded.
|
|
valid: bool,
|
|
error: Option<String>,
|
|
}
|
|
|
|
async fn update_session(
|
|
State(state): State<AppState>,
|
|
admin: RequireAdmin,
|
|
Json(body): Json<UpdateSessionRequest>,
|
|
) -> AppResult<Json<UpdateSessionResponse>> {
|
|
let c = require_crawler(&state)?;
|
|
// Fingerprint BEFORE move so the raw value never reaches tracing or
|
|
// the audit row. SHA256-prefix is opaque to anyone reading the audit
|
|
// log but deterministic enough to correlate two updates of the same
|
|
// session value.
|
|
let fingerprint = phpsessid_fingerprint(&body.phpsessid);
|
|
c.session
|
|
.update(&body.phpsessid)
|
|
.await
|
|
.map_err(|e| AppError::InvalidInput(format!("{e:#}")))?;
|
|
// Relaunch the browser so on_launch re-injects the new cookie and
|
|
// re-probes — the restart's success IS the session-validity signal.
|
|
let probe = c.browser_manager.coordinated_restart(c.drain_deadline).await;
|
|
// Session + browser state changed — push to live subscribers.
|
|
c.status.poke();
|
|
repo::admin_audit::insert(
|
|
&state.db,
|
|
admin.0.id,
|
|
"crawler_session_update",
|
|
"crawler",
|
|
None,
|
|
json!({ "valid": probe.is_ok(), "phpsessid_fingerprint": fingerprint }),
|
|
)
|
|
.await?;
|
|
Ok(Json(match probe {
|
|
Ok(()) => UpdateSessionResponse {
|
|
valid: true,
|
|
error: None,
|
|
},
|
|
Err(e) => UpdateSessionResponse {
|
|
valid: false,
|
|
error: Some(format!("{e:#}")),
|
|
},
|
|
}))
|
|
}
|
|
|
|
#[derive(Debug, Serialize)]
|
|
struct ClearExpiredResponse {
|
|
cleared: bool,
|
|
}
|
|
|
|
async fn clear_session_expired(
|
|
State(state): State<AppState>,
|
|
admin: RequireAdmin,
|
|
) -> AppResult<Json<ClearExpiredResponse>> {
|
|
let c = require_crawler(&state)?;
|
|
c.session.clear_expired();
|
|
// session.expired flipped — push to live subscribers.
|
|
c.status.poke();
|
|
repo::admin_audit::insert(
|
|
&state.db,
|
|
admin.0.id,
|
|
"crawler_session_clear_expired",
|
|
"crawler",
|
|
None,
|
|
json!({}),
|
|
)
|
|
.await?;
|
|
Ok(Json(ClearExpiredResponse { cleared: true }))
|
|
}
|
|
|
|
/// Opaque, short fingerprint of a PHPSESSID for the admin audit log.
|
|
/// The first 8 hex chars of SHA-256 — enough to correlate two updates
|
|
/// of the same value without revealing the raw cookie. Reading the audit
|
|
/// row does not give an operator anything that can re-construct the
|
|
/// session.
|
|
fn phpsessid_fingerprint(sid: &str) -> String {
|
|
use sha2::{Digest, Sha256};
|
|
let mut h = Sha256::new();
|
|
h.update(sid.as_bytes());
|
|
let digest = h.finalize();
|
|
let hex: String = digest.iter().take(4).map(|b| format!("{b:02x}")).collect();
|
|
hex
|
|
}
|