Admin-only crawler dashboard backed by an SSE live-status stream,
coordinated browser restart, runtime PHPSESSID refresh, dead-letter
requeue, and a batch of reliability fixes. Closes everything from
the two-pass audit (10 commits' worth) and bumps 0.52.0 -> 0.55.0.
Backend:
- New /admin/crawler/* surface (cookie-auth, RequireAdmin) split
into status / control / dead_jobs / backlog modules. SSE stream
composes in-memory status with DB-derived queue counts, memoizes
the counts for 1s and debounces watch pokes for 250ms (~10x QPS
reduction per subscriber). One-shot GET /admin/crawler shares the
same compose path.
- POST /admin/crawler/run gated by manual_pass_lock try_lock_owned
(409 Conflict on overlapping click); browser restart goes through
the coordinated_restart gate (drain + relaunch + auto-clear of the
sticky session_expired flag on Ok).
- Runtime PHPSESSID refresh via SessionController (allow-list
validation, never logged, audit row carries SHA-256 fingerprint).
Storage layer is repo::crawler::runtime_session_{load,persist}.
- Dead-letter requeue with four scopes (all/manga/chapter/job);
scope=all requires confirm:true; DISTINCT ON dedup keeps the
partial unique index from rejecting requeues for chapters with
multiple dead rows. SQL is four &'static str constants per scope.
- StatusHandle + ChapterGuard / CoverGuard RAII model survives
panics; last-writer-wins on cover so concurrent dispatches don't
clobber each other's slot. Pure functions (should_stop /
should_mark_clean_exit / should_abort_pass) with named regression
tests.
- Reliability bundle: per-lease heartbeat, jitter on retries,
per-job timeout, circuit breaker on consecutive failures, BrowserManager
coordinated restart gate, request fingerprint changes.
- Streaming page download: Storage::put_stream trait method,
LocalStorage impl atomic via temp + fsync + UUID-suffixed rename.
Pages stream through with peak memory ~one HTTP chunk + 64-byte
sniff prefix instead of one full image per dispatch.
- New partial indexes (migration 0022): mangas_missing_cover_idx
and crawler_jobs_dead_idx, both ordered by updated_at DESC to
match the dashboard's LIMIT/OFFSET reads.
- Security hardening: admin_csrf_guard (Origin/Referer allowlist
on /admin/* mutations, opt-in via ADMIN_ALLOWED_ORIGINS),
admin_no_store_guard (Cache-Control: no-store on admin
responses), audit rows carry per-scope target_id.
Frontend:
- /admin/crawler page decomposed into lib/components/crawler/
(11 components: ProgressBar, SearchBar, CrawlerHero,
CrawlerControls, ActiveChaptersCard, ActiveJobsTable,
MissingCoversTable, DeadJobsTable, RestartConfirmModal,
RequeueAllConfirmModal, SessionModal). Page is 532 LOC of
orchestration; each component 22-148 LOC.
- EventSource lifecycle wired to visibilitychange / pagehide /
pageshow (BFCache); after 5 consecutive errors probes the status
endpoint so a 401 routes through the global on401Hook instead of
infinite silent reconnects.
- Backlog $effect refetches debounced 500ms with per-loader
AbortControllers; refresh after a control action only runs when
the SSE stream is dead.
- Inline requeue button on /admin/mangas patches the affected row's
sync_state locally (no full chapter-list refetch); proper
aria-label. Requeue-all gets its own confirm modal; both confirm
modals autofocus Cancel.
- SvelteKit reverse proxy bypasses its 5-minute AbortController
for Accept: text/event-stream; pure shouldBypassProxyTimeout
helper covered by unit tests.
Config / docs:
- New env vars (.env.example): ADMIN_ALLOWED_ORIGINS,
CRAWLER_JOB_TIMEOUT_SECS, CRAWLER_METADATA_MAX_CONSECUTIVE_FAILURES,
CRAWLER_BROWSER_RESTART_THRESHOLD.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
176 lines
5.8 KiB
Rust
176 lines
5.8 KiB
Rust
//! Chapter persistence.
|
|
|
|
use sqlx::{PgExecutor, PgPool};
|
|
use uuid::Uuid;
|
|
|
|
use crate::domain::Chapter;
|
|
use crate::error::AppResult;
|
|
|
|
pub async fn list_for_manga(
|
|
pool: &PgPool,
|
|
manga_id: Uuid,
|
|
limit: i64,
|
|
offset: i64,
|
|
) -> AppResult<Vec<Chapter>> {
|
|
// Display order = source-site order reversed. The crawler stamps
|
|
// `source_index` = position in the source DOM (0 = first = newest
|
|
// on this site, see migration 0021), so DESC puts the oldest
|
|
// chapter first and keeps the site's variant grouping and the
|
|
// placement of non-numeric entries (e.g. "notice. : Officials")
|
|
// intact. NULLS LAST keeps user-uploaded chapters (no source row)
|
|
// and rows that pre-date the migration below crawled rows; the
|
|
// (number, created_at) tail then orders them deterministically.
|
|
let rows = sqlx::query_as::<_, Chapter>(
|
|
r#"
|
|
SELECT id, manga_id, number, title, page_count, created_at
|
|
FROM chapters
|
|
WHERE manga_id = $1
|
|
ORDER BY source_index DESC NULLS LAST, number ASC, created_at ASC
|
|
LIMIT $2 OFFSET $3
|
|
"#,
|
|
)
|
|
.bind(manga_id)
|
|
.bind(limit)
|
|
.bind(offset)
|
|
.fetch_all(pool)
|
|
.await?;
|
|
Ok(rows)
|
|
}
|
|
|
|
/// Look up a chapter by its UUID, scoped to its manga so a UUID guessed
|
|
/// from a different manga's URL doesn't accidentally resolve.
|
|
pub async fn find_by_id_in_manga(
|
|
pool: &PgPool,
|
|
manga_id: Uuid,
|
|
chapter_id: Uuid,
|
|
) -> AppResult<Option<Chapter>> {
|
|
let row = sqlx::query_as::<_, Chapter>(
|
|
r#"
|
|
SELECT id, manga_id, number, title, page_count, created_at
|
|
FROM chapters
|
|
WHERE manga_id = $1 AND id = $2
|
|
"#,
|
|
)
|
|
.bind(manga_id)
|
|
.bind(chapter_id)
|
|
.fetch_optional(pool)
|
|
.await?;
|
|
Ok(row)
|
|
}
|
|
|
|
/// Accepts any `PgExecutor` so the upload handler can run this inside a
|
|
/// transaction with the per-page inserts.
|
|
///
|
|
/// `uploaded_by` records who uploaded the chapter and feeds the
|
|
/// per-user upload history. `None` means "historical / API token with
|
|
/// no associated user" — kept nullable to support that case.
|
|
///
|
|
/// Chapter identity is the row UUID; the same (manga_id, number)
|
|
/// combination can repeat (multiple translations, re-uploads). The
|
|
/// 0013 migration dropped the (manga_id, number) UNIQUE, so duplicate
|
|
/// inserts succeed by design. If a future migration re-adds any
|
|
/// uniqueness, surface a 409 by adding a unique-violation arm here.
|
|
pub async fn create<'e, E: PgExecutor<'e>>(
|
|
executor: E,
|
|
manga_id: Uuid,
|
|
number: i32,
|
|
title: Option<&str>,
|
|
uploaded_by: Option<Uuid>,
|
|
) -> AppResult<Chapter> {
|
|
let row = sqlx::query_as::<_, Chapter>(
|
|
r#"
|
|
INSERT INTO chapters (manga_id, number, title, uploaded_by)
|
|
VALUES ($1, $2, $3, $4)
|
|
RETURNING id, manga_id, number, title, page_count, created_at
|
|
"#,
|
|
)
|
|
.bind(manga_id)
|
|
.bind(number)
|
|
.bind(title)
|
|
.bind(uploaded_by)
|
|
.fetch_one(executor)
|
|
.await?;
|
|
Ok(row)
|
|
}
|
|
|
|
/// Cross-link guard for `POST /bookmarks`: the bookmarks FK accepts
|
|
/// any valid chapter id, but a chapter must belong to the bookmark's
|
|
/// manga or the bookmark would dangle on a foreign manga. Handlers
|
|
/// call this before the insert and surface `NotFound` when it
|
|
/// returns `false`.
|
|
pub async fn belongs_to_manga(
|
|
pool: &PgPool,
|
|
chapter_id: Uuid,
|
|
manga_id: Uuid,
|
|
) -> AppResult<bool> {
|
|
let (exists,): (bool,) = sqlx::query_as(
|
|
"SELECT EXISTS(SELECT 1 FROM chapters WHERE id = $1 AND manga_id = $2)",
|
|
)
|
|
.bind(chapter_id)
|
|
.bind(manga_id)
|
|
.fetch_one(pool)
|
|
.await?;
|
|
Ok(exists)
|
|
}
|
|
|
|
/// Read just the page_count for a chapter. Used by the crawler
|
|
/// daemon's consumer-side dedup safety net so it can ack-done a job
|
|
/// whose chapter has already been fetched by a racing worker.
|
|
pub async fn page_count(pool: &PgPool, id: Uuid) -> sqlx::Result<Option<i32>> {
|
|
sqlx::query_scalar("SELECT page_count FROM chapters WHERE id = $1")
|
|
.bind(id)
|
|
.fetch_optional(pool)
|
|
.await
|
|
}
|
|
|
|
/// Look up the manga_id + most recent live source_url for a chapter.
|
|
/// Used by the daemon's chapter dispatcher to resolve the URL it needs
|
|
/// to hand to `content::sync_chapter_content`.
|
|
///
|
|
/// Skips soft-dropped sources (`cs.dropped_at IS NOT NULL`) and breaks
|
|
/// ties between multiple live sources by `last_seen_at DESC`, so the
|
|
/// freshest still-attached URL wins. Returns `None` when the chapter
|
|
/// is gone or all its source rows are dropped — callers in the
|
|
/// dispatcher treat `None` as "ack the job, skip the work."
|
|
///
|
|
/// The enqueue queries (`pipeline::enqueue_bookmarked_pending` and
|
|
/// `enqueue_pending_for_manga`) apply the same `dropped_at IS NULL`
|
|
/// filter — this resolver stays in lockstep so a chapter that was
|
|
/// dropped between enqueue and lease isn't dispatched against a stale
|
|
/// URL.
|
|
/// Returns `(manga_id, source_url, manga_title, chapter_number)`. The
|
|
/// title + number feed the live "currently crawling" status; the rest is
|
|
/// what the dispatcher needs to do the work.
|
|
pub async fn dispatch_target(
|
|
pool: &PgPool,
|
|
chapter_id: Uuid,
|
|
) -> sqlx::Result<Option<(Uuid, String, String, i32)>> {
|
|
sqlx::query_as(
|
|
"SELECT c.manga_id, cs.source_url, m.title, c.number \
|
|
FROM chapters c \
|
|
JOIN chapter_sources cs ON cs.chapter_id = c.id \
|
|
JOIN mangas m ON m.id = c.manga_id \
|
|
WHERE c.id = $1 \
|
|
AND cs.dropped_at IS NULL \
|
|
ORDER BY cs.last_seen_at DESC \
|
|
LIMIT 1",
|
|
)
|
|
.bind(chapter_id)
|
|
.fetch_optional(pool)
|
|
.await
|
|
}
|
|
|
|
pub async fn set_page_count<'e, E: PgExecutor<'e>>(
|
|
executor: E,
|
|
id: Uuid,
|
|
page_count: i32,
|
|
) -> AppResult<()> {
|
|
sqlx::query("UPDATE chapters SET page_count = $1 WHERE id = $2")
|
|
.bind(page_count)
|
|
.bind(id)
|
|
.execute(executor)
|
|
.await?;
|
|
Ok(())
|
|
}
|
|
|