fix(maintenance): reclaim the media of deliberately deleted uploads
The quota stopped bounding the disk. `soft_delete_in_event` stamps `deleted_at` and refunds `total_upload_bytes`, but nothing ever removed the bytes, and the hourly sweep reached only `compression_status = 'failed'`. Upload 500 MB, delete, quota back to zero, upload another 500 MB. Not an attack -- a guest curating their camera roll, which is what people do. The host then sees guests hitting "Du hast dein Upload-Limit erreicht" while the admin widget shows a disk full of files no upload row points at, and the quota message is actively misleading because the space really is gone, just not to anyone the accounting can name. Two retention windows, because the two deletes mean different things. A compression failure keeps its 14 days: the guest didn't ask for it and may not be able to retake the photo. A deliberate removal gets 24 hours -- 14 days outlives the whole event, so a deliberate delete would never reclaim anything while it mattered, and a day still covers a mis-tap. Wider than reported: ALL FOUR paths are reclaimed, not just the original. Preview, display and thumbnail are each a separate file, none counted in `original_size_bytes`, and nothing ever removed them either. That was invisible while the sweep only saw failed compressions (which produce no derivatives) and becomes three leaked files per upload the moment it reaches a successful one. A row is re-selected until every path is cleared, and the columns are cleared only once every file for that upload is gone -- clearing after a partial success would strand the survivors in exactly the unowned state this drains. `backfill_stale_derivatives` selects on `display_path IS NULL AND preview_path IS NOT NULL`, which is close enough to the post-sweep state to be worth pinning: it is guarded on `deleted_at IS NULL`, so it cannot re-decode an original that is no longer on disk. Covered. Residual, deliberately: within the 24h window the bytes are still spent and still unaccounted, so delete-and-re-upload through an eight-hour event can outrun the sweep. Bounding that means holding the quota until the file is reclaimed rather than refunding at `deleted_at`. The low-disk warning is the net under it. Tests: 6 DB-backed, replacing 3. The one asserting an owner-deleted upload IS reclaimed is the exact inverse of what this file used to assert. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,19 +1,27 @@
|
||||
//! DB-backed tests for the failed-original sweep (`services/maintenance.rs`).
|
||||
//! DB-backed tests for the deleted-media sweep (`services/maintenance.rs`).
|
||||
//!
|
||||
//! Context: the compression worker deliberately no longer deletes an upload's original when
|
||||
//! its transcode fails — a transient ENOSPC or a codec panic must never destroy the only
|
||||
//! copy of a photo a guest cannot retake. But the row is soft-deleted and the uploader's
|
||||
//! quota IS refunded, so those bytes become invisible, unowned and free. A guest hitting a
|
||||
//! reproducible codec failure could accumulate orphans at no personal cost, and since
|
||||
//! `active_uploaders` counts only users with non-deleted uploads, dropping out of that count
|
||||
//! actually RAISES everyone's per-user ceiling while the disk fills.
|
||||
//! Context, in two halves.
|
||||
//!
|
||||
//! The sweep reclaims them after a retention window. Its selection predicate is the whole
|
||||
//! safety argument — it must reach the give-up path's leftovers and nothing else — so that
|
||||
//! is what these tests pin, following the same "reproduce the SQL verbatim" pattern as
|
||||
//! `upload_concurrency.rs`.
|
||||
//! The compression worker deliberately no longer deletes an upload's original when its transcode
|
||||
//! fails — a transient ENOSPC or a codec panic must never destroy the only copy of a photo a guest
|
||||
//! cannot retake. But the row is soft-deleted and the uploader's quota IS refunded, so those bytes
|
||||
//! become invisible, unowned and free.
|
||||
//!
|
||||
//! `#[sqlx::test]` gives each test a fresh database with the real migrations applied.
|
||||
//! The SAME hole was reachable by the ordinary path, and that one is not an edge case at all:
|
||||
//! `soft_delete_in_event` refunds `total_upload_bytes` on every guest or host delete and nothing
|
||||
//! removed the files, so the quota stopped bounding the disk. Upload 500 MB, delete, quota back to
|
||||
//! zero, upload another 500 MB — a guest curating their camera roll, which is what people do. The
|
||||
//! sweep used to reach only `compression_status = 'failed'`, so it never touched this case; the
|
||||
//! test below that now asserts an owner-deleted upload IS reclaimed is the one that used to assert
|
||||
//! the opposite.
|
||||
//!
|
||||
//! Two windows, because the two deletes mean different things: 14 days for a failure an operator
|
||||
//! may want to investigate, 24 hours for a removal someone asked for (14 days outlives the whole
|
||||
//! event, so a deliberate delete would never reclaim anything while it mattered).
|
||||
//!
|
||||
//! The selection predicate is the whole safety argument — it must reach both leftovers and never a
|
||||
//! live upload — so that is what these pin, following the same "reproduce the SQL verbatim" pattern
|
||||
//! as `upload_concurrency.rs`. `#[sqlx::test]` gives each test a fresh, migrated database.
|
||||
|
||||
mod common;
|
||||
|
||||
@@ -21,195 +29,324 @@ use common::*;
|
||||
use sqlx::PgPool;
|
||||
use uuid::Uuid;
|
||||
|
||||
/// SRC: `services/maintenance.rs::cleanup_failed_originals` — the selection, verbatim.
|
||||
async fn sweep_selects(pool: &PgPool, retention_days: i64) -> Vec<Uuid> {
|
||||
sqlx::query_as::<_, (Uuid, String)>(
|
||||
"SELECT id, original_path FROM upload
|
||||
WHERE compression_status = 'failed'
|
||||
AND deleted_at IS NOT NULL
|
||||
AND deleted_at < NOW() - ($1 || ' days')::interval
|
||||
AND original_path <> ''",
|
||||
const FAILED_DAYS: i64 = 14;
|
||||
const DELETED_HOURS: i64 = 24;
|
||||
|
||||
/// SRC: `services/maintenance.rs::cleanup_deleted_media` — the selection, verbatim.
|
||||
async fn sweep_selects(pool: &PgPool, failed_days: i64, deleted_hours: i64) -> Vec<Uuid> {
|
||||
type Row = (Uuid, String, Option<String>, Option<String>, Option<String>);
|
||||
sqlx::query_as::<_, Row>(
|
||||
"SELECT id, original_path, preview_path, display_path, thumbnail_path FROM upload
|
||||
WHERE deleted_at IS NOT NULL
|
||||
AND CASE WHEN compression_status = 'failed'
|
||||
THEN deleted_at < NOW() - ($1 || ' days')::interval
|
||||
ELSE deleted_at < NOW() - ($2 || ' hours')::interval
|
||||
END
|
||||
AND (original_path <> '' OR preview_path IS NOT NULL
|
||||
OR display_path IS NOT NULL OR thumbnail_path IS NOT NULL)",
|
||||
)
|
||||
.bind(retention_days.to_string())
|
||||
.bind(failed_days.to_string())
|
||||
.bind(deleted_hours.to_string())
|
||||
.fetch_all(pool)
|
||||
.await
|
||||
.expect("sweep query")
|
||||
.into_iter()
|
||||
.map(|(id, _)| id)
|
||||
.map(|(id, ..)| id)
|
||||
.collect()
|
||||
}
|
||||
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
async fn seed_upload(
|
||||
/// Seed an upload aged `deleted_hours_ago` (None = live), with optional derivative paths.
|
||||
async fn seed_aged_upload(
|
||||
pool: &PgPool,
|
||||
event_id: Uuid,
|
||||
user_id: Uuid,
|
||||
status: &str,
|
||||
deleted_days_ago: Option<i64>,
|
||||
deleted_hours_ago: Option<i64>,
|
||||
original_path: &str,
|
||||
derivatives: bool,
|
||||
) -> Uuid {
|
||||
let id: Uuid = sqlx::query_scalar(
|
||||
sqlx::query_scalar(
|
||||
"INSERT INTO upload (event_id, user_id, original_path, mime_type, original_size_bytes,
|
||||
compression_status, deleted_at)
|
||||
compression_status, deleted_at,
|
||||
preview_path, display_path, thumbnail_path)
|
||||
VALUES ($1, $2, $3, 'image/jpeg', 1000, $4,
|
||||
CASE WHEN $5::bigint IS NULL THEN NULL
|
||||
ELSE NOW() - ($5::text || ' days')::interval END)
|
||||
ELSE NOW() - ($5::text || ' hours')::interval END,
|
||||
CASE WHEN $6 THEN 'previews/p.jpg' END,
|
||||
CASE WHEN $6 THEN 'displays/d.jpg' END,
|
||||
CASE WHEN $6 THEN 'thumbs/t.jpg' END)
|
||||
RETURNING id",
|
||||
)
|
||||
.bind(event_id)
|
||||
.bind(user_id)
|
||||
.bind(original_path)
|
||||
.bind(status)
|
||||
.bind(deleted_days_ago)
|
||||
.bind(deleted_hours_ago)
|
||||
.bind(derivatives)
|
||||
.fetch_one(pool)
|
||||
.await
|
||||
.expect("seed upload");
|
||||
id
|
||||
.expect("seed upload")
|
||||
}
|
||||
|
||||
/// A live upload is untouchable no matter how the windows are configured.
|
||||
///
|
||||
/// PREVENTS: the catastrophic loosening. Everything else here is about reclaiming more; this is the
|
||||
/// one assertion that must never bend.
|
||||
#[sqlx::test]
|
||||
async fn sweeps_only_long_failed_soft_deleted_uploads(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-event").await;
|
||||
async fn a_live_upload_is_never_selected(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-live").await;
|
||||
let user_id = seed_user(&pool, event_id, "Sweeper").await;
|
||||
|
||||
// The one and only thing the sweep may touch: the exact state the compression worker's
|
||||
// give-up path leaves behind, aged past the window.
|
||||
let target = seed_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
Some(30),
|
||||
"originals/e/target.jpg",
|
||||
)
|
||||
.await;
|
||||
|
||||
// Everything below is a near-miss that must survive.
|
||||
// A healthy live upload — the catastrophic case if the predicate were ever loosened.
|
||||
let live = seed_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"done",
|
||||
None,
|
||||
"originals/e/live.jpg",
|
||||
)
|
||||
.await;
|
||||
// Failed but still inside the retention window: the recovery window is the entire point
|
||||
// of keeping the file, so reclaiming it early would defeat the fix it protects.
|
||||
let recent = seed_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
Some(1),
|
||||
"originals/e/recent.jpg",
|
||||
)
|
||||
.await;
|
||||
// Failed but NOT soft-deleted — not the give-up path; something else set this status.
|
||||
let failed_live = seed_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
None,
|
||||
"originals/e/failed-live.jpg",
|
||||
)
|
||||
.await;
|
||||
// Soft-deleted by the OWNER, compression fine. Its file is already gone; this row must
|
||||
// never be re-processed.
|
||||
let owner_deleted = seed_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"done",
|
||||
Some(30),
|
||||
"originals/e/owner.jpg",
|
||||
)
|
||||
.await;
|
||||
// Already swept: `original_path` cleared. Re-selecting it every hour would log a
|
||||
// phantom reclaim forever.
|
||||
let already_swept = seed_upload(&pool, event_id, user_id, "failed", Some(30), "").await;
|
||||
|
||||
let selected = sweep_selects(&pool, 14).await;
|
||||
|
||||
assert_eq!(
|
||||
selected,
|
||||
vec![target],
|
||||
"the sweep must select exactly the aged give-up-path leftovers"
|
||||
);
|
||||
for (id, what) in [
|
||||
(live, "a live upload"),
|
||||
(recent, "a failure still inside the retention window"),
|
||||
(failed_live, "a failed but not soft-deleted upload"),
|
||||
(owner_deleted, "an owner-deleted upload"),
|
||||
(already_swept, "an already-swept row"),
|
||||
] {
|
||||
assert!(!selected.contains(&id), "the sweep must not touch {what}");
|
||||
for status in ["done", "failed", "processing", "pending"] {
|
||||
let live = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
status,
|
||||
None,
|
||||
"originals/e/live.jpg",
|
||||
true,
|
||||
)
|
||||
.await;
|
||||
assert!(
|
||||
!sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
|
||||
.await
|
||||
.contains(&live),
|
||||
"a non-deleted upload with status {status} must never be swept"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// THE FIX. An upload a guest or host deliberately deleted is reclaimed once past 24 hours.
|
||||
///
|
||||
/// PREVENTS: the regression back to a sweep scoped to `compression_status = 'failed'`, which is
|
||||
/// what let the quota stop bounding the disk. This assertion is the inverse of the one this file
|
||||
/// used to make.
|
||||
#[sqlx::test]
|
||||
async fn retention_window_is_honoured_at_the_boundary(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-boundary").await;
|
||||
let user_id = seed_user(&pool, event_id, "Boundary").await;
|
||||
async fn a_deliberately_deleted_upload_is_reclaimed_after_a_day(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-deleted").await;
|
||||
let user_id = seed_user(&pool, event_id, "Curator").await;
|
||||
|
||||
let inside = seed_upload(
|
||||
let deleted = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
Some(13),
|
||||
"originals/e/inside.jpg",
|
||||
"done",
|
||||
Some(48),
|
||||
"originals/e/owner.jpg",
|
||||
true,
|
||||
)
|
||||
.await;
|
||||
let outside = seed_upload(
|
||||
// Still inside the window — a mis-tap is recoverable for a day.
|
||||
let recent = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
Some(15),
|
||||
"originals/e/outside.jpg",
|
||||
"done",
|
||||
Some(2),
|
||||
"originals/e/recent.jpg",
|
||||
true,
|
||||
)
|
||||
.await;
|
||||
|
||||
let selected = sweep_selects(&pool, 14).await;
|
||||
let selected = sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await;
|
||||
assert!(
|
||||
selected.contains(&outside),
|
||||
"15 days old must be past a 14-day window"
|
||||
selected.contains(&deleted),
|
||||
"a deliberate delete past the window must be reclaimed — this is the leak"
|
||||
);
|
||||
assert!(
|
||||
!selected.contains(&inside),
|
||||
"13 days old must still be retained"
|
||||
!selected.contains(&recent),
|
||||
"a delete inside the window keeps its recovery grace"
|
||||
);
|
||||
}
|
||||
|
||||
/// The two windows are independent: a failure is retained far longer than a deliberate delete.
|
||||
///
|
||||
/// PREVENTS: collapsing them into one. Applying 24h to failures would destroy the recovery window
|
||||
/// the retained-original fix exists to provide; applying 14 days to deliberate deletes would mean
|
||||
/// nothing is ever reclaimed during an event.
|
||||
#[sqlx::test]
|
||||
async fn clearing_original_path_makes_the_sweep_idempotent(pool: PgPool) {
|
||||
// The sweep clears `original_path` after reclaiming the file. Without that, a row whose
|
||||
// file is already gone is re-selected on every hourly tick forever.
|
||||
let event_id = seed_event(&pool, "sweep-idempotent").await;
|
||||
let user_id = seed_user(&pool, event_id, "Idem").await;
|
||||
let id = seed_upload(
|
||||
async fn the_two_retention_windows_do_not_bleed_into_each_other(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-windows").await;
|
||||
let user_id = seed_user(&pool, event_id, "Windows").await;
|
||||
|
||||
// 48h old: past the deliberate window, nowhere near the failure window.
|
||||
let failed_recent = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
Some(30),
|
||||
"originals/e/once.jpg",
|
||||
Some(48),
|
||||
"originals/e/f-recent.jpg",
|
||||
false,
|
||||
)
|
||||
.await;
|
||||
let deleted_same_age = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"done",
|
||||
Some(48),
|
||||
"originals/e/d-same.jpg",
|
||||
false,
|
||||
)
|
||||
.await;
|
||||
// 30 days old: past both.
|
||||
let failed_old = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"failed",
|
||||
Some(30 * 24),
|
||||
"originals/e/f-old.jpg",
|
||||
false,
|
||||
)
|
||||
.await;
|
||||
|
||||
assert_eq!(sweep_selects(&pool, 14).await, vec![id]);
|
||||
let selected = sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await;
|
||||
assert!(
|
||||
!selected.contains(&failed_recent),
|
||||
"a 2-day-old compression failure is still inside its 14-day recovery window"
|
||||
);
|
||||
assert!(
|
||||
selected.contains(&deleted_same_age),
|
||||
"a deliberate delete of the same age is past its 24-hour window"
|
||||
);
|
||||
assert!(
|
||||
selected.contains(&failed_old),
|
||||
"a 30-day-old failure is past both windows"
|
||||
);
|
||||
}
|
||||
|
||||
/// Boundary behaviour on both windows.
|
||||
#[sqlx::test]
|
||||
async fn retention_windows_are_honoured_at_the_boundary(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-boundary").await;
|
||||
let user_id = seed_user(&pool, event_id, "Boundary").await;
|
||||
|
||||
let cases = [
|
||||
("failed", 13 * 24, false, "13 days"),
|
||||
("failed", 15 * 24, true, "15 days"),
|
||||
("done", 23, false, "23 hours"),
|
||||
("done", 25, true, "25 hours"),
|
||||
];
|
||||
for (status, hours, expected, label) in cases {
|
||||
let id = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
status,
|
||||
Some(hours),
|
||||
"originals/e/b.jpg",
|
||||
false,
|
||||
)
|
||||
.await;
|
||||
assert_eq!(
|
||||
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
|
||||
.await
|
||||
.contains(&id),
|
||||
expected,
|
||||
"a {status} upload deleted {label} ago: expected swept={expected}"
|
||||
);
|
||||
sqlx::query("DELETE FROM upload WHERE id = $1")
|
||||
.bind(id)
|
||||
.execute(&pool)
|
||||
.await
|
||||
.expect("clean up");
|
||||
}
|
||||
}
|
||||
|
||||
/// A row is re-selected until EVERY one of its paths is cleared.
|
||||
///
|
||||
/// PREVENTS: two failures at once. The sweep used to clear `original_path` alone, which was right
|
||||
/// for its only case (a failed compression produces no derivatives) but leaves preview, display and
|
||||
/// thumbnail on disk the moment it reaches a successfully processed upload — three files per
|
||||
/// upload, none of them counted in `original_size_bytes`, that nothing else ever removes. And a row
|
||||
/// whose paths are all cleared must stop coming back, or every hourly tick logs a phantom reclaim
|
||||
/// forever.
|
||||
#[sqlx::test]
|
||||
async fn a_row_is_reselected_until_every_path_is_cleared(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-idempotent").await;
|
||||
let user_id = seed_user(&pool, event_id, "Idem").await;
|
||||
let id = seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"done",
|
||||
Some(48),
|
||||
"originals/e/once.jpg",
|
||||
true,
|
||||
)
|
||||
.await;
|
||||
|
||||
assert_eq!(sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await, [id]);
|
||||
|
||||
// Clearing only the original is NOT enough — the derivatives are still on disk.
|
||||
sqlx::query("UPDATE upload SET original_path = '' WHERE id = $1")
|
||||
.bind(id)
|
||||
.execute(&pool)
|
||||
.await
|
||||
.expect("clear path");
|
||||
.expect("clear original");
|
||||
assert_eq!(
|
||||
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await,
|
||||
[id],
|
||||
"derivatives left behind must keep the row selected"
|
||||
);
|
||||
|
||||
sqlx::query(
|
||||
"UPDATE upload SET preview_path = NULL, display_path = NULL, thumbnail_path = NULL
|
||||
WHERE id = $1",
|
||||
)
|
||||
.bind(id)
|
||||
.execute(&pool)
|
||||
.await
|
||||
.expect("clear derivatives");
|
||||
assert!(
|
||||
sweep_selects(&pool, 14).await.is_empty(),
|
||||
"a swept row must not come back"
|
||||
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
|
||||
.await
|
||||
.is_empty(),
|
||||
"a fully swept row must not come back"
|
||||
);
|
||||
}
|
||||
|
||||
/// The derivative backfill must never resurrect what the sweep just reclaimed.
|
||||
///
|
||||
/// PREVENTS: an interaction, not a bug in either piece. The sweep nulls `preview_path`, and
|
||||
/// `backfill_stale_derivatives` selects on `display_path IS NULL AND preview_path IS NOT NULL` —
|
||||
/// close enough that a future edit to either could have the backfill re-decode an original that is
|
||||
/// no longer on disk, on every boot. `deleted_at IS NULL` is what keeps them apart.
|
||||
#[sqlx::test]
|
||||
async fn the_backfill_ignores_swept_rows(pool: PgPool) {
|
||||
let event_id = seed_event(&pool, "sweep-backfill").await;
|
||||
let user_id = seed_user(&pool, event_id, "Backfill").await;
|
||||
seed_aged_upload(
|
||||
&pool,
|
||||
event_id,
|
||||
user_id,
|
||||
"done",
|
||||
Some(48),
|
||||
"originals/e/gone.jpg",
|
||||
true,
|
||||
)
|
||||
.await;
|
||||
|
||||
// SRC: `services/compression.rs::backfill_stale_derivatives` — the selection, verbatim.
|
||||
let backfilled: Vec<(Uuid, String, String)> = sqlx::query_as(
|
||||
"SELECT id, original_path, mime_type FROM upload
|
||||
WHERE deleted_at IS NULL AND mime_type LIKE 'image/%'
|
||||
AND original_path IS NOT NULL
|
||||
AND (
|
||||
(display_path IS NULL AND preview_path IS NOT NULL)
|
||||
OR derivatives_rev < $1
|
||||
)",
|
||||
)
|
||||
.bind(1i16)
|
||||
.fetch_all(&pool)
|
||||
.await
|
||||
.expect("backfill query");
|
||||
|
||||
assert!(
|
||||
backfilled.is_empty(),
|
||||
"a soft-deleted row must be invisible to the backfill, before or after sweeping"
|
||||
);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user