Files
EventSnap/backend/tests/upload_concurrency.rs
fabi cfc8bd0016 test: replace coverage that could not fail with coverage that can
`backend/tests/` follows a house rule of copying production SQL character-for-
character rather than calling `src/`, because the crate is a binary and nothing
in it is importable from an integration test. For pinning behaviour that already
existed that is a defensible trade. Applied to a NEW fix whose only coverage is
the copy, it proves nothing: the fix and its test become two independent
implementations, and deleting the fix leaves the test green.

`audit_names.rs` did exactly that. It never called `audit::record` — it
reimplemented `resolve_names` and the INSERT inside the test file, down to a
hardcoded `.bind("host")`, and then asserted `actor_role == "host"` against its
own literal. That assertion could not fail for any change to the code it named,
and grep confirmed there was no other coverage of the audit-name work anywhere.

Moved into `#[cfg(test)]` inside `services/audit.rs`, where the real function IS
callable. CI already runs `cargo test --all-features` with a live DATABASE_URL,
so `#[sqlx::test]` works there; verified all four run and pass. The role
assertion now compares against `UserRole::as_str()` itself rather than a literal,
so it tracks a rename instead of pretending to, plus an explicit `assert_ne!`
against the Debug spelling.

Also:

- `retry-after-release.spec.ts` filtered the feed on `u.id === original.id` to
  prove "no second row was created". A duplicate gets a fresh uuid and could
  never match, so the filter yielded exactly 1 whether the gallery held one copy
  or five. Counts by uploader now, with the original's identity asserted
  separately. (The rest of that spec is sound — its 403 control and replay-id
  check both fail if the header fast-path is reverted.)

- `upload_after_release_commits_sees_the_lock_and_is_rejected` claimed the
  handler answers `UploadsLocked`. It answers `GalleryReleased` since the check
  order was inverted on this branch, and the test asserts no variant at all.
  Documented what it actually covers (the locked READ) and where the ordering IS
  covered (two e2e specs).

- Two `// SRC:` pointers had drifted ~130 lines into unrelated code, which is how
  a hand-copied fixture silently stops matching its original. Now named, not
  numbered.

- `emptyOutDir: false` claimed a failed viewer build "leaves the last good
  artifact in place". True for the `generateBundle` error, false for the newer
  `writeBundle` assertion, which fires after Vite has already written the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:46:04 +02:00

404 lines
17 KiB
Rust

//! DB-backed integration tests for the two concurrency guards in the upload commit path
//! (`handlers/upload.rs`). Both are SQL — a `FOR SHARE` row lock and an atomic compare-and-increment
//! — and both are load-bearing for things a user can actually lose: a wedding photo, or the disk.
//!
//! `#[sqlx::test]` gives each test a fresh database with the real migrations applied.
mod common;
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use std::time::Duration;
use chrono::{DateTime, Utc};
use common::*;
use sqlx::PgPool;
use uuid::Uuid;
/// SRC: `handlers/upload.rs::create_upload` — the guarded quota increment, verbatim.
/// (Named, not line-numbered: the previous pointer drifted by ~130 lines and landed in unrelated
/// code, which is how a hand-copied fixture silently stops matching its original.)
/// Returns `rows_affected()`; the handler aborts the whole upload tx when this is 0.
async fn quota_inc(exec: impl sqlx::PgExecutor<'_>, user_id: Uuid, size: i64, limit: i64) -> u64 {
sqlx::query(
"UPDATE \"user\" SET total_upload_bytes = total_upload_bytes + $2
WHERE id = $1 AND total_upload_bytes + $2 <= $3",
)
.bind(user_id)
.bind(size)
.bind(limit)
.execute(exec)
.await
.expect("quota_inc")
.rows_affected()
}
async fn total_bytes(pool: &PgPool, user_id: Uuid) -> i64 {
sqlx::query_scalar("SELECT total_upload_bytes FROM \"user\" WHERE id = $1")
.bind(user_id)
.fetch_one(pool)
.await
.expect("total_bytes")
}
// ─────────────────────────────────────────────────────────────────────────────
// 5. The atomic quota increment
// ─────────────────────────────────────────────────────────────────────────────
/// Two attempts sized off ONE stale snapshot, each of which "fits" on its own, cannot both commit.
/// The predicate re-reads `total_upload_bytes` inside the UPDATE, so the second matches 0 rows.
///
/// PREVENTS: one guest filling the disk. The handler's pre-flight quota check runs BEFORE the body is
/// streamed — minutes earlier, for a 500 MB video. If the commit trusted that snapshot, a guest could
/// start N uploads that each individually fit under the limit and land all N, blowing straight through
/// the quota and (in a 1 GB container) taking the event down for everyone.
#[sqlx::test]
async fn quota_two_attempts_from_one_stale_snapshot_cannot_both_commit(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user_id = seed_user(&pool, event_id, "Gierige Gudrun").await;
const LIMIT: i64 = 100;
const SIZE: i64 = 60;
// THE STALE SNAPSHOT: the pre-flight check both uploads were admitted on.
let snapshot = total_bytes(&pool, user_id).await;
assert_eq!(snapshot, 0);
// Each upload, judged against that snapshot alone, fits: 0 + 60 <= 100. Twice.
assert!(snapshot + SIZE <= LIMIT);
assert_eq!(
quota_inc(&pool, user_id, SIZE, LIMIT).await,
1,
"the first upload commits"
);
assert_eq!(
quota_inc(&pool, user_id, SIZE, LIMIT).await,
0,
"the second MUST affect 0 rows — it was admitted on a snapshot that is now a lie \
(60 + 60 = 120 > 100). rows_affected() == 0 is what makes the handler abort."
);
assert_eq!(
total_bytes(&pool, user_id).await,
SIZE,
"never 120 — the quota held"
);
}
/// The same, but genuinely CONCURRENT: two transactions that both read `total = 0`, then both try to
/// commit 60 bytes against a 100-byte limit. The second UPDATE blocks on the first's row lock and —
/// because every predicate is on the ROW BEING UPDATED — Postgres re-evaluates it against the
/// post-commit row (EPQ) rather than the statement's original snapshot. It matches nothing.
///
/// PREVENTS: exactly the same disk-filling overrun, on the path it actually happens — two uploads
/// in flight at once, which is the normal case at a party.
#[sqlx::test]
async fn quota_guard_is_atomic_under_concurrent_transactions(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user_id = seed_user(&pool, event_id, "Gierige Gudrun").await;
const LIMIT: i64 = 100;
const SIZE: i64 = 60;
let mut tx1 = pool.begin().await.unwrap();
let mut tx2 = pool.begin().await.unwrap();
// Both transactions read the same snapshot and both would pass a naive `total + size <= limit`
// check done in Rust.
for tx in [&mut tx1, &mut tx2] {
let seen: i64 = sqlx::query_scalar("SELECT total_upload_bytes FROM \"user\" WHERE id = $1")
.bind(user_id)
.fetch_one(&mut **tx)
.await
.unwrap();
assert_eq!(seen, 0, "both see an empty quota");
}
// tx1 takes the row lock and commits.
assert_eq!(quota_inc(&mut *tx1, user_id, SIZE, LIMIT).await, 1);
tx1.commit().await.unwrap();
// tx2's UPDATE was written against the stale snapshot but is evaluated against the row as it
// now stands.
assert_eq!(
quota_inc(&mut *tx2, user_id, SIZE, LIMIT).await,
0,
"the loser MUST see 0 rows affected — this is the entire quota guarantee"
);
tx2.rollback().await.unwrap();
assert_eq!(total_bytes(&pool, user_id).await, SIZE);
// And an upload that legitimately fits in what's left still succeeds — the guard rejects
// overruns, not everything.
assert_eq!(
quota_inc(&pool, user_id, 40, LIMIT).await,
1,
"0 + 60 + 40 == 100, exactly at the limit"
);
assert_eq!(total_bytes(&pool, user_id).await, LIMIT);
assert_eq!(
quota_inc(&pool, user_id, 1, LIMIT).await,
0,
"and one byte more is refused"
);
}
// ─────────────────────────────────────────────────────────────────────────────
// 6. The `FOR SHARE` upload lock vs. the release
// ─────────────────────────────────────────────────────────────────────────────
/// SRC: `handlers/upload.rs::create_upload` — the in-transaction `FOR SHARE` re-check, verbatim.
/// (Named, not line-numbered — see the note on `quota_inc`.)
async fn lock_and_read_event(
tx: &mut sqlx::PgConnection,
event_id: Uuid,
) -> (Option<DateTime<Utc>>, Option<DateTime<Utc>>) {
sqlx::query_as(
"SELECT uploads_locked_at, export_released_at FROM event WHERE id = $1 FOR SHARE",
)
.bind(event_id)
.fetch_one(tx)
.await
.expect("FOR SHARE re-check")
}
/// THE GUARD AGAINST SILENT, PERMANENT PHOTO LOSS.
///
/// An upload holding `FOR SHARE` on the event row must BLOCK the `UPDATE event SET
/// export_released_at = NOW()` in `release_gallery` until it commits. Either the upload commits first
/// — and the release (hence the export snapshot) is strictly ordered after it, so the keepsake
/// CONTAINS the photo — or the release commits first and the upload observes the lock and rejects
/// (reversibly: the client keeps the blob and resumes after a reopen).
///
/// PREVENTS: the lost wedding photo. Without this serialization: a guest starts a 500 MB video, the
/// pre-flight lock check passes, the host releases the gallery, the export workers snapshot the
/// uploads table, and THEN the upload commits. The photo appears in the live feed but is missing from
/// the downloaded keepsake, forever — nothing ever regenerates it and nobody ever notices.
#[sqlx::test]
async fn for_share_upload_lock_serializes_against_release(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user_id = seed_user(&pool, event_id, "Fotograf Fritz").await;
// ── The guest's upload transaction takes the share lock. ──
let mut upload_tx = pool.begin().await.unwrap();
let (locked, released) = lock_and_read_event(&mut upload_tx, event_id).await;
assert!(
locked.is_none() && released.is_none(),
"uploads are open, so we proceed to commit"
);
// ── Concurrently, the host hits "Galerie freigeben". ──
let release_done = Arc::new(AtomicBool::new(false));
let release_task = {
let pool = pool.clone();
let release_done = release_done.clone();
tokio::spawn(async move {
sqlx::query(
"UPDATE event
SET export_released_at = NOW(),
uploads_locked_at = COALESCE(uploads_locked_at, NOW()),
export_epoch = export_epoch + 1
WHERE id = $1 AND export_released_at IS NULL",
)
.bind(event_id)
.execute(&pool)
.await
.expect("release");
release_done.store(true, Ordering::SeqCst);
})
};
// The release MUST be stuck behind our `FOR SHARE` row lock. (`FOR SHARE` conflicts with the
// `FOR UPDATE` lock the UPDATE needs, so Postgres makes it wait — this is not a timing race,
// it is a lock-conflict guarantee; the sleep only gives it every chance to wrongly proceed.)
tokio::time::sleep(Duration::from_millis(750)).await;
assert!(
!release_done.load(Ordering::SeqCst),
"the release MUST block while an upload holds FOR SHARE — if it can slip past, the export \
snapshot is taken while a photo is still committing and that photo is lost forever"
);
// The photo commits. It is now unambiguously part of the upload set.
let upload_id: Uuid = sqlx::query_scalar(
"INSERT INTO upload (event_id, user_id, original_path, mime_type, original_size_bytes)
VALUES ($1, $2, 'originals/wedding/x.jpg', 'image/jpeg', 1234) RETURNING id",
)
.bind(event_id)
.bind(user_id)
.fetch_one(&mut *upload_tx)
.await
.unwrap();
upload_tx.commit().await.unwrap();
// Only now can the release proceed.
tokio::time::timeout(Duration::from_secs(5), release_task)
.await
.expect("the release must unblock once the upload commits")
.unwrap();
// THE PAYOFF: the export snapshot — the very query the ZIP worker runs — sees the photo. Order
// enforced by the lock: upload commit < release < snapshot.
let snapshot: Vec<Uuid> = sqlx::query_scalar(
"SELECT u.id FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE",
)
.bind(event_id)
.fetch_all(&pool)
.await
.unwrap();
assert_eq!(
snapshot,
vec![upload_id],
"the released keepsake CONTAINS the in-flight photo"
);
}
/// The other side of the same lock: once the release has COMMITTED, the next upload's `FOR SHARE`
/// re-read sees `export_released_at` set and the handler bails out.
///
/// PREVENTS: the same lost photo, on the losing side of the race — a photo committing AFTER the
/// export snapshot would be in the live feed but missing from the keepsake. Rejecting is the correct
/// outcome, and it is reversible: the client keeps the blob and resumes when the host reopens.
///
/// SCOPE, because the name overstates it: this asserts only what the LOCKED READ observes. It does
/// not go through the handler, so it says nothing about which error the handler picks. That
/// distinction is load-bearing — `create_upload` answers a released gallery with `GalleryReleased`
/// and a plain lock with `UploadsLocked`, in that order, and the two drive different client
/// behaviour (a `reopen` park vs. a retry). The ordering is covered end-to-end by
/// `e2e/specs/10-flow-review/upload-lock-code.spec.ts` and `02-upload/retry-after-release.spec.ts`;
/// this test's doc used to claim `UploadsLocked` outright and was simply wrong after that split.
#[sqlx::test]
async fn upload_after_release_commits_sees_the_lock_and_is_rejected(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
// Before the release, the re-check passes.
let mut tx = pool.begin().await.unwrap();
let (locked, released) = lock_and_read_event(&mut tx, event_id).await;
assert!(locked.is_none() && released.is_none());
tx.rollback().await.unwrap();
assert_eq!(release_gallery(&pool, "wedding").await, Some(1));
// After it, the identical re-check sees the release and the handler bails out.
let mut tx = pool.begin().await.unwrap();
let (locked, released) = lock_and_read_event(&mut tx, event_id).await;
assert!(
released.is_some(),
"the FOR SHARE re-read MUST observe the committed release"
);
assert!(
locked.is_some(),
"release locks uploads in the same statement (release ⇒ lock)"
);
tx.rollback().await.unwrap();
// And a reopen makes it uploadable again — the rejection was reversible, not terminal.
assert_eq!(open_event(&pool, "wedding").await, 1);
let mut tx = pool.begin().await.unwrap();
let (locked, released) = lock_and_read_event(&mut tx, event_id).await;
assert!(
locked.is_none() && released.is_none(),
"the guest can resume their upload"
);
tx.rollback().await.unwrap();
}
/// SRC: `handlers/me.rs::delete_account` — the last-operator guard, verbatim.
///
/// Returns the ids of the OTHER live operators, holding a row lock on each. The handler refuses the
/// deletion when this is empty.
async fn other_operators(tx: &mut sqlx::PgConnection, event_id: Uuid, self_id: Uuid) -> Vec<Uuid> {
sqlx::query("SELECT pg_advisory_xact_lock(4242, hashtext($1::text))")
.bind(event_id)
.execute(&mut *tx)
.await
.expect("advisory lock");
sqlx::query_scalar(
"SELECT id FROM \"user\"
WHERE event_id = $1 AND id != $2
AND role IN ('host', 'admin') AND is_banned = FALSE",
)
.bind(event_id)
.bind(self_id)
.fetch_all(tx)
.await
.expect("other_operators")
}
/// Two hosts deleting themselves at the same moment must not both succeed.
///
/// The guard used to run on the pool BEFORE the transaction opened, so each deleter saw the other,
/// both passed, and the event was left with no operator at all — nobody to moderate, nobody to
/// release the gallery, and no way to appoint anyone, because appointing requires a host. Not
/// recoverable from inside the app.
///
/// A transaction-scoped ADVISORY lock serialises them. A row lock on the other operators would
/// deadlock instead — each deleter locks the other's row and then tries to delete its own, so
/// Postgres kills one with a deadlock error; the invariant survives but the loser gets a 500. A
/// lock on the `event` row would serialise cleanly but inverts the order every moderation path
/// takes (upload/user rows first, event last). The advisory lock is a separate space, so it cannot
/// interact with the row-lock graph at all: the loser waits, then counts zero once the winner's row
/// is gone, and is refused with a sentence instead of an error.
#[sqlx::test]
async fn two_hosts_deleting_at_once_cannot_both_leave_the_event(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let a = seed_user(&pool, event_id, "Gastgeber Anton").await;
let b = seed_user(&pool, event_id, "Gastgeberin Berta").await;
for id in [a, b] {
sqlx::query("UPDATE \"user\" SET role = 'host' WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("promote");
}
// A opens first and takes the lock on B's row.
let mut tx_a = pool.begin().await.expect("tx a");
let a_sees = other_operators(&mut tx_a, event_id, a).await;
assert_eq!(a_sees, vec![b], "A must see B as the remaining operator");
// B now tries the same and blocks on A's row. Spawned, because it cannot return until A
// commits — which is precisely the serialisation under test.
let pool_b = pool.clone();
let b_task = tokio::spawn(async move {
let mut tx_b = pool_b.begin().await.expect("tx b");
let seen = other_operators(&mut tx_b, event_id, b).await;
tx_b.commit().await.expect("commit b");
seen
});
// Give B a moment to actually reach the lock rather than racing past it.
tokio::time::sleep(Duration::from_millis(300)).await;
// A completes its deletion.
sqlx::query("DELETE FROM \"user\" WHERE id = $1")
.bind(a)
.execute(&mut *tx_a)
.await
.expect("delete a");
tx_a.commit().await.expect("commit a");
let b_sees = b_task.await.expect("b task");
assert!(
b_sees.is_empty(),
"B unblocked and must now see NO remaining operator (A is gone), so its deletion is \
refused — it saw {b_sees:?}"
);
// The event still has exactly one operator: B.
let remaining: i64 = sqlx::query_scalar(
"SELECT COUNT(*) FROM \"user\" WHERE event_id = $1 AND role IN ('host','admin')",
)
.bind(event_id)
.fetch_one(&pool)
.await
.expect("count");
assert_eq!(
remaining, 1,
"the event must never be left without an operator"
);
}