fix(backend): three ways the end of the night could go wrong

**1. Releasing the gallery could arm the keepsake with no worker.**
`release_gallery` ran `tx.commit()` -> SSE `event-closed` -> `audit::record().await`
-> `spawn_export_jobs`. The audit write is two pool round-trips, each able to wait
the full 5s acquire timeout, and it runs in the same instant `event-closed` fans
out to ~100 phones whose upload queues all hit the API at once. Axum drops the
handler future when the client disconnects — the host taps "Freigeben" and
pockets the phone. The release has COMMITTED: event closed, uploads locked, epoch
bumped, both `export_job` rows pending, and no worker. `/export/*` 404s, the page
sits on "Wird vorbereitet…", `recover_exports` only runs at boot, and a second
release is refused. Every other regen call site spawns first; `me.rs` says so in
a comment. This was the sole violator, and the only path that arms the FIRST
build of the keepsake. Spawn moved immediately after the commit.

**2. The event could be left with no operator.**
`remaining_operators` was an unlocked pool COUNT followed by a separate UPDATE,
so `ban_user` and `set_role` raced each other and `DELETE /me`: an admin demotes
host B while host A deletes themselves, each check sees the other still present,
both commit, and nobody can moderate, release the gallery, or appoint anyone —
appointing requires being an operator. The count now runs inside the writing
transaction behind the same advisory lock `delete_account` uses, via one shared
helper so the key cannot drift between copies.

The lock is taken FIRST in all three, and the order is load-bearing:
`delete_account` previously took it last, after row locks on `upload` and
`event`, while the two new call sites take it before locking those same rows —
an ABBA that Postgres would resolve by killing one transaction with a 500. The
ordering rule is documented on the helper.

**3. The keepsake could become unbuildable the moment uploads stopped.**
The upload gate and the export preflight computed the IDENTICAL threshold
(`required_free_bytes(media, 2) + DISK_RESERVE_BYTES`), leaving zero margin
between them. Once the gate refused its first upload the preflight was already at
its own limit, so anything written afterwards decided the keepsake's fate: WAL up
to `max_wal_size`, 30 MB x 4 of container logs, and the compression backlog
draining at exactly that hour. The release commits before the workers bail, so
the failure lands at 01:00 with no second release possible. The gate now demands
`UPLOAD_GATE_HEADROOM_BYTES` more than the preflight, costing ~0.5 GB of media
ceiling — the trade README already argues for. The dashboard banner mirrors the
new threshold so its lead is unchanged, and a new test pins gate-before-preflight
at six gallery sizes.

Also: the global disk gate fails OPEN when the mount cannot be read, which is
deliberate, but did it SILENTLY — no log line at all, while the export preflight
warns on the identical condition. Inside a container `/` is an overlay rather
than a `/dev` device, so this is reachable, and when it happens the only global
disk bound is gone and the box fills until Postgres cannot write WAL.

README's sizing table was also arithmetically self-contradictory (it showed
~27 GB free against a ~27.6 GB requirement); recomputed for the new gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
fabi
2026-08-13 19:45:24 +02:00
parent 06e0bea0e9
commit 20c15c3500
4 changed files with 265 additions and 94 deletions

View File

@@ -311,19 +311,29 @@ to imply is gone. What bounds the disk is the **global gate in the upload handle
which refuses any upload that would leave too little room to build the keepsake: which refuses any upload that would leave too little room to build the keepsake:
``` ```
free_after_upload < media_after × 1.1 × 2 + DISK_RESERVE_BYTES → refused free_after_upload < media_after × 1.1 × 2 + DISK_RESERVE_BYTES
+ UPLOAD_GATE_HEADROOM_BYTES → refused
``` ```
Solving that for the gallery size gives the real ceiling. On the **40 GB box this runs That last term is what separates this gate from the export preflight, which bails at
on**, with ~5 GB for the OS, Docker images (the runbook pre-pulls the rollback tag too) `media × 1.1 × 2 + DISK_RESERVE_BYTES` — the same expression **minus** the headroom. The
and Postgres: two used to be identical, which meant the preflight was already sitting on its limit at
the exact moment uploads stopped: every byte written between the last refused upload and
the host tapping *Galerie freigeben* (Postgres WAL, container logs, the compression
backlog draining at precisely that hour) pushed it under, and the release commits before
the workers fail. The headroom buys 1.5 GB of slack so that cannot happen.
| Volume | Usable after baseline | Media ceiling | Free at release | Solving the gate for the gallery size gives the real ceiling — the gate's equilibrium is
|---|---|---|---| `3.2 × media`, so each GB of reserve or headroom costs ~0.31 GB of gallery. On the
| 40 GB | ~35 GB | **~8 GB** | ~27 GB → both archives fit | **40 GB box this runs on**, with ~5 GB for the OS, Docker images (the runbook pre-pulls
| 80 GB | ~70 GB | ~19 GB | ~51 GB → both archives fit | the rollback tag too) and Postgres:
**Uploads therefore stop at roughly 8 GB of media on a 40 GB box, not when the disk is | Volume | Usable after baseline | Media ceiling | Free at release | Preflight needs |
|---|---|---|---|---|
| 40 GB | ~35 GB | **~7.3 GB** | ~27.7 GB | ~26.2 GB → fits, 1.5 GB spare |
| 80 GB | ~70 GB | ~18.3 GB | ~51.7 GB | ~50.2 GB → fits, 1.5 GB spare |
**Uploads therefore stop at roughly 7 GB of media on a 40 GB box, not when the disk is
full.** That is deliberate. 1000 photos at ~3.5 MB is ~3.5 GB and fits comfortably; full.** That is deliberate. 1000 photos at ~3.5 MB is ~3.5 GB and fits comfortably;
video is what consumes the budget, so lower `max_video_size_mb` (seeded at 500) if you video is what consumes the budget, so lower `max_video_size_mb` (seeded at 500) if you
expect a lot of it. Refusing the 1001st upload is a far better outcome than accepting it expect a lot of it. Refusing the 1001st upload is a far better outcome than accepting it

View File

@@ -57,7 +57,8 @@ pub struct EventStatus {
/// into a decision someone can still make. /// into a decision someone can still make.
/// ///
/// IT MUST FIRE BEFORE THE UPLOAD GATE CLOSES, and that is why the reserve and the margin are /// IT MUST FIRE BEFORE THE UPLOAD GATE CLOSES, and that is why the reserve and the margin are
/// here. The gate in `handlers::upload` refuses at `free < keepsake_required + DISK_RESERVE_BYTES`; /// here. The gate in `handlers::upload` refuses at
/// `free < keepsake_required + DISK_RESERVE_BYTES + UPLOAD_GATE_HEADROOM_BYTES`;
/// warning at `free < keepsake_required` alone meant the two differed by the whole reserve, so /// warning at `free < keepsake_required` alone meant the two differed by the whole reserve, so
/// the wall was always hit FIRST. Every guest would be blocked from uploading while this /// the wall was always hit FIRST. Every guest would be blocked from uploading while this
/// dashboard showed a comfortable disk and no banner at all — on the shipped 40 GB box, uploads /// dashboard showed a comfortable disk and no banner at all — on the shipped 40 GB box, uploads
@@ -66,8 +67,14 @@ pub struct EventStatus {
/// The 25% margin makes it a warning rather than an obituary: the host sees it while there is /// The 25% margin makes it a warning rather than an obituary: the host sees it while there is
/// still room to act (delete a few large videos, which refunds immediately and reopens the gate). /// still room to act (delete a few large videos, which refunds immediately and reopens the gate).
fn disk_is_low(free: u64, keepsake_required: u64) -> bool { fn disk_is_low(free: u64, keepsake_required: u64) -> bool {
let gate_closes_at = // Mirrors the gate EXACTLY, headroom included. The gate now demands
keepsake_required.saturating_add(crate::handlers::upload::DISK_RESERVE_BYTES as u64); // `UPLOAD_GATE_HEADROOM_BYTES` more than the export preflight does, so that ordinary
// end-of-night writes cannot flip the preflight after uploads have already stopped. Leaving
// that term out here would shrink the warning's lead by 1.5 GB — and the whole point of this
// function is that the banner must appear while the host can still act.
let gate_closes_at = keepsake_required
.saturating_add(crate::handlers::upload::DISK_RESERVE_BYTES as u64)
.saturating_add(crate::handlers::upload::UPLOAD_GATE_HEADROOM_BYTES as u64);
let warn_at = gate_closes_at.saturating_add(gate_closes_at / 4); let warn_at = gate_closes_at.saturating_add(gate_closes_at / 4);
// No separate absolute-floor clause. There used to be `free < LOW_DISK_FLOOR_BYTES ||` here, // No separate absolute-floor clause. There used to be `free < LOW_DISK_FLOOR_BYTES ||` here,
// and it was unreachable: `gate_closes_at` is at least DISK_RESERVE_BYTES, so `warn_at` is at // and it was unreachable: `gate_closes_at` is at least DISK_RESERVE_BYTES, so `warn_at` is at
@@ -80,8 +87,15 @@ fn disk_is_low(free: u64, keepsake_required: u64) -> bool {
/// Count non-banned hosts/admins in the event OTHER than `excluding` — the operators /// Count non-banned hosts/admins in the event OTHER than `excluding` — the operators
/// who would remain if `excluding` were demoted or banned. Used to enforce the "an event /// who would remain if `excluding` were demoted or banned. Used to enforce the "an event
/// always keeps at least one operator" floor. /// always keeps at least one operator" floor.
///
/// Takes a CONNECTION, not the pool, and every caller passes the same transaction it is about to
/// write in — after taking [`lock_operator_floor`]. Read on the pool beforehand, this count was a
/// snapshot that any concurrent operator-removing action could invalidate before the UPDATE landed:
/// an admin demoting host B while host A calls `DELETE /me` saw two independent checks each observe
/// the other still present, both commit, and the event end up with zero operators — which is not
/// recoverable from inside the app, since appointing an operator requires being one.
async fn remaining_operators( async fn remaining_operators(
state: &AppState, conn: &mut sqlx::PgConnection,
event_id: Uuid, event_id: Uuid,
excluding: Uuid, excluding: Uuid,
) -> Result<i64, AppError> { ) -> Result<i64, AppError> {
@@ -92,11 +106,36 @@ async fn remaining_operators(
) )
.bind(event_id) .bind(event_id)
.bind(excluding) .bind(excluding)
.fetch_one(&state.pool) .fetch_one(conn)
.await?; .await?;
Ok(count) Ok(count)
} }
/// Serialise every action that can remove an operator from an event.
///
/// The same key `me::delete_account` takes — namespace 4242, `hashtext(event_id)` — and it MUST
/// stay identical, or the two families of caller lock against nothing. An advisory lock is used
/// rather than a row lock because it is a separate lock space and so cannot join the
/// `event`/`user` row-lock graph that moderation traffic already traverses in both directions;
/// it is released automatically when the transaction ends.
///
/// **Call this FIRST in the transaction, before taking any row lock.** Being a separate lock space
/// means it cannot form a cycle *with itself*, not that ordering is free: all three callers go on
/// to lock `user` and `event` rows, so a caller that took those rows first and reached for this
/// lock afterwards would deadlock against one that did it the other way round. Postgres would
/// break the tie by killing one transaction with a 500. Every caller acquires it first; keep it
/// that way.
pub(crate) async fn lock_operator_floor(
conn: &mut sqlx::PgConnection,
event_id: Uuid,
) -> Result<(), AppError> {
sqlx::query("SELECT pg_advisory_xact_lock(4242, hashtext($1::text))")
.bind(event_id)
.execute(conn)
.await?;
Ok(())
}
#[derive(Deserialize)] #[derive(Deserialize)]
pub struct SetRoleRequest { pub struct SetRoleRequest {
pub role: String, pub role: String,
@@ -193,14 +232,6 @@ pub async fn ban_user(
)); ));
} }
// Floor: never leave the event with zero operators. Banning removes the target from
// the active-operator pool, so refuse if they're the last non-banned host/admin.
if target.0 == "host" && remaining_operators(&state, auth.event_id, user_id).await? == 0 {
return Err(AppError::BadRequest(
"Der letzte Host kann nicht gesperrt werden.".into(),
));
}
// Ban ALWAYS hides: a banned user's content is "gone" everywhere. The visibility // Ban ALWAYS hides: a banned user's content is "gone" everywhere. The visibility
// views/queries now also filter on `is_banned` (defense in depth), and we set // views/queries now also filter on `is_banned` (defense in depth), and we set
// `uploads_hidden` so the existing `user-hidden` live-eviction path fires too. The old // `uploads_hidden` so the existing `user-hidden` live-eviction path fires too. The old
@@ -215,6 +246,21 @@ pub async fn ban_user(
// //
// The ban and the keepsake invalidation are ONE transaction — see `host_delete_upload`. // The ban and the keepsake invalidation are ONE transaction — see `host_delete_upload`.
let mut tx = state.pool.begin().await?; let mut tx = state.pool.begin().await?;
// Floor: never leave the event with zero operators. Banning removes the target from the
// active-operator pool, so refuse if they're the last non-banned host/admin.
//
// INSIDE the transaction and behind the operator lock — see `remaining_operators`. Checked on
// the pool beforehand, this raced `set_role` and `DELETE /me` into an event with no operator.
if target.0 == "host" {
lock_operator_floor(&mut tx, auth.event_id).await?;
if remaining_operators(&mut tx, auth.event_id, user_id).await? == 0 {
return Err(AppError::BadRequest(
"Der letzte Host kann nicht gesperrt werden.".into(),
));
}
}
sqlx::query( sqlx::query(
"UPDATE \"user\" "UPDATE \"user\"
SET is_banned = TRUE, uploads_hidden = TRUE, uploads_hidden_at = NOW() SET is_banned = TRUE, uploads_hidden = TRUE, uploads_hidden_at = NOW()
@@ -461,21 +507,26 @@ pub async fn set_role(
// Floor: demoting the last non-banned host/admin to guest would leave the event with // Floor: demoting the last non-banned host/admin to guest would leave the event with
// no operator. Refuse. // no operator. Refuse.
if new_role == "guest" //
&& target.0 == "host" // The check and the UPDATE are ONE transaction, behind the operator lock — see
&& remaining_operators(&state, auth.event_id, user_id).await? == 0 // `remaining_operators`. Split apart on the pool, this raced `ban_user` and `DELETE /me`.
{ let mut tx = state.pool.begin().await?;
return Err(AppError::BadRequest( if new_role == "guest" && target.0 == "host" {
"Der letzte Host kann nicht zum Gast gemacht werden.".into(), lock_operator_floor(&mut tx, auth.event_id).await?;
)); if remaining_operators(&mut tx, auth.event_id, user_id).await? == 0 {
return Err(AppError::BadRequest(
"Der letzte Host kann nicht zum Gast gemacht werden.".into(),
));
}
} }
sqlx::query("UPDATE \"user\" SET role = $2::user_role WHERE id = $1 AND event_id = $3") sqlx::query("UPDATE \"user\" SET role = $2::user_role WHERE id = $1 AND event_id = $3")
.bind(user_id) .bind(user_id)
.bind(new_role) .bind(new_role)
.bind(auth.event_id) .bind(auth.event_id)
.execute(&state.pool) .execute(&mut *tx)
.await?; .await?;
tx.commit().await?;
tracing::info!( tracing::info!(
actor_user_id = %auth.user_id, actor_user_id = %auth.user_id,
target_user_id = %user_id, target_user_id = %user_id,
@@ -938,9 +989,39 @@ pub async fn release_gallery(
// discovering it via a rejected upload. // discovering it via a rejected upload.
let _ = state.sse_tx.send(SseEvent::new("event-closed", "{}")); let _ = state.sse_tx.send(SseEvent::new("event-closed", "{}"));
// Detached — survives this handler being cancelled.
//
// SPAWNED IMMEDIATELY AFTER THE COMMIT, BEFORE ANY OTHER `.await`. Every `invalidate_and_arm`
// call site does this; `me::delete_account` carries the same note. The audit write below used
// to sit here, and it is two pool round-trips that can each wait up to the 5 s acquire timeout
// — right at the moment `event-closed` has just fanned out to ~100 phones whose queues all hit
// the API at once, so the pool is as contended as it ever gets. Drop the handler future during
// that suspension (the host's phone sleeps, the tab closes, Caddy times the request out) and
// the task never spawns: the event is released, uploads are locked, both `export_job` rows sit
// `pending` at the live epoch, and no worker exists. `/export/*` 404s, the page sits on "Wird
// vorbereitet…", `recover_exports` only runs at boot, and `release_gallery` refuses a retry
// because the gallery is already released.
//
// This is the one path that arms the FIRST build of the keepsake, so it is the worst possible
// place to reintroduce that window.
crate::services::export::spawn_export_jobs(
event_id,
event_name,
epoch,
state.config.comments_enabled,
std::time::Duration::ZERO,
state.pool.clone(),
state.config.media_path.clone(),
state.config.export_path.clone(),
state.sse_tx.clone(),
);
// Was logged NOWHERE at all before this — not even a tracing line. A host reading // Was logged NOWHERE at all before this — not even a tracing line. A host reading
// the record the morning after had no way to see when uploads were locked or the // the record the morning after had no way to see when uploads were locked or the
// gallery released, which are the two actions that change what every guest can do. // gallery released, which are the two actions that change what every guest can do.
//
// Last, deliberately: it is best-effort by design (it swallows its own errors), so nothing
// downstream may depend on it having completed.
crate::services::audit::record( crate::services::audit::record(
&state.pool, &state.pool,
auth.event_id, auth.event_id,
@@ -954,26 +1035,13 @@ pub async fn release_gallery(
) )
.await; .await;
// Detached — survives this handler being cancelled.
crate::services::export::spawn_export_jobs(
event_id,
event_name,
epoch,
state.config.comments_enabled,
std::time::Duration::ZERO,
state.pool.clone(),
state.config.media_path.clone(),
state.config.export_path.clone(),
state.sse_tx.clone(),
);
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::disk_is_low; use super::disk_is_low;
use crate::handlers::upload::DISK_RESERVE_BYTES; use crate::handlers::upload::{DISK_RESERVE_BYTES, UPLOAD_GATE_HEADROOM_BYTES};
use crate::services::export::required_free_bytes; use crate::services::export::required_free_bytes;
const GB: u64 = 1_000_000_000; const GB: u64 = 1_000_000_000;
@@ -1016,7 +1084,8 @@ mod tests {
// require it to be strictly above the level at which the gate closes, by a usable amount. // require it to be strictly above the level at which the gate closes, by a usable amount.
for media_gb in [0u64, 1, 4, 8, 16, 32] { for media_gb in [0u64, 1, 4, 8, 16, 32] {
let required = required_free_bytes(media_gb * GB, 2); let required = required_free_bytes(media_gb * GB, 2);
let gate_closes_at = required + DISK_RESERVE_BYTES as u64; let gate_closes_at =
required + DISK_RESERVE_BYTES as u64 + UPLOAD_GATE_HEADROOM_BYTES as u64;
// Just above the gate: guests can still upload, and the host must already be warned. // Just above the gate: guests can still upload, and the host must already be warned.
assert!( assert!(
@@ -1037,6 +1106,43 @@ mod tests {
} }
} }
/// The invariant the headroom exists for: uploads must stop while the keepsake can STILL be
/// built, with room to spare — not at the exact instant the preflight reaches its own limit.
///
/// Both thresholds used to be `required_free_bytes(media, 2) + DISK_RESERVE_BYTES`, identically.
/// So the moment the gate refused its first upload, the export preflight was already sitting on
/// its limit, and every byte written afterwards (WAL, container logs, the compression backlog
/// draining at exactly that hour) pushed it under. The release would then COMMIT — event closed,
/// uploads locked, epoch bumped, `event-closed` fanned out to every phone — and only then would
/// both workers bail, with no second release possible.
#[test]
fn the_upload_gate_closes_before_the_export_preflight_would_refuse() {
for media_gb in [0u64, 1, 4, 8, 16, 32] {
let required = required_free_bytes(media_gb * GB, 2);
// `services::export::preflight` bails below this.
let preflight_refuses_below = required + DISK_RESERVE_BYTES as u64;
// `handlers::upload` refuses below this.
let gate_refuses_below = preflight_refuses_below + UPLOAD_GATE_HEADROOM_BYTES as u64;
assert!(
gate_refuses_below > preflight_refuses_below,
"at media={media_gb}GB the gate and the preflight share a threshold, so the \
keepsake's fate rests on whatever is written after uploads stop"
);
// At the instant the last upload is refused, the preflight must still pass with the
// whole headroom to spare — that is the slack the night's remaining writes consume.
let free_when_gate_closes = gate_refuses_below;
assert!(
free_when_gate_closes
>= preflight_refuses_below + UPLOAD_GATE_HEADROOM_BYTES as u64,
"at media={media_gb}GB there is no slack between the gate closing and the \
preflight failing"
);
}
}
#[test] #[test]
fn plenty_of_space_is_still_low_when_the_keepsake_would_not_fit() { fn plenty_of_space_is_still_low_when_the_keepsake_would_not_fit() {
// THE case the fixed threshold misses, and the one that matters: 30 GB free is nowhere near // THE case the fixed threshold misses, and the one that matters: 30 GB free is nowhere near
@@ -1051,7 +1157,7 @@ mod tests {
// size — see `disk_is_low`. Warning at the bare size fired only after the gate had // size — see `disk_is_low`. Warning at the bare size fired only after the gate had
// already blocked every guest. // already blocked every guest.
let required = 20 * GB; let required = 20 * GB;
let gate = required + DISK_RESERVE_BYTES as u64; let gate = required + DISK_RESERVE_BYTES as u64 + UPLOAD_GATE_HEADROOM_BYTES as u64;
let warn_at = gate + gate / 4; let warn_at = gate + gate / 4;
assert!(!disk_is_low(warn_at, required), "exactly enough is enough"); assert!(!disk_is_low(warn_at, required), "exactly enough is enough");
assert!(disk_is_low(warn_at - 1, required)); assert!(disk_is_low(warn_at - 1, required));
@@ -1060,8 +1166,9 @@ mod tests {
#[test] #[test]
fn an_empty_gallery_still_reserves_room_for_postgres() { fn an_empty_gallery_still_reserves_room_for_postgres() {
// With no gallery the keepsake term is 0, so the warn threshold collapses to // With no gallery the keepsake term is 0, so the warn threshold collapses to
// 1.25 x DISK_RESERVE_BYTES (12.5 GB), which dominates the 10 GB absolute floor. // 1.25 x (DISK_RESERVE_BYTES + UPLOAD_GATE_HEADROOM_BYTES) = 1.25 x 11.5 GB = 14.375 GB,
assert!(!disk_is_low(13 * GB, 0)); // which dominates the 10 GB absolute floor.
assert!(!disk_is_low(15 * GB, 0));
assert!(disk_is_low(9 * GB, 0)); assert!(disk_is_low(9 * GB, 0));
} }
} }

View File

@@ -178,6 +178,54 @@ pub async fn delete_account(
.await?; .await?;
let mut tx = state.pool.begin().await?; let mut tx = state.pool.begin().await?;
// The last-host guard, AUTHORITATIVELY — inside the transaction, holding a lock.
//
// The pre-check further up runs on the pool before this transaction opens, so two hosts
// deleting themselves at the same moment each saw the other and both proceeded, leaving the
// event with NO operator: nobody to moderate, nobody to release the gallery, and no way to
// appoint anyone because appointing requires a host. Not recoverable from inside the app.
//
// Serialised with a transaction-scoped ADVISORY lock, not a row lock. `FOR UPDATE` on the
// other operators\' rows looks like the obvious answer and is the wrong one: each deleter would
// lock the OTHER\'s row and then try to delete its own, so the two block on each other and
// Postgres resolves it by killing one with a deadlock error — the invariant holds, but the
// loser gets a 500 instead of the sentence below. Locking the `event` row instead would
// serialise cleanly, but it inverts the lock order every moderation path uses (upload/user
// rows first, event last). An advisory lock is a separate lock space, so it cannot join the
// row-lock graph at all, and it is released automatically when this transaction ends.
//
// FIRST STATEMENT IN THE TRANSACTION, before any row lock — the ORDER matters as much as the
// lock. `ban_user` and `set_role` take this same lock and then go on to lock `user` and
// `event` rows. If this path grabbed those rows first and reached for the advisory lock
// afterwards, the two would deadlock, each holding what the other needs, and Postgres would
// kill one with a 500: the invariant would survive, but a host deleting their account would
// get an error page instead of the sentence below.
//
// Taking it up front also means the refusal path does no work at all before answering.
if matches!(user.role, UserRole::Host | UserRole::Admin) {
// Shared with `host::ban_user` and `host::set_role` — the same key, by construction rather
// than by two copies agreeing. All three remove an operator, so all three must serialise
// against each other or the floor is enforceable only against its own kind of caller.
crate::handlers::host::lock_operator_floor(&mut tx, auth.event_id).await?;
let others: Vec<uuid::Uuid> = sqlx::query_scalar(
"SELECT id FROM \"user\"
WHERE event_id = $1 AND id != $2
AND role IN ('host', 'admin') AND is_banned = FALSE",
)
.bind(auth.event_id)
.bind(auth.user_id)
.fetch_all(&mut *tx)
.await?;
if others.is_empty() {
return Err(AppError::BadRequest(
"Du bist der letzte Gastgeber. Ernenne zuerst einen anderen Gastgeber, bevor du \
dein Konto löschst."
.into(),
));
}
}
// Comments the guest wrote on OTHER people's photos. Hard delete, not `deleted_at`: this is // Comments the guest wrote on OTHER people's photos. Hard delete, not `deleted_at`: this is
// erasure, and a soft delete leaves the body in the table and in the keepsake's data.json. // erasure, and a soft delete leaves the body in the table and in the keepsake's data.json.
sqlx::query("DELETE FROM comment WHERE user_id = $1") sqlx::query("DELETE FROM comment WHERE user_id = $1")
@@ -199,48 +247,6 @@ pub async fn delete_account(
crate::services::export::Affects::Both, crate::services::export::Affects::Both,
) )
.await?; .await?;
// The last-host guard again, now AUTHORITATIVELY — inside the transaction, holding a lock.
//
// The check above runs on the pool before this transaction opens, so two hosts deleting
// themselves at the same moment each saw the other and both proceeded, leaving the event with
// NO operator: nobody to moderate, nobody to release the gallery, and no way to appoint anyone
// because appointing requires a host. That is not recoverable from inside the app.
//
// Serialised with a transaction-scoped ADVISORY lock, not a row lock.
//
// `FOR UPDATE` on the other operators' rows looks like the obvious answer and is the wrong one:
// each deleter would lock the OTHER's row and then try to delete its own, so the two block on
// each other and Postgres resolves it by killing one with a deadlock error — the invariant
// holds, but the loser gets a 500 instead of the sentence below. Locking the `event` row
// instead would serialise cleanly, but it inverts the lock order every moderation path uses
// (upload/user rows first, event last), which is an ABBA waiting to happen.
//
// An advisory lock has neither problem: it is a separate lock space, so it cannot interact with
// the row-lock graph at all, and it is released automatically when this transaction ends.
// 4242 is an arbitrary namespace to keep this key from colliding with any future advisory use.
if matches!(user.role, UserRole::Host | UserRole::Admin) {
sqlx::query("SELECT pg_advisory_xact_lock(4242, hashtext($1::text))")
.bind(auth.event_id)
.execute(&mut *tx)
.await?;
let others: Vec<uuid::Uuid> = sqlx::query_scalar(
"SELECT id FROM \"user\"
WHERE event_id = $1 AND id != $2
AND role IN ('host', 'admin') AND is_banned = FALSE",
)
.bind(auth.event_id)
.bind(auth.user_id)
.fetch_all(&mut *tx)
.await?;
if others.is_empty() {
return Err(AppError::BadRequest(
"Du bist der letzte Gastgeber. Ernenne zuerst einen anderen Gastgeber, bevor du \
dein Konto löschst."
.into(),
));
}
}
// And the account. `session`, `like` and `pin_reset_request` cascade from here. // And the account. `session`, `like` and `pin_reset_request` cascade from here.
sqlx::query("DELETE FROM \"user\" WHERE id = $1") sqlx::query("DELETE FROM \"user\" WHERE id = $1")
.bind(auth.user_id) .bind(auth.user_id)

View File

@@ -575,7 +575,12 @@ pub async fn upload(
.saturating_add(size); .saturating_add(size);
let keepsake_needs = let keepsake_needs =
crate::services::export::required_free_bytes(media_after.max(0) as u64, 2) as i64; crate::services::export::required_free_bytes(media_after.max(0) as u64, 2) as i64;
let required = keepsake_needs.saturating_add(DISK_RESERVE_BYTES); // Strictly more than the export preflight requires — see `UPLOAD_GATE_HEADROOM_BYTES`.
// Matching it exactly meant the preflight was already at its limit the moment uploads
// stopped, so the night's remaining writes decided whether the keepsake could be built.
let required = keepsake_needs
.saturating_add(DISK_RESERVE_BYTES)
.saturating_add(UPLOAD_GATE_HEADROOM_BYTES);
if free < required { if free < required {
tracing::error!( tracing::error!(
free_bytes = free, free_bytes = free,
@@ -583,6 +588,7 @@ pub async fn upload(
media_after, media_after,
keepsake_needs, keepsake_needs,
reserve = DISK_RESERVE_BYTES, reserve = DISK_RESERVE_BYTES,
headroom = UPLOAD_GATE_HEADROOM_BYTES,
"refusing upload: it would leave too little room to build the keepsake" "refusing upload: it would leave too little room to build the keepsake"
); );
return Err(AppError::QuotaExceeded( return Err(AppError::QuotaExceeded(
@@ -592,10 +598,27 @@ pub async fn upload(
.into(), .into(),
)); ));
} }
} else {
// Failing OPEN when the disk can't be read is deliberate and matches the per-user quota
// below: refusing every upload because a `statfs` failed would be a worse outage than the
// one being guarded against.
//
// But it must not be SILENT. `snapshot` returns `None` when `select_disk` finds neither a
// mount that prefixes the media path nor a `/` entry — and inside a container `/` is an
// overlay rather than a `/dev` device, so this is a real possibility rather than a
// theoretical one. When it happens, the ONLY global disk bound in the app is gone, the
// per-user quota fails open through the same `None`, and the box fills to 100% — at which
// point Postgres cannot write WAL and the whole event stops, with nothing having warned
// anyone. The export preflight already warns on the identical condition; this is the
// louder of the two paths and had no log line at all.
//
// Rate-limited by the disk cache's own TTL, so this cannot spam the log per upload.
tracing::warn!(
media_path = %state.config.media_path.display(),
"disk usage unreadable — the global free-space gate is INACTIVE and uploads are \
proceeding unbounded; check the admin stats page for a plausible free-space figure"
);
} }
// Failing OPEN when the disk can't be read is deliberate and matches the per-user quota
// below: refusing every upload because a `statfs` failed would be a worse outage than the
// one being guarded against.
// When quota is enforced, this holds the byte ceiling so the increment UPDATE below can // When quota is enforced, this holds the byte ceiling so the increment UPDATE below can
// enforce it atomically (`WHERE total + size <= limit`). Without that guard, two // enforce it atomically (`WHERE total + size <= limit`). Without that guard, two
@@ -1240,6 +1263,31 @@ const MIN_QUOTA_LIMIT_BYTES: i64 = 500 * 1024 * 1024;
/// the shared filesystem long after new uploads have been refused. /// the shared filesystem long after new uploads have been refused.
pub const DISK_RESERVE_BYTES: i64 = 10_000_000_000; pub const DISK_RESERVE_BYTES: i64 = 10_000_000_000;
/// Extra free space the UPLOAD gate demands on top of what the export preflight demands.
///
/// Both gates were computing the identical threshold — `required_free_bytes(media, 2) +
/// DISK_RESERVE_BYTES` — which left exactly zero margin between them. The moment the gate refused
/// its first upload, the preflight was already sitting on its own limit, so anything written
/// between that refusal and the host tapping "Galerie freigeben" pushed the preflight under:
///
/// * Postgres WAL, up to `max_wal_size` (1 GB by default) before a checkpoint reclaims it
/// * container logs, capped at 30 MB x 4 services by `docker-compose.yml`
/// * the compression backlog still draining — ~0.9 MB of derivatives per queued photo, and the
/// backlog is longest exactly at the end of the night
///
/// The failure that produces is the worst one in the app: the release COMMITS (event closed,
/// uploads locked, epoch bumped, `event-closed` fanned out to every phone) and only then do both
/// workers bail, at 01:00, with no second release possible and `rebuild_export` needing the same
/// space it just failed to find. Meanwhile ~10 GB of reserve sits unused — the preflight refused
/// on a threshold, not for want of room.
///
/// Giving the upload gate this much more to satisfy means it closes strictly earlier, so ordinary
/// end-of-night writes cannot flip the preflight. The cost is roughly 0.5 GB off the media ceiling
/// on a 40 GB box (the gate's equilibrium is `3.2 x media`, so headroom divides by 3.2), which is
/// the trade `README.md` already argues for: refusing the 1001st upload beats discovering at 01:00
/// that the archive can never be built.
pub const UPLOAD_GATE_HEADROOM_BYTES: i64 = 1_500_000_000;
/// Pure per-user quota formula: `max(floor((free_disk * tolerance) / divisor), MIN)`. /// Pure per-user quota formula: `max(floor((free_disk * tolerance) / divisor), MIN)`.
/// ///
/// `divisor` is the LARGER of the observed uploader count and the operator's /// `divisor` is the LARGER of the observed uploader count and the operator's