fix: gate uploads on keepsake headroom, and close five unattended-event gaps

The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit,
the committed comments and README's sizing section all assumed. That correction
is what the first change is about; the rest are the remaining pre-event items.

THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING

`required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each
gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The
upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let
uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free.
Every upload in that band succeeded and the keepsake could then never be built: the
product's entire promise, failing silently at the end of the night with nobody there.

The gate now enforces the invariant that actually matters — never accept an upload
that would make the keepsake unbuildable — sharing `required_free_bytes` with the
preflight so the two cannot drift into disagreeing about the same question. Uploads
stop at ~8 GB of media on this disk, with a German message naming the cause.
Refusing the 1001st photo beats losing all 1000.

`media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s,
rather than `estimate_export_bytes`'s join across every upload. It counts hidden and
banned users' bytes, which the export excludes — skew in the SAFE direction, so the
gate closes marginally early rather than late. Fails open on a query error.

A test pins the gate against the preflight across the whole gallery-size range, and
a second asserts the per-user floor alone would over-commit the volume — i.e. that
the global gate is what must bind.

THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED

`Date.now()` advances while a backgrounded phone is frozen but `setInterval` does
not, so the first tick after a screen lock read the whole sleep as silence and
aborted — re-sending a video from byte zero and burning one of five PERMANENT
auto-attempts. The interval is now its own suspension detector: a tick that arrives
125s late for a 5s schedule credits that window back, because a period the watchdog
could not observe is not evidence of silence.

Chosen over a `visibilitychange` listener, which only covers causes that fire that
event — a throttled-but-visible tab, a closed lid and an occluded window all freeze
timers without one — and which would have needed module state, an SSR guard and a
teardown for strictly less coverage. `performance.now()` was rejected because Safari
pauses it across system sleep on some paths and Chrome does not.

The credit buys one fresh window, not immunity: a socket iOS reaped while
backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout`
(5-60 min) with the queue's `processing` latch held.

Two latent leaks found while in there: `xhr.abort()` on a request already in
readyState DONE emits no `abort` event, so `settle()` never ran and the interval
re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button
silently stopped working); and a synchronous throw from `xhr.send` — a blob whose
backing store the OS purged — leaked the same way. Both closed.

OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER

red/amber/green were never in the @theme block and fell through to Tailwind v4's
`oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot
parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color`
falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to
Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already
in this pipeline — so modern browsers render exactly what they render today. Verified
against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal
(avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0.

Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on
the host dashboard was rendering default violet on EVERY browser, off-brand.

The keepsake viewer only picks this up on a rebuild, so its committed artefact is
rebuilt here too — still single-file, still zero external references.

A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER

With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a
redeploy or a dead uplink left the guest on the boot spinner with no message, no
reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing
nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a
timeout rather than feature detection: a SyntaxError in the bundle is invisible to any
capability check. Plus a <noscript>, since there was nothing at all to see without JS.

EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL

tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info.
If guests spend the evening hitting 429s or 413s, the post-event logs said nothing.
Now one WARN per client error; 5xx excluded because Internal already logs its source
chain and the pool-exhaustion 503 logs at construction.

A DEAD FRONTEND SERVED A BLANK 502

`handle_errors 5xx` with an inline German page (the caddy service mounts only the
Caddyfile, so there is no volume to ship a static file through). Verified empirically
against this config, not from documentation: an upstream 404 through `reverse_proxy`
still arrives as untouched `application/json`, and only a dial failure renders the
page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a
real 404/429 arriving, and swallowing those would have been worse than the blank 502.

CONFIG CORRECTIONS FOR THE REAL HARDWARE

DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count.
Since migration 024 a feed page costs well under a millisecond, so connections are no
longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host.
COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already
serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was
sized against is unreachable; dropping to 1 would halve light-path throughput and push
more feed tiles onto full-size originals. README's sizing section rewritten for the
actual disk.

Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest,
svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate,
compose YAML parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-08-08 22:07:56 +02:00
parent 1d9fb11c7b
commit eb0e405562
16 changed files with 624 additions and 64 deletions

View File

@@ -78,6 +78,26 @@ impl IntoResponse for AppError {
};
let message = self.message();
// Log every 4xx. Until now they were invisible at ANY log level: tower_http's
// `ServerErrorsAsFailures` classifier counts a 4xx as a *success*, so it goes to
// `DefaultOnResponse` at DEBUG, and production runs at `info`. The consequence is that a
// misconfigured limit leaves no trace at all — if guests spend the evening hitting 429s
// on `upload_rate_per_hour`, or 413s on the storage quota, `docker compose logs` after
// the event contains nothing about it and the cause is unknowable.
//
// WARN rather than INFO because every variant here is a request that did not do what
// the guest asked. 5xx is excluded: `Internal` already logs with its full source chain
// in `message()` above, and the pool-exhaustion 503 logs at construction — logging again
// here would double every server-side failure.
//
// No request context is available: `into_response` receives only the error, so there is
// no path, method or user id to attach. Status + code + message is what can honestly be
// reported from here, and it is enough to see the SHAPE of a bad evening. Raising
// `tower_http` to DEBUG instead was considered and rejected — see the note in main.rs.
if status.is_client_error() {
tracing::warn!(status = status.as_u16(), code, %message, "request rejected");
}
let mut body = serde_json::json!({
"error": code,
"message": message,
@@ -161,6 +181,43 @@ mod tests {
}
}
/// 4xx must be logged and 5xx must not be logged HERE — `Internal` logs its source chain in
/// `message()` and the pool-exhaustion 503 logs at construction, so a second line in
/// `into_response` would double every server-side failure in the post-event logs.
///
/// The guard is `status.is_client_error()`, so this pins the classification rather than the
/// logging itself (which needs a subscriber to observe).
#[test]
fn only_client_errors_are_in_the_logged_band() {
for err in [
AppError::BadRequest("x".into()),
AppError::Unauthorized("x".into()),
AppError::Forbidden("x".into()),
AppError::UploadsLocked("x".into()),
AppError::NotFound("x".into()),
AppError::Conflict("x".into()),
AppError::TooManyRequests("x".into(), Some(1)),
AppError::QuotaExceeded("x".into()),
] {
let (status, _) = err.status_and_code();
assert!(
status.is_client_error(),
"{status} should be in the 4xx band this logs"
);
}
for err in [
AppError::ServiceUnavailable("x".into(), Some(3)),
AppError::Internal(anyhow::anyhow!("boom")),
] {
let (status, _) = err.status_and_code();
assert!(
!status.is_client_error(),
"{status} logs elsewhere; logging it here would double it"
);
}
}
/// Pool saturation is load, not a bug. A 500 makes the frontend's retry classifier pile
/// straight back into the saturated pool with no backoff to pace it.
#[test]

View File

@@ -111,6 +111,12 @@ pub async fn truncate_all(
// steers the per-user limit off `free_disk_bytes`), i.e. two holes were masking each other.
state.disk_cache.invalidate();
// `media_total` caches SUM(user.total_upload_bytes) for the upload gate's keepsake-headroom
// check. TRUNCATE has just zeroed every one of those rows, so a surviving reading would make
// the next test's first upload measure its headroom against the previous test's gallery —
// and that gate REFUSES uploads, so the failure would look like a spurious quota rejection.
state.media_total.invalidate();
// `sse_tickets` maps a ticket to a session token hash. TRUNCATE deletes the sessions, so every
// surviving ticket is a dangling reference to a user that no longer exists.
state.sse_tickets.clear();

View File

@@ -438,43 +438,62 @@ pub async fn upload(
let quota_on = config::get_bool(&state.config_cache, "quota_enabled", true).await;
let storage_quota_on =
config::get_bool(&state.config_cache, "storage_quota_enabled", true).await;
// When quota is enforced, this holds the byte ceiling so the increment UPDATE below can
// enforce it atomically (`WHERE total + size <= limit`). Without that guard, two
// concurrent uploads from the same user (e.g. phone + laptop) both pass this stale
// pre-check and both increment, blowing past the quota. The pre-check stays as a
// fast path that avoids the disk write when the user is already clearly over.
// GLOBAL RESERVE, checked before the per-user ceiling and independent of every quota
// GLOBAL DISK GATE, checked before the per-user ceiling and independent of every quota
// toggle. The per-user quota is a fairness mechanism, not a disk guarantee — and since it
// now carries a floor (MIN_QUOTA_LIMIT_BYTES) so a guest's allowance stops shrinking as
// the party fills up, the aggregate ceiling it used to imply is gone entirely. Something
// has to own "do not fill the volume", because `postgres_data`, `media_data` and
// `exports_data` share one filesystem: the end state is not a degraded feature, it is
// Postgres unable to write WAL and the whole event down with nobody watching.
// carries a floor (MIN_QUOTA_LIMIT_BYTES) so a guest's allowance stops shrinking as the
// party fills up, the aggregate ceiling it used to imply is gone entirely. Something has to
// own "do not fill the volume", because `postgres_data`, `media_data` and `exports_data`
// share one filesystem: the end state is not a degraded feature, it is Postgres unable to
// write WAL and the whole event down with nobody watching.
//
// Deliberately NOT gated behind `quota_enabled`. That switch exists so an operator can
// stop rationing space between guests; it was never meant to authorise running the disk
// to zero, and an operator flipping it at 23:00 to unblock a guest should not silently
// disarm the last thing standing between the party and a dead database.
// WHAT IS RESERVED IS NOT A CONSTANT. A flat reserve answers "can Postgres still write",
// which is necessary and not sufficient: the keepsake needs room for BOTH halves at once —
// `required_free_bytes` is `media × 1.1 × 2`, since the ZIP and the HTML viewer are each
// gallery-sized. On the 40 GB box this runs on, a flat 10 GB reserve let uploads continue to
// roughly 25 GB of media while the release needed `2.2 × 25 + 10` = 65 GB free. Every upload
// in that band succeeded and then the archive could never be built — the product's entire
// promise, failing silently at the end of the night with nobody there to notice.
//
// So the gate enforces the invariant that actually matters: never accept an upload that
// would make the keepsake unbuildable. It shares `required_free_bytes` with the export
// preflight so the two cannot drift into disagreeing about the same question.
//
// Deliberately NOT gated behind `quota_enabled`. That switch exists so an operator can stop
// rationing space between guests; it was never meant to authorise running the disk to zero,
// and an operator flipping it at 23:00 to unblock a guest should not silently disarm the
// last thing standing between the party and a dead database.
if let Some(free) = crate::services::disk::free_bytes(&state.config.media_path) {
let remaining = (free as i64).saturating_sub(size);
if remaining < DISK_RESERVE_BYTES {
let media_after = state.media_total.get(&state.pool).await.saturating_add(size);
let keepsake_needs =
crate::services::export::required_free_bytes(media_after.max(0) as u64, 2) as i64;
let free_after = (free as i64).saturating_sub(size);
let required = keepsake_needs.saturating_add(DISK_RESERVE_BYTES);
if free_after < required {
tracing::error!(
free_bytes = free,
upload_size = size,
media_after,
keepsake_needs,
reserve = DISK_RESERVE_BYTES,
"refusing upload: it would take the media volume below the reserve"
"refusing upload: it would leave too little room to build the keepsake"
);
return Err(AppError::QuotaExceeded(
"Der Speicher des Events ist voll. Bitte sag einem Host Bescheid — neue \
Uploads sind vorübergehend nicht möglich."
"Der Speicher des Events ist fast voll — damit die Galerie am Ende noch als \
Download erstellt werden kann, sind neue Uploads jetzt gesperrt. Bitte sag \
einem Host Bescheid."
.into(),
));
}
}
// Failing OPEN when the disk can't be read is deliberate and matches the per-user quota
// above: refusing every upload because a `statfs` failed would be a worse outage than the
// below: refusing every upload because a `statfs` failed would be a worse outage than the
// one being guarded against.
// When quota is enforced, this holds the byte ceiling so the increment UPDATE below can
// enforce it atomically (`WHERE total + size <= limit`). Without that guard, two
// concurrent uploads from the same user (e.g. phone + laptop) both pass this stale
// pre-check and both increment, blowing past the quota. The pre-check stays as a
// fast path that avoids the disk write when the user is already clearly over.
let mut quota_limit: Option<i64> = None;
if quota_on && storage_quota_on {
let estimate = compute_storage_quota(&state).await;
@@ -1370,7 +1389,9 @@ pub async fn get_thumbnail(
#[cfg(test)]
mod tests {
use super::{MIN_QUOTA_LIMIT_BYTES, RangeSpec, parse_range, quota_limit_bytes};
use super::{
DISK_RESERVE_BYTES, MIN_QUOTA_LIMIT_BYTES, RangeSpec, parse_range, quota_limit_bytes,
};
// `Range` handling exists because iOS Safari probes every `<video>` with
// `Range: bytes=0-1` and abandons the load without a 206. These pin the forms a
@@ -1523,6 +1544,67 @@ mod tests {
);
}
/// THE INVARIANT THE UPLOAD GATE EXISTS FOR: if an upload is accepted, the keepsake must
/// still be buildable afterwards.
///
/// The gate and `ensure_export_space` answer the same question at different times, from the
/// same `required_free_bytes`. If they ever drift, the failure is silent and terminal — every
/// upload succeeds and the archive can never be built, discovered only when the host taps
/// release and there is nobody left to fix it. This pins the two together.
///
/// Models the real box: 40 GB volume, ~5 GB consumed by OS, images and Postgres.
#[test]
fn an_accepted_upload_always_leaves_room_to_build_the_keepsake() {
const USABLE: i64 = 35 * GB;
let reserve = DISK_RESERVE_BYTES;
// Walk the gallery upward in 250 MB steps and assert the two agree at every point.
let mut media: i64 = 0;
let step: i64 = 250 * 1024 * 1024;
let mut last_accepted = 0i64;
while media < USABLE {
let free = USABLE - media;
let media_after = media + step;
let free_after = free - step;
let required =
crate::services::export::required_free_bytes(media_after as u64, 2) as i64 + reserve;
let gate_accepts = free_after >= required;
if gate_accepts {
// The export preflight must agree, using the SAME arithmetic it will run later.
let preflight_needs =
crate::services::export::required_free_bytes(media_after as u64, 2) as i64
+ reserve;
assert!(
free_after >= preflight_needs,
"gate accepted at media={media_after} but the preflight would refuse"
);
last_accepted = media_after;
}
media = media_after;
}
// Sanity-check the ceiling is where the arithmetic says: 35 = 2.2·M + 10 ⇒ M ≈ 7.8 GB.
// Pinned loosely (69 GB) so a deliberate change to the overhead multiplier or the
// reserve fails this test loudly rather than silently moving the cliff.
assert!(
(6 * GB..=9 * GB).contains(&last_accepted),
"expected the gallery ceiling near 7.8 GB on a 35 GB volume, got {last_accepted} bytes"
);
}
/// The gate must be the binding constraint, not the per-user floor. With 100 guests each
/// allowed 500 MB, the per-user quota alone would authorise ~50 GB on a 40 GB disk.
#[test]
fn the_global_gate_binds_before_the_per_user_floor_can_overfill_the_disk() {
let per_user_total = MIN_QUOTA_LIMIT_BYTES * 100;
assert!(
per_user_total > 35 * GB,
"premise: the per-user floor alone over-commits the volume, so the global gate \
is what must stop it"
);
}
/// The floor must never write a cheque the volume cannot cash — otherwise a full disk
/// still hands out a 500 MB allowance and the filesystem Postgres needs fills up.
#[test]

View File

@@ -1354,7 +1354,7 @@ const EXPORT_SIZE_OVERHEAD_PCT: u64 = 110;
/// Computed in `u128` and clamped, NOT with `saturating_mul`: saturating first and then dividing by
/// 100 quietly turns an overflow into a number ~100x too small, which is the one direction that
/// matters here — an under-estimate authorises the very write the preflight exists to refuse.
fn required_free_bytes(media_bytes: u64, armed: i64) -> u64 {
pub(crate) fn required_free_bytes(media_bytes: u64, armed: i64) -> u64 {
let needed = media_bytes as u128 * EXPORT_SIZE_OVERHEAD_PCT as u128 / 100
* armed.max(1).min(i64::from(u32::MAX)) as u128;
needed.min(u64::MAX as u128) as u64

View File

@@ -0,0 +1,82 @@
//! Cached sum of all media bytes the event is holding.
//!
//! The upload gate needs to know "how big would the keepsake be if we accept this file", because
//! the archive needs room for BOTH halves at once (`export::required_free_bytes` is
//! `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized). Asking that question
//! per upload has to be cheap, and it has to be cheap on the busiest write path in the app.
//!
//! `export::estimate_export_bytes` answers the same question exactly, but it aggregates
//! `original_size_bytes` across every upload row joined to `user` — fine once per release,
//! wasteful per upload and growing all evening. This sums `user.total_upload_bytes` instead:
//! one row per guest (~100), already maintained transactionally by the quota path, already
//! refunded on delete.
//!
//! The two differ slightly — this one counts uploads belonging to banned or hidden users, which
//! the export filters out. That skew is in the SAFE direction: it over-estimates the archive, so
//! the gate closes marginally early rather than marginally late. Never swap it for a cheaper
//! query that could under-estimate; an under-estimate authorises the very upload that makes the
//! keepsake unbuildable, which is the failure this exists to prevent.
use std::sync::{Arc, RwLock};
use std::time::{Duration, Instant};
use sqlx::PgPool;
/// How long a reading is trusted. Shorter than [`crate::services::disk`]'s TTL because this
/// number only ever grows and does so in the same request path that reads it — a stale value
/// under-counts the newest uploads, and under-counting is the direction that matters.
const TTL: Duration = Duration::from_secs(5);
/// Cheap-to-clone cache of the event's total media bytes. Lives in `AppState`.
#[derive(Clone)]
pub struct MediaTotalCache {
inner: Arc<RwLock<Option<(i64, Instant)>>>,
}
impl MediaTotalCache {
pub fn new() -> Self {
Self {
inner: Arc::new(RwLock::new(None)),
}
}
/// Drop the cached reading so the next `get()` re-queries.
///
/// Used by the e2e TRUNCATE endpoint for the same reason `DiskCache::invalidate` exists:
/// truncation removes every upload, and a surviving reading would make the next test's
/// gate compute against the previous test's data.
pub fn invalidate(&self) {
*self.inner.write().unwrap() = None;
}
/// Total bytes of media the event is holding, cached for [`TTL`].
///
/// Returns 0 when the query fails. That is a deliberate FAIL-OPEN, consistent with the
/// quota path and the export preflight: a database blip must not turn into "every upload
/// refused". The disk-space half of the gate still applies, so a failure here degrades the
/// check to the old flat-reserve behaviour rather than disabling it.
pub async fn get(&self, pool: &PgPool) -> i64 {
if let Some((bytes, at)) = *self.inner.read().unwrap()
&& at.elapsed() < TTL
{
return bytes;
}
let bytes = sqlx::query_scalar::<_, Option<i64>>(
"SELECT SUM(total_upload_bytes)::bigint FROM \"user\"",
)
.fetch_one(pool)
.await
.ok()
.flatten()
.unwrap_or(0)
.max(0);
*self.inner.write().unwrap() = Some((bytes, Instant::now()));
bytes
}
}
impl Default for MediaTotalCache {
fn default() -> Self {
Self::new()
}
}

View File

@@ -4,6 +4,7 @@ pub mod disk;
pub mod export;
pub mod imaging;
pub mod maintenance;
pub mod media_total;
pub mod rate_limiter;
pub mod sse_tickets;
pub mod video;

View File

@@ -5,6 +5,7 @@ use crate::config::AppConfig;
use crate::services::compression::CompressionWorker;
use crate::services::config::ConfigCache;
use crate::services::disk::DiskCache;
use crate::services::media_total::MediaTotalCache;
use crate::services::rate_limiter::RateLimiter;
use crate::services::sse_tickets::SseTicketStore;
@@ -38,6 +39,8 @@ pub struct AppState {
pub config_cache: ConfigCache,
/// Cached total/free bytes for the media filesystem (quota + admin stats).
pub disk_cache: DiskCache,
/// Cached sum of all media bytes, for the upload gate's keepsake-headroom check.
pub media_total: MediaTotalCache,
}
impl AppState {
@@ -63,6 +66,7 @@ impl AppState {
sse_tickets: SseTicketStore::new(),
config_cache,
disk_cache: DiskCache::new(),
media_total: MediaTotalCache::new(),
}
}
}

File diff suppressed because one or more lines are too long