fix: gate uploads on keepsake headroom, and close five unattended-event gaps
The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit, the committed comments and README's sizing section all assumed. That correction is what the first change is about; the rest are the remaining pre-event items. THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING `required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free. Every upload in that band succeeded and the keepsake could then never be built: the product's entire promise, failing silently at the end of the night with nobody there. The gate now enforces the invariant that actually matters — never accept an upload that would make the keepsake unbuildable — sharing `required_free_bytes` with the preflight so the two cannot drift into disagreeing about the same question. Uploads stop at ~8 GB of media on this disk, with a German message naming the cause. Refusing the 1001st photo beats losing all 1000. `media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s, rather than `estimate_export_bytes`'s join across every upload. It counts hidden and banned users' bytes, which the export excludes — skew in the SAFE direction, so the gate closes marginally early rather than late. Fails open on a query error. A test pins the gate against the preflight across the whole gallery-size range, and a second asserts the per-user floor alone would over-commit the volume — i.e. that the global gate is what must bind. THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED `Date.now()` advances while a backgrounded phone is frozen but `setInterval` does not, so the first tick after a screen lock read the whole sleep as silence and aborted — re-sending a video from byte zero and burning one of five PERMANENT auto-attempts. The interval is now its own suspension detector: a tick that arrives 125s late for a 5s schedule credits that window back, because a period the watchdog could not observe is not evidence of silence. Chosen over a `visibilitychange` listener, which only covers causes that fire that event — a throttled-but-visible tab, a closed lid and an occluded window all freeze timers without one — and which would have needed module state, an SSR guard and a teardown for strictly less coverage. `performance.now()` was rejected because Safari pauses it across system sleep on some paths and Chrome does not. The credit buys one fresh window, not immunity: a socket iOS reaped while backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout` (5-60 min) with the queue's `processing` latch held. Two latent leaks found while in there: `xhr.abort()` on a request already in readyState DONE emits no `abort` event, so `settle()` never ran and the interval re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button silently stopped working); and a synchronous throw from `xhr.send` — a blob whose backing store the OS purged — leaked the same way. Both closed. OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER red/amber/green were never in the @theme block and fell through to Tailwind v4's `oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color` falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already in this pipeline — so modern browsers render exactly what they render today. Verified against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal (avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0. Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on the host dashboard was rendering default violet on EVERY browser, off-brand. The keepsake viewer only picks this up on a rebuild, so its committed artefact is rebuilt here too — still single-file, still zero external references. A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a redeploy or a dead uplink left the guest on the boot spinner with no message, no reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a timeout rather than feature detection: a SyntaxError in the bundle is invisible to any capability check. Plus a <noscript>, since there was nothing at all to see without JS. EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info. If guests spend the evening hitting 429s or 413s, the post-event logs said nothing. Now one WARN per client error; 5xx excluded because Internal already logs its source chain and the pool-exhaustion 503 logs at construction. A DEAD FRONTEND SERVED A BLANK 502 `handle_errors 5xx` with an inline German page (the caddy service mounts only the Caddyfile, so there is no volume to ship a static file through). Verified empirically against this config, not from documentation: an upstream 404 through `reverse_proxy` still arrives as untouched `application/json`, and only a dial failure renders the page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a real 404/429 arriving, and swallowing those would have been worse than the blank 502. CONFIG CORRECTIONS FOR THE REAL HARDWARE DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count. Since migration 024 a feed page costs well under a millisecond, so connections are no longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host. COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was sized against is unreachable; dropping to 1 would halve light-path throughput and push more feed tiles onto full-size originals. README's sizing section rewritten for the actual disk. Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest, svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate, compose YAML parse. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1354,7 +1354,7 @@ const EXPORT_SIZE_OVERHEAD_PCT: u64 = 110;
|
||||
/// Computed in `u128` and clamped, NOT with `saturating_mul`: saturating first and then dividing by
|
||||
/// 100 quietly turns an overflow into a number ~100x too small, which is the one direction that
|
||||
/// matters here — an under-estimate authorises the very write the preflight exists to refuse.
|
||||
fn required_free_bytes(media_bytes: u64, armed: i64) -> u64 {
|
||||
pub(crate) fn required_free_bytes(media_bytes: u64, armed: i64) -> u64 {
|
||||
let needed = media_bytes as u128 * EXPORT_SIZE_OVERHEAD_PCT as u128 / 100
|
||||
* armed.max(1).min(i64::from(u32::MAX)) as u128;
|
||||
needed.min(u64::MAX as u128) as u64
|
||||
|
||||
82
backend/src/services/media_total.rs
Normal file
82
backend/src/services/media_total.rs
Normal file
@@ -0,0 +1,82 @@
|
||||
//! Cached sum of all media bytes the event is holding.
|
||||
//!
|
||||
//! The upload gate needs to know "how big would the keepsake be if we accept this file", because
|
||||
//! the archive needs room for BOTH halves at once (`export::required_free_bytes` is
|
||||
//! `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized). Asking that question
|
||||
//! per upload has to be cheap, and it has to be cheap on the busiest write path in the app.
|
||||
//!
|
||||
//! `export::estimate_export_bytes` answers the same question exactly, but it aggregates
|
||||
//! `original_size_bytes` across every upload row joined to `user` — fine once per release,
|
||||
//! wasteful per upload and growing all evening. This sums `user.total_upload_bytes` instead:
|
||||
//! one row per guest (~100), already maintained transactionally by the quota path, already
|
||||
//! refunded on delete.
|
||||
//!
|
||||
//! The two differ slightly — this one counts uploads belonging to banned or hidden users, which
|
||||
//! the export filters out. That skew is in the SAFE direction: it over-estimates the archive, so
|
||||
//! the gate closes marginally early rather than marginally late. Never swap it for a cheaper
|
||||
//! query that could under-estimate; an under-estimate authorises the very upload that makes the
|
||||
//! keepsake unbuildable, which is the failure this exists to prevent.
|
||||
|
||||
use std::sync::{Arc, RwLock};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
use sqlx::PgPool;
|
||||
|
||||
/// How long a reading is trusted. Shorter than [`crate::services::disk`]'s TTL because this
|
||||
/// number only ever grows and does so in the same request path that reads it — a stale value
|
||||
/// under-counts the newest uploads, and under-counting is the direction that matters.
|
||||
const TTL: Duration = Duration::from_secs(5);
|
||||
|
||||
/// Cheap-to-clone cache of the event's total media bytes. Lives in `AppState`.
|
||||
#[derive(Clone)]
|
||||
pub struct MediaTotalCache {
|
||||
inner: Arc<RwLock<Option<(i64, Instant)>>>,
|
||||
}
|
||||
|
||||
impl MediaTotalCache {
|
||||
pub fn new() -> Self {
|
||||
Self {
|
||||
inner: Arc::new(RwLock::new(None)),
|
||||
}
|
||||
}
|
||||
|
||||
/// Drop the cached reading so the next `get()` re-queries.
|
||||
///
|
||||
/// Used by the e2e TRUNCATE endpoint for the same reason `DiskCache::invalidate` exists:
|
||||
/// truncation removes every upload, and a surviving reading would make the next test's
|
||||
/// gate compute against the previous test's data.
|
||||
pub fn invalidate(&self) {
|
||||
*self.inner.write().unwrap() = None;
|
||||
}
|
||||
|
||||
/// Total bytes of media the event is holding, cached for [`TTL`].
|
||||
///
|
||||
/// Returns 0 when the query fails. That is a deliberate FAIL-OPEN, consistent with the
|
||||
/// quota path and the export preflight: a database blip must not turn into "every upload
|
||||
/// refused". The disk-space half of the gate still applies, so a failure here degrades the
|
||||
/// check to the old flat-reserve behaviour rather than disabling it.
|
||||
pub async fn get(&self, pool: &PgPool) -> i64 {
|
||||
if let Some((bytes, at)) = *self.inner.read().unwrap()
|
||||
&& at.elapsed() < TTL
|
||||
{
|
||||
return bytes;
|
||||
}
|
||||
let bytes = sqlx::query_scalar::<_, Option<i64>>(
|
||||
"SELECT SUM(total_upload_bytes)::bigint FROM \"user\"",
|
||||
)
|
||||
.fetch_one(pool)
|
||||
.await
|
||||
.ok()
|
||||
.flatten()
|
||||
.unwrap_or(0)
|
||||
.max(0);
|
||||
*self.inner.write().unwrap() = Some((bytes, Instant::now()));
|
||||
bytes
|
||||
}
|
||||
}
|
||||
|
||||
impl Default for MediaTotalCache {
|
||||
fn default() -> Self {
|
||||
Self::new()
|
||||
}
|
||||
}
|
||||
@@ -4,6 +4,7 @@ pub mod disk;
|
||||
pub mod export;
|
||||
pub mod imaging;
|
||||
pub mod maintenance;
|
||||
pub mod media_total;
|
||||
pub mod rate_limiter;
|
||||
pub mod sse_tickets;
|
||||
pub mod video;
|
||||
|
||||
Reference in New Issue
Block a user