fix: close the nine ways an unattended event loses photos or dies
Every one of these was found in the pre-event audit, verified against source, and
survives to production on the current main. Grouped by what actually goes wrong.
PHOTOS DISAPPEAR
* compression.rs no longer soft-deletes on a failed derivative. The guest got a
201, watched the card appear, then watched it vanish — the row left v_feed,
find_visible_media and BOTH keepsakes, while its bytes sat on disk for 14 days
waiting for a cleanup nothing announced. No screen anywhere lists compression
failures, so recovery meant hand-written SQL that also had to re-add the
refunded quota. Now it does exactly what the ENOSPC arm beside it already did
and documented as correct: keep the row, serve the original, retry on the next
boot (bounded by derivative_attempts). `upload-deleted` is no longer emitted;
`upload-processed` is, so the card re-renders instead of sitting on a
placeholder.
* A 413 is now a reversible lock, so the blob survives. The quota moves — free
disk falls, uploader count rises — so a guest goes over it having done nothing,
and treating that as permanent meant a 400 MB video was pushed across cellular
in full and THEN deleted from IndexedDB. Gone on both sides, and unrecoverable
for an in-app camera capture that exists nowhere else.
* quota_limit_bytes gained a floor and a stable divisor. The ceiling used to
decrease monotonically all evening; it now settles at max(uploaders,
estimated_guest_count) — a config key that was seeded, validated in the admin
whitelist, and read by no code at all. The floor is clamped to what the disk
can actually back, so a full volume still yields zero rather than handing out
an allowance it cannot honour.
* Because that floor gives up the aggregate guarantee the formula used to imply,
uploads now check a hard 10 GB reserve first, independent of every quota
toggle. postgres_data, media_data and exports_data share one filesystem: the
end state was not a degraded feature, it was Postgres unable to write WAL.
THE ARCHIVE DISAPPEARS
* prune_superseded_archives runs only after the new generation lands. It ran
before the preflight, reasoning the old archive was already unreachable — true
of reachability, false of recoverability. An epoch is a value that can be
rolled back; deleted bytes cannot. Any failed rebuild left the event with NO
keepsake at all.
* The export preflight reserves the same 10 GB. `free < needed` authorised an
export sized at exactly free, which ran for half an hour and landed the box at
zero with the keepsake still unfinished.
THE APP DIES
* The feed reconcile re-reads the id set after its awaits instead of reusing one
captured up to three round-trips earlier. The new-upload SSE handler prepends
during exactly that window, so the row was both already present and absent from
the stale set — prepended twice, and a duplicate key in a keyed {#each} throws
in production, not just dev. The SSE handler and loadMore now dedupe too.
* Added routes/+error.svelte. Without it any uncaught error fell through to
SvelteKit's unstyled English 500 with no reload control — inside a chromeless
standalone PWA with no URL bar, for the rest of the evening.
THE OPERATOR IS LOCKED OUT
* admin_login verifies the password BEFORE charging the rate bucket, and a
correct password is never throttled. The old order made this a denial of
service against its own operator: every guest shares one NAT IP, the check ran
first, so five requests a minute from any phone in the room kept the bucket
full — and the escape hatch needed the admin session being blocked. A generous
separate ceiling still bounds bcrypt CPU.
THE PROJECTOR DIES
* The preload budget is now strictly inside the dwell. At the 3s option the 4s
budget could never land a commit on a slow uplink, so the wall froze on one
photo while the queue drained silently behind it.
* The wake lock retries every 30s while visible, and the page says so on screen
when the browser has no wake lock API. visibilitychange was the only retry
trigger and a kiosk never changes visibility, so one refusal — iOS in Low Power
Mode, say — was permanent.
* Caddy: /api/v1/upload/*/display joins the cacheable carve-out. The backend set
max-age=300 on it and the blanket no-store silently replaced it, so a projector
re-fetched a full-size JPEG per slide, ~2-4 GB over an evening on the uplink
the guests are uploading over.
Also removes Upload::soft_delete, now unreferenced and an unscoped footgun next
to soft_delete_in_event.
Verified: 146/146 backend tests against a live Postgres, clippy clean, 51/51
vitest, svelte-check 0 errors, eslint clean, vite build, caddy validate, compose
YAML parse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -409,6 +409,14 @@ pub struct AdminLoginResponse {
|
||||
pub display_name: String,
|
||||
}
|
||||
|
||||
/// Requests per minute per IP that may reach `verify_password` at all.
|
||||
///
|
||||
/// Not a security control — the failure bucket below is. This exists solely so an unauthenticated
|
||||
/// endpoint cannot burn the box's CPU on cost-12 bcrypt (~250 ms each) at line rate. Set far above
|
||||
/// anything a person typing a password can produce, because on venue NAT every guest shares the
|
||||
/// operator's IP and this ceiling, unlike the failure bucket, can still refuse a correct password.
|
||||
const ADMIN_LOGIN_CPU_CEILING: usize = 120;
|
||||
|
||||
pub async fn admin_login(
|
||||
State(state): State<AppState>,
|
||||
ConnectInfo(peer): ConnectInfo<SocketAddr>,
|
||||
@@ -421,21 +429,29 @@ pub async fn admin_login(
|
||||
));
|
||||
}
|
||||
|
||||
// Throttle password attempts. The admin password is bcrypt-hashed (slow to
|
||||
// verify) but with no IP-level limit a determined attacker can still mount
|
||||
// a long-running guess campaign. 5 attempts / minute / IP is plenty for
|
||||
// honest typos.
|
||||
// Throttling here is in two parts, and the ORDER is the whole point.
|
||||
//
|
||||
// A single tight IP-keyed bucket checked before the password was verified made this
|
||||
// endpoint a denial-of-service against its own operator. Every guest at the venue shares
|
||||
// one public IP behind NAT, `/admin/login` is a public linkable page, and the check ran
|
||||
// BEFORE `verify_password` — so five requests a minute from any phone in the room kept the
|
||||
// bucket permanently full and the admin, on that same IP, could never spend a slot.
|
||||
// Successful logins consumed budget too, so a typo plus a retry on two devices did it by
|
||||
// accident. And the escape hatch was circular: `admin_login_rate_enabled` can only be
|
||||
// flipped through `PATCH /admin/config`, which needs the session being blocked.
|
||||
let ip = client_ip(&headers, &peer.ip().to_string());
|
||||
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
|
||||
let admin_rate_on =
|
||||
config::get_bool(&state.config_cache, "admin_login_rate_enabled", true).await;
|
||||
// Stays keyed by IP on purpose: this guards a single shared credential, so a per-user
|
||||
// or per-name key would just hand an attacker a fresh bucket per guess.
|
||||
|
||||
// Part 1: a deliberately GENEROUS ceiling whose only job is to bound bcrypt CPU. Cost-12
|
||||
// verification is ~250 ms of a core, so an unbounded endpoint is a CPU exhaustion vector
|
||||
// regardless of whether anyone guesses right. No human typing a password reaches this.
|
||||
if rate_limits_on
|
||||
&& admin_rate_on
|
||||
&& let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
|
||||
format!("admin_login:{ip}"),
|
||||
5,
|
||||
format!("admin_login_cpu:{ip}"),
|
||||
ADMIN_LOGIN_CPU_CEILING,
|
||||
Duration::from_secs(60),
|
||||
)
|
||||
{
|
||||
@@ -452,6 +468,24 @@ pub async fn admin_login(
|
||||
.await;
|
||||
|
||||
if !valid {
|
||||
// Part 2: the tight bucket, charged ONLY on a wrong password. A correct password is
|
||||
// never rate-limited, so no amount of guessing by anyone else can lock the operator
|
||||
// out — which also dissolves the circular escape hatch above. Brute force is still
|
||||
// bounded: every wrong guess costs a slot, and slots are per-IP.
|
||||
if rate_limits_on
|
||||
&& admin_rate_on
|
||||
&& let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
|
||||
format!("admin_login_fail:{ip}"),
|
||||
5,
|
||||
Duration::from_secs(60),
|
||||
)
|
||||
{
|
||||
tracing::warn!(ip = %ip, "admin_login: wrong password, failure bucket exhausted");
|
||||
return Err(AppError::TooManyRequests(
|
||||
"Zu viele Anmeldeversuche. Bitte warte kurz und versuche es erneut.".into(),
|
||||
Some(retry_after_secs),
|
||||
));
|
||||
}
|
||||
tracing::warn!(ip = %ip, "admin_login: wrong password");
|
||||
return Err(AppError::Unauthorized("Falsches Passwort.".into()));
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user