fix: gate uploads on keepsake headroom, and close five unattended-event gaps
The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit, the committed comments and README's sizing section all assumed. That correction is what the first change is about; the rest are the remaining pre-event items. THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING `required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free. Every upload in that band succeeded and the keepsake could then never be built: the product's entire promise, failing silently at the end of the night with nobody there. The gate now enforces the invariant that actually matters — never accept an upload that would make the keepsake unbuildable — sharing `required_free_bytes` with the preflight so the two cannot drift into disagreeing about the same question. Uploads stop at ~8 GB of media on this disk, with a German message naming the cause. Refusing the 1001st photo beats losing all 1000. `media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s, rather than `estimate_export_bytes`'s join across every upload. It counts hidden and banned users' bytes, which the export excludes — skew in the SAFE direction, so the gate closes marginally early rather than late. Fails open on a query error. A test pins the gate against the preflight across the whole gallery-size range, and a second asserts the per-user floor alone would over-commit the volume — i.e. that the global gate is what must bind. THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED `Date.now()` advances while a backgrounded phone is frozen but `setInterval` does not, so the first tick after a screen lock read the whole sleep as silence and aborted — re-sending a video from byte zero and burning one of five PERMANENT auto-attempts. The interval is now its own suspension detector: a tick that arrives 125s late for a 5s schedule credits that window back, because a period the watchdog could not observe is not evidence of silence. Chosen over a `visibilitychange` listener, which only covers causes that fire that event — a throttled-but-visible tab, a closed lid and an occluded window all freeze timers without one — and which would have needed module state, an SSR guard and a teardown for strictly less coverage. `performance.now()` was rejected because Safari pauses it across system sleep on some paths and Chrome does not. The credit buys one fresh window, not immunity: a socket iOS reaped while backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout` (5-60 min) with the queue's `processing` latch held. Two latent leaks found while in there: `xhr.abort()` on a request already in readyState DONE emits no `abort` event, so `settle()` never ran and the interval re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button silently stopped working); and a synchronous throw from `xhr.send` — a blob whose backing store the OS purged — leaked the same way. Both closed. OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER red/amber/green were never in the @theme block and fell through to Tailwind v4's `oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color` falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already in this pipeline — so modern browsers render exactly what they render today. Verified against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal (avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0. Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on the host dashboard was rendering default violet on EVERY browser, off-brand. The keepsake viewer only picks this up on a rebuild, so its committed artefact is rebuilt here too — still single-file, still zero external references. A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a redeploy or a dead uplink left the guest on the boot spinner with no message, no reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a timeout rather than feature detection: a SyntaxError in the bundle is invisible to any capability check. Plus a <noscript>, since there was nothing at all to see without JS. EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info. If guests spend the evening hitting 429s or 413s, the post-event logs said nothing. Now one WARN per client error; 5xx excluded because Internal already logs its source chain and the pool-exhaustion 503 logs at construction. A DEAD FRONTEND SERVED A BLANK 502 `handle_errors 5xx` with an inline German page (the caddy service mounts only the Caddyfile, so there is no volume to ship a static file through). Verified empirically against this config, not from documentation: an upstream 404 through `reverse_proxy` still arrives as untouched `application/json`, and only a dial failure renders the page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a real 404/429 arriving, and swallowing those would have been worse than the blank 502. CONFIG CORRECTIONS FOR THE REAL HARDWARE DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count. Since migration 024 a feed page costs well under a millisecond, so connections are no longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host. COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was sized against is unreachable; dropping to 1 would halve light-path throughput and push more feed tiles onto full-size originals. README's sizing section rewritten for the actual disk. Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest, svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate, compose YAML parse. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
36
.env.example
36
.env.example
@@ -36,13 +36,20 @@ DATABASE_URL=postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/events
|
||||
POSTGRES_USER=eventsnap
|
||||
POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password
|
||||
POSTGRES_DB=eventsnap
|
||||
# Connection pool size. Default 10. For a busy event (~100 guests polling the feed
|
||||
# + SSE + uploads at once) raise to ~30 so requests don't queue on a pool permit.
|
||||
# PAIRED WITH THE DB CONTAINER'S MEMORY LIMIT: 30 backends plus Postgres 16's default
|
||||
# shared_buffers is already snug in the 1G that docker-compose.yml allots the `db`
|
||||
# service. If you raise this, raise `db.deploy.resources.limits.memory` with it — an
|
||||
# OOM in Postgres doesn't degrade one feature, it takes the whole event down.
|
||||
DATABASE_MAX_CONNECTIONS=30
|
||||
# Connection pool size. The code default is 10 (backend/src/db.rs) — set it explicitly,
|
||||
# because a `.env` written by hand from this file's secrets is otherwise silently on 10.
|
||||
#
|
||||
# SIZE IT TO THE CORES, NOT TO THE GUESTS. The earlier advice here was ~30, reasoned from
|
||||
# "~100 guests polling the feed at once" back when a feed page cost ~449 ms and connections
|
||||
# were spent waiting. Migration 024 replaced the feed view's GROUP BY with scalar subqueries
|
||||
# and a page now costs well under a millisecond, so concurrency is no longer where the time
|
||||
# goes. On a 2 vCPU box 30 simultaneous queries cannot run — they queue on the CPU instead of
|
||||
# on the pool, which is the same wait wearing a different hat, and 30 Postgres backends plus
|
||||
# shared_buffers is snug in the 1G that docker-compose.yml allots `db`.
|
||||
#
|
||||
# 15 on 2 vCPU / 4 GB. Raise toward 30 only alongside more cores AND a bigger `db` memory
|
||||
# limit — an OOM in Postgres doesn't degrade one feature, it takes the whole event down.
|
||||
DATABASE_MAX_CONNECTIONS=15
|
||||
|
||||
# Log level. `info` is the right production default: at `debug` the tower-http trace
|
||||
# layer writes a line per request AND per response, which on a busy event is a large
|
||||
@@ -124,12 +131,23 @@ EXPORT_PATH=/exports
|
||||
# display resize: ~145 MB at 12 MP, ~223 MB at 24 MP, ~354 MB at 48 MP.
|
||||
#
|
||||
# So on a 2 vCPU / 4 GB box (e.g. Hetzner CX22) KEEP THIS AT 2:
|
||||
# * concurrency 2, two 48 MP photos ≈ 800 MB against the 1G app limit — ~25% margin.
|
||||
# * concurrency 4, the same pair ≈ 1.5 GB — OOM.
|
||||
# * concurrency 4 would put two giants at ~1.5 GB against the 1G app limit — OOM.
|
||||
# * and app=2G + db=1G + frontend/caddy 256M each + ~370 MB of OS/Docker exceeds the
|
||||
# ~3910 MiB a "4 GB" VM actually reports. Raising the limit oversubscribes the host.
|
||||
# 4 is only reasonable on the 4 vCPU / 8 GB box README.md documents.
|
||||
#
|
||||
# The "two 48 MP photos at once" worst case this number used to be sized against is no
|
||||
# longer reachable: compression.rs takes an EXCLUSIVE `heavy` permit for any job whose
|
||||
# estimated peak exceeds HEAVY_IMAGE_BYTES (150 MiB), so two giants serialise no matter what
|
||||
# this is set to. What concurrency 2 now buys is two ORDINARY phone photos in parallel
|
||||
# (~145 MB peak each), which is both memory-safe and short enough not to starve the two
|
||||
# tokio worker threads a 2 vCPU box gets.
|
||||
#
|
||||
# Do NOT drop this to 1 hoping to protect the CPU. It halves throughput on the common light
|
||||
# path for a heavy path that is already serialised, and a longer compression backlog means
|
||||
# more feed tiles served from full-size originals (VirtualFeed falls back to /original while
|
||||
# derivatives are pending) — trading a little CPU for a lot of venue-wifi bandwidth.
|
||||
#
|
||||
# Throughput at 2 is not the bottleneck anyone thinks it is: ~2.5s per 12 MP photo, so
|
||||
# 100 photos is ~250 CPU-seconds spread over an entire evening.
|
||||
COMPRESSION_WORKER_CONCURRENCY=2
|
||||
|
||||
Reference in New Issue
Block a user