The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit, the committed comments and README's sizing section all assumed. That correction is what the first change is about; the rest are the remaining pre-event items. THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING `required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free. Every upload in that band succeeded and the keepsake could then never be built: the product's entire promise, failing silently at the end of the night with nobody there. The gate now enforces the invariant that actually matters — never accept an upload that would make the keepsake unbuildable — sharing `required_free_bytes` with the preflight so the two cannot drift into disagreeing about the same question. Uploads stop at ~8 GB of media on this disk, with a German message naming the cause. Refusing the 1001st photo beats losing all 1000. `media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s, rather than `estimate_export_bytes`'s join across every upload. It counts hidden and banned users' bytes, which the export excludes — skew in the SAFE direction, so the gate closes marginally early rather than late. Fails open on a query error. A test pins the gate against the preflight across the whole gallery-size range, and a second asserts the per-user floor alone would over-commit the volume — i.e. that the global gate is what must bind. THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED `Date.now()` advances while a backgrounded phone is frozen but `setInterval` does not, so the first tick after a screen lock read the whole sleep as silence and aborted — re-sending a video from byte zero and burning one of five PERMANENT auto-attempts. The interval is now its own suspension detector: a tick that arrives 125s late for a 5s schedule credits that window back, because a period the watchdog could not observe is not evidence of silence. Chosen over a `visibilitychange` listener, which only covers causes that fire that event — a throttled-but-visible tab, a closed lid and an occluded window all freeze timers without one — and which would have needed module state, an SSR guard and a teardown for strictly less coverage. `performance.now()` was rejected because Safari pauses it across system sleep on some paths and Chrome does not. The credit buys one fresh window, not immunity: a socket iOS reaped while backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout` (5-60 min) with the queue's `processing` latch held. Two latent leaks found while in there: `xhr.abort()` on a request already in readyState DONE emits no `abort` event, so `settle()` never ran and the interval re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button silently stopped working); and a synchronous throw from `xhr.send` — a blob whose backing store the OS purged — leaked the same way. Both closed. OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER red/amber/green were never in the @theme block and fell through to Tailwind v4's `oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color` falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already in this pipeline — so modern browsers render exactly what they render today. Verified against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal (avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0. Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on the host dashboard was rendering default violet on EVERY browser, off-brand. The keepsake viewer only picks this up on a rebuild, so its committed artefact is rebuilt here too — still single-file, still zero external references. A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a redeploy or a dead uplink left the guest on the boot spinner with no message, no reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a timeout rather than feature detection: a SyntaxError in the bundle is invisible to any capability check. Plus a <noscript>, since there was nothing at all to see without JS. EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info. If guests spend the evening hitting 429s or 413s, the post-event logs said nothing. Now one WARN per client error; 5xx excluded because Internal already logs its source chain and the pool-exhaustion 503 logs at construction. A DEAD FRONTEND SERVED A BLANK 502 `handle_errors 5xx` with an inline German page (the caddy service mounts only the Caddyfile, so there is no volume to ship a static file through). Verified empirically against this config, not from documentation: an upstream 404 through `reverse_proxy` still arrives as untouched `application/json`, and only a dial failure renders the page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a real 404/429 arriving, and swallowing those would have been worse than the blank 502. CONFIG CORRECTIONS FOR THE REAL HARDWARE DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count. Since migration 024 a feed page costs well under a millisecond, so connections are no longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host. COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was sized against is unreachable; dropping to 1 would halve light-path throughput and push more feed tiles onto full-size originals. README's sizing section rewritten for the actual disk. Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest, svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate, compose YAML parse. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
134 lines
7.0 KiB
Caddyfile
134 lines
7.0 KiB
Caddyfile
{$DOMAIN} {
|
|
# Compress everything EXCEPT the SSE stream — gzip buffering delays
|
|
# "real-time" likes/comments until the ~30s keep-alive tick.
|
|
@compressible not path /api/v1/stream
|
|
encode @compressible zstd gzip
|
|
|
|
# Site-wide security headers (defense-in-depth). HSTS is free since Caddy
|
|
# already terminates TLS. nosniff also covers all of /media/*.
|
|
header {
|
|
Strict-Transport-Security "max-age=31536000; includeSubDomains"
|
|
X-Content-Type-Options "nosniff"
|
|
Referrer-Policy "strict-origin-when-cross-origin"
|
|
}
|
|
|
|
# X-Frame-Options: DENY everywhere EXCEPT the keepsake download endpoints, which
|
|
# are navigated in a HIDDEN, SAME-ORIGIN iframe so a 404/429 can't unload the PWA
|
|
# (see frontend/src/routes/export/+page.svelte). WebKit enforces XFO *before*
|
|
# honouring Content-Disposition, so a blanket DENY makes the download silently do
|
|
# nothing on iOS Safari — the app's primary platform. SAMEORIGIN still blocks
|
|
# cross-origin framing.
|
|
#
|
|
# Split into two disjoint matchers rather than an override: Caddy applies the
|
|
# FIRST header directive outermost, so it wins on write — a later, more specific
|
|
# `header` would be silently ignored.
|
|
@framable path /api/v1/export/zip /api/v1/export/html
|
|
@not_framable not path /api/v1/export/zip /api/v1/export/html
|
|
header @framable X-Frame-Options "SAMEORIGIN"
|
|
header @not_framable X-Frame-Options "DENY"
|
|
|
|
# SvelteKit frontend — static assets with long-lived cache (content-hashed filenames)
|
|
@hashed_assets path_regexp hashed /_app/immutable/.*\.[a-f0-9]{8,}\.(js|css|woff2)$
|
|
header @hashed_assets Cache-Control "public, max-age=31536000, immutable"
|
|
|
|
# Preview/thumbnail/display images. These are served by the app through a
|
|
# visibility-checked alias (/api/v1/upload/{id}/{preview,thumbnail,display}) so
|
|
# moderation can revoke access; the app serves no /media route at all, so there is no
|
|
# direct path to the bytes. Privately cacheable for a short window (the app sets the
|
|
# same header; this is the edge carve-out from the blanket no-store below). Kept short
|
|
# so a moderated image stops being served to a direct-URL holder promptly.
|
|
#
|
|
# `display` was missing here while the backend set `private, max-age=300` on it, and
|
|
# because `header` REPLACES, the blanket no-store below silently won. That route is the
|
|
# ~2048px derivative the diashow uses exclusively, so a projector left running all
|
|
# evening re-fetched a full-size JPEG for every slide — roughly 2-4 GB pulled through
|
|
# the app over 8 hours, on the same venue uplink 100 guests are uploading over, and a
|
|
# blank frame on every network hiccup.
|
|
@media_api path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail /api/v1/upload/*/display
|
|
header @media_api Cache-Control "private, max-age=300"
|
|
|
|
# API and health — never cache, EXCEPT the gated image routes above. A cached health
|
|
# response would report the last known state rather than the current one.
|
|
@api {
|
|
path /api/* /health
|
|
not path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail /api/v1/upload/*/display
|
|
}
|
|
header @api Cache-Control "no-store"
|
|
|
|
# Route API and media requests to the Rust backend.
|
|
#
|
|
# The app serves no /media route at all (see the note in backend/src/main.rs) — media
|
|
# bytes are reachable only through the visibility-checked /api/v1/upload aliases, so
|
|
# /media/* forwards to a plain 404. The proxy line is kept deliberately: it means the
|
|
# edge faithfully hands /media to the app, so if a future change ever re-introduces a
|
|
# static media route the e2e gating specs see it here exactly as production would,
|
|
# instead of being masked by the SvelteKit 404 page.
|
|
reverse_proxy /api/* app:3000
|
|
reverse_proxy /media/* app:3000
|
|
|
|
# The backend registers /health on its ROOT router, not under /api/v1, so it needs its
|
|
# own line — without it the catch-all below hands /health to SvelteKit, which has no
|
|
# such route and returns its 404 page. That made the documented post-deploy check
|
|
# (`curl -fsS https://DOMAIN/health`) fail 100% of the time on a perfectly healthy
|
|
# stack. e2e/Caddyfile.test has always carried this line; production never did.
|
|
reverse_proxy /health app:3000
|
|
|
|
# Everything else goes to SvelteKit frontend
|
|
reverse_proxy frontend:3001
|
|
|
|
# Last-resort page for when Caddy itself cannot reach an upstream — the app or frontend
|
|
# container down, restarting, or still warming up after a host reboot. Without it a guest
|
|
# gets Caddy's bodiless 502: a completely blank page, which reads as "the whole thing is
|
|
# gone" rather than "try again in a moment".
|
|
#
|
|
# THIS DOES NOT TOUCH APPLICATION ERRORS. `handle_errors` fires only on errors CADDY
|
|
# generates; a status the app returns through `reverse_proxy` is written back verbatim and
|
|
# never reaches here. That distinction is load-bearing rather than incidental: the keepsake
|
|
# download navigates a HIDDEN IFRAME and depends on a real 404/429 arriving from the app
|
|
# (frontend/src/routes/export/+page.svelte), and every API route answers 403/404/429 as
|
|
# ordinary JSON that the client parses. Swallowing those into an HTML page would be a far
|
|
# worse regression than the blank 502 this fixes. Verified against this exact config: an
|
|
# upstream 404 through `reverse_proxy` still arrives as `Content-Type: application/json`
|
|
# with its body intact, while only a dial failure renders the page below.
|
|
#
|
|
# Scoped to 5xx so a hypothetical future Caddy-generated 4xx (there is none today) still
|
|
# returns plainly instead of claiming the server is restarting.
|
|
#
|
|
# The body is inline because the caddy service mounts ONLY ./Caddyfile and caddy_data —
|
|
# there is no volume to ship an HTML file through and the image has no build step, so a
|
|
# static file would mean changing the deployed stack's compose definition. No external
|
|
# font, stylesheet or image is referenced: the app may be exactly what is down.
|
|
#
|
|
# `handle_errors` has NO position in the directive order — Caddy hoists it into a separate
|
|
# `errors` route list — so it cannot disturb the "first `header` directive wins" hazard
|
|
# documented at the top of this file. The site-wide security headers still apply to it.
|
|
handle_errors 5xx {
|
|
header Content-Type "text/html; charset=utf-8"
|
|
header Cache-Control "no-store"
|
|
# {err.status_code} preserves the real status. Hardcoding 503 would mislabel a genuine
|
|
# 502 for anything watching from outside.
|
|
respond `<!doctype html>
|
|
<html lang="de">
|
|
<head>
|
|
<meta charset="utf-8">
|
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
|
<title>Gleich zurück</title>
|
|
<style>
|
|
html{background:#faf9f7;color:#1a1918;font-family:system-ui,-apple-system,"Segoe UI",Roboto,sans-serif}
|
|
body{margin:0;min-height:100vh;display:flex;align-items:center;justify-content:center;padding:2rem;text-align:center}
|
|
h1{font-family:Georgia,"Times New Roman",serif;font-weight:600;font-size:1.5rem;margin:0 0 .75rem}
|
|
p{margin:0;color:#545350;line-height:1.5}
|
|
@media (prefers-color-scheme:dark){html{background:#100f0f;color:#f5f4f2}p{color:#a6a4a1}}
|
|
</style>
|
|
</head>
|
|
<body>
|
|
<main>
|
|
<h1>Wir sind gleich zurück</h1>
|
|
<p>Die Seite wird gerade neu gestartet.<br>Bitte lade in einem Moment neu — deine Fotos bleiben gespeichert.</p>
|
|
</main>
|
|
</body>
|
|
</html>
|
|
` {err.status_code}
|
|
}
|
|
}
|