Files
EventSnap/Caddyfile
fabi f403222200 fix(deploy): a permanent upload outage, a dead-on-arrival Caddy, and 11pm commands that don't run
* Caddy had no read_body, on the reasoning that "a slow body still has to
  actually send bytes". That is an argument about disk, and disk is not the
  scarce resource: upload_admission budgets concurrent bodies at 4096 MiB and
  reserves the DECLARED cap, so a video/* upload reserves 500 MiB. Eight
  connections that stall mid-body hold the whole budget, every other guest
  waits 20s and gets a 503, and it never recovers on its own — the permit is
  held until the handler returns. No attacker needed: eight guests starting
  real videos and walking out of AP range does it, and TCP will not reap
  those sockets for hours. 30m carries a 500 MB upload at ~2.2 Mbit/s, so it
  does not fail the uploads this product exists to collect.

* APP_PORT is presented in .env.example as an ordinary editable line, while
  the healthcheck hardcodes 127.0.0.1:3000 and the Caddyfile hardcodes
  app:3000. Change it and the app boots and serves happily on the new port,
  the healthcheck fails forever, app never turns healthy — and because caddy
  is gated on service_healthy, CADDY NEVER STARTS. Port 443 dead for the
  whole event, sole diagnostic "dependency failed to start". Pinned in
  compose beside MEDIA_PATH and EXPORT_PATH, which are there for this reason.

* Runbook §12's recovery commands do not run as written: unwrapped
  "$POSTGRES_USER" is expanded by the operator's shell, which does not have
  it, so psql answers `FATAL: role "" does not exist`. §9 documents that trap
  two hundred lines earlier and wraps its own calls in sh -c; §12 did not.
  This is the block you run with the app crash-looping behind a live Caddy.
  Its DELETE also hard-coded versions 21,22,23 as if to be copied verbatim,
  on a tree that now has 31 migrations — now explicitly an example, with the
  instruction to take the numbers from the actual boot error.

* Migration counts corrected across the runbook and .env.example (22 -> 31,
  commit count 154 -> 196). All four were presented as literal command output
  the operator is invited to reproduce.
2026-08-12 09:15:55 +02:00

164 lines
8.5 KiB
Caddyfile

{
servers {
timeouts {
# Slowloris defence, at the layer that can actually apply it.
#
# There was no read or write timeout anywhere, so a client could open a POST, send one
# byte a minute, and hold a connection, a tokio task and a `.tmp` file indefinitely —
# and the upload sweeper is keyed on mtime precisely so a live upload never ages out,
# so ten such connections consumed disk the upload gate could not see.
#
# read_header is tight: a legitimate client sends its headers in one go.
read_header 10s
# read_body is GENEROUS but present. It was omitted on the reasoning that "a slow body
# still has to actually send bytes" — which is an argument about disk, and disk is not
# the scarce resource here. `upload_admission` budgets concurrent bodies at 4096 MiB and
# reserves the DECLARED cap, so a `video/*` upload reserves 500 MiB: eight connections
# that stall mid-body hold the entire budget, every other guest waits 20s and gets a
# 503, and it never recovers on its own because the permit is held until the handler
# returns. That needs no attacker — eight guests starting real videos and then walking
# out of AP range does it, and TCP will not reap those sockets for hours.
#
# 30m carries a 500 MB video at ~2.2 Mbit/s sustained, which is well under venue wifi
# and under most cellular, so it does not fail the uploads this product exists to
# collect. It does bound the leak to something that drains.
read_body 30m
idle 5m
}
}
}
{$DOMAIN} {
# Compress everything EXCEPT the SSE stream — gzip buffering delays
# "real-time" likes/comments until the ~30s keep-alive tick.
@compressible not path /api/v1/stream
encode @compressible zstd gzip
# Site-wide security headers (defense-in-depth). HSTS is free since Caddy
# already terminates TLS. nosniff also covers all of /media/*.
header {
Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Content-Type-Options "nosniff"
Referrer-Policy "strict-origin-when-cross-origin"
}
# X-Frame-Options: DENY everywhere EXCEPT the keepsake download endpoints, which
# are navigated in a HIDDEN, SAME-ORIGIN iframe so a 404/429 can't unload the PWA
# (see frontend/src/routes/export/+page.svelte). WebKit enforces XFO *before*
# honouring Content-Disposition, so a blanket DENY makes the download silently do
# nothing on iOS Safari — the app's primary platform. SAMEORIGIN still blocks
# cross-origin framing.
#
# Split into two disjoint matchers rather than an override: Caddy applies the
# FIRST header directive outermost, so it wins on write — a later, more specific
# `header` would be silently ignored.
@framable path /api/v1/export/zip /api/v1/export/html
@not_framable not path /api/v1/export/zip /api/v1/export/html
header @framable X-Frame-Options "SAMEORIGIN"
header @not_framable X-Frame-Options "DENY"
# SvelteKit frontend — static assets with long-lived cache (content-hashed filenames)
@hashed_assets path_regexp hashed /_app/immutable/.*\.[a-f0-9]{8,}\.(js|css|woff2)$
header @hashed_assets Cache-Control "public, max-age=31536000, immutable"
# Preview/thumbnail/display images. These are served by the app through a
# visibility-checked alias (/api/v1/upload/{id}/{preview,thumbnail,display}) so
# moderation can revoke access; the app serves no /media route at all, so there is no
# direct path to the bytes. Privately cacheable for a short window (the app sets the
# same header; this is the edge carve-out from the blanket no-store below). Kept short
# so a moderated image stops being served to a direct-URL holder promptly.
#
# `display` was missing here while the backend set `private, max-age=300` on it, and
# because `header` REPLACES, the blanket no-store below silently won. That route is the
# ~2048px derivative the diashow uses exclusively, so a projector left running all
# evening re-fetched a full-size JPEG for every slide — roughly 2-4 GB pulled through
# the app over 8 hours, on the same venue uplink 100 guests are uploading over, and a
# blank frame on every network hiccup.
@media_api path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail /api/v1/upload/*/display
header @media_api Cache-Control "private, max-age=300"
# API and health — never cache, EXCEPT the gated image routes above. A cached health
# response would report the last known state rather than the current one.
@api {
path /api/* /health
not path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail /api/v1/upload/*/display
}
header @api Cache-Control "no-store"
# Route API and media requests to the Rust backend.
#
# The app serves no /media route at all (see the note in backend/src/main.rs) — media
# bytes are reachable only through the visibility-checked /api/v1/upload aliases, so
# /media/* forwards to a plain 404. The proxy line is kept deliberately: it means the
# edge faithfully hands /media to the app, so if a future change ever re-introduces a
# static media route the e2e gating specs see it here exactly as production would,
# instead of being masked by the SvelteKit 404 page.
reverse_proxy /api/* app:3000
reverse_proxy /media/* app:3000
# The backend registers /health on its ROOT router, not under /api/v1, so it needs its
# own line — without it the catch-all below hands /health to SvelteKit, which has no
# such route and returns its 404 page. That made the documented post-deploy check
# (`curl -fsS https://DOMAIN/health`) fail 100% of the time on a perfectly healthy
# stack. e2e/Caddyfile.test has always carried this line; production never did.
reverse_proxy /health app:3000
# Everything else goes to SvelteKit frontend
reverse_proxy frontend:3001
# Last-resort page for when Caddy itself cannot reach an upstream — the app or frontend
# container down, restarting, or still warming up after a host reboot. Without it a guest
# gets Caddy's bodiless 502: a completely blank page, which reads as "the whole thing is
# gone" rather than "try again in a moment".
#
# THIS DOES NOT TOUCH APPLICATION ERRORS. `handle_errors` fires only on errors CADDY
# generates; a status the app returns through `reverse_proxy` is written back verbatim and
# never reaches here. That distinction is load-bearing rather than incidental: the keepsake
# download navigates a HIDDEN IFRAME and depends on a real 404/429 arriving from the app
# (frontend/src/routes/export/+page.svelte), and every API route answers 403/404/429 as
# ordinary JSON that the client parses. Swallowing those into an HTML page would be a far
# worse regression than the blank 502 this fixes. Verified against this exact config: an
# upstream 404 through `reverse_proxy` still arrives as `Content-Type: application/json`
# with its body intact, while only a dial failure renders the page below.
#
# Scoped to 5xx so a hypothetical future Caddy-generated 4xx (there is none today) still
# returns plainly instead of claiming the server is restarting.
#
# The body is inline because the caddy service mounts ONLY ./Caddyfile and caddy_data —
# there is no volume to ship an HTML file through and the image has no build step, so a
# static file would mean changing the deployed stack's compose definition. No external
# font, stylesheet or image is referenced: the app may be exactly what is down.
#
# `handle_errors` has NO position in the directive order — Caddy hoists it into a separate
# `errors` route list — so it cannot disturb the "first `header` directive wins" hazard
# documented at the top of this file. The site-wide security headers still apply to it.
handle_errors 5xx {
header Content-Type "text/html; charset=utf-8"
header Cache-Control "no-store"
# {err.status_code} preserves the real status. Hardcoding 503 would mislabel a genuine
# 502 for anything watching from outside.
respond `<!doctype html>
<html lang="de">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Gleich zurück</title>
<style>
html{background:#faf9f7;color:#1a1918;font-family:system-ui,-apple-system,"Segoe UI",Roboto,sans-serif}
body{margin:0;min-height:100vh;display:flex;align-items:center;justify-content:center;padding:2rem;text-align:center}
h1{font-family:Georgia,"Times New Roman",serif;font-weight:600;font-size:1.5rem;margin:0 0 .75rem}
p{margin:0;color:#545350;line-height:1.5}
@media (prefers-color-scheme:dark){html{background:#100f0f;color:#f5f4f2}p{color:#a6a4a1}}
</style>
</head>
<body>
<main>
<h1>Wir sind gleich zurück</h1>
<p>Die Seite wird gerade neu gestartet.<br>Bitte lade in einem Moment neu deine Fotos bleiben gespeichert.</p>
</main>
</body>
</html>
` {err.status_code}
}
}