Compare commits

...

103 Commits

Author SHA1 Message Date
fabi
9759c7c669 fix(ops): the hourly backup never ran, and an unset POSTGRES_USER crash-loops silently
**The backup that did not exist.**
`.env.example` ships `EVENT_NAME=Max & Maria's Wedding` and the runbook tells you
to `cp .env.example .env`. Compose's env_file parser reads that fine. POSIX `sh`
does not: `. ./.env` aborts with "Unterminated quoted string" (verified, rc=2),
and every variable defined after that line is left unset.

The §10.2 cron script is `#!/bin/sh` + `set -eu` + `. ./.env`, so it exited
before `pg_dump` — every hour, into a log nobody reads. The only automated backup
of the one thing the runbook calls irreconstructible produced nothing, and §10.2's
own "prove it works NOW" only catches it if `.env` is already final at that
moment.

The script reads NOTHING from `.env` — `POSTGRES_USER`/`POSTGRES_DB` are expanded
inside the db container by the single-quoted `sh -c`. The source line was pure
liability and is gone. `EVENT_NAME` is now double-quoted in `.env.example`, which
both parsers read identically (verified), and the three interactive sourcing
sites now read just `$DOMAIN` instead of sourcing the whole file. The verify step
also proves the dump is a non-empty valid gzip containing tables, rather than
that a file exists.

Also fixes the script's `cd /root/eventsnap`, which contradicts §5's non-root
deploy and §13's `~/eventsnap` — under a non-root deploy it failed the same way,
silently.

**The crash loop with no message.**
`docker-compose.yml` interpolated `POSTGRES_USER`/`POSTGRES_DB` with no default
and no `:?` guard, into `environment:`, which OVERRIDES `env_file`. Unset does
not fall back — it resolves to the empty string, initdb creates a role and
database named "", `DATABASE_URL` still says `eventsnap`, and the app hits
`FATAL: role "eventsnap" does not exist` forever. `pg_isready -U "" -d ""` never
passes, so `app` never turns healthy and Caddy — gated on `service_healthy` —
never starts: port 443 dead for the whole event, exit only via `down -v`.

Both now carry `:?` guards (verified they fire), and §3's ".env template — ALL of
them" list, which omitted both, now includes them.

Other runbook corrections: the backup/restore pointer named a line range that had
drifted into an unrelated section and stopped mid-restore, before the media
restore and the mandatory `chown` — now referenced by heading, which cannot go
stale. §7.3 told you to verify that `EXPORT_PATH` is not pinned when §3 correctly
says it is. Stale counts: rev-list 196 -> 217, "versions 007–022" -> 007–031,
`frontend/Dockerfile:9` -> :8, and the low-disk description now matches the code
(the 10 GB absolute floor was removed as unreachable).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:46:21 +02:00
fabi
cfc8bd0016 test: replace coverage that could not fail with coverage that can
`backend/tests/` follows a house rule of copying production SQL character-for-
character rather than calling `src/`, because the crate is a binary and nothing
in it is importable from an integration test. For pinning behaviour that already
existed that is a defensible trade. Applied to a NEW fix whose only coverage is
the copy, it proves nothing: the fix and its test become two independent
implementations, and deleting the fix leaves the test green.

`audit_names.rs` did exactly that. It never called `audit::record` — it
reimplemented `resolve_names` and the INSERT inside the test file, down to a
hardcoded `.bind("host")`, and then asserted `actor_role == "host"` against its
own literal. That assertion could not fail for any change to the code it named,
and grep confirmed there was no other coverage of the audit-name work anywhere.

Moved into `#[cfg(test)]` inside `services/audit.rs`, where the real function IS
callable. CI already runs `cargo test --all-features` with a live DATABASE_URL,
so `#[sqlx::test]` works there; verified all four run and pass. The role
assertion now compares against `UserRole::as_str()` itself rather than a literal,
so it tracks a rename instead of pretending to, plus an explicit `assert_ne!`
against the Debug spelling.

Also:

- `retry-after-release.spec.ts` filtered the feed on `u.id === original.id` to
  prove "no second row was created". A duplicate gets a fresh uuid and could
  never match, so the filter yielded exactly 1 whether the gallery held one copy
  or five. Counts by uploader now, with the original's identity asserted
  separately. (The rest of that spec is sound — its 403 control and replay-id
  check both fail if the header fast-path is reverted.)

- `upload_after_release_commits_sees_the_lock_and_is_rejected` claimed the
  handler answers `UploadsLocked`. It answers `GalleryReleased` since the check
  order was inverted on this branch, and the test asserts no variant at all.
  Documented what it actually covers (the locked READ) and where the ordering IS
  covered (two e2e specs).

- Two `// SRC:` pointers had drifted ~130 lines into unrelated code, which is how
  a hand-copied fixture silently stops matching its original. Now named, not
  numbered.

- `emptyOutDir: false` claimed a failed viewer build "leaves the last good
  artifact in place". True for the `generateBundle` error, false for the newer
  `writeBundle` assertion, which fires after Vite has already written the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:46:04 +02:00
fabi
ee70eec094 fix(frontend): three dead ends a guest cannot get out of
**1. The cached PIN could never be cleared after a host reset.**
`/recover` clears a rejected cached PIN only when the submitted name is the one
this device belongs to — narrowed on this branch so a guest who mistypes their
own name does not lose the only copy of their PIN (the server keeps just the
bcrypt). But it compared against `DISPLAY_NAME_KEY`, which `clearAuth` deletes
for shared-device privacy — one step BEFORE the guest ever reaches that screen:

  host taps "PIN zurücksetzen" -> the backend also revokes every session for that
  user -> the guest's next request 401s -> clearAuth -> redirect to /join -> they
  go to /recover, where the field is pre-filled with the dead PIN and the guard
  can never fire again

Since a 4-digit value auto-submits, every correction burns another of the four
wrong-PIN attempts the shared venue IP allows per 15 minutes. The PIN's owner is
now stored WITH the PIN and survives alongside it, with a fallback to the auth
display name for devices that cached a PIN before this key existed.

**2. Every layout-level SSE handler waited on the `/me/context` retry.**
The retry was awaited inside the same `onMount` that registers `pin-reset`,
`user-hidden`/`user-shown`, `event-closed`/`event-opened` and `event-updated`.
Worst case is a 20s timeout + 2s backoff + a second 20s timeout: ~42s with an
empty handler list, on exactly the wifi the retry exists for. Five of the six
self-heal; `pin-reset` does not, and a missed one leaves a dead PIN displayed in
"Mein Konto" and pre-filling /recover — the same state as (1), reached from the
other end. Detached, since nothing below reads its result.

**3. `crypto.randomUUID` was on the join critical path.**
It needs Safari >= 15.4 / Chrome >= 92 AND a secure context. The queue already
depended on it, so an old phone previously joined and browsed and only failed at
upload — degraded but survivable. Minting an idempotency key at join turned that
into a `TypeError` caught by the generic handler and rendered as "Ein Fehler ist
aufgetreten." on every retry: cannot join, cannot browse, and /recover is no help
because there is no account yet. The one screen where a hard failure has no way
out at all. Falls back to `crypto.getRandomValues` with the RFC 4122 version and
variant bits set; `Math.random` is deliberately NOT a further fallback, since a
collision between two guests would replay one guest's join or upload onto
another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:45:46 +02:00
fabi
9bae5d77ed fix(join): concurrent first joins no longer 500 on the QR-scan burst
`Event::find_or_create` was check-then-insert against a UNIQUE slug, and its only
callers are `/join` and `/admin/login` — both of which run before the row exists,
at the single most concurrent moment the app ever sees: the QR code goes up and
every phone in the room posts `/join` within the same second. All of them miss
the SELECT, all of them INSERT, one wins, and the rest get a bare unique
violation surfaced as a 500 on the very first screen of the event. There is no
retry on that path and nothing in the UI explains it.

In the documented timeline the host's T-5 admin login creates the row first, so
the blast radius is small — but it is one `down -v` or one `EVENT_SLUG` edit away
from being live on the night.

`ON CONFLICT (slug) DO UPDATE SET slug = EXCLUDED.slug` — a deliberate no-op
write, because `DO NOTHING` returns no row on conflict and would put the loser
back at square one. It touches only `slug`, so `name`, `export_epoch` and the
lock/release timestamps are never disturbed by a late arrival; a test pins that.
The read fast-path stays, so every join after the first is still a plain SELECT
and takes no row lock.

Tests live in `src/` rather than `tests/` because the crate is a binary and the
function is not importable from an integration test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:45:32 +02:00
fabi
20c15c3500 fix(backend): three ways the end of the night could go wrong
**1. Releasing the gallery could arm the keepsake with no worker.**
`release_gallery` ran `tx.commit()` -> SSE `event-closed` -> `audit::record().await`
-> `spawn_export_jobs`. The audit write is two pool round-trips, each able to wait
the full 5s acquire timeout, and it runs in the same instant `event-closed` fans
out to ~100 phones whose upload queues all hit the API at once. Axum drops the
handler future when the client disconnects — the host taps "Freigeben" and
pockets the phone. The release has COMMITTED: event closed, uploads locked, epoch
bumped, both `export_job` rows pending, and no worker. `/export/*` 404s, the page
sits on "Wird vorbereitet…", `recover_exports` only runs at boot, and a second
release is refused. Every other regen call site spawns first; `me.rs` says so in
a comment. This was the sole violator, and the only path that arms the FIRST
build of the keepsake. Spawn moved immediately after the commit.

**2. The event could be left with no operator.**
`remaining_operators` was an unlocked pool COUNT followed by a separate UPDATE,
so `ban_user` and `set_role` raced each other and `DELETE /me`: an admin demotes
host B while host A deletes themselves, each check sees the other still present,
both commit, and nobody can moderate, release the gallery, or appoint anyone —
appointing requires being an operator. The count now runs inside the writing
transaction behind the same advisory lock `delete_account` uses, via one shared
helper so the key cannot drift between copies.

The lock is taken FIRST in all three, and the order is load-bearing:
`delete_account` previously took it last, after row locks on `upload` and
`event`, while the two new call sites take it before locking those same rows —
an ABBA that Postgres would resolve by killing one transaction with a 500. The
ordering rule is documented on the helper.

**3. The keepsake could become unbuildable the moment uploads stopped.**
The upload gate and the export preflight computed the IDENTICAL threshold
(`required_free_bytes(media, 2) + DISK_RESERVE_BYTES`), leaving zero margin
between them. Once the gate refused its first upload the preflight was already at
its own limit, so anything written afterwards decided the keepsake's fate: WAL up
to `max_wal_size`, 30 MB x 4 of container logs, and the compression backlog
draining at exactly that hour. The release commits before the workers bail, so
the failure lands at 01:00 with no second release possible. The gate now demands
`UPLOAD_GATE_HEADROOM_BYTES` more than the preflight, costing ~0.5 GB of media
ceiling — the trade README already argues for. The dashboard banner mirrors the
new threshold so its lead is unchanged, and a new test pins gate-before-preflight
at six gallery sizes.

Also: the global disk gate fails OPEN when the mount cannot be read, which is
deliberate, but did it SILENTLY — no log line at all, while the export preflight
warns on the identical condition. Inside a container `/` is an overlay rather
than a `/dev` device, so this is reachable, and when it happens the only global
disk bound is gone and the box fills until Postgres cannot write WAL.

README's sizing table was also arithmetically self-contradictory (it showed
~27 GB free against a ~27.6 GB requirement); recomputed for the new gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:45:24 +02:00
fabi
06e0bea0e9 fix(export): a refused download mint no longer strands the running transfer
`/export/ticket` mints before charging the daily limit, deliberately: charging
first meant a store-capacity 503 — a server-side condition the guest cannot see
or cause — still cost one of their three daily downloads, with no refund path.

But a mint refused with 429 left its ticket in the store. Download tickets live
six hours (they must, so a 1.4 GB transfer can resume with `Range`), and the
per-session cap is four tickets OF THE SAME KIND. So:

  the transfer starts on ticket A -> the bar looks stuck on venue wifi -> the
  guest taps "Herunterladen" again -> mints 2 and 3 succeed, 4 and 5 return 429
  but still mint -> the fifth evicts the oldest download ticket for the session,
  which is A -> the transfer drops, resumes, and 401s -> re-minting is
  impossible, they are at the daily limit

The keepsake is unreachable until the next day, for tapping a button that
appeared to do nothing. This is the failure `sse_churn_cannot_evict_a_running_
download` was written to prevent, reintroduced through the one channel that test
does not cover: download tickets evicting each other.

Discarding the ticket on the refusal path keeps both properties that put the
mint first — a capacity 503 still costs no download, and a refused download now
costs no slot.

Also corrects three doc comments in this path still describing download tickets
as "single-use, 30s TTL", false since they were made resumable. Stale comments
here have already sent one review down the wrong path.

The new spec asserts the two 429s explicitly, so it cannot pass vacuously if the
daily limit stops being enforced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 19:45:08 +02:00
fabi
e6aeaa0a8b docs(compose): the app CPU cap cannot separate compression from the request path
The comment read "CPU ceiling for the two image workers + ffmpeg poster
extraction", which describes a separation Docker cannot make: `compression.rs`
runs that work in `tokio::task::spawn_blocking` — same process, same cgroup as
every Axum handler — and `cpus`/`cpu_shares` are per-container.

What actually happens is worth knowing when sizing this box: `cpu.max` is
`120000 100000`, so two CPU-pegged blocking workers exhaust the 120 ms quota
after ~60 ms of each 100 ms period and the kernel freezes the WHOLE cgroup —
uploads, feed and SSE included — for the rest of it. Over a 100-photo burst
that is ~210 s during which every request can eat up to 40 ms of throttle.

The value stays at 1.2: an app that can take both cores starves Postgres, and
every request path goes through Postgres. A slightly stalled request beats a
starved database. The comment now says which knob actually shortens the
backlog (COMPRESSION_WORKER_CONCURRENCY) rather than implying this one does.
2026-08-12 23:11:57 +02:00
fabi
ac04e27e34 fix(upload): a retry after release returns the stored photo instead of refusing it
The idempotency key was only readable as a multipart FIELD, and a field cannot
be read until the body is being parsed — which happens after the lock/release
pre-flight. So the replay was unreachable in exactly the case it exists for:

  the photo commits → the response is lost on the way back (the flaky-wifi
  failure the key was added for) → the host releases the gallery at the end of
  the night → the phone's retry answers `gallery_released`.

The guest is told a photo that is sitting in the gallery was never sent. And
the remedy the client offers is destructive: `open_event` clears
`export_released_at` AND bumps `export_epoch`, retiring the whole keepsake
generation and forcing a multi-GB rebuild on a 2-vCPU box at midnight — to
re-send a photo that was never missing. Several guests on one flaky evening
make this likely to happen at least once.

The key is now also sent as `X-Client-Upload-Id`, which arrives with the
request line, so the answer is knowable before anything is decided about
locks. The multipart field stays for the concurrent case and as a fallback.

Placed ahead of the hourly rate limiter too, which was the same mistake one
layer up: a 40-photo burst with two retries apiece exhausted the guest's hour
on uploads that had all committed the first time.

The body is still drained rather than abandoned — replying before reading it
makes the proxy see a broken pipe and turn a clean 200 into a 502.

The spec carries its own control: a DIFFERENT photo is asserted to still be
refused with `gallery_released` after the release, so the replay cannot be
green merely because the gate was open.
2026-08-12 23:11:57 +02:00
fabi
137b892480 docs(runbook): the .env template no longer implies edits that compose overrides
The template said `EXPORT_PATH=/exports  # NOT pinned by compose`, which has
been false for as long as the pin has existed — and it sits three lines under
MEDIA_PATH, which is correctly described as pinned, so the contrast reads as
deliberate. An operator moving exports to a separate volume (the remedy §13
now recommends for a full disk) would edit `.env`, see nothing change, and
have no reason to suspect the compose file.

All four pinned vars are now listed together with what the pin means: change
it there, not here. DATABASE_MAX_CONNECTIONS carries the extra note that its
value is boot-fatal when unparseable, which is why it is pinned at all.
2026-08-12 22:23:33 +02:00
fabi
b5a1580368 test(e2e): restore the download-side 404 coverage the mint pre-check displaced
Both tests in export.spec.ts are still named "ZIP download 404s…" but now
assert only that `/export/ticket` refuses. That move was right — the mint
pre-validates, and refusing there spends none of the guest's three daily
downloads — but it left `resolve_export_file` unasserted on the download path
itself, so deleting that check would not have turned anything red.

It cannot be covered by minting against an already-dead archive, because the
mint refuses first. The order has to be: real release → mint while healthy →
retire the generation → download. Which is also precisely what happens to a
ticket already in flight when the host takes a photo down mid-download.

Carries its own positive control: the ticket is asserted to serve a 200 before
the epoch moves, so the 404 afterwards cannot be green for the wrong reason —
an expired or never-valid ticket would 404 too.
2026-08-12 22:23:33 +02:00
fabi
6475199670 fix(feed): stop raising the stale pill on a list view that is not filtered
The merge gate tested the raw chip state (`selectedHashtag || activeFilters
.length`), but `filterParams()` — which decides what the server actually
returns — ignores `activeFilters` entirely in list view.

`switchView('list')` deliberately keeps a chip on `activeFilters` (an uploader
chip, or a second tag) while setting `selectedHashtag` to the first TAG
filter, which is null when the only chip was an uploader. So a guest who
filtered the grid by an uploader and then switched to list view had a
genuinely unfiltered view whose every delta raised "Neue Beiträge" instead of
merging the rows in — the same pill-for-rows-that-should-have-merged this
block was written to stop, one branch further in.

Nothing was lost (the pill always clears) but it trains guests to ignore the
one control that means something. Now gated on the effective filter, which is
exact by construction: it asks the same function the fetch does.
2026-08-12 21:40:29 +02:00
fabi
55b57fc037 fix(me): two hosts deleting at once can no longer leave the event with no operator
The last-operator guard ran on the pool, before the transaction opened. Two
hosts deleting themselves at the same moment each saw the other, both passed,
and the event was left with nobody who can moderate, nobody who can release
the gallery, and no way to appoint anyone — because appointing requires a
host. Not recoverable from inside the app.

The fix is a transaction-scoped ADVISORY lock, and the two obvious
alternatives are both worse:

* `FOR UPDATE` on the other operators' rows DEADLOCKS. Each deleter locks the
  other's row and then tries to delete its own, so Postgres resolves it by
  killing one. The invariant survives; the loser gets a 500 instead of the
  sentence explaining what to do. My first attempt did exactly this, and the
  test caught it.
* Locking the `event` row serialises cleanly but inverts the lock order every
  moderation path takes (upload/user rows first, event last). That is an ABBA
  against a path that runs constantly during the event, traded for one that
  runs approximately never.

An advisory lock is a separate lock space, so it cannot interact with the
row-lock graph at all, and it is released when the transaction ends. The loser
waits, counts zero once the winner's row is gone, and is refused with the
sentence it should have got.

The test is a genuine concurrency test — it spawns the second deleter and
asserts it unblocks to see no remaining operator. It fails against the
pre-check-outside-the-transaction version and against the FOR UPDATE version.
2026-08-12 21:40:29 +02:00
fabi
19b59d6fee docs(upload): stop claiming a proxy bandwidth control that does not exist
`get_original`'s comment said bandwidth abuse "belongs at the proxy, where
per-connection limits still work", which reads as though the removed per-IP
limiter had been replaced by something. It was not: the Caddyfile sets
timeouts and no rate or concurrency directive, and the tower stack is
TraceLayer alone.

Removing the limiter was right — the venue is one NAT address, so that bucket
throttled the whole party's feed — but the route is now unbounded, and the
comment should say so rather than imply cover. Records the actual cost
(no-store plus the derivative fallback plus the nonce'd retry, against a
15-slot pool that upload commits compete for) and the shape a real fix would
take: a concurrency semaphore over media streaming, not a request-rate bucket.
2026-08-12 20:51:51 +02:00
fabi
010bcc0e3c fix(build): stop a stale or missing keepsake viewer from shipping silently
Three ways the compiled-in viewer could be wrong, none of which anything would
have reported. Found by mutation-testing the guard added below — it failed
when it should have passed, and the reason was the second bullet.

* `include_dir!` registers NO rebuild dependency. Run `npm run build` in
  frontend/export-viewer, then `cargo build`, and cargo sees no source change
  and reuses the cached binary — carrying the PREVIOUS index.html. The file on
  disk and the file in the binary disagree, git is clean, every check passes,
  and Memories.zip ships a stale viewer. Confirmed empirically: after replacing
  the artifact the compiled-in copy did not change until a source file was
  touched. A build.rs now declares `rerun-if-changed` for
  `static/export-viewer` AND `migrations` — sqlx::migrate!() embeds its
  directory the same way, and there the stale snapshot is worse still: the
  binary boots against a database that already ran a newer migration and
  crash-loops with VersionMissing.

* `emptyOutDir: true` deleted the committed artifact BEFORE generating. That
  was safe while the build could not fail; it no longer is, because
  `inlineThemeFonts` now calls `this.error` on a keepsake that is not
  self-contained. A failed build left the directory empty — and include_dir!
  over an empty directory compiles fine, while `write_viewer_with_data`
  iterates zero files and returns Ok. The result is a valid archive with every
  photo and no viewer. The output is one overwritten file, so nothing
  accumulates without the wipe.

* Nothing asserted the viewer was there at all. Now asserted at the point of
  use (bail rather than write a viewer-less keepsake) and in a test that checks
  presence, plausible size, and that no `url(/...)` survived inlining — the
  three ways it can be present but useless.

The Dockerfile copies build.rs with the sources rather than with Cargo.toml, so
the dependency-cache layer stays byte-identical and the dummy build does not
run it.
2026-08-12 20:51:51 +02:00
fabi
8af8c4fab7 docs(runbook): validate the Caddyfile before the freeze, and pair down-migrations with a rollback
Nothing anywhere executes the production `Caddyfile` before the real deploy —
the e2e stack mounts `e2e/Caddyfile.test` — and a syntax error there is total:
Caddy exits, `restart: unless-stopped` loops, 443 is dead for the whole event,
and `docker compose up -d --force-recreate caddy` still exits 0 while it
crash-loops. Step zero now validates it. I ran it against the current file
(which I changed last commit, unexercised): "Valid configuration", and the new
`read_body 30m` adapts to `read_timeout: 1800000000000`ns as intended.

And a warning §9 needed: a down migration is not a standalone repair. Roll the
IMAGE back first. `Upload::create` sends an `ON CONFLICT ... WHERE` predicate
that must match the live partial index exactly and is not compile-checked, so
running 026's or 031's down against the current binary turns every upload
carrying a client_upload_id — i.e. every upload from the shipped client — into
a runtime 500. 026's down can also fail outright on any database where a guest
deleted and re-uploaded a photo; it rolls back cleanly, but you cannot go below
it. Both verified against a live Postgres.
2026-08-12 20:00:52 +02:00
fabi
301e6636a5 fix(audit): give the audit trail the names that make it readable
Migration 029 made `actor_id`/`target_id` non-FK on the stated grounds that
"the record must survive the actor's account being removed, which is exactly
when it is most likely to be wanted". All eleven call sites then passed None
for both name columns — so what survived a deletion was a bare uuid resolving
to nothing: the guarantee, minus the only thing that made it useful.

`record` now resolves whatever the caller omitted, in one query, so no call
site can forget. `me::delete_account` passes its names explicitly because it
has already hard-deleted the row by then — that is the one record a host is
most likely to be reading the next morning ("whose photos disappeared?").

Also: `actor_role` is written with `as_str()` rather than
`format!("{actor_role:?}")`. The Debug spelling is not a stable wire format,
so a derive change or a renamed variant would have silently started writing a
different string into a column nothing validates.

Migration 029's header lists three action slugs (`promote_user`,
`demote_user`, `delete_user`) that no call site has ever emitted, and it
cannot be corrected — editing an applied migration changes its checksum and
crash-loops every database that ran it. The real list, verified against the
call sites, is documented in this module instead, along with the fact that
there is no read endpoint and the query to use by hand.

A NULL name fails silently, so it is now asserted: names resolved from ids,
names surviving the row's deletion, and a row still written when neither can
be resolved (an audit write must never fail the action it records).
2026-08-12 20:00:52 +02:00
fabi
4916eed436 fix(deploy): ship the swap ceilings, pin the last boot-fatal env var, and correct docs that misdirect
* memswap_limit is now IN docker-compose.yml on all four services. Compose
  sets Memory but leaves MemorySwap unset, and Docker then permits swap equal
  to the memory limit — so following §5's "add 2 GB of swap" silently DOUBLED
  every ceiling, to ~5 GiB on a 3.82 GiB box. Nothing OOMs; instead Postgres's
  working set becomes swap-eligible on a shared-tenancy SSD, turning a bounded
  OOM-kill that restarts in seconds into unbounded latency with no signal but
  "everything is slow". The runbook told the operator to hand-add it, which
  also broke §0's own gate that docker-compose.yml must be unmodified.
  Verified rather than assumed: service-level memswap_limit does compose with
  deploy.resources.limits.memory (docker inspect → Memory=1073741824
  MemorySwap=1207959552).

* DATABASE_MAX_CONNECTIONS pinned in compose. It is the one env var that is
  now boot-FATAL when unparseable — the right call, but it means a stray quote
  or a trailing inline comment in .env crash-loops the app behind a live
  Caddy. MEDIA_PATH, EXPORT_PATH and APP_PORT are pinned for weaker reasons.

* .env.example's quota narrative was sized for a CX33: "~30 GB of a fresh
  70 GB" on a box with 40 GB. And on THIS box the fixed point never binds at
  all — ~210 MB/guest is below the 500 MiB floor, so everyone gets the floor
  and the per-user quota stops bounding aggregate growth. What actually stops
  uploads is the keepsake preflight at ~8 GB of media. That paragraph is what
  an operator reads when a guest is blocked, and it pointed at the wrong knob.

* The emergency card gains the one disk symptom that can appear mid-event,
  where `df -h` — its only disk instruction — actively misleads: the gate
  fires ~10 GB + 2.2x media BEFORE the disk is full, so df shows ~20 GB free
  at the moment uploads are being refused.

* Two code comments that now assert the opposite of the code: claim_job
  promised that "the update_progress liveness check bails such a worker out
  early" — it cannot, its predicate is on the job row, which a reopen does not
  touch, so a mid-export reopen grinds the whole gallery to completion on a
  2-vCPU box during the live event. And prune_superseded_archives still argued
  "deleted bytes cannot be rolled back" as an invariant, after the reclaim
  path was changed to prune even when that will not close the shortfall.
  Both now describe what the code does.

* Smaller corrections: runbook §3's "two 48 MP photos ≈ 800 MB" scenario is
  unreachable (compression.rs takes an exclusive heavy permit, so they
  serialise) and contradicted .env.example; "all four healthy" is wrong since
  caddy has no healthcheck; a README line reference pointed at a comment added
  by the same commit that broke it.
2026-08-12 19:10:45 +02:00
fabi
182e712a0e fix(export): a resumed download can no longer splice two archives together
`serve_file` emitted no validator — no ETag, no Last-Modified — and ignored
If-Range entirely, while `resolve_export_file` re-reads `export_current` on
EVERY request and a download ticket survives 20 redemptions over 6 hours.

So: a guest's 500 MB Gallery.zip drops at 500 MB. The host takes a photo down
— epoch bumps, the rebuild lands, the old generation is pruned. The client
resumes with `Range: bytes=500000000-`. The ticket and session are both still
valid, the handler resolves the NEW archive, seeks 500 MB into a different
file of a different length, and streams. The client concatenates the halves
into a structurally corrupt ZIP. Nothing logs an error anywhere; a 404 would
have been the correct answer.

Now every response carries an ETag over the generation-stamped filename plus
the length, and a partial is served only against a matching If-Range. A Range
with no validator — curl -C -, wget -c, the Android download manager, all of
which resume blindly — gets the whole file instead. Restarting a download is a
cost; a corrupt keepsake is not recoverable.

Browsers send If-Range, so this is also the first release where their resume
works at all: with no validator to send, they simply refused to try.
2026-08-12 19:10:45 +02:00
fabi
f403222200 fix(deploy): a permanent upload outage, a dead-on-arrival Caddy, and 11pm commands that don't run
* Caddy had no read_body, on the reasoning that "a slow body still has to
  actually send bytes". That is an argument about disk, and disk is not the
  scarce resource: upload_admission budgets concurrent bodies at 4096 MiB and
  reserves the DECLARED cap, so a video/* upload reserves 500 MiB. Eight
  connections that stall mid-body hold the whole budget, every other guest
  waits 20s and gets a 503, and it never recovers on its own — the permit is
  held until the handler returns. No attacker needed: eight guests starting
  real videos and walking out of AP range does it, and TCP will not reap
  those sockets for hours. 30m carries a 500 MB upload at ~2.2 Mbit/s, so it
  does not fail the uploads this product exists to collect.

* APP_PORT is presented in .env.example as an ordinary editable line, while
  the healthcheck hardcodes 127.0.0.1:3000 and the Caddyfile hardcodes
  app:3000. Change it and the app boots and serves happily on the new port,
  the healthcheck fails forever, app never turns healthy — and because caddy
  is gated on service_healthy, CADDY NEVER STARTS. Port 443 dead for the
  whole event, sole diagnostic "dependency failed to start". Pinned in
  compose beside MEDIA_PATH and EXPORT_PATH, which are there for this reason.

* Runbook §12's recovery commands do not run as written: unwrapped
  "$POSTGRES_USER" is expanded by the operator's shell, which does not have
  it, so psql answers `FATAL: role "" does not exist`. §9 documents that trap
  two hundred lines earlier and wraps its own calls in sh -c; §12 did not.
  This is the block you run with the app crash-looping behind a live Caddy.
  Its DELETE also hard-coded versions 21,22,23 as if to be copied verbatim,
  on a tree that now has 31 migrations — now explicitly an example, with the
  instruction to take the numbers from the actual boot error.

* Migration counts corrected across the runbook and .env.example (22 -> 31,
  commit count 154 -> 196). All four were presented as literal command output
  the operator is invited to reproduce.
2026-08-12 09:15:55 +02:00
fabi
4b61f4552b test(e2e): fix a flake that went red when the app behaved correctly
storage-purge failed roughly one run in three on two unrelated races, both of
which blamed whatever change happened to be in flight.

`page.goto('/admin')` rejected with "interrupted by another navigation" or
ERR_ABORTED when the admin layout redirected to /admin/login first — i.e. the
test went red precisely when the app did the right thing, quickly. The
assertion is the waitForURL that follows, which does not care how the
navigation ended, so the goto is now allowed to reject. (waitUntil: 'commit'
narrows the window but an abort can beat commit too.)

And the PIN test read the page's execution context while the layout's boot
hydration was still in flight, which surfaced as an intermittent "Execution
context was destroyed". Settles the page first.

Verified with 50 consecutive runs, previously ~1 in 3 red.
2026-08-12 09:15:40 +02:00
fabi
5aa2b2e886 fix(frontend): stop a parked photo being stranded for the session, and fix an SSE id
A parked upload has three ways to be released, and two were weaker than the
toast that promises "wird gesendet, sobald die Sperre aufgehoben ist":

* The live user-shown / event-opened events only reach a tab with an open
  stream, and streams are opened by /feed, /diashow, /export, /host and
  /admin — NOT /upload, which is exactly where the toast sends the guest to
  watch their queue.
* The boot-time release ran once and swallowed any failure, so a single
  failed request on venue wifi — the condition the whole parking mechanism
  exists for — skipped it for the entire session, leaving the row reading
  "Du bist gesperrt." after the ban was long lifted.

Now retried once. Deliberately NOT on a 401: api.get already answered that by
clearing auth and redirecting to /join, so a second attempt can only fire a
second redirect two seconds later, by which time the guest may have navigated
away. (That is not hypothetical — it made an existing browser-chaos spec fail
while I was writing this.) Guarded rather than an early return, so the SSE
listener registrations below still run.

Also: noteDelivered mapped upload-processed to p.id, but that payload carries
upload_id, so it recorded nothing — the docstring claimed a property the code
did not have. And it recorded user-shown against a `carried` clause that only
ever means "hidden", where it could only suppress a later genuine signal.
2026-08-12 09:15:40 +02:00
fabi
9f239882ac fix(export-viewer): make the self-contained guard, and its spec, actually load-bearing
The font-inlining guard only caught a RENAME. It asked "did /fonts/<listed
family>.woff2 disappear?", so an ADDITION walked straight past it — and an
addition is the likelier accident: someone doing ordinary app work adds a
display font or a decorative background to the shared theme, has no reason to
open a viewer build config, and ships a keepsake that reaches for
/fonts/Playfair.woff2 on the guest's own disk. font-display: swap hides it, so
the artifact looks right to everyone who happens to have the file locally and
renders in Times New Roman for the couple.

It now asserts the invariant instead of a list: nothing in the emitted
keepsake may reference an external URL. Self-maintaining, and it covers fonts,
images and stylesheets alike. In writeBundle rather than generateBundle —
generateBundle runs more than once and the stylesheet is not inlined on the
earlier pass, so asserting there fails a perfectly good build.

viewer-no-broken-tiles gets a positive anchor. Its "nothing is broken" check
filters img elements, so a viewer that rendered NOTHING yields [] and passes:
the one spec whose whole subject is that the images resolve was the one that
would have stayed green through a total viewer regression. Everything else it
checks comes from the backend and the classic head script, neither of which
needs the viewer bundle to have run.

And a CI job, because neither of the above fires on its own: no workflow,
Dockerfile or script built this viewer, so the guard could sit disarmed
indefinitely, and the committed artifact — compiled into the binary with
include_dir! — could drift from its source with nothing to say so.
2026-08-12 09:15:24 +02:00
fabi
a2b3cb0e8d fix(db): run migrations on their own connection, not a pooled one
after_connect puts lock_timeout = 5s on every pooled connection, and the
migrator inherited it. Migrations that take ACCESS EXCLUSIVE — 026's index
swap, 027's ADD COLUMN — then turn a short WAIT into a hard FAILURE.

The runbook installs an hourly pg_dump (§10.2) and tells the operator to back
up before deploying; pg_dump holds ACCESS SHARE on `upload` and `"user"` for
its whole run, and the runbook is full of psql snippets that do the same. Boot
into that window and the migration aborts, create_pool errors, main exits 1,
and `restart: unless-stopped` crash-loops the app behind a live Caddy. The
rollback is clean and a later retry succeeds, which is precisely what makes it
a baffling intermittent outage rather than an obvious one.

026's own comment reasons that "this runs at boot before the server accepts
requests, so the brief lock costs nothing" — true of the app's own sessions,
and it does not cover anything else on the database.

statement_timeout is dropped for the migrator too: a migration on a real table
can legitimately outlast the 15s a request is allowed.
2026-08-12 09:15:24 +02:00
fabi
a428fe6957 fix(social): make the counts clients patch with ban-aware, like the view
Migration 028 added `NOT is_banned` to v_feed.like_count and
v_feed.comment_count, but not to the two scalar counts in social.rs — which
are returned in the response AND broadcast over SSE, and which clients use to
patch a card in place rather than refetching.

So the two disagreed the moment anyone was banned: the host bans a guest, the
feed correctly drops to the lower number, and the very next like on that photo
pushes the unfiltered count back to every open client — including the host's,
who is watching that number to confirm the ban took. It stayed wrong until a
full page-1 refetch.

Both call sites carried comments asserting they mirror the view. 028 made
those comments false without touching them; this makes them true again.
2026-08-12 09:15:07 +02:00
fabi
6afb33e5b6 fix(export): four ways the keepsake could be lost, stranded, or published empty
* The HTML completeness guard counted manifest ROWS, and there are up to two
  per upload — a thumbnail and a full variant. Thumbnails are 400px JPEGs the
  export generates itself into its own temp dir, so they are no evidence that
  any original was captured. After the guard was relaxed to bail only on
  "nothing written at all", that case could no longer fire while thumbnails
  kept succeeding: if the media volume became unreadable after the stat pass,
  every original open failed, every thumb open succeeded, and a keepsake with
  100 thumbnails and ZERO full-resolution photos published green, done at the
  live epoch, with the download button lit. Boot recovery skips a done job,
  so nothing would ever have rebuilt it. Now counts photos, not files.

* A decoder panic failed the ENTIRE keepsake. The `?` was on the JoinError,
  not on the closure's Result, so a panic in the image crate propagated out
  where the same file merely failing costs one tile — and it was
  deterministic, because "Neu erzeugen" reads the same poison file and dies
  the same way. That is the exact failure shape the completeness guard was
  relaxed to eliminate, arriving through the other door.

* delete_account armed both export jobs and then spawned the workers AFTER an
  awaited file-removal loop. Axum drops a handler future on client
  disconnect, and every other invalidate_and_arm call site spawns with no
  intervening await. Dropped inside that loop, the keepsake is left with the
  epoch bumped, both rows pending at that epoch, and no worker: the downloads
  404 and the UI sits on "Wird vorbereitet..." until someone reboots the app.
  Deleting your account from a phone that walks out of range is enough.

* The daily download quota was charged before the ticket could fail, so a
  store-capacity 503 — a server-side condition the guest cannot see or cause
  — still cost one of their three downloads. There is no refund path.
2026-08-12 09:15:07 +02:00
fabi
9b38d31f97 fix(upload): stop a late retry from undoing a host takedown
Migration 026 freed the idempotency key as soon as deleted_at was set, so a
retry after a delete uploads afresh instead of 409ing forever. That rationale
only considered the GUEST deleting. deleted_at is also set by
host_delete_upload, and there the same rule reverses a moderation decision:

  1. Guest uploads; the row commits and the photo appears, but the response
     is lost on the way back — the flaky-wifi case the key exists for — so
     the phone keeps the queue item.
  2. The host takes the photo down. Epoch bumped, keepsake rebuilt without it.
  3. The phone reconnects ten minutes later and retries. The key is free, the
     INSERT succeeds, and the photo is back — in the feed and in the next
     keepsake, under a NEW uuid that matches nothing in the host's moderation
     history, with nothing logged to say a takedown was reversed.

Migration 031 keeps the key claimed for a host takedown and releases it only
for a guest's own delete, so the retry resolves to the duplicate path and is
refused. The refusal now says why ("von den Gastgebern entfernt") rather than
"already processed", which invites another try.

The index predicate and the ON CONFLICT arbiter are changed in lockstep;
these queries are not compile-checked, so a drift between them is a 500 on
exactly the retries the index exists to serve. Verified against a real
Postgres: live retry suppressed, host takedown holds the key, guest delete
releases it. The integration test's copy of the insert is updated too — it is
verbatim by design, and a stale copy would have kept passing.
2026-08-12 09:14:51 +02:00
fabi
1b3ca46f8a fix(auth): three ways one guest on the venue NAT could lock everyone else out
All three are the same mistake in different clothes: a limit keyed on an IP
that, behind the venue's NAT, is the entire party plus the host.

* join_ip_rate_per_min was raised 60 -> 300 last round and it never took
  effect. A config default is only a fallback for a MISSING key, and
  migration 017 seeds this one, so the seed won and the raise was dead code
  on every real install. Migration 030 raises the seeded value the way 015
  already did for upload_rate_per_hour. The e2e guard could not see this:
  it fires 12 concurrent joins, which is green at 60 and at 300 alike.

* /recover's per-(IP, name) bucket charged EVERY request, including
  successful ones, and refused before verifying the PIN. Its ceiling clamps
  to 4. So four POSTs naming "Braut Sophie" with PIN 0000, from any phone on
  the venue wifi, locked Sophie out of her own recovery for fifteen minutes
  WITH THE CORRECT PIN — and four more every fifteen minutes sustained it
  indefinitely, at a rate far under every volume ceiling above it. The benign
  version needs no attacker: the host mistypes their own PIN four times.
  Hosts are promoted guests whose only credential is that PIN, and /recover
  is their only way back after losing a session.

  Now it counts failures, and a spent budget changes what a FAILURE answers
  instead of refusing outright. Guessing is bounded exactly as before —
  wrong PINs are what spend it — with the per-account lockout underneath.

* /admin/login's pre-verify ceiling had the same shape, and the escape hatch
  was circular: admin_login_rate_enabled is only flippable through
  PATCH /admin/config, which needs the session being refused. One phone
  posting twice a minute cost the operator moderation, gallery release and
  every config key, including the ones that would undo it. Exceeding the
  ceiling now shortens the hash-permit wait rather than refusing: the CPU
  bound was always the semaphore, never this bucket, so a flood still sheds
  itself while a correct password gets a truthful answer.

Adds a regression test that reads the value a fresh database actually ends
up with, by replaying the migrations — the drift that made the first bullet
invisible is not otherwise detectable from the code.
2026-08-12 09:14:39 +02:00
fabi
0c0d5d5981 fix(deploy): give Postgres a CPU floor that Docker actually honours
`deploy.resources.reservations.cpus` was doing nothing. Outside Swarm, `docker compose up`
silently drops it — verified by inspecting a running container, where CpuShares, CpuQuota
and CpusetCpus were all unset while `limits.cpus` and `reservations.memory` came through as
NanoCpus and MemoryReservation. So the comment calling it "the piece that actually protects
the database" described a guarantee the box never had.

It matters on the CX22 the runbook targets: the ceilings sum to 1.2 + 0.6 + 0.5 = 2.3 on
2 vCPU, so the other services can oversubscribe the machine, and with every container on the
default weight Postgres competed on equal footing with two image resizes and an ffmpeg
poster. Replaced with `cpu_shares`, which does survive the translation — db 2048, caddy
1024, app 512, frontend 256 — so the weighting only binds when the CPU is actually
saturated, which is the moment the database must not lose.

The Caddyfile gains a 10s header-read timeout: there was no read timeout anywhere, so a
client could hold a connection, a tokio task and a `.tmp` file open indefinitely by sending
one byte a minute, and the upload sweeper is keyed on mtime precisely so a live upload never
ages out. Body reads stay unbounded — a 500 MB video over cellular legitimately takes
minutes, and a body timeout would fail exactly the uploads this product exists to collect.

.env.example documents that estimated_guest_count is a live input to the quota divisor
rather than the inert setting both it and the runbook previously implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:44:48 +02:00
fabi
32dfe6874a test(e2e): make nine red specs assert the contracts the code actually implements
The e2e suite had never been run during this audit. It failed 9 of 256; seven of those
predated the audit's changes, established by building a stack from a clean HEAD worktree
and running the same specs against it rather than guessing.

Most were stale assertions rather than product defects:

- quota.spec solved for a target limit using the observed uploader count, but the divisor is
  max(active, estimated_guest_count, 1) and that config seeds at 100 — so every limit it
  aimed for came out 100x small and every "within quota" upload 413'd.
- rate-limit-shared-nat destructured `ticket` from a 429 body and fetched with
  `ticket=undefined`, turning the 429 under test into an unrelated 401. It also faked a
  release with no archive on disk, so the mint's pre-check 404'd and the per-day limiter was
  never reached; it now does a real release and asserts 200 rather than "not 429".
- ddos allowed only [200,429] from ten concurrent streams, so it failed on the very defence
  it exercises: four tickets per session survive and the rest correctly 401. Now asserts
  exactly four, which a tightened cap or an inverted eviction order would catch.
- auth-tampering asserted a throttled IP is refused EVEN with the correct password. That
  contract was deliberately removed — it let any phone on the venue NAT lock the operator
  out of their own admin panel, with a circular escape hatch. Inverted, plus a new check
  that a success does not refill an attacker's bucket.
- moderation-ui assumed a ban leaves a comment "stuck on screen"; `list_for_upload` filters
  banned authors, so it is hidden from everyone including the host. Now pins the pair that
  matters — the ban hides it, and the host's permanent removal survives an unban — and the
  UI leg it used to own is restored as a separate test on a reachable comment.

The export specs mint with `?kind=` now that a download ticket is bound to one archive, and
four of them assert the mint's 404 rather than the download's: with the kind always known,
the pre-check refuses up front instead of after charging a daily download for an archive
that cannot be served.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:44:48 +02:00
fabi
a53729a704 fix(frontend): park uploads that cannot succeed, and stop two false signals
The upload queue gains `parkedFor`, so a photo rejected for a reason that cannot change on
its own stops re-pushing itself. A ban used to come back as a generic `forbidden`, which
purged the blob and moved the row to `blocked` — a terminal state with no retry button — so
lifting a ban restored everything except the photo actually in flight. Ban and release are
now distinct codes that keep the blob, charge no attempt, and tell the guest what has to
happen. `releaseResolvedParks` drains them at boot from /me/context, because the live
`user-shown` / `event-opened` events only reach a tab that was open when the host acted,
and the usual sequence is the other way round.

Two signals were firing on nothing. A filtered feed set `feedStale` on EVERY delta without
deduping — and the delta cursor boundary is inclusive while sse.ts deliberately rewinds
`lastEventTime`, so deltas routinely re-return rows already delivered. With the backstop
polling every 60-120s, a guest who tapped a hashtag got a "Neue Beiträge" pill they could
never clear, each tap costing a full filtered refetch. It now dedupes in both branches.

The SSE liveness backstop had the mirror problem: `noteDelivered` harvested id, upload_id
AND user_id from every payload, so by the time anything was deleted or anyone banned, their
ids were already marked delivered from ordinary traffic about live content. The
`deleted_ids` and `hidden_user_ids` clauses were false essentially always, leaving a
half-open socket undetected while a host moderated into a feed nobody was listening to.
Each event now records only the id its own clause tests.

Also: /admin no longer bounces to /join on a cleared session — AUTH_ROUTES had the `/admin`
prefix, which suppressed clearAuth() on the dashboard and let the login guard bounce back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:44:26 +02:00
fabi
f5c55d6f92 chore(backend): route wiring, error mapping, and a crossbeam-epoch bump
Cargo.lock moves crossbeam-epoch to 0.9.20, clearing RUSTSEC-2026-0204. Targeted rather
than a broad `cargo update` across 406 crates, which is not a change to make days before a
live event.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:44:26 +02:00
fabi
8720571beb fix(export-viewer): inline the webfonts so the offline keepsake is self-contained
The viewer inherits `src: url('/fonts/Inter.woff2')` from the shared theme CSS. That is
correct for the app, which serves `static/fonts/` from the site root — but the keepsake is
opened from file:// off a USB stick or a Downloads folder, where `/fonts/...` resolves to
the root of the guest's DISK. Both requests 404, and `font-display: swap` makes it silent:
the viewer renders in a fallback system font with nothing server-side able to report it.

Found by opening a real released keepsake in a browser and watching `requestfailed` — no
other signal exists, which is the recurring lesson about this artifact.

Fixing it in the shared CSS would inline ~154 KB into every app page load for nothing, and
shipping a `fonts/` folder beside index.html gives the guest a directory they can break by
moving one file. So the substitution belongs in the build that knows its output has no
origin. The plugin errors the build if the theme ever stops referencing those URLs, rather
than silently shipping another keepsake in Times New Roman.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:44:10 +02:00
fabi
c9a4d4a9c0 fix(export): stop the keepsake guards from destroying the keepsake
Two guards added to protect the archive each had a failure mode worse than the one they
prevented, and both were unrecoverable — which is what makes them worth reverting rather
than tuning.

The completeness gate refused to publish once skips passed max(2, 10% of expected). That
refusal is DETERMINISTIC ACROSS RETRIES: the unreadable files are still unreadable when the
host taps "Neu erzeugen", and the gallery is already released so the uploads cannot be
collected again. On a 30-photo event, four bad files meant nobody ever got the other 26.
That is precisely the "one-photo gap becomes total loss" outcome MAX_SKIPPED_FRACTION's own
comment says it exists to avoid. Anything short of an empty archive now publishes and logs
the counts at error level. `written == 0` stays fatal — a wrong MEDIA_PATH is a
misconfiguration the host CAN fix and retry, and it once shipped a few-hundred-byte ZIP
containing zero photos that passed every automated check.

The space reclaim refused to prune unless it freed the entire shortfall, to protect an
archive that no handler can serve: a download resolves through `export_current`, which
requires `job.epoch = event.export_epoch`, and the epoch only increments. Meanwhile
`reclaimable` is scoped to the caller's own prefix — one old archive — while `deficit` is
sized for both halves plus the reserve. So on a tight disk each worker measured its own
share as insufficient and neither pruned, though the two shares were jointly sufficient.
Every "Neu erzeugen" reran the identical arithmetic and refused identically: permanently
stuck, with dead archives nothing would reclaim and nothing could serve. It now prunes what
it can and lets the re-check decide, so the sibling's prune lets the host's retry converge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:44:10 +02:00
fabi
7154b3a810 fix(upload): remove the /original rate limit that would have broken the feed
The limiter added here was justified as bounding "100 guests occasionally tapping Original
anzeigen". That is not what this route is. `pickMediaUrl` resolves to
`preview_url ?? thumbnail_url ?? /original`, and a freshly committed upload has BOTH
derivatives null until the compression worker reaches it — at COMPRESSION_WORKER_CONCURRENCY=2
that is minutes during a post-ceremony burst. So /original is the feed's hot path for exactly
the newest photos, in a newest-first grid, at the busiest moment.

With every guest behind one NAT the 600/min bucket is venue-wide: six new photos fanned out
by `upload-new` to ~100 open feeds exhausts it, and then every original fetch from anyone
429s for the rest of the window. The tiles' own 4-second retry uses a fresh `?r=` nonce, so
the clients hold the bucket saturated themselves — the whole venue watching the newest
photos render as broken tiles while the projector skips slides.

A per-IP bucket cannot separate one scraper from the entire party when they share an
address, and these media routes are unauthenticated by design (an `<img>` cannot send a
bearer token), so there is no per-user key to move to. Bandwidth abuse belongs at the proxy.

Also here: the release/lock check order. `release ⇒ lock`, so testing the lock first made
the `GalleryReleased` arm unreachable dead code and every post-release upload answered
`uploads_locked`. The codes are not interchangeable to the client — `uploads_locked` charges
a retry attempt and re-pushes the whole photo on the backoff ladder against an answer that
cannot change, while `gallery_released` parks it and says the photo is safe but the hosts
must reopen. Both sites now test release first, so the fast path and the commit-time
re-check agree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:43:50 +02:00
fabi
ec7c7f18ca fix(auth): stop one guest on the venue NAT from locking everyone else out
Every guest at the venue shares one public IP, so an IP-keyed limiter throttles the whole
party as a single client. Three separate limits got that wrong, and the host — whose only
credential is a 4-digit PIN — was the one who could not absorb it.

/recover carried a cross-name failure budget checked BEFORE the account lookup, so it
refused a CORRECT PIN. Thirty POSTs with invented names spent the shared budget for fifteen
minutes and ~2 requests/minute sustained it indefinitely, denying PIN recovery to everyone
including a host locked out of their own event. The budget is now carried as a flag: a
correct PIN authenticates regardless, while wrong ones answer 429 instead of 401. Guessing
stays bounded where it always really was — the per-(IP,name) ceiling and the per-account
3-strike lockout, neither of which an attacker on any IP can evade.

join_ip went from 60/min to 300. A 100-guest wedding does not trickle in; it arrives when
the QR code goes up, all from one address, and guests 61-100 were turned away on the one
screen with no auto-retry. This limit only bounds raw volume — the per-name bucket is the
anti-spam control and BCRYPT_PERMITS is the CPU bound — so it can sit well above the peak.

Download tickets are now bound to ONE archive via `TicketKind::Download(ExportKind)`. Both
download routes share an authenticator, so a bare ticket opened either; combined with the
resume budget that made a single mint worth 40 transfers of a multi-GB keepsake while the
per-day limiter, charged only at mint, never moved. `kind` is consequently required at
/export/ticket; every shipped client already sends it.

The per-session ticket cap is now per-kind. A 6-hour download ticket is always the oldest
entry for its session, so ordinary SSE churn evicted it first — and /export opens its own
SSE connection on the session that just minted it. A couple of wifi flaps mid-transfer
killed the ticket, 401'd the resume, and cost the guest another of three daily downloads.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:43:29 +02:00
fabi
963f6449a1 feat(db): idempotency keys, ban-aware counts, and a host audit trail
Four migrations, all additive against a database that already has 001-025 applied.

026 narrows the client-upload idempotency index with `AND deleted_at IS NULL`. The old
index made a soft-deleted row keep its key forever, so a guest who deleted a photo and
re-sent the same one had the retry silently swallowed. The new indexed set is a strict
subset of the old, so it cannot fail on existing rows.

027 adds `client_join_id`, which lets a join retry after a lost response resume the same
account instead of 409ing on a name the caller itself owns. Every existing row gets NULL
and the partial index excludes NULLs, so it indexes nothing at creation.

028 brings the feed view's like/comment counts in line with what the feed actually renders:
a banned guest's rows were still counted, so a card showed "3 comments" above two.
`comment.rs` gets the matching `NOT u.is_banned` on the live read path — the export and
hashtag queries already filtered it, so the two views of one moderation action disagreed.

029 records host moderation actions, which were previously invisible after the fact.

Verified by applying 001-029 to a real Postgres against seeded data, including a
soft-deleted row holding a key and a banned user's like and comment. 026's down-migration
legitimately fails where a deleted and a live row share a key — that is inherent to the
direction, documented in the file, and sqlx never runs downs at boot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:43:11 +02:00
MechaCat02
ef6d3a077a fix: close what nine adversarial reviews found, most of it mine
Some checks failed
Checks / Backend — cargo test + clippy + fmt (push) Failing after 1m5s
Checks / Frontend — vitest + svelte-check (push) Failing after 5m55s
Checks / E2E — typecheck + lint (push) Failing after 49s
E2E / Playwright E2E (chromium + webkit) (push) Failing after 10m42s
E2E / Cross-UA smoke matrix (push) Failing after 7m57s
Audit / cargo audit (backend) (push) Failing after 11m12s
Audit / npm audit (frontend) (push) Successful in 44s
Nine focused reviews (export state machine, upload path, auth/abuse, client
queue, guest UI, database, deploy/ops, regression hunt, test honesty). Every
finding below was re-verified against the code before being acted on; several
plausible-sounding ones were checked and rejected.

## Data loss and denial of service

**One request could OOM-kill the app container.** `client_upload_id` was read
with `Field::text()` — axum builds its multipart reader with no SizeLimit, so
the only bound was the route's 576 MiB body limit, then decoded into a second
full String. `caption` and `hashtags` go through `read_text_field_bounded` for
exactly this reason; this field arrived later and missed it. Any guest, one
request, and every SSE stream drops and every in-flight temp file is stranded.

**Nothing bounded concurrent upload bodies.** The headroom gate can only refuse
to COMMIT — the body is already streamed to a temp file by the time it runs, and
neither axum, the tower stack nor Caddy limits how many stream at once. ~100
guests tapping "upload all" after the ceremony puts 10-20 GB of .tmp on a 40 GB
volume, invisible to the gate, eating the reserve that keeps Postgres able to
write WAL. New `UploadAdmission` budgets bytes (not requests, so one video and
two hundred photos coexist) via a permit that releases on drop, so every exit
path returns it.

**The export decode bypassed the memory permit the compression path takes.**
Same class of work — decode + resize every image in the gallery — in a bare
spawn_blocking. A release fired while the last photos were still compressing put
both in the same 1 GiB cgroup; the OOM kill marks the export failed and
`recover_exports` re-spawns it into the same conditions on the next boot. The
permit is now process-wide in `imaging`, because the constraint it expresses is
the container's memory, not one worker's.

**`MediaTotalCache` cached its own failure as 0.** For the whole TTL the gate
then saw an empty event and collapsed to the flat reserve — the behaviour the
two-halves design replaced — with no log line. And the trigger correlates with
the danger: with max_connections 10 the query fails exactly during a burst. Now
falls back to the last good reading and says so.

**V8's heap ceiling sat above the frontend container's entire budget** (measured:
259 MB inside a 256M limit), so GC could never intervene and the only
backpressure was SIGKILL under an arrival burst.

## Guest-visible

**The feed stopped being newest-first after the first reconcile.** It fetches
whole 100-item server pages while `uploads` grows in 20s, so everything in the
gap was absent from `present`, classified as new, and prepended — ~80 photos
from earlier in the evening above the newest ones. It also stalled infinite
scroll, since the cursor still pointed at item 20 and the observer only re-fires
on a change. The union is now sorted on the server's own (created_at, id) key,
which additionally places an SSE arrival correctly.

**A stale `loadMoreError` outlived every refresh and filter change**, leaving a
false error above a button that returns immediately on `!nextCursor`.

**A failed derivative toasted "Ein Upload konnte nicht verarbeitet werden."** for
a photo sitting right there on screen — the handler still assumed 1d9fb11's
pre-fix behaviour (row deleted, quota refunded, card evicted), none of which is
true any more. It was the last surviving route for the "your photo is gone"
signal that fix set out to remove.

## Enforcement that existed only in comments

`recover_name_rate_per_15min` is clamped at the point of use: the ordering
`3 x ceiling <= PIN_LOCK_THRESHOLD` is the whole control against one source
locking any guest whose name is on the feed, it was asserted in a comment, and
`patch_config` accepted 1..100_000. The test pinned the default constant rather
than the enforced bound; it now pins the bound.

## Tests that could not fail

- The gate test asserted only its own premise (`500MB x 100 > 35GB`) and never
  touched the gate. It now checks both controls against the same state and
  requires them to disagree in the right direction.
- `the_banner_always_fires_before_the_upload_gate_closes` reduced to
  `G < G + G/4` — true for any margin, including zero, so it could not detect
  the banner moving to exactly the gate. It now pins the gap.
- `disk_is_low`'s `free < LOW_DISK_FLOOR_BYTES` clause was unreachable (warn_at
  is always >= 12.5 GB against a 10 GB floor). Two tests were named after it and
  neither could fail if it were deleted. Clause and constant removed.
- The suspension test I added last commit hard-coded the credit cap instead of
  importing it, so changing STALL_TIMEOUT_MS would leave it passing against a
  system that no longer exists. Now imports MAX_SUSPEND_CREDIT_MS.

## Stale comments corrected

The prune doc still argued at length for the pre-build ordering that 1d9fb11
reversed — a reader trusting it would reopen the blocker 0506369 fixed.
DISK_RESERVE_BYTES claimed to equal the banner threshold that 0506369
deliberately offset by 25%. And host.rs kept its own duplicate 10 GB literal
instead of importing the constant.

154/154 backend, 59/59 vitest, clippy clean, svelte-check 0 errors, eslint
clean, both builds, compose + caddy validate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 14:51:58 +02:00
MechaCat02
214f9e3062 fix: close four confirmed defects an adversarial review found
Findings from a multi-angle review, most of them in code I wrote in the last
few commits. Each was verified against the code before being acted on.

## Backend

**The export daily limit was bypassable ~60x/minute.** `SseTicketStore` is
untyped, and the export download quietly started reusing it. `POST
/stream/ticket` is free and rate-limited at 60/min per user; `POST
/export/ticket` charges one of three PER-DAY downloads. So a guest could mint
at the cheap endpoint and redeem at the expensive one, each redemption
streaming the whole multi-GB keepsake, `no-store`, off the same filesystem
Postgres writes WAL to. Tickets now carry a `TicketKind` and `consume` requires
it to match, asserted in both directions. The comment claiming "one mint is at
most one download" was simply false.

**`export_ticket` answered 200 `{"ticket": null}` when the store was full** —
after charging a daily slot. `issue` returns `Option`; `sse.rs` handles the
None with a 503 and this call site unwrapped it into the JSON body. The page
toasted success, the iframe navigated to `?ticket=null`, and one of three
downloads was gone. That is the phantom-success failure the pre-validation in
5b70531 exists to prevent, arriving through the other door.

**`finalize_job` collapsed a DB error into "we lost the epoch race."** At that
point the archive is built, fsynced and renamed, so the caller deleted the
finished multi-GB file and returned the Superseded sentinel — which
`abandon_if_superseded` swallows into Ok, so `mark_failed` never ran either.
The row stayed `running` at 99% at the LIVE epoch: "Wird erstellt (99 %)",
download disabled, forever. No sweep re-examines `running` rows and
`recover_exports` runs only at boot. `claim_job`'s own doc comment says errors
are distinguished there precisely because of this failure shape. A pool timeout
is not exotic: max_connections 10, acquire_timeout 5s, firing at the end of a
full-gallery export while 100 guests upload.

**`PATCH {"hashtags": []}` was a free keepsake-retire loop.** The no-op guard
only compared captions, and my comment defended the gap by claiming an
identical hashtag list "is not a free loop". It is exactly one. Each request
bumped the epoch, retiring the HTML keepsake; REGEN_DEBOUNCE throttles when a
rebuild may start, not the bump, so at 30/min no rebuild ever gets a quiet
window and /export/html 404s all event. Now compares against the stored tags.

Also: four config keys migration 025 inserts (and the handlers read) were
missing from `patch_config`'s allowlist, so `GET /admin/config` listed them
while `PATCH` answered "Unbekannter Konfigurationsschlüssel" — the rate limits
an operator reaches for while abuse is happening.

## Client upload queue

**The ✕ was cosmetic.** A cancel deliberately charges no attempt and sets no
backoff — so `requeueRetriable` matched it on both counts and restarted the
upload from byte zero within ~120s (an `online` event, or the SSE backstop's
`feed-delta` poll). It then restarted forever, because a path that never
charges an attempt can never exhaust the budget that would stop it. The row
read "Abgebrochen. Tippe auf „Erneut“." throughout. Cancels are now explicitly
terminal until the guest taps Erneut.

**The retry budget was a lifetime quota, not a rate.** Five attempts on a
5/10/20/40s ladder is ~75 seconds, so any outage longer than that — a venue AP
brownout, a captive portal re-arming, an `app` restart, all with
`navigator.onLine` still true — permanently parked every in-flight photo
behind a per-row button three taps deep. It now refills after 10 quiet minutes,
which still forbids a hot loop re-sending a 200 MB video over a shared uplink.

**A test asserted a property the code does not have.** The suspension test
omitted the MAX_SUSPEND_CREDIT_MS clamp the production tick applies, so it
could not fail. Replaced with a helper that replays the real tick loop, and the
true bound is now asserted: a 60s lock survives, a 3-minute lock aborts.

152/152 backend, 59/59 vitest, clippy clean, svelte-check 0 errors, eslint
clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 14:39:54 +02:00
MechaCat02
253878e027 fix(export): give the keepsake viewer the two-phase preflight it was meant to get
The two-phase preflight from eb0e405 landed in ONE place and was spliced inside
the other. `run_zip_export` ended up containing both blocks nested, so the
Gallery path pruned "Memories" archives that were not its to reclaim, while
`run_html_export` silently kept the single-phase form.

That left the exact deadlock the two-phase preflight exists to break, still
open on half the product. At a gallery size where a rebuild needs the previous
generation's bytes: the ZIP prunes its own superseded archive and rebuilds, and
the HTML preflight fails against a Memories archive still on disk. The prune
that would free it runs only after a success that can never happen, and any
epoch bump — a guest deleting one photo — retires the current viewer
immediately. Permanently stuck, unreachable from any handler, discovered at the
end of the night with nobody there.

Both halves now call one `ensure_export_space_reclaiming`, keyed on the
caller's OWN prefix, so they cannot drift again.

Three smaller things found in the same pass:

- The boot-failure panel hardcoded light-mode colours, and its heading set none
  at all — the UA default black on the `#100f0f` dark background. On the one
  screen whose entire job is to be readable, and in the failure mode where the
  app's own stylesheet may be what did not load. Moved to classes in the inline
  <style> so the `html.dark` variants apply.

- `.env.example` assigned RUST_LOG twice, 120 lines apart. Compose takes the
  last one, so an operator raising the level mid-event to chase a problem would
  have changed nothing, silently.

- The comment justifying `detail = ?message` in error.rs still claimed
  `validate_display_name` allows newlines. It rejects control characters now —
  but that is one input against every 4xx message in the app, so the escaping
  is what makes the guarantee general. Said so.

151/151 backend, 58/58 vitest, clippy clean, svelte-check 0 errors, both
builds, caddy validate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:48:22 +02:00
MechaCat02
5b705317ef fix: close the last three guest-facing dead ends (items 7-9)
C1 — a failed page-append silently ended infinite scroll

`loadMore`'s catch showed a toast and changed no state, unlike every sibling error
path in the file. `nextCursor` survived so the feed was still technically paginable,
but the IntersectionObserver only fires on a CHANGE: after a failed append nothing
scrolls and no rows are added, so it never re-fires. One 429 or wifi blip and the
guest concluded the gallery was 20 photos. Now leaves a retry control at the sentinel
— a toast that fades in 5s is not an affordance — resuming from the untouched cursor.

C2 — the export page reported downloads that never happened

`downloadFile` toasted 'Download gestartet' the instant it assigned the iframe's src,
before a single byte existed. Since the iframe swallows errors BY DESIGN (a top-level
navigation to a 404 would unload the PWA), a failure produced a green success message,
a consumed single-use ticket, and one of only three daily slots spent — repeatable
until the day's allowance was gone, on the screen that is the whole point of the app.

Root cause is two sources of truth: `export_status` reports `done` from `export_job`
and enables the button, while the download resolves through `export_current.file_path`
plus a `Path::exists()`. They can legitimately disagree. `export_ticket` now takes a
`kind` and calls the existing `resolve_export_file` BEFORE charging the rate slot, so
a missing archive fails honestly on a plain fetch that `toastError` already renders.
Not the HEAD probe ruled out elsewhere: it reads the same indexed row the download
will read and touches no ticket, so it cannot consume anything. The parameter is
optional, so an older client degrades to today's behaviour rather than breaking.

C3 — the WhatsApp journey could dead-end with no error at all

The join link travels through guest group chats, and a link tapped inside one opens in
that app's browser, where the file picker and getUserMedia both depend on the host app
having wired them up. When they aren't, the buttons do nothing — no error, nothing to
act on. Two targeted changes rather than a UI rebuild: the camera error panel now
offers "Aus Galerie wählen" (its advice to change "Browsereinstellungen" refers to
settings that do not exist in a webview, so retrying could never help those guests),
and the sheet carries a standing one-line hint to open the link in Safari or Chrome.

Deliberately no user-agent sniffing: a sniff list is wrong for every browser it has
not heard of, while a quiet standing hint costs one line and is never wrong. The hint
lives in UploadSheet rather than the root layout because both layout banners are gated
on `$showBottomNav`, which `/upload` turns off — one there would never render on the
composer.

Verified: 151/151 backend tests against a live Postgres, clippy clean, 58/58 vitest,
svelte-check 0 errors, eslint clean, both builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:38:07 +02:00
MechaCat02
05063694d2 fix: close eight regressions the audit pass found, five of them mine
Two adversarial reviews over 61119be, 1d9fb11 and eb0e405. The merge itself came
back clean — client_upload_id end to end, TempFileGuard's arm/retarget/disarm, the
supervised sweep wiring and v_feed's column parity were all verified sound. What
follows is what my own three commits broke.

BLOCKER — a post-release rebuild was permanently impossible, and it 404'd the keepsake

1d9fb11 deferred prune_superseded_archives to run only on success, so a failed
rebuild could no longer destroy the last good archive. It did not follow that
through: ensure_export_space runs BEFORE the prune, so at rebuild time the previous
generation is still on disk and counted against free. That halves the gallery a
rebuild can survive (~4.6 GB) relative to what the upload gate accepts (~7.8 GB) —
and it self-locks, because invalidate_and_arm bumps the epoch on COMMIT, which 404s
both download routes immediately, while the only code that could free the space now
runs only after a success that can never happen. A guest deleting their own photo is
enough to trigger it. Recovery needed `docker exec rm`.

Now two-phase: try to build while preserving the old generation; if that genuinely
does not fit, reclaim it and try once more. Strictly better than both the original
ordering and my change — the old archive is sacrificed only when it is the only way
to get a new one.

BLOCKER — the deferred prune could delete the last archive when a worker LOST the race

run_*_export_inner returned Ok(()) on the superseded/discard path, so `res.is_ok()`
fired the prune with the worker's own RETIRED epoch as keep_seq. At that moment the
winning generation is still `pending` with no file, so protected_files is empty and
the last good archive was deleted with no replacement. Exactly the invariant
deferring the prune was meant to establish. Returns Err(Superseded) now, which
abandon_if_superseded already swallows for the caller.

BLOCKER — the low-disk banner could never fire before the wall

eb0e405's gate refuses at `free < keepsake + DISK_RESERVE`, while disk_is_low warned
at `free < keepsake`. The two differ by the whole reserve, so the wall always came
first: every guest blocked from uploading while the host dashboard showed ~27 GB free
and no banner, with nobody on site. disk_is_low now shares the gate's expression plus
a 25% margin, and a test asserts the banner fires at the gate threshold across the
whole gallery-size range.

BLOCKER — I raised the unauthenticated bcrypt ceiling 24x on a 2 vCPU box

1d9fb11 moved admin_login's tight bucket after verify_password (correct — that is what
stops a guest locking the operator out) but replaced the incidental 5/min bound on
bcrypt with 120/min and nothing global. bcrypt is on spawn_blocking, but tokio's
blocking pool is 512 threads, so "off the runtime" is not "bounded": enough concurrent
verifies preempt both async workers and uploads, feed and SSE stall. Three
unauthenticated endpoints reach bcrypt and every guest shares one NAT IP, so per-IP
limits bound nothing globally. Adds a process-wide semaphore of `cores - 1` around both
verify and hash, and drops the ceiling to 30.

Also correcting my own claim: "a correct password is never throttled" was wrong. The
failure bucket cannot block it, but the CPU ceiling still can. The code comment said so;
the commit message did not.

BLOCKER — migration 025 could crash-loop the app on boot

Its UPDATE derives `Name (8hex)` with no guard against idx_user_event_name_ci. A guest
who had already joined as exactly that string makes the migration fail, which
propagates out of create_pool, exits main, and `restart: unless-stopped` turns it into
a permanent loop — a worse version of the lockout the migration exists to clean up.
Now skips colliding rows (create_admin_user already falls back to Admin-<8hex>, so the
cleanup is convenience, not load-bearing). Also `role = 'guest'` rather than
`<> 'admin'`, which was renaming legitimately promoted hosts named "Host".

DEGRADATION — the watchdog's suspension credit was unbounded

Background tabs are throttled to ~1 tick/min WITHOUT the network stack pausing, and the
tick gap cannot tell that from a freeze. Crediting every late tick grew the observed
silence by only one interval per real minute, so a dead socket took ~18 minutes to
detect while holding the queue's processing latch. Credit is now capped at one stall
window and REFILLS on real progress: an upload that is moving survives any number of
screen locks, while one that is silent and suspended is detected within ~3 minutes.

DEGRADATION — the 4xx log line was an unauthenticated log-injection vector

validate_display_name allowed newlines, several 4xx messages interpolate the name, and
%message wrote it unescaped. Two unauthenticated /join requests could forge arbitrary
lines in the only forensic record an unattended event has. Fixed at both ends: control
characters rejected at the door, and `detail = ?message` escapes on the way out (which
also stops colliding with tracing's reserved `message` field). 401/404 drop to DEBUG —
they carry no operator signal and were the cheapest lines for a scanner to use to roll
the 30 MB log window in minutes.

DEGRADATION — the quota floor was inverted exactly where it mattered

`computed.max(MIN.min(budget))`: `budget` is the whole disk's share, so below 500 MiB
the "floor" became the entire remaining budget and EVERY uploader was authorised all of
it — 400 MB free, 3 uploaders, 300 MB each. A test pinned that as correct under the name
`the_floor_never_exceeds_what_the_disk_can_back`. Both fixed.

Also replaces the headline gate test, which asserted its own precondition inside an `if`
on that precondition and could not fail. It now pins what actually binds the gate to the
preflight — that required_free_bytes charges for both halves — plus the ceiling band.

Verified: 151/151 backend tests against a live Postgres, clippy clean, 58/58 vitest,
svelte-check 0 errors, eslint clean, both builds, caddy validate, and the migration
collision reproduced against Postgres 16 before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:33:43 +02:00
MechaCat02
23e2f485dd fix(upload): correct two errors in the keepsake headroom gate
Both found reviewing my own change rather than by a test, which is the point.

DOUBLE-SUBTRACTION. The gate computed `free - size`, but the body is streamed to
its temp file during multipart parsing — far above the gate — so the free-space
reading already excludes those bytes. Subtracting again refused uploads a full
file-size early; with max_video_size_mb at 500 that is half a gigabyte of phantom
pressure. `media_total` genuinely does need `+ size` (its row is not committed
yet), which is what made the asymmetry easy to miss.

BLOCKING SCAN ON THE HOTTEST PATH. It called `disk::free_bytes`, whose doc comment
says it deliberately bypasses DiskCache — but that rationale is the export
preflight's: a rare, high-stakes decision where a sibling worker can move free space
by tens of GB inside the TTL. Per upload it means sysinfo re-scanning every mount,
synchronously, on the async runtime, on a 2 vCPU box with two worker threads. Now
uses the cached snapshot, the same 15s staleness the quota check immediately below
already accepts for the same question.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:10:40 +02:00
MechaCat02
eb0e405562 fix: gate uploads on keepsake headroom, and close five unattended-event gaps
The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit,
the committed comments and README's sizing section all assumed. That correction
is what the first change is about; the rest are the remaining pre-event items.

THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING

`required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each
gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The
upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let
uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free.
Every upload in that band succeeded and the keepsake could then never be built: the
product's entire promise, failing silently at the end of the night with nobody there.

The gate now enforces the invariant that actually matters — never accept an upload
that would make the keepsake unbuildable — sharing `required_free_bytes` with the
preflight so the two cannot drift into disagreeing about the same question. Uploads
stop at ~8 GB of media on this disk, with a German message naming the cause.
Refusing the 1001st photo beats losing all 1000.

`media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s,
rather than `estimate_export_bytes`'s join across every upload. It counts hidden and
banned users' bytes, which the export excludes — skew in the SAFE direction, so the
gate closes marginally early rather than late. Fails open on a query error.

A test pins the gate against the preflight across the whole gallery-size range, and
a second asserts the per-user floor alone would over-commit the volume — i.e. that
the global gate is what must bind.

THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED

`Date.now()` advances while a backgrounded phone is frozen but `setInterval` does
not, so the first tick after a screen lock read the whole sleep as silence and
aborted — re-sending a video from byte zero and burning one of five PERMANENT
auto-attempts. The interval is now its own suspension detector: a tick that arrives
125s late for a 5s schedule credits that window back, because a period the watchdog
could not observe is not evidence of silence.

Chosen over a `visibilitychange` listener, which only covers causes that fire that
event — a throttled-but-visible tab, a closed lid and an occluded window all freeze
timers without one — and which would have needed module state, an SSR guard and a
teardown for strictly less coverage. `performance.now()` was rejected because Safari
pauses it across system sleep on some paths and Chrome does not.

The credit buys one fresh window, not immunity: a socket iOS reaped while
backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout`
(5-60 min) with the queue's `processing` latch held.

Two latent leaks found while in there: `xhr.abort()` on a request already in
readyState DONE emits no `abort` event, so `settle()` never ran and the interval
re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button
silently stopped working); and a synchronous throw from `xhr.send` — a blob whose
backing store the OS purged — leaked the same way. Both closed.

OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER

red/amber/green were never in the @theme block and fell through to Tailwind v4's
`oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot
parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color`
falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to
Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already
in this pipeline — so modern browsers render exactly what they render today. Verified
against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal
(avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0.

Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on
the host dashboard was rendering default violet on EVERY browser, off-brand.

The keepsake viewer only picks this up on a rebuild, so its committed artefact is
rebuilt here too — still single-file, still zero external references.

A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER

With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a
redeploy or a dead uplink left the guest on the boot spinner with no message, no
reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing
nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a
timeout rather than feature detection: a SyntaxError in the bundle is invisible to any
capability check. Plus a <noscript>, since there was nothing at all to see without JS.

EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL

tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info.
If guests spend the evening hitting 429s or 413s, the post-event logs said nothing.
Now one WARN per client error; 5xx excluded because Internal already logs its source
chain and the pool-exhaustion 503 logs at construction.

A DEAD FRONTEND SERVED A BLANK 502

`handle_errors 5xx` with an inline German page (the caddy service mounts only the
Caddyfile, so there is no volume to ship a static file through). Verified empirically
against this config, not from documentation: an upstream 404 through `reverse_proxy`
still arrives as untouched `application/json`, and only a dial failure renders the
page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a
real 404/429 arriving, and swallowing those would have been worse than the blank 502.

CONFIG CORRECTIONS FOR THE REAL HARDWARE

DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count.
Since migration 024 a feed page costs well under a millisecond, so connections are no
longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host.
COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already
serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was
sized against is unreachable; dropping to 1 would halve light-path throughput and push
more feed tiles onto full-size originals. README's sizing section rewritten for the
actual disk.

Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest,
svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate,
compose YAML parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:07:56 +02:00
MechaCat02
1d9fb11c7b fix: close the nine ways an unattended event loses photos or dies
Every one of these was found in the pre-event audit, verified against source, and
survives to production on the current main. Grouped by what actually goes wrong.

PHOTOS DISAPPEAR

* compression.rs no longer soft-deletes on a failed derivative. The guest got a
  201, watched the card appear, then watched it vanish — the row left v_feed,
  find_visible_media and BOTH keepsakes, while its bytes sat on disk for 14 days
  waiting for a cleanup nothing announced. No screen anywhere lists compression
  failures, so recovery meant hand-written SQL that also had to re-add the
  refunded quota. Now it does exactly what the ENOSPC arm beside it already did
  and documented as correct: keep the row, serve the original, retry on the next
  boot (bounded by derivative_attempts). `upload-deleted` is no longer emitted;
  `upload-processed` is, so the card re-renders instead of sitting on a
  placeholder.

* A 413 is now a reversible lock, so the blob survives. The quota moves — free
  disk falls, uploader count rises — so a guest goes over it having done nothing,
  and treating that as permanent meant a 400 MB video was pushed across cellular
  in full and THEN deleted from IndexedDB. Gone on both sides, and unrecoverable
  for an in-app camera capture that exists nowhere else.

* quota_limit_bytes gained a floor and a stable divisor. The ceiling used to
  decrease monotonically all evening; it now settles at max(uploaders,
  estimated_guest_count) — a config key that was seeded, validated in the admin
  whitelist, and read by no code at all. The floor is clamped to what the disk
  can actually back, so a full volume still yields zero rather than handing out
  an allowance it cannot honour.

* Because that floor gives up the aggregate guarantee the formula used to imply,
  uploads now check a hard 10 GB reserve first, independent of every quota
  toggle. postgres_data, media_data and exports_data share one filesystem: the
  end state was not a degraded feature, it was Postgres unable to write WAL.

THE ARCHIVE DISAPPEARS

* prune_superseded_archives runs only after the new generation lands. It ran
  before the preflight, reasoning the old archive was already unreachable — true
  of reachability, false of recoverability. An epoch is a value that can be
  rolled back; deleted bytes cannot. Any failed rebuild left the event with NO
  keepsake at all.

* The export preflight reserves the same 10 GB. `free < needed` authorised an
  export sized at exactly free, which ran for half an hour and landed the box at
  zero with the keepsake still unfinished.

THE APP DIES

* The feed reconcile re-reads the id set after its awaits instead of reusing one
  captured up to three round-trips earlier. The new-upload SSE handler prepends
  during exactly that window, so the row was both already present and absent from
  the stale set — prepended twice, and a duplicate key in a keyed {#each} throws
  in production, not just dev. The SSE handler and loadMore now dedupe too.

* Added routes/+error.svelte. Without it any uncaught error fell through to
  SvelteKit's unstyled English 500 with no reload control — inside a chromeless
  standalone PWA with no URL bar, for the rest of the evening.

THE OPERATOR IS LOCKED OUT

* admin_login verifies the password BEFORE charging the rate bucket, and a
  correct password is never throttled. The old order made this a denial of
  service against its own operator: every guest shares one NAT IP, the check ran
  first, so five requests a minute from any phone in the room kept the bucket
  full — and the escape hatch needed the admin session being blocked. A generous
  separate ceiling still bounds bcrypt CPU.

THE PROJECTOR DIES

* The preload budget is now strictly inside the dwell. At the 3s option the 4s
  budget could never land a commit on a slow uplink, so the wall froze on one
  photo while the queue drained silently behind it.

* The wake lock retries every 30s while visible, and the page says so on screen
  when the browser has no wake lock API. visibilitychange was the only retry
  trigger and a kiosk never changes visibility, so one refusal — iOS in Low Power
  Mode, say — was permanent.

* Caddy: /api/v1/upload/*/display joins the cacheable carve-out. The backend set
  max-age=300 on it and the blanket no-store silently replaced it, so a projector
  re-fetched a full-size JPEG per slide, ~2-4 GB over an evening on the uplink
  the guests are uploading over.

Also removes Upload::soft_delete, now unreferenced and an unscoped footgun next
to soft_delete_in_event.

Verified: 146/146 backend tests against a live Postgres, clippy clean, 51/51
vitest, svelte-check 0 errors, eslint clean, vite build, caddy validate, compose
YAML parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:32:52 +02:00
MechaCat02
61119be817 Merge branch 'fix/deploy-unattended-blockers' into main
Two independent lines of production hardening diverged at 7d0334b and attacked
overlapping problems. Neither was a superset, so this is a merge of substance
rather than a fast-forward: every conflict was resolved on the merits, and the
losing side's intent was re-checked against the winner rather than assumed.

MIGRATIONS. The branch's 021/022/023 collided with main's already-DEPLOYED
021_hashtag_counts_respect_bans and 022_client_upload_idempotency. Renumbered to
023/024/025 in a prior commit — main's versions are applied in production, so
their version numbers are immutable and the branch's had to move. Verified by
running the full sqlx::test suite, which applies the whole chain from scratch.

RESOLVED IN MAIN'S FAVOUR (the branch would have regressed these):
  * upload-queue.ts wholesale — the branch's copy has ZERO client_upload_id
    references, so taking it would have silently destroyed end-to-end upload
    idempotency, the one thing standing between a lost response and a duplicate
    photo charged twice against the guest's quota.
  * maintenance.rs supervisor — the branch replaced it with a bare tokio::spawn,
    where one panic silently stops session pruning, media reclaim, the temp
    sweep and both HashMap prunes, permanently and with no log line.
  * The decode-budget probe on spawn_blocking, not inline on the async runtime.
  * feed/+page.svelte's 8s debounce + jitter + max-wait + hidden-tab deferral,
    against the branch's naive 800ms — at 100 guests the branch's version walks
    straight into the per-user feed rate limit.
  * db.rs pool tuning, /uploaders, and the docker-compose deployment story.
  * ONE /health, still DB-backed. The branch's split (dependency-free liveness +
    DB-backed readiness) is defensible, but a constant-"ok" /health is the exact
    defect faea555 fixed and verified live, its motive (Caddy's boot gate) is
    already covered by app depends_on db: service_healthy, and the two handlers
    were the same SELECT 1 under two names.

TAKEN FROM THE BRANCH:
  * The large-PNG OOM guard and its bounded-retry counter (023). Together these
    turn a single upload that can OOM-kill a 1G container into a bounded failure
    instead of an infinite restart loop under `restart: unless-stopped`.
  * 024_feed_scalar_counts — the feed no longer aggregates the whole event per
    page. Pure SQL; column names, order and types are unchanged by design.
  * The admin-lockout fix: look the admin up BY ROLE, never by name. 025 also
    frees any guest already squatting on a reserved name.
  * PIN lockout tier ordering, bounded caption/hashtag reads, SSE ticket caps,
    PoolTimedOut -> 503 + Retry-After, and the ffmpeg stderr drain.
  * backfill_video_posters, which main lacked entirely.
  * TempFileGuard, plus sweep_orphan_originals wired into main's SUPERVISED loop
    (not the branch's bare one) — it reclaims final-named originals whose commit
    never happened, a class main's .tmp-only sweep structurally cannot see.
  * shouldAbortForStall, hand-ported into main's upload-queue.ts since that file
    was resolved to main. Widens the watchdog at loadend instead of disarming it,
    bounding a half-open socket at 2 minutes rather than handing the window to
    xhr.timeout (5-60 min) with the whole queue's `processing` latch held.

ALSO: RUST_LOG and EXPORT_PATH pinned in compose. The code fallback was
`debug` (a line per request, all night) and EXPORT_PATH was the one path with a
mount-shaped default that nothing validated.

Verified: cargo check --all-targets, cargo clippy (clean), 144/144 backend tests
against a live Postgres including upload_idempotency and upload_concurrency,
51/51 vitest, svelte-check 0 errors, eslint clean, vite build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:16:31 +02:00
MechaCat02
e1c689d1a7 chore(migrations): renumber 021-023 to 023-025 to clear the collision with main
main independently shipped 021_hashtag_counts_respect_bans and
022_client_upload_idempotency. sqlx::migrate! embeds ./migrations and refuses two
files per version, so these three had to move before the branch could merge at all.

Contents are untouched — only the version prefixes change. Renumbering (rather than
renumbering main's) is the safe direction: main is already deployed, so its 021 and
022 are applied in production and their versions are now immutable.
2026-08-08 21:04:58 +02:00
Fabian Hamm (Privat)
d5b4bf0ac1 fix(camera): get the upload sheet out from in front of the shutter button
Some checks failed
Audit / cargo audit (backend) (push) Failing after 12m25s
Audit / npm audit (frontend) (push) Failing after 23m21s
Checks / Backend — cargo test + clippy + fmt (push) Failing after 55s
Checks / Frontend — vitest + svelte-check (push) Failing after 42m2s
Checks / E2E — typecheck + lint (push) Failing after 20m58s
E2E / Playwright E2E (chromium + webkit) (push) Failing after 19m54s
E2E / Cross-UA smoke matrix (push) Failing after 20m27s
Tapping "Kamera" opened the viewfinder with the Galerie/Kamera sheet still sitting over
the bottom of it, covering the capture controls. The sheet is `fixed`, so it could not
be scrolled out of the way: the only route to the shutter was the phone's back button,
which is not a discoverable step and is one most guests would read as "the camera is
broken".

Two independent causes, both fixed, because either one alone leaves a gap.

The sheet never closed. It stays mounted for its translate-y animation and nothing told
it the camera had taken over, so it kept its panel, its backdrop and its `aria-modal`
while a full-screen overlay was up. `CameraCapture` now reports when its preview is
live and the sheet dismisses itself on that signal.

Deliberately on the preview, not on the tap. Closing when "Kamera" is pressed would
dismiss the sheet before we know the camera works at all — and it often does not: a
denied permission, no camera, or any non-secure context (where `navigator.mediaDevices`
is simply absent) all end at the error panel. Closing early would leave the guest
looking at that error with nothing behind it. Gated on `loadedmetadata`, the sheet is
still there when the camera fails, so "Schließen" returns them to where they were. The
signal is one-shot, because flipping the lens or switching photo/video re-acquires the
stream and re-announcing "ready" would ask the caller to redo a dismissal it has
already done.

And the stacking was ambiguous. Both elements were `z-50` and the sheet is rendered
after the camera, so it won on paint order. The overlay moves to `z-[60]` — the tier
the Toaster already occupies, so toasts still surface above the viewfinder on DOM
order. This is the part that holds regardless of timing: the controls are now reachable
during the permission prompt and on the error panel, before anything has been
dismissed.

Focus follows the same reasoning. When the camera closes the sheet, restoring focus
immediately would put it on the FAB *behind* the overlay, where a Tab could walk the
page underneath; it is restored when the overlay goes away instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 19:38:49 +02:00
Fabian Hamm (Privat)
0ae5a64e77 Merge branch 'chore/production-readiness'
Some checks failed
Audit / cargo audit (backend) (push) Failing after 12m13s
Audit / npm audit (frontend) (push) Successful in 44s
Checks / Backend — cargo test + clippy + fmt (push) Failing after 37s
Checks / Frontend — vitest + svelte-check (push) Successful in 10m49s
Checks / E2E — typecheck + lint (push) Failing after 42s
E2E / Playwright E2E (chromium + webkit) (push) Failing after 9m27s
E2E / Cross-UA smoke matrix (push) Failing after 6m24s
2026-08-03 18:37:31 +02:00
Fabian Hamm (Privat)
edc5f1f62c chore(export-viewer): rebuild the embedded bundle, and fix the lint ignores
The offline keepsake viewer had the same two filter defects as the app: typed
suggestions were capped so a matching tag could be unselectable, and the dropdown used
`onmousedown` with a backdrop that swallowed the selection. Rebuilt into
`backend/static/export-viewer/index.html`, which `include_dir!` embeds in the binary.

The eslint ignores were unanchored, so once the export-viewer's dependencies were
installed its nested `.svelte-kit` output was linted as source. Anchored with `**/`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:37:09 +02:00
Fabian Hamm (Privat)
46bb2e5174 docs: correct the claims that no longer match the code
Checked each against the implementation and fixed the doc, never the code:

- FEATURES claimed the ban modal offers a choice about hiding existing uploads. There
  is no such choice — a ban always hides. USER_JOURNEYS §9 was already right.
- FEATURES claimed hosts may demote other hosts. It is admin-only, enforced in the
  backend, and the two documents contradicted each other on it.
- FEATURES showed the quota widget as guest-facing; it is deliberately staff-only.
- The first-visit tour has six steps, not four.
- USER_JOURNEYS §12.7 said export downloads are rate-limited per IP. They are per USER
  — a materially different thing at a shared-NAT venue, where per-IP would have locked
  out the fourth guest to fetch their keepsake.
- §15 still described "Event verlassen"; that button is now Abmelden / Auf allen
  Geräten abmelden.
- §4, §13, §14, §16 and §18 were marked "(planned)" and have shipped.
- The lightbox row now describes what exists after this branch: prev/next controls,
  arrow keys and swipe.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:37:09 +02:00
Fabian Hamm (Privat)
2b1500e624 fix(ui): make the dashboards agree with each other and with what the code does
Host and admin implement the same four operations with independently written copy, and
admin was stale or wrong in every case. Its release button was always enabled and always
read "Galerie freigeben", so a second tap returned a 409; it showed no release state, no
keepsake progress, no failure reason, no rebuild, and never refreshed after releasing.
It now matches the host page.

Both dashboards subscribed to SSE and never opened the connection — `onSseEvent` only
registers a handler. Every subscription was inert, so the keepsake progress bar sat
frozen after a release and PIN requests appeared only on a manual refresh. It happened
to work when arriving straight from /feed, which connects, and /feed disconnects on
destroy, so navigating to the dashboard killed it again.

The unban confirm named neither of the two things a host most needs to know: unbanning
also restores ALL of that guest's previously hidden photos to the gallery, diashow and
export, and it retires and rebuilds a released keepsake, during which every guest's
download is briefly unavailable. The ban modal warns that uploads vanish; nothing said
they come back. Both now do, gated on the gallery actually being released.

"Event verlassen" implied the account was being deleted, then the dialog said the guest
could log back in. It calls `DELETE /session` — this device only, nothing deleted — so
it is "Abmelden" now. Gallery release now states it locks uploads and is reversible; PIN
reset states the guest is signed out on all devices.

The keepsake download failed silently: nothing inspected the iframe result and the
ticket POST always succeeded, so an over-limit tap did nothing at all. It now surfaces
the (newly visible) 429 and confirms the download started. `/export` rendered "Export
noch nicht verfügbar / Schau nach der Veranstaltung noch einmal vorbei" when the status
request had merely FAILED — telling a guest to come back after an event that already
happened. Both dashboards' error states gained a retry, which a host on a PWA with no
URL bar otherwise has no way to reach.

Modals were centred with no max-height, so on a short viewport the join PIN dialog
clipped equally top and bottom — potentially putting "Weiter zur Galerie" off-screen at
the moment a first-time guest must proceed. The ten moderation buttons were ~28px tall
side by side, on the screen where a mis-tap bans the wrong guest; they are 44px now.
Six German quotation marks paired the opening „ with an ASCII straight quote.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:36:57 +02:00
Fabian Hamm (Privat)
89ca819529 fix(diashow): never project a video, however its poster turned out
`merge` decided by whether a usable still existed, so the same clip was included or
dropped depending on whether ffmpeg happened to extract a poster — a video WITH a
thumbnail was queued and shown as a frozen frame, one without was skipped. Keyed on
the mime type instead: the projector shows stills only.

The test factory casts through `as unknown as FeedUpload`, so adding a field the queue
reads does not fail typechecking — it fails at runtime, which is how this surfaced as
ten broken tests rather than a compile error. The factory now carries `mime_type` and
the comment says why keeping it in step matters.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:36:57 +02:00
Fabian Hamm (Privat)
fffa2d556c fix(upload): stop the queue from wedging, and tell the guest when it fails
A STALLED UPLOAD BLOCKED EVERYTHING, FOREVER. The XHR set no timeout and had no stall
detection, so a half-open connection from an AP roam left the item `uploading`
indefinitely — which kept `processQueue`'s `processing` flag set, so the whole rest of
the queue stopped draining. The UI offered no control at all for an `uploading` item.
The guest saw "Wird hochgeladen 43%" all evening with four photos stuck behind it and
no button to press; the only escape was force-quitting the PWA, which nobody guesses.
Now: a watchdog aborts when no bytes move for 90s, disarmed on `loadend` so the server
may take its time storing a file it already has; a size-scaled timeout as a generous
backstop that will not kill slow-but-progressing LTE; and a cancel button.

FAILURES WERE INVISIBLE. `handleSubmit` navigates to /feed immediately, and the queue
component is mounted only on /upload — so a 5xx, a captive-portal error or an
uploads-locked 403 wrote a German message into an item that nothing ever rendered. The
guest believed the photo was uploading; it never appeared. Same for the documented
rate-limit countdown banner, which lives in that same unreachable component and is now
also rendered from the layout.

RETRIES WERE UNCAPPED. `requeueRetriable` flipped every errored item back to pending on
the `online` event AND on every `feed-delta` — i.e. every SSE reconnect — with no
attempt counter and no backoff. On a flapping network a large failing video was
re-uploaded from byte zero all evening, saturating the AP for everyone. Now a persisted
attempt count, exponential backoff and a cap of five.

INDEXEDDB COULD STRAND THE COMPOSER. `openDB` had no `blocked` handler, so a second tab
holding an older version made it never settle, and it rejects outright on iOS private
mode; `handleSubmit` had no try/catch and never reset `submitting`, so both buttons
stayed disabled reading "Wird hochgeladen…" permanently, with no error and nothing
queued. There is now a `blocked` handler plus a settle timeout, an in-memory fallback
so uploading still works when persistence is unavailable, and a `finally`.

A 401 during a background upload cleared the session without redirecting — and api.ts
documents exactly why that strands a guest: the nav and FAB are gated on
`isAuthenticated` so they vanish, route guards only run on mount, and a standalone PWA
has no URL bar. Three early-return paths wrote status only to memory and never to
IndexedDB, leaving blob-less error rows that could never be evicted and held the red
FAB badge lit all night.

A banned guest was offered the entire upload flow — FAB, camera, staging — and only the
POST 403'd, while the new read-only banner told them uploading was disabled. The sheet
now consults the ban, the layout subscribes to `user-hidden` so a live ban reaches the
UI instead of arriving as a stream of 403 toasts, and the banner clears
`env(safe-area-inset-bottom)` so the bottom nav stops covering it on notched iPhones.
The /upload submit bar gets the same inset — it sat in the home-indicator zone, where
the system swallows the first tap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:36:30 +02:00
Fabian Hamm (Privat)
51e55b1ace fix(feed): survive a bad network, and let the lightbox actually browse
Four things a guest on congested venue wifi would have hit, and one they would have
hit immediately.

A FAILED FEED LOAD CLAIMED THE GALLERY WAS EMPTY. `loadFeed` caught, toasted for five
seconds and left `uploads` empty, so the page fell through to "Noch keine Fotos. Tippe
auf den Kamera-Button unten!" — the most likely first impression at the party, and a
lie. There is now a distinct error state with "Erneut laden". Refreshes suppressed the
toast entirely, so pull-to-refresh and the "Neue Beiträge" pill failed in total
silence; they now report, and the pill survives its own failure instead of clearing
before the request.

THE FILTER-EMPTY STATE WAS DEAD CODE. With filtering server-side `displayUploads` is a
plain alias of `uploads`, so the grid's "Keine Treffer für die gewählten Filter." plus
its reset button sat behind an identical earlier branch and could never render — a
guest tapping a chip with no matches was told to go take a photo.

SSE COULD FREEZE THE FEED FOR THE WHOLE EVENING. Nothing in the feed ever refetched on
a timer; every update path was triggered exclusively by a stream event. Behind a proxy
that buffers `text/event-stream` `onopen` never fires, so the guest saw only the photos
that were on screen when they arrived; and a socket left half-open by an AP roam is
worse, because `connectSse` early-returns on a non-null EventSource and nothing ever
reconnects. A pure silence timer is not implementable — the backend sends keep-alives
as SSE comments, which the EventSource parser discards without dispatching — so
liveness is established on evidence instead: a jittered 60-120s `/feed/delta` backstop
that reconnects when a poll returns content the stream never delivered. The ticket
round-trip also seeds the delta cursor before the EventSource is created, so the
backstop has a `since` even if `onopen` never fires.

THE PILL COLLAPSED A DEEPLY-SCROLLED FEED to 20 items and dumped the guest at an
arbitrary scroll position — the exact yank the pill exists to avoid. It merges now.

The refresh debounce was 800ms + jitter, which during a burst is roughly one feed query
per client every two seconds; at 100 guests that approaches the 60/min per-user limit,
and the resulting 429s were swallowed by a bare `catch {}`, so the feed would simply
stop updating with no signal. Now 8s + jitter, coalescing, and skipped entirely while
the page is hidden.

Not one `<img>` in the app had an `onerror`. `pickMediaUrl` falls back to the original
whenever preview and thumbnail are null — i.e. for everything still compressing, which
during a burst is the top of the feed — so a 404 there rendered an empty grey box with
`alt=""`, not even a message. Each now retries once, then shows the placeholder.

The lightbox had no swipe, no prev/next and no arrow keys, so browsing 300 photos meant
closing and reopening the modal for every one — while FEATURES.md and USER_JOURNEYS
both claimed swipe shipped. It now has chevrons (44px, German aria-labels, hidden at
the ends), arrow keys, and horizontal swipe, with focus handed to the surviving control
so a disappearing chevron can't drop focus to `<body>`. Comment deletion was a ~14px
`✕` four pixels from the text that deleted permanently on one tap, while deleting a
POST two components away goes through a ConfirmSheet; it now matches.

`feed-filter.ts` and its test are deleted — with the server filtering, they were dead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:36:05 +02:00
Fabian Hamm (Privat)
87d01a8a26 fix(export): charge the download limit where the client can see the answer
The keepsake download is an iframe navigation, so its response is invisible to the
page. The rate limit was enforced inside the zip/html handler — i.e. inside that
navigation — while the ticket POST in front of it always returned 200. A guest over the
limit therefore tapped "Herunterladen" and absolutely nothing happened, forever, with
no explanation, on the one screen that is the emotional payoff of the whole app. With
the default of 3/day, ZIP + HTML costs 2 and one retry locks them out until tomorrow.

Minting is a normal `fetch`, so the limit moves there and the 429 reaches the user. The
limit is not weakened: tickets are single-use with a 30s TTL and can only be obtained
from that authenticated endpoint, so one mint is at most one download — and charging it
in both places would have cost every download two slots.

The message named the wrong timescale too. It shared the generic "warte kurz" wording
with the per-minute limiters, but this bucket is a DAY, so a guest was told to wait a
moment for something that could not work again until tomorrow.

Verified live: three mints succeed, the fourth returns 429 in German; raising
`export_rate_per_day` through the admin API takes effect on the next request with no
restart, and the HTML keepsake then downloads.

Also here, from the same pass:

- `looks_bcrypt` checks the SHAPE of ADMIN_PASSWORD_HASH, not just placeholder-ness. A
  hash corrupted by shell or Compose escaping is not a placeholder, so the app booted
  green, `/health` said ok, and every admin login 401'd — unrecoverable mid-event,
  because the Admin row is only created BY a successful admin login and promoting a
  host requires one.
- The rate limiter indexed `timestamps[0]` while holding its mutex, so a `max == 0`
  configuration panicked and poisoned the lock process-wide. Uses `first()`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:35:39 +02:00
Fabian Hamm (Privat)
496dba5a1f fix(feed): filter on the server, exactly, and keep banned uploads out of the chips
Filtering was split across two independent client-side states and applied to whatever
page 1 happened to hold, by caption SUBSTRING. So a tag chip selected in the list view
was silently still applied in the grid without being shown; a filter matched photos
whose caption merely contained the text; and anything past the first page was invisible
to it. Verified against the seeded data: `hashtag=tanz` returned 6 photos by substring,
1 by tag.

`FeedQuery` now carries `hashtag` (single, list view), `hashtags` (CSV, OR'd, grid
chips) and `uploader` (exact, AND'd), normalised through one function that trims,
strips `#`, lowercases and dedupes, and yields None when empty — so an empty filter
means "no filter", never "match nothing". The two SQL branches collapse into one with
`h.tag = ANY($4)`. Tag-OR plus tag+user-AND is a specified feature, not an accident:
`e2e/specs/03-feed/filter-search.spec.ts` and USER_JOURNEYS §8 pin it, which is why the
semantics moved to the server rather than being simplified away.

Tags travel as CSV safely because the backend restricts them to ASCII alphanumerics and
`_`; `uploader` stays a single exact parameter because a display name can contain a
comma.

New `GET /api/v1/uploaders` reads `v_feed`, so banned and hidden uploaders are excluded
for free.

Migration 021 gives `v_hashtag_counts` the same treatment. It counted every upload
regardless of the uploader's ban state, so banning a guest left their tags in the chip
list as ghost filters that lead to an empty feed. Verified: after banning the guest who
owned all six `tanz*` photos, the chips went 6 -> 0.

`?limit=-5` returned a 500 — only the upper bound was clamped, so Postgres was asked
for `LIMIT -4`. Clamped at both ends.

`is_banned` is added to `/me/context` so the client can show a read-only notice instead
of letting a banned guest discover the ban one 403 toast at a time. `add_comment` sorts
and dedupes hashtags on the normalised key, matching the upload path — the two disagreed,
which is a lock-ordering deadlock between concurrent upserts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:35:23 +02:00
Fabian Hamm (Privat)
1d0df3ebf6 feat(upload): make uploads idempotent so a lost response cannot duplicate a photo
The ordinary mobile failure, not an exotic one: the server receives the body,
validates it, commits the row — and the response is lost on the way back because the
guest walked out of range or the AP dropped the connection. The client sees a network
error with the blob still in hand and re-sends it, both when the guest taps "Erneut"
and automatically when the queue requeues on reconnect. Every attempt minted a fresh
`Uuid::new_v4()` server-side, so the same photo landed in the gallery two or three
times and was charged against the guest's storage quota each time.

The client already has a stable per-queue-item UUID, so it costs nothing to send.
Migration 022 adds `client_upload_id` with a partial unique index — partial so the
NULLs of every pre-022 upload, and of any caller that doesn't send one, keep working
untouched.

Two paths, because there are two races:

- Sequential retry: a lookup before the transaction finds the stored row, deletes the
  re-sent bytes and replays the original response as 200. The body has necessarily
  already been streamed, since the key arrives as a multipart field — re-sending is the
  client's cost and is already paid by the time we see it. What must be prevented is a
  second ROW.
- Concurrent retry: two attempts in flight at once. `ON CONFLICT DO NOTHING` returns no
  row to the loser, which abandons its transaction (quota increment included) and
  replays the winner. Letting the unique index raise instead would only surface after
  the transaction had aborted, as an opaque error the caller would have to string-match.

The replay reads live state rather than assuming a fresh row: a reconnect can be
minutes later, by which time the derivatives may exist and the photo may have been
liked. Every read there fails soft — the upload is already safely stored, so a sparser
response is fine and failing the request is not.

Verified live: the same photo sent three times returns 201, 200, 200 with one id, one
row, and the quota charged exactly once.

Also in this file: the two image-header probes at admission now run on `spawn_blocking`.
Both open the file and run the codec's header parse synchronously, and `#[tokio::main]`
gives two worker threads on a 2-vCPU box — so every upload stalled half the runtime's
request-serving capacity. Everything else that blocks here (image encode, bcrypt) was
already offloaded; this was the one that wasn't.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:35:05 +02:00
Fabian Hamm (Privat)
faea555967 fix(ops): make /health a readiness probe, and bound how long the pool waits
`/health` returned the literal string "ok" and touched nothing. Every request in this
app needs the database, so that answered a question nobody asked: the container
reported healthy while every request 500'd, and with no operator watching during the
event there was no signal at all. It now runs `SELECT 1` with a 2s timeout.

Verified live against the production stack: 200, stop Postgres, 503 "database timeout",
start Postgres, 200 again — with no app restart, because sqlx revalidates on acquire.

Deliberately NOT wired to automatic recovery. Compose's `restart: unless-stopped` does
not react to healthcheck state anyway, and an autoheal sidecar would be actively wrong
here: it would truncate every in-flight upload to "fix" an outage that, as the test
above shows, clears on its own. This is a diagnostic — including for the runbook's
event-day `curl`.

The pool had only `max_connections` set. Three additions:

- `acquire_timeout(5s)`. sqlx defaults to 30, so a DB blip parked every request AND all
  ~100 SSE session revalidations for half a minute before erroring — the app looked
  hung rather than degraded, and the backlog outlived the blip.
- `min_connections(2)`, so the first request after the setup-to-guests-arriving gap
  doesn't pay TCP + auth.
- `statement_timeout=15s` / `lock_timeout=5s` per connection. Without them a single
  pathological query holds a pool slot indefinitely and no client-side timeout can take
  it back, because the slot is only released when Postgres finishes.

Those two SETs are sent as two statements. `sqlx::query` uses the extended query
protocol, which permits exactly one per call — as `SET a; SET b` every new connection
failed, which surfaced as the pool never opening one and `create_pool` reporting a
connect timeout. Caught by booting against an empty database rather than a warm one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:34:46 +02:00
Fabian Hamm (Privat)
157499d493 fix(ops): reclaim abandoned upload temp files, and supervise the task that does
`stream_field_to_file` removes its `.tmp` on every error return, which covers
everything the handler can see. It cannot cover what actually happens at a party: the
client goes away — a phone sleeps, a guest walks out of range, the PWA is evicted
mid-video — and axum DROPS the handler future rather than returning an error, so no
cleanup runs at all. The shutdown backstop force-exits in-flight handlers for the same
net effect.

Nothing else reclaimed them. `cleanup_deleted_media` only visits rows with
`deleted_at`, and an abandoned upload never got a row; `export::sweep_orphan_temps` is
only ever pointed at the exports volume. `grep -rn read_dir src/` had three hits, all
in export.rs — the media tree was never read by anything.

So every abandonment stranded up to `max_video_size_mb` of unowned bytes permanently,
on the same 40 GB filesystem as `postgres_data`. Worse than a leak: the per-user quota
is computed from live free disk, so those bytes were also subtracted from what everyone
else was allowed to upload. An evening of flaky venue wifi could take the event down.

The threshold is on modification time, not creation time, which is what makes an hour
safe: a live upload is written to continuously so its mtime keeps advancing and it can
never age into the sweep no matter how slow the connection. The clock only starts once
the writer stops.

The periodic task is now supervised. It carries every piece of recurring hygiene in the
app — session pruning, media reclamation, this sweep, and the rate-limiter and
SSE-ticket maps — as a bare `tokio::spawn` with no retained handle, so a single panic
anywhere inside it stopped all five permanently and silently. No log line, no symptom
until the disk or a HashMap grew into one.

Tested: an hour-old temp is reclaimed, a temp still being written to is not (deleting
that one destroys a live upload), a committed `.jpg` is never touched, and a media tree
that does not exist yet is a silent no-op rather than an error logged 24 times a day.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:34:46 +02:00
Fabian Hamm (Privat)
2f952494c2 fix(media): stop a poster-frame failure from deleting the guest's video
Reproduced live, by accident, while smoke-testing on a machine with no ffmpeg: the
clip uploaded fine, returned 201, and roughly six seconds later had `deleted_at` set
and was gone from the feed.

The `Ok(None)` "this clip yields no frame" case was already handled — that fix landed
when sub-second clips were being destroyed. But the `?` on the call itself still routed
every OTHER failure into the same give-up path, which soft-deletes: ffmpeg missing from
the image, ffmpeg hanging on a truncated `.mov` and tripping the timeout, an ENOSPC on
`thumbnails/`, or a DB blip in `set_thumbnail_path`. None of those says anything about
the video, and `get_original` serves the file byte-for-byte, so a post that merely
lacks a poster is fully watchable. No failure in the video branch may fail the upload.

iPhone `.mov` is exactly the input most likely to trip it, and a wedding clip is not
retakeable.

ENOSPC gets its own classifier. It was the one failure the retry loop actively made
worse: a disk does not drain during six seconds of backoff, so all three attempts
failed identically while holding a compression permit that photos were queued behind —
and the give-up path then refunded the quota and soft-deleted the row while
deliberately KEEPING the original. That freed nothing, removed the photo seconds after
a 201, and handed the guest the allowance to upload it again into the same full disk.
Now: no retry, no refund, no delete. The row stays live and the photo is served from
its original, and `backfill_stale_derivatives` regenerates the derivatives on the next
start once there is room. `is_storage_full_error` has to look inside
`ImageError::IoError` as well as at bare io errors, because `image` wraps rather than
sources it and a plain chain walk would miss every derivative-write failure.

FFMPEG_TIMEOUT drops 120s -> 45s. It was never a budget for honest work — a poster from
a phone clip takes well under a second, and `-ss` before `-i` means even a 500 MB file
seeks rather than scans. It is the ceiling on how long a pathological input holds a
permit that guests' photos are waiting behind, so it should be as tight as it can be
without cutting off real work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:34:23 +02:00
Fabian Hamm (Privat)
43d37269b6 deploy: pull prebuilt images instead of building on the event server
The production compose still carried `build:` keys and no `image:` keys, so a
`git clone` onto the CX22 followed by `docker compose up -d` would have started a
fat-LTO release build of 427 crates on a 2-vCPU/4 GB box — the outcome the whole
build-on-the-Mac decision exists to avoid, reached silently because `pull` skips a
service it is told to build rather than failing.

Both services now pull `registry.mc02.dev/eventsnap/*:${EVENTSNAP_VERSION}` with the
`:?` form, so a missing tag fails the command instead of resolving to an empty one.
`docker-compose.build.yml` restores the `build:` keys for the workstation that
produces the images, from the same context paths.

Also here:

- `DOMAIN` gets the same `:?` guard. Blank did not fail — it produced `https://` for
  the frontend's ORIGIN and collapsed the Caddyfile's site block into a malformed
  global block, so the stack came up with no TLS and no site.
- `stop_grace_period: 20s` on the app. Docker's default stop timeout is 10s, exactly
  the app's own drain budget, so a redeploy could SIGKILL the process at the moment it
  was finishing — truncating the in-flight upload the graceful shutdown protects.
- `COMMENTS_ENABLED` is pinned "false" alongside MEDIA_PATH. It is a product decision
  for this event, and `.env.example` ships the generic `true`; pinning it means an
  operator who copies the example and edits only the secrets cannot ship comments on.
- The frontend runtime stage now copies the lockfile and uses `npm ci`. Without it the
  three `^`-ranged deps re-resolved at build time, so an image rebuilt days later could
  differ from the one that was tested. Image also drops 120 MB -> 65 MB.
- `docker-compose.dev.yml` told the operator that production had the same `$`-eating
  bug and to escape the hash as `$$` in `.env`. That is wrong and it breaks a working
  deployment: Compose uses single-quoted env_file values literally, and doubling
  produces a 74-character string `looks_bcrypt` rejects. Verified with `printenv`.

The runbook's rollback pointed at `v0.12.0`, which has 6 migrations against HEAD's 22
and was never built or pushed — running the emergency card's rollback line would have
crash-looped the app with `VersionMissing` during the event. §9 now has you tag one
build twice so the rollback target is bit-identical, and says plainly what that can and
cannot fix. Every `$DOMAIN` command gained the `set -a; . ./.env` it needs, the
down-migration psql commands are wrapped in `sh -c` so the container expands the
credentials rather than sending `-U ""`, and the advice to lower `max_video_size_mb`
is withdrawn: the client guard it was premised on does exist, but is pinned to a
compile-time constant, so lowering the DB value only moves failures later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:34:07 +02:00
MechaCat02
9b90929269 test(e2e): re-point the PIN lockout specs at the property that matters
Three specs asserted the OLD policy — that three wrong PINs lock an account —
which is exactly the behaviour the previous commit removed, because that threshold
sat below the per-(IP, name) throttle ceiling and so let any single IP lock any
guest whose display name is readable off the feed.

Rewritten to assert the distinction the fix introduces, which a status code alone
cannot show: both tiers answer 429, but only the account lock costs the VICTIM.
The new specs read the row via db.isPinLocked rather than the response, so:

- one IP hammering /recover is throttled and the account stays UNLOCKED;
- a distributed guesser (counter preloaded via db.setFailedPinAttempts, since no
  single source can reach the threshold any more) still trips the lock, and it
  holds even against the correct PIN;
- concurrent wrong PINs are all counted — the atomicity property the old parallel
  test was really about, now asserted on the counter instead of inferred from a
  429 that the throttle could equally have produced.

The UI spec asserts the user-visible half: after four wrong PINs Dave can still
get into his own account. It also now types the PIN digit by digit rather than
filling and clicking, because the 4th digit auto-submits (pin-auto-submit.spec.ts)
and doing both raced the button's disabled state.

The adversarial spec enables rate_limits_enabled for its own run — it is off by
default in this environment, so without that the throttle tier would silently not
be exercised — and restores it in afterEach so it cannot leak into other specs
sharing the stack.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 23:58:17 +02:00
MechaCat02
117e67fa80 fix(auth,upload): close the admin lockout and four unbounded-input paths
ADMIN LOCKOUT. admin_login looked its user up BY NAME. Migration 007 makes
display_name unique per event case-insensitively and join had no reserved-name
guard, so any guest joining as "admin"/"Admin"/"ADMIN" before the operator's first
login made find(role == Admin) miss, the fallback create("Admin") violate that
index, and `?` return a 500 — permanently, with no in-app recovery. Moderation,
config and gallery release all gone; the fix was hand-editing the database.

The root cause is the lookup key, not the creation. The name was never the
identity. User::find_admin_for_event resolves by role, which makes the whole class
of name collisions irrelevant — including the homoglyph bypasses of the new
reserved-name list, which is now defence in depth rather than the control.

Promoting the squatting row would be the obvious fix and is a serious mistake: it
carries a recovery_pin_hash the guest knows, so it would hand them the admin
dashboard via /recover, permanently, through a path needing no password. A
separate row under a fallback name is worse UX and much better security. Verified
against the real schema — the guest keeps their uploads, PIN and session under a
freed name, and the role lookup then finds exactly one admin.

Second, independent bug in that block: create() followed by a SEPARATE UPDATE ...
SET role = 'admin' manufactures the same poisoned state if anything fails between
them. Collapsed into create_with_role.

UNBOUNDED INPUTS — one root cause, four places: validation ran after the
allocation.
- upload caption/hashtags used Field::text(), which buffers the whole field, on
  the one route whose DefaultBodyLimit is 576 MiB — so 576 MiB of heap per
  concurrent request in a 1 GiB container, with the length check running
  afterwards on a string already built. Now refused mid-read.
- the hashtag CSV was never length-checked at all and was upserted tag by tag
  INSIDE the commit transaction, which holds FOR SHARE on the event row — one
  request could stall every other upload behind tens of thousands of round trips.
  Capped at 30 tags of <=50 chars.
- /recover and /recover/request built rate-limiter keys by format!() from an
  unvalidated, unbounded display name, retained up to 24h in a map pruned hourly:
  the limiter itself became the memory-exhaustion primitive it exists to prevent.
  join validated first; that check is now shared by all three. /recover/request
  also had no per-IP ceiling at all — /join got one in 017, /recover in 019, and
  019's own comment describes exactly this attack. It returns 204 rather than 400
  on a bad name, because a 400 would be a new signal on an endpoint whose contract
  is that it cannot enumerate guests.
- the SSE ticket store had no size cap, no per-session cap and no rate limit on
  its endpoint, while prune ran hourly against a 30s TTL. Now pruned on issue,
  capped, and rate-limited. At capacity it REFUSES rather than evicting a
  stranger's ticket — evicting would let one client deny SSE to the venue. Not
  one-ticket-per-session either: two tabs open their EventSources concurrently.

PATCH /upload/{id} had no rate limit, no validation, and called
invalidate_and_arm unconditionally — outside both `if let Some` guards. So
PATCH {} bumped export_epoch and armed a fresh pair of full-gallery export workers
every call; REGEN_DEBOUNCE bounds the rate of that, not the total work, so a guest
could keep the keepsake permanently un-downloadable. All three fixed. The
validation also resolves a divergence: upload normalised tags while edit stored
them raw, so #Party via edit and party via upload became two hashtag rows.

PIN LOCKOUT was an ordering bug before a policy one: the account-lock threshold
(3) sat BELOW the per-(IP, name) ceiling (5), so three requests from one IP locked
any guest whose name is on the feed, every 15 minutes, forever. The tier meant to
protect a guest was the cheapest way to attack them. Ceiling drops to 4, threshold
rises to 12, so locking a victim now needs at least three distinct sources.
Brute-force cost is unchanged — 48 attempts/hour means 10k PINs still take ~208h
regardless of IP count — and increment_failed_pin now decays the streak after 15
minutes, since the counter previously only cleared on success and honest typos
accumulated across days. Both invariants are pinned by tests rather than comments.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 23:09:20 +02:00
MechaCat02
f275de5c8f perf(feed): stop every feed page from aggregating the whole event
v_feed computed like_count/comment_count with LEFT JOINs and a GROUP BY. Postgres
CAN push `event_id = $1` and the keyset predicate through the view — verified with
EXPLAIN, it uses idx_upload_event_created_id — but it CANNOT push ORDER BY ...
LIMIT across a GroupAggregate. So every request aggregated every upload in the
event, times its likes and comments, and only then sorted and took 21 rows. Cost
grew with the event, not with the page.

Measured on a throwaway database seeded to a real reception (1000 uploads, 100
guests, 27k likes, 10k comments), same query, same data:

  before   GroupAggregate (actual rows=1001) -> Sort -> Limit    449 ms
  after    Index Scan (actual rows=21) -> Limit -> SubPlans      0.58 ms

migration 022 replaces the joins with correlated scalar subqueries, which puts the
counts ABOVE the Limit so they run 21 times instead of 1001. Exactly equivalent,
not merely close: "like" is keyed (upload_id, user_id) so COUNT(DISTINCT user_id)
== count(*), comment.id is the PK so COUNT(DISTINCT c.id) == count(*), and the
GROUP BY was on u.id so it was already one row per upload. Column names, order and
types are unchanged, so no Rust changes. No new index needed — idx_like_upload and
idx_comment_upload already match the subqueries.

Note the existing load harness cannot see any of this: e2e/loadtest/driver.mjs
creates no likes and no comments, so the expensive path had never been exercised.

The amplifier, feed/+page.svelte: every open feed subscribed to `upload-processed`
and refetched page 1 — the most expensive page — so ~100 open feeds each fired one
per completed upload. Now gated on whether this client actually shows the card
that changed, and the debounce is jittered, because a fixed delay just moves a
simultaneous herd 800 ms later. Nothing is lost by skipping: a client without the
card also missed its `new-upload`, and the reconnect `feed-delta` already
schedules a refresh.

Load shedding, because the above reduces the risk rather than removing it: db.rs
set no acquire_timeout, so sqlx's 30 s default applied — longer than the
frontend's own 20 s fetch timeout, meaning the browser gave up while the server
kept holding the slot and the work was done for nobody. Now 5 s, and PoolTimedOut
maps to 503 + Retry-After instead of a generic 500. That mattered because the
upload queue classifies 5xx as transient and retries: a 500 sent the retries
straight back into the saturated pool with nothing to pace them. PoolClosed stays
Internal — it only occurs during shutdown, where a 503 would invite a retry
against a server that is going away.

The Retry-After extraction in into_response matches on variants, so unlike
message() a missing arm is not a compile error — it would silently drop the
header. Pinned by a test covering both retry-carrying variants.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 23:01:42 +02:00
MechaCat02
3f0f9c098b fix(upload-queue): rehydrate the persisted queue app-wide, not only on /upload
loadQueue() had exactly one call site in the entire frontend: the /upload route's
onMount. So after a reload, an iOS tab discard or a PWA relaunch, staged photos
sat in IndexedDB while the badge read 0 and nothing sent them — unless the guest
happened to navigate back to /upload, which they have no reason to do, having
already been shown a success. The photo never leaves the phone and the guest is
never told.

The root cause was narrower than "loadQueue isn't called enough".
requeueRetriable() read IndexedDB but only .map()'d over whatever the in-memory
store already held, so it could reset statuses and never ADD an entry — and
processQueue reads only that store. That is why the `online` listener and the SSE
resume hooks, which both call it, could not recover a cold start either. It now
REBUILDS the store from IndexedDB, which makes all three resume paths work.

Rebuilding needs one guard: entryToQueueItem downgrades `uploading` to `pending`
with progress 0, and this runs on every `online` event and every SSE reconnect,
so a blind rebuild would visibly reset the progress bar of a request still on the
wire. In-flight items are carried over by id.

Hydration is module-level, SSR-guarded and idempotent, re-armed via
onSetAuth/onClearAuth because login is a client-side goto() — no module
re-import, no onMount re-run — so a hydration that no-oped for lack of a token
gets a second chance. Module level rather than a layout onMount because
+layout.svelte already imports this module on every entry point, it matches the
file's own bindOnline()/bindSse() pattern, and a store owning its own persistence
keeps the layout free of a concern it cannot test. auth.ts does not import this
module, so no cycle.

The burst-queue e2e test no longer navigates to /upload after its reload — it now
asserts the resume happens wherever the reload lands, which is the actual
regression guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:58:26 +02:00
MechaCat02
e52b2f1cd1 fix(upload): reclaim the bytes of uploads that never finish
Reclaim was a dozen explicit remove_file calls on the handler's return paths.
That covers every way the handler can FINISH and none of the ways it can simply
STOP. When a client disconnects mid-body — a phone leaving wifi, iOS killing a
backgrounded PWA, the user hitting back — axum drops the handler future at an
.await inside field.chunk() and no return path runs at all. The partial file then
survives forever: it has no upload row, so cleanup_deleted_media (row-driven)
can never see it, and no sweeper covered the originals directory. The triggers
are routine rather than adversarial, and the client keeps the blob and auto-
retries on every `online` event, so one large video over bad wifi leaves several
copies.

Those bytes are also invisible to the quota while still consuming the free disk
that compute_storage_quota divides among guests — so orphans silently shrink
every guest's ceiling while the admin widget under-reports. All three volumes
share one filesystem; the end state is Postgres unable to write WAL.

TempFileGuard is an RAII guard, because dropping the future is exactly what runs
Drop — it is the only construct that survives cancellation. Armed before the file
can exist, disarmed only after tx.commit() succeeds. The twelve explicit cleanups
are deleted so one owner holds the rule.

The subtler half is the rename. It happens BEFORE the commit, so between them the
file exists under its final name with no row pointing at it — an orphan that
looks legitimate. The guard is RETARGETED there rather than disarmed, and the
retarget sits on the same poll as the rename with no .await between, which is
what makes that window uncancellable.

sweep_orphan_originals is the backstop for the process that was killed, where no
Drop can run at all. Hourly, alongside the existing sweeps: .tmp files past the
window go unconditionally (a .tmp never has a row by construction), other files
are batched 500 at a time through a single NOT EXISTS query.

Two things that look like oversights and are not, both commented in place:
- the 6h window is what makes the sweep safe against the rename-before-commit
  ordering, since a committing upload is briefly indistinguishable from an
  orphan. It must not be shortened to speed up a test.
- the NOT EXISTS deliberately does NOT filter deleted_at IS NULL. A soft-deleted
  row still points at its file during its retention window, and reclaiming that
  is cleanup_deleted_media's job; filtering here would race the two sweeps and
  destroy the files the recovery window exists to preserve.

Verified against the real schema: given a live original, a soft-deleted one and a
true orphan, the query returns only the orphan.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:56:26 +02:00
MechaCat02
5969ec74ea fix(compression): bound derivative retries so one bad upload can't loop forever
The OOM in the previous commit was survivable; what made it an outage was that it
repeated. The upload row is committed before compression starts, derivatives_rev
defaults to 0, and set_derivatives_rev only runs on success — so a row whose
processing killed the container survived at rev 0, and backfill_stale_derivatives
(called unconditionally at every boot) re-selected it and re-ran the identical
workload. With restart: unless-stopped that is an infinite kill loop, and every
cycle also drops every SSE stream and truncates every in-flight upload.

Verified end to end against the real schema in a scratch database: with the new
guard the backfill selects the row on boots 1-3 and zero rows from boot 4 on,
and a later success resets the counter.

migration 021 adds derivative_attempts and derivative_last_error.

The counter is incremented WRITE-AHEAD, before the work is attempted. This is the
whole design: the failure being bounded is a cgroup SIGKILL, so no Err is
returned, no error handler runs and no Drop fires. A counter bumped in a failure
path increments zero times per crash and the loop would be unchanged.
set_derivatives_rev clears it, so success is the only reset and both the live
path and the backfill get it without a new call site to forget.

Also in the backfill:
- one task walking the rows sequentially instead of one task per row. A large
  backlog used to spawn thousands of tasks, each holding a pool handle and
  queueing on the same two permits, competing with live uploads for a whole boot.
- LIMIT 200 per boot, and original_path <> '' replacing an IS NOT NULL that was
  dead (the column is NOT NULL; cleanup_deleted_media blanks it instead).
- a once-per-boot error log naming how many uploads have given up. Without it the
  give-up is invisible — the loop stops, which is the point, but the photos keep
  a stale derivative forever with nothing to notice.

Adds backfill_video_posters for the mirror-image gap: a video interrupted by a
restart has its compression_status flipped processing -> failed by
startup_recovery and is never re-enqueued, so thumbnail_path stays NULL for the
rest of the event while the clip itself plays fine. It shares the same attempt
budget, which means a genuinely posterless sub-second clip (Live Photo, mis-tap)
stops being re-ffmpeg'd after three boots. That is intended, not a bug to fix
later — Ok(false) is a normal permanent outcome there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:52:48 +02:00
MechaCat02
a4bac03628 fix(compression): stop a single large PNG from OOM-killing the container
An 8000x8000 RGBA PNG passes admission — 256,000,000 bytes is just under the
256 MiB max_alloc, and smooth content is under 3 MB on disk, far below any size
cap. Processing it peaked at ~1250 MiB inside a 1 GiB cgroup. Measured, not
argued: the new (ignored) test builds exactly that image and reads VmHWM around
the pipeline, resetting the watermark via /proc/self/clear_refs so the number
covers only the code under test.

Three independent causes, all of which had to go:

1. The decode outlived everything. `resize` takes &self, and the no-downscale
   arm bound `img` into `display`, so the ~244 MiB buffer was still alive when
   oxipng ran — and oxipng decodes the PNG *again*, holding a full-size buffer
   per filter trial. The decode now lives in a block that yields the display
   derivative; the else arm moves `img` out, which is what makes "the block's
   value is the only survivor" true in both arms.

2. oxipng was unbounded in every dimension: preset 2 with timeout: None, and the
   default features pull in rayon, which evaluates filter trials concurrently
   with a full-size buffer each and has no Options knob to cap it. Now gated at
   8 MP, given a 20 s timeout, and built with default-features = false so
   oxipng's own sequential shim is used. "filetime" is kept — without it
   preserve_attrs silently no-ops. Dropping "binary" also stops compiling
   clap/glob/env_logger (a CLI's deps) into the server image.

3. Even with those fixed it still measured 516 MiB, and compression_concurrency
   defaults to 2 — so two guests uploading big photos at once was another OOM,
   1032 MiB against a 1 GiB limit. The cost is dominated by image's Lanczos3
   resize, which accumulates in f32: the intermediate is new_width * old_height
   * 16 bytes, i.e. 262 MiB for this image — larger than the decode itself, and
   invisible to max_alloc. Two changes: the preview now derives from the 2048px
   display instead of re-resizing the original (one full-size pass, not two),
   and a job whose header-estimated peak exceeds 150 MiB takes an exclusive
   permit so two giants can never overlap. Ordinary photos (a 12 MP JPEG
   estimates ~50 MiB) never touch that permit, so throughput is unchanged for
   everything except the case that must not run in parallel.

Chaining 8000 -> 2048 -> 800 for the preview is not a quality trade: a staged
Lanczos3 downscale is standard for large ratios and is visually
indistinguishable at 800px.

The blocking half is now a free function so its memory behaviour is testable —
the fix is a scoping property a future edit could silently undo.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:42:03 +02:00
MechaCat02
6fd75adb27 fix(ops): bound logs, drain ffmpeg's stderr, and stop three silent hangs
Six independent operational defects, none of which needed a new feature to fix.

Log rotation. Docker's json-file driver is unbounded by default, and those files
land on the HOST filesystem — outside every deploy.resources.limits in the compose
file, and on the same disk as postgres_data and media_data. A full disk stops
Postgres writing WAL, which takes the event down. Capped at 10m x 3 per service.

Log level. RUST_LOG was set in neither .env.example nor docker-compose.yml, so the
code fallback WAS the production level — and it was `debug`, with tower_http=debug
emitting a line per request and per response into that unrotated file. Now info,
with tower_http=warn to state that those spans are diagnostics, not an access log.

ffmpeg pipe deadlock. run_ffmpeg piped stdout and stderr and then called wait(),
which drains neither. Once the ~64 KiB pipe buffer filled, ffmpeg blocked writing
and wait() never returned — burning the full 120s timeout, twice per seek position,
three times per compression attempt. And the timeout is an Err, so the end state was
a soft-deleted upload: a guest's playable video destroyed by a poster-frame failure.
Now stdout is null (nothing ever read it) and stderr is drained by wait_with_output,
whose tail is logged on a non-zero exit. Note wait_with_output consumes the child, so
the old kill-on-timeout is gone; kill_on_drop(true) already covers it.

Readiness probe. /health never touched the pool, so the disk-full endgame above
stayed green all the way down. Adds /health/ready (SELECT 1 under 2s) as a SECOND
route — the compose healthcheck deliberately keeps pointing at /health, because
caddy gates its startup on it and a DB-dependent probe would turn a Postgres blip
into the reverse proxy refusing to start.

api.ts request timeout. The abort timer was cleared in a finally around fetch(),
which resolves on the response HEAD — leaving res.text() uncovered and no longer
abortable. An upstream that sends headers then stalls the body hung the call
forever. The timer now lives until the body is read, including the 204 path (which
otherwise leaked a live 20s timer per no-content request).

Upload XHR watchdog. The XHR had no timeout while processQueue held the
isProcessing latch across it; on a half-open socket neither error nor abort ever
fires, so the latch pinned and the queue wedged. Bounds SILENCE rather than total
duration — a 500 MB video over a venue uplink legitimately runs 30+ minutes while
making steady progress. Rejects as NetworkError, which is already the retryable
branch, so a stalled upload now recovers like any network blip.

IndexedDB failures. addToQueue called getDb() unguarded and handleSubmit had no
catch, so a private-mode refusal or a QuotaExceededError on a large blob left a
permanent "Wird hochgeladen…" spinner, no toast, and — for an in-app camera
capture — the only copy of the photo gone. Now reported as 'failed', which keeps
the staged files on screen and stays on the page.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:25:23 +02:00
fabi
7d0334bf22 Merge branch 'fix/video-thumbnail-seek'
Some checks failed
Checks / Backend — cargo test + clippy + fmt (push) Failing after 45s
Checks / Frontend — vitest + svelte-check (push) Failing after 6m31s
Checks / E2E — typecheck + lint (push) Failing after 47s
E2E / Playwright E2E (chromium + webkit) (push) Failing after 9m31s
E2E / Cross-UA smoke matrix (push) Failing after 5m36s
Audit / cargo audit (backend) (push) Failing after 11m42s
Audit / npm audit (frontend) (push) Successful in 43s
2026-07-30 19:40:27 +02:00
fabi
402215d405 fix(video): extract a real poster frame, and stop claiming one that isn't there
Both the compression worker and the HTML export ran the same invocation:

    ffmpeg -i <src> -vframes 1 -ss 00:00:01 -vf scale=… -y <out>

`-ss` AFTER `-i` is an output-side seek. Against a clip of a second or less ffmpeg
exits 0 and writes nothing, and both call sites gated on the exit status:

- the worker wrote `thumbnail_path` and logged "thumbnail generated" for a file that
  was never created, so GET /upload/{id}/thumbnail 404s in the live feed;
- the export listed media/<id>_thumb.jpg in data.json while the ZIP writer skipped the
  unopenable file, so the keepsake drew a broken image tile.

Any clip at or under a second, which phones produce constantly — mis-taps, Live
Photos, boomerangs. Not data loss; the .mp4 is in both archives and plays. Every
server-side signal stayed green.

New `services/video.rs` owns the extraction for both callers, mirroring the imaging.rs
precedent (created for the same duplication, and it paid off when the max_alloc fix
landed in both workers at once). Three changes in it:

- `-ss` before `-i`, an input-side seek. NOT sufficient alone: verified against the
  production image, seeking to 1s in a 1.000s clip is still past the last frame and
  still exits 0 with no file. The 0s fallback is what actually fixes this, and 1s is
  tried first only because an opening frame makes a poor poster.
- Verify the artifact, not the exit status. This is the check both sites were missing.
- Carry compression.rs's 120s timeout. export.rs had NONE — a hung ffmpeg there would
  strand the job at `running` and the keepsake would never complete.

The worker's call used `?`. Tightening the check without also making a missing poster
non-fatal would have been far worse than the bug: every sub-second clip would fail
compression, exhaust its retries and be soft-deleted. It now logs a warning and leaves
`thumbnail_path` NULL, which FeedListCard, VirtualFeed and LightboxModal already
handle.

The export now sets `thumb: ""` and skips the manifest entry when there is no poster —
and does the same for the IMAGE branch, whose decode failure left the identical
dangling reference. No viewer change was needed: +page.svelte already guards
`{#if post.media.thumb}` and falls back to a video tile with a play glyph. The comment
claiming "viewer handles missing thumbs gracefully" was true about the viewer and false
about what the backend sent — the guard never fired because the string was never empty.

e2e/specs/06-export/export-video.spec.ts had DOCUMENTED this as intended behaviour
("the fixture clip is <1s, so ffmpeg extracts no thumbnail frame — but exits 0 …
that's the intended shape here"). sample.mp4 is exactly 1.000s, so every video test in
the suite ran at that boundary and none ever fetched the poster. That comment is now
corrected to say what it actually was.

Tests: 2 unit; a new sample-5s.mp4 fixture so the ordinary first-seek path is covered
at all; video-playback now FETCHES the poster rather than asserting the attribute (the
one extra request that nine rounds of green never made); a new spec covering both
fixtures plus the mirror that a posterless video still uploads and plays; and a
keepsake spec asserting every <img> in the opened viewer resolves — naturalWidth === 0
is exactly the broken-tile case, whatever produced it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 19:40:27 +02:00
fabi
2b313e67e0 Merge branch 'fix/track-e2e-fixtures' 2026-07-30 19:39:59 +02:00
fabi
2551c25436 fix(e2e): track the test fixtures the suite reads
`.gitignore` carried an unanchored `media/`, commented "Media uploads (mounted volume in
production)". Production media lives in the `media_data` Docker volume and never
appears in the working tree, so that rule guarded nothing real — but unanchored it
matches a directory of that name at ANY depth, and the only one in the repo is
`e2e/fixtures/media/`.

Its entire practical effect was to keep every E2E fixture untracked. A fresh clone got
the specs and none of the images or videos they read, and
`.github/workflows/e2e.yml` does a plain `actions/checkout` and generates nothing — so
the committed CI job could not have run the upload, video, EXIF, quota, oversized-image
or export suites at all. It only ever looked green locally, where the fixtures survive
as untracked leftovers from whoever first created them.

The `origin` remote is git.mc02.dev rather than GitHub, so those workflows have most
likely never executed against this repo, which is consistent with nobody noticing. That
also makes this a live blocker for the standing "rotate the token and push" item: the
first time CI runs, it fails on ENOENT in a way that looks like a broken suite rather
than a missing file.

Anchors the pattern to `/media/` and commits the six fixtures (672 KB total). Verified
`e2e/fixtures/media` is the ONLY directory the old pattern matched, so nothing else
changes visibility.

Found while adding `sample-5s.mp4` for the video-poster work: the new fixture would
have been invisible to every other machine, which is how the existing ones got here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 19:39:59 +02:00
fabi
d51c6b8c4b Merge branch 'fix/postgres-password-deploy-trap' 2026-07-30 19:28:34 +02:00
fabi
3bcb7c6a76 fix(deploy): close the POSTGRES_PASSWORD trap that broke a fresh deploy either way
README step 2 said to set "DOMAIN, JWT_SECRET, ADMIN_PASSWORD_HASH, EVENT_NAME, etc."
-- POSTGRES_PASSWORD was not in that list. .env.example said to set it and keep it in
sync with DATABASE_URL. The two documents disagreed, and both readings ended badly.

Branch A, README verbatim: the stack came up GREEN and healthy on
CHANGE_ME_use_a_strong_password -- a database credential published in the public repo.
The production secret guard covered JWT_SECRET and ADMIN_PASSWORD_HASH, and nothing
anywhere looked at the Postgres password.

Branch B, .env.example verbatim: a permanent restart loop, "password authentication
failed for user eventsnap".

The guard's own design made Branch B near-certain. It stops the APP on the first
`docker compose up -d` -- but not the `db` service in that same command, which
initialises its data directory and bakes in whatever password was in .env at that
moment. POSTGRES_PASSWORD is honoured ONLY at initdb. So the intended recovery -- see
the refusal, fix your secrets, boot again -- was exactly the sequence that broke it.
Nothing in the error named the cause, and the remedy (`down -v`) is both unguessable
and the one command you must never run once real data exists.

Three changes, which have to ship together: the guard alone would just move operators
out of Branch A and into Branch B.

- The guard now rejects a placeholder DATABASE_URL in production (the password rides
  in that URL, which is what the app actually reads). Branch A can no longer boot.
- It reports EVERY unset secret in one message instead of returning on the first.
  Fixing two secrets used to cost two boot cycles, on a stack where Caddy waits on the
  unhealthy app throughout, and each avoidable cycle is another chance to reach for -v.
- A 28P01 handler in db.rs turns the unguessable failure into a self-explaining one:
  it names the initdb semantics, gives `down -v` with an explicit "deletes db + media +
  exports, no undo", and gives the ALTER ROLE alternative for when data already exists.

Docs: step 2 now names POSTGRES_PASSWORD and says every secret must be set BEFORE the
first up; the troubleshooting block covers the auth-failure loop and both remedies.

Also replaces htpasswd (apache2-utils -- not on a stock VPS) with
`docker run --rm caddy:2-alpine caddy hash-password`, an image the stack already pulls,
in README, .env.example and the guard's own message. Verified the output ($2a$14)
against the shipped $2y$12 example.

Verified on a real Postgres, not just in tests: Branch A refuses and names DATABASE_URL;
all three placeholders report in one boot; a volume initialised with one password and
connected to with another prints the diagnostic; an unrelated connect failure (dead
port) stays silent. 8 unit tests, including that non-prod ignores all of it so the e2e
stack is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 19:28:34 +02:00
fabi
08d92b7531 fix(audit-7): keepsake viewer integrity, readable archives, quota_tolerance floor
Seventh round, both defects in the keepsake — the one artifact that leaves the
system entirely, and so the one where a green server-side signal carries no
information at all.

Squashed from 2 commits, original messages preserved below.

──────── fix(export): stop a caption bricking the viewer, and ship readable archives

Two defects in the keepsake, both silent server-side and both only visible by
extracting the real artifact and trying to use it.

1. A CAPTION COULD BRICK THE VIEWER.

The viewer's data is inlined as `<script>window.__EXPORT_DATA__={…}</script>` -- it has
to be, since guests open index.html over file:// where fetching a sibling data.json is
blocked. The escape was `</` -> `<\/`. Against XSS that holds; I fired
`</script><img src=x onerror=…>` through a real Chromium parser and it round-trips
inert.

It does not stop the caption steering the HTML TOKENIZER. `<!--<script` with no later
`-->` drives the parser into script-data-double-escaped state, where the template's own
`</script>` only steps back to script-data-escaped instead of closing the element.
Everything after it -- including the viewer bundle -- is swallowed as script data.
Nothing executes and nothing leaks: `__EXPORT_DATA__` is simply never assigned and the
keepsake opens blank. A denial of the deliverable, not an XSS.

Reproduced in Chromium before changing anything, and the near-miss is worth recording:
`<!--<script>alert(1)</script>-->` comes back CLEAN, because the trailing `-->` returns
the parser to script-data state. A probe using the terminated form quietly repairs the
thing it is testing for.

Fix: escape every `<` as `<`, not just `</`. `<` never appears in JSON structural
syntax -- only inside string values -- so a global replace is sound, and one rule covers
`</script`, `<!--` and `<script` together. That is the point: the old escape was named
for the single case it handled. Only the INLINED copy is escaped; data.json is written
separately, in no HTML context, and stays literal.

2. EVERY ENTRY IN BOTH ARCHIVES WAS STORED MODE 0000.

`ZipEntryBuilder::new` leaves the external file attribute at zero and async_zip's host
compatibility defaults to Unix, so `unzip -Z` showed `?---------` on every line of both
Gallery.zip and Memories.zip. Windows Explorer ignores Unix modes, which is why this
survived; on Linux and macOS `unzip` faithfully applies what the archive asks for and
the guest gets a folder of photos none of which they can open.

Unconditional -- every keepsake ever produced, no hostile input required -- and
invisible server-side: the export succeeds, the ZIP is well-formed, the job writes
`done`, /export/status is green.

Found by accident. The browser test for defect 1 failed with ERR_ACCESS_DENIED on
file://, which looked exactly like a Playwright sandbox quirk; I twice "worked around"
it (fresh context, then a separately launched browser) before checking the extracted
files and finding mode 000. The workaround was suppressing a real bug. Both workarounds
are gone -- the ordinary `page` fixture loads the archive fine now.

Fix: all six ZipEntryBuilder sites route through one `keepsake_entry` helper stamping
`S_IFREG | 0644`, so the mode cannot be forgotten at a call site.

Tests: 3 unit (no `<` survives; the payload still decodes to the original value, because
this is a transport encoding and not a sanitiser; a clean payload is untouched) and 2
e2e that release for real, download the real archives, and check them from outside the
app -- one opening index.html over file:// in Chromium and asserting the viewer booted,
the captions came back verbatim and nothing executed; one asserting every stored mode
and every extracted file is readable. Both assertions verified to FAIL against the
pre-fix artifacts.

──────── fix(admin): reject quota_tolerance = 0 instead of silently blocking every upload

Zero is inside the documented 0–1 range and catastrophic. The per-user limit is
`free_disk * tolerance / active_uploaders`, so a tolerance of 0 makes every limit 0 and
refuses EVERY upload -- mid-event, with "Du hast dein Upload-Limit für dieses Event
erreicht", an error naming the wrong cause entirely. An admin reaching for an off-switch
wants `storage_quota_enabled`; the rejection now says so.

Rejecting the value rather than raising the floor. A floor of 0.01 was the obvious fix
and it is wrong: very small tolerances are legitimate -- they are how a large disk is
throttled down to a sensible per-guest ceiling, and how the quota specs steer it
(tolerance = target * active / free lands around 1e-5 on the 174 GB volume this suite
runs on). A floor would forbid real configurations, and would have broken the entire
storage-quota describe block, to prevent one typo. Verified: those four tests still pass.

Tests: the rejection, that the stored value is untouched (validation fully precedes any
write), and the mirror -- 0.00001 still round-trips -- so the guard can't quietly become
a floor later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 21:37:39 +02:00
fabi
06bc9ddcb3 refactor(export): share the visibility filter between the row query and the estimate
`query_uploads` selects the rows the archives are built from; `estimate_export_bytes`
sizes them for the disk preflight. They stated the same WHERE clause separately, and
the direction of drift matters: an estimate that MISSES rows the archive writes
under-reserves, which is precisely the ENOSPC the preflight exists to prevent.

The integration test claimed to guard this and cannot. Both sides of
`the_estimate_sums_exactly_the_rows_the_archive_will_contain` are `SRC:`-marked
hand-copies in tests/common/mod.rs -- neither is production code -- so drift means
production moved while both copies sat still, and the test goes on passing. The
convention is sound for pinning behaviour; it is structurally incapable of detecting
divergence from the thing it copies.

So fix it where it can be fixed. One `export_visibility_where!()` fragment,
`concat!`-ed into both queries at compile time (still `&'static str`, no allocation),
with the `u`/`usr` alias contract stated. Divergence is now impossible by
construction rather than watched for.

The tests keep their value and lose the overclaim: the docstrings now say they pin
WHICH uploads may be counted -- each excluded row in the fixture is excluded by a
different predicate, so weakening any one of them still fails here -- and say plainly
that they do not detect drift, with a pointer to what does.

No behaviour change. The filters were verified identical before the hoist
(`u.event_id = $1 AND u.deleted_at IS NULL AND usr.uploads_hidden = FALSE AND
usr.is_banned = FALSE`); 99 backend tests still pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 20:50:04 +02:00
fabi
5f702f2b40 fix(audit-5): export disk preflight, media reclaim, low-disk warning, restore docs
Fifth round, all storage. The keepsake could not fit on the documented hardware
and failed halfway through a multi-GB write, leaving the deliverable stuck; the
quota had stopped bounding the disk; and the backup had no restore procedure.

Squashed from 6 commits, original messages preserved below.

──────── fix(export): refuse an export that cannot fit, and stop peaking at two generations

Nothing in export.rs ever asked whether the keepsake would fit. Both archives write
their media `Compression::Stored`, so each is essentially a byte-for-byte second copy
of the originals -- Gallery.zip always, and Memories.zip for every video and every
image at or under 5 MB. On the documented CX33 (80 GB, all three volumes on one
filesystem) the upload quota's fixed point leaves ~40 GB free, and a release spawns
BOTH halves concurrently against it.

The failure is not "the export failed", it is "the deliverable is stuck":

  1. ENOSPC lands partway through a multi-GB write.
  2. The epoch has already moved, so the job row is `failed` at the CURRENT
     generation and readiness (epoch = event.export_epoch AND status = 'done') is
     false -- GET /export/zip 404s.
  3. The last good archive sits on disk, unreferenced and unreachable.
  4. POST /host/export/rebuild, the only escape, re-arms the same doomed write.

Three changes.

Reclaim before building. `prune_stale_export_files` ran only after the new archive
was written, renamed and finalised. That reads as durability but buys nothing: the
moment `invalidate_and_arm` bumps the epoch the old archive is ALREADY unreachable,
so keeping it reserves gigabytes for a download nobody can perform -- and for a
takedown it is content someone explicitly asked to have removed. Peak usage is now
one generation. Narrower than the post-finalize prune on purpose: final archives
only, never a `.tmp` or a `viewer_tmp_` dir, since a superseded worker can still be
streaming into those and at build START is far more likely to be alive.

Preflight the space. SUM(original_size_bytes) over exactly `query_uploads`'
visibility filter, +10% for ZIP overhead, multiplied by the number of armed jobs --
without that multiplier each of the two concurrent halves independently sees "it
fits" and together they don't. Runs AFTER claim_job, not before as reported: bailing
before the claim leaves the row `pending` with no worker and no error, the
spinner-forever state `mark_failed`'s status guard exists to prevent. Fails open when
the mount can't be read, exactly as the upload quota does.

Show the host the reason. /export/status returned {status, progress_pct} and nothing
else, so the host dashboard could only render "fehlgeschlagen" next to the retry
button. The message was written to the row and surfaced solely in the ADMIN job list
-- a different screen, possibly a different person. It now travels with the status,
and only on a failure, so a message left on a since-succeeded row can't appear beside
a green "ist bereit".

Tests: 10 unit (the u128 clamp caught a real bug in the first draft -- saturating_mul
then /100 turns an overflow into a number ~100x too small, the one direction that
authorises the write being guarded against; the carried-forward archive must survive
its own older epoch in the filename), 4 DB-backed (the estimate is asserted against
the row set the archive actually contains, not against a restatement of the WHERE
clause, so the two queries cannot drift), 3 e2e over the four-hop plumbing.

──────── fix(maintenance): reclaim the media of deliberately deleted uploads

The quota stopped bounding the disk. `soft_delete_in_event` stamps `deleted_at` and
refunds `total_upload_bytes`, but nothing ever removed the bytes, and the hourly
sweep reached only `compression_status = 'failed'`. Upload 500 MB, delete, quota back
to zero, upload another 500 MB. Not an attack -- a guest curating their camera roll,
which is what people do. The host then sees guests hitting "Du hast dein Upload-Limit
erreicht" while the admin widget shows a disk full of files no upload row points at,
and the quota message is actively misleading because the space really is gone, just
not to anyone the accounting can name.

Two retention windows, because the two deletes mean different things. A compression
failure keeps its 14 days: the guest didn't ask for it and may not be able to retake
the photo. A deliberate removal gets 24 hours -- 14 days outlives the whole event, so
a deliberate delete would never reclaim anything while it mattered, and a day still
covers a mis-tap.

Wider than reported: ALL FOUR paths are reclaimed, not just the original. Preview,
display and thumbnail are each a separate file, none counted in
`original_size_bytes`, and nothing ever removed them either. That was invisible while
the sweep only saw failed compressions (which produce no derivatives) and becomes
three leaked files per upload the moment it reaches a successful one. A row is
re-selected until every path is cleared, and the columns are cleared only once every
file for that upload is gone -- clearing after a partial success would strand the
survivors in exactly the unowned state this drains.

`backfill_stale_derivatives` selects on `display_path IS NULL AND preview_path IS NOT
NULL`, which is close enough to the post-sweep state to be worth pinning: it is
guarded on `deleted_at IS NULL`, so it cannot re-decode an original that is no longer
on disk. Covered.

Residual, deliberately: within the 24h window the bytes are still spent and still
unaccounted, so delete-and-re-upload through an eight-hour event can outrun the
sweep. Bounding that means holding the quota until the file is reclaimed rather than
refunding at `deleted_at`. The low-disk warning is the net under it.

Tests: 6 DB-backed, replacing 3. The one asserting an owner-deleted upload IS
reclaimed is the exact inverse of what this file used to assert.

──────── feat(host): warn about low disk before it becomes unrecoverable

Storage visibility existed in exactly one place: a passive Speicherauslastung widget
on the ADMIN dashboard. A host who isn't the admin had no view of it, and nothing
warned anyone. README carried "Low-disk alert (< 10 GB free)" under Planned since v1.

Two things make this a safety net rather than a nice-to-have. postgres_data,
media_data and exports_data are all Docker named volumes on ONE filesystem, so
running out doesn't degrade a subsystem -- Postgres stops being able to write and the
whole event goes down. And the keepsake needs room for two gallery-sized archives,
which the export preflight can only ever refuse AFTER the release, when the event is
over and every remedy is harder.

So the threshold is not a fixed number alone. It fires on the 10 GB floor the README
always named, OR on "you could not build the keepsake right now" -- the trigger a
host can still act on, computed with the same arithmetic the preflight uses. Unknown
free space is NOT low: it fails open like the upload quota and the preflight do,
because a banner that cries wolf on an unreadable mount is a banner nobody reads.

Carried on GET /host/event, which the dashboard already fetches on load and on every
reload -- no new endpoint, no new poll. Rendered above everything else including the
PIN-reset queue, and it names the consequence (the event, not just the download)
rather than only the number.

Also fixes the host page's formatBytes, which topped out at MB: 30 GB free would have
rendered as "30720.0 MB", and a guest with 2 GB of uploads was already being shown
that way in the user list.

Tests: 5 unit on the threshold (including that plenty of free space is still low when
the keepsake wouldn't fit -- the case a fixed threshold misses entirely), 3 e2e.
The e2e drives it through `original_size_bytes` rather than a genuinely full disk:
the estimate is pure SQL over that column, so overstating one row moves the
accounting without touching a byte on disk.

──────── docs: add a restore procedure, fix the backup cadence, and correct quota_tolerance

Four things, all found by the same question: what does an operator standing at the
venue actually need?

A RESTORE PROCEDURE. There was none anywhere, and a backup you have never restored
isn't a backup. Two hazards worth writing down: media must be extracted preserving
ownership (the app runs as uid 100 / gid 101, and a root-owned restore makes every
upload fail with EACCES surfacing as a generic 500), and the app must be STOPPED
first, because migrations run on boot and a live pool will fight the restore.

Both the backup and the restore commands were run against the real stack before being
written down, which caught two that would have failed:

  - The plain `pg_dump` did not restore: `psql` aborted on `ERROR: schema
    "_sqlx_test" already exists`. pg_dump emits no DROPs without --clean --if-exists,
    so the documented dump could only ever be restored into an empty database. Fixed
    at the source (the dump is now self-cleaning) and verified end to end: 16 tables
    back, exit 0.
  - `--same-owner` does not exist in BusyBox tar, which is what `alpine` ships, so
    the extract aborted before unpacking anything. `--numeric-owner` plus the
    explicit chown, verified to land 100:101.

BACKUP CADENCE. "Weekly offsite" is the wrong shape when every irreplaceable byte is
created in one eight-hour window and nobody can retake a wedding. The backup that
matters runs that night, and again after the release so the keepsake is captured.
Also: take the DB dump and the media tarball back to back, or you get rows pointing
at files the dump doesn't know about.

quota_tolerance WAS DOCUMENTED AS SOMETHING IT ISN'T. .env.example called it "fraction
of disk that triggers the low-storage warning". It is the multiplier in
`floor(free_disk * tolerance / active_uploaders)` -- so an operator who wants "warn me
later" and sets 0.95 is actually authorising guests to fill 95% of the disk, moving
the fixed point from 43% to ~49% and eating the export headroom. The admin UI labelled
it "Toleranz (0-1)" with no explanation at all, which invites exactly that reading;
it is now "Speicher-Anteil für Gäste" with the formula in the hint. Wrong docs on a
tuning knob are worse than no docs.

SIZING. New section with the arithmetic: three volumes on one filesystem, the quota
fixed point at tolerance/(1+tolerance), and the fact the 80 GB baseline does not cover
the keepsake -- both archives are built concurrently and each is roughly a second copy
of every original. Provision ~3x expected media, or give exports its own volume.

Also ticks the low-disk alert off the roadmap, since it now exists.

──────── chore: raise the db memory limit and rate-limit social writes

Two smaller operational items.

POSTGRES 512M -> 1G. DATABASE_MAX_CONNECTIONS is 30 for a ~100-guest event (feed
polling + SSE + uploads at once), and 30 backends plus Postgres 16's default
shared_buffers leaves very little headroom at 512M. An OOM here doesn't degrade one
feature -- every request path touches the database, so it takes the event down.
Memory is the cheaper knob than shrinking the pool back and reintroducing the
queueing it was raised to fix. .env.example now names the pairing explicitly, the way
it already does for COMPRESSION_WORKER_CONCURRENCY.

SOCIAL WRITES WERE UNTHROTTLED. toggle_like, add_comment and delete_comment were the
only mutating endpoints in the app with no limit at all -- upload, join, recover,
export and admin login all carry one. Asymmetric coverage rather than a deliberate
decision.

Low severity, and honestly so: a like fans an SSE broadcast to every client, but the
export regeneration a comment deletion triggers is contained (REGEN_DEBOUNCE 20s,
workers born with their epoch, superseded ones inert). So the ceiling is 120/min --
far above anything a real guest produces. This bounds a script, not an enthusiastic
double-tapper.

ONE bucket across all three actions: separate buckets would let a caller triple the
aggregate write rate by alternating between them. Keyed per USER, matching the feed
and upload limits -- at a venue every guest is behind one NAT, and an IP key is what
made the /join and /feed limits turn guests away in the first place.

Migration 020 seeds both keys, and both are wired into the admin allowlist, the
config UI and the e2e reseed -- the step two earlier per-area toggles missed, which
left switches that existed in code and could never be flipped.

Tests: 4 e2e, including that the shared bucket really is shared (the part most likely
to be lost in a refactor) and that one guest hitting the ceiling doesn't block
another behind the same IP.

──────── fix(e2e): stop the video poster assertion racing the ffmpeg thumbnail

Pre-existing, and it fired for real during the full-suite run on a cold stack.

The lightbox binds `poster={upload.thumbnail_url ?? undefined}`, so the attribute is
absent until compression produces the thumbnail. This test asserted on it immediately
after seeding, never waiting for the worker -- unlike the Range test further down the
same file, which does poll. Against a warm stack the worker usually wins; against a
freshly rebuilt one (`stack:down -v`, cold ffmpeg) it doesn't.

That is the worst possible time for a false failure: the first run after a rebuild is
exactly when you are trying to establish whether a change broke something. Poll for
`compression_status = 'done'` before the poster assertion. The `src` assertion needs
no wait and keeps none.

Verified with --repeat-each=3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 20:10:54 +02:00
fabi
31faccfdf8 fix(audit-4): refuse undecodable images at admission, stop retrying permanent failures
Fourth round: reject at the door what the worker can only fail on, and stop
retrying errors that cannot succeed — while keeping IoError retryable so ENOSPC
still gets another attempt.

Squashed from 2 commits, original messages preserved below.

──────── fix(upload): refuse undecodable images at the door, and stop retrying them

Two halves of the same complaint: an oversized photo was accepted with a 201 and
then silently soft-deleted minutes later, after the worker had burned six seconds
of backoff re-reaching a conclusion it could not change.

Admission. The compression budget now runs at upload time, against the header
only, so a guest is told immediately and told why:

  "Bild hat zu viele Bildpunkte (ca. 99 Megapixel) und kann nicht verarbeitet
   werden. Bitte verkleinere es und lade es erneut hoch."

instead of watching the photo vanish behind a vague "could not be processed" —
which arrived only if they happened to still be on the feed with that card
loaded. Nothing is stored, so there is no row to soft-delete and no orphan for
the sweep to reclaim.

Admission and the worker share ONE function (`decoder_within_budget`), so they
cannot drift apart and start disagreeing about what is acceptable — a photo
accepted at the door and rejected by the worker would be worse than either
behaviour alone. The worker keeps its own check: the backfill decodes files that
predate this check, and defence in depth is the whole reason the budget exists.

Retries. The loop retried every failure, including ones that are a property of
the input. An image over the budget, a corrupt file, an unsupported format: each
fails identically on all three attempts, so the only effect was 2s + 4s of sleep
and three near-identical warnings before the same outcome. `is_permanent_image_error`
classifies the `ImageError` variants that cannot change between attempts — Limits,
Unsupported, Decoding — and the loop gives up on those at once. `IoError` is
deliberately excluded: an ENOSPC while writing a derivative is exactly the
transient case the retry exists for, and misclassifying it would turn a blip back
into the data loss round 1 fixed. Measured: retry log lines went from 3 per
oversized upload to 0.

Tests: unit tests for both sides of the classifier (a Limits error is permanent, a
missing file is not) and for admission agreeing with the decoder on accept AND
reject. The e2e spec is rewritten for the new contract — 400 with an actionable
message, nothing stored, backend alive after a burst of four — plus a mirror
asserting an ordinary photo still uploads and processes, since a budget that
rejected everything would satisfy the other two.

──────── fix(upload): narrow the admission check to the memory budget only

The admission check I just added rejected ANY image the decoder couldn't build —
corrupt, truncated, or unsupported, not only over-budget. That broke two
adversarial tests, and they were right to break.

07-adversarial/file-upload-attacks pins, deliberately, that acceptance follows the
MAGIC BYTES: a payload whose first three bytes are a JPEG header is accepted
regardless of what follows, because the security property under test is that the
client-declared Content-Type has no influence. Both failing cases upload 1024
bytes of JPEG magic followed by zeros. Rejecting those at admission is a
different, broader contract than the one asked for, and rewriting an adversarial
test to match new behaviour is precisely the thing that needs justifying rather
than doing quietly.

So admission now checks only what it was meant to: `exceeds_decode_budget`
returns true solely for `ImageError::Limits`. A corrupt file goes to the
compression worker exactly as before — which handles it gracefully and, since the
retry classifier in the previous commit, no longer burns backoff on it. The
resource guard is the part that had to move earlier; nothing else did.

Tests: the size agreement between admission and the worker is still asserted in
both directions, plus a new one writing a magic-bytes-only stub and asserting
admission accepts it WHILE the worker still rejects it — pinning the boundary
between the two checks so a future widening fails here rather than in the
adversarial suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:55:56 +02:00
fabi
06ade4e158 fix(audit-3): restore the decode allocation guard, route /health in production
Third round. The decode allocation guard is a regression from round 1: swapping
reader.decode() for into_decoder() silently dropped the max_alloc enforcement
while keeping the comment that claimed it held.

Squashed from 3 commits, original messages preserved below.

──────── fix(imaging): restore the decode allocation guard I removed in round 1

This is a regression I introduced, not a pre-existing gap. Before 05948d8 the
compression worker used `ImageReader::decode()`, which does:

    let mut decoder = Self::make_decoder(format, self.inner, limits.clone())?;
    limits.reserve(decoder.total_bytes())?;   // enforces max_alloc
    decoder.set_limits(limits)?;

Reading the EXIF orientation tag needs `into_decoder()` instead, and that skips
the reserve entirely — the crate's own FIXME concedes `from_decoder` doesn't
compensate. Nothing else enforces `max_alloc`: the JPEG decoder's `set_limits`
only checks support and dimensions. So the 256 MiB budget has been inert since
that commit, and round 2 then propagated the weakened path into export.rs through
the shared helper, in a commit whose message claimed the helper "carries" the
decompression-bomb cap. It didn't, and the comment saying max_alloc "hard-caps
the decode allocation" was simply false.

What was left was only the per-axis cap, which permits 12000x12000 — 412 MiB
decoded, 824 MiB for the two concurrent decodes the worker runs by default,
against a 1 GiB container. Deploy-blocking right now because bumping
DERIVATIVES_REV makes the first boot after a deploy re-decode the entire gallery
two at a time: an OOM kill there restarts the container, which re-runs the
backfill. A boot loop, on the first deploy of these fixes.

Re-add the reserve exactly as `decode()` does it. Per the budget decision it stays
at 256 MiB (~89 MP for RGB8, above any mainstream phone's real output); two
concurrent decodes now peak at 512 MiB. Oversized images take the graceful path
from round 1 — original retained, quota refunded, upload-error toast — and fail
after the header parse but BEFORE any pixels are read, so they cost a header read
rather than an allocation. Measured peak during a concurrent oversized burst: 3.0
MiB.

Test parity is the other half, and the reason this was invisible: the e2e app
container had NO memory limit while production is capped at 1 GiB, so a decode
that would OOM-kill production simply succeeded in CI. Mirror the 1 GiB cap in
docker-compose.test.yml. That is the third divergence of this shape, after WebKit
missing from CI and /health existing only in Caddyfile.test.

Tests: a fixture that is 568 KiB on disk and 283 MiB decoded (11000x9000 = 99 MP,
deliberately UNDER the per-axis cap so the axis check cannot be what rejects it).
A unit test asserts the refusal — it fails against the old code, which decoded it
into an 11000x9000 buffer — with a companion asserting an ordinary photo still
decodes AND still gets its orientation applied, so the guard didn't become a
blanket refusal. An e2e test uploads it singly and as a concurrent pair, asserting
compression lands in 'failed' and the backend is still serving and still
processing afterwards.

──────── fix(deploy): route /health in production, and actually apply Caddyfile changes

Two defects in the update procedure I wrote last round, both of which make a
successful-looking deploy a lie.

1. The documented health check could never pass.

`curl -fsS https://DOMAIN/health` 404s against a perfectly healthy production
stack. The backend registers /health on its ROOT router, not under /api/v1, and
the production Caddyfile proxies only /api/* and /media/* — so /health fell
through to the SvelteKit catch-all, which has no such route and returns its 404
page. With -f, curl exits 22 and the `&& echo` never runs. My own gloss
("Anything other than ok means check the logs") then sent the operator chasing a
phantom outage.

e2e/Caddyfile.test has carried `reverse_proxy /health app:3000` since it was
written — precisely because the catch-all would otherwise swallow it. Production
never did. Per the fix-the-gap-not-the-doc call, production gets the same line,
and /health joins the no-store matcher so a cached response can't report the last
known state instead of the current one. Verified by running the production
Caddyfile against the real backend: /health -> 200 "ok", Cache-Control: no-store,
with /api/v1/event and / unaffected.

2. The sequence never reloaded Caddy, so a Caddyfile-only change was dropped.

`--build` only rebuilds services with a `build:` section, and caddy is a pinned
upstream image. Compose decides whether to recreate a container from its config
hash, which covers the mount SPECIFICATION but not the mounted file's CONTENTS —
so a git pull that changes ./Caddyfile produces no delta, Compose reports
`Running`, and Caddy serves its old config indefinitely. Exit code 0 throughout.

Round 1's iOS download fix (137c4ee) is exactly this shape: Caddyfile plus four
e2e files, so 100% of its production effect is in that one file. Following the
README to the letter deployed it, showed both image IDs changing, and left iOS
downloads broken.

Demonstrated rather than assumed — added a probe header to a Caddyfile, ran the
old sequence (`up -d --build`): header absent, change silently dropped. Ran the
new step 4 (`up -d --force-recreate caddy`): header served.

`--force-recreate` rather than `restart` or `caddy reload` because the bind mount
is resolved to an inode at container-create time and git pull replaces the file
rather than editing in place, so a restart can re-read the stale content — the
exact failure I hit in round 1 when `caddy reload` didn't pick up an edit.

Also rewrites the "db and caddy are untouched … so data volumes survive" sentence.
I wrote it as reassurance; "caddy is untouched" was the bug.

──────── chore: take the Bash(*) permission change back out of the shared settings

`.claude/settings.json` is committed and applies to anyone who clones. Fabi's
local `allow: ["Bash(*)"]` plus deny list ended up in it, inside f0d69f1 — a
commit about the image decode guard, which has nothing to do with permissions.

That was my mistake, twice. The file was already modified when I started the
round: my `git status --short` check printed "(clean)" from an unconditional
`echo` rather than from the status output, so I read a dirty tree as clean. Then
`git add -A` swept it into an unrelated commit, and I reported afterwards that I
had left it untouched. Neither the check nor the claim was true.

Restores the shared file to its previous three narrow entries. The permission
setup itself is preserved, moved to `.claude/settings.local.json`, which
`.gitignore:34` covers precisely so per-user permissions stay per-user — the
existing 442 entries there are kept alongside it.

Not rewriting f0d69f1 to erase this: main is unpushed so it would be safe, but a
visible correction is worth more than a tidy history, and a rebase across the
merge commits carries more risk than the mistake does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:18:08 +02:00
fabi
b601c062bd fix(audit-2): video playback, role identity, keepsake EXIF, recover ceiling, WebKit CI
Second round: the two regressions the first round introduced, the remaining
gaps it left, and the CI change that makes the iOS guarantees actually gate a
change rather than being asserted.

Squashed from 9 commits, original messages preserved below.

──────── docs(deploy): document the update path — `up -d` alone ships nothing

The README only ever described a fresh install. There was no update section
anywhere, and `--build` appeared nowhere in the docs.

That matters because `app` and `frontend` are `build:` services with no published
image tag, and Compose has no source-change detection: if an image by that name
exists it is reused. So the natural `git pull && docker compose up -d` reports
"Container app-1 Running", rebuilds nothing, and exits 0. A deploy that shipped
none of the new code is indistinguishable from a successful one — which is how
eleven merged fixes can sit in the repo and never reach the box.

Verified both halves against a real stack rather than asserting them: with a
source change staged, `up -d` left the image ID untouched; `up -d --build`
produced a new image ID and a healthy /health.

Adds an "Updating an existing deployment" section covering backup-before-migrate,
pull, rebuild, health check, and an image-ID comparison to prove a build actually
happened. Also spells out the rollback trap: migrations run on boot and are not
undone by checking out an older commit, so rolling back code without restoring
the snapshot leaves the schema ahead of the binary and the app refusing to start.

──────── fix(video): play the actual video, and answer Range requests

Every video in the app was unplayable. Two independent defects, either one
sufficient on its own, and nothing in the suite covered either — no test
anywhere played media or asserted a `<video>` src.

1. The lightbox handed `<video>` a JPEG.

`pickMediaUrl` is mime-agnostic, and compression only ever produces a THUMBNAIL
for a video (one `ffmpeg -vframes 1` frame) — no preview, no display. So in the
DEFAULT saver mode the element's src resolved to `/api/v1/upload/{id}/thumbnail`,
served as `image/jpeg` with `nosniff` so the browser can't even sniff its way
out. Chromium reports DEMUXER_ERROR_COULD_NOT_OPEN.

Fixed in the lightbox rather than in `pickMediaUrl`: FeedListCard shares that
helper and legitimately wants the thumbnail for its `<img>` poster, so a central
mime branch would break the feed. This mirrors the rule the diashow already
applies ("videos play the original file directly"). Added `preload="none"` so
saver-mode guests on cellular still fetch nothing until they press play — there
is no smaller video derivative to offer them — plus `playsinline`, without which
iOS hijacks playback into fullscreen.

2. `stream_media_file` ignored Range entirely.

It took no request headers, so it could not see `Range`; it always returned 200
with the whole body and never sent Accept-Ranges or Content-Range. iOS Safari
opens every `<video>` with a `Range: bytes=0-1` probe and abandons the load
without a 206 — so video failed on the app's primary platform even in `original`
mode, where the src was already correct.

Adds single-range support (`bytes=N-`, `bytes=N-M`, `bytes=-S`) with 206 +
Content-Range, 416 + `bytes */len` past EOF, and Accept-Ranges advertised on
every response. Anything it won't handle — multi-range, non-bytes units, garbage
— falls back to a full 200, which RFC 9110 explicitly permits and which is safer
than guessing. All four media routes share the helper, so seeking works
uniformly.

`get_original` now serves `inline` instead of `attachment`. An attachment
disposition is hostile to a `<video>` element, and this route is the only source
of playable video bytes; it also matches what the UI promises, since the action
is labelled "Original anzeigen" — view, not download. `no-store` is deliberately
kept so a takedown still revokes access promptly; ranges work fine under it, the
client just re-fetches.

Tests: 11 unit tests pin the parser (the iOS `bytes=0-1` probe, inclusive ends,
suffix ranges, clamping past EOF, 416 vs 200, malformed fallbacks). A new
03-feed/video-playback spec asserts the src is the original and not the
thumbnail, that the browser accepts the bytes as media (readyState > 0, no
MediaError), that no video bytes are delivered before play, and that Range
returns the correct 206 slices and a 416 past EOF — verified on both Chromium
and WebKit.

The "not downloaded before play" test asserts no *delivered body* rather than no
request: WebKit opens a connection for a preload="none" video and immediately
aborts it (GET, no Range, status 0, nothing transferred) while Chromium issues
nothing at all. The portable guarantee is that no response carrying bytes
completes.

──────── fix(auth): bind the role store to the identity, not to the tab

The role store I added in the moderation work is a module-level singleton seeded
ONCE at import. `goto()` is a client-side navigation, so leaving and re-joining in
the same tab re-imports no module and re-runs no onMount — the previous user's
role simply stayed resident. Nothing reset it: not join, recover, admin login,
"Event verlassen", `clearAuth`, nor the api.ts 401 auto-clear.

So a host who left, followed by a guest joining on the same phone, left that guest
with `isStaff === true` and a "🚫 Beitrag entfernen" action on other people's
photos. The backend 403s the delete, so this was a false affordance rather than a
privilege escalation — but `/feed` never fetched `/me/context`, so unlike every
other route it never self-corrected either. It survived until a hard reload.

The mirror case was equally broken and easier to overlook: a guest who recovered
into a host account got NO host affordances.

`clearAuth` already had a hook registry for exactly this shape of problem, with a
comment explaining it exists to avoid circular imports. Add the missing mirror,
`onSetAuth`, fired by both `setAuth` and `setAdminAuth` after the new token is
resident, and have the role store register on both sides: clear to null on
logout, re-seed from the new token on login. That also gives
`syncRoleFromToken` — dead code with zero callers since I introduced it — its
intended purpose.

Seeding from the claim fixes the reported bug, but the claim is frozen for the
token's 30-day life, so a promotion or demotion still wouldn't reach the feed.
`/feed` now calls the existing `refreshEventState()` on mount, which fetches
`/me/context` and applies both the authoritative role and the lock/release state
in one request. The feed is the one route gating a destructive action on the role,
so it should not be the only route running on a stale claim.

Tests: 04-host/role-identity-reset drives the real flows. The first asserts the
host DOES see the action before asserting the newcomer does not — a negative
assertion alone would pass against a build that shipped no moderation at all. The
second covers the mirror, promoting a guest server-side while their resident token
still claims `role: guest`, so a fix that only cleared the role would fail it.

──────── fix(compression): reclaim failed originals instead of leaking them

Round 1 stopped the compression worker deleting an upload's original on failure —
a transient ENOSPC or a codec panic must never destroy the only copy of a photo a
guest cannot retake. But it left `Upload::soft_delete`'s quota refund in place, so
the bytes stayed on disk while the uploader was charged nothing for them.

That is worse than it first looks. The row is soft-deleted, so the file is
invisible and unowned; a guest hitting a reproducible codec failure can accumulate
orphans indefinitely at zero personal cost. And `active_uploaders` counts only
users with non-deleted uploads, so dropping out of that count RAISES everyone's
per-user ceiling — the leak loosens the very quota meant to contain it.

Keep the refund: the uploader didn't cause the failure and shouldn't silently lose
quota to it. Bound the leak instead, with an hourly sweep alongside the existing
session cleanup in `spawn_periodic_tasks`, reclaiming failed originals older than
14 days — comfortably longer than any single event, so an operator investigating a
failed upload still has the file.

The selection predicate is the entire safety argument, so it is deliberately
narrow: `compression_status = 'failed'` AND soft-deleted AND past the window AND
`original_path <> ''`. That is exactly the state the give-up path leaves behind,
and it cannot reach a live upload, an owner-deleted one, or a failure still inside
its recovery window. `original_path` is cleared after a successful reclaim, which
makes the sweep idempotent — otherwise a row whose file is already gone is
re-selected on every tick forever. The row itself is kept as the audit trail.

Tests reproduce the selection verbatim (same pattern as upload_concurrency) and
assert it against five near-misses that must survive, both sides of the retention
boundary, and the idempotence property.

Also fixes two comments in export.rs still claiming "the compression worker
hard-deletes an original when its transcode fails" — no longer true, and the
defensive handling they justify is now justified by this sweep and by ordinary
deletes instead.

──────── fix(export): apply EXIF orientation in the keepsake too

Round 1 fixed EXIF orientation in the compression worker, which corrected the live
app — feed preview and diashow display. The export worker was missed, and it does
not reuse those derivatives: it re-decodes the originals itself with `image::open`,
which ignores the orientation tag, then re-encodes to JPEG, which drops the tag —
so the viewer has no way to recover it.

The damage was oddly shaped, which is exactly why it reads as a viewer bug:

  Gallery.zip originals              correct  (byte-copied, EXIF intact)
  Memories viewer grid thumbnails    SIDEWAYS (always)
  Memories viewer full image >5 MB   SIDEWAYS (re-encoded at 2000px)
  Memories viewer full image ≤5 MB   correct  (streamed byte-for-byte)

So in the keepsake people actually keep, every portrait photo in the grid was on
its side, and clicking through silently "fixed" small photos but not large ones.

Rather than paste the decoder dance a third time, extract `services::imaging::
decode_oriented` and route both workers through it, so there is exactly one way to
turn a file on disk into a DynamicImage. It carries a second invariant that had
also drifted: `image::open` applies NO decode limits, so the export path was
decoding arbitrary user-supplied images unbounded — the decompression-bomb cap
existed only in the compression worker. Both now come as a pair, which is the
point of having one function.

Not done: switching export to consume the existing `display` derivative. It would
fix orientation and drop a redundant full-resolution decode per photo, but it
would also replace the pristine ≤5 MB originals in the keepsake with 2048px
re-encodes — a real quality regression in the one artefact people keep forever.

Test uploads the round-1 fixture (40x20 landscape tagged Orientation=6), runs a
real export, pulls the thumbnail out of Memories.zip and asserts it came back
portrait — with a sanity check that the source really is stored landscape, so the
test can't pass against a pipeline that does nothing.

──────── fix(recover): cap name cycling, and stop bcrypt blocking the runtime

Round 1 gave /join a per-IP ceiling and left /recover with only its
`recover:{ip}:{name}` bucket. That key is right for the job it was written for —
stopping someone who knows a display name (they're listed on the feed) from
burning the victim's 3-strike PIN counter and locking them out on repeat. But the
name is ATTACKER-CHOSEN, so cycling names mints a fresh 5-attempt bucket every
time and the per-IP cost is unbounded.

What sits behind that limiter makes it worse than a normal flood: every call runs
a cost-12 bcrypt verify, including an UNCONDITIONAL throwaway verify for names
that don't exist — added deliberately to close a timing oracle. So an unknown name
is the single cheapest way to make the server do ~200ms of hashing.

Adds `recover_ip_rate_per_min` (default 30, migration 019), checked BEFORE the
per-name bucket so a name generator can't walk past it. 30/min is far above any
real recovery attempt while capping a flood. The per-name bucket is untouched and
remains the anti-guessing control.

The second half matters as much as the first: bcrypt was running inline on the
async runtime everywhere. At cost 12 that pins a tokio worker thread for ~200ms,
and there is only one per core — so a login flood stalled every other request on
the box, including the feed. There was no spawn_blocking anywhere in the auth
module, despite SECURITY-BACKLOG claiming bcrypt had been offloaded.

Route all of it through `verify_password` / `hash_password` on the blocking pool.
That covers /recover, /admin/login, the host PIN reset, and — the one most likely
to bite at a real event — the PIN hash minted on every single /join. Saturating
the blocking pool degrades logins; saturating the worker threads degrades
everything.

Tests: cycling distinct names from one IP now hits the ceiling with a Retry-After,
and — the assertion that keeps the fix honest — repeated wrong PINs against ONE
name are still throttled with the ceiling set generously high, so the ceiling
added protection rather than replacing it.

──────── ci(e2e): run WebKit, so the iOS guarantees actually gate a PR

The workflow installed only Chromium and ran chromium-desktop + chromium-mobile.
iOS Safari is the app's stated primary user — a wedding guest opening a QR link —
and WebKit is the only engine in the matrix that reproduces two of its behaviours:

  - it enforces X-Frame-Options on the hidden download iframe, so a site-wide DENY
    makes the keepsake download silently do nothing. Blink hands attachments to
    the download manager before the frame check and never notices.
  - it abandons a <video> load unless its Range probe gets a 206.

Both of those shipped. Adding 06-export to the webkit project in the round-1 fix
bought nothing on a PR, because CI never ran that project at all — the regression
test written specifically to catch the blocker only ever executed locally.

Runs 71 tests (67 pass, 4 skip on the documented IndexedDB-blob harness
limitation) in ~1.5 minutes locally, using the exact command added here.

──────── chore(backend): satisfy cargo fmt

`checks.yml` runs `cargo fmt --check`, and it has been failing since the round-1
audit fixes: I gated those on `cargo build` and `cargo clippy` but never ran fmt,
so three files drifted then and eight more this round. Pure formatting — no
behaviour change; clippy stays at zero and all 70 backend tests still pass.

Worth noting for next time: clippy passing is not evidence fmt does.

──────── chore: satisfy prettier in frontend and e2e

`checks.yml` runs `npm run format:check` for both projects and both were failing.

- frontend/src/lib/ui-store.ts is mine, unformatted since the round-1 upload-queue
  badge fix — the same miss as the rustfmt one: I gated on svelte-check and eslint
  but never on format:check.
- e2e/loadtest/* and e2e/shots.mjs have been unformatted since 7758270 and are
  unrelated to the audit work. Fixed here because they block the same gate and the
  fix is mechanical; no behaviour change in either project.

Still red and deliberately NOT fixed here: `npm run lint` in the frontend reports
`svelte/prefer-svelte-reactivity` on routes/diashow/+page.svelte:208 (a mutable
`Set` where the rule wants `SvelteSet`), pre-existing since 5009590. That one is a
real reactivity change in code I have no test coverage for, so it belongs in its
own change rather than smuggled into a formatting commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 21:28:11 +02:00
fabi
0c0eed885a fix(audit-1): media gating, per-user limits, moderation UI, upload integrity
First audit round: a blocker and eight high-severity findings across the media
access gate, the guest-facing rate limiters, host moderation, and the upload
pipeline — plus the backup commands and the e2e/production divergences that hid
several of them.

Squashed from 7 commits, original messages preserved below.

──────── fix(export): let the keepsake download through X-Frame-Options on iOS

The keepsake download navigates a hidden, same-origin iframe (deliberately:
a top-level navigation to a 404/429 would unload the PWA). Caddy stamped a
site-wide `X-Frame-Options: DENY` that also covered the proxied `/api/*`.

Blink hands a `Content-Disposition: attachment` response to the download
manager at the network layer, so Chromium never noticed. WebKit enforces XFO
on the frame navigation first and aborts the load — so on iOS Safari, the
app's primary platform, tapping Download did nothing at all, silently.

Carve the two export endpoints out to SAMEORIGIN, which still blocks
cross-origin framing. Implemented as two disjoint matchers rather than an
override: Caddy applies the FIRST `header` directive outermost, so it wins on
write and a later, more specific `header` is silently ignored (verified
against the running test stack).

Also close the test gap that let this ship:

- `06-export` ran on chromium-desktop only; add it to `webkit-iphone`, the
  only engine that enforces XFO on the download frame.
- No test in the suite ever clicked a download button — every archive
  assertion used Node `fetch`, which has no frame and no XFO enforcement.
  Add a spec that clicks it and awaits a real `download` event. Verified
  falsifiable: with the blanket DENY reinstated it fails and reports the
  WebKit refusal as the cause.
- Fix `ExportPage`'s card-scoped locators, which matched nothing: the cards
  carry `class="card p-5"` (a Tailwind `@apply` component class), never the
  `rounded-xl` the page object looked for. This had left the "shows enabled
  download buttons" test red on main.

──────── fix(media): close the percent-escape bypass of the media gate

`/media/%70reviews/{id}.jpg` served a taken-down photo to anyone,
unauthenticated. Verified against the running stack: the literal path 404s,
the escaped one returned 200 with the full image. Same for displays,
thumbnails and originals, and any escaped byte in any position works.

Cause: the block was four `nest_service("/media/previews", 404)` route
matches sitting above a `ServeDir` on `/media`. axum matches on the RAW path
(matchit does no percent-decoding), while `ServeDir` percent-decodes when it
resolves the file. So `%70reviews` missed every blocker, fell through to the
ServeDir, and was decoded back to `previews/` on disk — reaching the bytes
with no soft-delete and no ban-hide check. That defeats a host takedown,
which is the entire point of the gate.

Remove the `/media` route tree outright instead of racing the decoder.
Nothing needs it: every media URL the backend emits is already a gated
`/api/v1/upload/{id}/{original,preview,display,thumbnail}` alias
(handlers::feed), the frontend contains zero `/media/` references, and the
`/media` in config.rs/disk.rs is the filesystem path while `media/` in
export.rs is a path inside the zip. `/media/**` now 404s regardless of
encoding. The route's own comment already said it "serves nothing" — it
wasn't a backstop, it was the vector.

Caddy keeps proxying /media/* deliberately: the app 404s it, and forwarding
means the e2e gating specs exercise the app's refusal exactly as production
would rather than being masked by the SvelteKit 404 page.

Extend the gating spec with the encoded variants — asserting only the literal
spelling is what let this sit undetected.

──────── fix(rate-limit): key the guest-facing limiters per user, not per IP

At a venue every guest is behind one NAT, so an IP-keyed limiter hands the
whole party a single bucket. On a fresh deploy 12 guests arriving together
meant 5 joined and 7 were turned away, with no Retry-After telling them when
to retry. `/feed` (60/min) and `/export` (3 per DAY — the fourth guest to
fetch their keepsake locked out until tomorrow) had the same defect.

`feed_delta` was already keyed per-user and its comment states the exact
rationale ("so one client can't starve others behind a shared NAT"); this
makes its siblings match.

- feed:{ip}   -> feed:{user_id}    (auth was already in scope)
- export:{ip} -> export:{user_id}  (resolved from the download ticket's
  session, which was previously looked up and discarded)
- join:{ip}: pre-auth, so there is no user to key on. Split in two — a loose
  per-IP ceiling that only bounds raw volume (new `join_ip_rate_per_min`,
  default 60, migration 017), plus the real 5/60s anti-spam bucket keyed
  per (ip, name), mirroring the existing `recover:{ip}:{name}`.

admin_login / recover / pin_reset_req stay IP-keyed on purpose and are now
commented as such: they guard credential guessing, where a per-user or
per-name key would just hand an attacker a fresh bucket per guess.

Retry-After: the machinery existed but 7 of 8 sites called `check()` and
hard-coded `None`, so a throttled client was told to back off but never for
how long. Delete the bool `check()` wrapper entirely so `check_with_retry`
is the only entry point and the delay cannot be discarded by accident. Also
surface it for the PIN lockout, where the deadline was already known.

Fix the "unknown" fallback while here: every client_ip() caller passed that
literal, so any request without X-Forwarded-For — anything reaching the app
directly rather than through Caddy — shared ONE global bucket. Serve with
connect-info and use the peer address.

Tests: the reseed forces every limiter toggle off before each test, which is
why this whole class was invisible. Add 01-auth/rate-limit-shared-nat, which
enables them and asserts 12 guests share an IP without collision, that one
guest hammering their own name IS still throttled (so the fix re-keys rather
than removes the limit), and that feed/export buckets are per-user. Retarget
the ddos join test at the new per-IP ceiling — it asserted the defect.

Also seed `admin_login_rate_enabled` (read by the handler, seeded by no
migration and no reseed) and register `join_ip_rate_per_min` in the admin
config allowlist. Unrelated pre-existing red test fixed: 01-auth/join
asserted a "Willkommen!" heading the wedding redesign removed.

──────── feat(moderation): let a host remove a guest's photo or comment from the UI

`DELETE /host/upload/{id}` and `DELETE /host/comment/{id}` were complete on the
backend — transactional, SSE-broadcasting, audit-logged — and had zero frontend
callers. The feed context sheet offered "Löschen" only when
`target.user_id === myUserId`, so the only lever a host actually had against an
unwanted photo was banning the uploader.

That is both disproportionate and ineffective. A ban doesn't retract what was
already posted, and it makes things strictly worse for comments: the ban check
runs BEFORE the ownership check on the guest delete route, so banning an abusive
author leaves their comment on screen and permanently undeletable by them. With
no host affordance, nobody could remove it at all.

- feed: hosts/admins get "Beitrag entfernen" on other people's posts, routed to
  the host endpoint (the guest route 403s anything the caller doesn't own) with
  moderation-specific confirm copy. Own-post "Löschen" is unchanged.
- lightbox: same for comments, via /host/comment/{id}.
- Ban semantics are deliberately untouched (USER_JOURNEYS §10 — banned users keep
  read access and cannot write). The deadlock is broken by giving the host a way
  in, not by loosening the ban.

Live role (this had to come first). `getRole()` decodes the JWT claim, but the
token is never reissued — the backend slides the session row forward and treats
the DB row as authoritative. The claim is therefore frozen for the token's
lifetime: up to 30 days. A guest promoted at the party saw no Host-Dashboard and
no moderation actions until they signed out and back in, even though
`/me/context` had been returning their real role on every page load and 4 of its
6 call sites dropped the field on the floor.

Add `role-store.ts`: seeded from the claim so there's no flash of the wrong nav,
then corrected by every `/me/context` response. Point the ad-hoc `getRole()`
callers at it (account, upload, host, admin, and the new feed gate). The host and
admin dashboards now derive `myRole` reactively, so a demotion disables their
controls immediately instead of at next login.

Tests: 04-host/moderation-ui drives the real UI — host removes a guest photo and
it's gone from /feed server-side; a plain guest is offered nothing on someone
else's post (the mirror that keeps the first test honest); a promoted guest gains
the dashboard on reload while their token still carries `role: guest`; and a host
removes the comment of an already-banned guest, asserting first that the author's
own delete 403s so the deadlock is real.

──────── fix(upload): stop destroying originals, apply EXIF orientation, surface rejections

Three defects in the same pipeline, each of which loses a photo or misrepresents
one.

1. A transient error destroyed the guest's only copy.

`process`'s error arm unconditionally `remove_file`d the original. Every failure
routed there: `create_dir_all`, both derivative `save_with_format` calls (disk
full is the canonical case, and it arrives exactly when many guests upload at
once), a panic inside the image codec, or a momentary DB-pool exhaustion. The
row is only SOFT-deleted, so the bytes were the sole unrecoverable part — and
they were the part we deleted. The author already knew this was wrong next door:
`backfill_missing_display` says it "must NEVER soft-delete an upload that already
has a working preview".

Retry up to 3 times with backoff (re-checking the e2e generation guard after each
sleep), and on final failure keep the refund + soft-delete but leave the original
on disk, logging its path. A failed upload is now recoverable instead of gone.

2. Every portrait photo was stored sideways.

Phones don't rotate sensor data — they record the camera orientation in EXIF and
store the pixels as shot. `decode()` returns those raw pixels and the JPEG
re-encode writes no EXIF, so the 800px preview, the 2048px diashow display and
the keepsake were all rotated 90°, while "Original anzeigen" rendered upright
because the original keeps its tag. That asymmetry is why it reads as a viewer
bug. There was no EXIF handling anywhere in the repo and no exif crate.

Read the tag via `into_decoder()` (which carries the decode Limits through, so
the decompression-bomb cap is untouched) and apply it. Missing/malformed tags
fall back to NoTransforms — most images have none.

Existing derivatives are already baked wrong, so migration 018 adds
`derivatives_rev` and `backfill_missing_display` becomes
`backfill_stale_derivatives`: it now also picks up anything below the current rev
and regenerates it once from the original, which still carries its EXIF. Videos
are marked current in the migration — ffmpeg already honours the rotation matrix.
Bump DERIVATIVES_REV for any future change that invalidates derivatives.

3. A rejected upload vanished without a word.

`UploadQueue.svelte` — 162 lines holding the ONLY renderer of an item's error
text, the only "Erneut" retry button and the only rate-limit countdown — was
never imported anywhere, so `retryItem`, `removeItem` and `clearCompleted` were
unreachable at runtime. On a terminal rejection the store purged the blob and
wrote a clear German reason into `entry.error` "so the UI shows a clear reason".
There was no such UI. And `uploadBadgeCount` counted only pending/uploading, so
the badge decremented exactly as if the upload had succeeded.

Mount the queue on /upload, toast the reason immediately (the flow sends the user
to /feed straight after staging, so the list alone would still miss them), and
count blocked/error in the badge so a failure can't read as success.

Tests: 02-upload/exif-orientation uploads a 40x20 fixture tagged Orientation=6
and asserts both derivatives come back PORTRAIT, with a sanity check that the
source really is stored landscape. 02-upload/rejection-visible bans the uploader
between staging and sending, then asserts the toast, the queue row with the
server's reason, and that the item is still counted.

Note: 02-upload/quota's 4 failures are pre-existing and unrelated — see the next
commit.

──────── docs(backup): make the backup commands work; fix the e2e/prod divergences

Backup. Both documented commands failed on the shipped stack, and the sentence
explaining them was wrong too:

- `pg_dump $DATABASE_URL` — `DATABASE_URL` is only ever in the compose
  environment, never an operator's shell, and it points at `db:5432`, which is
  compose-internal DNS. The app image has no postgres client either.
- `> /media/backups/…` — `/media` is a named volume mounted inside the app
  container, not a host path, and nothing ever creates a `backups` subdirectory.
- `rsync /opt/eventsnap/media/` — that path does not exist anywhere.
- "a single path to back up" — false, and dangerously so: exports were moved to
  their own `exports_data` volume precisely so a keepsake (which contains every
  photo in the event) can't be served off the media tree. Backing up only
  `media_data` silently loses every generated keepsake.

Rewritten as three commands — db via `docker compose exec -T db pg_dump`, and one
`docker run … tar` per volume — all verified against the running stack. The
volume mounts use `/src`, not `/media`: I hit the footgun while testing this.
Docker pre-populates an EMPTY volume from the image's own directory and chowns it
to match, so `-v media_data:/media alpine` tars alpine's cdrom/floppy/usb, writes
them into the volume, and leaves it root-owned so the non-root app can no longer
write. Mounting where the image has nothing avoids all of it. Documented inline
so the next person doesn't rediscover it.

Also correct the architecture notes: `/media/*` no longer routes to the backend
(that static tree was removed as a gating bypass), and `exports_data` was missing
from the volume list — the one volume an operator most needs to know about.

e2e stack: add the `EXPORT_PATH` + `/exports` volume it was missing. The file
says "mirrors production layout"; without these, exports landed on the container's
writable layer at the default path, so export-leak and export-video wrote real
archives into ephemeral storage and the "exports live outside media" invariant
was never actually exercised.

Pre-existing red test, unrelated to the audit: all four 02-upload/quota tests
have been failing since 4464147 "stop /me/quota leaking raw disk to guests"
(2026-07-19), which post-dates the spec's last edit. `setLimitTo` calibrated
`quota_tolerance` from `free_disk_bytes` read through the GUEST's token — a field
that commit deliberately zeroes for non-staff. Dividing by it yields a NaN
tolerance, so every test in the block died in the helper. Read the calibration
inputs through a staff token and keep reading the ceiling back through the guest,
whose limit is the thing under test.

──────── test(e2e): document why WebKit can't run the client-queue upload tests

Three 02-upload tests have been failing on `webkit-iphone` on main — `02-upload`
was already in that project's testMatch, so this is pre-existing red, not
something the audit work introduced.

Root cause is the harness, not the app. Playwright's Linux WebKit build cannot
store Blobs in IndexedDB at all: `put()` fails with "UnknownError: Error
preparing Blob/File data to be stored in object store". Confirmed it is not
about how Playwright delivers files — a Blob constructed in-page with
`new Blob([bytes])` fails identically, while Chromium stores both that and a
`setInputFiles` File without complaint.

That breaks every test driving the composer (FAB → UploadSheet → /upload →
submit), because `handleSubmit` awaits `addToQueue`, which persists the file
before it can navigate. The symptom is a submit button stuck on "Wird
hochgeladen…" and a timeout waiting for /feed — which reads like an app hang and
cost real time to run down.

Skip those four (the three above plus the new rejection-visible) on WebKit only,
behind a named helper carrying the full explanation, so the next person gets the
answer instead of the investigation. Deliberately narrow: WebKit still runs every
API-driven upload test, all of 01-auth, 03-feed and 06-export — including the
keepsake download, which only WebKit can meaningfully verify. Chromium continues
to run all 14.

Worth being explicit, since these tests exist to protect iOS: real Safari
supports Blobs in IndexedDB, so this is NOT evidence that the offline upload
queue is broken on the platform. It does mean that guarantee is currently
unverifiable in CI and rests on Chromium coverage plus manual device testing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 19:14:37 +02:00
MechaCat02
a77c2ddc00 Merge branch 'fix/prod-readiness-healthcheck-domain'
Some checks failed
Checks / Backend — cargo test + clippy + fmt (push) Failing after 1m14s
Checks / Frontend — vitest + svelte-check (push) Failing after 5m50s
Checks / E2E — typecheck + lint (push) Failing after 48s
E2E / Playwright E2E (chromium-desktop) (push) Failing after 6m44s
E2E / Cross-UA smoke matrix (push) Failing after 4m32s
Audit / cargo audit (backend) (push) Failing after 12m17s
Audit / npm audit (frontend) (push) Successful in 43s
2026-07-19 18:42:36 +02:00
MechaCat02
40c6fd2ccb fix(deploy): unblock production bring-up (healthchecks + Caddy DOMAIN)
Two issues would each stop a clean production `docker compose up` for the event:

1. Healthchecks probed http://localhost:{3000,3001}, but the app and frontend
   bind IPv4 (0.0.0.0) while `localhost` resolves to ::1 (IPv6) first inside the
   container — so the probe got "connection refused" and neither container ever
   turned healthy. Caddy is gated on `condition: service_healthy` for both, so on
   a fresh boot it would block forever and nothing gets served. Switch both probes
   to 127.0.0.1. (Verified: both containers now report healthy.)

2. The prod caddy service never received DOMAIN, so the Caddyfile's `{$DOMAIN}`
   site address expanded to empty — malformed site block, no TLS, no serving. Add
   `environment: { DOMAIN: ${DOMAIN} }` to the caddy service.

Also make .env.example honest and event-ready:
- Add DATABASE_MAX_CONNECTIONS (real env lever; recommend 30 for ~100 guests).
- The DEFAULT_* upload/rate/capacity vars are NOT read from env — they are seeded
  into the DB config table and managed at runtime via the admin dashboard. Replace
  the misleading entries (e.g. upload rate showed 10; live value is 100 via
  migration 015) with a note pointing to the admin UI and the real seeded defaults.
- Document that raising COMPRESSION_WORKER_CONCURRENCY also needs the app memory
  limit raised (ffmpeg), so a video burst can't OOM the box.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 18:42:32 +02:00
MechaCat02
e69ec4d736 Merge branch 'feat/hide-comments-when-disabled' 2026-07-19 18:27:05 +02:00
MechaCat02
3fb1b5d80d feat(comments): hide every comment mention when COMMENTS_ENABLED=false
The kill-switch previously left comment UI/text visible in several places.
Sweep the whole surface so a comments-off instance shows no trace:

Frontend (main app):
- VirtualFeed grid tile: gate the comment button/count on $commentsEnabled
  (was ungated — the only feed surface still showing it).
- admin stats: hide the "Kommentare" count card.
- UploadSheet "Uploads geschlossen" notice, host ban-modal description, and
  host/admin unban confirmations: drop the "…und kommentieren" wording.
- export page: drop "Kommentaren" from the keepsake description.

Keepsake export (had no concept of the flag):
- export.rs: thread comments_enabled into the exported data (ViewerEvent),
  wired through spawn_export_jobs/recover_exports and their call sites
  (host.rs, main.rs).
- export-viewer: gate comment counts (list + grid) and the lightbox comments
  section; older exports without the field default to enabled (?? true).

Backend still 403s comment writes when disabled (unchanged) — this is the UI
half so stale clients and archives match.

Verified on the running stack (COMMENTS_ENABLED=false): /event reports
comments_enabled=false, a regenerated keepsake embeds "comments_enabled": false
with no comment UI, and the uploads-closed notice renders "…ansehen und liken."

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 18:26:59 +02:00
MechaCat02
e1ca9d192f Merge branch 'feat/diashow-completeness-display' 2026-07-19 17:53:57 +02:00
MechaCat02
5009590882 feat(diashow): guarantee all eligible photos shown + 2048px display derivative
Diashow completeness rewrite so every eligible upload is shown regardless of
bursts, disconnects, or library size:

- queue.ts: SlideQueue with live/shuffle queues, allKnown map, recentlyShown
  ring; merge(dedup, live-first), remove/removeByUser (prunes recentlyShown),
  knownIds for reconcile-eviction. Adds queue.test.ts (burst/completeness/race).
- diashow/+page.svelte: reconcile (full paginate + evict, pre-scan snapshot to
  spare concurrent uploads) on mount/reconnect/periodic; catchUpNew paginate-
  until-known for bursts with debounced maxWait; hard-cut removals; decode
  timeout + candidate fallback + bounded skip so a broken image never stalls.

New ~2048px "display" derivative for big-screen sharpness, decoupled from the
data-saver preview (800px) used on phones:

- migration 016: upload.display_path + v_feed rebuilt (DROP+CREATE, not REPLACE,
  to slot the column beside preview/thumbnail).
- compression: generate_image_derivatives emits preview+display (downscale-only
  guard, no upscaling); backfill_missing_display regenerates on startup (safe:
  logs on error, never soft-deletes).
- upload.rs/main.rs: GET /upload/{id}/display (mirrors preview auth/cache),
  /media/displays direct-serve blocked.
- feed.rs + types.ts: display_url in feed/delta DTOs.
- diashow candidate chain: display -> original -> preview.

Verified on the running stack: migration applied, 10/10 existing images
backfilled (2048px cap honoured, small images not upscaled), /display serves
200, /feed returns display_url, diashow cycles.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 17:53:53 +02:00
MechaCat02
d9738a4cb9 Merge branch 'fix/prod-env-hardening'
Harden production env handling: single-quote the bcrypt ADMIN_PASSWORD_HASH so
Compose/dotenvy don't corrupt it, and pin MEDIA_PATH to the /media mount.
2026-07-19 16:46:55 +02:00
MechaCat02
669a191968 fix(deploy): harden prod env for bcrypt hash and media path
Two latent production pitfalls, both proven by probing Compose's env_file resolution:

- A bcrypt ADMIN_PASSWORD_HASH is full of $; Compose env_file interpolation AND
  dotenvy variable substitution both eat the $… segments (as unset vars), corrupting
  the hash so every admin login 401s. Single-quote it so both read it literally —
  verified correct for env_file (container) and dotenvy 0.15.7 strong-quote (native).
  Updated .env.example with the quotes + a warning comment.

- MEDIA_PATH: pin it to the /media mount in the app 'environment:' block so a stray
  host path in .env can't leak in and make every upload 500 with EACCES. environment
  overrides env_file, so the container is always correct.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 16:46:38 +02:00
MechaCat02
a1733b03d5 Merge branch 'feature/diashow-theming-config'
Runtime colour theme + COMMENTS_ENABLED switch, diashow transition/fullscreen/wake-lock
work, quota disk-leak fix, and docker-compose.dev.yml container-env corrections.
2026-07-19 15:23:04 +02:00
MechaCat02
4026648f98 fix(diashow): working transitions + slide/push, fullscreen, activity controls
Transitions never actually animated: Svelte scopes @keyframes names but does NOT
rewrite animation references in inline style attributes, so the inline
'animation: crossfade-in ...' never matched its scoped keyframe — a hard cut, no fade
(Ken Burns' zoom was dead too). Move the animation into scoped classes and pass dynamic
values via CSS custom properties.

- Real two-layer crossfade: hold the outgoing frame opaque beneath the incoming one
  (which is decode-gated), so there's no black flash and no decode pop.
- New slide transitions (from below/left/right) as one reusable component; the previous
  frame is pushed off the opposite edge in step (a conveyor/push), easeInOutSine 1s.
- Fullscreen toggle button (Fullscreen API) + 'f' shortcut; Esc in fullscreen no longer
  also navigates away.
- Controls now reveal on pointer/keyboard activity and fade when idle (cursor hides too)
  instead of being always visible.
- wakelock: null the sentinel on the OS 'release' event so re-acquire-on-visible
  actually fires (previously the screen could sleep after the tab was first hidden).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:22:15 +02:00
MechaCat02
44641473ea fix(quota): stop /me/quota leaking raw disk to guests
The per-user quota widget was shown to everyone and the /me/quota payload returned
free_disk_bytes (raw server free space) and active_uploaders to any authenticated
guest. Gate the widget to staff (host/admin) on the upload and account pages, and zero
the server-wide telemetry fields for non-staff in the handler. Guests still get their
own used/limit so enforcement stays transparent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:22:03 +02:00
MechaCat02
57a907eca5 feat(theme+comments): runtime colour theme and COMMENTS_ENABLED switch
Colour theme is configurable at runtime from two seed colours (brand + accent);
neutrals stay fixed for contrast safety. Tailwind v4 var()-based tokens let a
:root:root override recolour everything with no rebuild; the 50->950 ramps are
derived via a color-mix ladder. Config lives in the DB config table (admin UI:
Config > Farbschema, presets + custom pickers + live preview), served on the public
/event endpoint with env defaults (THEME_PRESET/PRIMARY/ACCENT), propagated live via
event-updated SSE, and cached in localStorage for a no-flash boot. The keepsake export
mirrors the same ladder in Rust so offline archives match the event theme.

COMMENTS_ENABLED (env, default true) is a boot-time kill-switch: the backend rejects
new comments with 403 and the frontend hides the comment button (feed card) and
panel/composer (lightbox). Existing comments stay in the DB, hidden, and return when
re-enabled.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:21:55 +02:00
MechaCat02
6e0a760271 chore(dev): correct container env in docker-compose.dev.yml
Running in a container picked up host-oriented values from .env via env_file:
- MEDIA_PATH pointed at a host path that doesn't exist in the container, so every
  upload 500'd with EACCES; point it at the /media volume mount.
- ADMIN_PASSWORD_HASH lost its $-delimited bcrypt segments to Compose interpolation
  (admin login 401'd); re-supply it with $$ escaping.
- COMMENTS_ENABLED=false to smoke-test the comment kill-switch.

NOTE: production compose + .env carry the same MEDIA_PATH / admin-hash pitfalls.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:21:44 +02:00
MechaCat02
0abf413693 Merge branch 'test/burst-queue'
Add an e2e regression covering the client upload queue under a 10-20 photo
burst: serial drain, IndexedDB persistence, all-land, and reload-resume.
2026-07-18 19:54:44 +02:00
MechaCat02
002355ba40 test(e2e): client upload-queue under a realistic burst
Cover the scenario the server-side load test could not — a guest multi-selecting
10-20 photos at once — by driving the real client path (UploadSheet → /upload →
addToQueue → processQueue → XHR) rather than hitting POST /upload directly.

Two tests assert the properties that keep a guest's photos from being lost:
serial per-device draining (one upload in flight at a time), IndexedDB
persistence (blobs kept until upload, dropped after), every file landing
server-side, and a hard reload mid-burst resuming the rest via loadQueue()
with no photo lost. Injects small per-upload latency via route interception so
the serial drain is observable and the mid-burst reload window is reliable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 19:54:36 +02:00
MechaCat02
6155b4123d Merge branch 'release/v0.16.0'
Wedding redesign, single-file offline keepsake, diashow SSE fix,
upload-rate 10→100, and a 100-guest load-test harness.
2026-07-18 17:29:30 +02:00
MechaCat02
7758270cac chore: shared permission allowlist, screenshot script, gitignore
Add project .claude/settings.json (read-only cargo check/clippy + git diff
allowlist; personal settings.local.json stays gitignored). Add e2e/shots.mjs
(one-off mobile screenshot seeder). Ignore load-test run artifacts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 17:29:20 +02:00
MechaCat02
9b8698f86b test(loadtest): 100-guest / 1000-image stress harness
HTTP-level load driver simulating ~100 guests uploading ~1000 images in bursts
over a window, plus SSE viewers and one real browser on /diashow. Correlates
upload→upload-processed (pipeline latency), waits for the compression backlog
to drain against DB ground truth, and emits per-status/latency metrics with
pass/fail flags. Includes a realistic-JPEG generator and a diashow-SSE
regression check (confirm-diashow-fix.mjs). Run artifacts are gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 17:29:20 +02:00
MechaCat02
3c3a7d0082 fix(diashow): open SSE on mount; raise upload rate 10→100/hour
diashow only subscribed to SSE events but never called connectSse(), so a
kiosk/projector opening /diashow directly (not via /feed) never opened the
EventSource — the showcase display got a one-time /feed snapshot and no live
updates, showing "Noch keine Beiträge" forever when turned on before any
photos. Open the stream in onMount (idempotent) and close it in onDestroy,
mirroring the feed page.

Raise the default upload_rate_per_hour from 10 to 100 (migration 015, scoped
to installs still on the old default so admin overrides are preserved). Guests
routinely upload bursts of 10-20 photos; the old default throttled the first
burst. Also update the code fallback and the test-mode reseed.

Both verified end-to-end against the docker test stack.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 17:29:08 +02:00
MechaCat02
3654aca18b feat(export): single-file offline keepsake viewer + guest download
Rebuild the guest "keepsake" (offline HTML gallery) so it renders when the
extracted index.html is opened over file://. Browsers block external ES-module
scripts and fetch() at origin null, so a normal multi-file SvelteKit build shows
a blank window. Build the viewer as one self-contained index.html (inlined
JS/CSS via vite-plugin-singlefile, standalone non-SvelteKit entry) and inject
the data as window.__EXPORT_DATA__; the backend embeds the built viewer via
include_dir! and writes the data global into index.html when zipping.

Also surface the guest-facing download: a feed banner and the BottomNav Export
tab, both shown once the host releases the export.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 17:28:54 +02:00
MechaCat02
f243bfe89a feat(ui): elegant wedding white/silver/gold redesign
Reskin the whole app via design tokens in tailwind-theme.css (remapped
color ramps) plus a define-once component layer in lib/styles/components.css
(.btn, .card, .input, .chip, .badge, .sheet, …). Buttons are muted/outlined
gold rather than flat fills. Self-host Inter + Fraunces (woff2) under the
existing font-src 'self' CSP. Restyle the shared components and the account,
admin, host, join, recover and upload screens against the new tokens.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 17:28:42 +02:00
233 changed files with 24793 additions and 2631 deletions

9
.claude/settings.json Normal file
View File

@@ -0,0 +1,9 @@
{
"permissions": {
"allow": [
"Bash(cargo check *)",
"Bash(cargo clippy *)",
"Bash(git --no-pager diff *)"
]
}
}

View File

@@ -1,7 +1,28 @@
# ── Domain ──────────────────────────────────────────────────────────────────── # ── Domain ────────────────────────────────────────────────────────────────────
# Public domain Caddy will serve and obtain a TLS certificate for. # Public domain Caddy will serve and obtain a TLS certificate for.
#
# The DNS A record must already point at this server BEFORE the first `up -d`: Caddy
# requests a certificate on boot, and Let's Encrypt allows only 5 failed validations per
# hostname per hour. Never delete the caddy_data volume — it holds the certificate and
# the ACME account key.
DOMAIN=my-event.example.com DOMAIN=my-event.example.com
# ── Image version ─────────────────────────────────────────────────────────────
# Tag pulled for the `app` and `frontend` services (docker-compose.yml). Production runs
# prebuilt images from the registry and never compiles — see DEPLOYMENT_RUNBOOK.md.
# Always an immutable tag, never `latest`: rollback is `EVENTSNAP_VERSION=<previous>`
# + `docker compose up -d`, which works offline if that image is still resident locally.
#
# ⚠ THIS TAG DOES NOT EXIST YET. The newest git tag is v0.12.0; v0.13.0 is the release you
# cut for the event. Build and push it (plus its identical rollback twin v0.13.0-a) BEFORE
# the first `docker compose up -d` — see DEPLOYMENT_RUNBOOK.md §6 (build) and §9 (rollback).
# Copying this file and starting the stack without that step fails with `manifest unknown`.
#
# Do NOT "fix" this by dropping back to v0.12.0: no image was ever built for it, and a
# 6-migration tree booting against a 31-migration database returns VersionMissing and
# crash-loops forever behind a live Caddy. §9 covers this in full.
EVENTSNAP_VERSION=v0.13.0
# ── App server ──────────────────────────────────────────────────────────────── # ── App server ────────────────────────────────────────────────────────────────
APP_PORT=3000 APP_PORT=3000
# Set to `production` in real deployments. This activates the secret guard that # Set to `production` in real deployments. This activates the secret guard that
@@ -12,10 +33,43 @@ APP_ENV=production
# ── Database ────────────────────────────────────────────────────────────────── # ── Database ──────────────────────────────────────────────────────────────────
# Set a strong password and keep it in sync between DATABASE_URL and # Set a strong password and keep it in sync between DATABASE_URL and
# POSTGRES_PASSWORD. Generate one with: openssl rand -hex 24 # POSTGRES_PASSWORD. Generate one with: openssl rand -hex 24
#
# SET THIS BEFORE THE FIRST `docker compose up -d`. Postgres reads POSTGRES_PASSWORD
# only when it initialises its data directory, on that very first boot. Change it
# afterwards and the app authenticates with the new password against a volume still
# holding the old one — a permanent restart loop ("password authentication failed").
# The only ways out are restoring the old password or `docker compose down -v`, which
# deletes the database, the media and the exports. In production the app refuses to
# boot while this is still the placeholder below, so it cannot be missed by accident.
DATABASE_URL=postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap DATABASE_URL=postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap
POSTGRES_USER=eventsnap POSTGRES_USER=eventsnap
POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password
POSTGRES_DB=eventsnap POSTGRES_DB=eventsnap
# Connection pool size. The code default is 15 (DEFAULT_MAX_CONNECTIONS in backend/src/db.rs),
# and docker-compose.yml pins this value in `app.environment` so an edit here cannot reach the
# container. That pin is deliberate: since the value became boot-FATAL when unparseable — so an
# operator tuning a knob that never took effect gets told, instead of silently staying on the
# default — a stray quote or a trailing inline comment in `.env` would crash-loop the app behind
# a live Caddy. Change the pin in compose, not this line.
#
# SIZE IT TO THE CORES, NOT TO THE GUESTS. The earlier advice here was ~30, reasoned from
# "~100 guests polling the feed at once" back when a feed page cost ~449 ms and connections
# were spent waiting. Migration 024 replaced the feed view's GROUP BY with scalar subqueries
# and a page now costs well under a millisecond, so concurrency is no longer where the time
# goes. On a 2 vCPU box 30 simultaneous queries cannot run — they queue on the CPU instead of
# on the pool, which is the same wait wearing a different hat, and 30 Postgres backends plus
# shared_buffers is snug in the 1G that docker-compose.yml allots `db`.
#
# 15 on 2 vCPU / 4 GB. Raise toward 30 only alongside more cores AND a bigger `db` memory
# limit — an OOM in Postgres doesn't degrade one feature, it takes the whole event down.
DATABASE_MAX_CONNECTIONS=15
# Log level: see the "Logging" section near the bottom of this file.
#
# Defined THERE and nowhere else, deliberately. This file used to assign RUST_LOG twice —
# once here and once there — and Compose takes the LAST assignment, so editing this line to
# `debug` to chase a problem during the event changed nothing at all, silently. A key that
# appears twice in a .env is a trap regardless of which value is better.
# ── Authentication ──────────────────────────────────────────────────────────── # ── Authentication ────────────────────────────────────────────────────────────
# Generate with: openssl rand -hex 64 # Generate with: openssl rand -hex 64
@@ -23,11 +77,25 @@ JWT_SECRET=change_me_to_a_random_64_byte_hex_string
SESSION_EXPIRY_DAYS=30 SESSION_EXPIRY_DAYS=30
# Admin dashboard password (bcrypt hash). # Admin dashboard password (bcrypt hash).
# Generate with: htpasswd -bnBC 12 "" yourpassword | tr -d ':\n' # Generate with an image the stack already pulls (htpasswd needs apache2-utils, which
ADMIN_PASSWORD_HASH=$2y$12$placeholder_replace_me # a stock VPS does not have):
# docker run --rm caddy:2-alpine caddy hash-password --plaintext 'yourpassword'
# IMPORTANT: keep the SINGLE QUOTES. A bcrypt hash is full of `$` (e.g. $2b$12$…$…),
# and both Docker Compose's env_file interpolation and dotenvy's variable substitution
# would otherwise eat the `$…` segments (reading them as unset vars) and corrupt the
# hash — every admin login then 401s. Single quotes make both read it literally.
ADMIN_PASSWORD_HASH='$2y$12$placeholder_replace_me'
# ── Event ───────────────────────────────────────────────────────────────────── # ── Event ─────────────────────────────────────────────────────────────────────
EVENT_NAME=Max & Maria's Wedding # DOUBLE-QUOTED, and it matters. Compose's env_file parser reads `Max & Maria's Wedding`
# unquoted just fine — but the runbook also tells you to `set -a; . ./.env; set +a` in a plain
# shell, and POSIX `sh` aborts on the apostrophe with "Unterminated quoted string" (rc=2).
# Everything defined BELOW this line is then left unset, silently: the hourly pg_dump cron in
# §10.2 does exactly this, so it would exit before ever writing a backup, every hour, into a log
# nobody reads. Double quotes are read identically by both parsers (verified) — keep them, and
# keep them double, since single quotes would make a literal `$` in a name survive but are what
# `ADMIN_PASSWORD_HASH` above needs for the opposite reason.
EVENT_NAME="Max & Maria's Wedding"
EVENT_SLUG=max-maria-2026 EVENT_SLUG=max-maria-2026
# ── Storage ─────────────────────────────────────────────────────────────────── # ── Storage ───────────────────────────────────────────────────────────────────
@@ -36,19 +104,113 @@ MEDIA_PATH=/media
# /media is publicly served, so exports here would be downloadable without auth. # /media is publicly served, so exports here would be downloadable without auth.
EXPORT_PATH=/exports EXPORT_PATH=/exports
# ── Upload limits ───────────────────────────────────────────────────────────── # ── Runtime settings (upload limits, rate limits, capacity) ───────────────────
DEFAULT_MAX_IMAGE_SIZE_MB=20 # NOTE: These are NOT environment variables. Upload size caps, rate limits, guest
DEFAULT_MAX_VIDEO_SIZE_MB=500 # count and quota tolerance are stored in the database `config` table (seeded once
# at first boot) and changed at runtime from the ADMIN DASHBOARD — the backend does
# ── Rate limiting ───────────────────────────────────────────────────────────── # not read them from .env. Setting them here has no effect. Current seeded defaults:
DEFAULT_UPLOAD_RATE_PER_HOUR=10 # upload rate 100 / hour / guest (raised from 10 by migration 015)
DEFAULT_FEED_RATE_PER_MIN=60 # feed rate 60 / minute
DEFAULT_EXPORT_RATE_PER_DAY=3 # export rate 3 / day
# max image size 20 MB
# ── Capacity ────────────────────────────────────────────────────────────────── # max video size 500 MB
DEFAULT_ESTIMATED_GUEST_COUNT=100 # estimated guests 100
# Fraction of total storage that triggers the "low storage" warning (0.01.0) # quota tolerance 0.75 (see below — NOT a warning threshold)
DEFAULT_QUOTA_TOLERANCE=0.75 # Adjust these in the admin UI before the event if needed.
#
# quota_tolerance is the MULTIPLIER IN THE PER-USER QUOTA FORMULA, not the point at
# which anything warns you:
#
# divisor = max(active_uploaders, estimated_guest_count, 1)
# per_user_limit = max(floor(free_disk * quota_tolerance / divisor), 500 MiB)
#
# estimated_guest_count is a FLOOR ON THE DIVISOR, not decoration — it is a live knob
# (upload::quota_limit_bytes). Earlier drafts of this file and the runbook both omitted
# it and told operators it was inert; it is not.
#
# It is recomputed against LIVE free space on every upload, so in principle it self-
# throttles: guests converge on a fixed point at tolerance/(1+tolerance) of the free space
# you started with — 43% at 0.75.
#
# ON THIS BOX THAT FIXED POINT NEVER BINDS, and it is worth knowing which knob actually
# stops the disk filling. The arithmetic above used to be quoted as "~30 GB of a fresh
# 70 GB", which is an 80 GB CX33; this deploys to a CX22 with 40 GB. At ~28 GB free and
# estimated_guest_count = 100 flooring the divisor, the formula yields ~210 MB per guest —
# BELOW the 500 MiB floor — so every guest is granted the floor and the per-user quota
# stops bounding aggregate growth at all.
#
# What actually bounds it is the keepsake preflight in upload.rs: uploads are refused once
# free < media x 1.1 x 2 + 10 GB, which on 40 GB lands at ~8 GB of media (README, "Sizing
# the disk"). So if a guest reports being blocked, the number to look at is total media,
# not this one.
#
# Raising this still AUTHORISES GUESTS TO FILL MORE OF THE DISK on a larger box, and it
# still eats the headroom the keepsake needs — Gallery.zip and Memories.zip are each
# roughly a second copy of every original (both store media uncompressed). Budget for
# media + 2x media, or move exports to their own volume.
#
# 0.75 is the tested default. Lower it if the box is tight; raise it only if you have
# provisioned export headroom separately.
# ── Workers ─────────────────────────────────────────────────────────────────── # ── Workers ───────────────────────────────────────────────────────────────────
# Number of parallel media compression workers. Default 2. Boot-time only.
#
# CORRECTION TO EARLIER GUIDANCE: this used to say "each worker can run an ffmpeg
# transcode, so raise the app memory limit to ~2G if you set 4". There is NO video
# transcode anywhere in this codebase — services/video.rs runs
# `ffmpeg -ss <t> -i <src> -vframes 1 -vf scale=...`, a single poster frame, and video
# originals are stored and served byte-for-byte. Poster extraction costs ~150-250 MB
# for a moment; it is not the constraint.
#
# The real memory consumer is the IMAGE path. `image` 0.25's resize builds an Rgba32F
# intermediate at 16 BYTES PER PIXEL, sized (source_width x target_height) — which the
# 256 MiB decode guard in imaging.rs does NOT cover. Peak per photo, decode + the 2048px
# display resize: ~145 MB at 12 MP, ~223 MB at 24 MP, ~354 MB at 48 MP.
#
# So on a 2 vCPU / 4 GB box (e.g. Hetzner CX22) KEEP THIS AT 2:
# * concurrency 4 would put two giants at ~1.5 GB against the 1G app limit — OOM.
# * and app=2G + db=1G + frontend/caddy 256M each + ~370 MB of OS/Docker exceeds the
# ~3910 MiB a "4 GB" VM actually reports. Raising the limit oversubscribes the host.
# 4 is only reasonable on the 4 vCPU / 8 GB box README.md documents.
#
# The "two 48 MP photos at once" worst case this number used to be sized against is no
# longer reachable: compression.rs takes an EXCLUSIVE `heavy` permit for any job whose
# estimated peak exceeds HEAVY_IMAGE_BYTES (150 MiB), so two giants serialise no matter what
# this is set to. What concurrency 2 now buys is two ORDINARY phone photos in parallel
# (~145 MB peak each), which is both memory-safe and short enough not to starve the two
# tokio worker threads a 2 vCPU box gets.
#
# Do NOT drop this to 1 hoping to protect the CPU. It halves throughput on the common light
# path for a heavy path that is already serialised, and a longer compression backlog means
# more feed tiles served from full-size originals (VirtualFeed falls back to /original while
# derivatives are pending) — trading a little CPU for a lot of venue-wifi bandwidth.
#
# Throughput at 2 is not the bottleneck anyone thinks it is: ~2.5s per 12 MP photo, so
# 100 photos is ~250 CPU-seconds spread over an entire evening.
COMPRESSION_WORKER_CONCURRENCY=2 COMPRESSION_WORKER_CONCURRENCY=2
# ── Comments ──────────────────────────────────────────────────────────────────
# Master switch for the comment feature. Boot-time only (NOT in the admin UI), so it
# needs a `docker compose up -d` to apply. Anything other than false/0/no/off is on.
#
# When false the backend rejects NEW comments with 403 and the frontend hides the whole
# comment UI, including in the offline keepsake viewer. Likes and captions are entirely
# separate features and are unaffected. Existing comments stay in the database (hidden),
# so flipping it back restores them.
#
# Note it gates POSTING only: GET /upload/{id}/comments still serves already-existing
# comments, and the keepsake's data.json still embeds their text. Irrelevant if the flag
# is off from the first boot, since no comment can ever have been written.
# NOTE for the current deployment: `docker-compose.yml` PINS this to "false" on the app
# service, and `environment` overrides `env_file` — so changing it here has no effect in
# production. Remove that line from the compose file first if you want comments back.
COMMENTS_ENABLED=true
# ── Logging ───────────────────────────────────────────────────────────────────
# SET THIS IN PRODUCTION. Without it the app falls back to
# `eventsnap_backend=debug,tower_http=debug` (see main.rs), and with TraceLayer that is a
# debug line per HTTP request — including every preview and thumbnail fetch. Combined with
# Docker's json-file driver it writes to the same filesystem as the database and the media.
# docker-compose.yml caps each service's logs at 30 MB; this keeps the volume sane in the
# first place. The e2e stack has always used exactly this value.
RUST_LOG=eventsnap_backend=info,tower_http=warn

View File

@@ -102,6 +102,44 @@ jobs:
working-directory: ./frontend working-directory: ./frontend
run: npm run format:check run: npm run format:check
export-viewer:
name: Keepsake viewer — builds, self-contained, committed artifact in sync
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
cache: 'npm'
cache-dependency-path: 'frontend/export-viewer/package-lock.json'
- name: Install deps
working-directory: ./frontend/export-viewer
run: npm ci || npm install
# Two things nothing else in CI covered, both of which ship a broken keepsake silently.
#
# 1. The build's own self-contained guard (`inlineThemeFonts`) is the only thing standing
# between an added theme asset and a viewer that reaches for files on the guest's disk.
# It is a build-time `this.error`, so it only fires when somebody runs this build — and
# no workflow, Dockerfile or script did. It could sit disarmed indefinitely.
#
# 2. `backend/static/export-viewer/index.html` is COMMITTED and compiled into the binary with
# `include_dir!`. A viewer source change merged without a manual rebuild ships the stale
# artifact, and nothing anywhere would say so. `git diff --exit-code` is the check.
- name: Build the standalone viewer
working-directory: ./frontend/export-viewer
run: npm run build
- name: Committed artifact matches a clean rebuild
run: |
if ! git diff --exit-code -- backend/static/export-viewer/; then
echo "::error::backend/static/export-viewer/ is out of date with frontend/export-viewer/."
echo "Run 'npm run build' in frontend/export-viewer and commit the result."
exit 1
fi
e2e-typecheck: e2e-typecheck:
name: E2E — typecheck + lint name: E2E — typecheck + lint
runs-on: ubuntu-latest runs-on: ubuntu-latest

View File

@@ -7,7 +7,7 @@ on:
jobs: jobs:
e2e: e2e:
name: Playwright E2E (chromium-desktop) name: Playwright E2E (chromium + webkit)
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30 timeout-minutes: 30
steps: steps:
@@ -25,7 +25,7 @@ jobs:
- name: Install Playwright browsers - name: Install Playwright browsers
working-directory: ./e2e working-directory: ./e2e
run: npx playwright install --with-deps chromium run: npx playwright install --with-deps chromium webkit
- name: Bring up the test stack - name: Bring up the test stack
working-directory: ./e2e working-directory: ./e2e
@@ -54,6 +54,22 @@ jobs:
working-directory: ./e2e working-directory: ./e2e
run: npm run test:e2e -- --project=chromium-mobile run: npm run test:e2e -- --project=chromium-mobile
# iOS Safari is the app's stated primary user (a wedding guest opening a QR link), and
# WebKit is the ONLY engine here that reproduces two of its behaviours:
# - it enforces X-Frame-Options on the download iframe, so a site-wide `DENY` makes the
# keepsake download silently do nothing. Blink hands attachments to the download
# manager first and never notices. That shipped once already.
# - it abandons a <video> load without a 206 response to its Range probe.
# Both regressions are invisible to every Chromium project, so running WebKit is what
# actually gates them on a PR rather than on someone remembering to test locally.
#
# The project is scoped in playwright.config.ts to the journeys a guest walks
# (01-auth, 02-upload, 03-feed, 06-export); four IndexedDB-blob tests skip themselves
# there — see helpers/webkit.ts for why that is the harness and not the app.
- name: Run E2E tests (webkit / iOS)
working-directory: ./e2e
run: npm run test:e2e -- --project=webkit-iphone
- name: Upload Playwright report - name: Upload Playwright report
if: failure() if: failure()
uses: actions/upload-artifact@v4 uses: actions/upload-artifact@v4

15
.gitignore vendored
View File

@@ -13,8 +13,16 @@ frontend/build/
frontend/export-viewer/node_modules/ frontend/export-viewer/node_modules/
frontend/export-viewer/.svelte-kit/ frontend/export-viewer/.svelte-kit/
# Media uploads (mounted volume in production) # Media uploads. In production these live in the `media_data` DOCKER VOLUME, never in the
media/ # working tree — so this pattern is anchored to the repo root and exists only for a local
# bind-mount experiment.
#
# It used to read `media/`, unanchored, which matches a directory of that name at ANY depth.
# The only one in the repo is `e2e/fixtures/media/`, so the rule's entire practical effect was
# to keep every E2E fixture untracked: a fresh clone got the specs and none of the images or
# videos they read. `.github/workflows/e2e.yml` does a plain checkout and generates nothing, so
# the committed CI job could not have run the upload, video or export suites at all.
/media/
# Playwright E2E suite — runtime artifacts (the suite itself is committed) # Playwright E2E suite — runtime artifacts (the suite itself is committed)
e2e/node_modules/ e2e/node_modules/
@@ -29,3 +37,6 @@ e2e/.env.test
# OS # OS
.DS_Store .DS_Store
Thumbs.db Thumbs.db
# Claude Code personal (per-user) settings — shared settings.json IS committed
.claude/settings.local.json

145
Caddyfile
View File

@@ -1,3 +1,33 @@
{
servers {
timeouts {
# Slowloris defence, at the layer that can actually apply it.
#
# There was no read or write timeout anywhere, so a client could open a POST, send one
# byte a minute, and hold a connection, a tokio task and a `.tmp` file indefinitely —
# and the upload sweeper is keyed on mtime precisely so a live upload never ages out,
# so ten such connections consumed disk the upload gate could not see.
#
# read_header is tight: a legitimate client sends its headers in one go.
read_header 10s
# read_body is GENEROUS but present. It was omitted on the reasoning that "a slow body
# still has to actually send bytes" — which is an argument about disk, and disk is not
# the scarce resource here. `upload_admission` budgets concurrent bodies at 4096 MiB and
# reserves the DECLARED cap, so a `video/*` upload reserves 500 MiB: eight connections
# that stall mid-body hold the entire budget, every other guest waits 20s and gets a
# 503, and it never recovers on its own because the permit is held until the handler
# returns. That needs no attacker — eight guests starting real videos and then walking
# out of AP range does it, and TCP will not reap those sockets for hours.
#
# 30m carries a 500 MB video at ~2.2 Mbit/s sustained, which is well under venue wifi
# and under most cellular, so it does not fail the uploads this product exists to
# collect. It does bound the leak to something that drains.
read_body 30m
idle 5m
}
}
}
{$DOMAIN} { {$DOMAIN} {
# Compress everything EXCEPT the SSE stream — gzip buffering delays # Compress everything EXCEPT the SSE stream — gzip buffering delays
# "real-time" likes/comments until the ~30s keep-alive tick. # "real-time" likes/comments until the ~30s keep-alive tick.
@@ -9,34 +39,125 @@
header { header {
Strict-Transport-Security "max-age=31536000; includeSubDomains" Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Content-Type-Options "nosniff" X-Content-Type-Options "nosniff"
X-Frame-Options "DENY"
Referrer-Policy "strict-origin-when-cross-origin" Referrer-Policy "strict-origin-when-cross-origin"
} }
# X-Frame-Options: DENY everywhere EXCEPT the keepsake download endpoints, which
# are navigated in a HIDDEN, SAME-ORIGIN iframe so a 404/429 can't unload the PWA
# (see frontend/src/routes/export/+page.svelte). WebKit enforces XFO *before*
# honouring Content-Disposition, so a blanket DENY makes the download silently do
# nothing on iOS Safari — the app's primary platform. SAMEORIGIN still blocks
# cross-origin framing.
#
# Split into two disjoint matchers rather than an override: Caddy applies the
# FIRST header directive outermost, so it wins on write — a later, more specific
# `header` would be silently ignored.
@framable path /api/v1/export/zip /api/v1/export/html
@not_framable not path /api/v1/export/zip /api/v1/export/html
header @framable X-Frame-Options "SAMEORIGIN"
header @not_framable X-Frame-Options "DENY"
# SvelteKit frontend — static assets with long-lived cache (content-hashed filenames) # SvelteKit frontend — static assets with long-lived cache (content-hashed filenames)
@hashed_assets path_regexp hashed /_app/immutable/.*\.[a-f0-9]{8,}\.(js|css|woff2)$ @hashed_assets path_regexp hashed /_app/immutable/.*\.[a-f0-9]{8,}\.(js|css|woff2)$
header @hashed_assets Cache-Control "public, max-age=31536000, immutable" header @hashed_assets Cache-Control "public, max-age=31536000, immutable"
# Preview/thumbnail images. These are now served by the app through a # Preview/thumbnail/display images. These are served by the app through a
# visibility-checked alias (/api/v1/upload/{id}/{preview,thumbnail}) so moderation # visibility-checked alias (/api/v1/upload/{id}/{preview,thumbnail,display}) so
# can revoke access; direct /media/previews|thumbnails is 404-blocked at the app. # moderation can revoke access; the app serves no /media route at all, so there is no
# Privately cacheable for a short window (the app sets the same header; this is the # direct path to the bytes. Privately cacheable for a short window (the app sets the
# edge carve-out from the blanket no-store below). Kept short so a moderated image # same header; this is the edge carve-out from the blanket no-store below). Kept short
# stops being served to a direct-URL holder promptly. # so a moderated image stops being served to a direct-URL holder promptly.
@media_api path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail #
# `display` was missing here while the backend set `private, max-age=300` on it, and
# because `header` REPLACES, the blanket no-store below silently won. That route is the
# ~2048px derivative the diashow uses exclusively, so a projector left running all
# evening re-fetched a full-size JPEG for every slide — roughly 2-4 GB pulled through
# the app over 8 hours, on the same venue uplink 100 guests are uploading over, and a
# blank frame on every network hiccup.
@media_api path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail /api/v1/upload/*/display
header @media_api Cache-Control "private, max-age=300" header @media_api Cache-Control "private, max-age=300"
# API — never cache, EXCEPT the gated image routes above. # API and health — never cache, EXCEPT the gated image routes above. A cached health
# response would report the last known state rather than the current one.
@api { @api {
path /api/* path /api/* /health
not path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail not path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail /api/v1/upload/*/display
} }
header @api Cache-Control "no-store" header @api Cache-Control "no-store"
# Route API and media requests to the Rust backend # Route API and media requests to the Rust backend.
#
# The app serves no /media route at all (see the note in backend/src/main.rs) — media
# bytes are reachable only through the visibility-checked /api/v1/upload aliases, so
# /media/* forwards to a plain 404. The proxy line is kept deliberately: it means the
# edge faithfully hands /media to the app, so if a future change ever re-introduces a
# static media route the e2e gating specs see it here exactly as production would,
# instead of being masked by the SvelteKit 404 page.
reverse_proxy /api/* app:3000 reverse_proxy /api/* app:3000
reverse_proxy /media/* app:3000 reverse_proxy /media/* app:3000
# The backend registers /health on its ROOT router, not under /api/v1, so it needs its
# own line — without it the catch-all below hands /health to SvelteKit, which has no
# such route and returns its 404 page. That made the documented post-deploy check
# (`curl -fsS https://DOMAIN/health`) fail 100% of the time on a perfectly healthy
# stack. e2e/Caddyfile.test has always carried this line; production never did.
reverse_proxy /health app:3000
# Everything else goes to SvelteKit frontend # Everything else goes to SvelteKit frontend
reverse_proxy frontend:3001 reverse_proxy frontend:3001
# Last-resort page for when Caddy itself cannot reach an upstream — the app or frontend
# container down, restarting, or still warming up after a host reboot. Without it a guest
# gets Caddy's bodiless 502: a completely blank page, which reads as "the whole thing is
# gone" rather than "try again in a moment".
#
# THIS DOES NOT TOUCH APPLICATION ERRORS. `handle_errors` fires only on errors CADDY
# generates; a status the app returns through `reverse_proxy` is written back verbatim and
# never reaches here. That distinction is load-bearing rather than incidental: the keepsake
# download navigates a HIDDEN IFRAME and depends on a real 404/429 arriving from the app
# (frontend/src/routes/export/+page.svelte), and every API route answers 403/404/429 as
# ordinary JSON that the client parses. Swallowing those into an HTML page would be a far
# worse regression than the blank 502 this fixes. Verified against this exact config: an
# upstream 404 through `reverse_proxy` still arrives as `Content-Type: application/json`
# with its body intact, while only a dial failure renders the page below.
#
# Scoped to 5xx so a hypothetical future Caddy-generated 4xx (there is none today) still
# returns plainly instead of claiming the server is restarting.
#
# The body is inline because the caddy service mounts ONLY ./Caddyfile and caddy_data —
# there is no volume to ship an HTML file through and the image has no build step, so a
# static file would mean changing the deployed stack's compose definition. No external
# font, stylesheet or image is referenced: the app may be exactly what is down.
#
# `handle_errors` has NO position in the directive order — Caddy hoists it into a separate
# `errors` route list — so it cannot disturb the "first `header` directive wins" hazard
# documented at the top of this file. The site-wide security headers still apply to it.
handle_errors 5xx {
header Content-Type "text/html; charset=utf-8"
header Cache-Control "no-store"
# {err.status_code} preserves the real status. Hardcoding 503 would mislabel a genuine
# 502 for anything watching from outside.
respond `<!doctype html>
<html lang="de">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Gleich zurück</title>
<style>
html{background:#faf9f7;color:#1a1918;font-family:system-ui,-apple-system,"Segoe UI",Roboto,sans-serif}
body{margin:0;min-height:100vh;display:flex;align-items:center;justify-content:center;padding:2rem;text-align:center}
h1{font-family:Georgia,"Times New Roman",serif;font-weight:600;font-size:1.5rem;margin:0 0 .75rem}
p{margin:0;color:#545350;line-height:1.5}
@media (prefers-color-scheme:dark){html{background:#100f0f;color:#f5f4f2}p{color:#a6a4a1}}
</style>
</head>
<body>
<main>
<h1>Wir sind gleich zurück</h1>
<p>Die Seite wird gerade neu gestartet.<br>Bitte lade in einem Moment neu deine Fotos bleiben gespeichert.</p>
</main>
</body>
</html>
` {err.status_code}
}
} }

974
DEPLOYMENT_RUNBOOK.md Normal file
View File

@@ -0,0 +1,974 @@
# EventSnap — Production Deployment Runbook
Target: **Hetzner CX22** (2 vCPU, 4 GB RAM, 40 GB disk), single host, docker compose.
Event profile: ~100 guests, ~100 photos + a few videos, one evening, **operator attending and unavailable to troubleshoot**.
Everything here is written for that last constraint. Where a choice trades throughput for
"cannot need attention on the night", it takes the stable option.
> **Note on this repo's own docs.** `README.md:62` and `PROJECT.md:359` specify a **CX33
> (4 vCPU / 8 GB / 80 GB)**. You are deploying to half of that on every axis. Two pieces of
> tuning advice in `.env.example` are calibrated for the CX33 and are actively wrong for a
> CX22 — they are called out in §3. Trust this file over `.env.example` for sizing.
---
## 0. Timeline — the single most important control
### Step zero: verify the deployment files are committed, before you build anything
Everything the server clones must be in git — §7 tells you to `git clone` onto the box, so
anything living only in your working tree is not part of the deployment. Two failure modes if it
is not:
- If the committed `docker-compose.yml` still carried `build:` keys and no `image:` keys, then on
that clone `docker compose pull` would skip both services and `docker compose up -d` would start
**a fat-LTO release build of 427 crates on the CX22** — the exact scenario §1 rules out as an
expected OOM.
- `sqlx::migrate!()` embeds `./migrations` **at compile time**. An image built from a working tree
with uncommitted migrations bakes them in and applies them on first boot; any later rebuild from
a clean clone produces an image that lacks them and crash-loops with `VersionMissing` against its
own database.
**As of this writing all of these are committed and the check below passes.** Run it anyway — it
costs a second and it is the difference between finding this now and finding it at T5.
```bash
# Every deployment file must be tracked. Prints nothing and exits 0 when correct;
# names the offender and exits non-zero otherwise.
git ls-files --error-unmatch \
docker-compose.yml docker-compose.dev.yml docker-compose.build.yml .env.example \
backend/.dockerignore frontend/.dockerignore frontend/Dockerfile \
DEPLOYMENT_RUNBOOK.md Caddyfile >/dev/null
# No uncommitted edits to them.
git status --porcelain -- docker-compose.yml .env.example Caddyfile DEPLOYMENT_RUNBOOK.md
# The COMMITTED compose must pull, not build: 4 `image:` lines, zero `build:` lines.
git show HEAD:docker-compose.yml | grep -cE '^[[:space:]]*image:' # must be 4
git show HEAD:docker-compose.yml | grep -cE '^[[:space:]]*build:' # must be 0
# Every migration in the tree is committed — a build from a dirty tree bakes in extras.
git status --porcelain -- backend/migrations/ # must print nothing
# The Caddyfile PARSES. Nothing else checks it: the e2e stack mounts `e2e/Caddyfile.test`,
# so the production file is never executed until the real deploy — and a syntax error there
# is total. Caddy exits, `restart: unless-stopped` loops, 443 is dead for the whole event,
# and `docker compose up -d --force-recreate caddy` still exits 0 while it crash-loops.
docker run --rm -v "$PWD/Caddyfile:/etc/caddy/Caddyfile:ro" -e DOMAIN=example.com \
caddy:2-alpine caddy validate --config /etc/caddy/Caddyfile # must end "Valid configuration"
```
| When | What |
|---|---|
| **T7 days** | Commit and push everything above. Registry + DNS pre-flight (§5). Build and push images (§6). |
| **T5 days** | First deploy to the server (§7). Verify admin login. Leave it running. |
| **T5 days** | ⚠ **Enable Hetzner automated snapshots** (§10.1) and **point an uptime monitor at `/health`** (§10.4). Two console checkboxes, ~10 minutes total. Without them a failure during the event is both total and unnoticed. |
| **T3 days** | **Freeze migrations.** No further code deploys unless something is broken. |
| **T2 days** | Pre-pull current *and* previous image tags (§9). Install the hourly DB dump and prove it runs (§10.2). Run the backup rehearsal (§10.3). |
| **Event day** | Change nothing. Configuration tweaks via the admin dashboard only (§4). |
**Why the freeze matters more than anything else here.** Migrations run automatically at boot
(`backend/src/db.rs`, `create_pool`) and `sqlx::migrate!()` is used *without* `set_ignore_missing`. An older
image booted against a newer schema does not degrade — it **crash-loops** with `VersionMissing`.
The repo documents this itself in `backend/migrations/014_export_epoch.up.sql:3-8`. Once the
migration set is frozen, rollback is a one-line `.env` edit; before it is frozen, rollback means
hand-running down-SQL under pressure.
---
## 1. Deployment strategy — build on the Mac, push, pull
**Decision: build `linux/amd64` images on the M3 Pro, push to `registry.mc02.dev`, pull on the server.**
Rejected alternatives, with the reason each loses:
| Option | Why not |
|---|---|
| **Build on the CX22** | `backend/Cargo.toml` sets `lto = true` + `codegen-units = 1` over **427 crates**. Fat LTO links the whole program in one single-threaded process; peak RSS is estimated at 2.54 GB against ~1.52.5 GB free with the stack running. OOM is the expected outcome, not a tail risk. And *rollback would also be a build* — 3560 min under pressure. |
| **CI (Gitea on Pi 5)** | Ruled out by you, and correct: a Pi 5 is strictly worse than the CX22 for a fat-LTO link. |
| **True cross-compilation** (`--target x86_64-unknown-linux-musl`) | The dependency graph has **four C-compiling crates**`ring` (per-arch assembly), `zstd-sys`, `libdeflate-sys`, `libsqlite3-sys` — so it needs a real musl cross-toolchain. Alpine has no such package for an arm64 host; the working answer is `cargo-zigbuild`, which means a new base image and a new toolchain days before an unrepeatable event. **Slow but certain beats fast but novel.** |
Emulated build is the right call because the cost lands where it is free: on your laptop, days
early, with nothing depending on it. Estimated wall-clock for the backend is **2560 min with
Rosetta enabled** (hours without it) — start it and walk away.
> **Escape hatch if emulation is unbearable:** spin up a temporary Hetzner CPX41 (8 vCPU/16 GB) in
> the same account, build natively, push, destroy it. Costs cents, same commands, no repo change.
---
## 2. Repo changes — ALREADY APPLIED
> **Status: these are done, in the working tree.** They are documented here so you know what
> changed and why, not as work to repeat. Run `git diff` to review before committing.
### 2.1 `.dockerignore` files (were absent)
`backend/.dockerignore`:
```
target/
```
`frontend/.dockerignore`:
```
node_modules/
.svelte-kit/
build/
```
Without these, the first time you run `cargo build` or `npm install` locally, every image build
ships a multi-GB context to an *emulated* builder. Worse: `frontend/Dockerfile:8` does `COPY . .`
**after** `npm ci`, so a macOS `node_modules/` would be merged over the container's Linux one.
### 2.2 Switch compose from `build:` to `image:`
In `docker-compose.yml`, replace the `build:` block on `app` and `frontend`:
```yaml
app:
image: registry.mc02.dev/eventsnap/app:${EVENTSNAP_VERSION:?set EVENTSNAP_VERSION in .env}
```
```yaml
frontend:
image: registry.mc02.dev/eventsnap/frontend:${EVENTSNAP_VERSION:?set EVENTSNAP_VERSION in .env}
```
Two deliberate properties:
- The `:?` form **fails loudly** on an unset variable instead of resolving to an empty tag.
- **Removing `build:` entirely is a safety feature.** With no `build:` key on the server, a wrong
tag is an instant `manifest unknown` — never a surprise 45-minute compile on the production box.
New `docker-compose.build.yml` (opt-in, Mac only — follows the convention stated in
the header of `docker-compose.dev.yml` that overlays are never auto-loaded):
```yaml
# Build overlay. NOT loaded automatically. Used only where images are BUILT — never on the
# production server, which pulls.
# docker compose -f docker-compose.yml -f docker-compose.build.yml build
services:
app:
build: { context: ./backend, dockerfile: Dockerfile }
frontend:
build: { context: ./frontend, dockerfile: Dockerfile }
```
`e2e/docker-compose.test.yml` keeps its own `build:` blocks, so the e2e gate is unaffected.
### 2.3 Add log rotation to every service
`docker-compose.yml` currently sets **no logging config**, so the default `json-file` driver keeps
logs forever, on the same filesystem as Postgres and the media. Add to each of the four services:
```yaml
logging:
driver: json-file
options: { max-size: "10m", max-file: "3" }
```
(Equivalently, set it once in `/etc/docker/daemon.json` — but that restarts the Docker daemon,
so do it *before* the stack is live, not after.)
---
## 3. `.env` — production values
```bash
# ── Identity ──────────────────────────────────────────────────────────────
DOMAIN=<your domain>
EVENT_NAME=<...>
EVENT_SLUG=<...>
# ── Image version (NEW — drives the image: tags in docker-compose.yml) ────
EVENTSNAP_VERSION=v0.13.0
# ── Secrets — ALL of them, before the first `up -d` ───────────────────────
JWT_SECRET=<openssl rand -hex 64>
# These two are NOT optional and have no defaults. docker-compose.yml interpolates them into
# `environment:`, which overrides `env_file`, so leaving them out does not fall back — it creates
# a Postgres role and database named "" while DATABASE_URL still says `eventsnap`. The result is
# a permanent crash loop whose only clean exit is `down -v`. Compose now refuses to start without
# them, but write them here anyway: the three values below must agree with each other.
POSTGRES_USER=eventsnap
POSTGRES_DB=eventsnap
POSTGRES_PASSWORD=<openssl rand -hex 24>
DATABASE_URL=postgres://eventsnap:<SAME PASSWORD>@db:5432/eventsnap
ADMIN_PASSWORD_HASH='<docker run --rm caddy:2-alpine caddy hash-password --plaintext "pw">'
# ── Paths — must match the volume mounts ──────────────────────────────────
# All four of these are PINNED in docker-compose.yml under `app.environment`, which overrides
# `env_file`. Keep them consistent here for readability, but understand that editing them in
# `.env` changes nothing — the pin is what the container gets. Change the pin.
MEDIA_PATH=/media
EXPORT_PATH=/exports
APP_PORT=3000
# ── Sizing (see the two corrections below) ────────────────────────────────
# 15, matching .env.example, the `db` sizing comment in docker-compose.yml and the code
# default. An earlier draft of this runbook said 30: that does not fit the 1G memory limit
# compose allots `db`, and 30 simultaneous queries cannot run on 2 vCPU anyway — they queue
# on the CPU instead of on the pool. Raise it only alongside more cores AND a bigger limit.
#
# ALSO PINNED IN COMPOSE (see above), and pinned for a sharper reason than the paths: an
# unparseable value here is boot-FATAL rather than falling back to the default, so a stray
# quote or a trailing inline comment in `.env` would crash-loop the app behind a live Caddy.
DATABASE_MAX_CONNECTIONS=15
COMPRESSION_WORKER_CONCURRENCY=2
# ── Comments off, likes + captions on ─────────────────────────────────────
COMMENTS_ENABLED=false
# ── Logging (NEW — production currently defaults to DEBUG) ────────────────
RUST_LOG=eventsnap_backend=info,tower_http=warn
```
### Two sizing decisions worth understanding before you touch them
`.env.example` now agrees with this section on both — it carries the same reasoning inline and
self-corrects the old advice. Kept here because these are the two knobs an operator is most
tempted to raise under pressure.
**`COMPRESSION_WORKER_CONCURRENCY`: keep `2`. Do NOT raise to 4, and do NOT raise the app memory
limit to 2G.** An earlier draft justified both on the premise that "each worker can run an
ffmpeg transcode". **There is no transcode anywhere in this codebase.** `services/video.rs::run_ffmpeg`
runs `ffmpeg -ss <t> -i <src> -vframes 1 -vf scale=…` — a single poster frame. Video originals are
stored and served byte-for-byte.
The real memory consumer is the **image** path: `image` 0.25's `resize` builds an `Rgba32F`
intermediate at **16 bytes/px**, sized `source_width × target_height`, which the 256 MiB decode
guard in `imaging::decode_limits` does not cover. Estimated peak per photo:
| Source | Peak (decode + display resize) |
|---|---|
| 12 MP (typical phone) | ~145 MB |
| 24 MP (iPhone Pro default) | ~223 MB |
| 48 MP ("Max" mode) | ~354 MB |
Those are per-photo peaks, and the "two 48 MP photos at once" pair this limit used to be sized
against **is no longer reachable**: `compression.rs` takes an EXCLUSIVE `heavy` permit for a large
decode, so two giants serialise no matter what `COMPRESSION_WORKER_CONCURRENCY` is set to (see
`.env.example`, which makes the same point). The binding case is now one giant (~354 MB) plus the
ordinary working set against the 1 GiB cap, which is comfortable.
What has not changed is the reason to keep concurrency at 2 and `app` at 1G: at concurrency 4 the
memory arithmetic stops working (app=2G + db=1G + 256M + 256M + ~370 MB OS/Docker ≈ 3954 MiB
against ~3910 MiB MemTotal — oversubscribed before a single photo arrives).
**`quota_tolerance`: keep `0.75`. Raising it does not make anything more generous for a real guest.**
See §4.
---
## 4. Admin dashboard configuration
These live in the DB `config` table, not `.env`. Changes take effect on the **next request**
`patch_config` invalidates the cache synchronously (`admin::patch_config`). No restart needed. Set them
after the first deploy, before the event.
| Key | Default | **Set to** | Why |
|---|---|---|---|
| `upload_rate_per_hour` | 100 | **1000** | A guest multi-selecting 100 photos hits exactly 100. Client backs off and resumes, but this makes it a non-issue. |
| `feed_rate_per_min` | 60 | **240** | Headroom for reconnect bursts. |
| `social_rate_per_min` | 120 | **600** | This is the **likes** limiter — the interaction you're keeping. |
| `export_rate_per_day` | 3 | **20** | **One shared bucket across both archives**`enforce_export_rate` keys on `export:{user_id}` regardless of which archive is being fetched. Downloading Gallery.zip + Memories.zip costs 2 of 3; one retry locks a guest out for 24 h. The limit is now charged when the download *ticket* is minted, so being over it produces a visible German error instead of a tap that silently does nothing. |
| `max_video_size_mb` | 500 | **leave at 500** | See below. |
| `quota_tolerance` | 0.75 | **leave at 0.75** | See below. |
| `quota_enabled`, `storage_quota_enabled`, `rate_limits_enabled` | true | **leave on** | This is the only disk-full safety net. |
> **Correction (was wrong in an earlier draft).** This section used to say "**Ignore
> `estimated_guest_count`** … read by no code at all". That is **false** — it is a live tuning
> knob and it is the dominant term in the quota divisor for a normal event. An operator who
> believed the old text and changed it would have moved every guest's ceiling. `upload_count_quota_enabled`
> genuinely is inert.
### Why quotas are already as generous as you want
```
divisor = max(active_uploaders, estimated_guest_count, 1)
per_user_limit = max(floor(free_disk × quota_tolerance / divisor), 500 MiB)
```
(`upload::quota_limit_bytes`. The 500 MiB floor applies only when the whole budget can back it —
below that the divided value stands, so the quota cannot promise space the disk does not have.)
`active_uploaders` is `SELECT COUNT(DISTINCT user_id) FROM upload WHERE deleted_at IS NULL`
**people who actually uploaded**, not guests who joined. But it is a `max`, not the sole divisor:
`estimated_guest_count` (default **100**) acts as a **floor on the divisor**, so the ceiling settles
at its final value early instead of sliding down all evening as guests arrive. It also blunts the
abuse case where the divisor was attacker-controlled — ~1000 throwaway accounts once drove every
real guest's ceiling to ~52 MB.
With ~28 GB free, `quota_tolerance` 0.75 and a realistic 30 people actually uploading, the divisor
is **100** (not 30, because `estimated_guest_count` floors it), giving 28 GB × 0.75 / 100 ≈ 210 MB —
which is below the floor, so **every guest is granted the 500 MiB minimum**. Against an expected
~1.25 GB for the *entire event*, nobody will be blocked. An earlier draft computed "~700 MB each"
by dividing by 30; that ignored the floor on the divisor and was wrong.
Raising `quota_tolerance` would only raise the **saturation ceiling** (media converges to
`t/(1+t)` of free space: 43% at 0.75, 50% at 1.0). It does nothing for a real guest at your volume,
and it eats the headroom the keepsake needs — the export preflight reserves `media × 1.10 × 2`
because `Gallery.zip` and `Memories.zip` each store media **uncompressed** (`export.rs` — both archives store media uncompressed).
**Why `max_video_size_mb` stays at 500.** An earlier draft of this runbook said to lower it to 250,
on the grounds that there is no client-side size check. That was wrong on both halves:
- A client guard **does** exist — `frontend/src/routes/upload/+page.svelte` rejects anything over
`HARD_MAX_UPLOAD_BYTES` (576 MiB) before a byte leaves the phone, alongside the HEIC reject.
- That guard is a **compile-time constant**, not the DB value. Lowering `max_video_size_mb` therefore
changes nothing on the client: the picker still accepts the clip, the phone still starts sending
it, and the *server* aborts it partway (`stream_field_to_file` does stop at the cap, so the AP is
not saturated for the full file — but the guest still gets a mid-upload failure where 500 MB would
simply have worked).
So lowering it is a pure capacity decision with no UX upside, and at your volume there is no capacity
problem to solve. Leave it. If you ever do want a smaller ceiling to bite on the phone, the constant
above has to move with it.
---
## 5. Pre-flight — T7 days (do NOT leave this to event week)
### Registry
From **both** the Mac and the server, as **the user who will deploy** (root's `docker login` does
not help a non-root deploy — credentials go to `~/.docker/config.json`):
```bash
getent hosts registry.mc02.dev # DNS resolves from the server
curl -fsSI https://registry.mc02.dev/v2/ # TLS must be PUBLICLY trusted; Docker rejects
# self-signed without extra daemon config
docker login registry.mc02.dev
```
Also confirm: the `eventsnap` repository/namespace exists if your registry requires pre-creation
(Harbor does; plain `registry:2` does not), and the registry host has ~500 MB free per release.
### DNS and TLS — get this wrong and you are locked out for an hour
The DNS **A record must point at the CX22 before the first `docker compose up -d`**, because Caddy
attempts certificate issuance on boot. Let's Encrypt allows only **5 failed validations per hostname
per hour** (refilling 1 per 12 min). A misconfigured DNS record plus a few impatient restarts will
rate-limit you out of getting a certificate at all.
- Deploy days early so issuance happens with no time pressure.
- **Never delete the `caddy_data` volume** — it holds the certificate and the ACME account key.
- Never run `docker compose down -v`. It destroys `postgres_data`, `media_data`, `exports_data`
*and* `caddy_data`.
### Host preparation
```bash
docker compose version # must be v2.x — see below, this one is not optional
free -h && swapon --show # Hetzner images ship no swap
df -h /var/lib/docker # want ≥ 25 GB free
```
> **If `docker compose version` reports v1 (or `docker-compose` is a separate Python binary), STOP
> and install the v2 plugin before deploying.** This check previously had no failure action, which
> made it decorative — and it is the single check that the whole sizing argument rests on.
>
> On Compose v1, `deploy.resources.limits` is **silently ignored** outside Swarm: no warning, no
> error, `up -d` exits 0. Every memory and CPU limit in `docker-compose.yml` evaporates, and §1's
> arithmetic (`app` 1G + `db` 1G + 256M + 256M inside ~3910 MiB) becomes fiction — the first 48 MP
> photo takes the box out via the OOM killer instead of being bounded. On v2 the limits are real
> (verified empirically: `memory: 1G` produces `HostConfig.Memory=1073741824`).
>
> ```bash
> # Debian/Ubuntu, with Docker's official repo already configured:
> apt-get update && apt-get install -y docker-compose-plugin
> docker compose version # must now print v2.x
> ```
>
> Verify the limits actually landed, once the stack is up — this is the check that matters, not the
> version string:
>
> ```bash
> docker inspect eventsnap-app-1 --format '{{.HostConfig.Memory}} {{.HostConfig.NanoCpus}}'
> # Must print two NON-ZERO numbers. `0 0` means the limits were dropped.
> ```
**Add 2 GB of swap** as an OOM cushion — a compression spike that would otherwise kill the container
instead swaps out cold pages and merely runs slowly:
```bash
fallocate -l 2G /swapfile && chmod 600 /swapfile && mkswap /swapfile && swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab
sysctl -w vm.swappiness=10 && echo 'vm.swappiness=10' > /etc/sysctl.d/99-swap.conf
```
> **Already handled — do not hand-edit compose.** Compose sets each container's `Memory` limit but
> leaves `MemorySwap` unset, and Docker then allows swap equal to the memory limit, so adding host
> swap would silently **double** every container ceiling (to ~5 GiB of ceilings on a 3.82 GiB box).
> `docker-compose.yml` now ships `memswap_limit` on all four services — 1152m on `app` and `db`,
> 320m on `frontend` and `caddy` — so this step is safe as written.
>
> This used to say "add it yourself", which also broke §0's own gate that
> `git status --porcelain -- docker-compose.yml` must print nothing. Confirm it is still there:
>
> ```bash
> docker inspect eventsnap-app-1 --format '{{.HostConfig.Memory}} {{.HostConfig.MemorySwap}}'
> # 1073741824 1207959552 — the second number MUST be larger than the first but not double it.
> ```
---
## 6. Build and push — on the Mac
Enable **Docker Desktop → Settings → General → "Use Rosetta for x86_64/amd64 emulation"**, and set
**Resources → Memory ≥ 8 GB** (the fat-LTO step will OOM inside the VM otherwise, even with 18 GB on
the host).
```bash
docker run --rm --platform linux/amd64 alpine uname -m # must print x86_64
cd /Users/fabianhammprivat/Projects/EventSnap
VERSION=v0.13.0 # latest existing tag is v0.12.0 — see §9 before reusing it
ROLLBACK=v0.13.0-a # the SAME source, tagged twice; §9 explains why
SHA=$(git rev-parse --short HEAD)
docker buildx create --name eventsnap --use 2>/dev/null || docker buildx use eventsnap
docker buildx build --platform linux/amd64 \
-t registry.mc02.dev/eventsnap/app:$VERSION \
-t registry.mc02.dev/eventsnap/app:$ROLLBACK \
-t registry.mc02.dev/eventsnap/app:$SHA \
--push ./backend
docker buildx build --platform linux/amd64 \
-t registry.mc02.dev/eventsnap/frontend:$VERSION \
-t registry.mc02.dev/eventsnap/frontend:$ROLLBACK \
-t registry.mc02.dev/eventsnap/frontend:$SHA \
--push ./frontend
```
### Verify the architecture — never skip this
An arm64 image pulls fine and then dies with `exec format error`. Catch it here, not on the server:
```bash
docker buildx imagetools inspect registry.mc02.dev/eventsnap/app:$VERSION
docker buildx imagetools inspect registry.mc02.dev/eventsnap/frontend:$VERSION
# Both MUST report Platform: linux/amd64
```
Then prove the binary actually executes. `AppConfig::from_env()` runs before any DB connection
(`main.rs` builds `AppConfig` before touching the pool), so this needs no database:
```bash
docker run --rm --platform linux/amd64 \
-e APP_ENV=production -e JWT_SECRET=x -e DATABASE_URL=x -e EVENT_SLUG=x \
registry.mc02.dev/eventsnap/app:$VERSION
```
Expect the "Refusing to start in production … placeholder" message. **That message means the amd64
binary ran.** `exec format error` means the architecture is wrong.
**Tag policy: never `latest` in production.** With `latest` you cannot tell what is running,
rollback becomes a registry re-push (impossible if the registry is down — exactly when you need it),
and `restart: unless-stopped` after a reboot is ambiguous. Two immutable tags per build: semver and
short SHA. Deploy by semver.
---
## 7. First deploy — on the server
```bash
# Directory name matters: volumes are prefixed with it, and README's backup commands
# hardcode the `eventsnap_` prefix.
git clone <repo> eventsnap && cd eventsnap
cp .env.example .env && nano .env # every value from §3
docker login registry.mc02.dev
docker compose pull # must fully succeed before anything starts
```
### 7.1 The secret pre-flight — run this before the first `up -d`, always
```bash
docker compose run --rm --no-deps app
```
`--no-deps` means `db` never starts, so **no Postgres data directory is initialised**.
`AppConfig::from_env()` runs first and reports *every* placeholder at once (`config::validate_secrets`).
When the only remaining complaint is a database *connection* failure, the secrets are good.
**Why this step exists.** `POSTGRES_PASSWORD` is applied **only at initdb**. `docker-compose.yml`
starts `db` in the same command as `app`, so a single `up -d` with a placeholder bakes the wrong
password in permanently — the app then loops on `password authentication failed`, and the only exits
are `ALTER ROLE` or `down -v`, which deletes the database, the media and the exports. The repo
describes this trap at `backend/src/db.rs` (`explain_auth_failure`); this command is what avoids it.
### 7.2 Bring it up
```bash
# `.env` is consumed by docker compose, not by your shell — read $DOMAIN out of it first.
# Reads that ONE variable rather than sourcing the file: `.env` legitimately holds values with
# apostrophes (EVENT_NAME), and `. ./.env` aborts on one with "Unterminated quoted string".
DOMAIN=$(sed -n 's/^DOMAIN=//p' .env | tr -d "\"'")
docker compose up -d
docker compose logs -f app # wait for "database connected and migrations applied"
curl -fsS https://$DOMAIN/health # → ok (503 means the app is up but the DB is not)
```
### 7.3 Verify what the container actually received
`MEDIA_PATH`, `EXPORT_PATH` and `APP_PORT` are all **pinned** on the `app` service in
`docker-compose.yml`, exactly as §3 says — editing them in `.env` changes nothing. This step is
not about whether they are pinned; it is about confirming the container got the values you think
it did, including the two that genuinely do come from `.env`:
```bash
docker compose exec app printenv DATABASE_URL EXPORT_PATH ADMIN_PASSWORD_HASH
```
1. **`DATABASE_URL`** (from `.env`) must contain `@db:5432`. A dev `.env` points it at
`@localhost`, which inside the container is the app itself.
2. **`EXPORT_PATH`** (pinned) must read `/exports`. If it does not, the pin has been edited —
anywhere else and the keepsake archives are written to the container's writable layer and
**vanish on the next `up -d`**, including on a rollback.
3. **`ADMIN_PASSWORD_HASH`** (from `.env`) must match `.env` **byte for byte.**
**Then actually log in to `/admin` with the real password.** This is not optional politeness:
- The production secret guard only rejects *placeholders* (`config::looks_placeholder`). A hash **corrupted**
by shell or Compose escaping is not a placeholder — the app boots green, `/health` says `ok`, and
every admin login 401s.
- The Admin user row is created **by a successful admin login** (`auth::handlers::admin_login`), and
promoting anyone to Host requires an Admin or Host (`auth::middleware`'s role guard). **No admin login
⇒ no host, ever** ⇒ you cannot close the event, release the gallery, ban anyone, reset a PIN, or
change any limit — for the whole event.
Per the Compose spec, single-quoted `env_file` values *are* used literally and the quotes are
stripped, so the single-quote form in `.env.example` is correct. The comment in
`docker-compose.dev.yml` claiming production has the same bug is **stale**. Verify anyway —
the cost of checking is 10 seconds; the cost of being wrong is the whole event.
**Promote a second person to Host** once you are in, so a single lost session is not fatal.
---
## 8. Post-deploy verification
```bash
DOMAIN=$(sed -n 's/^DOMAIN=//p' .env | tr -d "\"'") # $DOMAIN comes from .env, not your shell
docker compose ps # db, app, frontend healthy; caddy has
# no healthcheck and shows only "running"
docker inspect -f '{{.HostConfig.Memory}}' eventsnap-app-1 # must be 1073741824, not 0
curl -fsS https://$DOMAIN/health # ok — now a real DB check, not a constant
docker compose exec app printenv COMMENTS_ENABLED RUST_LOG
docker builder prune -af && docker image prune -f # reclaim build cache
df -h /var/lib/docker
```
Then, from a phone on cellular (not the office wifi):
- [ ] Join as a guest with a PIN
- [ ] Upload a photo → appears in the feed within a few seconds
- [ ] Upload a video → plays back (iOS Safari range requests)
- [ ] Add a caption, add a like — **no comment UI anywhere**
- [ ] Admin login works; host dashboard reachable
- [ ] Release the gallery on a test event and download both archives
---
## 9. Rollback
### There is no older image you can roll back to. Build the rollback target yourself.
Read this before the event, not during it. The obvious move — drop `EVENTSNAP_VERSION` back to the
previous released tag — **takes the app down permanently** and looks like a crash loop with no
explanation:
```
$ git ls-tree --name-only v0.12.0 backend/migrations/ | wc -l
12 # 6 migrations. HEAD has 31.
$ git rev-list --count v0.12.0..HEAD
217
```
`db.rs` runs `sqlx::migrate!()` with no `set_ignore_missing`, so an image built from a 6-migration
tree, booting against a database that already carries versions 007031, returns `VersionMissing`.
`create_pool` errors, `main` exits 1, and `restart: unless-stopped` restarts it forever — with Caddy
still routing traffic to it. (`014_export_epoch.up.sql` documents this failure mode; §0 restates it.)
No `v0.12.0` image was ever built or pushed either, so the pre-pull would fail with
`manifest unknown` before you ever got that far.
**So: at build time, tag the SAME frozen commit twice.** Two identical images, two names. The
rollback then swaps to a binary that is bit-for-bit what you tested and carries the identical
migration set, which makes it a genuine no-op rather than a gamble:
```bash
# In §6, push both tags from the one build:
VERSION=v0.13.0
ROLLBACK=v0.13.0-a # same source, different name — the rollback target
docker buildx build --platform linux/amd64 \
-t registry.mc02.dev/eventsnap/app:$VERSION \
-t registry.mc02.dev/eventsnap/app:$ROLLBACK \
--push ./backend
# ...and the same two tags for ./frontend
```
**Rolling back — ~30 seconds, no network:**
```bash
sed -i 's/^EVENTSNAP_VERSION=.*/EVENTSNAP_VERSION=v0.13.0-a/' .env
docker compose up -d app frontend
```
This only works offline if both are already resident. **Pre-pull all four at T2:**
```bash
docker pull registry.mc02.dev/eventsnap/app:v0.13.0
docker pull registry.mc02.dev/eventsnap/frontend:v0.13.0
docker pull registry.mc02.dev/eventsnap/app:v0.13.0-a
docker pull registry.mc02.dev/eventsnap/frontend:v0.13.0-a
docker image ls | grep eventsnap # confirm all four
```
Once resident, `up -d`, reboots, restarts and rollbacks need **zero** registry contact. That single
step makes a registry outage on event day irrelevant.
Be clear-eyed about what this buys you: an identical image cannot undo a bad *release*, only an
image that got corrupted or a container that wedged — and `docker compose restart app` already
covers both. It exists so that the rollback line in the emergency card is safe to run rather than
catastrophic. **If you genuinely need to undo a code change during the event, you cannot; freeze
early enough that you never have to.**
**Across a migration boundary — avoid by freezing.** If you must: all 31 migrations have paired
`.down.sql` files, but **none of them removes its own `_sqlx_migrations` row**, so that second step
is mandatory and undocumented:
```bash
docker compose stop app
# `sh -c` so the CONTAINER expands the credentials. They live in the container's environment
# and in .env — not in your shell — so an unwrapped `-U "$POSTGRES_USER"` sends `-U ""` and
# psql answers `FATAL: role "" does not exist`.
docker compose exec -T db sh -c \
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -v ON_ERROR_STOP=1' \
< backend/migrations/0NN_x.down.sql
docker compose exec -T db sh -c \
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "DELETE FROM _sqlx_migrations WHERE version = NN;"'
```
> **A down migration is only valid PAIRED WITH A CODE ROLLBACK — it is not a standalone repair.**
> `Upload::create` sends an `ON CONFLICT ... WHERE` predicate that must match the live partial
> index exactly, and these queries are not compile-checked. Run **026**'s or **031**'s down against
> the current binary and every upload carrying a `client_upload_id` — i.e. every upload from the
> shipped client — becomes a runtime 500. Roll the image back first, then the migration.
>
> **026's down can also fail outright, and that is expected.** It restores a wider unique index, so
> it aborts with `could not create unique index ... is duplicated` on any database where a guest
> ever deleted a photo and re-uploaded it. The transaction rolls back cleanly and the narrow index
> survives intact — no half-state — but you cannot go below 026 on a database that has seen real
> use. Verified against a live Postgres.
Migration **014** is the only destructive one on the way up, and the only one with a rehearsal
harness — `backend/scripts/rehearse-014.sh`. Run it once against a real dump before the event.
**Registry-down transport fallback** (layers are already compressed — do not add gzip):
```bash
docker save registry.mc02.dev/eventsnap/app:v0.13.0 \
registry.mc02.dev/eventsnap/frontend:v0.13.0 | ssh root@SERVER 'docker load'
```
---
## 10. Backup
Full commands are in README's **`## Backup`** and **`## Restore`** sections — read `## Restore` to
its END (through the media *and* exports restore, and the `chown` that follows), not just the
database step. Referenced by heading, not by line number: the previous pointer named a line range
that had drifted to the middle of an unrelated section and stopped mid-way through restore step 2,
which would have restored the database and no media. They are correct — `pg_dump --clean --if-exists`, plus
`alpine tar` out of `eventsnap_media_data` and `eventsnap_exports_data`, mounted at `/src` (not
`/media`), with `chown -R 100:101` on restore because the app runs non-root and BusyBox tar has no
`--same-owner`. There is deliberately no script.
Three gaps the README does not cover:
1. **Nothing backs up `.env`**, which holds the only copy of `POSTGRES_PASSWORD`. A dump you cannot
authenticate against is not a backup. Copy `.env` off the box, encrypted, once it is final.
2. **The final dump is not a backup — it is an archive.** Taking it *after* locking uploads gives
you a consistent pair, and that is the right way to archive the finished event. But it means
that until the host locks uploads there is **no copy of anything anywhere**. All four volumes
sit on the same 40 GB filesystem, on one VPS, with no redundancy. A disk or host failure at
23:00 — the fullest the gallery will ever be — loses **100% of the event**, permanently, with
the guests still in the room. Both of the mitigations below are required.
3. **Nothing is watching.** See §10.2.
### 10.1 Snapshots — do this once, before the event
> ⚠ **ACTION REQUIRED — Hetzner Cloud console, ~5 minutes, one checkbox.**
> Server → **Backups** → enable. Costs ~20% of the server price and needs no operator action
> ever again.
This is the single highest-value item in this runbook. It converts "total, permanent loss" into
"lose at most the hours since the last snapshot", automatically, with nobody awake. It covers the
whole volume set at once — database, media, exports and `.env` — which the `pg_dump` path does not.
It does **not** replace §10's archive: snapshots are whole-disk and crash-consistent, so restoring
one gives you the box back, not a portable copy of the photos. Do both.
### 10.2 A mid-event database dump — cheap, and the only thing cron should do
The database is small (a few MB — it holds rows, not pixels) and it is the part that cannot be
reconstructed: media files on disk without their `upload` rows are anonymous UUIDs with no
uploader, caption, hashtag or timestamp. Dumping it hourly costs essentially nothing and is safe
while uploads are live, because a `pg_dump` is transactionally consistent on its own.
Media is the bulk and *is* recoverable from guests' phones in the worst case, so it stays on the
event-night schedule below.
Install this as **the same user you deployed as** (§5) — not root. `docker compose` needs that
user's docker group membership and its compose project, and a root crontab has neither.
```bash
# On the server, before the event. Hourly DB-only dump, keeping the last 48.
mkdir -p ~/eventsnap-dumps
cat >~/eventsnap-dump.sh <<'SH'
#!/bin/sh
set -eu
# Must match your deploy directory from §5. cron starts in $HOME, so this cannot be relative.
cd "$HOME/eventsnap"
# NOTE: deliberately does NOT source .env. Nothing here reads it — POSTGRES_USER and POSTGRES_DB
# are expanded INSIDE the db container by the single-quoted sh -c below, using the values compose
# already injected. Sourcing it was actively harmful: `.env` legitimately contains values with
# apostrophes (EVENT_NAME="Max & Maria's Wedding"), and POSIX sh aborts on one with
# "Unterminated quoted string". Under `set -eu` this script would exit before pg_dump — every
# hour, silently, leaving the only automated backup of the irreplaceable table permanently empty.
OUT="$HOME/eventsnap-dumps/db-$(date -u +%Y%m%dT%H%M%SZ).sql.gz"
docker compose exec -T db sh -c \
'pg_dump --clean --if-exists -U "$POSTGRES_USER" "$POSTGRES_DB"' | gzip >"$OUT.tmp"
mv "$OUT.tmp" "$OUT" # atomic: never leave a truncated dump looking complete
ls -1t "$HOME"/eventsnap-dumps/db-*.sql.gz | tail -n +49 | xargs -r rm
SH
chmod +x ~/eventsnap-dump.sh
( crontab -l 2>/dev/null; echo "17 * * * * $HOME/eventsnap-dump.sh >>$HOME/eventsnap-dumps/dump.log 2>&1" ) | crontab -
# Prove it works NOW, not at 23:00 — and prove it produced a NON-EMPTY dump, since the failure
# this replaces produced a zero-byte file and a clean exit code.
~/eventsnap-dump.sh && ls -lh ~/eventsnap-dumps/
gzip -t ~/eventsnap-dumps/db-*.sql.gz && echo "dump is a valid gzip"
zcat ~/eventsnap-dumps/db-*.sql.gz | grep -c 'CREATE TABLE' # must be > 0, not just "a file exists"
```
These land on the same filesystem, so they do **not** survive a disk loss — that is what §10.1 is
for. They protect against the far more likely failure: a bad migration, an accidental host action,
or a corrupted table.
### 10.3 The event-night archive — unchanged
Take the dump and the media tarball back-to-back **the night of the event, after locking uploads
from the host dashboard**, so the pair is consistent. Copy both off the box before you sleep.
### 10.4 Monitoring — something has to be able to wake you
> ⚠ **ACTION REQUIRED — external uptime monitor, ~5 minutes.**
> Point any free monitor (UptimeRobot, Better Stack, Healthchecks.io — all have free tiers with
> SMS or push) at `https://$DOMAIN/health`, 15 minute interval, **alerting to a phone that will
> be on you during the event.**
There is otherwise **no** metrics collection, no alerting, no log shipping and no external check
anywhere in this deployment. Without this step, none of the following reaches a human: a crash
loop, a full disk, a dead database, an expired certificate, or the box being off. The host is at
a party and is not watching a dashboard.
`/health` is already built for exactly this and nothing currently consumes it:
| Response | Meaning | Action |
|---|---|---|
| `200 ok` | App **and** database are answering | — |
| `503 database timeout` / `database unavailable` | App is up, Postgres is not | §13 emergency card |
| Connection refused / TLS error | App container or Caddy is down | `docker compose ps`, then §13 |
| Timeout | Box is gone, or the disk is full enough to wedge it | §10.1 snapshot restore |
It runs a real `SELECT 1` against the pool with a 2 s timeout — a green check means the request
path guests use is genuinely working, not merely that a process is listening.
**The one signal this does not give you is disk.** The low-disk banner on `/host` requires the
host to open a dashboard during their own party and does not refresh without a manual reload, so
treat it as a pre-event check, not an alert. Before the event, confirm headroom with §11's
numbers; the export preflight and the upload quota are the automated backstops.
---
## 11. Disk — the numbers for 40 GB
Expected event (100 photos @ ~4 MB, 5 videos @ ~150 MB). Originals are **always kept** and never
lossily recompressed; derivatives are a ≤800 px preview and a ≤2048 px display JPEG.
```
media (originals + derivatives) ~1.25 GB
peak during keepsake build (+ Gallery + Memories) ~3.5 GB
OS + Docker images + Postgres + logs ~4.5 GB
────────────────────────────────────────────────────────────
total ~1113 GB of ~36 GB usable
```
**40 GB fits with roughly 3× headroom**, provided you build elsewhere (a server-side build adds
35 GB of cache that permanently shrinks the guest quota, because the quota is recomputed against
*live* free space on every upload).
The "ENOSPC" projection in README's **`### Sizing the disk`** discussion models guests
**saturating the quota** (~12 GB of media), not 100 photos. That scenario needs ~10× your expected
volume — and it degrades gracefully: the export preflight refuses up front rather than hitting
ENOSPC mid-write, and the host dashboard warns while free space is still **1.25× above the level at
which uploads stop** (`handlers::host::disk_is_low`). Note that is the only trigger: the separate
10 GB absolute floor this used to describe was removed as unreachable, because the derived
threshold is always higher.
---
## 12. Known issues you are shipping with
None of these has a fix in this runbook; they are listed so nothing is a surprise. Severity is
scored purely by "would this interrupt you during the event".
### Fixed in this pass
| Was | Fix |
|---|---|
| **Queued uploads never rehydrated**`loadQueue()` had one call site, so a guest whose PWA was evicted mid-upload and reopened onto `/feed` had pending photos in IndexedDB that nothing ever read. Silent photo loss. | `loadQueue()` now runs on every authenticated boot in `+layout.svelte`. |
| **A corrupted `ADMIN_PASSWORD_HASH` booted green** and only failed at admin login — unrecoverable mid-event. | `config.rs` now validates the bcrypt *shape*, not just placeholder-ness, and refuses to start. |
| **SSE reconnect thundering herd** — flat 500 ms jitter regardless of backoff. | Jitter now scales with the delay; the feed's in-place refresh is spread over 8002800 ms. |
| **No client-side size/HEIC pre-check** — a doomed 600 MB upload saturated the venue AP before being rejected. | Certain-rejects are refused before any bytes leave the phone. |
| **Upload-queue UI unreachable** — the badged FAB opened the picker, not the queue, so a failed upload had no retry button. | The upload sheet now shows a queue entry whenever the badge is non-zero. |
| **A mid-event 401 left a dead screen** with no nav and no URL bar. | `api.ts` now returns the guest to `/join` (skipping the auth routes themselves). |
| **Rate limiter could panic while holding its mutex** (`timestamps[0]` with `max == 0`), poisoning it process-wide. | Uses `first()` with a fail-open guard. |
| **Hashtag deadlock**`edit_upload` and `add_comment` upserted tags in client/text order. | Both now sort + dedup on the normalised key, matching the upload path. |
| **Unbounded Docker logs + debug-level app logging.** | `logging:` caps every service at 30 MB; `RUST_LOG` documented in `.env.example`. |
### Fixed in the readiness pass — behaviour you should know about
These change what you will observe on the night, so they are listed separately from the table above.
| Was | Now |
|---|---|
| **Any ffmpeg-level failure DESTROYED the video.** A poster-frame error — ffmpeg hanging on a truncated `.mov`, an ENOSPC on `thumbnails/`, a DB blip — propagated into the give-up path, which soft-deleted the upload. Reproduced live: on a host with no ffmpeg the clip was gone ~6 s after its `201 Created`. | No failure in the video branch can fail the upload. The clip stays in the feed and plays from its original; only the poster is missing. |
| **A lost response duplicated the photo.** Every retry minted a fresh upload id, so a phone that lost the reply and re-sent — manually, or automatically on reconnect — stored the same photo two or three times and paid quota for each. | The client sends a `client_upload_id`; migration 022 makes it unique. A retry returns `200` with the original row. Verified live: three sends → one row, quota charged once. |
| **`/health` returned a constant `"ok"`.** The container reported healthy while every request 500'd. | Runs `SELECT 1` with a 2 s timeout: `200 ok` / `503`. Verified live: 200 → stop Postgres → 503 → start Postgres → 200, **with no app restart** (the pool revalidates on acquire). Note that Compose does not restart on an unhealthy probe — this is a diagnostic, deliberately not wired to automatic recovery, because a restart would truncate in-flight uploads to "fix" an outage that clears on its own. |
| **Abandoned `.tmp` uploads were never reclaimed** (see the old issue #2 below — it is now fixed, not deferred). | Swept at boot and hourly, keyed on modification time so a live upload can never age into it. |
| **A full disk made itself worse.** ENOSPC was retried three times, then refunded the quota, soft-deleted the row and *kept* the bytes — freeing nothing and inviting an immediate re-upload into the same full disk. | ENOSPC is classified separately: no retry, no refund, no delete. The row stays live and the photo is served from its original until there is room to compress it. |
| **The keepsake download failed invisibly.** The rate limit was enforced inside the iframe navigation, so an over-limit guest tapped and *nothing happened*, forever. | Charged when the ticket is minted — a normal `fetch` — so it surfaces as a German message naming the daily window. Raise `export_rate_per_day` per §4 anyway. |
| **`/host` and `/admin` subscribed to SSE but never opened the connection**, so the keepsake progress bar froze after a release. | Both connect on mount and disconnect on destroy, as `/export` already did. |
| **`?limit=-5` on the feed returned a 500.** | Clamped at both ends. |
| **All recurring hygiene lived in one unsupervised task** — one panic silently stopped session pruning, media reclamation and the temp sweep for the rest of the event. | Supervised and re-spawned, with an error log. |
| **No pool acquire or statement timeout.** A DB blip parked every request for 30 s. | 5 s acquire, `statement_timeout=15s`, `lock_timeout=5s`. |
### Still open — accepted for this event
| # | Issue | Impact |
|---|---|---|
| 1 | **Guest delete/caption-edit after release invalidates both keepsake archives** (`upload.rs`), forcing a rebuild with a 20 s debounce. **No-op before release**, so it cannot bite during the event. | Post-event only. Left alone because blocking guest deletes has real privacy downsides — that is a product call, not a bug fix. |
| 2 | **Video poster frames do not regenerate after a restart** mid-compression (the backfill filters `mime_type LIKE 'image/%'`). | Cosmetic — the video still plays; only its poster is missing. |
| 3 | **SSE keep-alives are sent as SSE comments** (`:ping`), which the browser's EventSource parser discards without dispatching. A client therefore cannot implement a pure silence timer to detect a half-open socket. | Worked around client-side: the feed runs a jittered 60120 s `/feed/delta` backstop and reconnects when a poll returns content the stream never delivered. A cleaner fix is to emit keep-alives as a *named* event; that is a coordinated backend+frontend change, not worth making during a freeze. |
| 4 | **The lightbox stops at the end of the loaded page** — stepping past the last loaded photo does not fetch the next one. | The guest scrolls the feed (which does page) and re-opens. |
### Migration checksum mismatch — `VersionMissing` / "previously applied but has been modified"
`sqlx` compares **checksums**, so renaming or renumbering a migration file is indistinguishable
from editing one. If a box ever booted an image built from a branch that numbered migrations
differently, the next boot aborts with *"migration 21 was previously applied but has been
modified"*, `main` exits non-zero, and `restart: unless-stopped` makes it **permanent** — with
Caddy still routing traffic to the dead container.
Every main-line migration is byte-identical to what shipped, so a box that only ever ran tagged
releases is unaffected. **Verify rather than assume** — run this against the server before any
deploy.
Note the `sh -c` wrapping, for the same reason as §9: `POSTGRES_USER` and `POSTGRES_DB` live in
`.env`, which Compose reads and **your shell does not**. Unwrapped, `-U "$POSTGRES_USER"` sends
`-U ""` and psql answers `FATAL: role "" does not exist` — at 11pm, with the app crash-looping.
```bash
docker compose exec -T db sh -c \
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "SELECT version, description, success FROM _sqlx_migrations ORDER BY version;"'
```
If the app is already crash-looping on a renumbered migration, and **only** if you have confirmed
the SQL in the new file is equivalent to what was actually applied. Take the version numbers from
the crash message and the query above — do **not** copy the ones below, which are an example:
```bash
docker compose stop app
# Replace 21,22,23 with the versions the boot error actually named.
docker compose exec -T db sh -c \
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "DELETE FROM _sqlx_migrations WHERE version IN (21,22,23);"'
docker compose start app # re-applies exactly those, then continues
```
This re-runs those migrations. They must be idempotent (`IF NOT EXISTS` / `IF EXISTS`) or this
fails differently. Take a `pg_dump` first — see §10.
### Schema changes are **not** compile-time checked
All ~120 queries use the runtime `sqlx::query()` API. There are no `query!` macros and no `.sqlx`
cache, and the backend **compiles with no `DATABASE_URL` at all**. Consequences:
- A migration that renames or drops a column **compiles clean** and fails in production as a
runtime 500 on whichever request path touches it first.
- `cargo build` succeeding tells you nothing about schema/query agreement. Only `cargo test`
(which runs against a real Postgres) and manual exercise of the affected route do.
So: after any migration that touches an existing column, exercise the routes that read it before
you consider the deploy done. README and PROJECT previously claimed compile-time checking; they
have been corrected.
---
## 13. Event-day emergency card
Assume you have a phone and two minutes.
```bash
cd ~/eventsnap
# FIRST LINE, ALWAYS. `.env` is read by docker compose, NOT by your shell — without this,
# every `$DOMAIN` below expands to nothing and `curl https:///health` reads like an outage
# when the site is fine. Reads the one variable instead of sourcing the file, because `.env`
# legitimately contains an apostrophe (EVENT_NAME) and `. ./.env` dies on it — which at 11pm
# looks exactly like the outage you came here to diagnose.
DOMAIN=$(sed -n 's/^DOMAIN=//p' .env | tr -d "\"'")
# Is it alive? (200 = app AND database are answering; 503 = the app is up, the DB is not)
curl -fsS https://$DOMAIN/health
# What is broken?
docker compose ps
docker compose logs --tail=100 app
# Nuclear option that is SAFE (keeps all data):
docker compose restart app
# Roll back to the identically-built sibling image — see §9 for what this can and cannot fix.
# Do NOT substitute an older release tag here; it will crash-loop on the migration set.
sed -i 's/^EVENTSNAP_VERSION=.*/EVENTSNAP_VERSION=v0.13.0-a/' .env && docker compose up -d app frontend
# Disk check
df -h /var/lib/docker
```
### "Der Speicher des Events ist fast voll" — guests cannot upload
**`df -h` will look fine, and that is not a contradiction.** The upload gate refuses long before the
disk fills: it reserves room for the keepsake, which is roughly a second copy of every original, plus
a 10 GB floor. Uploads stop at **~8 GB of media** on a 40 GB box, when `df` still shows ~20 GB free.
Check the number that actually binds, not free space:
```bash
docker compose exec -T db sh -c \
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -tAc "SELECT pg_size_pretty(sum(original_size_bytes)) FROM upload WHERE deleted_at IS NULL;"'
```
Mid-event, in order of preference: delete the largest videos from the host dashboard (each frees its
own bytes immediately), or move `exports_data` to a separate volume. Raising `quota_tolerance` will
**not** help — on this box every guest is already on the 500 MiB floor, so that knob is not what is
refusing them (see §4 and `.env.example`).
**NEVER** run `docker compose down -v`. It deletes the database, all media, all exports and the TLS
certificate. There is no undo.
Most limits are changeable from the **admin dashboard without a restart** — reach for that before
touching the shell.

View File

@@ -344,7 +344,7 @@ COMPRESSION_WORKER_CONCURRENCY=2
| Styling | Tailwind CSS | Utility-first, mobile-first; zero runtime CSS overhead | | Styling | Tailwind CSS | Utility-first, mobile-first; zero runtime CSS overhead |
| Backend | Rust + Axum | Developer preference; memory safety, single-binary deploy | | Backend | Rust + Axum | Developer preference; memory safety, single-binary deploy |
| Async Runtime | Tokio | De-facto Rust async runtime; Axum is built on it | | Async Runtime | Tokio | De-facto Rust async runtime; Axum is built on it |
| Database Driver | SQLx | Async PostgreSQL with compile-time query checking; automatic prepared statements | | Database Driver | SQLx | Async PostgreSQL; automatic prepared statements. **Queries use the runtime `sqlx::query()` API, not the checked macros** — see the note under "Schema changes" |
| Database | PostgreSQL 16 | Robust, relational; straightforward to back up | | Database | PostgreSQL 16 | Robust, relational; straightforward to back up |
| Auth | Custom JWT (`jsonwebtoken` crate) | No external service needed; name + PIN is the full auth model | | Auth | Custom JWT (`jsonwebtoken` crate) | No external service needed; name + PIN is the full auth model |
| Image Compression | `image` crate + `oxipng` | Lossless PNG compression; JPEG preview generation | | Image Compression | `image` crate + `oxipng` | Lossless PNG compression; JPEG preview generation |
@@ -709,7 +709,7 @@ CREATE TABLE config (
INSERT INTO config (key, value) VALUES INSERT INTO config (key, value) VALUES
('max_image_size_mb', '20'), ('max_image_size_mb', '20'),
('max_video_size_mb', '500'), ('max_video_size_mb', '500'),
('upload_rate_per_hour', '10'), ('upload_rate_per_hour', '100'), -- raised from 10 in migration 015 (guests upload bursts of 10-20)
('feed_rate_per_min', '60'), ('feed_rate_per_min', '60'),
('export_rate_per_day', '3'), ('export_rate_per_day', '3'),
('quota_tolerance', '0.75'), ('quota_tolerance', '0.75'),
@@ -1133,16 +1133,19 @@ eventsnap/
### Backup Strategy ### Backup Strategy
```bash Three artefacts in three places: the database, the `media_data` volume
# Daily (e.g. as a separate Compose service or cron on the VPS) (originals + derivatives), and the **separate** `exports_data` volume. See
pg_dump $DATABASE_URL | gzip > /media/backups/db_$(date +%Y-%m-%d).sql.gz [README.md](README.md#backup) for the exact commands.
# Weekly: rsync /media volume to Hetzner Storage Box Everything runs through `docker compose` / `docker run`, because `DATABASE_URL`
rsync -az /opt/eventsnap/media/ \ and the `/media` and `/exports` paths only exist inside the compose network —
user@u123456.your-storagebox.de:backup/eventsnap/ they are not host paths, and `DATABASE_URL` is never exported into an operator's
``` shell.
The `/media` volume contains originals, previews, thumbnails, generated exports, and DB backups — a single volume to back up. Export archives are deliberately outside `MEDIA_PATH` (`EXPORT_PATH=/exports`): a
keepsake contains every photo in the event, and keeping it off the media tree is
what stops it being reachable except through the ticket-gated handler. A backup
of the media volume alone silently loses every generated keepsake.
--- ---
@@ -1203,7 +1206,7 @@ The `/media` volume contains originals, previews, thumbnails, generated exports,
|-------|---------| |-------|---------|
| `axum` | Web framework | | `axum` | Web framework |
| `tokio` | Async runtime | | `tokio` | Async runtime |
| `sqlx` | Async PostgreSQL driver; compile-time query checking; prepared statements; migrations | | `sqlx` | Async PostgreSQL driver; prepared statements; migrations embedded at compile time. Queries are runtime-checked (`sqlx::query()`), **not** macro-checked |
| `jsonwebtoken` | JWT sign / verify | | `jsonwebtoken` | JWT sign / verify |
| `bcrypt` | PIN + admin password hashing | | `bcrypt` | PIN + admin password hashing |
| `uuid` | UUID v7 (time-sortable) | | `uuid` | UUID v7 (time-sortable) |

328
README.md
View File

@@ -34,7 +34,6 @@ A guest scans the QR code on their way in, types their name, and is immediately
### Planned (v1.x) ### Planned (v1.x)
- Individual file download button - Individual file download button
- Low-disk alert (< 10 GB free)
- Event banner / cover image - Event banner / cover image
- Chunked resumable upload for large videos - Chunked resumable upload for large videos
- Host-curated story highlights - Host-curated story highlights
@@ -50,7 +49,7 @@ A guest scans the QR code on their way in, types their name, and is immediately
| Styling | Tailwind CSS v4 | | Styling | Tailwind CSS v4 |
| Backend | Rust + Axum | | Backend | Rust + Axum |
| Async | Tokio | | Async | Tokio |
| Database | PostgreSQL 16 via SQLx (compile-time query checking) | | Database | PostgreSQL 16 via SQLx (runtime query API; migrations embedded at compile time) |
| Auth | Custom JWT (`jsonwebtoken`) + bcrypt PINs | | Auth | Custom JWT (`jsonwebtoken`) + bcrypt PINs |
| Image processing | `image` crate + `oxipng` (lossless compression) | | Image processing | `image` crate + `oxipng` (lossless compression) |
| Video processing | ffmpeg via `tokio::process::Command` | | Video processing | ffmpeg via `tokio::process::Command` |
@@ -98,33 +97,153 @@ eventsnap/
git clone https://git.mc02.dev/fabi/EventSnap.git eventsnap git clone https://git.mc02.dev/fabi/EventSnap.git eventsnap
cd eventsnap cd eventsnap
# 2. Configure environment # 2. Configure environment — set EVERY secret NOW, before step 3.
cp .env.example .env cp .env.example .env
nano .env # set DOMAIN, JWT_SECRET, ADMIN_PASSWORD_HASH, EVENT_NAME, etc. nano .env # DOMAIN, EVENT_NAME, EVENT_SLUG,
# JWT_SECRET, ADMIN_PASSWORD_HASH,
# POSTGRES_PASSWORD *and* the same password inside DATABASE_URL
# (see "Generate required secrets" below)
# 3. Start the stack # 3. Start the stack
docker compose up -d docker compose up -d
``` ```
> **Set every secret before step 3 — `POSTGRES_PASSWORD` especially.** Postgres reads it **only
> when it initialises its data directory**, which happens on the very first `docker compose up -d`.
> Changing it in `.env` afterwards does not change the stored password: the app then authenticates
> with the new one against a volume holding the old one, and you get a permanent restart loop with
> `password authentication failed for user "eventsnap"`. The only fixes are restoring the old
> password or `docker compose down -v`, which **deletes the database, the media and the exports**.
> Getting it right once, up front, costs nothing; getting it wrong costs the volume.
Caddy automatically obtains a Let's Encrypt certificate on first start. The app is live at `https://DOMAIN` within ~30 seconds. Caddy automatically obtains a Let's Encrypt certificate on first start. The app is live at `https://DOMAIN` within ~30 seconds.
> **If the site never comes up:** with `APP_ENV=production` the backend **refuses to boot** while `JWT_SECRET`/`ADMIN_PASSWORD_HASH` still hold the `.env.example` placeholders (this is deliberate — a publicly-known signing key is worse than downtime). Caddy then waits on the unhealthy `app` container and never serves. Check `docker compose logs app` — a "Refusing to start … placeholder …" line means you skipped step 2. Rotate the secrets (see below) and restart. > **If the site never comes up:** with `APP_ENV=production` the backend **refuses to boot** while
> `JWT_SECRET`, `ADMIN_PASSWORD_HASH` or the password inside `DATABASE_URL` still hold the
> `.env.example` placeholders (this is deliberate — a publicly-known signing key or database
> password is worse than downtime). Caddy then waits on the unhealthy `app` container and never
> serves. Check `docker compose logs app` — a "Refusing to start … placeholder …" line lists
> **every** unset secret at once, so one edit fixes them all.
>
> **If it comes up but keeps restarting with `password authentication failed for user
> "eventsnap"`:** `POSTGRES_PASSWORD` was changed after the database volume was created. Postgres
> applies that variable only at initialisation, so `.env` and the stored password have drifted
> apart permanently. `docker compose logs app` spells this out. Before the event, with nothing
> worth keeping:
>
> ```bash
> docker compose down -v && docker compose up -d # -v DELETES db + media + exports. No undo.
> ```
>
> **Once the event has real data, never do that.** Put the original password back into
> `DATABASE_URL`, or change the stored one instead:
>
> ```bash
> docker compose exec db psql -U "$POSTGRES_USER" -c \
> "ALTER ROLE eventsnap WITH PASSWORD 'the-password-now-in-your-.env';"
> ```
> **Production note:** `docker compose up -d` does **not** expose the database — Postgres is reachable only on the internal Docker network. For local development where you need host access to Postgres, opt into the dev overlay explicitly: > **Production note:** `docker compose up -d` does **not** expose the database — Postgres is reachable only on the internal Docker network. For local development where you need host access to Postgres, opt into the dev overlay explicitly:
> ```bash > ```bash
> docker compose -f docker-compose.yml -f docker-compose.dev.yml up > docker compose -f docker-compose.yml -f docker-compose.dev.yml up
> ``` > ```
### Updating an existing deployment
> **The event server never compiles.** `app` and `frontend` have **no `build:` key** — they
> pull an immutable tag from the registry (the `app` service in `docker-compose.yml` says so explicitly, so that
> a wrong tag fails instantly with `manifest unknown` instead of silently starting a 45-minute
> compile on the box guests are using). A `git pull` therefore deploys **nothing** on its own,
> and `docker compose up -d --build` **errors** — there is nothing to build. Deploying means
> pushing a new tag from a workstation and pointing `EVENTSNAP_VERSION` at it.
```bash
# ── On your workstation: build and push the new tag ───────────────────────────
# Push the rollback twin at the same time, from the same source — see
# DEPLOYMENT_RUNBOOK.md §9 for why an identical second tag is the rollback target.
VERSION=v0.13.1
docker buildx build --platform linux/amd64 \
-t registry.mc02.dev/eventsnap/app:$VERSION \
-t registry.mc02.dev/eventsnap/app:$VERSION-a --push ./backend
docker buildx build --platform linux/amd64 \
-t registry.mc02.dev/eventsnap/frontend:$VERSION \
-t registry.mc02.dev/eventsnap/frontend:$VERSION-a --push ./frontend
# ── On the server ─────────────────────────────────────────────────────────────
cd /path/to/eventsnap
# 1. Back up first — migrations run automatically on boot and are not reversible in place.
# (See "Backup" below; the database dump is the one that matters here.)
# 2. Fetch the new compose/Caddyfile. This does NOT change which image runs.
git pull
# 3. Point the stack at the new tag.
sed -i 's/^EVENTSNAP_VERSION=.*/EVENTSNAP_VERSION=v0.13.1/' .env
# 4. Pull explicitly, BEFORE restarting. A failure here (bad tag, registry down) leaves the
# running stack untouched; letting `up -d` discover it takes the app down first.
docker compose pull app frontend
# 5. Restart onto the new images.
docker compose up -d app frontend
# 6. Apply any Caddyfile change. Step 5 does NOT do this — see the warning below.
docker compose up -d --force-recreate caddy
# 7. Confirm the app came back up. Anything other than "ok" means check the logs.
curl -fsS https://DOMAIN/health && echo
# 8. Confirm the running containers are actually on the new tag.
docker compose images app frontend
```
Migrations are applied by the backend on startup, so step 5 covers them. If `app` stays
unhealthy afterwards, `docker compose logs app` will name the failing migration — and note
that a migration applied by a *newer* build is not removed by rolling the tag back, so
reverting `EVENTSNAP_VERSION` without restoring the database snapshot from step 1 leaves the
schema ahead of the binary and the app refusing to boot. **This is why the rollback target is
an identical twin tag rather than an older release** — see `DEPLOYMENT_RUNBOOK.md` §9.
> **Why step 6 exists.** Steps 45 only touch `app` and `frontend`; `caddy` is a separate
> pinned upstream image. Compose decides whether to recreate a container from its
> *config hash*, which covers the mount **specification** (`./Caddyfile:/etc/caddy/Caddyfile:ro`)
> but **not the file's contents** — so a `git pull` that changes `./Caddyfile` produces no
> delta, Compose reports `Running`, and Caddy keeps serving its old config indefinitely. Exit
> code 0 throughout.
>
> That is not hypothetical: the fix that made the keepsake download work on iOS
> (`137c4ee`) touched the Caddyfile and four e2e files and nothing else, so **all** of its
> production effect lives in that one file. Without step 6 you deploy it, watch both image IDs
> change, and iOS downloads stay broken.
>
> `--force-recreate` rather than `restart` or `caddy reload`: the bind mount is resolved to an
> **inode** when the container is created, and `git pull` replaces the file instead of editing
> it in place, so the container can still be bound to the old, now-unlinked inode. A restart
> then re-reads the stale content. Recreating the container re-resolves the path.
`db` is never touched, and recreating `caddy` does not disturb the `caddy_data` volume, so the
TLS certificate and all data volumes survive.
### Generate required secrets ### Generate required secrets
```bash ```bash
# JWT secret (64 random bytes) # JWT secret (64 random bytes)
openssl rand -hex 64 openssl rand -hex 64
# Admin password hash (bcrypt, cost 12) # Database password (goes in BOTH DATABASE_URL and POSTGRES_PASSWORD)
htpasswd -bnBC 12 "" yourpassword | tr -d ':\n' openssl rand -hex 24
# Admin password hash (bcrypt). Uses an image the stack already pulls, so it needs
# nothing installed on the host — `htpasswd` lives in apache2-utils, which a stock
# VPS does not have. Emits cost 14 rather than 12; that is fine (admin login is
# rate-limited and hashed off the async runtime), and any $2a/$2b/$2y hash verifies.
docker run --rm caddy:2-alpine caddy hash-password --plaintext 'yourpassword'
``` ```
Wrap the resulting hash in **single quotes** in `.env` — see the note there; a bcrypt
hash is full of `$`, and both Compose and dotenvy would otherwise eat those segments.
### Environment Variables ### Environment Variables
See [.env.example](.env.example) for the full list with descriptions and defaults. Key variables: See [.env.example](.env.example) for the full list with descriptions and defaults. Key variables:
@@ -162,23 +281,200 @@ See [.env.example](.env.example) for the full list with descriptions and default
└────────┘ └────────┘
``` ```
- `/api/*` and `/media/*` → Rust backend - `/api/*` → Rust backend
- Everything else → SvelteKit frontend (`adapter-node`) - Everything else → SvelteKit frontend (`adapter-node`)
- Named volumes: `postgres_data`, `media_data`, `caddy_data` - Named volumes: `postgres_data`, `media_data`, `exports_data`, `caddy_data`
Media is **not** served as static files. Every image goes through a
visibility-checked alias (`/api/v1/upload/{id}/{preview,display,thumbnail,original}`)
so a host takedown or a ban actually revokes access to the bytes.
---
## Sizing the disk
`postgres_data`, `media_data` and `exports_data` are all Docker named volumes under
`/var/lib/docker/volumes`, so **they share one filesystem**. Filling it does not
degrade one subsystem — Postgres stops being able to write and the whole event goes
down.
**`Gallery.zip` and `Memories.zip` are each roughly a second copy of every original.**
Both write their media `Compression::Stored`, and `Memories.zip` streams the untouched
original for every video and for every image at or under 5 MB. So a release wants room
for **two more copies of the gallery** on top of the gallery itself — which is what
`required_free_bytes` encodes as `media × 1.1 × 2`.
The per-user quota does **not** bound this. It is a fairness mechanism that divides
free space between guests, and since it carries a floor (`MIN_QUOTA_LIMIT_BYTES`, so a
guest's allowance stops shrinking as the party fills up) the aggregate ceiling it used
to imply is gone. What bounds the disk is the **global gate in the upload handler**,
which refuses any upload that would leave too little room to build the keepsake:
```
free_after_upload < media_after × 1.1 × 2 + DISK_RESERVE_BYTES
+ UPLOAD_GATE_HEADROOM_BYTES → refused
```
That last term is what separates this gate from the export preflight, which bails at
`media × 1.1 × 2 + DISK_RESERVE_BYTES` — the same expression **minus** the headroom. The
two used to be identical, which meant the preflight was already sitting on its limit at
the exact moment uploads stopped: every byte written between the last refused upload and
the host tapping *Galerie freigeben* (Postgres WAL, container logs, the compression
backlog draining at precisely that hour) pushed it under, and the release commits before
the workers fail. The headroom buys 1.5 GB of slack so that cannot happen.
Solving the gate for the gallery size gives the real ceiling — the gate's equilibrium is
`3.2 × media`, so each GB of reserve or headroom costs ~0.31 GB of gallery. On the
**40 GB box this runs on**, with ~5 GB for the OS, Docker images (the runbook pre-pulls
the rollback tag too) and Postgres:
| Volume | Usable after baseline | Media ceiling | Free at release | Preflight needs |
|---|---|---|---|---|
| 40 GB | ~35 GB | **~7.3 GB** | ~27.7 GB | ~26.2 GB → fits, 1.5 GB spare |
| 80 GB | ~70 GB | ~18.3 GB | ~51.7 GB | ~50.2 GB → fits, 1.5 GB spare |
**Uploads therefore stop at roughly 7 GB of media on a 40 GB box, not when the disk is
full.** That is deliberate. 1000 photos at ~3.5 MB is ~3.5 GB and fits comfortably;
video is what consumes the budget, so lower `max_video_size_mb` (seeded at 500) if you
expect a lot of it. Refusing the 1001st upload is a far better outcome than accepting it
and discovering at 01:00 that the archive can never be built.
Two ways to buy headroom:
- **Provision ~3× your expected media** on one volume (media + two archives), or
- **give `exports_data` its own volume** so a full export cannot reach Postgres, and
size that one at ~2× expected media.
None of this is silent. The upload gate refuses with a German message naming the cause,
the export preflight refuses up front with both numbers rather than hitting ENOSPC
halfway through a multi-GB write, a rebuild only reclaims the superseded generation
**after** the new one lands (so a failed rebuild can never leave you with no archive at
all), and the host dashboard warns as soon as the keepsake would not fit — which is the
only point at which anyone can still do something about it.
--- ---
## Backup ## Backup
```bash There are **three** things to back up, and they live in three different places.
# Database snapshot `DATABASE_URL` and the container paths (`/media`, `/exports`) are meaningful only
pg_dump $DATABASE_URL | gzip > /media/backups/db_$(date +%Y-%m-%d).sql.gz *inside* the compose network — they are not host paths, and `DATABASE_URL` is
never exported into an operator's shell — so every command below runs through
`docker compose` from the repo directory.
# Weekly offsite sync (Hetzner Storage Box or similar) ```bash
rsync -az /opt/eventsnap/media/ user@storagebox.example.com:backup/eventsnap/ # 1. Database snapshot. Runs pg_dump inside the db container (the app image has no
# postgres client), reading credentials from the compose environment.
# --clean --if-exists makes the dump SELF-CLEANING: without it the restore below
# aborts on the first "already exists" against a database that has ever booted,
# which is every database you would actually want to restore over.
mkdir -p ./backups
docker compose exec -T db \
sh -c 'pg_dump --clean --if-exists -U "$POSTGRES_USER" "$POSTGRES_DB"' \
| gzip > ./backups/db_$(date +%Y-%m-%d).sql.gz
# 2. Uploaded media (originals + derivatives) out of the named volume.
# NOTE the mountpoint is /src, not /media: if the volume is ever empty, Docker
# pre-populates a fresh mount from the image's own directory, and alpine ships a
# /media containing cdrom/floppy/usb. Mounting somewhere the image has nothing
# avoids silently tarring (and polluting the volume with) those.
docker run --rm \
-v eventsnap_media_data:/src:ro -v "$PWD/backups":/backup \
alpine tar czf /backup/media_$(date +%Y-%m-%d).tar.gz -C /src .
# 3. Export archives — a SEPARATE volume (see the security note below).
docker run --rm \
-v eventsnap_exports_data:/src:ro -v "$PWD/backups":/backup \
alpine tar czf /backup/exports_$(date +%Y-%m-%d).tar.gz -C /src .
# Offsite sync of the three artefacts above.
rsync -az ./backups/ user@storagebox.example.com:backup/eventsnap/
``` ```
The `/media` volume holds originals, previews, thumbnails, exports, and DB backups — a single path to back up. Volume names are prefixed with the compose project name — `eventsnap_` if you run
from a directory called `eventsnap`. Confirm yours with `docker volume ls`.
> **Exports are deliberately NOT under `/media`.** They live on their own
> `exports_data` volume (`EXPORT_PATH=/exports`) because a keepsake archive
> contains every photo in the event; keeping it outside the media tree is what
> stops it being reachable except through the ticket-gated download handler.
> Backing up only the media volume therefore loses every generated keepsake.
### When to run it
**A nightly cron is the wrong shape for this app.** Every irreplaceable byte is
created inside one eight-hour window, and nobody can retake a wedding. Run the three
commands above:
1. **The night of the event**, once uploads have stopped. This is the backup that
matters; everything else is a formality.
2. **After the host releases the gallery**, so the generated keepsake is captured too.
3. Weekly thereafter, until the event is archived and torn down.
Take the DB dump and the media tarball **back to back**, without uploads in flight
between them. Upload rows reference files by path — a database from 22:00 and a media
volume from 23:00 gives you rows pointing at files the dump doesn't know about, and
rows whose files aren't in the tarball. Locking uploads from the host dashboard first
(**Uploads sperren**) makes the pair genuinely consistent.
---
## Restore
An untested backup is not a backup. Run this once against a scratch host **before**
the event — it is roughly ten minutes, and it is the only way to find out that your
tarball is empty or your dump is truncated while that is still a small problem.
```bash
# 0. Stop the app FIRST. Migrations run on boot and a live pool will fight the
# restore — a booting app against a half-restored schema can leave the migration
# table and the schema disagreeing, which is its own recovery problem.
# Leave `db` running: the dump is restored through it.
docker compose stop app caddy
# 1. Database. The dump carries its own DROPs (step 1 of Backup), so this replaces
# rather than collides. A dump taken WITHOUT --clean --if-exists will abort here
# on the first "already exists" — restore that one into a fresh empty database
# instead.
gunzip -c ./backups/db_2026-07-29.sql.gz \
| docker compose exec -T db \
sh -c 'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" --set ON_ERROR_STOP=1'
# 2. Media. NOTE the `--numeric-owner` and the chown: the app runs as a
# NON-ROOT user (uid 100, gid 101 — `addgroup -S app && adduser -S app`), and a
# restore that lands root-owned files makes every upload fail with EACCES deep in
# the write path, surfacing to the guest as a generic 500 with nothing in the UI
# to suggest permissions. The explicit chown is what guarantees it — BusyBox tar
# (which is what `alpine` ships) has no --same-owner, and restores ownership only
# because it runs as root here.
docker run --rm \
-v eventsnap_media_data:/dst -v "$PWD/backups":/backup:ro \
alpine sh -c 'tar xzf /backup/media_2026-07-29.tar.gz -C /dst \
--numeric-owner && chown -R 100:101 /dst'
# 3. Exports. Same volume-name caveat, same ownership rules.
docker run --rm \
-v eventsnap_exports_data:/dst -v "$PWD/backups":/backup:ro \
alpine sh -c 'tar xzf /backup/exports_2026-07-29.tar.gz -C /dst \
--numeric-owner && chown -R 100:101 /dst'
# 4. Back up. Migrations run, then export recovery re-arms any keepsake whose file
# didn't come back with the volume.
docker compose up -d app caddy
docker compose logs -f app # watch for "migrations applied"
# 5. Verify — all three, not just the first.
curl -fsS https://DOMAIN/health && echo # → ok
# … then sign in as host and confirm the feed renders images (proves the media
# volume restored AND is readable by uid 100), and that the keepsake downloads.
```
If the media volume restored but images 404 while the feed lists them, the paths are
there and the bytes aren't — check `docker compose exec app ls -ln /media/originals`
and confirm both the files and the `100:101` ownership.
The restore is deliberately **not** automated. It is rare, destructive, and the one
operation where a script that half-works is worse than a checklist someone reads.
--- ---
@@ -254,7 +550,7 @@ Open:
- [ ] SSE delta-fetch on foreground reconnect (scaffolded in [sse.ts](frontend/src/lib/sse.ts), not wired) - [ ] SSE delta-fetch on foreground reconnect (scaffolded in [sse.ts](frontend/src/lib/sse.ts), not wired)
- [ ] Live diashow / slideshow mode — see [docs/CONCEPT_DIASHOW.md](docs/CONCEPT_DIASHOW.md) - [ ] Live diashow / slideshow mode — see [docs/CONCEPT_DIASHOW.md](docs/CONCEPT_DIASHOW.md)
- [ ] Individual file download button per post - [ ] Individual file download button per post
- [ ] Low-disk alert (< 10 GB free) - [x] Low-disk alert — host dashboard warns below 10 GB free, or whenever the keepsake would not fit
- [ ] Event banner / cover image - [ ] Event banner / cover image
- [ ] Chunked resumable upload for files > 100 MB - [ ] Chunked resumable upload for files > 100 MB
- [ ] Shared Tailwind config between main app and export-viewer - [ ] Shared Tailwind config between main app and export-viewer

12
backend/.dockerignore Normal file
View File

@@ -0,0 +1,12 @@
# Build context exclusions.
#
# `target/` does not exist on a clean checkout, which is why builds have worked without this
# file — but the moment anyone runs `cargo build` locally it becomes multi-GB, and every
# `docker build` would ship all of it to the daemon for nothing (the Dockerfile only COPYs
# Cargo.toml, Cargo.lock, src, static and migrations). On an EMULATED amd64 builder that
# transfer is the slowest part of the build.
target/
# Never let a real .env reach an image layer.
.env
.env.*

196
backend/Cargo.lock generated
View File

@@ -65,56 +65,6 @@ dependencies = [
"libc", "libc",
] ]
[[package]]
name = "anstream"
version = "1.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "824a212faf96e9acacdbd09febd34438f8f711fb84e09a8916013cd7815ca28d"
dependencies = [
"anstyle",
"anstyle-parse",
"anstyle-query",
"anstyle-wincon",
"colorchoice",
"is_terminal_polyfill",
"utf8parse",
]
[[package]]
name = "anstyle"
version = "1.0.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "940b3a0ca603d1eade50a4846a2afffd5ef57a9feac2c0e2ec2e14f9ead76000"
[[package]]
name = "anstyle-parse"
version = "1.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "52ce7f38b242319f7cabaa6813055467063ecdc9d355bbb4ce0c68908cd8130e"
dependencies = [
"utf8parse",
]
[[package]]
name = "anstyle-query"
version = "1.1.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "40c48f72fd53cd289104fc64099abca73db4166ad86ea0b4341abe65af83dadc"
dependencies = [
"windows-sys 0.61.2",
]
[[package]]
name = "anstyle-wincon"
version = "3.0.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "291e6a250ff86cd4a820112fb8898808a366d8f9f58ce16d1f538353ad55747d"
dependencies = [
"anstyle",
"once_cell_polyfill",
"windows-sys 0.61.2",
]
[[package]] [[package]]
name = "anyhow" name = "anyhow"
version = "1.0.102" version = "1.0.102"
@@ -554,46 +504,12 @@ dependencies = [
"inout", "inout",
] ]
[[package]]
name = "clap"
version = "4.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b193af5b67834b676abd72466a96c1024e6a6ad978a1f484bd90b85c94041351"
dependencies = [
"clap_builder",
]
[[package]]
name = "clap_builder"
version = "4.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "714a53001bf66416adb0e2ef5ac857140e7dc3a0c48fb28b2f10762fc4b5069f"
dependencies = [
"anstream",
"anstyle",
"clap_lex",
"strsim",
"terminal_size",
]
[[package]]
name = "clap_lex"
version = "1.1.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9"
[[package]] [[package]]
name = "color_quant" name = "color_quant"
version = "1.1.0" version = "1.1.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3d7b894f5411737b7867f4827955924d7c254fc9f4d91a6aad6b097804b1018b" checksum = "3d7b894f5411737b7867f4827955924d7c254fc9f4d91a6aad6b097804b1018b"
[[package]]
name = "colorchoice"
version = "1.0.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1d07550c9036bf2ae0c684c4297d503f838287c83c53686d05370d0e139ae570"
[[package]] [[package]]
name = "compression-codecs" name = "compression-codecs"
version = "0.4.37" version = "0.4.37"
@@ -677,15 +593,6 @@ dependencies = [
"cfg-if", "cfg-if",
] ]
[[package]]
name = "crossbeam-channel"
version = "0.5.15"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "82b8f8f868b36967f9606790d1903570de9ceaf870a7bf9fbbd3016d636a2cb2"
dependencies = [
"crossbeam-utils",
]
[[package]] [[package]]
name = "crossbeam-deque" name = "crossbeam-deque"
version = "0.8.6" version = "0.8.6"
@@ -698,9 +605,9 @@ dependencies = [
[[package]] [[package]]
name = "crossbeam-epoch" name = "crossbeam-epoch"
version = "0.9.18" version = "0.9.20"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5b82ac4a3c2ca9c3460964f020e1402edd5753411d7737aa39c3714ad1b5420e" checksum = "2d6914041f254d6e9176c01941b21115dcfb7089e55135a35411081bd106ef3f"
dependencies = [ dependencies = [
"crossbeam-utils", "crossbeam-utils",
] ]
@@ -816,27 +723,6 @@ dependencies = [
"cfg-if", "cfg-if",
] ]
[[package]]
name = "env_filter"
version = "1.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32e90c2accc4b07a8456ea0debdc2e7587bdd890680d71173a15d4ae604f6eef"
dependencies = [
"log",
]
[[package]]
name = "env_logger"
version = "0.11.10"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0621c04f2196ac3f488dd583365b9c09be011a4ab8b9f37248ffcc8f6198b56a"
dependencies = [
"anstream",
"anstyle",
"env_filter",
"log",
]
[[package]] [[package]]
name = "equator" name = "equator"
version = "0.4.2" version = "0.4.2"
@@ -1229,12 +1115,6 @@ dependencies = [
"weezl", "weezl",
] ]
[[package]]
name = "glob"
version = "0.3.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0cc23270f6e1808e30a928bdc84dea0b9b4136a8bc82338574f23baf47bbd280"
[[package]] [[package]]
name = "governor" name = "governor"
version = "0.6.3" version = "0.6.3"
@@ -1622,7 +1502,6 @@ checksum = "7714e70437a7dc3ac8eb7e6f8df75fd8eb422675fc7678aff7364301092b1017"
dependencies = [ dependencies = [
"equivalent", "equivalent",
"hashbrown 0.16.1", "hashbrown 0.16.1",
"rayon",
"serde", "serde",
"serde_core", "serde_core",
] ]
@@ -1656,12 +1535,6 @@ dependencies = [
"syn", "syn",
] ]
[[package]]
name = "is_terminal_polyfill"
version = "1.70.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a6cb138bb79a146c1bd460005623e142ef0181e3d0219cb493e02f7d08a35695"
[[package]] [[package]]
name = "itertools" name = "itertools"
version = "0.14.0" version = "0.14.0"
@@ -1795,12 +1668,6 @@ dependencies = [
"vcpkg", "vcpkg",
] ]
[[package]]
name = "linux-raw-sys"
version = "0.12.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53"
[[package]] [[package]]
name = "litemap" name = "litemap"
version = "0.8.1" version = "0.8.1"
@@ -2089,12 +1956,6 @@ version = "1.21.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50" checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
[[package]]
name = "once_cell_polyfill"
version = "1.70.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]] [[package]]
name = "oxipng" name = "oxipng"
version = "9.1.5" version = "9.1.5"
@@ -2102,18 +1963,12 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "26c613f0f566526a647c7473f6a8556dbce22c91b13485ee4b4ec7ab648e4973" checksum = "26c613f0f566526a647c7473f6a8556dbce22c91b13485ee4b4ec7ab648e4973"
dependencies = [ dependencies = [
"bitvec", "bitvec",
"clap",
"crossbeam-channel",
"env_logger",
"filetime", "filetime",
"glob",
"indexmap", "indexmap",
"libdeflater", "libdeflater",
"log", "log",
"rayon",
"rgb", "rgb",
"rustc-hash", "rustc-hash",
"zopfli",
] ]
[[package]] [[package]]
@@ -2607,19 +2462,6 @@ version = "2.1.2"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "94300abf3f1ae2e2b8ffb7b58043de3d399c73fa6f4b73826402a5c457614dbe" checksum = "94300abf3f1ae2e2b8ffb7b58043de3d399c73fa6f4b73826402a5c457614dbe"
[[package]]
name = "rustix"
version = "1.1.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b6fe4565b9518b83ef4f91bb47ce29620ca828bd32cb7e408f0062e9930ba190"
dependencies = [
"bitflags",
"errno",
"libc",
"linux-raw-sys",
"windows-sys 0.61.2",
]
[[package]] [[package]]
name = "rustversion" name = "rustversion"
version = "1.0.22" version = "1.0.22"
@@ -3060,12 +2902,6 @@ dependencies = [
"unicode-properties", "unicode-properties",
] ]
[[package]]
name = "strsim"
version = "0.11.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7da8b5736845d9f2fcb837ea5d9e2628564b3b043a70948a3f0b778838c5fb4f"
[[package]] [[package]]
name = "subtle" name = "subtle"
version = "2.6.1" version = "2.6.1"
@@ -3120,16 +2956,6 @@ version = "1.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "55937e1799185b12863d447f42597ed69d9928686b8d88a1df17376a097d8369" checksum = "55937e1799185b12863d447f42597ed69d9928686b8d88a1df17376a097d8369"
[[package]]
name = "terminal_size"
version = "0.4.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "230a1b821ccbd75b185820a1f1ff7b14d21da1e442e22c0863ea5f08771a8874"
dependencies = [
"rustix",
"windows-sys 0.61.2",
]
[[package]] [[package]]
name = "thiserror" name = "thiserror"
version = "1.0.69" version = "1.0.69"
@@ -3505,12 +3331,6 @@ version = "1.0.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b6c140620e7ffbb22c2dee59cafe6084a59b5ffc27a8859a5f0d494b5d52b6be" checksum = "b6c140620e7ffbb22c2dee59cafe6084a59b5ffc27a8859a5f0d494b5d52b6be"
[[package]]
name = "utf8parse"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "06abde3611657adf66d383f00b093d7faecc7fa57071cce2578660c9f1010821"
[[package]] [[package]]
name = "uuid" name = "uuid"
version = "1.23.0" version = "1.23.0"
@@ -4187,18 +4007,6 @@ version = "1.0.21"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8848ee67ecc8aedbaf3e4122217aff892639231befc6a1b58d29fff4c2cabaa" checksum = "b8848ee67ecc8aedbaf3e4122217aff892639231befc6a1b58d29fff4c2cabaa"
[[package]]
name = "zopfli"
version = "0.8.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f05cd8797d63865425ff89b5c4a48804f35ba0ce8d125800027ad6017d2b5249"
dependencies = [
"bumpalo",
"crc32fast",
"log",
"simd-adler32",
]
[[package]] [[package]]
name = "zstd" name = "zstd"
version = "0.13.3" version = "0.13.3"

View File

@@ -27,7 +27,17 @@ tracing-subscriber = { version = "0.3", features = ["env-filter"] }
dotenvy = "0.15" dotenvy = "0.15"
sysinfo = "0.32" sysinfo = "0.32"
image = "0.25" image = "0.25"
oxipng = "9" # default-features = false drops "parallel", which is what actually bounds oxipng's memory:
# with rayon it evaluates row filters concurrently, each trial holding its own full-size
# buffer, and there is no Options knob to cap that. Without the feature, lib.rs swaps in a
# sequential shim (oxipng's own supported path) so peak scales with ONE trial, not N.
# PNG optimisation gets slower; it is a background, best-effort, lossless size saving.
#
# "filetime" must be KEPT: without it OutFile::Path { preserve_attrs: true } silently no-ops.
# Dropping "binary" also removes clap/glob/env_logger — a CLI's dependencies that were being
# compiled into a server image — and "zopfli", which preset 2 does not use (it selects
# Deflaters::Libdeflater, which is not feature-gated).
oxipng = { version = "9", default-features = false, features = ["filetime"] }
async_zip = { version = "0.0.17", features = ["tokio", "deflate"] } async_zip = { version = "0.0.17", features = ["tokio", "deflate"] }
include_dir = "0.7" include_dir = "0.7"
infer = "0.15" infer = "0.15"

View File

@@ -13,6 +13,13 @@ RUN mkdir src && echo "fn main(){}" > src/main.rs && \
COPY src ./src COPY src ./src
COPY static ./static COPY static ./static
COPY migrations ./migrations COPY migrations ./migrations
# Copied WITH the sources, not with Cargo.toml above: cargo auto-detects `build.rs` by presence, so
# putting it in the dependency-cache layer would make the dummy build run it too and invalidate a
# layer that is otherwise stable. Copied at all because without it the image builds a subtly
# DIFFERENT package from the one developers build — no build script, hence none of the
# rerun-if-changed tracking for `static/export-viewer` and `migrations`. Harmless here (every image
# build is clean, so there is no stale cache to reuse) and confusing everywhere else.
COPY build.rs ./
RUN touch src/main.rs && cargo build --release RUN touch src/main.rs && cargo build --release
# --- Runtime stage --- # --- Runtime stage ---

25
backend/build.rs Normal file
View File

@@ -0,0 +1,25 @@
//! Tell cargo which non-Rust inputs are baked into the binary.
//!
//! `include_dir!` and `sqlx::migrate!()` both embed directory contents at COMPILE time, and neither
//! registers a rebuild dependency on its own. Cargo therefore reuses a cached binary when only
//! those directories changed — the source files are untouched, so as far as cargo is concerned
//! nothing happened.
//!
//! For the keepsake viewer that is a silent, shippable defect: run `npm run build` in
//! `frontend/export-viewer`, then `cargo build`, and the resulting binary still carries the
//! PREVIOUS `static/export-viewer/index.html`. The artifact on disk and the artifact in the binary
//! disagree, `git status` is clean, and every check passes — while `Memories.zip` ships a stale
//! viewer. Confirmed empirically: after replacing the file, the compiled-in copy did not change
//! until a source file was touched.
//!
//! Production is mostly insulated because images are built from a clean context (no cache to
//! reuse), but every incremental build — i.e. all local development and any test run that follows
//! a viewer rebuild — hits it, and that includes the test that asserts the viewer is present.
fn main() {
// The compiled-in keepsake viewer (services/export.rs: `include_dir!`).
println!("cargo:rerun-if-changed=static/export-viewer");
// The embedded migration set (db.rs: `sqlx::migrate!()`). Same mechanism, and the failure is
// worse: a binary built from a stale snapshot boots against a database that has already run a
// newer migration and crash-loops with VersionMissing.
println!("cargo:rerun-if-changed=migrations");
}

View File

@@ -0,0 +1,3 @@
-- Revert the default upload rate to 10/hour for installs still on the raised
-- default (preserves any explicit admin override at another value).
UPDATE config SET value = '10' WHERE key = 'upload_rate_per_hour' AND value = '100';

View File

@@ -0,0 +1,10 @@
-- Raise the default per-guest upload rate from 10/hour to 100/hour.
--
-- Rationale: guests routinely upload a burst of 10-20 photos at once (phone
-- multi-select). At the old default of 10/hour a real guest's first burst was
-- throttled — surfaced by the 2026-07-18 load test. 100/hour comfortably covers
-- several bursts across an event while still bounding abuse.
--
-- Only bump installs still on the old default; an admin who deliberately set a
-- different value keeps it (migration 005 seeded 10; this UPDATE is scoped to '10').
UPDATE config SET value = '100' WHERE key = 'upload_rate_per_hour' AND value = '10';

View File

@@ -0,0 +1,28 @@
-- Drop the view (frees the column dependency), remove the column, then restore the
-- pre-016 view definition (matches migration 011).
DROP VIEW IF EXISTS v_feed;
ALTER TABLE upload DROP COLUMN display_path;
CREATE VIEW v_feed AS
SELECT
u.id,
u.event_id,
u.user_id,
usr.display_name AS uploader_name,
usr.is_banned,
usr.uploads_hidden,
u.preview_path,
u.thumbnail_path,
u.mime_type,
u.caption,
u.created_at,
COUNT(DISTINCT l.user_id) AS like_count,
COUNT(DISTINCT c.id) AS comment_count
FROM upload u
JOIN "user" usr ON u.user_id = usr.id
LEFT JOIN "like" l ON l.upload_id = u.id
LEFT JOIN comment c ON c.upload_id = u.id AND c.deleted_at IS NULL
WHERE u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE
GROUP BY u.id, usr.display_name, usr.is_banned, usr.uploads_hidden;

View File

@@ -0,0 +1,34 @@
-- Display derivative: a big-screen-quality image (~2048px long edge) for the diashow.
-- The 800px `preview_path` is sized for phone feeds (data saver); upscaled on a projector
-- it looks soft. The diashow uses `display_path` instead — bounded in size (safe to decode
-- on weak kiosk hardware) yet sharp on 1080p/4K. NULL until the compression worker (or the
-- one-time backfill) generates it; consumers fall back to the original when absent.
ALTER TABLE upload ADD COLUMN display_path TEXT;
-- Recreate (not CREATE OR REPLACE, which only allows appending columns at the end) so the
-- new column can sit alongside preview_path/thumbnail_path.
DROP VIEW IF EXISTS v_feed;
CREATE VIEW v_feed AS
SELECT
u.id,
u.event_id,
u.user_id,
usr.display_name AS uploader_name,
usr.is_banned,
usr.uploads_hidden,
u.preview_path,
u.thumbnail_path,
u.display_path,
u.mime_type,
u.caption,
u.created_at,
COUNT(DISTINCT l.user_id) AS like_count,
COUNT(DISTINCT c.id) AS comment_count
FROM upload u
JOIN "user" usr ON u.user_id = usr.id
LEFT JOIN "like" l ON l.upload_id = u.id
LEFT JOIN comment c ON c.upload_id = u.id AND c.deleted_at IS NULL
WHERE u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE
GROUP BY u.id, usr.display_name, usr.is_banned, usr.uploads_hidden;

View File

@@ -0,0 +1 @@
DELETE FROM config WHERE key IN ('join_ip_rate_per_min', 'admin_login_rate_enabled');

View File

@@ -0,0 +1,18 @@
-- Per-IP flood ceiling for /join, and the `admin_login_rate_enabled` toggle that
-- every prior migration forgot to seed.
--
-- Rationale: /join was throttled at 5 requests per 60s keyed on the client IP. At a
-- venue every guest is behind one NAT, so the whole party shared a single bucket —
-- 12 guests scanning the QR code within a few seconds meant 5 got in and 7 were
-- turned away. The handler now keys the real anti-spam bucket per (ip, name), the
-- same shape as `recover:{ip}:{name}`, and keeps only a loose per-IP ceiling to bound
-- raw volume. 60/min comfortably covers a whole wedding arriving at once while still
-- capping a flood from a single source.
--
-- `admin_login_rate_enabled` is read by auth::handlers::admin_login with a code
-- default of `true`, but no migration ever inserted it, so it was invisible to the
-- admin config UI and to the e2e reseed. Seed it explicitly.
INSERT INTO config (key, value) VALUES
('join_ip_rate_per_min', '60'),
('admin_login_rate_enabled', 'true')
ON CONFLICT (key) DO NOTHING;

View File

@@ -0,0 +1,2 @@
DROP INDEX IF EXISTS idx_upload_derivatives_rev;
ALTER TABLE upload DROP COLUMN IF EXISTS derivatives_rev;

View File

@@ -0,0 +1,18 @@
-- Track which revision of the derivative pipeline produced an upload's preview/display.
--
-- Rev 1 applies the EXIF orientation tag. Everything generated before it decoded the raw
-- sensor pixels and re-encoded to JPEG (which writes no EXIF), so every portrait phone photo
-- was stored sideways in the feed preview, the diashow display and the keepsake — while the
-- untouched original still rendered upright.
--
-- Existing rows default to 0 so the startup backfill can find and re-generate them exactly
-- once; bump the constant in services/compression.rs if the pipeline ever changes again.
ALTER TABLE upload ADD COLUMN IF NOT EXISTS derivatives_rev SMALLINT NOT NULL DEFAULT 0;
-- Only image derivatives are affected — video thumbnails are extracted by ffmpeg, which
-- already honours the rotation matrix. Mark them current so the backfill skips them.
UPDATE upload SET derivatives_rev = 1 WHERE mime_type NOT LIKE 'image/%';
CREATE INDEX IF NOT EXISTS idx_upload_derivatives_rev
ON upload (derivatives_rev)
WHERE deleted_at IS NULL;

View File

@@ -0,0 +1 @@
DELETE FROM config WHERE key = 'recover_ip_rate_per_min';

View File

@@ -0,0 +1,19 @@
-- Per-IP flood ceiling for /recover, mirroring the one migration 017 added for /join.
--
-- Rationale: /recover is keyed `recover:{ip}:{name}` at 5 per 15 minutes. That is the
-- right shape for its actual job — stopping someone who knows a display name (they are
-- visible on the feed) from burning the victim's 3-strike PIN counter and locking them
-- out repeatedly. But the name is attacker-chosen, so cycling names mints a fresh bucket
-- every time and the per-IP cost is unbounded.
--
-- Behind that limiter sits a cost-12 bcrypt verify, including an UNCONDITIONAL throwaway
-- verify for names that don't exist — deliberately, to close a timing oracle. So an
-- unknown name is the cheapest possible way to make the server do ~200ms of hashing.
-- Without a ceiling, one client can saturate the box's CPU with a name generator.
--
-- 30/min is far above any real recovery attempt (a guest tries their PIN a handful of
-- times) while capping a name-cycling flood. The per-(ip, name) bucket is unchanged and
-- remains the anti-guessing control.
INSERT INTO config (key, value) VALUES
('recover_ip_rate_per_min', '30')
ON CONFLICT (key) DO NOTHING;

View File

@@ -0,0 +1 @@
DELETE FROM config WHERE key IN ('social_rate_per_min', 'social_rate_enabled');

View File

@@ -0,0 +1,16 @@
-- Per-user rate limit for social writes (likes, comments, comment deletions).
--
-- These were the only writes in the app with no limit at all. Every other mutating
-- path -- upload, join, recover, export, admin login -- carries one; social.rs
-- carried none, so the coverage was asymmetric rather than deliberately open.
--
-- Severity is genuinely low for an invited-guest event, and the amplification worry
-- turned out to be contained: a like fans an SSE broadcast to ~100 clients, but the
-- export regeneration it could otherwise trigger is debounced (REGEN_DEBOUNCE 20s)
-- and superseded workers are inert. So this closes the gap for symmetry, not urgency,
-- and the ceiling is set high enough that no real guest will ever meet it -- a
-- double-tapping enthusiast at a wedding is not the thing being defended against.
INSERT INTO config (key, value) VALUES
('social_rate_per_min', '120'),
('social_rate_enabled', 'true')
ON CONFLICT (key) DO NOTHING;

View File

@@ -0,0 +1,13 @@
-- Restore the pre-021 definition (no ban/hide filtering) exactly as 004 created it.
DROP VIEW IF EXISTS v_hashtag_counts;
CREATE VIEW v_hashtag_counts AS
SELECT
h.event_id,
h.tag,
COUNT(uh.upload_id) AS upload_count
FROM hashtag h
JOIN upload_hashtag uh ON uh.hashtag_id = h.id
JOIN upload u ON u.id = uh.upload_id AND u.deleted_at IS NULL
GROUP BY h.event_id, h.id, h.tag
ORDER BY upload_count DESC;

View File

@@ -0,0 +1,31 @@
-- v_hashtag_counts: apply the same visibility rules as v_feed.
--
-- The chip row and the grid's tag picker are both fed by this view, but it has only ever
-- filtered `u.deleted_at IS NULL`. Migration 011 added ban/hide filtering to the feed and
-- never reached here, so the two disagreed about which uploads exist:
--
-- * A host bans a guest who posted 3 of the 12 `#tanz` photos. The chip keeps reading
-- "#tanz 12"; tapping it returns 9. The count is presented as authoritative and is not.
-- * A tag used ONLY by a banned or hidden guest stays in the chip row and in the tag
-- picker as a selectable option that leads to an empty feed — a ghost filter that
-- cannot be cleared because there is nothing wrong with it to see.
--
-- Bans are exactly the moment a host is watching these numbers to confirm the moderation
-- took effect, so a stale count reads as "the ban didn't work".
--
-- Same predicate as v_feed (see 016_display_derivative.up.sql), joined through `user`.
DROP VIEW IF EXISTS v_hashtag_counts;
CREATE VIEW v_hashtag_counts AS
SELECT
h.event_id,
h.tag,
COUNT(uh.upload_id) AS upload_count
FROM hashtag h
JOIN upload_hashtag uh ON uh.hashtag_id = h.id
JOIN upload u ON u.id = uh.upload_id AND u.deleted_at IS NULL
JOIN "user" usr ON usr.id = u.user_id
WHERE usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE
GROUP BY h.event_id, h.id, h.tag
ORDER BY upload_count DESC;

View File

@@ -0,0 +1,2 @@
DROP INDEX IF EXISTS upload_client_upload_id_key;
ALTER TABLE upload DROP COLUMN IF EXISTS client_upload_id;

View File

@@ -0,0 +1,25 @@
-- Idempotency key for uploads, supplied by the client.
--
-- The failure this closes is the ordinary one on a phone, not an exotic race: the server
-- receives the body, validates it, commits the row, and the response is lost on the way back
-- because the guest walked out of range or the AP dropped the connection. The client sees a
-- network error with the blob still in hand, marks the item retryable, and re-sends it — both
-- when the guest taps "Erneut" and automatically when the queue requeues on reconnect. Every
-- attempt minted a fresh `Uuid::new_v4()` server-side, so the same photo landed in the gallery
-- two or three times and was charged against the guest's storage quota each time.
--
-- The client already has a stable per-queue-item UUID, so it costs nothing to send. NULL is
-- allowed and unconstrained: uploads that predate this column, and any client that doesn't send
-- one, keep working exactly as before.
ALTER TABLE upload ADD COLUMN client_upload_id UUID;
-- Partial rather than a plain UNIQUE. Postgres would tolerate the NULLs either way, but indexing
-- only the rows that carry a key keeps it small and states the rule exactly: uniqueness applies
-- where a key exists, and nowhere else.
--
-- Scoped globally rather than per user or per event. The key is a client-generated v4 UUID, so a
-- collision between two different photos is not a real possibility, and a single-column index
-- means the uniqueness check cannot be wrong about which event or user a retry belongs to.
CREATE UNIQUE INDEX upload_client_upload_id_key
ON upload (client_upload_id)
WHERE client_upload_id IS NOT NULL;

View File

@@ -0,0 +1,3 @@
DROP INDEX IF EXISTS idx_upload_derivative_backfill;
ALTER TABLE upload DROP COLUMN IF EXISTS derivative_last_error;
ALTER TABLE upload DROP COLUMN IF EXISTS derivative_attempts;

View File

@@ -0,0 +1,23 @@
-- Bound how many times a permanently-failing upload can be re-processed.
--
-- Without this, one poisoned row is an outage. The upload row is committed BEFORE compression
-- starts, `derivatives_rev` defaults to 0, and `set_derivatives_rev` only runs on success — so
-- a row whose processing kills the container survives at rev 0, the unconditional startup
-- backfill re-selects it on the next boot, and `restart: unless-stopped` turns that into an
-- infinite kill loop. Every restart also drops every SSE stream and truncates every in-flight
-- upload. That was reachable via a single large PNG (see services/compression.rs), but the
-- shape is general: any input that can kill or hang the worker repeats forever.
--
-- The counter is incremented WRITE-AHEAD, before the work is attempted, because the failure
-- mode being defended against is a SIGKILL — no error is returned, no handler runs, no Drop
-- fires. A counter bumped in an error path increments zero times per crash and changes nothing.
ALTER TABLE upload ADD COLUMN IF NOT EXISTS derivative_attempts SMALLINT NOT NULL DEFAULT 0;
-- Last failure text, so a row that has given up can be diagnosed without reproducing it.
-- Nothing reads this in code; it exists for the operator.
ALTER TABLE upload ADD COLUMN IF NOT EXISTS derivative_last_error TEXT;
-- Serves the backfill selection, which now filters on both columns.
CREATE INDEX IF NOT EXISTS idx_upload_derivative_backfill
ON upload (derivatives_rev, derivative_attempts)
WHERE deleted_at IS NULL;

View File

@@ -0,0 +1,26 @@
-- Restore the 016 definition verbatim.
DROP VIEW IF EXISTS v_feed;
CREATE VIEW v_feed AS
SELECT
u.id,
u.event_id,
u.user_id,
usr.display_name AS uploader_name,
usr.is_banned,
usr.uploads_hidden,
u.preview_path,
u.thumbnail_path,
u.display_path,
u.mime_type,
u.caption,
u.created_at,
COUNT(DISTINCT l.user_id) AS like_count,
COUNT(DISTINCT c.id) AS comment_count
FROM upload u
JOIN "user" usr ON u.user_id = usr.id
LEFT JOIN "like" l ON l.upload_id = u.id
LEFT JOIN comment c ON c.upload_id = u.id AND c.deleted_at IS NULL
WHERE u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE
GROUP BY u.id, usr.display_name, usr.is_banned, usr.uploads_hidden;

View File

@@ -0,0 +1,55 @@
-- Make a feed page cost a page, not the whole event.
--
-- The previous definition (016) computed like_count/comment_count with LEFT JOINs and a
-- GROUP BY. Postgres CAN push the `event_id = $1` qual and the keyset predicate through the
-- view — verified with EXPLAIN, it uses idx_upload_event_created_id — but it CANNOT push
-- ORDER BY ... LIMIT across a GroupAggregate. So every feed request aggregated every upload in
-- the event (times its likes and comments) and only then sorted and took 21 rows. The cost
-- grew with the event, not with the page, and page 1 — the most expensive one — is exactly
-- what refreshFeedInPlace refetches on every completed upload, from every open feed in the
-- venue.
--
-- Correlated scalar subqueries move the counts ABOVE the Limit in the plan: they are evaluated
-- once per returned row, so 21 index lookups instead of a full aggregation.
--
-- The rewrite is EXACTLY equivalent, not merely close:
-- * "like" is keyed (upload_id, user_id), so COUNT(DISTINCT l.user_id) == count(*).
-- * comment.id is the primary key, so COUNT(DISTINCT c.id) == count(*).
-- * one row per upload either way — the GROUP BY was on u.id.
-- Column names, order and types are unchanged (count(*) and COUNT(DISTINCT ...) are both
-- bigint), so no Rust code changes.
--
-- No new index needed: idx_like_upload plus the (upload_id, user_id) PK serve the like
-- subquery, and idx_comment_upload ... WHERE deleted_at IS NULL matches the comment
-- subquery's predicate exactly.
--
-- One thing a future editor needs to know: the hashtag-filtered feed joins upload_hashtag
-- against this view. That was safe before only because the GROUP BY collapsed the join
-- fan-out; it is safe now because the view is one row per upload and upload_hashtag is keyed
-- (upload_id, hashtag_id) with a single tag filtered. Adding a second tag filter would need
-- fresh thought.
-- Not CASCADE: if something ever comes to depend on this view, the migration should fail
-- loudly rather than silently drop it.
DROP VIEW IF EXISTS v_feed;
CREATE VIEW v_feed AS
SELECT
u.id,
u.event_id,
u.user_id,
usr.display_name AS uploader_name,
usr.is_banned,
usr.uploads_hidden,
u.preview_path,
u.thumbnail_path,
u.display_path,
u.mime_type,
u.caption,
u.created_at,
(SELECT count(*) FROM "like" l WHERE l.upload_id = u.id) AS like_count,
(SELECT count(*) FROM comment c WHERE c.upload_id = u.id AND c.deleted_at IS NULL) AS comment_count
FROM upload u
JOIN "user" usr ON u.user_id = usr.id
WHERE u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE;

View File

@@ -0,0 +1,11 @@
-- NOTE: the reserved-name rename in the up migration is NOT reversible. The original names
-- are not recorded anywhere, and reversing it would in any case re-create the state that
-- bricked admin login. Rolling back the schema does not roll back that data change.
DELETE FROM config WHERE key IN (
'recover_name_rate_per_15min',
'pin_reset_ip_rate_per_min',
'upload_edit_rate_per_min',
'upload_edit_rate_enabled'
);
ALTER TABLE "user" DROP COLUMN IF EXISTS last_failed_pin_at;

View File

@@ -0,0 +1,70 @@
-- Two independent auth defects that share a migration because they share a table.
-- 1. RESERVED NAMES — free any guest squatting on a name the admin path used to depend on.
--
-- Migration 007 made display_name unique per event case-insensitively, and `join` had no
-- reserved-name guard. So any guest could join as "admin"/"Admin"/"ADMIN" before the operator's
-- first admin login; admin_login then looked its user up BY NAME, missed (wrong role), fell
-- through to creating "Admin", violated that unique index, and returned a 500 — permanently,
-- with no in-app recovery. Moderation, config and gallery release all gone, fixed only by SQL.
--
-- The real fix is in code (look the admin up by role, never by name — see auth/handlers.rs).
-- This clears the state an already-deployed database may be carrying.
--
-- RENAMED, NEVER DELETED: the guest keeps their uploads, their PIN and their session. Only
-- non-admin rows are touched — a real admin row named "Admin" is the expected state.
-- Two guards that are not optional, because this statement runs INSIDE the migration
-- transaction on boot and a failure here exits the process — `restart: unless-stopped` then
-- turns it into a crash loop with no in-app recovery. That is a strictly worse version of the
-- lockout this migration exists to clean up after.
--
-- * role = 'guest', not role <> 'admin'. The enum also has 'host' (001), and hosts are
-- promoted from guests at runtime — so <> 'admin' renamed a legitimately promoted staff
-- member whose name happens to be "Host".
-- * NOT EXISTS. The target name is derived, not unique: `idx_user_event_name_ci` (007) is a
-- UNIQUE index on (event_id, lower(display_name)), and nothing stopped a second guest from
-- having already joined as exactly "Admin (a1b2c3d4)" — the old code had no reserved-name
-- guard and the join response hands each guest their own id. Rare, but the cost of losing
-- that bet is the whole event.
--
-- A row that collides is simply left alone: `create_admin_user` already handles a name clash by
-- falling back to `Admin-<8hex>`, and `admin_login` no longer resolves by name at all, so this
-- cleanup is convenience rather than load-bearing.
UPDATE "user" u
SET display_name = u.display_name || ' (' || left(u.id::text, 8) || ')'
WHERE u.role = 'guest'
AND lower(u.display_name) IN ('admin', 'administrator', 'host', 'eventsnap')
AND NOT EXISTS (
SELECT 1 FROM "user" x
WHERE x.event_id = u.event_id
AND lower(x.display_name) = lower(u.display_name || ' (' || left(u.id::text, 8) || ')')
);
-- 2. PIN LOCKOUT DECAY.
--
-- failed_pin_attempts only ever cleared on a successful recovery or after a lockout expired, so
-- honest typos accumulated across days: a guest who fat-fingered their PIN twice last night
-- arrives today already two-thirds of the way to being locked out. With the threshold now
-- raised (see below) a decay window is what keeps that raise safe rather than merely lenient.
ALTER TABLE "user" ADD COLUMN IF NOT EXISTS last_failed_pin_at TIMESTAMPTZ;
-- Rate-limit knobs introduced with this release.
--
-- recover_name_rate_per_15min (4, was a hardcoded 5): the per-(IP, name) ceiling. It MUST stay
-- below the account-lock threshold, which is the whole defect — at 5-per-IP against a 3-strike
-- lock, three requests from one IP locked any guest whose name is visible on the feed, every 15
-- minutes, forever. The lock threshold moves to 12 in code, so locking a victim now needs at
-- least three distinct sources while an honest guest never comes close.
--
-- pin_reset_ip_rate_per_min (30): /recover/request was the one unauthenticated endpoint with no
-- per-IP ceiling at all — /join got one in 017 and /recover in 019, and this third one was
-- simply missed. Its per-name key is attacker-chosen, so cycling names minted a fresh bucket
-- every time and the per-IP cost was unbounded.
--
-- upload_edit_rate_per_min (30): PATCH /upload/{id} had no rate limit of any kind.
INSERT INTO config (key, value) VALUES
('recover_name_rate_per_15min', '4'),
('pin_reset_ip_rate_per_min', '30'),
('upload_edit_rate_per_min', '30'),
('upload_edit_rate_enabled', 'true')
ON CONFLICT (key) DO NOTHING;

View File

@@ -0,0 +1,13 @@
-- Restore migration 022's wider index (which also covered soft-deleted rows).
--
-- Note this can FAIL where the up-migration succeeded: once retries-after-delete have been
-- allowed, two rows may legitimately share a `client_upload_id` (one deleted, one live), and
-- the wider unique index cannot be rebuilt over them. That is inherent to reverting this
-- direction, not a defect in the down-migration. If it fails, the live-only index is still
-- correct and should simply be kept.
DROP INDEX IF EXISTS upload_client_upload_id_key;
CREATE UNIQUE INDEX upload_client_upload_id_key
ON upload (client_upload_id)
WHERE client_upload_id IS NOT NULL;

View File

@@ -0,0 +1,35 @@
-- Narrow the client-upload idempotency index so it stops covering soft-deleted rows.
--
-- The bug (H9). Migration 022 created the index partial on `client_upload_id IS NOT NULL`
-- only, so a soft-deleted row kept occupying its key. But `find_by_client_upload_id` filters
-- `deleted_at IS NULL` — deliberately, and its doc comment says so: "if the guest deleted the
-- photo and their queue later retries, they should get a fresh upload rather than a
-- resurrection of a deleted one." The index and the lookup therefore disagreed, and the
-- disagreement is reachable by an ordinary guest:
--
-- 1. guest uploads a photo, then deletes it (soft delete — the row stays, `deleted_at` set)
-- 2. their queue retries the same item (reconnect requeue, or they tap "Erneut")
-- 3. the whole body is re-streamed and re-validated, then `ON CONFLICT DO NOTHING` matches
-- the DEAD row and inserts nothing
-- 4. the replay lookup filters that row out and finds nothing, so the handler returns 409
-- 5. the client classifies 409 as terminal and DELETES the blob from IndexedDB
--
-- The photo is now gone from the device with no row in the gallery, and there is no path back.
-- Re-selecting the same file from the camera roll mints a new `client_upload_id`, so that does
-- work — but the guest has no way to know that is what is required.
--
-- Adding `deleted_at IS NULL` makes the index agree with the lookup: a key is claimed only
-- while a LIVE row holds it, so step 3 inserts a fresh row and the retry succeeds.
--
-- Uniqueness among live rows is what the feature actually needs. The property migration 022
-- was protecting — "the same photo must not land in the gallery twice" — is about rows the
-- guest can see, and a soft-deleted row is not one of those.
DROP INDEX IF EXISTS upload_client_upload_id_key;
-- CONCURRENTLY is deliberately NOT used: sqlx runs each migration inside a transaction, and
-- CREATE INDEX CONCURRENTLY cannot run in one. The table is small (one event's uploads) and
-- this runs at boot before the server accepts requests, so the brief lock costs nothing.
CREATE UNIQUE INDEX upload_client_upload_id_key
ON upload (client_upload_id)
WHERE client_upload_id IS NOT NULL AND deleted_at IS NULL;

View File

@@ -0,0 +1,2 @@
DROP INDEX IF EXISTS user_client_join_id_key;
ALTER TABLE "user" DROP COLUMN IF EXISTS client_join_id;

View File

@@ -0,0 +1,34 @@
-- Idempotency key for /join, supplied by the client.
--
-- The failure this closes (H16) is the single most likely failure of the evening, on step one
-- of the product. `/join` commits the user row AND the bcrypt hash of the PIN, but the PLAINTEXT
-- PIN exists nowhere except the HTTP response body. So:
--
-- 1. guest scans the QR in the venue car park, taps "Beitreten"
-- 2. the server creates the account and hashes the PIN
-- 3. the response is lost on the way back — the 5G-to-nothing transition every wedding venue
-- has, or the AP handing off
-- 4. the client retries; the name is now taken, so it 409s
-- 5. the client shows a PIN entry form for a PIN THAT WAS NEVER DISPLAYED
--
-- The guest is locked out of their own brand-new account, and the only recovery is finding a
-- host with a dashboard open. `/upload` already solved exactly this with `client_upload_id`;
-- join never got the same treatment.
--
-- With a key, a retry is recognised as the same join and answered with a usable PIN. We do NOT
-- store the plaintext to replay it — see the handler: a retry ROTATES the PIN. That is sound
-- precisely because the original was never shown to anybody, so there is nothing to preserve,
-- and it keeps this table free of recoverable credentials.
ALTER TABLE "user" ADD COLUMN client_join_id UUID;
-- Partial, for the same reasons as `upload_client_upload_id_key`: index only the rows that
-- carry a key, and state the rule exactly. NULL is allowed and unconstrained, so any client
-- that does not send one (and every row that predates this column) behaves exactly as before.
--
-- Scoped per event as well as per key. The key is a client-generated v4 UUID so a cross-event
-- collision is not realistic, but a reused install genuinely has two events in one table and
-- "this join belongs to that event" is the property we actually mean.
CREATE UNIQUE INDEX user_client_join_id_key
ON "user" (event_id, client_join_id)
WHERE client_join_id IS NOT NULL;

View File

@@ -0,0 +1,22 @@
-- Restore migration 024's counts (which included banned users' likes and comments).
CREATE OR REPLACE VIEW v_feed AS
SELECT
u.id,
u.event_id,
u.user_id,
usr.display_name AS uploader_name,
usr.is_banned,
usr.uploads_hidden,
u.preview_path,
u.thumbnail_path,
u.display_path,
u.mime_type,
u.caption,
u.created_at,
(SELECT count(*) FROM "like" l WHERE l.upload_id = u.id) AS like_count,
(SELECT count(*) FROM comment c WHERE c.upload_id = u.id AND c.deleted_at IS NULL) AS comment_count
FROM upload u
JOIN "user" usr ON u.user_id = usr.id
WHERE u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE;

View File

@@ -0,0 +1,44 @@
-- Exclude banned users' likes and comments from the feed's scalar counts (H11).
--
-- `v_feed` already excludes banned UPLOADERS (`usr.is_banned = FALSE` on the join), but the two
-- correlated subqueries added by migration 024 counted every like and every non-deleted comment
-- regardless of who wrote it. So after a ban:
--
-- * the banned guest's own photos disappear from the feed (correct), but
-- * their likes still inflate the counter on everyone else's photos, and
-- * their comments still contribute to `comment_count` — and, until the change to
-- `Comment::list_for_upload` that ships with this migration, were still RENDERED in the
-- lightbox on the most-viewed photo of the evening.
--
-- The host's mental model of "ban" is "this person's contributions are gone". Photos honoured it;
-- likes and comments did not. Migration 021 already applied exactly this reasoning to hashtag
-- counts, and the export query filters `is_banned` too — this brings the last read path in line.
--
-- Derived at read time, so `unban_user` restores the counts with no extra work, exactly as it
-- already restores the photos.
CREATE OR REPLACE VIEW v_feed AS
SELECT
u.id,
u.event_id,
u.user_id,
usr.display_name AS uploader_name,
usr.is_banned,
usr.uploads_hidden,
u.preview_path,
u.thumbnail_path,
u.display_path,
u.mime_type,
u.caption,
u.created_at,
(SELECT count(*) FROM "like" l
JOIN "user" lu ON lu.id = l.user_id
WHERE l.upload_id = u.id AND NOT lu.is_banned) AS like_count,
(SELECT count(*) FROM comment c
JOIN "user" cu ON cu.id = c.user_id
WHERE c.upload_id = u.id AND c.deleted_at IS NULL AND NOT cu.is_banned) AS comment_count
FROM upload u
JOIN "user" usr ON u.user_id = usr.id
WHERE u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE
AND usr.is_banned = FALSE;

View File

@@ -0,0 +1,2 @@
DROP INDEX IF EXISTS host_action_audit_event_created_idx;
DROP TABLE IF EXISTS host_action_audit;

View File

@@ -0,0 +1,41 @@
-- An audit trail for privileged actions (H17).
--
-- What existed before: nothing. `grep -i audit` across `handlers/host.rs` and `handlers/admin.rs`
-- returned no hits. Individual actions logged a `tracing::info!` line, but config changes, gallery
-- release and event lock/unlock logged nothing at all — and the "audit trail" as a whole was a
-- 30 MB rotating Docker log that the runbook's own retention settings will discard.
--
-- Why it matters here specifically: a host is a promoted GUEST, and `reset_pin` overwrites another
-- guest's credential and returns the new PIN in the clear. So a host can take over any guest's
-- account and post as them, and nothing in the record showed it happened (only /recover FAILURES
-- were logged). At a wedding the people involved know each other; the point is not catching a
-- villain, it is being able to answer "what happened to my photo?" the next morning without
-- guessing.
--
-- Deliberately append-only in practice: no UPDATE or DELETE path is written for it anywhere. Small
-- (a few hundred rows for a real event), so no partitioning or retention job.
CREATE TABLE host_action_audit (
id BIGSERIAL PRIMARY KEY,
event_id UUID NOT NULL REFERENCES event(id) ON DELETE CASCADE,
-- The privileged caller. NOT a FK with ON DELETE CASCADE: the record must survive the actor's
-- account being removed, which is exactly when it is most likely to be wanted.
actor_id UUID,
actor_name TEXT,
actor_role TEXT NOT NULL,
-- Short stable slug: 'ban_user', 'unban_user', 'reset_pin', 'delete_upload',
-- 'delete_comment', 'release_gallery', 'lock_uploads', 'unlock_uploads', 'patch_config',
-- 'promote_user', 'demote_user', 'delete_user'.
action TEXT NOT NULL,
-- The guest or object acted upon, when there is one.
target_id UUID,
target_name TEXT,
-- Free-form context: the config key and its old/new value, the caption that was removed, etc.
-- Never credentials — a reset PIN must not be recoverable from this table.
detail JSONB,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
-- The only query shape this needs: "what happened at this event, newest first".
CREATE INDEX host_action_audit_event_created_idx
ON host_action_audit (event_id, created_at DESC);

View File

@@ -0,0 +1,3 @@
-- Revert the join ceiling to 60/min for installs still on the raised default
-- (preserves any explicit admin override at another value).
UPDATE config SET value = '60' WHERE key = 'join_ip_rate_per_min' AND value = '300';

View File

@@ -0,0 +1,15 @@
-- Raise the per-IP join ceiling from 60/min to 300/min.
--
-- Rationale: every guest at the venue arrives through one NAT'd public address,
-- so `join_ip:{ip}` is not a per-guest limit at all — it is a ceiling on the
-- whole party. The QR code goes up once and is scanned in a burst: at 60/min,
-- guest 61 onwards is refused on the join screen, which is the one screen with
-- no auto-retry, and every manual retry spends another slot.
--
-- The code default was already raised to 300 (auth/handlers.rs), but a default
-- only applies when the key is ABSENT, and migration 017 seeds it. Without this
-- UPDATE the raise is dead code on every existing install.
--
-- Only bump installs still on the seeded default; an admin who deliberately set
-- a different value keeps it (migration 017 seeded 60; this UPDATE is scoped to '60').
UPDATE config SET value = '300' WHERE key = 'join_ip_rate_per_min' AND value = '60';

View File

@@ -0,0 +1,10 @@
-- Restore migration 026's predicate, then drop the column it depended on.
--
-- Note the same pairing caveat 026's own down carries: this is only valid alongside a code
-- rollback. `Upload::create` sends an ON CONFLICT predicate that must match the live index, so
-- running this down against the current binary makes every keyed upload a runtime 500.
DROP INDEX IF EXISTS upload_client_upload_id_key;
CREATE UNIQUE INDEX upload_client_upload_id_key ON upload (client_upload_id)
WHERE client_upload_id IS NOT NULL AND deleted_at IS NULL;
ALTER TABLE upload DROP COLUMN IF EXISTS taken_down_by_host;

View File

@@ -0,0 +1,28 @@
-- Keep a client upload key CLAIMED when the deletion was a host takedown.
--
-- Migration 026 narrowed `upload_client_upload_id_key` to live rows so that a guest who deletes
-- their own photo and whose queue later retries gets a fresh upload instead of a permanent 409.
-- That rationale reasoned only about the GUEST deleting. `deleted_at` is also set by
-- `host_delete_upload`, and for that case the same rule undoes a moderation decision:
--
-- 1. Guest uploads. The row commits and the photo appears in the feed, but the response is lost
-- on the way back (the flaky-wifi case this whole feature exists for), so the phone keeps the
-- queue item.
-- 2. The host sees the photo and takes it down. `deleted_at` is stamped, the keepsake epoch is
-- bumped, and the archive is rebuilt without it.
-- 3. Ten minutes later the phone reconnects and retries. The key is no longer claimed, the
-- INSERT succeeds, and the photo is BACK — in the feed, in the next keepsake, under a NEW
-- uuid that matches nothing in the host's moderation history, with nothing logged to say a
-- takedown was undone.
--
-- So the key stays claimed for a host takedown and is released only for a guest's own delete. The
-- retry then resolves to the duplicate path and is refused, which is the correct answer: the photo
-- was deliberately removed, and re-sending the bytes must not bring it back.
ALTER TABLE upload ADD COLUMN taken_down_by_host BOOLEAN NOT NULL DEFAULT FALSE;
-- KEEP THE PREDICATE IN LOCKSTEP WITH `Upload::create`'s ON CONFLICT clause (models/upload.rs).
-- A drift between the two is not a compile error here — queries are checked at runtime — it is a
-- 500 on every upload that carries a key, i.e. on exactly the retries this index exists to serve.
DROP INDEX IF EXISTS upload_client_upload_id_key;
CREATE UNIQUE INDEX upload_client_upload_id_key ON upload (client_upload_id)
WHERE client_upload_id IS NOT NULL AND (deleted_at IS NULL OR taken_down_by_host);

File diff suppressed because it is too large Load Diff

View File

@@ -17,24 +17,92 @@ fn looks_placeholder(s: &str) -> bool {
|| lower.contains("placeholder") || lower.contains("placeholder")
} }
/// A bcrypt hash is exactly 60 characters and opens with `$2<variant>$<cost>$`.
///
/// Checking the SHAPE, not just placeholder-ness, is what catches a hash silently mangled in
/// transit. The `$` segments are variable-expansion bait for both shell quoting and Docker
/// Compose's `env_file` parsing, and a mangled hash is not a placeholder — so without this it
/// passes every other guard here, the app boots green, `/health` returns `ok`, and every admin
/// login 401s.
///
/// That failure is unrecoverable mid-event, which is why it is worth a hard fail at boot: the
/// Admin row is created BY a successful admin login (`auth/handlers.rs`), and only an Admin or
/// Host can promote a Host. No admin login therefore means no host at all — the event cannot be
/// closed, the gallery cannot be released, and nothing can be moderated.
fn looks_bcrypt(s: &str) -> bool {
let b = s.as_bytes();
s.len() == 60
&& b[0] == b'$'
&& b[1] == b'2'
&& matches!(b[2], b'a' | b'b' | b'x' | b'y')
&& b[3] == b'$'
&& b[4].is_ascii_digit()
&& b[5].is_ascii_digit()
&& b[6] == b'$'
}
/// Enforce secret hygiene. In production every guard is hard-fail: a booting app /// Enforce secret hygiene. In production every guard is hard-fail: a booting app
/// with a publicly-known signing key is worse than one that refuses to start. /// with a publicly-known signing key is worse than one that refuses to start.
/// Outside production the dev sentinel is tolerated (warned) so local dev is frictionless. /// Outside production the dev sentinel is tolerated (warned) so local dev is frictionless.
fn validate_secrets(is_prod: bool, jwt_secret: &str, admin_password_hash: &str) -> Result<()> { ///
/// EVERY failure is collected and reported together. Returning on the first one made fixing two
/// secrets cost two boot cycles — the operator rotates JWT_SECRET, restarts, and only then learns
/// about ADMIN_PASSWORD_HASH. Restarting this stack is not free (Caddy waits on the unhealthy app),
/// and each avoidable cycle is another chance to reach for `down -v`.
fn validate_secrets(
is_prod: bool,
jwt_secret: &str,
admin_password_hash: &str,
database_url: &str,
) -> Result<()> {
if is_prod { if is_prod {
let mut problems: Vec<&str> = Vec::new();
if looks_placeholder(jwt_secret) { if looks_placeholder(jwt_secret) {
return Err(anyhow!( problems.push(
"Refusing to start in production with a placeholder JWT_SECRET — \ "JWT_SECRET is still the .env.example placeholder — rotate it \
rotate it (openssl rand -hex 64)." (openssl rand -hex 64).",
)); );
} } else if jwt_secret.len() < 32 {
if jwt_secret.len() < 32 { problems.push("JWT_SECRET must be at least 32 characters.");
return Err(anyhow!("JWT_SECRET must be at least 32 characters."));
} }
if admin_password_hash.is_empty() || looks_placeholder(admin_password_hash) { if admin_password_hash.is_empty() || looks_placeholder(admin_password_hash) {
problems.push(
"ADMIN_PASSWORD_HASH is unset or still the .env.example placeholder — generate one \
(docker run --rm caddy:2-alpine caddy hash-password --plaintext '<password>').",
);
} else if !looks_bcrypt(admin_password_hash) {
problems.push(
"ADMIN_PASSWORD_HASH is not a well-formed bcrypt hash: expected exactly 60 \
characters starting `$2b$12$…`. First look at the value that actually reached \
the app — `docker compose exec app printenv ADMIN_PASSWORD_HASH` — and compare \
it to .env character for character. In .env, SINGLE-QUOTE the hash \
('$2b$12$…'): Compose uses single-quoted env_file values literally, so the `$` \
segments survive. Double them to `$$` ONLY when setting the value under \
`environment:` in docker-compose.yml — doing that in .env corrupts a hash that \
would otherwise have worked. Regenerate with: \
docker run --rm caddy:2-alpine caddy hash-password --plaintext '<password>'",
);
}
// The DATABASE_URL carries the Postgres password, so a placeholder here means the stack is
// running on `CHANGE_ME_use_a_strong_password` — a credential published in the repo. The
// app used to boot green on it, because this guard only ever covered the two secrets it
// was written for and nothing else looked at POSTGRES_PASSWORD at all.
//
// Read POSTGRES_PASSWORD's docs before changing this: it is applied ONLY at initdb, so the
// remedy is not "edit .env and restart" — see the 28P01 diagnostic in db.rs.
if looks_placeholder(database_url) {
problems.push(
"DATABASE_URL still carries the .env.example placeholder password — set a strong \
one (openssl rand -hex 24) in BOTH DATABASE_URL and POSTGRES_PASSWORD.",
);
}
if !problems.is_empty() {
return Err(anyhow!( return Err(anyhow!(
"Refusing to start in production without a real ADMIN_PASSWORD_HASH — \ "Refusing to start in production — {} secret(s) still unset or placeholder:\n - {}\n\
generate one (htpasswd -bnBC 12 '' <password> | tr -d ':\\n')." ALL secrets must be set BEFORE the first `docker compose up -d`: Postgres bakes \
POSTGRES_PASSWORD into its data directory on first boot and ignores later changes.",
problems.len(),
problems.join("\n - ")
)); ));
} }
} else if jwt_secret == DEV_JWT_SECRET_SENTINEL { } else if jwt_secret == DEV_JWT_SECRET_SENTINEL {
@@ -63,6 +131,66 @@ pub struct AppConfig {
pub app_port: u16, pub app_port: u16,
/// Number of concurrent media compression workers (read once at boot). /// Number of concurrent media compression workers (read once at boot).
pub compression_concurrency: usize, pub compression_concurrency: usize,
/// Master switch for the comment feature (env `COMMENTS_ENABLED`, default true).
/// When false the backend rejects new comments and the frontend hides the whole
/// comment UI. Existing comments stay in the DB (hidden), so flipping it back
/// restores them. Boot-time immutable, like `compression_concurrency`.
pub comments_enabled: bool,
/// Default colour theme, used as the fallback when the DB config keys are unset.
/// Runtime overrides live in the `config` table (admin UI); these env vars only
/// seed the initial default. `preset` is an id the frontend knows (e.g.
/// "champagne-gold", "rose", … or "custom"); the two seeds are `#rrggbb` brand +
/// accent colours the whole palette is derived from.
pub default_theme_preset: String,
pub default_theme_primary: String,
pub default_theme_accent: String,
}
/// The shipped default brand/accent seed (champagne gold — matches the hand-tuned
/// ramp in tailwind-theme.css). Kept here so an unset env still yields the current look.
const DEFAULT_THEME_SEED: &str = "#8a6a2b";
/// Upper bound on `SESSION_EXPIRY_DAYS`. ~10 years — absurdly generous for a one-evening event,
/// and low enough that `chrono::Duration::days` cannot overflow downstream.
const MAX_SESSION_EXPIRY_DAYS: i64 = 3650;
/// Parse and RANGE-CHECK `SESSION_EXPIRY_DAYS`. Refusing to boot is the whole point.
///
/// This was `.parse().context(...)` with no bounds, and both ends of the range were live faults
/// that a green health check hid completely (H7):
///
/// * A huge value made `chrono::Duration::days` PANIC on every `/join`, `/recover` and
/// `/admin/login`. There is no `CatchPanicLayer`, so the client got a connection reset with no
/// HTTP response at all — the app was up, healthy, and unable to authenticate anybody.
/// * Zero or negative created every session already-expired: `/join` returns 201 with a token,
/// and then every authenticated request 401s. A guest joins successfully and the app
/// immediately behaves as though they never did.
///
/// Both booted green because `/health` only probes the database. A bad value must stop the
/// container instead, where the operator sees it.
fn parse_session_expiry_days(raw: Option<&str>) -> Result<i64> {
let Some(raw) = raw else { return Ok(30) };
let trimmed = raw.trim();
if trimmed.is_empty() {
return Ok(30);
}
let days: i64 = trimmed
.parse()
.with_context(|| format!("SESSION_EXPIRY_DAYS must be a whole number (got {trimmed:?})"))?;
if days < 1 {
return Err(anyhow!(
"SESSION_EXPIRY_DAYS must be at least 1 (got {days}). Zero or negative makes every \
session expire the moment it is created: /join succeeds and every request after it \
returns 401."
));
}
if days > MAX_SESSION_EXPIRY_DAYS {
return Err(anyhow!(
"SESSION_EXPIRY_DAYS must be at most {MAX_SESSION_EXPIRY_DAYS} (got {days}). Larger \
values overflow the token-expiry arithmetic and panic on every auth request."
));
}
Ok(days)
} }
impl AppConfig { impl AppConfig {
@@ -72,16 +200,16 @@ impl AppConfig {
let jwt_secret = std::env::var("JWT_SECRET").context("JWT_SECRET must be set")?; let jwt_secret = std::env::var("JWT_SECRET").context("JWT_SECRET must be set")?;
let admin_password_hash = std::env::var("ADMIN_PASSWORD_HASH").unwrap_or_default(); let admin_password_hash = std::env::var("ADMIN_PASSWORD_HASH").unwrap_or_default();
let database_url = std::env::var("DATABASE_URL").context("DATABASE_URL must be set")?;
validate_secrets(is_prod, &jwt_secret, &admin_password_hash)?; validate_secrets(is_prod, &jwt_secret, &admin_password_hash, &database_url)?;
Ok(Self { Ok(Self {
database_url: std::env::var("DATABASE_URL").context("DATABASE_URL must be set")?, database_url,
jwt_secret, jwt_secret,
session_expiry_days: std::env::var("SESSION_EXPIRY_DAYS") session_expiry_days: parse_session_expiry_days(
.unwrap_or_else(|_| "30".to_string()) std::env::var("SESSION_EXPIRY_DAYS").ok().as_deref(),
.parse() )?,
.context("SESSION_EXPIRY_DAYS must be a number")?,
admin_password_hash, admin_password_hash,
event_name: std::env::var("EVENT_NAME").unwrap_or_else(|_| "EventSnap".to_string()), event_name: std::env::var("EVENT_NAME").unwrap_or_else(|_| "EventSnap".to_string()),
event_slug: std::env::var("EVENT_SLUG").context("EVENT_SLUG must be set")?, event_slug: std::env::var("EVENT_SLUG").context("EVENT_SLUG must be set")?,
@@ -100,6 +228,20 @@ impl AppConfig {
.and_then(|v| v.parse().ok()) .and_then(|v| v.parse().ok())
.filter(|&n| n >= 1) .filter(|&n| n >= 1)
.unwrap_or(2), .unwrap_or(2),
comments_enabled: std::env::var("COMMENTS_ENABLED")
.map(|v| {
!matches!(
v.trim().to_ascii_lowercase().as_str(),
"false" | "0" | "no" | "off"
)
})
.unwrap_or(true),
default_theme_preset: std::env::var("THEME_PRESET")
.unwrap_or_else(|_| "champagne-gold".to_string()),
default_theme_primary: std::env::var("THEME_PRIMARY")
.unwrap_or_else(|_| DEFAULT_THEME_SEED.to_string()),
default_theme_accent: std::env::var("THEME_ACCENT")
.unwrap_or_else(|_| DEFAULT_THEME_SEED.to_string()),
}) })
} }
} }
@@ -109,13 +251,22 @@ mod tests {
use super::*; use super::*;
const REAL_SECRET: &str = "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2"; const REAL_SECRET: &str = "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2";
const REAL_HASH: &str = "$2y$12$abcdefghijklmnopqrstuv.wxyzABCDEFGHIJKLMNOPQRSTUVWXYZ012"; // A structurally valid bcrypt hash: exactly 60 chars, `$2y$12$` + 53 of salt/digest.
// The shape matters — `looks_bcrypt` enforces it, so a fixture of the wrong length
// would assert the opposite of what these tests claim.
const REAL_HASH: &str = "$2y$12$abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXY01";
const REAL_DB_URL: &str = "postgres://eventsnap:7f3a9c1e5b2d8a4f@db:5432/eventsnap";
#[test] #[test]
fn prod_rejects_shipped_placeholder_secret() { fn prod_rejects_shipped_placeholder_secret() {
// The exact string shipped in `.env` — >32 chars, so it must be caught by // The exact string shipped in `.env` — >32 chars, so it must be caught by
// the substring guard, not the length check. // the substring guard, not the length check.
let err = validate_secrets(true, "change_me_to_a_random_64_byte_hex_string", REAL_HASH); let err = validate_secrets(
true,
"change_me_to_a_random_64_byte_hex_string",
REAL_HASH,
REAL_DB_URL,
);
assert!( assert!(
err.is_err(), err.is_err(),
"placeholder JWT_SECRET must be rejected in prod" "placeholder JWT_SECRET must be rejected in prod"
@@ -124,29 +275,133 @@ mod tests {
#[test] #[test]
fn prod_rejects_dev_sentinel_and_short_secret() { fn prod_rejects_dev_sentinel_and_short_secret() {
assert!(validate_secrets(true, DEV_JWT_SECRET_SENTINEL, REAL_HASH).is_err()); assert!(validate_secrets(true, DEV_JWT_SECRET_SENTINEL, REAL_HASH, REAL_DB_URL).is_err());
assert!(validate_secrets(true, "tooshort", REAL_HASH).is_err()); assert!(validate_secrets(true, "tooshort", REAL_HASH, REAL_DB_URL).is_err());
}
/// The failure this guards is silent and unrecoverable mid-event: a hash whose `$`
/// segments were eaten by shell or Compose interpolation is NOT a placeholder, so every
/// other guard passes, the app boots green and `/health` reports ok — and then every
/// admin login 401s, which (because the Admin row is created by a successful login, and
/// only an Admin/Host can promote a Host) means no host exists for the whole event.
#[test]
fn prod_rejects_mangled_admin_hash() {
// What `$2y$12$…` degrades to once `$2y`/`$12` are read as unset variables.
assert!(validate_secrets(true, REAL_SECRET, "abcdefghijklmnop", REAL_DB_URL).is_err());
// Right prefix, truncated body — still not a usable hash.
assert!(validate_secrets(true, REAL_SECRET, "$2y$12$tooshort", REAL_DB_URL).is_err());
// Correct length but no bcrypt prefix at all.
assert!(validate_secrets(true, REAL_SECRET, &"x".repeat(60), REAL_DB_URL).is_err());
// All three shipped bcrypt variants stay acceptable.
let body = "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXY01";
for variant in ["2a", "2b", "2y"] {
let hash = format!("${variant}$12${body}");
assert_eq!(hash.len(), 60);
assert!(
validate_secrets(true, REAL_SECRET, &hash, REAL_DB_URL).is_ok(),
"bcrypt variant ${variant}$ must be accepted"
);
}
} }
#[test] #[test]
fn prod_rejects_missing_or_placeholder_admin_hash() { fn prod_rejects_missing_or_placeholder_admin_hash() {
assert!(validate_secrets(true, REAL_SECRET, "").is_err()); assert!(validate_secrets(true, REAL_SECRET, "", REAL_DB_URL).is_err());
assert!(validate_secrets(true, REAL_SECRET, "$2y$12$placeholder_replace_me").is_err()); assert!(
validate_secrets(
true,
REAL_SECRET,
"$2y$12$placeholder_replace_me",
REAL_DB_URL
)
.is_err()
);
}
/// The stack used to come up GREEN on the database password published in the repo: this guard
/// covered the two secrets it was written for, and nothing anywhere looked at the Postgres
/// credential. README step 2 doesn't name POSTGRES_PASSWORD either, so following the
/// documented procedure verbatim shipped it.
#[test]
fn prod_rejects_the_shipped_placeholder_database_password() {
let shipped = "postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap";
let err = validate_secrets(true, REAL_SECRET, REAL_HASH, shipped).unwrap_err();
assert!(
err.to_string().contains("DATABASE_URL"),
"the refusal must name DATABASE_URL, not just fail: {err}"
);
// And it must point at the initdb trap, or the operator edits .env, restarts, and lands
// in a permanent auth-failure loop instead.
assert!(
err.to_string().contains("POSTGRES_PASSWORD"),
"the refusal must name POSTGRES_PASSWORD as the other half: {err}"
);
}
/// Every problem in ONE message. Reporting them one per boot made fixing two secrets cost two
/// restart cycles, on a stack where Caddy waits on the unhealthy app the whole time.
#[test]
fn prod_reports_every_placeholder_at_once() {
let err = validate_secrets(
true,
"change_me_to_a_random_64_byte_hex_string",
"$2y$12$placeholder_replace_me",
"postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap",
)
.unwrap_err()
.to_string();
for expected in ["JWT_SECRET", "ADMIN_PASSWORD_HASH", "DATABASE_URL"] {
assert!(err.contains(expected), "{expected} missing from: {err}");
}
assert!(
err.contains("3 secret(s)"),
"the count must match what is listed: {err}"
);
} }
#[test] #[test]
fn prod_accepts_real_secrets() { fn prod_accepts_real_secrets() {
assert!(validate_secrets(true, REAL_SECRET, REAL_HASH).is_ok()); assert!(validate_secrets(true, REAL_SECRET, REAL_HASH, REAL_DB_URL).is_ok());
}
/// A real password that happens to contain no placeholder substring must pass — including one
/// with URL-ish punctuation, so the guard can't be mistaken for a URL validator.
#[test]
fn prod_accepts_a_real_database_url_with_awkward_punctuation() {
assert!(
validate_secrets(
true,
REAL_SECRET,
REAL_HASH,
"postgres://eventsnap:aB3%24xY9-_.qW@db:5432/eventsnap"
)
.is_ok()
);
}
/// The e2e stack runs without APP_ENV=production, so none of this applies there — but assert
/// it, because a guard that tripped in e2e would be found the hard way.
#[test]
fn non_prod_ignores_a_placeholder_database_url() {
assert!(
validate_secrets(
false,
REAL_SECRET,
"",
"postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap"
)
.is_ok()
);
} }
#[test] #[test]
fn non_prod_tolerates_dev_sentinel() { fn non_prod_tolerates_dev_sentinel() {
assert!(validate_secrets(false, DEV_JWT_SECRET_SENTINEL, "").is_ok()); assert!(validate_secrets(false, DEV_JWT_SECRET_SENTINEL, "", REAL_DB_URL).is_ok());
} }
#[test] #[test]
fn non_prod_still_rejects_short_non_sentinel_secret() { fn non_prod_still_rejects_short_non_sentinel_secret() {
assert!(validate_secrets(false, "tooshort", "").is_err()); assert!(validate_secrets(false, "tooshort", "", REAL_DB_URL).is_err());
} }
#[test] #[test]
@@ -154,9 +409,23 @@ mod tests {
// looks_placeholder lowercases before matching — an upper/mixed-case // looks_placeholder lowercases before matching — an upper/mixed-case
// placeholder must still be rejected in prod. // placeholder must still be rejected in prod.
assert!( assert!(
validate_secrets(true, "CHANGE_ME_TO_A_RANDOM_64_BYTE_HEX_STRING", REAL_HASH).is_err() validate_secrets(
true,
"CHANGE_ME_TO_A_RANDOM_64_BYTE_HEX_STRING",
REAL_HASH,
REAL_DB_URL
)
.is_err()
);
assert!(
validate_secrets(
true,
REAL_SECRET,
"$2Y$12$PLACEHOLDER_replace_me",
REAL_DB_URL
)
.is_err()
); );
assert!(validate_secrets(true, REAL_SECRET, "$2Y$12$PLACEHOLDER_replace_me").is_err());
} }
#[test] #[test]
@@ -166,7 +435,7 @@ mod tests {
const LEN_31: &str = "abcdefghijklmnopqrstuvwxyz01234"; const LEN_31: &str = "abcdefghijklmnopqrstuvwxyz01234";
assert_eq!(LEN_32.len(), 32); assert_eq!(LEN_32.len(), 32);
assert_eq!(LEN_31.len(), 31); assert_eq!(LEN_31.len(), 31);
assert!(validate_secrets(true, LEN_32, REAL_HASH).is_ok()); assert!(validate_secrets(true, LEN_32, REAL_HASH, REAL_DB_URL).is_ok());
assert!(validate_secrets(true, LEN_31, REAL_HASH).is_err()); assert!(validate_secrets(true, LEN_31, REAL_HASH, REAL_DB_URL).is_err());
} }
} }

View File

@@ -2,24 +2,131 @@ use anyhow::{Context, Result};
use sqlx::PgPool; use sqlx::PgPool;
use sqlx::postgres::PgPoolOptions; use sqlx::postgres::PgPoolOptions;
const DEFAULT_MAX_CONNECTIONS: u32 = 10; /// Keep in step with `.env.example` and the `db` sizing comment in `docker-compose.yml`.
/// These three drifted apart once (code 10 / `.env.example` 15 / runbook 30) and the runbook
/// presented its number as authoritative, so the contradiction was invisible at deploy time.
/// 15 is sized to 2 vCPU and the 1G `db` memory limit — raise it only alongside both.
const DEFAULT_MAX_CONNECTIONS: u32 = 15;
/// SQLSTATE for `invalid_password`.
const PG_INVALID_PASSWORD: &str = "28P01";
/// Turn the one connect failure with an unguessable cause into a self-explaining one.
///
/// `POSTGRES_PASSWORD` is honoured ONLY when Postgres initialises its data directory. Change it in
/// `.env` afterwards and the app authenticates with the new password against a volume that still
/// holds the old one — a permanent restart loop whose only symptom is
/// `password authentication failed`.
///
/// The production secret guard makes that sequence NEARLY CERTAIN rather than rare: it stops the
/// app on the first `docker compose up -d`, but not the `db` service in that same command, which
/// initialises and bakes in whatever password was in `.env` at that moment. So the intended
/// recovery — see the refusal, fix your secrets, boot again — is exactly the sequence that breaks
/// it. Nothing in the error names the cause, and the remedy destroys data, so it is the last thing
/// an operator should guess at.
fn explain_auth_failure(err: &sqlx::Error) {
let is_auth_failure = match err {
sqlx::Error::Database(db) => db.code().as_deref() == Some(PG_INVALID_PASSWORD),
_ => false,
};
if !is_auth_failure {
return;
}
tracing::error!(
"Postgres rejected the credentials in DATABASE_URL (SQLSTATE {PG_INVALID_PASSWORD}).\n\
\n\
This almost always means POSTGRES_PASSWORD was changed AFTER the database volume was \
first created. Postgres applies that variable only when it initialises its data \
directory; editing .env and restarting does not change the stored password, so the two \
drift apart permanently.\n\
\n\
If the event has NOT started and you have no data worth keeping:\n\n \
docker compose down -v && docker compose up -d\n\n\
(-v DELETES the database, the uploaded media and the exports. There is no undo.)\n\
\n\
If you DO have data: restore the old password into DATABASE_URL instead, or change the \
stored one with ALTER ROLE inside the running db container. Never reach for -v to fix a \
login problem on a live event."
);
}
pub async fn create_pool(database_url: &str) -> Result<PgPool> { pub async fn create_pool(database_url: &str) -> Result<PgPool> {
let max_connections = std::env::var("DATABASE_MAX_CONNECTIONS") // A malformed value must not silently become the default: an operator who typed
.ok() // `DATABASE_MAX_CONNECTIONS=3O` (letter O) would otherwise get 15 with no indication,
.and_then(|s| s.parse::<u32>().ok()) // and would keep tuning a knob that never took effect.
.unwrap_or(DEFAULT_MAX_CONNECTIONS); let max_connections = match std::env::var("DATABASE_MAX_CONNECTIONS") {
Err(_) => DEFAULT_MAX_CONNECTIONS,
Ok(raw) => match raw.trim().parse::<u32>() {
Ok(0) => {
anyhow::bail!("DATABASE_MAX_CONNECTIONS must be at least 1 (got 0)");
}
Ok(n) => n,
Err(e) => {
anyhow::bail!(
"DATABASE_MAX_CONNECTIONS must be a positive integer (got {raw:?}): {e}"
);
}
},
};
let pool = PgPoolOptions::new() let pool = match PgPoolOptions::new()
.max_connections(max_connections) .max_connections(max_connections)
// Fail fast instead of parking. sqlx's default is 30s, which on a DB blip means every
// request AND all ~100 SSE session revalidations sit on the pool for half a minute
// before erroring — the app looks hung rather than degraded, and the backlog outlives
// the blip. Five seconds is far longer than a healthy acquire ever takes.
.acquire_timeout(std::time::Duration::from_secs(5))
// Keep a couple of connections warm so the first request after an idle stretch (the gap
// between setting the venue up and the guests arriving) doesn't pay TCP + auth.
.min_connections(2)
// Bound every statement server-side. Without this a single pathological query holds a
// pool slot indefinitely and no client-side timeout can take it back — the slot is only
// released when Postgres finishes. `lock_timeout` covers the same hazard for a row lock
// contended by, say, a release running against an in-flight upload.
.after_connect(|conn, _meta| {
Box::pin(async move {
// Two statements, two round-trips, deliberately. `sqlx::query` uses the extended
// query protocol, which permits exactly ONE statement per call — sending them as
// `SET a; SET b` makes every new connection fail, which surfaces as the pool
// never opening one at all and `create_pool` reporting a connect timeout.
sqlx::query("SET statement_timeout = '15s'")
.execute(&mut *conn)
.await?;
sqlx::query("SET lock_timeout = '5s'").execute(conn).await?;
Ok(())
})
})
.connect(database_url) .connect(database_url)
.await .await
.context("failed to connect to database")?; {
Ok(pool) => pool,
Err(e) => {
explain_auth_failure(&e);
return Err(e).context("failed to connect to database");
}
};
// Migrations run on their OWN connection, deliberately NOT from the pool.
//
// `after_connect` above puts `lock_timeout = 5s` on every pooled connection, and the migrator
// would inherit it. Migrations that take ACCESS EXCLUSIVE (026's index swap, 027's ADD COLUMN)
// then turn a short WAIT into a hard FAILURE: anything holding ACCESS SHARE on `upload` or
// `"user"` for more than five seconds — the hourly `pg_dump` the runbook installs in §10.2, or
// an operator's open `psql` transaction — aborts the migration, `create_pool` returns an
// error, `main` exits 1, and `restart: unless-stopped` crash-loops the app behind a live Caddy.
// The rollback is clean and a later retry succeeds, which is exactly what makes it a confusing
// intermittent outage rather than an obvious one.
//
// `statement_timeout` is left off here too: a migration on a real table can legitimately run
// longer than the 15s a request is allowed.
let mut migrator_conn = <sqlx::PgConnection as sqlx::Connection>::connect(database_url)
.await
.context("failed to open a connection for migrations")?;
sqlx::migrate!() sqlx::migrate!()
.run(&pool) .run(&mut migrator_conn)
.await .await
.context("failed to run database migrations")?; .context("failed to run database migrations")?;
let _ = sqlx::Connection::close(migrator_conn).await;
tracing::info!(max_connections, "database connected and migrations applied"); tracing::info!(max_connections, "database connected and migrations applied");
Ok(pool) Ok(pool)

View File

@@ -11,6 +11,27 @@ pub enum AppError {
/// (banned user, quota): the queued blob is kept and retried if the host reopens, /// (banned user, quota): the queued blob is kept and retried if the host reopens,
/// instead of being purged like a genuinely-terminal rejection. /// instead of being purged like a genuinely-terminal rejection.
UploadsLocked(String), UploadsLocked(String),
/// The gallery has been RELEASED — the keepsake was snapshotted, so a late upload could
/// never appear in it. Mechanically this is still reversible (a host reopen clears
/// `export_released_at` and bumps the epoch), which is why the blob must still be kept.
///
/// Distinct from `UploadsLocked` because the two differ in *expectation*, and the client's
/// retry policy has to differ with them. A closed event is a pause the host means to undo;
/// a released gallery is the end of the event, and nobody reopens it. Under one shared code
/// the queue kept auto-retrying a released event forever — re-streaming a multi-megabyte
/// photo over cellular on every budget refill, for a request whose answer will not change,
/// while telling the guest to tap a camera button that 403s. `gallery_released` lets the
/// client park the item visibly and wait for an actual `event-opened` instead of guessing.
GalleryReleased(String),
/// The uploader is banned. A 403 like `Forbidden`, but tagged `user_banned` so the client
/// keeps the queued blob instead of purging it.
///
/// A ban is reversible — `unban_user` exists, and the host's own confirm copy promises the
/// photos come back — but the client classified the generic `forbidden` code as permanent,
/// deleted the blob from IndexedDB, and moved the row to `blocked`, which has no retry
/// button. So an unban could restore everything except the photos that were in flight when
/// the ban landed, and a ban issued by mistake destroyed them with no way back.
UserBanned(String),
NotFound(String), NotFound(String),
Conflict(String), Conflict(String),
/// Second field: optional retry-after seconds to include in the response. /// Second field: optional retry-after seconds to include in the response.
@@ -19,6 +40,12 @@ pub enum AppError {
/// the client can treat it as *terminal* (413, no retry) instead of backing off and /// the client can treat it as *terminal* (413, no retry) instead of backing off and
/// retrying a permanently-failing upload forever. /// retrying a permanently-failing upload forever.
QuotaExceeded(String), QuotaExceeded(String),
/// The server is temporarily unable to serve this request — currently only pool
/// saturation. Distinct from `Internal` because it is TRANSIENT and the client should be
/// told so: a 500 reads as "this request is broken", while a 503 + Retry-After reads as
/// "come back shortly", which is what the upload queue's retry classifier needs to make
/// the right call. Second field: optional retry-after seconds.
ServiceUnavailable(String, Option<u64>),
Internal(anyhow::Error), Internal(anyhow::Error),
} }
@@ -29,10 +56,15 @@ impl AppError {
Self::Unauthorized(_) => (StatusCode::UNAUTHORIZED, "unauthorized"), Self::Unauthorized(_) => (StatusCode::UNAUTHORIZED, "unauthorized"),
Self::Forbidden(_) => (StatusCode::FORBIDDEN, "forbidden"), Self::Forbidden(_) => (StatusCode::FORBIDDEN, "forbidden"),
Self::UploadsLocked(_) => (StatusCode::FORBIDDEN, "uploads_locked"), Self::UploadsLocked(_) => (StatusCode::FORBIDDEN, "uploads_locked"),
Self::GalleryReleased(_) => (StatusCode::FORBIDDEN, "gallery_released"),
Self::UserBanned(_) => (StatusCode::FORBIDDEN, "user_banned"),
Self::NotFound(_) => (StatusCode::NOT_FOUND, "not_found"), Self::NotFound(_) => (StatusCode::NOT_FOUND, "not_found"),
Self::Conflict(_) => (StatusCode::CONFLICT, "conflict"), Self::Conflict(_) => (StatusCode::CONFLICT, "conflict"),
Self::TooManyRequests(..) => (StatusCode::TOO_MANY_REQUESTS, "too_many_requests"), Self::TooManyRequests(..) => (StatusCode::TOO_MANY_REQUESTS, "too_many_requests"),
Self::QuotaExceeded(_) => (StatusCode::PAYLOAD_TOO_LARGE, "quota_exceeded"), Self::QuotaExceeded(_) => (StatusCode::PAYLOAD_TOO_LARGE, "quota_exceeded"),
Self::ServiceUnavailable(..) => {
(StatusCode::SERVICE_UNAVAILABLE, "service_unavailable")
}
Self::Internal(_) => (StatusCode::INTERNAL_SERVER_ERROR, "internal_error"), Self::Internal(_) => (StatusCode::INTERNAL_SERVER_ERROR, "internal_error"),
} }
} }
@@ -43,9 +75,12 @@ impl AppError {
| Self::Unauthorized(msg) | Self::Unauthorized(msg)
| Self::Forbidden(msg) | Self::Forbidden(msg)
| Self::UploadsLocked(msg) | Self::UploadsLocked(msg)
| Self::GalleryReleased(msg)
| Self::UserBanned(msg)
| Self::NotFound(msg) | Self::NotFound(msg)
| Self::Conflict(msg) => msg.clone(), | Self::Conflict(msg) => msg.clone(),
Self::TooManyRequests(msg, _) => msg.clone(), Self::TooManyRequests(msg, _) => msg.clone(),
Self::ServiceUnavailable(msg, _) => msg.clone(),
Self::QuotaExceeded(msg) => msg.clone(), Self::QuotaExceeded(msg) => msg.clone(),
Self::Internal(err) => { Self::Internal(err) => {
tracing::error!("internal error: {err:#}"); tracing::error!("internal error: {err:#}");
@@ -58,13 +93,61 @@ impl AppError {
impl IntoResponse for AppError { impl IntoResponse for AppError {
fn into_response(self) -> Response { fn into_response(self) -> Response {
let (status, code) = self.status_and_code(); let (status, code) = self.status_and_code();
let retry_after_secs = if let Self::TooManyRequests(_, Some(secs)) = &self { // BOTH retry-carrying variants must be matched here. `message()` would fail to
Some(*secs) // compile on a missing arm; this one would not — it would silently drop the header and
} else { // the `retry_after_secs` body field, which is exactly the sort of omission that only
None // shows up under the load the 503 exists for.
let retry_after_secs = match &self {
Self::TooManyRequests(_, secs) | Self::ServiceUnavailable(_, secs) => *secs,
_ => None,
}; };
let message = self.message(); let message = self.message();
// Log every 4xx. Until now they were invisible at ANY log level: tower_http's
// `ServerErrorsAsFailures` classifier counts a 4xx as a *success*, so it goes to
// `DefaultOnResponse` at DEBUG, and production runs at `info`. The consequence is that a
// misconfigured limit leaves no trace at all — if guests spend the evening hitting 429s
// on `upload_rate_per_hour`, or 413s on the storage quota, `docker compose logs` after
// the event contains nothing about it and the cause is unknowable.
//
// WARN rather than INFO because every variant here is a request that did not do what
// the guest asked. 5xx is excluded: `Internal` already logs with its full source chain
// in `message()` above, and the pool-exhaustion 503 logs at construction — logging again
// here would double every server-side failure.
//
// No request context is available: `into_response` receives only the error, so there is
// no path, method or user id to attach. Status + code + message is what can honestly be
// reported from here, and it is enough to see the SHAPE of a bad evening. Raising
// `tower_http` to DEBUG instead was considered and rejected — see the note in main.rs.
//
// `detail = ?message`, NOT `%message`. Two reasons, both learned the hard way:
//
// * `message` is tracing's own reserved field for an event's format literal, so `%message`
// printed unlabelled and would collide under a JSON layer.
// * Debug formatting QUOTES AND ESCAPES the string, and several 4xx messages interpolate
// attacker-chosen text — the guest's name in `Der Name "X" ist bereits vergeben.`, and
// multipart/parse errors that echo their input. With Display formatting, a value
// carrying a newline plus a plausible log prefix lets two unauthenticated requests
// forge lines in the only forensic record an unattended event has.
// `validate_display_name` now rejects control characters, so the name route is closed
// at the source as well — but that is ONE input, and this line formats every 4xx
// message in the app. Escaping here is what makes the guarantee general; do not
// "simplify" it to `%message` on the grounds that names are already validated.
//
// 401 and 404 are logged at DEBUG rather than WARN. They carry no operator signal (an
// expired session, a mistyped URL) and they are the cheapest lines for a scanner to
// generate — at ~260 bytes each against the 30 MB the json-file driver retains
// (docker-compose.yml), a sustained flood could otherwise roll the whole window in
// minutes and destroy the post-event forensics this logging exists to provide.
if status.is_client_error() {
let noisy = status == StatusCode::UNAUTHORIZED || status == StatusCode::NOT_FOUND;
if noisy {
tracing::debug!(status = status.as_u16(), code, detail = ?message, "request rejected");
} else {
tracing::warn!(status = status.as_u16(), code, detail = ?message, "request rejected");
}
}
let mut body = serde_json::json!({ let mut body = serde_json::json!({
"error": code, "error": code,
"message": message, "message": message,
@@ -93,6 +176,120 @@ impl From<anyhow::Error> for AppError {
impl From<sqlx::Error> for AppError { impl From<sqlx::Error> for AppError {
fn from(err: sqlx::Error) -> Self { fn from(err: sqlx::Error) -> Self {
Self::Internal(err.into()) match err {
// Pool saturation is load, not a bug. Reporting it as a 500 was actively harmful:
// the frontend's upload-queue classifier treats 5xx as transient and retries, so
// the retries piled straight back into the saturated pool with no Retry-After to
// pace them. A 503 says the same thing honestly and carries the backoff.
//
// `PoolClosed` stays `Internal` — it only happens during shutdown, where a 503
// would invite a retry against a server that is going away.
sqlx::Error::PoolTimedOut => {
tracing::warn!("database pool exhausted; shedding a request with 503");
Self::ServiceUnavailable(
"Server ist gerade ausgelastet. Bitte versuche es in ein paar Sekunden erneut."
.into(),
Some(POOL_TIMEOUT_RETRY_AFTER_SECS),
)
}
other => Self::Internal(other.into()),
}
}
}
/// Retry-After for a shed request. Short: pool saturation clears in seconds once the queue
/// drains, and a long value would make a brief spike feel like an outage.
const POOL_TIMEOUT_RETRY_AFTER_SECS: u64 = 3;
#[cfg(test)]
mod tests {
use super::*;
/// `into_response` extracts `retry_after_secs` by MATCHING ON VARIANTS, so unlike
/// `message()` a missing arm is not a compile error — it silently drops the header. Pin the
/// behaviour for both retry-carrying variants.
#[test]
fn both_retry_carrying_variants_emit_retry_after() {
for err in [
AppError::TooManyRequests("slow down".into(), Some(42)),
AppError::ServiceUnavailable("busy".into(), Some(3)),
] {
let expected = match &err {
AppError::TooManyRequests(_, Some(s))
| AppError::ServiceUnavailable(_, Some(s)) => s.to_string(),
_ => unreachable!(),
};
let resp = err.into_response();
assert_eq!(
resp.headers()
.get(axum::http::header::RETRY_AFTER)
.and_then(|v| v.to_str().ok()),
Some(expected.as_str()),
"a shed/throttled client must be told when to come back"
);
}
}
/// 4xx must be logged and 5xx must not be logged HERE — `Internal` logs its source chain in
/// `message()` and the pool-exhaustion 503 logs at construction, so a second line in
/// `into_response` would double every server-side failure in the post-event logs.
///
/// The guard is `status.is_client_error()`, so this pins the classification rather than the
/// logging itself (which needs a subscriber to observe).
#[test]
fn only_client_errors_are_in_the_logged_band() {
for err in [
AppError::BadRequest("x".into()),
AppError::Unauthorized("x".into()),
AppError::Forbidden("x".into()),
AppError::UploadsLocked("x".into()),
AppError::NotFound("x".into()),
AppError::Conflict("x".into()),
AppError::TooManyRequests("x".into(), Some(1)),
AppError::QuotaExceeded("x".into()),
] {
let (status, _) = err.status_and_code();
assert!(
status.is_client_error(),
"{status} should be in the 4xx band this logs"
);
}
for err in [
AppError::ServiceUnavailable("x".into(), Some(3)),
AppError::Internal(anyhow::anyhow!("boom")),
] {
let (status, _) = err.status_and_code();
assert!(
!status.is_client_error(),
"{status} logs elsewhere; logging it here would double it"
);
}
}
/// Pool saturation is load, not a bug. A 500 makes the frontend's retry classifier pile
/// straight back into the saturated pool with no backoff to pace it.
#[test]
fn pool_exhaustion_sheds_with_503_but_shutdown_does_not() {
let shed: AppError = sqlx::Error::PoolTimedOut.into();
assert_eq!(
shed.status_and_code(),
(StatusCode::SERVICE_UNAVAILABLE, "service_unavailable")
);
// PoolClosed only happens during shutdown; a 503 there would invite a retry against a
// server that is going away.
let closing: AppError = sqlx::Error::PoolClosed.into();
assert_eq!(
closing.status_and_code(),
(StatusCode::INTERNAL_SERVER_ERROR, "internal_error")
);
// Everything else must keep its existing mapping.
let missing: AppError = sqlx::Error::RowNotFound.into();
assert_eq!(
missing.status_and_code(),
(StatusCode::INTERNAL_SERVER_ERROR, "internal_error")
);
} }
} }

View File

@@ -3,13 +3,14 @@ use std::time::Duration;
use axum::Json; use axum::Json;
use axum::extract::{Query, State}; use axum::extract::{Query, State};
use axum::http::{HeaderMap, StatusCode}; use axum::http::StatusCode;
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use uuid::Uuid;
use crate::auth::middleware::RequireAdmin; use crate::auth::middleware::RequireAdmin;
use crate::error::AppError; use crate::error::AppError;
use crate::services::config; use crate::services::config;
use crate::services::rate_limiter::client_ip; use crate::services::sse_tickets::TicketKind;
use crate::state::AppState; use crate::state::AppState;
// ── DTOs ───────────────────────────────────────────────────────────────────── // ── DTOs ─────────────────────────────────────────────────────────────────────
@@ -104,7 +105,7 @@ pub struct PatchConfigRequest(pub HashMap<String, String>);
pub async fn patch_config( pub async fn patch_config(
State(state): State<AppState>, State(state): State<AppState>,
RequireAdmin(_auth): RequireAdmin, RequireAdmin(auth): RequireAdmin,
Json(body): Json<HashMap<String, String>>, Json(body): Json<HashMap<String, String>>,
) -> Result<StatusCode, AppError> { ) -> Result<StatusCode, AppError> {
// Numeric keys validated as f64; boolean keys validated as truthy strings; the // Numeric keys validated as f64; boolean keys validated as truthy strings; the
@@ -120,8 +121,28 @@ pub async fn patch_config(
("upload_rate_per_hour", true, 1.0, 100_000.0), ("upload_rate_per_hour", true, 1.0, 100_000.0),
("feed_rate_per_min", true, 1.0, 100_000.0), ("feed_rate_per_min", true, 1.0, 100_000.0),
("export_rate_per_day", true, 1.0, 100_000.0), ("export_rate_per_day", true, 1.0, 100_000.0),
// Loose per-IP ceiling on /join. The real anti-spam bucket is per (ip, name); this
// only bounds raw volume from one source, so it must stay well above the size of a
// party arriving at once (see migration 017).
("join_ip_rate_per_min", true, 1.0, 100_000.0),
// Same shape for /recover: the per-(ip, name) bucket is the anti-guessing control,
// this only bounds a name-cycling flood in front of a cost-12 bcrypt (migration 019).
("recover_ip_rate_per_min", true, 1.0, 100_000.0),
// Aggregate ceiling on likes + comments + comment deletions, per user per minute.
// These were the only mutating endpoints with no limit at all (migration 020).
("social_rate_per_min", true, 1.0, 100_000.0),
("quota_tolerance", false, 0.0, 1.0), ("quota_tolerance", false, 0.0, 1.0),
("estimated_guest_count", true, 1.0, 1_000_000.0), ("estimated_guest_count", true, 1.0, 1_000_000.0),
// The three limiters migration 025 introduced. All are READ at runtime
// (`upload.rs` for the edit limiter, `auth/handlers.rs` for the other two) and 025
// INSERTs all of them into `config`, so `GET /admin/config` listed them while
// `PATCH /admin/config` answered "Unbekannter Konfigurationsschlüssel" — the same
// dead-key defect the comment under BOOL_KEYS says was fixed for the two login
// toggles. These are precisely the knobs an operator reaches for while abuse is
// happening, which is the one moment a restart to change them is unaffordable.
("upload_edit_rate_per_min", true, 1.0, 100_000.0),
("recover_name_rate_per_15min", true, 1.0, 100_000.0),
("pin_reset_ip_rate_per_min", true, 1.0, 100_000.0),
]; ];
const BOOL_KEYS: &[&str] = &[ const BOOL_KEYS: &[&str] = &[
"rate_limits_enabled", "rate_limits_enabled",
@@ -134,14 +155,35 @@ pub async fn patch_config(
// missing from this allowlist — so the switch existed in code and could never be flipped. // missing from this allowlist — so the switch existed in code and could never be flipped.
"admin_login_rate_enabled", "admin_login_rate_enabled",
"recover_rate_enabled", "recover_rate_enabled",
"social_rate_enabled",
// Read by `upload::edit_upload`, inserted by migration 025, and until now unreachable
// from this endpoint — see the note in NUMERIC_SPECS.
"upload_edit_rate_enabled",
"quota_enabled", "quota_enabled",
"storage_quota_enabled", "storage_quota_enabled",
"upload_count_quota_enabled", "upload_count_quota_enabled",
]; ];
const TEXT_KEYS: &[&str] = &["privacy_note"]; const TEXT_KEYS: &[&str] = &[
"privacy_note",
"theme_preset",
"theme_primary",
"theme_accent",
];
const PRIVACY_NOTE_MAX_LEN: usize = 16 * 1024; // 16 KiB free text is plenty const PRIVACY_NOTE_MAX_LEN: usize = 16 * 1024; // 16 KiB free text is plenty
// Preset ids the frontend knows how to render (mirror of PRESETS in
// frontend/src/lib/theme/palette.ts). "custom" means "use the theme_primary/accent
// seeds verbatim". Kept in sync by hand — a new preset must be added in both places.
const THEME_PRESETS: &[&str] = &[
"champagne-gold",
"rose",
"sage",
"dusk-blue",
"classic-silver",
"custom",
];
let mut privacy_note_changed = false; let mut privacy_note_changed = false;
let mut theme_changed = false;
// Validate every key first so a bad value in the batch can't leave a partial // Validate every key first so a bad value in the batch can't leave a partial
// update behind — validation must fully precede any write. // update behind — validation must fully precede any write.
@@ -169,6 +211,23 @@ pub async fn patch_config(
"Wert für {key} liegt außerhalb des zulässigen Bereichs ({min}{max})." "Wert für {key} liegt außerhalb des zulässigen Bereichs ({min}{max})."
))); )));
} }
// Zero is in range and catastrophic. `quota_tolerance` is the multiplier in
// `free_disk * tolerance / active_uploaders`, so 0 makes every per-user limit 0 and
// refuses EVERY upload — mid-event, with "Du hast dein Upload-Limit für dieses Event
// erreicht", an error naming the wrong cause entirely. `storage_quota_enabled` is the
// intended off-switch.
//
// Rejecting the value rather than raising the floor: very small tolerances are
// legitimate (they are how a large disk is throttled down to a sensible per-guest
// ceiling, and how the e2e quota tests steer it — around 1e-5 on a 174 GB volume), so
// a floor of, say, 0.01 would forbid real configurations to prevent one typo.
if key_str == "quota_tolerance" && n == 0.0 {
return Err(AppError::BadRequest(
"quota_tolerance = 0 würde jeden Upload blockieren. Zum Abschalten der \
Speicher-Quote stattdessen „Speicher-Quote aktiv“ ausschalten."
.into(),
));
}
} else if BOOL_KEYS.contains(&key_str) { } else if BOOL_KEYS.contains(&key_str) {
match value.trim().to_ascii_lowercase().as_str() { match value.trim().to_ascii_lowercase().as_str() {
"true" | "false" | "1" | "0" | "yes" | "no" | "on" | "off" => {} "true" | "false" | "1" | "0" | "yes" | "no" | "on" | "off" => {}
@@ -186,8 +245,23 @@ pub async fn patch_config(
"Wert für {key} ist zu lang (max. {PRIVACY_NOTE_MAX_LEN} Zeichen)." "Wert für {key} ist zu lang (max. {PRIVACY_NOTE_MAX_LEN} Zeichen)."
))); )));
} }
if key_str == "privacy_note" { match key_str {
privacy_note_changed = true; "privacy_note" => privacy_note_changed = true,
"theme_preset" => {
if !THEME_PRESETS.contains(&value.trim()) {
return Err(AppError::BadRequest(format!("Ungültiges Theme: {value}.")));
}
theme_changed = true;
}
"theme_primary" | "theme_accent" => {
if !is_hex_color(value.trim()) {
return Err(AppError::BadRequest(format!(
"Ungültige Farbe für {key}: muss #rrggbb sein."
)));
}
theme_changed = true;
}
_ => {}
} }
} else { } else {
return Err(AppError::BadRequest(format!( return Err(AppError::BadRequest(format!(
@@ -215,18 +289,49 @@ pub async fn patch_config(
// the TTL is only a backstop and must not be relied on for correctness. // the TTL is only a backstop and must not be relied on for correctness.
state.config_cache.invalidate(); state.config_cache.invalidate();
// Config changes were logged NOWHERE. They are the actions most likely to be blamed the
// morning after ("why did uploads stop?") and the hardest to reconstruct, because the value
// that caused the problem has since been changed back. Record the keys and their new values;
// these are operational settings, not credentials, so the payload is safe to keep.
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"patch_config",
None,
None,
serde_json::to_value(&body).ok(),
)
.await;
// Notify all clients that a publicly-readable config value changed so their stores // Notify all clients that a publicly-readable config value changed so their stores
// (e.g. the privacy note in My Account) refresh without a manual reload. // (e.g. the privacy note in My Account) refresh without a manual reload.
if privacy_note_changed || theme_changed {
let mut keys: Vec<&str> = Vec::new();
if privacy_note_changed { if privacy_note_changed {
keys.push("privacy_note");
}
if theme_changed {
keys.push("theme");
}
let _ = state.sse_tx.send(crate::state::SseEvent::new( let _ = state.sse_tx.send(crate::state::SseEvent::new(
"event-updated", "event-updated",
serde_json::json!({ "keys": ["privacy_note"] }).to_string(), serde_json::json!({ "keys": keys }).to_string(),
)); ));
} }
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
/// A strict `#rrggbb` hex-colour check (6 hex digits, leading `#`). Deliberately not
/// accepting shorthand/`#rgba` so the value is safe to drop straight into CSS.
fn is_hex_color(s: &str) -> bool {
let bytes = s.as_bytes();
bytes.len() == 7 && bytes[0] == b'#' && bytes[1..].iter().all(|b| b.is_ascii_hexdigit())
}
pub async fn get_export_jobs( pub async fn get_export_jobs(
State(state): State<AppState>, State(state): State<AppState>,
RequireAdmin(_auth): RequireAdmin, RequireAdmin(_auth): RequireAdmin,
@@ -259,44 +364,178 @@ pub struct DownloadQuery {
/// is a top-level navigation so the multi-GB ZIP streams straight to disk instead /// is a top-level navigation so the multi-GB ZIP streams straight to disk instead
/// of being buffered in memory by `fetch()` + `blob()` — but a navigation can't /// of being buffered in memory by `fetch()` + `blob()` — but a navigation can't
/// carry an `Authorization` header, so the client exchanges its Bearer token for /// carry an `Authorization` header, so the client exchanges its Bearer token for
/// an opaque ticket here, then hits `/export/zip?ticket=...`. Reuses the same /// an opaque ticket here, then hits `/export/zip?ticket=...`. Uses the same store as the SSE
/// single-use, 30s-TTL store as the SSE stream. /// stream, but NOT the same lifetime: a download ticket lives `DOWNLOAD_TTL` (6 h) and is
/// redeemable up to `MAX_DOWNLOAD_REDEMPTIONS` times, because a multi-GB transfer over venue wifi
/// has to survive being resumed with `Range`.
#[derive(serde::Deserialize)]
pub struct ExportTicketQuery {
/// Which archive the ticket is for — `zip` or `html`.
///
/// REQUIRED. It used to be optional "so an older client keeps working", but the ticket is now
/// bound to the archive it was minted for (see `TicketKind::Download`), and a ticket with no
/// archive would either have to be valid for both — the abuse this closes — or be issued for a
/// guess that 401s at the other endpoint. Every shipped client sends it.
#[serde(default)]
pub kind: Option<String>,
}
pub async fn export_ticket( pub async fn export_ticket(
State(state): State<AppState>, State(state): State<AppState>,
axum::extract::Query(q): axum::extract::Query<ExportTicketQuery>,
auth: crate::auth::middleware::AuthUser, auth: crate::auth::middleware::AuthUser,
) -> Json<serde_json::Value> { ) -> Result<Json<serde_json::Value>, AppError> {
// NOTE: intentionally NOT gated on `is_banned`. A banned user keeps *read* access // NOTE: intentionally NOT gated on `is_banned`. A banned user keeps *read* access
// by design (USER_JOURNEYS §10.3, FEATURES: "Can still download the export once // by design (USER_JOURNEYS §10.3, FEATURES: "Can still download the export once
// released — Spec design choice"). The export is read-only, so it stays available // released — Spec design choice"). The export is read-only, so it stays available
// to them, consistent with the read-only-ban model. // to them, consistent with the read-only-ban model.
let ticket = state.sse_tickets.issue(auth.token_hash);
Json(serde_json::json!({ "ticket": ticket })) // The rate limit is enforced HERE rather than on the download itself, and that placement is
// the whole point: the download is an iframe navigation, so its response is invisible to the
// page. Limiting it there meant a guest over the limit tapped "Herunterladen", the ticket
// POST returned 200, the iframe silently received a 429, and absolutely nothing happened —
// forever, with no explanation, on the one screen that is the emotional payoff of the app.
// Minting is a normal `fetch`, so a 429 here reaches the user as a German message.
//
// Moving it does not weaken the limit meaningfully: a ticket can only be obtained from this
// authenticated endpoint, is bound to one archive, and — since downloads must be resumable —
// is worth at most `MAX_DOWNLOAD_REDEMPTIONS` transfers rather than exactly one. The daily
// limit is therefore a bound on mints, not on bytes; see `MAX_DOWNLOAD_REDEMPTIONS` for why
// charging per redemption would re-break resumption.
// Confirm the archive actually EXISTS before spending anything on it.
//
// `export_status` — which is what enables the Download button — reports `done` from
// `export_job`, while the download resolves through `export_current.file_path` plus a
// `Path::exists()`. Those are different sources of truth and can legitimately disagree: a
// row can say done while the file is gone, or an epoch bump can retire it between the page
// rendering and the guest tapping. When they disagreed the guest got the worst possible
// shape of failure — a green "Download gestartet" toast, a consumed single-use ticket, one
// of only three daily slots spent, and nothing in their Downloads folder, repeatable until
// the day's allowance was gone.
//
// Checking here, before `enforce_export_rate`, turns that into an honest error on a plain
// `fetch` that the existing `toastError` path already renders. This is NOT the HEAD probe
// ruled out elsewhere: it reads the same indexed row the download will read and touches no
// ticket, so it cannot consume anything.
let export_kind = match q.kind.as_deref() {
Some("zip") => crate::services::sse_tickets::ExportKind::Zip,
Some("html") => crate::services::sse_tickets::ExportKind::Html,
Some(other) => {
return Err(AppError::BadRequest(format!(
"Unbekannter Export-Typ: {other}"
)));
}
None => {
return Err(AppError::BadRequest(
"Es fehlt die Angabe, welches Archiv geladen werden soll. Bitte lade die Seite \
neu und versuche es erneut."
.into(),
));
}
};
{
let export_type = match export_kind {
crate::services::sse_tickets::ExportKind::Zip => "zip",
crate::services::sse_tickets::ExportKind::Html => "html",
};
let msg = if export_type == "zip" {
"Der ZIP-Export ist noch nicht verfügbar."
} else {
"Der HTML-Export ist noch nicht verfügbar."
};
resolve_export_file(&state, export_type, msg).await?;
} }
/// Validate a download ticket (single-use) and confirm its session still exists. // `issue` returns None when the ticket store is at capacity. Unwrapping it into the JSON body
async fn authenticate_download_ticket(state: &AppState, ticket: &str) -> Result<(), AppError> { // serialized `{"ticket": null}` with a 200 — so `api.post` resolved happily, the page toasted
// success, and the iframe navigated to `?ticket=null`. 503 + Retry-After, matching how
// `sse::issue_ticket` answers the identical condition.
//
// Minted BEFORE the rate limit is charged. Charging first meant a store-capacity 503 — a
// server-side condition the guest did nothing to cause and cannot see — still cost one of
// their three DAILY downloads. There is no refund path, so the only fix is not to charge until
// the thing being charged for actually exists.
let ticket = state
.sse_tickets
.issue(auth.token_hash, TicketKind::Download(export_kind))
.ok_or_else(|| {
AppError::ServiceUnavailable(
"Server ist gerade ausgelastet. Bitte versuch es in einem Moment erneut.".into(),
Some(30),
)
})?;
// A refused mint must not leave its ticket behind. The per-session cap is FOUR tickets of the
// same kind, and a download ticket now lives six hours instead of being consumed on first use —
// so every abandoned one occupies a slot until it expires. A guest whose 1.4 GB transfer looks
// stuck and who taps "Herunterladen" a few more times spends mints 1-3 legitimately, then gets
// a 429 on taps 4 and 5 — but both still minted, and the fifth evicted the OLDEST download
// ticket for the session: the one the running transfer is holding. The next `Range` resume then
// 401s, and re-minting is impossible because they are at the daily limit. The keepsake is gone
// until tomorrow, having done nothing worse than tapping a button that appeared to do nothing.
//
// Discarding here keeps both properties that put the mint first: a store-capacity 503 still
// costs no download, and a refused download costs no slot.
if let Err(e) = enforce_export_rate(&state, auth.user_id).await {
let _ = state
.sse_tickets
.consume(&ticket, TicketKind::Download(export_kind));
return Err(e);
}
Ok(Json(serde_json::json!({ "ticket": ticket })))
}
/// Validate a download ticket and confirm its session still exists, resolving it to the user who
/// minted it. Deliberately NOT single-use — see the note on `redeem_download` below.
async fn authenticate_download_ticket(
state: &AppState,
ticket: &str,
want: crate::services::sse_tickets::ExportKind,
) -> Result<Uuid, AppError> {
// Non-consuming: a keepsake download must survive being resumed with `Range`, and a
// single-use ticket meant the resume 401'd and cost the guest another of their three daily
// downloads. `redeem_download` bounds it by DOWNLOAD_TTL instead, and the session check
// below still runs on every request.
let token_hash = state let token_hash = state
.sse_tickets .sse_tickets
.consume(ticket) .redeem_download(ticket, want)
.ok_or_else(|| AppError::Unauthorized("Ticket ungültig oder abgelaufen.".into()))?; .ok_or_else(|| AppError::Unauthorized("Ticket ungültig oder abgelaufen.".into()))?;
crate::models::session::Session::find_by_token_hash(&state.pool, &token_hash) let session = crate::models::session::Session::find_by_token_hash(&state.pool, &token_hash)
.await .await
.map_err(|e| AppError::Internal(e.into()))? .map_err(|e| AppError::Internal(e.into()))?
.ok_or_else(|| AppError::Unauthorized("Sitzung nicht gefunden.".into()))?; .ok_or_else(|| AppError::Unauthorized("Sitzung nicht gefunden.".into()))?;
Ok(()) Ok(session.user_id)
} }
pub async fn download_zip( pub async fn download_zip(
State(state): State<AppState>, State(state): State<AppState>,
headers: axum::http::HeaderMap,
Query(q): Query<DownloadQuery>, Query(q): Query<DownloadQuery>,
headers: HeaderMap,
) -> Result<axum::response::Response, AppError> { ) -> Result<axum::response::Response, AppError> {
authenticate_download_ticket(&state, &q.ticket).await?; // Ticket validation only — the rate limit was charged at mint time, where a 429 is visible
enforce_export_rate(&state, &headers).await?; // to the page. Charging it again here would cost every download two slots.
authenticate_download_ticket(
&state,
&q.ticket,
crate::services::sse_tickets::ExportKind::Zip,
)
.await?;
let path = let path =
resolve_export_file(&state, "zip", "Der ZIP-Export ist noch nicht verfügbar.").await?; resolve_export_file(&state, "zip", "Der ZIP-Export ist noch nicht verfügbar.").await?;
serve_file(path, "Gallery.zip", "application/zip").await serve_file(
path,
"Gallery.zip",
"application/zip",
headers
.get(axum::http::header::RANGE)
.and_then(|v| v.to_str().ok()),
headers
.get(axum::http::header::IF_RANGE)
.and_then(|v| v.to_str().ok()),
)
.await
} }
/// Resolve the on-disk path of the CURRENT export generation — readiness check and path lookup in /// Resolve the on-disk path of the CURRENT export generation — readiness check and path lookup in
@@ -342,46 +581,130 @@ async fn resolve_export_file(
pub async fn download_html( pub async fn download_html(
State(state): State<AppState>, State(state): State<AppState>,
headers: axum::http::HeaderMap,
Query(q): Query<DownloadQuery>, Query(q): Query<DownloadQuery>,
headers: HeaderMap,
) -> Result<axum::response::Response, AppError> { ) -> Result<axum::response::Response, AppError> {
authenticate_download_ticket(&state, &q.ticket).await?; // See `download_zip`: the limit is charged at ticket mint, where the client can see it.
enforce_export_rate(&state, &headers).await?; authenticate_download_ticket(
&state,
&q.ticket,
crate::services::sse_tickets::ExportKind::Html,
)
.await?;
let path = let path =
resolve_export_file(&state, "html", "Der HTML-Export ist noch nicht verfügbar.").await?; resolve_export_file(&state, "html", "Der HTML-Export ist noch nicht verfügbar.").await?;
serve_file(path, "Memories.zip", "application/zip").await serve_file(
path,
"Memories.zip",
"application/zip",
headers
.get(axum::http::header::RANGE)
.and_then(|v| v.to_str().ok()),
headers
.get(axum::http::header::IF_RANGE)
.and_then(|v| v.to_str().ok()),
)
.await
} }
/// Stream a keepsake archive, honouring `Range`.
///
/// Range support is not a nicety here. The keepsake is the emotional payoff of the product and can
/// be ~1.4 GB; without `Accept-Ranges` a download that dies at 90% over hotel wifi restarts at byte
/// zero. Worse, the 3/day limit is charged when the download TICKET is minted and ZIP+HTML already
/// costs 2 — so one dropped connection locked a guest out of their own wedding photos for ~24h.
///
/// Reuses `upload::parse_range`, which already implements exactly the forms a client sends and is
/// unit-tested there. The media routes have always done this correctly; this route was the outlier.
async fn serve_file( async fn serve_file(
path: std::path::PathBuf, path: std::path::PathBuf,
filename: &str, filename: &str,
content_type: &str, content_type: &str,
range_header: Option<&str>,
if_range_header: Option<&str>,
) -> Result<axum::response::Response, AppError> { ) -> Result<axum::response::Response, AppError> {
use crate::handlers::upload::{RangeSpec, parse_range};
use axum::body::Body; use axum::body::Body;
use axum::http::{Response, StatusCode, header}; use axum::http::{Response, StatusCode, header};
use tokio::io::{AsyncReadExt, AsyncSeekExt};
use tokio_util::io::ReaderStream; use tokio_util::io::ReaderStream;
let file = tokio::fs::File::open(&path) let mut file = tokio::fs::File::open(&path)
.await .await
.map_err(|e| AppError::Internal(e.into()))?; .map_err(|e| AppError::Internal(e.into()))?;
let metadata = file let len = file
.metadata() .metadata()
.await .await
.map_err(|e| AppError::Internal(e.into()))?; .map_err(|e| AppError::Internal(e.into()))?
let stream = ReaderStream::new(file); .len();
let disposition = format!("attachment; filename=\"{filename}\""); let disposition = format!("attachment; filename=\"{filename}\"");
let response = Response::builder() // A validator that CHANGES when the archive does, so a resume cannot splice two generations.
.status(StatusCode::OK) //
.header(header::CONTENT_TYPE, content_type) // The on-disk name is `{prefix}.{event_id}.{epoch}.zip`, so it already identifies the exact
.header(header::CONTENT_DISPOSITION, disposition) // generation; length distinguishes a rebuild at the same epoch. Together they are a strong
.header(header::CONTENT_LENGTH, metadata.len()) // validator.
.body(Body::from_stream(stream)) //
.map_err(|e| AppError::Internal(e.into()))?; // Why this matters: `resolve_export_file` re-reads `export_current` on EVERY request, and a
// download ticket outlives several redemptions. So a guest whose 500 MB download drops at
// 500 MB, while the host takes a photo down (epoch bumps, rebuild lands, the old generation is
// pruned), used to resume with `Range: bytes=500000000-` against a DIFFERENT FILE of a
// different length — and the server would happily seek 500 MB into it and stream. The client
// concatenated the two halves into a structurally corrupt ZIP, with nothing logged anywhere.
let etag = format!(
"\"{}-{len}\"",
path.file_name()
.and_then(|n| n.to_str())
.unwrap_or(filename)
);
Ok(response) let base = |status: StatusCode| {
Response::builder()
.status(status)
.header(header::CONTENT_TYPE, content_type)
.header(header::CONTENT_DISPOSITION, disposition.clone())
// Advertised on EVERY response, including the 200. A client only knows it may resume
// if the first (unranged) response says so.
.header(header::ACCEPT_RANGES, "bytes")
.header(header::ETAG, etag.clone())
};
// Serve a partial ONLY when the client proves it is resuming the same bytes.
//
// `If-Range` matching our ETag is that proof. A client that sends `Range` with no `If-Range`
// at all (curl -C -, wget -c, most download managers) cannot be given a partial safely — it
// has no way to notice the archive changed underneath it — so it gets a 200 and starts over.
// Restarting a download is a cost; a silently corrupt keepsake is not recoverable. Browsers
// send `If-Range`, so the ordinary resume path is unaffected, and this is the first release
// where their resume works at all: without a validator they simply refused to try.
let resume_is_safe = if_range_header.is_some_and(|v| v.trim() == etag);
let effective_range = if resume_is_safe { range_header } else { None };
match parse_range(effective_range, len) {
RangeSpec::Full => base(StatusCode::OK)
.header(header::CONTENT_LENGTH, len)
.body(Body::from_stream(ReaderStream::new(file)))
.map_err(|e| AppError::Internal(e.into())),
RangeSpec::Partial { start, end } => {
file.seek(std::io::SeekFrom::Start(start))
.await
.map_err(|e| AppError::Internal(e.into()))?;
let span = end - start + 1;
base(StatusCode::PARTIAL_CONTENT)
.header(header::CONTENT_LENGTH, span)
.header(header::CONTENT_RANGE, format!("bytes {start}-{end}/{len}"))
.body(Body::from_stream(ReaderStream::new(file.take(span))))
.map_err(|e| AppError::Internal(e.into()))
}
RangeSpec::Unsatisfiable => base(StatusCode::RANGE_NOT_SATISFIABLE)
.header(header::CONTENT_RANGE, format!("bytes */{len}"))
.body(Body::empty())
.map_err(|e| AppError::Internal(e.into())),
}
} }
/// Also expose export status to all authenticated users (guests need it for the export page) /// Also expose export status to all authenticated users (guests need it for the export page)
@@ -403,8 +726,13 @@ pub async fn export_status(
// worker superseded mid-run) is meaningless — surfacing its frozen `running`/77% would show a // worker superseded mid-run) is meaningless — surfacing its frozen `running`/77% would show a
// progress bar that never moves for a keepsake nobody is building. It reads as "locked" (no // progress bar that never moves for a keepsake nobody is building. It reads as "locked" (no
// current job), which is exactly what it is. // current job), which is exactly what it is.
let jobs: Vec<(String, String, i16)> = sqlx::query_as( // `error_message` is carried here, not just on the admin dashboard's job list. The host is the
"SELECT j.type::text, j.status::text, j.progress_pct // one who releases the keepsake and the one who owns the "Erneut versuchen" button, but this
// endpoint used to hand them a bare `failed` — so a fully actionable reason (notably the disk
// preflight's "needs X GB, Y GB free") was written to the row and then shown to nobody who
// could act on it. An admin-only diagnostic is not a diagnostic for the person on the spot.
let jobs: Vec<(String, String, i16, Option<String>)> = sqlx::query_as(
"SELECT j.type::text, j.status::text, j.progress_pct, j.error_message
FROM export_job j FROM export_job j
JOIN event e ON e.id = j.event_id JOIN event e ON e.id = j.event_id
WHERE e.id = $1 AND j.epoch = e.export_epoch", WHERE e.id = $1 AND j.epoch = e.export_epoch",
@@ -415,9 +743,21 @@ pub async fn export_status(
let job_status = |type_name: &str| { let job_status = |type_name: &str| {
jobs.iter() jobs.iter()
.find(|(t, _, _)| t == type_name) .find(|(t, _, _, _)| t == type_name)
.map(|(_, status, pct)| serde_json::json!({ "status": status, "progress_pct": pct })) .map(|(_, status, pct, err)| {
.unwrap_or_else(|| serde_json::json!({ "status": "locked", "progress_pct": 0 })) serde_json::json!({
"status": status,
"progress_pct": pct,
// Only on a failure. A stale message left on a row that has since been re-armed
// would otherwise show an error next to a running progress bar.
"error_message": if status == "failed" { err.clone() } else { None },
})
})
.unwrap_or_else(|| {
serde_json::json!({
"status": "locked", "progress_pct": 0, "error_message": null,
})
})
}; };
Ok(Json(serde_json::json!({ Ok(Json(serde_json::json!({
@@ -430,21 +770,29 @@ pub async fn export_status(
/// Centralised guard for the export rate limit. Same pattern as upload/feed: master /// Centralised guard for the export rate limit. Same pattern as upload/feed: master
/// switch + per-endpoint switch + numeric value, all stored in `config` and read on /// switch + per-endpoint switch + numeric value, all stored in `config` and read on
/// each request. /// each request.
async fn enforce_export_rate(state: &AppState, headers: &HeaderMap) -> Result<(), AppError> { async fn enforce_export_rate(state: &AppState, user_id: Uuid) -> Result<(), AppError> {
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await; let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let export_rate_on = config::get_bool(&state.config_cache, "export_rate_enabled", true).await; let export_rate_on = config::get_bool(&state.config_cache, "export_rate_enabled", true).await;
if !(rate_limits_on && export_rate_on) { if !(rate_limits_on && export_rate_on) {
return Ok(()); return Ok(());
} }
let ip = client_ip(headers, "unknown");
let limit = config::get_usize(&state.config_cache, "export_rate_per_day", 3).await; let limit = config::get_usize(&state.config_cache, "export_rate_per_day", 3).await;
if !state // Keyed per-user. This was the worst of the IP-keyed limiters: 3 downloads per DAY
.rate_limiter // shared across every guest behind the venue's public IP, so the fourth person to
.check(format!("export:{ip}"), limit, Duration::from_secs(86400)) // fetch their keepsake was locked out until the next day.
{ if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("export:{user_id}"),
limit,
Duration::from_secs(86400),
) {
// Names the real window. The generic "warte kurz" wording this used to share with the
// per-minute limiters is actively wrong here — the bucket is a DAY, so a guest told to
// wait a moment would keep tapping a button that cannot work again until tomorrow.
return Err(AppError::TooManyRequests( return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(), "Du hast das Tageslimit für Downloads erreicht. Versuch es später noch einmal — \
None, deine Galerie bleibt gespeichert."
.into(),
Some(retry_after_secs),
)); ));
} }
Ok(()) Ok(())

View File

@@ -2,7 +2,6 @@ use std::time::Duration;
use axum::Json; use axum::Json;
use axum::extract::{Query, State}; use axum::extract::{Query, State};
use axum::http::HeaderMap;
use chrono::{DateTime, Utc}; use chrono::{DateTime, Utc};
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use uuid::Uuid; use uuid::Uuid;
@@ -10,14 +9,49 @@ use uuid::Uuid;
use crate::auth::middleware::AuthUser; use crate::auth::middleware::AuthUser;
use crate::error::AppError; use crate::error::AppError;
use crate::services::config; use crate::services::config;
use crate::services::rate_limiter::client_ip;
use crate::state::AppState; use crate::state::AppState;
#[derive(Deserialize)] #[derive(Deserialize)]
pub struct FeedQuery { pub struct FeedQuery {
pub cursor: Option<Uuid>, pub cursor: Option<Uuid>,
pub limit: Option<i64>, pub limit: Option<i64>,
/// Single tag (list view). Kept alongside `hashtags` so existing callers keep working.
pub hashtag: Option<String>, pub hashtag: Option<String>,
/// Comma-separated tags, combined with **OR** — the grid's chip semantics
/// (USER_JOURNEYS §8). Filtering moved server-side because the client could only ever
/// filter the pages it had already loaded: with page 1 = 20 items out of a 1000-photo
/// event, selecting a tag showed a handful of tiles and looked complete. Matching is
/// now the exact `hashtag` row in both views, so the grid and the list can no longer
/// disagree about which photos carry a tag (the client matched a caption SUBSTRING, so
/// `#tanz` also matched `#tanzflaeche`).
pub hashtags: Option<String>,
/// Exact uploader display name, combined with the tag group using **AND**.
pub uploader: Option<String>,
}
/// Merge the single-tag and CSV tag params into one normalised, de-duplicated list.
///
/// Normalisation mirrors `Hashtag::upsert` exactly (trim, drop a leading `#`, lowercase), so
/// a chip built from a display string like `#Tanz` matches the stored `tanz` row. Returns
/// `None` when no usable tag was supplied, which makes the SQL predicate a no-op — an empty
/// list must mean "no tag filter", never "match nothing".
fn normalize_tags(single: Option<&str>, csv: Option<&str>) -> Option<Vec<String>> {
let mut out: Vec<String> = Vec::new();
let mut push = |raw: &str| {
let t = raw.trim().trim_start_matches('#').to_lowercase();
if !t.is_empty() && !out.contains(&t) {
out.push(t);
}
};
if let Some(s) = single {
push(s);
}
if let Some(s) = csv {
for part in s.split(',') {
push(part);
}
}
if out.is_empty() { None } else { Some(out) }
} }
#[derive(Serialize)] #[derive(Serialize)]
@@ -27,6 +61,8 @@ pub struct FeedUpload {
pub uploader_name: String, pub uploader_name: String,
pub preview_url: Option<String>, pub preview_url: Option<String>,
pub thumbnail_url: Option<String>, pub thumbnail_url: Option<String>,
/// Big-screen (~2048px) variant for the diashow. Absent until the derivative exists.
pub display_url: Option<String>,
pub mime_type: String, pub mime_type: String,
pub caption: Option<String>, pub caption: Option<String>,
pub like_count: i64, pub like_count: i64,
@@ -48,6 +84,7 @@ struct FeedRow {
uploader_name: String, uploader_name: String,
preview_path: Option<String>, preview_path: Option<String>,
thumbnail_path: Option<String>, thumbnail_path: Option<String>,
display_path: Option<String>,
mime_type: String, mime_type: String,
caption: Option<String>, caption: Option<String>,
like_count: i64, like_count: i64,
@@ -58,26 +95,30 @@ struct FeedRow {
pub async fn feed( pub async fn feed(
State(state): State<AppState>, State(state): State<AppState>,
auth: AuthUser, auth: AuthUser,
headers: HeaderMap,
Query(q): Query<FeedQuery>, Query(q): Query<FeedQuery>,
) -> Result<Json<FeedResponse>, AppError> { ) -> Result<Json<FeedResponse>, AppError> {
let ip = client_ip(&headers, "unknown");
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await; let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let feed_rate_on = config::get_bool(&state.config_cache, "feed_rate_enabled", true).await; let feed_rate_on = config::get_bool(&state.config_cache, "feed_rate_enabled", true).await;
if rate_limits_on && feed_rate_on { if rate_limits_on && feed_rate_on {
let rate_limit = config::get_usize(&state.config_cache, "feed_rate_per_min", 60).await; let rate_limit = config::get_usize(&state.config_cache, "feed_rate_per_min", 60).await;
if !state // Keyed per-user, exactly like `feed_delta` below: at a venue every guest shares
.rate_limiter // one public IP, so an IP key gave the whole party a single 60/min bucket and the
.check(format!("feed:{ip}"), rate_limit, Duration::from_secs(60)) // fastest scroller starved everyone else.
{ if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("feed:{}", auth.user_id),
rate_limit,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests( return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(), "Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None, Some(retry_after_secs),
)); ));
} }
} }
let limit = q.limit.unwrap_or(20).min(100); // Clamped at BOTH ends: only the upper bound was enforced, so `?limit=-5` reached Postgres
// as `LIMIT -4` and answered a hand-written URL with a 500.
let limit = q.limit.unwrap_or(20).clamp(1, 100);
// Resolve the cursor to a (created_at, id) position. The pair is compared as a // Resolve the cursor to a (created_at, id) position. The pair is compared as a
// tuple so ties on created_at break on id — keyset pagination on created_at // tuple so ties on created_at break on id — keyset pagination on created_at
@@ -90,43 +131,42 @@ pub async fn feed(
None => (None, None), None => (None, None),
}; };
let rows = if let Some(hashtag) = &q.hashtag { // Tags from either param, normalised the same way `Hashtag::upsert` stores them
let tag = hashtag.trim().trim_start_matches('#').to_lowercase(); // (trimmed, leading `#` dropped, lowercased) so the comparison is exact.
sqlx::query_as::<_, FeedRow>( let tags = normalize_tags(q.hashtag.as_deref(), q.hashtags.as_deref());
let uploader = q
.uploader
.as_deref()
.map(str::trim)
.filter(|s| !s.is_empty());
// ONE statement for every combination, rather than a branch per filter. `EXISTS` with
// `= ANY($4)` gives OR across the tag group without the row multiplication a JOIN would
// cause when a photo carries two selected tags; the uploader predicate ANDs on top. Both
// are no-ops when NULL, so the unfiltered feed takes the same path.
let rows = sqlx::query_as::<_, FeedRow>(
"SELECT v.id, v.user_id, v.uploader_name, v.preview_path, v.thumbnail_path, "SELECT v.id, v.user_id, v.uploader_name, v.preview_path, v.thumbnail_path,
v.mime_type, v.caption, v.like_count, v.comment_count, v.created_at v.display_path, v.mime_type, v.caption, v.like_count, v.comment_count,
v.created_at
FROM v_feed v FROM v_feed v
JOIN upload_hashtag uh ON uh.upload_id = v.id WHERE v.event_id = $1
JOIN hashtag h ON h.id = uh.hashtag_id AND h.tag = $1 AND ($2::timestamptz IS NULL OR (v.created_at, v.id) < ($2, $3))
WHERE v.event_id = $2 AND ($4::text[] IS NULL OR EXISTS (
AND ($3::timestamptz IS NULL OR (v.created_at, v.id) < ($3, $4)) SELECT 1 FROM upload_hashtag uh
JOIN hashtag h ON h.id = uh.hashtag_id
WHERE uh.upload_id = v.id AND h.tag = ANY($4)))
AND ($5::text IS NULL OR v.uploader_name = $5)
ORDER BY v.created_at DESC, v.id DESC ORDER BY v.created_at DESC, v.id DESC
LIMIT $5", LIMIT $6",
)
.bind(&tag)
.bind(auth.event_id)
.bind(cursor_time)
.bind(cursor_id)
.bind(limit + 1)
.fetch_all(&state.pool)
.await?
} else {
sqlx::query_as::<_, FeedRow>(
"SELECT id, user_id, uploader_name, preview_path, thumbnail_path,
mime_type, caption, like_count, comment_count, created_at
FROM v_feed
WHERE event_id = $1
AND ($2::timestamptz IS NULL OR (created_at, id) < ($2, $3))
ORDER BY created_at DESC, id DESC
LIMIT $4",
) )
.bind(auth.event_id) .bind(auth.event_id)
.bind(cursor_time) .bind(cursor_time)
.bind(cursor_id) .bind(cursor_id)
.bind(tags.as_deref())
.bind(uploader)
.bind(limit + 1) .bind(limit + 1)
.fetch_all(&state.pool) .fetch_all(&state.pool)
.await? .await?;
};
let has_more = rows.len() as i64 > limit; let has_more = rows.len() as i64 > limit;
let rows: Vec<FeedRow> = rows.into_iter().take(limit as usize).collect(); let rows: Vec<FeedRow> = rows.into_iter().take(limit as usize).collect();
@@ -154,6 +194,10 @@ pub async fn feed(
.thumbnail_path .thumbnail_path
.as_ref() .as_ref()
.map(|_| format!("/api/v1/upload/{}/thumbnail", r.id)); .map(|_| format!("/api/v1/upload/{}/thumbnail", r.id));
let display_url = r
.display_path
.as_ref()
.map(|_| format!("/api/v1/upload/{}/display", r.id));
FeedUpload { FeedUpload {
liked_by_me: liked_set.contains(&r.id), liked_by_me: liked_set.contains(&r.id),
id: r.id, id: r.id,
@@ -161,6 +205,7 @@ pub async fn feed(
uploader_name: r.uploader_name, uploader_name: r.uploader_name,
preview_url, preview_url,
thumbnail_url, thumbnail_url,
display_url,
mime_type: r.mime_type, mime_type: r.mime_type,
caption: r.caption, caption: r.caption,
like_count: r.like_count, like_count: r.like_count,
@@ -216,14 +261,14 @@ pub async fn feed_delta(
let feed_rate_on = config::get_bool(&state.config_cache, "feed_rate_enabled", true).await; let feed_rate_on = config::get_bool(&state.config_cache, "feed_rate_enabled", true).await;
if rate_limits_on && feed_rate_on { if rate_limits_on && feed_rate_on {
let rate_limit = config::get_usize(&state.config_cache, "feed_rate_per_min", 60).await; let rate_limit = config::get_usize(&state.config_cache, "feed_rate_per_min", 60).await;
if !state.rate_limiter.check( if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("feed_delta:{}", auth.user_id), format!("feed_delta:{}", auth.user_id),
rate_limit, rate_limit,
Duration::from_secs(60), Duration::from_secs(60),
) { ) {
return Err(AppError::TooManyRequests( return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(), "Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None, Some(retry_after_secs),
)); ));
} }
} }
@@ -247,7 +292,7 @@ pub async fn feed_delta(
// response's `server_time`, so this doesn't re-fetch on every subsequent delta. // response's `server_time`, so this doesn't re-fetch on every subsequent delta.
let rows = sqlx::query_as::<_, FeedRow>( let rows = sqlx::query_as::<_, FeedRow>(
"SELECT id, user_id, uploader_name, preview_path, thumbnail_path, "SELECT id, user_id, uploader_name, preview_path, thumbnail_path,
mime_type, caption, like_count, comment_count, created_at display_path, mime_type, caption, like_count, comment_count, created_at
FROM v_feed FROM v_feed
WHERE event_id = $1 AND created_at >= $2 WHERE event_id = $1 AND created_at >= $2
ORDER BY created_at DESC, id DESC ORDER BY created_at DESC, id DESC
@@ -304,6 +349,10 @@ pub async fn feed_delta(
.thumbnail_path .thumbnail_path
.as_ref() .as_ref()
.map(|_| format!("/api/v1/upload/{}/thumbnail", r.id)), .map(|_| format!("/api/v1/upload/{}/thumbnail", r.id)),
display_url: r
.display_path
.as_ref()
.map(|_| format!("/api/v1/upload/{}/display", r.id)),
mime_type: r.mime_type, mime_type: r.mime_type,
caption: r.caption, caption: r.caption,
like_count: r.like_count, like_count: r.like_count,
@@ -344,6 +393,32 @@ pub async fn hashtags(
)) ))
} }
/// Every uploader who has at least one visible upload, for the grid's "Nutzer suchen" picker.
///
/// The picker used to derive names from the uploads currently in memory — page 1, 20 items —
/// so typing a guest's name found nothing whenever their photos happened to sit below the
/// fold, which reads as "search is broken". This is the authoritative list.
///
/// Reads `v_feed`, so it inherits exactly the feed's visibility rules: soft-deleted uploads,
/// banned uploaders and hidden uploaders are all excluded, and a guest who has not uploaded
/// anything never appears. Uncapped on purpose — one short string per uploader, bounded by
/// the guest count, and truncating it would reintroduce the very bug this replaces.
/// Deliberately NOT the host-only `/host/users` route: that one lists every joined guest and
/// exposes moderation state.
pub async fn uploaders(
State(state): State<AppState>,
auth: AuthUser,
) -> Result<Json<Vec<String>>, AppError> {
let rows: Vec<(String,)> = sqlx::query_as(
"SELECT DISTINCT uploader_name FROM v_feed WHERE event_id = $1 ORDER BY uploader_name",
)
.bind(auth.event_id)
.fetch_all(&state.pool)
.await?;
Ok(Json(rows.into_iter().map(|(name,)| name).collect()))
}
/// Resolve a cursor id to its `(created_at, id)` position. Both are needed: /// Resolve a cursor id to its `(created_at, id)` position. Both are needed:
/// `created_at` alone isn't unique, so pagination must break ties on `id` to /// `created_at` alone isn't unique, so pagination must break ties on `id` to
/// avoid silently dropping rows that share a timestamp across a page boundary. /// avoid silently dropping rows that share a timestamp across a page boundary.
@@ -375,3 +450,41 @@ async fn get_liked_set(
rows.into_iter().map(|r| r.0).collect() rows.into_iter().map(|r| r.0).collect()
} }
#[cfg(test)]
mod tests {
use super::normalize_tags;
/// The chips carry display strings (`#Tanz`), the `hashtag` table stores `tanz`. If these
/// two drift the filter silently returns nothing, which is indistinguishable from "no
/// photos have this tag" — so pin the normalisation to `Hashtag::upsert`'s rule.
#[test]
fn tags_are_normalised_like_upsert_stores_them() {
assert_eq!(
normalize_tags(Some("#Tanz"), None),
Some(vec!["tanz".to_string()])
);
assert_eq!(
normalize_tags(None, Some(" #Buffet , reden ")),
Some(vec!["buffet".to_string(), "reden".to_string()])
);
}
/// An empty list must mean "no filter", never "match nothing" — returning `Some(vec![])`
/// would make `= ANY('{}')` false for every row and blank the feed.
#[test]
fn blank_input_disables_the_filter() {
assert_eq!(normalize_tags(None, None), None);
assert_eq!(normalize_tags(Some(" "), Some(" , ,#")), None);
}
/// Both params feed one list, de-duplicated: the list view sends `hashtag`, the grid sends
/// `hashtags`, and carrying a filter across views can legitimately set both to the same tag.
#[test]
fn single_and_csv_merge_without_duplicates() {
assert_eq!(
normalize_tags(Some("tanz"), Some("tanz,buffet")),
Some(vec!["tanz".to_string(), "buffet".to_string()])
);
}
}

View File

@@ -35,13 +35,67 @@ pub struct EventStatus {
pub is_active: bool, pub is_active: bool,
pub uploads_locked: bool, pub uploads_locked: bool,
pub export_released: bool, pub export_released: bool,
/// Free space on the volume the keepsake is written to. `None` when the mount can't be
/// resolved — the UI hides the widget rather than rendering a confident zero.
pub disk_free_bytes: Option<u64>,
/// What a full keepsake build would need right now (both halves).
pub keepsake_required_bytes: u64,
/// Whether the host should be warned. See [`disk_is_low`].
pub disk_low: bool,
}
/// Is free space low enough that the host needs to know?
///
/// Two triggers, because a fixed threshold answers the wrong question. `postgres_data`,
/// `media_data` and `exports_data` are all Docker named volumes on one filesystem, so a full disk
/// does not degrade one subsystem — it stops Postgres writing and takes the event down. That is
/// what the absolute floor is for.
///
/// The second trigger is the one that actually earns its place: the keepsake needs room for two
/// gallery-sized archives, and the only moment a host can do anything about that is BEFORE they
/// release. Warning at "you could not build the keepsake right now" turns a post-event dead end
/// into a decision someone can still make.
///
/// IT MUST FIRE BEFORE THE UPLOAD GATE CLOSES, and that is why the reserve and the margin are
/// here. The gate in `handlers::upload` refuses at
/// `free < keepsake_required + DISK_RESERVE_BYTES + UPLOAD_GATE_HEADROOM_BYTES`;
/// warning at `free < keepsake_required` alone meant the two differed by the whole reserve, so
/// the wall was always hit FIRST. Every guest would be blocked from uploading while this
/// dashboard showed a comfortable disk and no banner at all — on the shipped 40 GB box, uploads
/// stopping with ~27 GB free and nothing on screen to explain it, with no operator present.
///
/// The 25% margin makes it a warning rather than an obituary: the host sees it while there is
/// still room to act (delete a few large videos, which refunds immediately and reopens the gate).
fn disk_is_low(free: u64, keepsake_required: u64) -> bool {
// Mirrors the gate EXACTLY, headroom included. The gate now demands
// `UPLOAD_GATE_HEADROOM_BYTES` more than the export preflight does, so that ordinary
// end-of-night writes cannot flip the preflight after uploads have already stopped. Leaving
// that term out here would shrink the warning's lead by 1.5 GB — and the whole point of this
// function is that the banner must appear while the host can still act.
let gate_closes_at = keepsake_required
.saturating_add(crate::handlers::upload::DISK_RESERVE_BYTES as u64)
.saturating_add(crate::handlers::upload::UPLOAD_GATE_HEADROOM_BYTES as u64);
let warn_at = gate_closes_at.saturating_add(gate_closes_at / 4);
// No separate absolute-floor clause. There used to be `free < LOW_DISK_FLOOR_BYTES ||` here,
// and it was unreachable: `gate_closes_at` is at least DISK_RESERVE_BYTES, so `warn_at` is at
// least 1.25x it (12.5 GB) — always above the 10 GB floor. Two tests were named after that
// clause and neither could fail if it were deleted. Keeping dead code that tests claim to
// cover is worse than not having it.
free < warn_at
} }
/// Count non-banned hosts/admins in the event OTHER than `excluding` — the operators /// Count non-banned hosts/admins in the event OTHER than `excluding` — the operators
/// who would remain if `excluding` were demoted or banned. Used to enforce the "an event /// who would remain if `excluding` were demoted or banned. Used to enforce the "an event
/// always keeps at least one operator" floor. /// always keeps at least one operator" floor.
///
/// Takes a CONNECTION, not the pool, and every caller passes the same transaction it is about to
/// write in — after taking [`lock_operator_floor`]. Read on the pool beforehand, this count was a
/// snapshot that any concurrent operator-removing action could invalidate before the UPDATE landed:
/// an admin demoting host B while host A calls `DELETE /me` saw two independent checks each observe
/// the other still present, both commit, and the event end up with zero operators — which is not
/// recoverable from inside the app, since appointing an operator requires being one.
async fn remaining_operators( async fn remaining_operators(
state: &AppState, conn: &mut sqlx::PgConnection,
event_id: Uuid, event_id: Uuid,
excluding: Uuid, excluding: Uuid,
) -> Result<i64, AppError> { ) -> Result<i64, AppError> {
@@ -52,11 +106,36 @@ async fn remaining_operators(
) )
.bind(event_id) .bind(event_id)
.bind(excluding) .bind(excluding)
.fetch_one(&state.pool) .fetch_one(conn)
.await?; .await?;
Ok(count) Ok(count)
} }
/// Serialise every action that can remove an operator from an event.
///
/// The same key `me::delete_account` takes — namespace 4242, `hashtext(event_id)` — and it MUST
/// stay identical, or the two families of caller lock against nothing. An advisory lock is used
/// rather than a row lock because it is a separate lock space and so cannot join the
/// `event`/`user` row-lock graph that moderation traffic already traverses in both directions;
/// it is released automatically when the transaction ends.
///
/// **Call this FIRST in the transaction, before taking any row lock.** Being a separate lock space
/// means it cannot form a cycle *with itself*, not that ordering is free: all three callers go on
/// to lock `user` and `event` rows, so a caller that took those rows first and reached for this
/// lock afterwards would deadlock against one that did it the other way round. Postgres would
/// break the tie by killing one transaction with a 500. Every caller acquires it first; keep it
/// that way.
pub(crate) async fn lock_operator_floor(
conn: &mut sqlx::PgConnection,
event_id: Uuid,
) -> Result<(), AppError> {
sqlx::query("SELECT pg_advisory_xact_lock(4242, hashtext($1::text))")
.bind(event_id)
.execute(conn)
.await?;
Ok(())
}
#[derive(Deserialize)] #[derive(Deserialize)]
pub struct SetRoleRequest { pub struct SetRoleRequest {
pub role: String, pub role: String,
@@ -72,11 +151,29 @@ pub async fn get_event_status(
.await? .await?
.ok_or_else(|| AppError::NotFound("Event nicht gefunden.".into()))?; .ok_or_else(|| AppError::NotFound("Event nicht gefunden.".into()))?;
// Measured on the EXPORT volume, not the media one: that is where the cliff is, and it is a
// distinct mount point even when both are backed by the same filesystem. The cached reading is
// right here — this is advisory, polled on every dashboard load, and a 15s-stale number costs
// nothing (unlike the export preflight, which reads uncached because it is about to write).
let free = state
.disk_cache
.snapshot(&state.config.export_path)
.map(|d| d.free);
let keepsake_required_bytes =
crate::services::export::keepsake_space_required(&state.pool, event.id)
.await
.unwrap_or(0);
Ok(Json(EventStatus { Ok(Json(EventStatus {
name: event.name, name: event.name,
is_active: event.is_active, is_active: event.is_active,
uploads_locked: event.uploads_locked_at.is_some(), uploads_locked: event.uploads_locked_at.is_some(),
export_released: event.export_released_at.is_some(), export_released: event.export_released_at.is_some(),
disk_free_bytes: free,
keepsake_required_bytes,
// Unknown free space is NOT low. Fails open, exactly as the upload quota and the export
// preflight do: a scary banner on an unreadable mount would train the host to ignore it.
disk_low: free.is_some_and(|f| disk_is_low(f, keepsake_required_bytes)),
})) }))
} }
@@ -135,14 +232,6 @@ pub async fn ban_user(
)); ));
} }
// Floor: never leave the event with zero operators. Banning removes the target from
// the active-operator pool, so refuse if they're the last non-banned host/admin.
if target.0 == "host" && remaining_operators(&state, auth.event_id, user_id).await? == 0 {
return Err(AppError::BadRequest(
"Der letzte Host kann nicht gesperrt werden.".into(),
));
}
// Ban ALWAYS hides: a banned user's content is "gone" everywhere. The visibility // Ban ALWAYS hides: a banned user's content is "gone" everywhere. The visibility
// views/queries now also filter on `is_banned` (defense in depth), and we set // views/queries now also filter on `is_banned` (defense in depth), and we set
// `uploads_hidden` so the existing `user-hidden` live-eviction path fires too. The old // `uploads_hidden` so the existing `user-hidden` live-eviction path fires too. The old
@@ -157,6 +246,21 @@ pub async fn ban_user(
// //
// The ban and the keepsake invalidation are ONE transaction — see `host_delete_upload`. // The ban and the keepsake invalidation are ONE transaction — see `host_delete_upload`.
let mut tx = state.pool.begin().await?; let mut tx = state.pool.begin().await?;
// Floor: never leave the event with zero operators. Banning removes the target from the
// active-operator pool, so refuse if they're the last non-banned host/admin.
//
// INSIDE the transaction and behind the operator lock — see `remaining_operators`. Checked on
// the pool beforehand, this raced `set_role` and `DELETE /me` into an event with no operator.
if target.0 == "host" {
lock_operator_floor(&mut tx, auth.event_id).await?;
if remaining_operators(&mut tx, auth.event_id, user_id).await? == 0 {
return Err(AppError::BadRequest(
"Der letzte Host kann nicht gesperrt werden.".into(),
));
}
}
sqlx::query( sqlx::query(
"UPDATE \"user\" "UPDATE \"user\"
SET is_banned = TRUE, uploads_hidden = TRUE, uploads_hidden_at = NOW() SET is_banned = TRUE, uploads_hidden = TRUE, uploads_hidden_at = NOW()
@@ -198,6 +302,19 @@ pub async fn ban_user(
"host: ban_user" "host: ban_user"
); );
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"ban_user",
Some(user_id),
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
@@ -256,12 +373,40 @@ pub async fn unban_user(
start_regen(&state, r); start_regen(&state, r);
} }
// The exact mirror of `ban_user`'s `user-hidden`, and it was missing entirely: every open
// feed and the unattended projector kept the guest evicted until somebody reloaded the page
// by hand. Meanwhile the host's own confirm copy promises the photos "come back to the
// gallery, die Diashow und den Export" — so the one surface that would have shown the host
// their action had worked showed the opposite.
//
// Also the signal a banned guest's upload queue waits on: their queued photos parked with
// the blob intact rather than being purged (see `AppError::UserBanned`), and this is what
// releases them.
let _ = state.sse_tx.send(SseEvent::new(
"user-shown",
serde_json::json!({ "user_id": user_id }).to_string(),
));
tracing::info!( tracing::info!(
actor_user_id = %auth.user_id, actor_user_id = %auth.user_id,
target_user_id = %user_id, target_user_id = %user_id,
event_id = %auth.event_id, event_id = %auth.event_id,
"host: unban_user" "host: unban_user"
); );
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"unban_user",
Some(user_id),
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
@@ -302,6 +447,7 @@ pub async fn rebuild_export(
r.event_id, r.event_id,
r.event_name, r.event_name,
r.epoch, r.epoch,
state.config.comments_enabled,
std::time::Duration::ZERO, std::time::Duration::ZERO,
state.pool.clone(), state.pool.clone(),
state.config.media_path.clone(), state.config.media_path.clone(),
@@ -361,21 +507,26 @@ pub async fn set_role(
// Floor: demoting the last non-banned host/admin to guest would leave the event with // Floor: demoting the last non-banned host/admin to guest would leave the event with
// no operator. Refuse. // no operator. Refuse.
if new_role == "guest" //
&& target.0 == "host" // The check and the UPDATE are ONE transaction, behind the operator lock — see
&& remaining_operators(&state, auth.event_id, user_id).await? == 0 // `remaining_operators`. Split apart on the pool, this raced `ban_user` and `DELETE /me`.
{ let mut tx = state.pool.begin().await?;
if new_role == "guest" && target.0 == "host" {
lock_operator_floor(&mut tx, auth.event_id).await?;
if remaining_operators(&mut tx, auth.event_id, user_id).await? == 0 {
return Err(AppError::BadRequest( return Err(AppError::BadRequest(
"Der letzte Host kann nicht zum Gast gemacht werden.".into(), "Der letzte Host kann nicht zum Gast gemacht werden.".into(),
)); ));
} }
}
sqlx::query("UPDATE \"user\" SET role = $2::user_role WHERE id = $1 AND event_id = $3") sqlx::query("UPDATE \"user\" SET role = $2::user_role WHERE id = $1 AND event_id = $3")
.bind(user_id) .bind(user_id)
.bind(new_role) .bind(new_role)
.bind(auth.event_id) .bind(auth.event_id)
.execute(&state.pool) .execute(&mut *tx)
.await?; .await?;
tx.commit().await?;
tracing::info!( tracing::info!(
actor_user_id = %auth.user_id, actor_user_id = %auth.user_id,
target_user_id = %user_id, target_user_id = %user_id,
@@ -384,6 +535,19 @@ pub async fn set_role(
new_role, new_role,
"host: set_role" "host: set_role"
); );
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"set_role",
Some(user_id),
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
@@ -432,7 +596,7 @@ pub async fn reset_user_pin(
} }
let pin: String = format!("{:04}", rand::rng().random_range(0..10000u32)); let pin: String = format!("{:04}", rand::rng().random_range(0..10000u32));
let pin_hash = bcrypt::hash(&pin, 12).map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?; let pin_hash = crate::auth::handlers::hash_password(pin.clone(), 12).await?;
sqlx::query( sqlx::query(
"UPDATE \"user\" "UPDATE \"user\"
@@ -476,6 +640,19 @@ pub async fn reset_user_pin(
"host: reset_user_pin" "host: reset_user_pin"
); );
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"reset_pin",
Some(user_id),
None,
None,
)
.await;
Ok(Json(PinResetResponse { pin })) Ok(Json(PinResetResponse { pin }))
} }
@@ -549,10 +726,16 @@ pub fn start_regen(state: &AppState, regen: crate::services::export::PendingRege
regen.event_id, regen.event_id,
regen.event_name, regen.event_name,
regen.epoch, regen.epoch,
state.config.comments_enabled,
// Debounced: a takedown pass is a burst, and each request retires the last generation. The // Debounced: a takedown pass is a burst, and each request retires the last generation. The
// delay lets superseded workers fail their claim and do zero work instead of each building // delay lets superseded workers fail their claim and do zero work instead of each building
// a full archive. See export::REGEN_DEBOUNCE. // a full archive.
crate::services::export::REGEN_DEBOUNCE, //
// Measured from the START of the burst, not from this request — a fixed per-request delay
// meant a steady stream of invalidations faster than one per 20s deferred the build
// forever, leaving the keepsake permanently 404 and the UI stuck on "Wird vorbereitet…".
// See export::regen_delay_for.
crate::services::export::regen_delay_for(regen.event_id),
state.pool.clone(), state.pool.clone(),
state.config.media_path.clone(), state.config.media_path.clone(),
state.config.export_path.clone(), state.config.export_path.clone(),
@@ -573,7 +756,9 @@ pub async fn host_delete_upload(
// invalidation didn't, the taken-down photo would stay downloadable forever and nothing would // invalidation didn't, the taken-down photo would stay downloadable forever and nothing would
// notice (the keepsake still looks complete, and the host can no longer find the upload to retry). // notice (the keepsake still looks complete, and the host can no longer find the upload to retry).
let mut tx = state.pool.begin().await?; let mut tx = state.pool.begin().await?;
let deleted = Upload::soft_delete_in_event(&mut tx, upload_id, auth.event_id).await?; // `by_host: true` — the takedown holds the uploader's idempotency key so a late retry from
// their queue cannot resurrect the photo. See migration 031.
let deleted = Upload::soft_delete_in_event(&mut tx, upload_id, auth.event_id, true).await?;
if !deleted { if !deleted {
return Err(AppError::NotFound("Upload nicht gefunden.".into())); return Err(AppError::NotFound("Upload nicht gefunden.".into()));
} }
@@ -600,6 +785,19 @@ pub async fn host_delete_upload(
"host: host_delete_upload" "host: host_delete_upload"
); );
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"delete_upload",
Some(upload_id),
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
@@ -637,12 +835,25 @@ pub async fn host_delete_comment(
comment_id = %comment_id, comment_id = %comment_id,
"host: host_delete_comment" "host: host_delete_comment"
); );
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"delete_comment",
Some(comment_id),
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
pub async fn close_event( pub async fn close_event(
State(state): State<AppState>, State(state): State<AppState>,
RequireHost(_auth): RequireHost, RequireHost(auth): RequireHost,
) -> Result<StatusCode, AppError> { ) -> Result<StatusCode, AppError> {
let result = sqlx::query( let result = sqlx::query(
"UPDATE event SET uploads_locked_at = NOW() WHERE slug = $1 AND uploads_locked_at IS NULL", "UPDATE event SET uploads_locked_at = NOW() WHERE slug = $1 AND uploads_locked_at IS NULL",
@@ -657,12 +868,27 @@ pub async fn close_event(
let _ = state.sse_tx.send(SseEvent::new("event-closed", "{}")); let _ = state.sse_tx.send(SseEvent::new("event-closed", "{}"));
} }
// Was logged NOWHERE at all before this — not even a tracing line. A host reading
// the record the morning after had no way to see when uploads were locked or the
// gallery released, which are the two actions that change what every guest can do.
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"lock_uploads",
None,
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
pub async fn open_event( pub async fn open_event(
State(state): State<AppState>, State(state): State<AppState>,
RequireHost(_auth): RequireHost, RequireHost(auth): RequireHost,
) -> Result<StatusCode, AppError> { ) -> Result<StatusCode, AppError> {
// Reopening invalidates any prior release: the keepsake was snapshotted at release time, so // Reopening invalidates any prior release: the keepsake was snapshotted at release time, so
// allowing new uploads afterwards would silently diverge the live feed from the frozen export. // allowing new uploads afterwards would silently diverge the live feed from the frozen export.
@@ -688,12 +914,27 @@ pub async fn open_event(
let _ = state.sse_tx.send(SseEvent::new("event-opened", "{}")); let _ = state.sse_tx.send(SseEvent::new("event-opened", "{}"));
} }
// Was logged NOWHERE at all before this — not even a tracing line. A host reading
// the record the morning after had no way to see when uploads were locked or the
// gallery released, which are the two actions that change what every guest can do.
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"unlock_uploads",
None,
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
pub async fn release_gallery( pub async fn release_gallery(
State(state): State<AppState>, State(state): State<AppState>,
RequireHost(_auth): RequireHost, RequireHost(auth): RequireHost,
) -> Result<StatusCode, AppError> { ) -> Result<StatusCode, AppError> {
// The release claim, the epoch bump, the upload lock and BOTH job rows are written in ONE // The release claim, the epoch bump, the upload lock and BOTH job rows are written in ONE
// transaction. Two reasons, both of which were live bugs: // transaction. Two reasons, both of which were live bugs:
@@ -749,10 +990,25 @@ pub async fn release_gallery(
let _ = state.sse_tx.send(SseEvent::new("event-closed", "{}")); let _ = state.sse_tx.send(SseEvent::new("event-closed", "{}"));
// Detached — survives this handler being cancelled. // Detached — survives this handler being cancelled.
//
// SPAWNED IMMEDIATELY AFTER THE COMMIT, BEFORE ANY OTHER `.await`. Every `invalidate_and_arm`
// call site does this; `me::delete_account` carries the same note. The audit write below used
// to sit here, and it is two pool round-trips that can each wait up to the 5 s acquire timeout
// — right at the moment `event-closed` has just fanned out to ~100 phones whose queues all hit
// the API at once, so the pool is as contended as it ever gets. Drop the handler future during
// that suspension (the host's phone sleeps, the tab closes, Caddy times the request out) and
// the task never spawns: the event is released, uploads are locked, both `export_job` rows sit
// `pending` at the live epoch, and no worker exists. `/export/*` 404s, the page sits on "Wird
// vorbereitet…", `recover_exports` only runs at boot, and `release_gallery` refuses a retry
// because the gallery is already released.
//
// This is the one path that arms the FIRST build of the keepsake, so it is the worst possible
// place to reintroduce that window.
crate::services::export::spawn_export_jobs( crate::services::export::spawn_export_jobs(
event_id, event_id,
event_name, event_name,
epoch, epoch,
state.config.comments_enabled,
std::time::Duration::ZERO, std::time::Duration::ZERO,
state.pool.clone(), state.pool.clone(),
state.config.media_path.clone(), state.config.media_path.clone(),
@@ -760,5 +1016,159 @@ pub async fn release_gallery(
state.sse_tx.clone(), state.sse_tx.clone(),
); );
// Was logged NOWHERE at all before this — not even a tracing line. A host reading
// the record the morning after had no way to see when uploads were locked or the
// gallery released, which are the two actions that change what every guest can do.
//
// Last, deliberately: it is best-effort by design (it swallows its own errors), so nothing
// downstream may depend on it having completed.
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
None,
auth.role.clone(),
"release_gallery",
None,
None,
None,
)
.await;
Ok(StatusCode::NO_CONTENT) Ok(StatusCode::NO_CONTENT)
} }
#[cfg(test)]
mod tests {
use super::disk_is_low;
use crate::handlers::upload::{DISK_RESERVE_BYTES, UPLOAD_GATE_HEADROOM_BYTES};
use crate::services::export::required_free_bytes;
const GB: u64 = 1_000_000_000;
#[test]
fn a_healthy_disk_with_room_for_the_keepsake_is_not_low() {
// Room for the keepsake AND the reserve the upload gate holds back, with margin.
assert!(!disk_is_low(60 * GB, 25 * GB));
}
/// Renamed from `the_absolute_floor_fires_...`: there is no separate floor clause any more
/// (see `disk_is_low`). What still has to hold is the behaviour the floor was there FOR — a
/// nearly-empty disk is low even when the gallery is small enough that the keepsake term
/// alone would clear it, because all three volumes share one filesystem and Postgres needs
/// room to write.
#[test]
fn a_nearly_empty_disk_is_low_even_when_the_gallery_is_tiny() {
// All three volumes share one filesystem, so running out doesn't degrade one subsystem —
// Postgres stops being able to write and the event goes down. A 1 GB gallery would clear
// the keepsake test comfortably; the floor is what catches this.
assert!(disk_is_low(5 * GB, GB));
assert!(disk_is_low(9 * GB, 0));
assert!(!disk_is_low(20 * GB, 0), "a roomy empty disk is not low");
}
/// THE PROPERTY THIS EXISTS FOR: the host must be warned BEFORE guests are blocked.
///
/// `handlers::upload` refuses at `free < keepsake_required + DISK_RESERVE_BYTES`. If the
/// banner fires only at or below that, the host's first signal is 100 guests being unable
/// to upload while the dashboard shows a comfortable disk — with nobody on site to ask.
#[test]
fn the_banner_always_fires_before_the_upload_gate_closes() {
// Asserting `disk_is_low(gate_closes_at, required)` is what this used to do, and it was a
// tautology: `disk_is_low` recomputes the same `gate_closes_at` internally and compares
// against `gate + gate/4`, so the assertion reduced to `G < G + G/4` — true for every G,
// for any margin, even a margin of zero. It could not detect the banner being moved to
// exactly the gate, which is the regression it is named for.
//
// So pin the GAP instead: find the free-space level at which the banner starts, and
// require it to be strictly above the level at which the gate closes, by a usable amount.
for media_gb in [0u64, 1, 4, 8, 16, 32] {
let required = required_free_bytes(media_gb * GB, 2);
let gate_closes_at =
required + DISK_RESERVE_BYTES as u64 + UPLOAD_GATE_HEADROOM_BYTES as u64;
// Just above the gate: guests can still upload, and the host must already be warned.
assert!(
disk_is_low(gate_closes_at + 1, required),
"at media={media_gb}GB the banner is not yet showing while the gate still allows uploads"
);
// The warning must lead by a margin the host can act inside, not by one byte.
let mut warn_starts_at = gate_closes_at;
while disk_is_low(warn_starts_at, required) {
warn_starts_at += GB / 10;
}
assert!(
warn_starts_at >= gate_closes_at + gate_closes_at / 5,
"at media={media_gb}GB the banner leads the gate by only {} bytes",
warn_starts_at - gate_closes_at
);
}
}
/// The invariant the headroom exists for: uploads must stop while the keepsake can STILL be
/// built, with room to spare — not at the exact instant the preflight reaches its own limit.
///
/// Both thresholds used to be `required_free_bytes(media, 2) + DISK_RESERVE_BYTES`, identically.
/// So the moment the gate refused its first upload, the export preflight was already sitting on
/// its limit, and every byte written afterwards (WAL, container logs, the compression backlog
/// draining at exactly that hour) pushed it under. The release would then COMMIT — event closed,
/// uploads locked, epoch bumped, `event-closed` fanned out to every phone — and only then would
/// both workers bail, with no second release possible.
#[test]
fn the_upload_gate_closes_before_the_export_preflight_would_refuse() {
for media_gb in [0u64, 1, 4, 8, 16, 32] {
let required = required_free_bytes(media_gb * GB, 2);
// `services::export::preflight` bails below this.
let preflight_refuses_below = required + DISK_RESERVE_BYTES as u64;
// `handlers::upload` refuses below this.
let gate_refuses_below = preflight_refuses_below + UPLOAD_GATE_HEADROOM_BYTES as u64;
assert!(
gate_refuses_below > preflight_refuses_below,
"at media={media_gb}GB the gate and the preflight share a threshold, so the \
keepsake's fate rests on whatever is written after uploads stop"
);
// At the instant the last upload is refused, the preflight must still pass with the
// whole headroom to spare — that is the slack the night's remaining writes consume.
let free_when_gate_closes = gate_refuses_below;
assert!(
free_when_gate_closes
>= preflight_refuses_below + UPLOAD_GATE_HEADROOM_BYTES as u64,
"at media={media_gb}GB there is no slack between the gate closing and the \
preflight failing"
);
}
}
#[test]
fn plenty_of_space_is_still_low_when_the_keepsake_would_not_fit() {
// THE case the fixed threshold misses, and the one that matters: 30 GB free is nowhere near
// any floor, but a 30 GB gallery needs room for TWO archives. The host can act on this
// before releasing; after releasing, they cannot.
assert!(disk_is_low(30 * GB, 66 * GB));
}
#[test]
fn the_keepsake_trigger_is_exact_at_the_boundary() {
// The boundary is the UPLOAD GATE's threshold plus a 25% margin, not the bare keepsake
// size — see `disk_is_low`. Warning at the bare size fired only after the gate had
// already blocked every guest.
let required = 20 * GB;
let gate = required + DISK_RESERVE_BYTES as u64 + UPLOAD_GATE_HEADROOM_BYTES as u64;
let warn_at = gate + gate / 4;
assert!(!disk_is_low(warn_at, required), "exactly enough is enough");
assert!(disk_is_low(warn_at - 1, required));
}
#[test]
fn an_empty_gallery_still_reserves_room_for_postgres() {
// With no gallery the keepsake term is 0, so the warn threshold collapses to
// 1.25 x (DISK_RESERVE_BYTES + UPLOAD_GATE_HEADROOM_BYTES) = 1.25 x 11.5 GB = 14.375 GB,
// which dominates the 10 GB absolute floor.
assert!(!disk_is_low(15 * GB, 0));
assert!(disk_is_low(9 * GB, 0));
}
}

View File

@@ -14,7 +14,7 @@ use serde::Serialize;
use crate::auth::middleware::AuthUser; use crate::auth::middleware::AuthUser;
use crate::error::AppError; use crate::error::AppError;
use crate::handlers::upload::compute_storage_quota; use crate::handlers::upload::compute_storage_quota;
use crate::models::user::User; use crate::models::user::{User, UserRole};
use crate::services::config; use crate::services::config;
use crate::state::AppState; use crate::state::AppState;
@@ -37,12 +37,26 @@ pub async fn get_quota(
let estimate = compute_storage_quota(&state).await; let estimate = compute_storage_quota(&state).await;
// Raw server telemetry (free disk, concurrent uploader count) is staff-only — it
// must never reach a guest, even though the guest upload UI no longer renders it.
// A guest still gets their own `used`/`limit` so enforcement stays transparent to
// the code paths that consume it; only the server-wide fields are zeroed.
let is_staff = matches!(auth.role, UserRole::Host | UserRole::Admin);
Ok(Json(QuotaDto { Ok(Json(QuotaDto {
enabled: estimate.limit_bytes.is_some(), enabled: estimate.limit_bytes.is_some(),
used_bytes: user.total_upload_bytes, used_bytes: user.total_upload_bytes,
limit_bytes: estimate.limit_bytes, limit_bytes: estimate.limit_bytes,
active_uploaders: estimate.active_uploaders, active_uploaders: if is_staff {
free_disk_bytes: estimate.free_disk_bytes, estimate.active_uploaders
} else {
0
},
free_disk_bytes: if is_staff {
estimate.free_disk_bytes
} else {
0
},
})) }))
} }
@@ -61,6 +75,14 @@ pub struct MeContextDto {
/// The gallery has been released and the export snapshotted — uploads are permanently /// The gallery has been released and the export snapshotted — uploads are permanently
/// closed for this run (release ⇒ lock, and reopening regenerates). /// closed for this run (release ⇒ lock, and reopening regenerates).
pub gallery_released: bool, pub gallery_released: bool,
/// This guest is banned: a deliberately READ-ONLY ban (see `handlers/host.rs`) — they keep
/// the feed and the keepsake, but every write is refused.
///
/// Exposed so the UI can SAY so. Without it the client had no idea, so the upload button,
/// the like button and "Löschen" all rendered enabled and returned 403 "Du bist gesperrt."
/// on every tap — a guest tapping upload repeatedly with nobody to ask. The lock case
/// (`uploads_locked`) has always been surfaced for exactly this reason; a ban was not.
pub is_banned: bool,
} }
pub async fn get_context( pub async fn get_context(
@@ -96,5 +118,203 @@ pub async fn get_context(
storage_quota_enabled, storage_quota_enabled,
uploads_locked, uploads_locked,
gallery_released, gallery_released,
is_banned: user.is_banned,
})) }))
} }
/// `(original_path, preview_path, thumbnail_path, display_path)` for one upload.
type UploadFilePaths = (String, Option<String>, Option<String>, Option<String>);
/// Delete the caller's own account and everything attached to it.
///
/// The erasure path (H18). There was no user-deletion route at ANY role, so honouring a "please
/// remove my photos and my name" request meant hand-written SQL against production — during or
/// after a wedding, by whoever happened to have psql access. Deletion also never removed text:
/// captions, comment bodies and hashtag links survived indefinitely by design, so even the
/// existing per-photo delete left the guest's words in the database and in the keepsake.
///
/// Self-service on purpose. The alternative (host-initiated only) puts a guest's erasure request
/// through a third party who is at a party, and the join page's data notice now promises this.
///
/// ORDER MATTERS. `upload.user_id` and `comment.user_id` are plain FKs with NO `ON DELETE CASCADE`
/// (migration 002), so deleting the user first fails on a constraint violation. Children first,
/// then the row itself — at which point `session`, `like` and `pin_reset_request` do cascade.
pub async fn delete_account(
State(state): State<AppState>,
auth: AuthUser,
) -> Result<axum::http::StatusCode, AppError> {
// The last host/admin may not erase themselves: it would leave the event with no operator and
// no way to appoint one. Mirrors the floor `set_role` and `ban_user` already enforce.
let user = User::find_by_id(&state.pool, auth.user_id)
.await?
.ok_or_else(|| AppError::NotFound("Benutzer nicht gefunden.".into()))?;
if matches!(user.role, UserRole::Host | UserRole::Admin) {
let others = sqlx::query_scalar::<_, i64>(
"SELECT COUNT(*) FROM \"user\"
WHERE event_id = $1 AND id != $2
AND role IN ('host', 'admin') AND is_banned = FALSE",
)
.bind(auth.event_id)
.bind(auth.user_id)
.fetch_one(&state.pool)
.await?;
if others == 0 {
return Err(AppError::BadRequest(
"Du bist der letzte Gastgeber. Ernenne zuerst einen anderen Gastgeber, bevor du \
dein Konto löschst."
.into(),
));
}
}
// Collect the file paths BEFORE the rows go, or they are unrecoverable. Every derivative, not
// just the original: a preview left behind is still the guest's photo.
let files: Vec<UploadFilePaths> = sqlx::query_as(
"SELECT original_path, preview_path, thumbnail_path, display_path
FROM upload WHERE user_id = $1",
)
.bind(auth.user_id)
.fetch_all(&state.pool)
.await?;
let mut tx = state.pool.begin().await?;
// The last-host guard, AUTHORITATIVELY — inside the transaction, holding a lock.
//
// The pre-check further up runs on the pool before this transaction opens, so two hosts
// deleting themselves at the same moment each saw the other and both proceeded, leaving the
// event with NO operator: nobody to moderate, nobody to release the gallery, and no way to
// appoint anyone because appointing requires a host. Not recoverable from inside the app.
//
// Serialised with a transaction-scoped ADVISORY lock, not a row lock. `FOR UPDATE` on the
// other operators\' rows looks like the obvious answer and is the wrong one: each deleter would
// lock the OTHER\'s row and then try to delete its own, so the two block on each other and
// Postgres resolves it by killing one with a deadlock error — the invariant holds, but the
// loser gets a 500 instead of the sentence below. Locking the `event` row instead would
// serialise cleanly, but it inverts the lock order every moderation path uses (upload/user
// rows first, event last). An advisory lock is a separate lock space, so it cannot join the
// row-lock graph at all, and it is released automatically when this transaction ends.
//
// FIRST STATEMENT IN THE TRANSACTION, before any row lock — the ORDER matters as much as the
// lock. `ban_user` and `set_role` take this same lock and then go on to lock `user` and
// `event` rows. If this path grabbed those rows first and reached for the advisory lock
// afterwards, the two would deadlock, each holding what the other needs, and Postgres would
// kill one with a 500: the invariant would survive, but a host deleting their account would
// get an error page instead of the sentence below.
//
// Taking it up front also means the refusal path does no work at all before answering.
if matches!(user.role, UserRole::Host | UserRole::Admin) {
// Shared with `host::ban_user` and `host::set_role` — the same key, by construction rather
// than by two copies agreeing. All three remove an operator, so all three must serialise
// against each other or the floor is enforceable only against its own kind of caller.
crate::handlers::host::lock_operator_floor(&mut tx, auth.event_id).await?;
let others: Vec<uuid::Uuid> = sqlx::query_scalar(
"SELECT id FROM \"user\"
WHERE event_id = $1 AND id != $2
AND role IN ('host', 'admin') AND is_banned = FALSE",
)
.bind(auth.event_id)
.bind(auth.user_id)
.fetch_all(&mut *tx)
.await?;
if others.is_empty() {
return Err(AppError::BadRequest(
"Du bist der letzte Gastgeber. Ernenne zuerst einen anderen Gastgeber, bevor du \
dein Konto löschst."
.into(),
));
}
}
// Comments the guest wrote on OTHER people's photos. Hard delete, not `deleted_at`: this is
// erasure, and a soft delete leaves the body in the table and in the keepsake's data.json.
sqlx::query("DELETE FROM comment WHERE user_id = $1")
.bind(auth.user_id)
.execute(&mut *tx)
.await?;
// Their uploads. Cascades comments and likes ON those uploads, plus upload_hashtag links.
sqlx::query("DELETE FROM upload WHERE user_id = $1")
.bind(auth.user_id)
.execute(&mut *tx)
.await?;
// Invalidate the keepsake inside the same transaction — an already-released archive still
// contains this guest's photos and captions, and erasure that leaves them in the downloadable
// ZIP has not happened. Returns None when the event isn't released, in which case there is
// nothing to rebuild.
let regen = crate::services::export::invalidate_and_arm(
&mut tx,
&state.config.event_slug,
crate::services::export::Affects::Both,
)
.await?;
// And the account. `session`, `like` and `pin_reset_request` cascade from here.
sqlx::query("DELETE FROM \"user\" WHERE id = $1")
.bind(auth.user_id)
.execute(&mut *tx)
.await?;
tx.commit().await?;
// IMMEDIATELY after the commit, before any other `.await`. Every other `invalidate_and_arm`
// call site does this; this one used to spawn the workers *after* the file-removal loop below,
// and axum drops a handler future the moment the client disconnects. Drop it inside that loop
// and the keepsake is left with the epoch bumped, both `export_job` rows armed `pending` at
// that epoch, and NO WORKER: `/export/zip` and `/export/html` 404, the UI sits on
// "Wird vorbereitet…" forever, and `recover_exports` only runs at boot. Deleting your account
// from a phone that walks out of wifi range is enough to do it.
if let Some(r) = regen {
crate::handlers::host::start_regen(&state, r);
}
// Best effort, after the commit. Anything missed here is an orphan with no row pointing at it,
// which `sweep_orphan_originals` reclaims on its next pass — so a failure delays reclamation
// rather than leaving the file referenced.
for (original, preview, thumbnail, display) in &files {
for rel in [
Some(original),
preview.as_ref(),
thumbnail.as_ref(),
display.as_ref(),
]
.into_iter()
.flatten()
{
let abs = state.config.media_path.join(rel);
if let Err(e) = tokio::fs::remove_file(&abs).await
&& e.kind() != std::io::ErrorKind::NotFound
{
tracing::warn!(error = ?e, path = %abs.display(), "account deletion: could not remove media file");
}
}
}
// Evict their content from every open feed and the projector. `user-hidden` is exactly the
// right signal — it already means "this user's cards must go" — and reusing it means every
// client already handles this with no new event type.
let _ = state.sse_tx.send(crate::state::SseEvent::new(
"user-hidden",
serde_json::json!({ "user_id": auth.user_id }).to_string(),
));
// Audited like the host actions it resembles, with the actor and target being the same person.
//
// The names are passed EXPLICITLY here, unlike every other call site. `audit::record` resolves
// a missing name by looking the id up in `"user"` — and this handler has just hard-deleted that
// row, so the lookup would find nothing and write the NULL that makes the record unreadable.
// This is the row most likely to be read later ("whose photos disappeared?"), and migration 029
// made these columns non-FK precisely so it would survive the deletion.
crate::services::audit::record(
&state.pool,
auth.event_id,
auth.user_id,
Some(&user.display_name),
user.role.clone(),
"delete_account",
Some(auth.user_id),
Some(&user.display_name),
Some(serde_json::json!({ "uploads_removed": files.len() })),
)
.await;
tracing::info!(user_id = %auth.user_id, uploads = files.len(), "account deleted by its owner");
Ok(axum::http::StatusCode::NO_CONTENT)
}

View File

@@ -4,21 +4,51 @@ use axum::Json;
use axum::extract::State; use axum::extract::State;
use serde::Serialize; use serde::Serialize;
use crate::services::config;
use crate::state::AppState; use crate::state::AppState;
#[derive(Serialize)] #[derive(Serialize)]
pub struct PublicEventDto { pub struct PublicEventDto {
pub name: String, pub name: String,
pub slug: String, pub slug: String,
/// Whether the comment feature is on (env `COMMENTS_ENABLED`). The frontend hides
/// the whole comment UI when false; exposed here so even the pre-auth shell knows.
pub comments_enabled: bool,
/// Active colour theme. `preset` is an id the frontend maps to a palette (or
/// "custom"); `primary`/`accent` are the `#rrggbb` seeds the ramps derive from.
/// Resolved as DB-config override → env default. Public so the theme applies on
/// the very first (pre-auth) paint without a flash.
pub theme_preset: String,
pub theme_primary: String,
pub theme_accent: String,
/// The operator's data notice, if they set one. Empty string when unset (migration 009
/// defaults it to `''`).
///
/// Exposed PUBLICLY — it was only on `/me/context`, which requires a token, so the one place a
/// notice actually has to appear (before a name is collected) could not read it. The join page
/// pairs this with a baseline notice of its own, precisely because this can be empty: relying
/// on an operator-supplied string meant a stock deploy collected ~100 EU guests' photos of
/// identifiable people, including children, with no notice at the point of collection at all.
pub privacy_note: String,
} }
/// Public event identity, used by the pre-auth join/recover screens so a guest can /// Public event identity + presentation config, used by the pre-auth join/recover
/// see *which* event they're joining. Only the display name and slug are exposed — /// screens (which event am I joining, what does it look like). Only non-user-scoped
/// nothing user-scoped — so this is safe without a token. Served straight from the /// fields are exposed, so this is safe without a token. Identity comes straight from
/// instance config (no DB round-trip needed). /// instance config; the theme is resolved from the runtime `config` table (admin UI)
/// falling back to the env-seeded default.
pub async fn get_public_event(State(state): State<AppState>) -> Json<PublicEventDto> { pub async fn get_public_event(State(state): State<AppState>) -> Json<PublicEventDto> {
let cache = &state.config_cache;
Json(PublicEventDto { Json(PublicEventDto {
name: state.config.event_name.clone(), name: state.config.event_name.clone(),
slug: state.config.event_slug.clone(), slug: state.config.event_slug.clone(),
comments_enabled: state.config.comments_enabled,
theme_preset: config::get_str(cache, "theme_preset", &state.config.default_theme_preset)
.await,
theme_primary: config::get_str(cache, "theme_primary", &state.config.default_theme_primary)
.await,
theme_accent: config::get_str(cache, "theme_accent", &state.config.default_theme_accent)
.await,
privacy_note: config::get_str(cache, "privacy_note", "").await,
}) })
} }

View File

@@ -10,8 +10,40 @@ use crate::error::AppError;
use crate::models::comment::{Comment, CommentDto}; use crate::models::comment::{Comment, CommentDto};
use crate::models::hashtag::{self, Hashtag}; use crate::models::hashtag::{self, Hashtag};
use crate::models::upload::Upload; use crate::models::upload::Upload;
use crate::services::config;
use crate::state::AppState; use crate::state::AppState;
/// Throttle a social write. Keyed PER USER, like the feed and upload limits and for the same
/// reason: at a venue every guest sits behind one NAT, so an IP key hands the whole party a
/// single bucket and the most active guest starves everyone else.
///
/// These were the only mutating endpoints in the app with no limit at all — the coverage was
/// asymmetric, not deliberately open. The ceiling is set well above anything a real guest
/// produces; this bounds a script, not an enthusiastic double-tapper.
async fn check_social_rate(state: &AppState, user_id: Uuid) -> Result<(), AppError> {
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let social_rate_on = config::get_bool(&state.config_cache, "social_rate_enabled", true).await;
if !(rate_limits_on && social_rate_on) {
return Ok(());
}
let rate_limit = config::get_usize(&state.config_cache, "social_rate_per_min", 120).await;
// ONE bucket across likes, comments and comment deletions. Separate buckets would let a
// caller triple the aggregate write rate just by alternating between them.
state
.rate_limiter
.check_with_retry(
format!("social:{user_id}"),
rate_limit,
std::time::Duration::from_secs(60),
)
.map_err(|retry_after_secs| {
AppError::TooManyRequests(
"Zu viele Aktionen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
)
})
}
#[derive(Serialize)] #[derive(Serialize)]
pub struct LikeResponse { pub struct LikeResponse {
/// The caller's like state *after* this toggle. The client sets `liked_by_me` from /// The caller's like state *after* this toggle. The client sets `liked_by_me` from
@@ -35,6 +67,7 @@ pub async fn toggle_like(
if user.is_banned { if user.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into())); return Err(AppError::Forbidden("Du bist gesperrt.".into()));
} }
check_social_rate(&state, auth.user_id).await?;
// Event-scope: the upload must belong to the caller's event (404 otherwise), // Event-scope: the upload must belong to the caller's event (404 otherwise),
// matching the host handlers' find_by_id_and_event pattern. // matching the host handlers' find_by_id_and_event pattern.
@@ -71,8 +104,17 @@ pub async fn toggle_like(
// itself is already committed, so a failed count must not fail the request — but we // itself is already committed, so a failed count must not fail the request — but we
// also must NOT broadcast/return a bogus 0 (that would push like_count: 0 to every // also must NOT broadcast/return a bogus 0 (that would push like_count: 0 to every
// client until the next event). On error we skip the broadcast and return null. // client until the next event). On error we skip the broadcast and return null.
// The `NOT u.is_banned` join is what makes "mirrors v_feed.like_count" true. Migration 028
// added it to the view and not here, so the two disagreed the moment anyone was banned: the
// host bans a guest, the feed correctly drops to the lower number, and then the very next like
// on that photo broadcasts the UNFILTERED count back to every open client — including the
// host's, who is watching that number to confirm the ban took. It stayed wrong until a full
// page-1 refetch. `like.user_id` is NOT NULL REFERENCES "user"(id), so the inner join can
// neither drop nor duplicate a row.
let like_count = sqlx::query_scalar::<_, i64>( let like_count = sqlx::query_scalar::<_, i64>(
"SELECT COUNT(DISTINCT user_id) FROM \"like\" WHERE upload_id = $1", "SELECT COUNT(DISTINCT l.user_id) FROM \"like\" l \
JOIN \"user\" u ON u.id = l.user_id \
WHERE l.upload_id = $1 AND NOT u.is_banned",
) )
.bind(upload_id) .bind(upload_id)
.fetch_one(&state.pool) .fetch_one(&state.pool)
@@ -129,12 +171,19 @@ pub async fn add_comment(
Path(upload_id): Path<Uuid>, Path(upload_id): Path<Uuid>,
Json(body): Json<AddCommentRequest>, Json(body): Json<AddCommentRequest>,
) -> Result<(StatusCode, Json<CommentDto>), AppError> { ) -> Result<(StatusCode, Json<CommentDto>), AppError> {
// Comments can be disabled instance-wide (env COMMENTS_ENABLED). The frontend hides
// the UI, but gate the API too so a stale client or direct call can't slip one in.
if !state.config.comments_enabled {
return Err(AppError::Forbidden("Kommentare sind deaktiviert.".into()));
}
let user = crate::models::user::User::find_by_id(&state.pool, auth.user_id) let user = crate::models::user::User::find_by_id(&state.pool, auth.user_id)
.await? .await?
.ok_or_else(|| AppError::NotFound("Benutzer nicht gefunden.".into()))?; .ok_or_else(|| AppError::NotFound("Benutzer nicht gefunden.".into()))?;
if user.is_banned { if user.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into())); return Err(AppError::Forbidden("Du bist gesperrt.".into()));
} }
check_social_rate(&state, auth.user_id).await?;
// Event-scope: only comment on an upload that belongs to the caller's event. // Event-scope: only comment on an upload that belongs to the caller's event.
Upload::find_by_id_and_event(&state.pool, upload_id, auth.event_id) Upload::find_by_id_and_event(&state.pool, upload_id, auth.event_id)
@@ -155,7 +204,14 @@ pub async fn add_comment(
// Insert the comment and link its hashtags atomically, so a crash mid-loop // Insert the comment and link its hashtags atomically, so a crash mid-loop
// can't leave a committed comment with only some of its tags indexed. // can't leave a committed comment with only some of its tags indexed.
let tags = hashtag::extract_hashtags(text); let mut tags = hashtag::extract_hashtags(text);
// Deterministic lock order, matching the upload path. `Hashtag::upsert` takes row locks,
// so two transactions touching the same two tags in OPPOSITE order deadlock — Postgres
// aborts one after ~1s and that guest's comment 500s. `extract_hashtags` returns them in
// text order, which is exactly the unordered case. Sort on the NORMALISED form, because
// that is the key `upsert` locks on.
tags.sort_by_key(|t| t.trim().trim_start_matches('#').to_lowercase());
tags.dedup_by_key(|t| t.trim().trim_start_matches('#').to_lowercase());
let mut tx = state.pool.begin().await?; let mut tx = state.pool.begin().await?;
let comment = Comment::create(&mut *tx, upload_id, auth.user_id, text).await?; let comment = Comment::create(&mut *tx, upload_id, auth.user_id, text).await?;
for tag in &tags { for tag in &tags {
@@ -175,8 +231,13 @@ pub async fn add_comment(
// over the same deleted_at filter is identical since comment.id is the PK). The // over the same deleted_at filter is identical since comment.id is the PK). The
// count + broadcast are a UI optimisation — the comment is already committed, so a // count + broadcast are a UI optimisation — the comment is already committed, so a
// failure here must not fail the request. Swallow the error and skip the broadcast. // failure here must not fail the request. Swallow the error and skip the broadcast.
// `NOT u.is_banned` for the same reason as `like_count` above — see that comment. Migration
// 028 put this filter in `v_feed.comment_count` and `Comment::list_for_upload`, but not here,
// so posting a comment pushed the pre-ban total back to every client.
if let Ok(comment_count) = sqlx::query_scalar::<_, i64>( if let Ok(comment_count) = sqlx::query_scalar::<_, i64>(
"SELECT COUNT(*) FROM comment WHERE upload_id = $1 AND deleted_at IS NULL", "SELECT COUNT(*) FROM comment c \
JOIN \"user\" u ON u.id = c.user_id \
WHERE c.upload_id = $1 AND c.deleted_at IS NULL AND NOT u.is_banned",
) )
.bind(upload_id) .bind(upload_id)
.fetch_one(&state.pool) .fetch_one(&state.pool)
@@ -210,6 +271,7 @@ pub async fn delete_comment(
if auth.is_banned { if auth.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into())); return Err(AppError::Forbidden("Du bist gesperrt.".into()));
} }
check_social_rate(&state, auth.user_id).await?;
let comment = Comment::find_by_id(&state.pool, comment_id) let comment = Comment::find_by_id(&state.pool, comment_id)
.await? .await?
.ok_or_else(|| AppError::NotFound("Kommentar nicht gefunden.".into()))?; .ok_or_else(|| AppError::NotFound("Kommentar nicht gefunden.".into()))?;

View File

@@ -13,6 +13,7 @@ use tokio_stream::wrappers::errors::BroadcastStreamRecvError;
use crate::auth::middleware::AuthUser; use crate::auth::middleware::AuthUser;
use crate::error::AppError; use crate::error::AppError;
use crate::models::session::Session; use crate::models::session::Session;
use crate::services::sse_tickets::TicketKind;
use crate::state::AppState; use crate::state::AppState;
#[derive(Deserialize)] #[derive(Deserialize)]
@@ -38,7 +39,29 @@ pub async fn issue_ticket(
State(state): State<AppState>, State(state): State<AppState>,
auth: AuthUser, auth: AuthUser,
) -> Result<Json<StreamTicketResponse>, AppError> { ) -> Result<Json<StreamTicketResponse>, AppError> {
let ticket = state.sse_tickets.issue(auth.token_hash); // The endpoint had no rate limit at all. Authentication is not a bound here: one valid
// session could loop it freely. 60/min is far above a real client (one ticket per SSE
// (re)connect, and reconnects are backed off) while capping a loop.
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("sse_ticket:{}", auth.user_id),
60,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Verbindungsversuche. Bitte warte kurz.".into(),
Some(retry_after_secs),
));
}
let ticket = state
.sse_tickets
.issue(auth.token_hash, TicketKind::Sse)
.ok_or_else(|| {
AppError::ServiceUnavailable(
"Server ist gerade ausgelastet. Live-Updates folgen in Kürze.".into(),
Some(30),
)
})?;
let server_time = sqlx::query_scalar("SELECT NOW()") let server_time = sqlx::query_scalar("SELECT NOW()")
.fetch_one(&state.pool) .fetch_one(&state.pool)
.await?; .await?;
@@ -48,6 +71,57 @@ pub async fn issue_ticket(
})) }))
} }
/// Live SSE streams one session may hold OPEN at once.
///
/// The ticket store's `MAX_TICKETS_PER_SESSION` bounds UNCONSUMED tickets, not open streams — so it
/// never bounded this at all: mint a ticket, redeem it (freeing the slot), repeat. At the 60/min
/// ticket ceiling one guest could accumulate 60 new live streams per minute indefinitely, each
/// holding a broadcast receiver, a tokio task and a 60-second DB revalidation ticker.
///
/// 6 rather than 2: a guest legitimately has the feed in one tab, the diashow on a laptop, and both
/// may briefly double during a reconnect before the old socket's `Drop` lands. Well above real use,
/// far below anything that hurts.
const MAX_OPEN_STREAMS_PER_SESSION: usize = 6;
/// Open stream count per session token hash.
type OpenStreams = std::collections::HashMap<String, usize>;
static OPEN_STREAMS: std::sync::LazyLock<std::sync::Mutex<OpenStreams>> =
std::sync::LazyLock::new(|| std::sync::Mutex::new(OpenStreams::new()));
/// Decrements the open-stream count for its session when the stream is dropped.
///
/// A `Drop` guard is the only thing that works here: a client vanishing off wifi never runs any
/// cleanup path we write, but dropping the response future is exactly what happens.
struct StreamSlot(String);
impl Drop for StreamSlot {
fn drop(&mut self) {
if let Ok(mut map) = OPEN_STREAMS.lock()
&& let Some(n) = map.get_mut(&self.0)
{
*n = n.saturating_sub(1);
if *n == 0 {
map.remove(&self.0);
}
}
}
}
/// Claim one of this session's stream slots, or `None` when it is already at the cap.
fn claim_stream_slot(token_hash: &str) -> Option<StreamSlot> {
let mut map = match OPEN_STREAMS.lock() {
Ok(m) => m,
// Never let a poisoned lock take live updates down for the whole venue.
Err(e) => e.into_inner(),
};
let n = map.entry(token_hash.to_string()).or_insert(0);
if *n >= MAX_OPEN_STREAMS_PER_SESSION {
return None;
}
*n += 1;
Some(StreamSlot(token_hash.to_string()))
}
/// SSE stream endpoint. Authenticates via a single-use ticket (see /// SSE stream endpoint. Authenticates via a single-use ticket (see
/// [`issue_ticket`]) — never the raw JWT. /// [`issue_ticket`]) — never the raw JWT.
pub async fn stream( pub async fn stream(
@@ -56,7 +130,7 @@ pub async fn stream(
) -> Result<Sse<impl Stream<Item = Result<Event, Infallible>>>, AppError> { ) -> Result<Sse<impl Stream<Item = Result<Event, Infallible>>>, AppError> {
let token_hash = state let token_hash = state
.sse_tickets .sse_tickets
.consume(&q.ticket) .consume(&q.ticket, TicketKind::Sse)
.ok_or_else(|| AppError::Unauthorized("Ticket ungültig oder abgelaufen.".into()))?; .ok_or_else(|| AppError::Unauthorized("Ticket ungültig oder abgelaufen.".into()))?;
// NOTE: this authenticates via ticket→session, not the `AuthUser` extractor. The // NOTE: this authenticates via ticket→session, not the `AuthUser` extractor. The
@@ -68,6 +142,17 @@ pub async fn stream(
.map_err(|e| AppError::Internal(e.into()))? .map_err(|e| AppError::Internal(e.into()))?
.ok_or_else(|| AppError::Unauthorized("Sitzung nicht gefunden.".into()))?; .ok_or_else(|| AppError::Unauthorized("Sitzung nicht gefunden.".into()))?;
// Bound how many streams this session holds open — see MAX_OPEN_STREAMS_PER_SESSION. Refuse
// rather than evict: closing somebody's live feed to make room for their own reconnect loop
// reads exactly like the flakiness it would be trying to fix.
let slot = claim_stream_slot(&token_hash).ok_or_else(|| {
tracing::warn!("session at its open-SSE-stream cap; refusing another");
AppError::TooManyRequests(
"Zu viele offene Verbindungen. Bitte schließe andere Tabs.".into(),
Some(10),
)
})?;
let rx = state.sse_tx.subscribe(); let rx = state.sse_tx.subscribe();
let events = BroadcastStream::new(rx).filter_map(|msg| match msg { let events = BroadcastStream::new(rx).filter_map(|msg| match msg {
Ok(sse_event) => Some(Ok(Event::default() Ok(sse_event) => Some(Ok(Event::default()
@@ -93,6 +178,10 @@ pub async fn stream(
let pool = state.pool.clone(); let pool = state.pool.clone();
let session_hash = token_hash.clone(); let session_hash = token_hash.clone();
let session_gone = async move { let session_gone = async move {
// Owns the slot guard, and this future is owned by the returned stream — so the slot is
// released exactly when the stream is dropped, including when the client simply walks out
// of range and no cleanup code of ours ever runs.
let _slot = slot;
let mut ticker = tokio::time::interval(Duration::from_secs(60)); let mut ticker = tokio::time::interval(Duration::from_secs(60));
ticker.tick().await; // consume the immediate first tick ticker.tick().await; // consume the immediate first tick
loop { loop {

View File

@@ -13,9 +13,10 @@ use crate::auth::middleware::RequireAdmin;
use crate::error::AppError; use crate::error::AppError;
use crate::state::AppState; use crate::state::AppState;
/// Truncates every event-scoped table, wipes media on disk, and reseeds the /// Truncates every event-scoped table, wipes media on disk, and reseeds the `config`
/// `config` table from migration defaults. Requires an admin JWT — even with /// table: numeric values from the migration defaults, but every feature toggle forced
/// `EVENTSNAP_TEST_MODE=1` it cannot be hit anonymously. /// OFF (production seeds them ON — see the note at the reseed below). Requires an admin
/// JWT — even with `EVENTSNAP_TEST_MODE=1` it cannot be hit anonymously.
pub async fn truncate_all( pub async fn truncate_all(
State(state): State<AppState>, State(state): State<AppState>,
RequireAdmin(_auth): RequireAdmin, RequireAdmin(_auth): RequireAdmin,
@@ -40,15 +41,29 @@ pub async fn truncate_all(
.execute(&state.pool) .execute(&state.pool)
.await?; .await?;
// Reseed config mirrors migrations 005 and 009. Kept in sync by hand // Reseed config. The NUMERIC values mirror migrations 005/015/016/017/019; the BOOLEAN
// because pulling SQL out of the migration files at runtime is fragile. // toggles deliberately do NOT — migration 009 seeds every one of them `true`
// (production), and this forces them `false` so the suite isn't fighting rate limits
// and quotas it isn't testing.
//
// Be aware of what that costs: this runs as an auto-fixture before EVERY test, so no
// test starts from production's config unless it explicitly turns a toggle back on
// (02-upload/rate-limit, 07-adversarial/ddos, 01-auth/rate-limit-nat, …). That blind
// spot is exactly why an entire class of per-IP limiter bugs went unnoticed: the
// limiters were simply off. When adding a limiter or quota, add a spec that enables it.
//
// Kept in sync by hand because pulling SQL out of the migration files at runtime is
// fragile — if you add a config key in a migration, add it here too.
sqlx::query( sqlx::query(
r#"INSERT INTO config (key, value) VALUES r#"INSERT INTO config (key, value) VALUES
('max_image_size_mb', '20'), ('max_image_size_mb', '20'),
('max_video_size_mb', '500'), ('max_video_size_mb', '500'),
('upload_rate_per_hour', '10'), ('upload_rate_per_hour', '100'),
('feed_rate_per_min', '60'), ('feed_rate_per_min', '60'),
('export_rate_per_day', '3'), ('export_rate_per_day', '3'),
('join_ip_rate_per_min', '300'),
('recover_ip_rate_per_min', '30'),
('social_rate_per_min', '120'),
('quota_tolerance', '0.75'), ('quota_tolerance', '0.75'),
('estimated_guest_count', '100'), ('estimated_guest_count', '100'),
('compression_concurrency', '2'), ('compression_concurrency', '2'),
@@ -57,6 +72,8 @@ pub async fn truncate_all(
('feed_rate_enabled', 'false'), ('feed_rate_enabled', 'false'),
('export_rate_enabled', 'false'), ('export_rate_enabled', 'false'),
('join_rate_enabled', 'false'), ('join_rate_enabled', 'false'),
('social_rate_enabled', 'false'),
('admin_login_rate_enabled', 'false'),
('quota_enabled', 'false'), ('quota_enabled', 'false'),
('storage_quota_enabled', 'false'), ('storage_quota_enabled', 'false'),
('upload_count_quota_enabled', 'false'), ('upload_count_quota_enabled', 'false'),
@@ -94,6 +111,12 @@ pub async fn truncate_all(
// steers the per-user limit off `free_disk_bytes`), i.e. two holes were masking each other. // steers the per-user limit off `free_disk_bytes`), i.e. two holes were masking each other.
state.disk_cache.invalidate(); state.disk_cache.invalidate();
// `media_total` caches SUM(user.total_upload_bytes) for the upload gate's keepsake-headroom
// check. TRUNCATE has just zeroed every one of those rows, so a surviving reading would make
// the next test's first upload measure its headroom against the previous test's gallery —
// and that gate REFUSES uploads, so the failure would look like a spurious quota rejection.
state.media_total.invalidate();
// `sse_tickets` maps a ticket to a session token hash. TRUNCATE deletes the sessions, so every // `sse_tickets` maps a ticket to a session token hash. TRUNCATE deletes the sessions, so every
// surviving ticket is a dangling reference to a user that no longer exists. // surviving ticket is a dangling reference to a user that no longer exists.
state.sse_tickets.clear(); state.sse_tickets.clear();

File diff suppressed because it is too large Load Diff

View File

@@ -2,7 +2,6 @@ use anyhow::Result;
use axum::Router; use axum::Router;
use axum::extract::DefaultBodyLimit; use axum::extract::DefaultBodyLimit;
use axum::routing::{delete, get, patch, post}; use axum::routing::{delete, get, patch, post};
use tower_http::services::ServeDir;
use tower_http::trace::TraceLayer; use tower_http::trace::TraceLayer;
use tracing_subscriber::{layer::SubscriberExt, util::SubscriberInitExt}; use tracing_subscriber::{layer::SubscriberExt, util::SubscriberInitExt};
@@ -28,13 +27,34 @@ async fn main() -> Result<()> {
tracing_subscriber::registry() tracing_subscriber::registry()
.with( .with(
// `info`, not `debug`. A stock deploy sets RUST_LOG nowhere (it is absent from
// .env.example and was absent from docker-compose.yml), so this fallback IS the
// production level — and at `debug` the TraceLayer below emits a line per request
// AND per response, into a log file that had no rotation. `tower_http=warn`
// rather than `info` states the intent: those spans are diagnostics, not an
// access log, and a future `DefaultOnResponse::new().level(Level::INFO)` should
// not silently re-enable them.
tracing_subscriber::EnvFilter::try_from_default_env() tracing_subscriber::EnvFilter::try_from_default_env()
.unwrap_or_else(|_| "eventsnap_backend=debug,tower_http=debug".into()), .unwrap_or_else(|_| "eventsnap_backend=info,tower_http=warn".into()),
) )
.with(tracing_subscriber::fmt::layer()) .with(tracing_subscriber::fmt::layer())
.init(); .init();
let config = AppConfig::from_env()?; let config = AppConfig::from_env()?;
// Prove both media directories are writable BEFORE anything else runs. This is first
// because everything downstream — the derivative backfill, export recovery, every upload —
// assumes it silently.
//
// This used to be `create_dir_all(&media_path).await.ok()` far below, which discarded the
// only signal there was, and EXPORT_PATH was never created or probed at all. The failure
// mode that produced: a wrong bind mount or a root-owned volume left the app booting
// *green* — `/health` only probes the database — so Caddy routed traffic to it, guests
// joined, and every single upload failed with EACCES. Existence is not the property we
// need; writability is, and the only way to know is to write.
ensure_writable_dir(&config.media_path, "MEDIA_PATH").await?;
ensure_writable_dir(&config.export_path, "EXPORT_PATH").await?;
let pool = db::create_pool(&config.database_url).await?; let pool = db::create_pool(&config.database_url).await?;
// Reset any rows left mid-flight by a previous (possibly crashed) instance — // Reset any rows left mid-flight by a previous (possibly crashed) instance —
@@ -45,6 +65,19 @@ async fn main() -> Result<()> {
let state = AppState::new(pool.clone(), config.clone()); let state = AppState::new(pool.clone(), config.clone());
// Regenerate image derivatives an older pipeline produced: the big-screen display for
// uploads processed before it existed (v0.17.x), and anything predating the current
// DERIVATIVES_REV (rev 1 applies the EXIF orientation, without which every portrait
// phone photo is stored sideways). Fire-and-forget behind the compression semaphore;
// originals are never touched, so a failure just retries on the next start.
state.compression.backfill_stale_derivatives().await;
// Re-extract poster frames for videos a restart interrupted. `startup_recovery` above
// marks their compression `failed` but nothing re-enqueued them, so `thumbnail_path`
// stayed NULL for the rest of the event. Shares the attempt budget with the image
// backfill, so a clip that genuinely yields no frame stops being retried.
state.compression.backfill_video_posters().await;
// Re-spawn exports for events that were released but whose keepsake never finished // Re-spawn exports for events that were released but whose keepsake never finished
// (crash mid-export). Needs the media/export paths + SSE sender, so it runs here // (crash mid-export). Needs the media/export paths + SSE sender, so it runs here
// rather than inside `startup_recovery`. Fire-and-forget: the workers run in the // rather than inside `startup_recovery`. Fire-and-forget: the workers run in the
@@ -53,6 +86,7 @@ async fn main() -> Result<()> {
pool.clone(), pool.clone(),
config.media_path.clone(), config.media_path.clone(),
config.export_path.clone(), config.export_path.clone(),
config.comments_enabled,
state.sse_tx.clone(), state.sse_tx.clone(),
) )
.await; .await;
@@ -63,11 +97,9 @@ async fn main() -> Result<()> {
pool, pool,
state.rate_limiter.clone(), state.rate_limiter.clone(),
state.sse_tickets.clone(), state.sse_tickets.clone(),
config.media_path.clone(),
); );
// Ensure media directories exist
tokio::fs::create_dir_all(&config.media_path).await.ok();
let api = Router::new() let api = Router::new()
// Auth // Auth
.route("/api/v1/event", get(handlers::public::get_public_event)) .route("/api/v1/event", get(handlers::public::get_public_event))
@@ -106,6 +138,10 @@ async fn main() -> Result<()> {
"/api/v1/upload/{id}/preview", "/api/v1/upload/{id}/preview",
get(handlers::upload::get_preview), get(handlers::upload::get_preview),
) )
.route(
"/api/v1/upload/{id}/display",
get(handlers::upload::get_display),
)
.route( .route(
"/api/v1/upload/{id}/thumbnail", "/api/v1/upload/{id}/thumbnail",
get(handlers::upload::get_thumbnail), get(handlers::upload::get_thumbnail),
@@ -113,10 +149,15 @@ async fn main() -> Result<()> {
// Current-user endpoints (live quota estimate, profile + privacy note bundle) // Current-user endpoints (live quota estimate, profile + privacy note bundle)
.route("/api/v1/me/context", get(handlers::me::get_context)) .route("/api/v1/me/context", get(handlers::me::get_context))
.route("/api/v1/me/quota", get(handlers::me::get_quota)) .route("/api/v1/me/quota", get(handlers::me::get_quota))
// Self-service erasure. There was no user-deletion route at any role, so an erasure
// request could only be honoured with hand-written SQL against production — and the join
// page's data notice now promises this exists. See `me::delete_account`.
.route("/api/v1/me", delete(handlers::me::delete_account))
// Feed // Feed
.route("/api/v1/feed", get(handlers::feed::feed)) .route("/api/v1/feed", get(handlers::feed::feed))
.route("/api/v1/feed/delta", get(handlers::feed::feed_delta)) .route("/api/v1/feed/delta", get(handlers::feed::feed_delta))
.route("/api/v1/hashtags", get(handlers::feed::hashtags)) .route("/api/v1/hashtags", get(handlers::feed::hashtags))
.route("/api/v1/uploaders", get(handlers::feed::uploaders))
// Social // Social
.route( .route(
"/api/v1/upload/{id}/like", "/api/v1/upload/{id}/like",
@@ -219,45 +260,154 @@ async fn main() -> Result<()> {
api api
}; };
// Serve media files from disk // NOTE: media is deliberately NOT served over HTTP.
let media_service = ServeDir::new(&config.media_path); //
// Files live under `media_path` so the compression worker and the export job can read
// them off disk, but nothing may pull them straight from `/media/**` — that bypasses
// the visibility checks (soft-delete + ban-hide) that make a host takedown stick.
// Every legitimate fetch goes through `/api/v1/upload/{id}/{original,preview,display,
// thumbnail}`, which filter via `find_visible_media`; those are the only media URLs the
// backend ever emits (see `handlers::feed`).
//
// This used to be a `ServeDir` on `/media` with four `nest_service` blockers on the
// subtrees above it. That was bypassable: axum routes on the RAW path while `ServeDir`
// percent-decodes afterwards, so `/media/%70reviews/{id}.jpg` missed every blocker,
// fell through to the `ServeDir`, and was decoded back to `previews/` on disk — serving
// a taken-down photo to anyone, unauthenticated. Any single escaped byte worked, in all
// four subtrees. Deleting the route removes the vector outright rather than racing the
// decoder; `/media/**` now 404s regardless of encoding.
let router = Router::new() let router = Router::new()
.route("/health", get(|| async { "ok" })) // ONE probe, and it touches the database. The merge of the unattended-blockers work
// brought a competing design — a dependency-free `/health` for the compose gate plus
// a DB-backed `/health/ready` for an external monitor. That split is defensible, and
// it was rejected deliberately:
//
// * `/health` returning a constant "ok" is the exact defect faea555 fixed and
// verified live (stop Postgres → 503 → start → 200, with no app restart). Every
// request path touches the database, so a constant probe reports healthy while
// the app is useless — the disk-full endgame stayed green all the way down.
// * The split's motive was that Caddy's `depends_on: app: service_healthy` would
// be blocked by a Postgres hiccup at boot. But `app` itself already gates on
// `db: service_healthy`, so the DB is up before this probe ever runs, and the
// healthcheck carries a 20s start_period plus 5 retries on top.
// * The two handlers were the same `SELECT 1` with the same 2s timeout under two
// names, so keeping both bought nothing.
//
// The external uptime monitor points at this route — DEPLOYMENT_RUNBOOK.md §10.4,
// which documents the response table and is the only thing in this deployment that
// can page a human. (That section previously did not exist and this comment claimed
// it did; if you are removing §10.4, this route loses its only consumer.)
.route("/health", get(health))
.merge(api) .merge(api)
// Block direct HTTP access to ALL media subtrees. The files live under
// `media_path` (so the compression worker and export can read them off disk) but
// must NOT be pullable straight from `/media/**` — that bypasses the visibility
// checks (soft-delete + ban-hide) in the gated handlers. Every legitimate fetch
// goes through `/api/v1/upload/{id}/{original,preview,thumbnail}`, which filter
// via `find_visible_media`. The more specific nests take precedence over the
// `/media` ServeDir below (which, with all three subtrees blocked, now serves
// nothing — kept as a backstop).
.nest_service(
"/media/originals",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service(
"/media/previews",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service(
"/media/thumbnails",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service("/media", media_service)
.layer(TraceLayer::new_for_http()) .layer(TraceLayer::new_for_http())
.with_state(state); .with_state(state);
let listener = tokio::net::TcpListener::bind(("0.0.0.0", config.app_port)).await?; let listener = tokio::net::TcpListener::bind(("0.0.0.0", config.app_port)).await?;
tracing::info!("listening on {}", listener.local_addr()?); tracing::info!("listening on {}", listener.local_addr()?);
axum::serve(listener, router) // `into_make_service_with_connect_info` is required by the pre-auth handlers, which
// extract `ConnectInfo<SocketAddr>` to use the peer address as the rate-limit key when
// X-Forwarded-For is absent. Without it those extractors fail at runtime.
axum::serve(
listener,
router.into_make_service_with_connect_info::<std::net::SocketAddr>(),
)
.with_graceful_shutdown(shutdown_signal()) .with_graceful_shutdown(shutdown_signal())
.await?; .await?;
Ok(()) Ok(())
} }
/// Create `dir` if absent, then prove we can actually write inside it. Hard error otherwise.
///
/// `create_dir_all` succeeding proves nothing: it is a no-op on an existing directory, so a
/// root-owned volume, a read-only bind mount and a full filesystem all "succeed". The probe
/// below is the only thing that distinguishes them, and it is worth the two syscalls once per
/// boot to turn a silent evening of failed uploads into a container that refuses to start.
///
/// `label` is the env var name so the operator gets the name of the knob to fix, not a path
/// they then have to trace back to a variable.
async fn ensure_writable_dir(dir: &std::path::Path, label: &str) -> anyhow::Result<()> {
use anyhow::Context;
tokio::fs::create_dir_all(dir)
.await
.with_context(|| format!("{label}: cannot create {}", dir.display()))?;
// A fixed name is fine: this runs once, before the server accepts requests, and two
// instances sharing one volume would be a misconfiguration in its own right. Removed on
// both the success and failure paths so a crashed boot cannot leave litter behind.
let probe = dir.join(".eventsnap-write-probe");
let result = async {
let mut f = tokio::fs::File::create(&probe)
.await
.with_context(|| format!("{label}: cannot create a file in {}", dir.display()))?;
// Write and fsync rather than just create: a full filesystem lets the create succeed
// and fails at the first byte, which is exactly the disk-full endgame this guards.
tokio::io::AsyncWriteExt::write_all(&mut f, b"ok")
.await
.with_context(|| format!("{label}: cannot write to {}", dir.display()))?;
f.sync_all()
.await
.with_context(|| format!("{label}: cannot flush to {}", dir.display()))?;
anyhow::Ok(())
}
.await;
let _ = tokio::fs::remove_file(&probe).await;
result.with_context(|| {
format!(
"{label} ({}) is not writable. The app refuses to start rather than accept uploads \
it cannot store — check the bind mount and that the volume is owned by the \
container's non-root user.",
dir.display()
)
})?;
tracing::info!(path = %dir.display(), "{label} is writable");
Ok(())
}
/// How long `/health` waits for the database before calling the app unhealthy. Deliberately
/// short: the point is to answer "can this process actually serve a request right now", and a
/// probe that blocks for the acquire timeout is itself a symptom.
const HEALTH_DB_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(2);
/// Readiness probe — the Docker healthcheck, Caddy's `depends_on` gate, and the runbook's
/// event-day `curl` all hit this.
///
/// It used to return the literal string `"ok"` and touch nothing. Every request in the app needs
/// the database, so that answered a question nobody asked: the container reported healthy while
/// every real request 500'd, and with no operator watching during the event there was no signal
/// at all. Note what this does NOT buy: Compose's `restart: unless-stopped` does not react to
/// healthcheck state, so nothing restarts on a red probe — this is a diagnostic, and it is
/// deliberately not wired to automatic recovery, because the pool already heals itself across a
/// Postgres restart (sqlx revalidates on acquire) and an auto-restart would truncate every
/// in-flight upload to "fix" an outage that was about to clear on its own.
async fn health(
axum::extract::State(state): axum::extract::State<AppState>,
) -> impl axum::response::IntoResponse {
use axum::http::StatusCode;
match tokio::time::timeout(
HEALTH_DB_TIMEOUT,
sqlx::query("SELECT 1").execute(&state.pool),
)
.await
{
Ok(Ok(_)) => (StatusCode::OK, "ok"),
Ok(Err(e)) => {
tracing::error!(error = ?e, "health check: database query failed");
(StatusCode::SERVICE_UNAVAILABLE, "database unavailable")
}
Err(_) => {
tracing::error!(
timeout_s = HEALTH_DB_TIMEOUT.as_secs(),
"health check: database did not respond"
);
(StatusCode::SERVICE_UNAVAILABLE, "database timeout")
}
}
}
/// Hard cap on how long we wait for in-flight connections to drain after a shutdown /// Hard cap on how long we wait for in-flight connections to drain after a shutdown
/// signal. Uploads (streamed to disk) finish in well under this; the cap exists because /// signal. Uploads (streamed to disk) finish in well under this; the cap exists because
/// long-lived SSE streams never end on their own and would otherwise keep the graceful /// long-lived SSE streams never end on their own and would otherwise keep the graceful

View File

@@ -62,13 +62,27 @@ impl Comment {
) -> Result<Vec<CommentDto>, sqlx::Error> { ) -> Result<Vec<CommentDto>, sqlx::Error> {
// Two-step: pick the newest `limit` rows older than `before`, then flip // Two-step: pick the newest `limit` rows older than `before`, then flip
// them back into ascending order so the caller can render top-to-bottom. // them back into ascending order so the caller can render top-to-bottom.
// `AND NOT u.is_banned` — the filter that was missing (H11).
//
// Only `deleted_at` was checked, so a banned guest's comments stayed on the live feed
// forever: the host bans somebody for an abusive comment, watches every photo of theirs
// vanish, and the comment is still sitting there on the most-viewed photo of the evening.
// Nothing on the client evicted them either.
//
// The tell that this was an oversight rather than a decision: the EXPORT query already
// filters `is_banned`, so the comment disappeared from the keepsake but not from the app —
// the two views of the same moderation action disagreed. Migration 021 did the same for
// hashtag counts. This brings the live read path in line with both.
//
// A ban is reversible and this is derived at read time, so `unban_user` restores the
// comments with no extra work.
sqlx::query_as::<_, CommentDto>( sqlx::query_as::<_, CommentDto>(
"SELECT * FROM ( "SELECT * FROM (
SELECT c.id, c.upload_id, c.user_id, u.display_name AS uploader_name, SELECT c.id, c.upload_id, c.user_id, u.display_name AS uploader_name,
c.body, c.created_at c.body, c.created_at
FROM comment c FROM comment c
JOIN \"user\" u ON u.id = c.user_id JOIN \"user\" u ON u.id = c.user_id
WHERE c.upload_id = $1 AND c.deleted_at IS NULL WHERE c.upload_id = $1 AND c.deleted_at IS NULL AND NOT u.is_banned
AND ($2::timestamptz IS NULL OR c.created_at < $2) AND ($2::timestamptz IS NULL OR c.created_at < $2)
ORDER BY c.created_at DESC ORDER BY c.created_at DESC
LIMIT $3 LIMIT $3

View File

@@ -32,14 +32,32 @@ impl Event {
.await .await
} }
/// Insert the event, or return the existing row if another request won the race.
///
/// `ON CONFLICT`, not a bare INSERT. `slug` is UNIQUE (migration 002), and the only callers are
/// `/join` and `/admin/login` — both of which run before the row exists, at the one moment the
/// app is most concurrent: the QR code goes up and every phone in the room posts `/join` within
/// the same second. A check-then-insert loses that race by construction, and the losers got a
/// bare unique violation surfaced as a 500 on the very first screen of the event.
///
/// `DO UPDATE SET slug = EXCLUDED.slug` is a deliberate no-op write: `DO NOTHING` returns no
/// row on conflict, which would put the loser right back at square one. It touches only `slug`,
/// so `name`, `export_epoch` and the lock/release timestamps are never disturbed by a late
/// arrival.
pub async fn create(pool: &PgPool, slug: &str, name: &str) -> Result<Self, sqlx::Error> { pub async fn create(pool: &PgPool, slug: &str, name: &str) -> Result<Self, sqlx::Error> {
sqlx::query_as::<_, Self>("INSERT INTO event (slug, name) VALUES ($1, $2) RETURNING *") sqlx::query_as::<_, Self>(
"INSERT INTO event (slug, name) VALUES ($1, $2)
ON CONFLICT (slug) DO UPDATE SET slug = EXCLUDED.slug
RETURNING *",
)
.bind(slug) .bind(slug)
.bind(name) .bind(name)
.fetch_one(pool) .fetch_one(pool)
.await .await
} }
/// Reads first so the common case (the row already exists, i.e. every join after the first)
/// stays a plain SELECT and never takes a row lock.
pub async fn find_or_create( pub async fn find_or_create(
pool: &PgPool, pool: &PgPool,
slug: &str, slug: &str,
@@ -51,3 +69,71 @@ impl Event {
Self::create(pool, slug, name).await Self::create(pool, slug, name).await
} }
} }
#[cfg(test)]
mod tests {
use super::*;
/// The QR code goes up and every phone posts `/join` in the same second, before the event row
/// exists. `find_or_create` reads first, so all of them miss, and all of them insert.
///
/// With a bare `INSERT`, exactly one wins and the rest get a unique violation on `slug` —
/// surfaced as a 500 on the first screen of the event, for everyone but the winner. There is no
/// retry on that path and nothing in the UI explains it.
#[sqlx::test]
async fn concurrent_first_joins_all_get_the_same_event(pool: PgPool) {
let racers: Vec<_> = (0..16)
.map(|_| {
let pool = pool.clone();
tokio::spawn(
async move { Event::find_or_create(&pool, "wedding", "Hochzeit").await },
)
})
.collect();
let mut ids = Vec::new();
for r in racers {
let event = r
.await
.expect("task panicked")
.expect("a concurrent first join must not fail — this is the QR-scan burst");
ids.push(event.id);
}
assert_eq!(ids.len(), 16);
assert!(
ids.iter().all(|id| *id == ids[0]),
"every racer must land on ONE event row, not create rivals"
);
let count: i64 = sqlx::query_scalar("SELECT COUNT(*) FROM event WHERE slug = 'wedding'")
.fetch_one(&pool)
.await
.expect("count");
assert_eq!(count, 1, "exactly one event row may exist for a slug");
}
/// A late arrival must not clobber the row it collides with — the no-op `DO UPDATE` exists to
/// return the loser a row, not to let it rewrite one mid-event.
#[sqlx::test]
async fn a_late_create_does_not_disturb_the_existing_row(pool: PgPool) {
let first = Event::find_or_create(&pool, "wedding", "Hochzeit")
.await
.expect("first");
sqlx::query("UPDATE event SET name = $1, export_epoch = 7 WHERE id = $2")
.bind("Anna und Ben")
.bind(first.id)
.execute(&pool)
.await
.expect("simulate a live event");
let late = Event::create(&pool, "wedding", "Hochzeit")
.await
.expect("a colliding insert must still return the row");
assert_eq!(late.id, first.id);
assert_eq!(late.name, "Anna und Ben", "the name must survive");
assert_eq!(late.export_epoch, 7, "and so must the export epoch");
}
}

View File

@@ -63,6 +63,33 @@ impl Hashtag {
.await?; .await?;
Ok(()) Ok(())
} }
/// The upload's current tags, lowercased and sorted — the comparable form.
///
/// Exists so `edit_upload` can tell a real hashtag change from a re-send of the same list.
/// Without it, `PATCH {"hashtags": []}` in a loop retired the HTML keepsake on every request
/// (readiness is derived from `event.export_epoch`), and every armed rebuild was superseded
/// before the debounce let it start — so the viewer 404'd for the rest of the event at zero
/// cost to the client. One indexed lookup on `upload_hashtag(upload_id)` is a fair price for
/// closing that.
pub async fn normalized_for_upload<'e, E>(
executor: E,
upload_id: Uuid,
) -> Result<Vec<String>, sqlx::Error>
where
E: sqlx::PgExecutor<'e>,
{
let rows: Vec<(String,)> = sqlx::query_as(
"SELECT lower(h.tag) FROM hashtag h
JOIN upload_hashtag uh ON uh.hashtag_id = h.id
WHERE uh.upload_id = $1
ORDER BY lower(h.tag)",
)
.bind(upload_id)
.fetch_all(executor)
.await?;
Ok(rows.into_iter().map(|(t,)| t).collect())
}
} }
/// Extract `#hashtags` from text (caption or body). Tags are restricted to /// Extract `#hashtags` from text (caption or body). Tags are restricted to

View File

@@ -46,12 +46,16 @@ pub struct VisibleMedia {
pub original_path: String, pub original_path: String,
pub preview_path: Option<String>, pub preview_path: Option<String>,
pub thumbnail_path: Option<String>, pub thumbnail_path: Option<String>,
pub display_path: Option<String>,
pub mime_type: String, pub mime_type: String,
} }
impl Upload { impl Upload {
/// Takes any executor so the caller can run it inside a transaction (atomic /// Takes any executor so the caller can run it inside a transaction (atomic
/// quota + insert) or standalone against the pool. /// quota + insert) or standalone against the pool.
// Eight arguments, one per column the INSERT writes, with exactly one call site. A params
// struct here would restate the column list a second time and buy nothing.
#[allow(clippy::too_many_arguments)]
pub async fn create<'e, E>( pub async fn create<'e, E>(
executor: E, executor: E,
event_id: Uuid, event_id: Uuid,
@@ -60,13 +64,34 @@ impl Upload {
mime_type: &str, mime_type: &str,
original_size_bytes: i64, original_size_bytes: i64,
caption: Option<&str>, caption: Option<&str>,
) -> Result<Self, sqlx::Error> client_upload_id: Option<Uuid>,
) -> Result<Option<Self>, sqlx::Error>
where where
E: sqlx::PgExecutor<'e>, E: sqlx::PgExecutor<'e>,
{ {
// `Ok(None)` means this exact `client_upload_id` is already stored — the caller's request
// is a retry of one that already succeeded, and it must replay the original row rather
// than create a second. Letting the unique index raise instead would work, but only after
// the whole transaction had aborted, and it would arrive as an opaque database error the
// caller would have to string-match to recognise.
//
// The conflict target repeats the index's `WHERE` clause because it is a partial index;
// without it Postgres cannot prove which index to use and rejects the statement.
//
// KEEP THIS IN LOCKSTEP WITH `upload_client_upload_id_key` (migrations 026 and 031). The
// predicate here must match the index's, or the arbiter cannot be inferred and every
// upload that carries a `client_upload_id` fails as a runtime 500 — queries in this
// codebase are not compile-time checked, so nothing catches a drift at build time.
//
// `deleted_at IS NULL` is what makes a retry-after-delete work instead of 409ing forever:
// the key is claimed only while a LIVE row holds it, which is what
// `find_by_client_upload_id` below has always assumed. `OR taken_down_by_host` carves the
// moderation case back out — see migration 031: releasing the key for a HOST takedown let
// a late retry resurrect a photo the host had deliberately removed.
sqlx::query_as::<_, Self>( sqlx::query_as::<_, Self>(
"INSERT INTO upload (event_id, user_id, original_path, mime_type, original_size_bytes, caption) "INSERT INTO upload (event_id, user_id, original_path, mime_type, original_size_bytes, caption, client_upload_id)
VALUES ($1, $2, $3, $4, $5, $6) VALUES ($1, $2, $3, $4, $5, $6, $7)
ON CONFLICT (client_upload_id) WHERE client_upload_id IS NOT NULL AND (deleted_at IS NULL OR taken_down_by_host) DO NOTHING
RETURNING *", RETURNING *",
) )
.bind(event_id) .bind(event_id)
@@ -75,7 +100,52 @@ impl Upload {
.bind(mime_type) .bind(mime_type)
.bind(original_size_bytes) .bind(original_size_bytes)
.bind(caption) .bind(caption)
.fetch_one(executor) .bind(client_upload_id)
.fetch_optional(executor)
.await
}
/// Look up a live upload by the idempotency key its client sent.
///
/// Scoped to the user as well as the key: the key alone is unique, but a lookup that ignored
/// ownership would let one guest's retry return another guest's row if a key ever repeated.
/// Soft-deleted rows are excluded on purpose — if the guest deleted the photo and their queue
/// later retries, they should get a fresh upload rather than a resurrection of a deleted one.
pub async fn find_by_client_upload_id(
pool: &sqlx::PgPool,
user_id: Uuid,
client_upload_id: Uuid,
) -> Result<Option<Self>, sqlx::Error> {
sqlx::query_as::<_, Self>(
"SELECT * FROM upload
WHERE client_upload_id = $1 AND user_id = $2 AND deleted_at IS NULL",
)
.bind(client_upload_id)
.bind(user_id)
.fetch_optional(pool)
.await
}
/// Was this key claimed by a row the HOST took down?
///
/// Only used to answer a refused retry honestly. Without it the guest's queue shows
/// "Dieser Upload wurde bereits verarbeitet." for a photo that was in fact removed by the
/// hosts — technically true, actively misleading, and it invites them to try again.
pub async fn taken_down_by_client_upload_id(
pool: &sqlx::PgPool,
user_id: Uuid,
client_upload_id: Uuid,
) -> Result<bool, sqlx::Error> {
sqlx::query_scalar::<_, bool>(
"SELECT EXISTS (
SELECT 1 FROM upload
WHERE client_upload_id = $1 AND user_id = $2
AND deleted_at IS NOT NULL AND taken_down_by_host
)",
)
.bind(client_upload_id)
.bind(user_id)
.fetch_one(pool)
.await .await
} }
@@ -93,7 +163,7 @@ impl Upload {
id: Uuid, id: Uuid,
) -> Result<Option<VisibleMedia>, sqlx::Error> { ) -> Result<Option<VisibleMedia>, sqlx::Error> {
sqlx::query_as::<_, VisibleMedia>( sqlx::query_as::<_, VisibleMedia>(
"SELECT up.original_path, up.preview_path, up.thumbnail_path, up.mime_type "SELECT up.original_path, up.preview_path, up.thumbnail_path, up.display_path, up.mime_type
FROM upload up FROM upload up
JOIN \"user\" u ON u.id = up.user_id JOIN \"user\" u ON u.id = up.user_id
WHERE up.id = $1 AND up.deleted_at IS NULL WHERE up.id = $1 AND up.deleted_at IS NULL
@@ -134,6 +204,94 @@ impl Upload {
Ok(()) Ok(())
} }
pub async fn set_display_path(
pool: &PgPool,
id: Uuid,
display_path: &str,
) -> Result<(), sqlx::Error> {
sqlx::query("UPDATE upload SET display_path = $2 WHERE id = $1")
.bind(id)
.bind(display_path)
.execute(pool)
.await?;
Ok(())
}
/// Stamp which revision of the derivative pipeline produced this row's preview/display,
/// so the startup backfill can find rows generated by an older one exactly once.
///
/// Also clears the attempt counter: success is the only thing that resets it, and folding
/// the reset in here means both the live path and the backfill get it with no extra call
/// site to forget.
pub async fn set_derivatives_rev(pool: &PgPool, id: Uuid, rev: i16) -> Result<(), sqlx::Error> {
sqlx::query(
"UPDATE upload
SET derivatives_rev = $2, derivative_attempts = 0, derivative_last_error = NULL
WHERE id = $1",
)
.bind(id)
.bind(rev)
.execute(pool)
.await?;
Ok(())
}
/// Record that derivative processing is ABOUT to be attempted, returning the new count.
///
/// WRITE-AHEAD ON PURPOSE. The failure this bounds is a cgroup SIGKILL: the process
/// vanishes mid-work, so no `Err` is returned, no error handler runs and no `Drop` fires.
/// A counter incremented after a failure would increment zero times per crash and the
/// boot loop would be unchanged. Counting the ATTEMPT is the only thing that survives the
/// process dying. The cost is that a genuinely transient failure also burns an attempt —
/// acceptable, because the retry budget is per-boot-loop, not per-request, and success
/// resets it to zero.
/// `None` when the row no longer exists (hard-deleted, or an e2e TRUNCATE landed while the
/// task waited on the semaphore) — the caller should abandon quietly rather than treat a
/// missing row as a processing failure.
pub async fn begin_derivative_attempt(
pool: &PgPool,
id: Uuid,
) -> Result<Option<i16>, sqlx::Error> {
sqlx::query_scalar(
"UPDATE upload
SET derivative_attempts = derivative_attempts + 1
WHERE id = $1
RETURNING derivative_attempts",
)
.bind(id)
.fetch_optional(pool)
.await
}
/// Read the lifetime derivative-attempt counter WITHOUT charging it.
///
/// Used by the in-request retries after the first: those re-enter `do_process` but must not
/// spend the lifetime budget again (see `charge_lifetime_attempt`). `None` still means the
/// row vanished, so the caller's "nothing to do" branch keeps working unchanged.
pub async fn derivative_attempts(pool: &PgPool, id: Uuid) -> Result<Option<i16>, sqlx::Error> {
sqlx::query_scalar("SELECT derivative_attempts FROM upload WHERE id = $1")
.bind(id)
.fetch_optional(pool)
.await
}
/// Store why the last derivative attempt failed. Diagnostics only — nothing branches on it.
pub async fn record_derivative_failure(
pool: &PgPool,
id: Uuid,
error: &str,
) -> Result<(), sqlx::Error> {
// Bounded: an anyhow chain can be long, and this is written on a failure path that may
// repeat across every row of a bad batch.
let truncated: String = error.chars().take(500).collect();
sqlx::query("UPDATE upload SET derivative_last_error = $2 WHERE id = $1")
.bind(id)
.bind(truncated)
.execute(pool)
.await?;
Ok(())
}
pub async fn set_thumbnail_path( pub async fn set_thumbnail_path(
pool: &PgPool, pool: &PgPool,
id: Uuid, id: Uuid,
@@ -147,40 +305,8 @@ impl Upload {
Ok(()) Ok(())
} }
/// Soft-deletes the upload and decrements the uploader's `total_upload_bytes`. /// Soft-deletes an upload within its event and refunds the uploader's
/// Done in a single transaction so a crash between the two writes can't leave /// `total_upload_bytes`, in one transaction. Returns `false` if no row
/// the quota counter pointing at bytes the user has already deleted (which would
/// silently lock them out of future uploads).
///
/// No-op if the row is already deleted — protects against a double-tap on the
/// delete action double-decrementing the counter.
pub async fn soft_delete(pool: &PgPool, id: Uuid) -> Result<(), sqlx::Error> {
let mut tx = pool.begin().await?;
let row: Option<(Uuid, i64)> = sqlx::query_as(
"UPDATE upload
SET deleted_at = NOW()
WHERE id = $1 AND deleted_at IS NULL
RETURNING user_id, original_size_bytes",
)
.bind(id)
.fetch_optional(&mut *tx)
.await?;
if let Some((user_id, bytes)) = row {
sqlx::query(
"UPDATE \"user\"
SET total_upload_bytes = GREATEST(0, total_upload_bytes - $2)
WHERE id = $1",
)
.bind(user_id)
.bind(bytes)
.execute(&mut *tx)
.await?;
}
tx.commit().await?;
Ok(())
}
/// Event-scoped variant of [`Self::soft_delete`]. Returns `false` if no row
/// matched (already deleted, wrong event, or unknown id) so host handlers /// matched (already deleted, wrong event, or unknown id) so host handlers
/// can return a clean 404 instead of silently no-op'ing. /// can return a clean 404 instead of silently no-op'ing.
/// Executor-generic so a caller can run the delete and the keepsake regeneration in ONE /// Executor-generic so a caller can run the delete and the keepsake regeneration in ONE
@@ -188,20 +314,27 @@ impl Upload {
/// dropped handler future, a failed second tx), the taken-down photo stays in the downloadable /// dropped handler future, a failed second tx), the taken-down photo stays in the downloadable
/// archive forever, and recovery can't tell — the keepsake still looks complete at the current /// archive forever, and recovery can't tell — the keepsake still looks complete at the current
/// epoch, and the host can no longer even find the upload to retry. /// epoch, and the host can no longer even find the upload to retry.
///
/// `by_host` records WHO removed it, which decides whether the row keeps holding its
/// idempotency key — see migration 031. A host takedown holds it, so a late retry from the
/// uploader's queue cannot bring the photo back; a guest deleting their own photo releases it,
/// so their next upload of the same queue item succeeds.
pub async fn soft_delete_in_event( pub async fn soft_delete_in_event(
conn: &mut sqlx::PgConnection, conn: &mut sqlx::PgConnection,
id: Uuid, id: Uuid,
event_id: Uuid, event_id: Uuid,
by_host: bool,
) -> Result<bool, sqlx::Error> { ) -> Result<bool, sqlx::Error> {
let tx = conn; let tx = conn;
let row: Option<(Uuid, i64)> = sqlx::query_as( let row: Option<(Uuid, i64)> = sqlx::query_as(
"UPDATE upload "UPDATE upload
SET deleted_at = NOW() SET deleted_at = NOW(), taken_down_by_host = $3
WHERE id = $1 AND event_id = $2 AND deleted_at IS NULL WHERE id = $1 AND event_id = $2 AND deleted_at IS NULL
RETURNING user_id, original_size_bytes", RETURNING user_id, original_size_bytes",
) )
.bind(id) .bind(id)
.bind(event_id) .bind(event_id)
.bind(by_host)
.fetch_optional(&mut *tx) .fetch_optional(&mut *tx)
.await?; .await?;
let deleted = if let Some((user_id, bytes)) = row { let deleted = if let Some((user_id, bytes)) = row {

View File

@@ -47,19 +47,89 @@ impl User {
event_id: Uuid, event_id: Uuid,
display_name: &str, display_name: &str,
pin_hash: &str, pin_hash: &str,
client_join_id: Option<Uuid>,
) -> Result<Self, sqlx::Error> { ) -> Result<Self, sqlx::Error> {
sqlx::query_as::<_, Self>( sqlx::query_as::<_, Self>(
"INSERT INTO \"user\" (event_id, display_name, recovery_pin_hash) "INSERT INTO \"user\" (event_id, display_name, recovery_pin_hash, client_join_id)
VALUES ($1, $2, $3) VALUES ($1, $2, $3, $4)
RETURNING *", RETURNING *",
) )
.bind(event_id) .bind(event_id)
.bind(display_name) .bind(display_name)
.bind(pin_hash) .bind(pin_hash)
.bind(client_join_id)
.fetch_one(pool) .fetch_one(pool)
.await .await
} }
/// Look up a join that already succeeded, by the idempotency key its client sent.
///
/// The retry path for H16: the account was created but the response never arrived, so the
/// client re-sends the same `client_join_id`. Finding a row here means "this join already
/// happened" — the caller rotates the PIN and answers with a usable one rather than 409ing
/// on a name the caller itself owns.
pub async fn find_by_client_join_id(
pool: &PgPool,
event_id: Uuid,
client_join_id: Uuid,
) -> Result<Option<Self>, sqlx::Error> {
sqlx::query_as::<_, Self>(
"SELECT * FROM \"user\" WHERE event_id = $1 AND client_join_id = $2",
)
.bind(event_id)
.bind(client_join_id)
.fetch_optional(pool)
.await
}
/// Create a user with an explicit role, in ONE statement.
///
/// `create` + a separate `UPDATE ... SET role` is not equivalent: a crash or a pool error
/// between the two leaves a GUEST row holding a reserved name, which is exactly the
/// poisoned state that bricked admin login — now self-inflicted, and invisible to a
/// role-based lookup, so the next login would create yet another.
pub async fn create_with_role(
pool: &PgPool,
event_id: Uuid,
display_name: &str,
pin_hash: &str,
role: UserRole,
) -> Result<Self, sqlx::Error> {
sqlx::query_as::<_, Self>(
"INSERT INTO \"user\" (event_id, display_name, recovery_pin_hash, role)
VALUES ($1, $2, $3, $4)
RETURNING *",
)
.bind(event_id)
.bind(display_name)
.bind(pin_hash)
.bind(role)
.fetch_one(pool)
.await
}
/// The event's admin, looked up BY ROLE.
///
/// The name is not the identity and never was. Looking the admin up by `display_name`
/// meant any guest who joined as "Admin" first made the lookup miss, and the fallback
/// `create` then violated the case-insensitive unique index from migration 007 — a
/// permanent 500 on admin login, recoverable only by hand-editing the database.
///
/// `ORDER BY created_at` so a database that somehow acquired two admin rows resolves to a
/// stable one rather than alternating between them.
pub async fn find_admin_for_event(
pool: &PgPool,
event_id: Uuid,
) -> Result<Option<Self>, sqlx::Error> {
sqlx::query_as::<_, Self>(
"SELECT * FROM \"user\" WHERE event_id = $1 AND role = 'admin'
ORDER BY created_at ASC LIMIT 1",
)
.bind(event_id)
.fetch_optional(pool)
.await
}
pub async fn find_by_id(pool: &PgPool, id: Uuid) -> Result<Option<Self>, sqlx::Error> { pub async fn find_by_id(pool: &PgPool, id: Uuid) -> Result<Option<Self>, sqlx::Error> {
sqlx::query_as::<_, Self>("SELECT * FROM \"user\" WHERE id = $1") sqlx::query_as::<_, Self>("SELECT * FROM \"user\" WHERE id = $1")
.bind(id) .bind(id)
@@ -96,14 +166,31 @@ impl User {
Ok(row.0) Ok(row.0)
} }
/// Window after which a failed-PIN streak is forgotten. Matches the lockout duration, so
/// "wait out the cooldown" and "start clean" are the same interval to a guest.
const PIN_ATTEMPT_DECAY_MINUTES: i64 = 15;
/// Record a wrong PIN and return the CURRENT streak length.
///
/// The counter decays: before this, it only ever cleared on a successful recovery or after
/// a lockout expired, so ordinary typos accumulated across days and a guest could arrive at
/// an event already most of the way to being locked out by mistakes made the night before.
/// Decay is what makes the raised lock threshold safe rather than merely lenient.
pub async fn increment_failed_pin(pool: &PgPool, id: Uuid) -> Result<i16, sqlx::Error> { pub async fn increment_failed_pin(pool: &PgPool, id: Uuid) -> Result<i16, sqlx::Error> {
let row: (i16,) = sqlx::query_as( let row: (i16,) = sqlx::query_as(
"UPDATE \"user\" "UPDATE \"user\"
SET failed_pin_attempts = failed_pin_attempts + 1 SET failed_pin_attempts = CASE
WHEN last_failed_pin_at IS NULL
OR last_failed_pin_at < NOW() - ($2 || ' minutes')::interval
THEN 1
ELSE failed_pin_attempts + 1
END,
last_failed_pin_at = NOW()
WHERE id = $1 WHERE id = $1
RETURNING failed_pin_attempts", RETURNING failed_pin_attempts",
) )
.bind(id) .bind(id)
.bind(Self::PIN_ATTEMPT_DECAY_MINUTES.to_string())
.fetch_one(pool) .fetch_one(pool)
.await?; .await?;
Ok(row.0) Ok(row.0)
@@ -124,7 +211,9 @@ impl User {
pub async fn reset_pin_attempts(pool: &PgPool, id: Uuid) -> Result<(), sqlx::Error> { pub async fn reset_pin_attempts(pool: &PgPool, id: Uuid) -> Result<(), sqlx::Error> {
sqlx::query( sqlx::query(
"UPDATE \"user\" SET failed_pin_attempts = 0, pin_locked_until = NULL WHERE id = $1", "UPDATE \"user\"
SET failed_pin_attempts = 0, pin_locked_until = NULL, last_failed_pin_at = NULL
WHERE id = $1",
) )
.bind(id) .bind(id)
.execute(pool) .execute(pool)

View File

@@ -0,0 +1,314 @@
//! Append-only record of privileged actions. See migration 029 for why it exists.
//!
//! Design constraints, both learned from the rest of this codebase:
//!
//! * **Never fail the action.** An audit write that can turn a successful ban into a 500 makes
//! moderation less reliable than no audit at all. Every failure here is logged and swallowed.
//! * **Never store a credential.** `reset_pin` is the action most worth recording and the one
//! whose payload must never be in `detail` — a table that could hand back a guest's PIN would
//! be a worse privacy problem than the gap it closes.
//!
//! **Action slugs actually written**, since migration 029's header lists three (`promote_user`,
//! `demote_user`, `delete_user`) that no call site has ever emitted, and the migration file cannot
//! be corrected without changing its checksum and crash-looping every database that ran it:
//!
//! `ban_user`, `unban_user`, `set_role`, `reset_pin`, `delete_upload`, `delete_comment`,
//! `lock_uploads`, `unlock_uploads`, `release_gallery`, `delete_account`, `patch_config`.
//! Eleven, one per `audit::record` call site — grep for it if this list ages.
//!
//! **There is deliberately no read endpoint.** The table is queried by hand:
//!
//! ```sql
//! SELECT created_at, actor_name, actor_role, action, target_name, detail
//! FROM host_action_audit ORDER BY created_at DESC LIMIT 50;
//! ```
use serde_json::Value;
use sqlx::PgPool;
use uuid::Uuid;
use crate::models::user::UserRole;
/// Record one privileged action.
///
/// Takes `&PgPool` rather than a transaction on purpose: the audit row is not part of the action's
/// atomicity. If the action commits and the audit write fails we want the action to stand (and a
/// loud log line); if the action rolls back, an orphan audit row saying "someone tried" is more
/// useful than silence.
#[allow(clippy::too_many_arguments)]
pub async fn record(
pool: &PgPool,
event_id: Uuid,
actor_id: Uuid,
actor_name: Option<&str>,
actor_role: UserRole,
action: &str,
target_id: Option<Uuid>,
target_name: Option<&str>,
detail: Option<Value>,
) {
// Resolve whatever names the caller did not supply.
//
// Migration 029 made `actor_id`/`target_id` deliberately non-FK so "the record survives the
// actor's account being removed, which is exactly when it is most likely to be wanted". Every
// caller passed None for both names, so what survived was a bare uuid resolving to nothing —
// the guarantee the column exists for, minus the only thing that made it readable.
//
// Resolved HERE rather than at eleven call sites so none can be missed. The one caller that
// destroys the row it is recording — `me::delete_account` — must still pass the name in, since
// by the time this runs there is nothing left to look up, and that is precisely the row a host
// will be reading the next morning ("whose photos disappeared?").
let (actor_name, target_name) =
resolve_names(pool, actor_id, actor_name, target_id, target_name).await;
let result = sqlx::query(
"INSERT INTO host_action_audit
(event_id, actor_id, actor_name, actor_role, action, target_id, target_name, detail)
VALUES ($1, $2, $3, $4, $5, $6, $7, $8)",
)
.bind(event_id)
.bind(actor_id)
.bind(actor_name.as_deref())
// `as_str()`, not `format!("{actor_role:?}")`: the Debug spelling is not a stable wire format,
// so a `#[derive(Debug)]` change or a renamed variant would silently start writing a different
// string into a column nothing validates. `as_str` is the one the rest of the codebase uses.
.bind(actor_role.as_str())
.bind(action)
.bind(target_id)
.bind(target_name.as_deref())
.bind(detail)
.execute(pool)
.await;
match result {
Ok(_) => {}
Err(e) => {
// `error`, not `warn`: losing an audit row is the kind of thing that should show up in
// whatever is watching the logs, even though it must not fail the request.
tracing::error!(
error = ?e, action, %actor_id, ?target_id,
"failed to write host action audit row"
);
}
}
}
/// Fill in any name the caller left as `None`, in ONE query.
///
/// Best-effort by the same rule as the insert: a failed lookup writes NULL rather than failing the
/// action, and it is one round-trip whether zero, one or both names are missing.
async fn resolve_names(
pool: &PgPool,
actor_id: Uuid,
actor_name: Option<&str>,
target_id: Option<Uuid>,
target_name: Option<&str>,
) -> (Option<String>, Option<String>) {
let need_actor = actor_name.is_none();
let need_target = target_name.is_none() && target_id.is_some();
if !need_actor && !need_target {
return (
actor_name.map(str::to_owned),
target_name.map(str::to_owned),
);
}
let mut wanted: Vec<Uuid> = Vec::with_capacity(2);
if need_actor {
wanted.push(actor_id);
}
if let Some(t) = target_id
&& need_target
{
wanted.push(t);
}
let rows: Vec<(Uuid, String)> =
sqlx::query_as("SELECT id, display_name FROM \"user\" WHERE id = ANY($1)")
.bind(&wanted)
.fetch_all(pool)
.await
.unwrap_or_default();
let lookup = |id: Uuid| rows.iter().find(|(i, _)| *i == id).map(|(_, n)| n.clone());
(
actor_name.map(str::to_owned).or_else(|| lookup(actor_id)),
target_name
.map(str::to_owned)
.or_else(|| target_id.and_then(lookup)),
)
}
/// These live HERE, not in `tests/`, and that is the entire point.
///
/// `backend/` is a binary crate, so an integration test cannot import `record`. The house rule in
/// `tests/common/mod.rs` — copy the production SQL character-for-character — works for pinning
/// behaviour that already existed, but applied to a NEW fix whose only coverage is the copy it
/// proves nothing: the fix and its test become two independent implementations, and deleting the
/// fix leaves the test green. The previous `tests/audit_names.rs` did exactly that, down to
/// asserting `actor_role == "host"` against its own hardcoded `.bind("host")` — an assertion that
/// could not fail for any change to the code it named.
///
/// A `#[cfg(test)]` module inside the binary can call the real function, so these do.
#[cfg(test)]
mod tests {
use super::*;
async fn seed_event(pool: &PgPool, slug: &str) -> Uuid {
sqlx::query_scalar("INSERT INTO event (slug, name) VALUES ($1, $2) RETURNING id")
.bind(slug)
.bind("Hochzeit")
.fetch_one(pool)
.await
.expect("seed event")
}
async fn seed_user(pool: &PgPool, event_id: Uuid, name: &str) -> Uuid {
sqlx::query_scalar(
"INSERT INTO \"user\" (event_id, display_name, recovery_pin_hash)
VALUES ($1, $2, 'x') RETURNING id",
)
.bind(event_id)
.bind(name)
.fetch_one(pool)
.await
.expect("seed user")
}
async fn audit_row(pool: &PgPool, action: &str) -> Option<(Option<String>, Option<String>)> {
sqlx::query_as(
"SELECT actor_name, target_name FROM host_action_audit
WHERE action = $1 ORDER BY created_at DESC LIMIT 1",
)
.bind(action)
.fetch_optional(pool)
.await
.expect("audit lookup")
}
/// The ordinary case: the caller supplies no names and `record` resolves both from the ids.
/// This is what nine of the eleven call sites do. Revert `resolve_names` and both names go NULL.
#[sqlx::test]
async fn a_recorded_action_carries_both_names_without_the_caller_supplying_them(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let host = seed_user(&pool, event_id, "Gastgeberin Greta").await;
let guest = seed_user(&pool, event_id, "Gesperrter Gustav").await;
record(
&pool,
event_id,
host,
None,
UserRole::Host,
"ban_user",
Some(guest),
None,
None,
)
.await;
let (actor_name, target_name) = audit_row(&pool, "ban_user").await.expect("a row");
assert_eq!(actor_name.as_deref(), Some("Gastgeberin Greta"));
assert_eq!(target_name.as_deref(), Some("Gesperrter Gustav"));
}
/// `as_str()`, not the `Debug` spelling. Asserted against `UserRole::as_str` itself rather than
/// a literal, so it tracks a rename instead of pretending to: what must hold is that the column
/// carries the SAME string the rest of the codebase uses, whatever that string is. Swap line 75
/// back to `format!("{actor_role:?}")` and this goes red on the `Host`/`host` casing.
#[sqlx::test]
async fn the_role_column_carries_the_canonical_spelling(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let host = seed_user(&pool, event_id, "Gastgeberin Greta").await;
record(
&pool,
event_id,
host,
None,
UserRole::Host,
"release_gallery",
None,
None,
None,
)
.await;
let role: String = sqlx::query_scalar(
"SELECT actor_role FROM host_action_audit WHERE action = 'release_gallery'",
)
.fetch_one(&pool)
.await
.expect("role");
assert_eq!(role, UserRole::Host.as_str());
assert_ne!(
role,
format!("{:?}", UserRole::Host),
"the Debug spelling is not a wire format"
);
}
/// The case the columns exist for. `delete_account` hard-deletes the user row, so a name
/// resolved AFTER the fact would be NULL — the caller has to pass it in.
#[sqlx::test]
async fn a_name_supplied_by_the_caller_survives_the_row_being_deleted(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let leaver = seed_user(&pool, event_id, "Abschied Anke").await;
// Exactly the order `me::delete_account` runs in: the row goes first, the audit row second.
sqlx::query("DELETE FROM \"user\" WHERE id = $1")
.bind(leaver)
.execute(&pool)
.await
.expect("delete user");
record(
&pool,
event_id,
leaver,
Some("Abschied Anke"),
UserRole::Guest,
"delete_account",
Some(leaver),
Some("Abschied Anke"),
None,
)
.await;
let (actor_name, target_name) = audit_row(&pool, "delete_account").await.expect("a row");
assert_eq!(
actor_name.as_deref(),
Some("Abschied Anke"),
"the audit row must name the deleted account — resolving it later is impossible"
);
assert_eq!(target_name.as_deref(), Some("Abschied Anke"));
}
/// And the failure mode that made this worth testing: with nothing supplied and nothing to look
/// up, the write must still succeed (an audit row must never fail an action) and carry NULLs.
#[sqlx::test]
async fn an_unresolvable_name_writes_the_row_anyway(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let ghost = Uuid::new_v4();
record(
&pool,
event_id,
ghost,
None,
UserRole::Host,
"reset_pin",
Some(ghost),
None,
None,
)
.await;
let (actor_name, target_name) = audit_row(&pool, "reset_pin")
.await
.expect("the row must be written even when no name can be resolved");
assert_eq!(actor_name, None);
assert_eq!(target_name, None);
}
}

View File

@@ -48,6 +48,31 @@ impl CompressionWorker {
self.generation.fetch_add(1, Ordering::SeqCst); self.generation.fetch_add(1, Ordering::SeqCst);
} }
/// How many times `do_process` is attempted before an upload is given up on. The
/// give-up path is user-visible (the photo disappears), so transient infrastructure
/// errors must not reach it.
const MAX_PROCESS_ATTEMPTS: u32 = 3;
/// Revision of the image-derivative pipeline. Bump this whenever a change makes existing
/// previews/displays wrong, so `backfill_stale_derivatives` regenerates them once on the
/// next start. Rev 1 = EXIF orientation is applied.
const DERIVATIVES_REV: i16 = 1;
/// How many times derivative generation may be ATTEMPTED for one upload before it is left
/// alone. Counted write-ahead and reset on success — see `Upload::begin_derivative_attempt`.
///
/// This is what turns a fatal input from an outage into a blemish. The startup backfill
/// runs unconditionally on every boot, so before this bound a row whose processing killed
/// the process was re-selected and re-run forever, and `restart: unless-stopped` made that
/// an infinite loop that also dropped every SSE stream and truncated every in-flight
/// upload on each cycle. Three attempts absorbs genuinely transient infrastructure
/// failures (an ENOSPC spike, a pool blip) without ever becoming unbounded.
const MAX_DERIVATIVE_ATTEMPTS: i16 = 3;
/// Rows regenerated per boot. Bounds both the query and the amount of work a single start
/// can queue; whatever is left is picked up on the next boot.
const BACKFILL_BATCH: i64 = 200;
/// Spawn a background task to process an uploaded file. /// Spawn a background task to process an uploaded file.
pub fn process(&self, upload_id: Uuid, original_path: String, mime_type: String) { pub fn process(&self, upload_id: Uuid, original_path: String, mime_type: String) {
let worker = self.clone(); let worker = self.clone();
@@ -60,10 +85,44 @@ impl CompressionWorker {
if worker.generation.load(Ordering::SeqCst) != born_at { if worker.generation.load(Ordering::SeqCst) != born_at {
return; return;
} }
// Retry before giving up. Most failures here are transient and self-clearing —
// an ENOSPC spike while several guests upload at once, a momentary DB-pool
// exhaustion, a panic inside the image codec — and the give-up path is
// user-visible data loss, so it is worth a few seconds to avoid entering it.
//
// But only for failures that CAN clear. An image that exceeds the decode budget,
// is corrupt, or is in an unsupported format fails identically on every attempt,
// so retrying it just burns 2s + 4s of backoff and writes three near-identical
// warnings before reaching the same conclusion. Give up on those immediately.
let mut attempt = 1u32;
let outcome = loop {
match worker match worker
.do_process(upload_id, &original_path, &mime_type) // Charge the lifetime budget once per episode, on the first attempt only.
.do_process(upload_id, &original_path, &mime_type, attempt == 1)
.await .await
{ {
Ok(v) => break Ok(v),
Err(e)
if attempt < Self::MAX_PROCESS_ATTEMPTS
&& !crate::services::imaging::is_permanent_image_error(&e)
&& !crate::services::imaging::is_storage_full_error(&e) =>
{
tracing::warn!(
error = ?e, %upload_id, attempt,
"compression attempt failed; retrying"
);
tokio::time::sleep(std::time::Duration::from_secs(2u64.pow(attempt))).await;
attempt += 1;
// The data may have been reset while we slept (e2e TRUNCATE).
if worker.generation.load(Ordering::SeqCst) != born_at {
return;
}
}
Err(e) => break Err(e),
}
};
match outcome {
Ok(_) => { Ok(_) => {
tracing::info!("compression completed for upload {upload_id}"); tracing::info!("compression completed for upload {upload_id}");
let _ = worker.sse_tx.send(SseEvent { let _ = worker.sse_tx.send(SseEvent {
@@ -71,29 +130,75 @@ impl CompressionWorker {
data: serde_json::json!({ "upload_id": upload_id }).to_string(), data: serde_json::json!({ "upload_id": upload_id }).to_string(),
}); });
} }
Err(e) => { Err(e) if crate::services::imaging::is_storage_full_error(&e) => {
tracing::error!("compression failed for upload {upload_id}: {e:#}"); // Out of disk. Keep the row AND the original — the opposite of the branch
// Auto-cleanup: a failed transcode would otherwise leave a // below, and for the same reason it retains the file: nothing here is the
// permanently broken feed card, silently charge the uploader's // guest's fault and nothing about the photo is wrong.
// quota, and orphan the original on disk. Refund + soft-delete //
// (one tx, so v_feed excludes it), remove the orphan file, then // Soft-deleting on ENOSPC was strictly harmful. It refunded the quota while
// tell the uploader (upload-error toast) and evict the card // keeping the bytes, so it freed nothing, removed the photo from the feed
// everywhere (upload-deleted, already handled by the feed). // seconds after a `201 Created`, and handed the guest the allowance to
// upload it again into the same full disk. Leaving the row live costs
// nothing instead: every client already falls back to the original when
// `preview_url` and `thumbnail_url` are NULL, so the photo stays visible —
// just uncompressed — and `backfill_stale_derivatives` regenerates the
// derivatives on the next start, once there is room for them.
tracing::error!(
%upload_id,
"compression failed: the media filesystem is out of space. The upload is \
kept and served from its original; free disk space and restart to \
regenerate derivatives: {e:#}"
);
let _ = Upload::set_compression_status(&worker.pool, upload_id, "failed").await; let _ = Upload::set_compression_status(&worker.pool, upload_id, "failed").await;
if let Err(del) = Upload::soft_delete(&worker.pool, upload_id).await { // Not an "error" event: nothing was lost and there is nothing for the guest
tracing::warn!(error = ?del, %upload_id, "failed to soft-delete after compression failure"); // to act on. Clients treat this purely as "refetch me", which is what makes
} // the card appear with its original as the image source.
let orphan = worker.media_path.join(&original_path); let _ = worker.sse_tx.send(SseEvent {
if let Err(rm) = tokio::fs::remove_file(&orphan).await { event_type: "upload-processed".to_string(),
tracing::warn!(error = ?rm, path = %orphan.display(), "failed to remove orphaned original"); data: serde_json::json!({ "upload_id": upload_id }).to_string(),
});
} }
Err(e) => {
tracing::error!(
"compression failed for upload {upload_id} after {attempt} attempt(s): {e:#}"
);
// KEEP THE ROW. This used to soft-delete, which made a derivative failure
// indistinguishable — to the guest — from their photo being deleted: they
// got a `201 Created`, watched the card appear, and then watched it vanish.
// The row left `v_feed`, `find_visible_media` and BOTH keepsake archives,
// so the photo was gone from the product's core promise while its bytes sat
// on disk for 14 days waiting for a `cleanup_deleted_media` that nothing
// told anyone about. There is no host or admin screen listing compression
// failures, so recovery meant hand-written SQL that also had to re-add the
// refunded quota bytes. Against "0 lost uploads", that was silent per-photo
// loss on any error the ENOSPC arm above doesn't catch — a HEIC that slipped
// the allowlist, a truncated frame, an ffmpeg hiccup, a pool blip.
//
// This is exactly what the ENOSPC arm already does and documents as correct:
// every client falls back to the original when `preview_url` and
// `thumbnail_url` are NULL, so the photo stays visible and downloadable —
// just uncompressed — and `backfill_stale_derivatives` retries it on the
// next boot, now bounded by `derivative_attempts` so a poisoned row cannot
// loop. The quota stays charged, which is correct: the bytes are still on
// disk and still the guest's.
let _ = Upload::set_compression_status(&worker.pool, upload_id, "failed").await;
tracing::warn!(
%upload_id,
path = %worker.media_path.join(&original_path).display(),
"derivatives failed; the upload is kept and served from its original"
);
// `upload-error` still fires so the uploader learns the photo will look
// uncompressed. `upload-deleted` deliberately does NOT — nothing was
// deleted, and evicting the card was the visible half of the data loss.
let _ = worker.sse_tx.send(SseEvent { let _ = worker.sse_tx.send(SseEvent {
event_type: "upload-error".to_string(), event_type: "upload-error".to_string(),
data: serde_json::json!({ "upload_id": upload_id, "error": e.to_string() }) data: serde_json::json!({ "upload_id": upload_id, "error": e.to_string() })
.to_string(), .to_string(),
}); });
// Tell every client to refetch, so the card re-renders from the original
// instead of sitting on a stale "processing" placeholder forever.
let _ = worker.sse_tx.send(SseEvent { let _ = worker.sse_tx.send(SseEvent {
event_type: "upload-deleted".to_string(), event_type: "upload-processed".to_string(),
data: serde_json::json!({ "upload_id": upload_id }).to_string(), data: serde_json::json!({ "upload_id": upload_id }).to_string(),
}); });
} }
@@ -101,140 +206,675 @@ impl CompressionWorker {
}); });
} }
/// `charge_lifetime_attempt` is true only for the FIRST `do_process` of a given
/// `process()` call, so the two budgets stay independent.
///
/// They were not. `MAX_PROCESS_ATTEMPTS` (in-request retries, 3) and
/// `MAX_DERIVATIVE_ATTEMPTS` (lifetime, 3) are equal, and every retry re-entered here and
/// charged the lifetime counter — so one request's three retries, six seconds apart,
/// exhausted the entire lifetime budget. A ten-second pool blip during the arrival burst
/// therefore stranded every photo whose worker was inside that window with no preview and no
/// display derivative, permanently, recoverable by nothing: the boot backfill re-selects them
/// and immediately gives up on the same exhausted counter.
///
/// The two exist to bound different things — "this request is flapping" versus "this INPUT is
/// poison" — and only the second should survive across requests.
async fn do_process( async fn do_process(
&self, &self,
upload_id: Uuid, upload_id: Uuid,
original_path: &str, original_path: &str,
mime_type: &str, mime_type: &str,
charge_lifetime_attempt: bool,
) -> Result<()> { ) -> Result<()> {
Upload::set_compression_status(&self.pool, upload_id, "processing").await?; Upload::set_compression_status(&self.pool, upload_id, "processing").await?;
let original = self.media_path.join(original_path); let original = self.media_path.join(original_path);
if mime_type.starts_with("image/") { if mime_type.starts_with("image/") {
let preview_rel = self // Count the attempt BEFORE doing the work — see `begin_derivative_attempt`. If this
.generate_image_preview(upload_id, &original, mime_type) // input is the one that kills the container, this write is the only record that
// survives, and it is what stops the boot backfill replaying it forever. Charging on
// the first attempt preserves that: a container-killing input never reaches a second.
let charged = if charge_lifetime_attempt {
Upload::begin_derivative_attempt(&self.pool, upload_id).await?
} else {
// Already charged for this episode. Re-read the row only to notice it vanished.
Upload::derivative_attempts(&self.pool, upload_id).await?
};
match charged {
Some(attempts) if attempts > Self::MAX_DERIVATIVE_ATTEMPTS => {
anyhow::bail!(
"derivative generation gave up after {} attempt(s)",
attempts - 1
);
}
Some(_) => {}
// The row vanished while this task waited on the semaphore. Nothing to do, and
// reporting a failure would broadcast into a stream that no longer has a card.
None => return Ok(()),
}
let (preview_rel, display_rel) = self
.generate_image_derivatives(upload_id, &original, mime_type)
.await?; .await?;
Upload::set_preview_path(&self.pool, upload_id, &preview_rel).await?; Upload::set_preview_path(&self.pool, upload_id, &preview_rel).await?;
tracing::info!("preview generated for upload {upload_id}"); Upload::set_display_path(&self.pool, upload_id, &display_rel).await?;
Upload::set_derivatives_rev(&self.pool, upload_id, Self::DERIVATIVES_REV).await?;
tracing::info!("preview + display generated for upload {upload_id}");
} else if mime_type.starts_with("video/") { } else if mime_type.starts_with("video/") {
let thumb_rel = self.generate_video_thumbnail(upload_id, &original).await?; // A missing poster must NOT fail the upload. `set_thumbnail_path` is only reached when
Upload::set_thumbnail_path(&self.pool, upload_id, &thumb_rel).await?; // a file really exists, so `thumbnail_path` stays NULL otherwise — which every consumer
tracing::info!("thumbnail generated for upload {upload_id}"); // already handles (FeedListCard, VirtualFeed, LightboxModal are all null-safe).
//
// The `?` here used to hide the defect; making the check strict without also making
// this non-fatal would have been far worse than the bug. Every clip of a second or less
// would fail compression, exhaust its retries and be soft-deleted — a cosmetic defect
// turned into data loss, on exactly the mis-tap/Live-Photo clips guests produce most.
// Handling only the `Ok(None)` arm was not enough: the `?` on the call itself still
// routed every OTHER poster failure into the give-up path. `extract_poster_frame`
// returns `Err` when ffmpeg is missing from the image, when it hangs on a truncated
// `.mov` and trips FFMPEG_TIMEOUT, or when `thumbnails/` can't be created — and
// `set_thumbnail_path` returns `Err` on any DB blip. None of those say anything about
// the video itself, yet each one destroyed it. Confirmed live: on a box with no ffmpeg
// the spawn error propagated, exhausted all three attempts and soft-deleted the clip.
//
// Nothing about a video post depends on the poster — `get_original` serves the file
// byte-for-byte and the tile falls back to the video element — so no failure in this
// branch may fail the upload.
match self.generate_video_thumbnail(upload_id, &original).await {
Ok(Some(thumb_rel)) => {
match Upload::set_thumbnail_path(&self.pool, upload_id, &thumb_rel).await {
Ok(()) => tracing::info!("thumbnail generated for upload {upload_id}"),
Err(e) => tracing::warn!(
error = ?e, %upload_id,
"poster extracted but could not be recorded; the video keeps its own tile"
),
}
}
Ok(None) => {
tracing::warn!(
%upload_id,
"no poster frame could be extracted; the video keeps its own tile"
);
}
Err(e) => {
tracing::warn!(
error = ?e, %upload_id,
"poster extraction failed; the video keeps its own tile"
);
}
}
} }
Upload::set_compression_status(&self.pool, upload_id, "done").await?; Upload::set_compression_status(&self.pool, upload_id, "done").await?;
Ok(()) Ok(())
} }
async fn generate_image_preview( /// Longest edge of the big-screen "display" derivative used by the diashow. Sized to be
/// sharp on 1080p/4K while staying bounded (a ~2048px JPEG decodes to ~16 MB — trivial
/// for any kiosk, unlike a raw multi-thousand-pixel original).
const DISPLAY_MAX_EDGE: u32 = 2048;
/// Longest edge of the phone-feed "preview" (data-saver default).
const PREVIEW_MAX_EDGE: u32 = 800;
/// Above this pixel count the PNG original is stored as uploaded, unoptimised.
///
/// oxipng's peak memory scales with PIXELS, not file size: it decodes the PNG itself and
/// then evaluates row filters, each trial holding a full-size buffer. That is why a 2.82
/// MiB file could measure 1250 MiB of peak RSS inside a 1 GiB container — smooth,
/// synthetic content compresses to almost nothing on disk while still being 8000x8000.
/// 8 MP covers every real phone photo; beyond it we decline the (lossless, cosmetic)
/// saving rather than risk the OOM kill.
const OXIPNG_MAX_PIXELS: u64 = 8_000_000;
/// Wall-clock ceiling for one oxipng run.
///
/// Bounds TIME, NOT MEMORY — oxipng checks the deadline between trials, so a single trial
/// still allocates in full. The pixel gate above and the sequential build (see
/// `default-features = false` in Cargo.toml) are what bound memory. Do not treat this
/// constant as the OOM fix.
const OXIPNG_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(20);
/// Decode the image ONCE and emit both derivatives — the 800px `preview` (phone feed)
/// and the 2048px `display` (diashow). Returns `(preview_rel, display_rel)`.
async fn generate_image_derivatives(
&self, &self,
upload_id: Uuid, upload_id: Uuid,
original: &Path, original: &Path,
mime_type: &str, mime_type: &str,
) -> Result<String> { ) -> Result<(String, String)> {
let previews_dir = self.media_path.join("previews"); let previews_dir = self.media_path.join("previews");
let displays_dir = self.media_path.join("displays");
tokio::fs::create_dir_all(&previews_dir).await?; tokio::fs::create_dir_all(&previews_dir).await?;
tokio::fs::create_dir_all(&displays_dir).await?;
let preview_filename = format!("{upload_id}.jpg"); let filename = format!("{upload_id}.jpg");
let preview_path = previews_dir.join(&preview_filename); let preview_path = previews_dir.join(&filename);
let display_path = displays_dir.join(&filename);
let original = original.to_path_buf(); let original = original.to_path_buf();
let preview_path_clone = preview_path.clone();
let mime_owned = mime_type.to_string(); let mime_owned = mime_type.to_string();
// Run blocking image operations in a spawn_blocking task // Estimate the peak from the HEADER (no pixels decoded — the same kind of cheap probe
tokio::task::spawn_blocking(move || -> Result<()> { // the upload handler already does via `exceeds_decode_budget`) and, if this job is a
// Reject decompression bombs *before* fully decoding: the upload body // giant, take the exclusive permit so it cannot overlap another giant. Held for the
// cap bounds the file size on disk, but a small file can still decode to // whole blocking section, released on drop including on error.
// enormous dimensions (e.g. a ~1 MB image expanding to 50k×50k px → let estimate = crate::services::imaging::estimated_processing_peak_bytes(
// gigabytes), OOM-ing the box during decode/resize. 12000×12000 covers &original,
// any real phone photo; max_alloc hard-caps the decode allocation. Self::DISPLAY_MAX_EDGE,
let mut reader = image::ImageReader::open(&original)
.context("failed to open image")?
.with_guessed_format()
.context("failed to read image header")?;
let mut limits = image::Limits::default();
limits.max_image_width = Some(12_000);
limits.max_image_height = Some(12_000);
limits.max_alloc = Some(256 * 1024 * 1024);
reader.limits(limits);
let img = reader.decode().context("failed to decode image")?;
// Resize to max 800px wide, preserving aspect ratio
let preview = img.resize(800, 800, image::imageops::FilterType::Lanczos3);
preview
.save_with_format(&preview_path_clone, image::ImageFormat::Jpeg)
.context("failed to save preview")?;
// If the original is PNG, try lossless compression in-place
if mime_owned == "image/png" {
let opts = oxipng::Options::from_preset(2);
let _ = oxipng::optimize(
&oxipng::InFile::Path(original),
&oxipng::OutFile::Path {
path: None,
preserve_attrs: true,
},
&opts,
); );
let _heavy_permit = match estimate {
Some(bytes) if bytes > crate::services::imaging::HEAVY_IMAGE_BYTES => {
tracing::debug!(
%upload_id,
estimated_mib = bytes / (1024 * 1024),
"waiting for the heavy-image permit"
);
Some(
crate::services::imaging::HEAVY_IMAGE_PERMITS
.acquire()
.await,
)
} }
_ => None,
};
Ok(()) // Run blocking image operations in a spawn_blocking task
tokio::task::spawn_blocking(move || {
write_image_derivatives(
upload_id,
&original,
&mime_owned,
&preview_path,
&display_path,
)
}) })
.await??; .await??;
Ok(format!("previews/{preview_filename}")) Ok((
format!("previews/{filename}"),
format!("displays/{filename}"),
))
} }
async fn generate_video_thumbnail(&self, upload_id: Uuid, original: &Path) -> Result<String> { /// Regenerate image derivatives that an older pipeline produced. Fire-and-forget from
/// startup; picks up two cases, both of which leave the ORIGINAL untouched:
///
/// - uploads processed before the `display` derivative existed (preview but no
/// `display_path`), and
/// - uploads whose derivatives predate `DERIVATIVES_REV` — currently rev 1, which applies
/// the EXIF orientation. Everything generated before it is stored sideways for any
/// portrait phone photo.
///
/// Unlike the failure path in `process`, a backfill error is logged and skipped — it must
/// NEVER destroy or soft-delete an upload that already has a working preview.
///
/// Bounded in three ways, all of them load-bearing on a box that restarts itself:
/// `derivative_attempts` stops a fatal row being replayed on every boot, `BACKFILL_BATCH`
/// stops one start queueing unbounded work, and the whole thing runs as ONE task walking
/// the rows sequentially rather than N tasks racing for the same semaphore.
pub async fn backfill_stale_derivatives(&self) {
// `original_path IS NOT NULL` was dead — the column is NOT NULL. What actually needs
// excluding is the blanked path `cleanup_deleted_media` leaves behind.
let rows = sqlx::query_as::<_, (Uuid, String, String)>(
"SELECT id, original_path, mime_type FROM upload
WHERE deleted_at IS NULL AND mime_type LIKE 'image/%'
AND original_path <> ''
AND derivative_attempts < $2
AND (
(display_path IS NULL AND preview_path IS NOT NULL)
OR derivatives_rev < $1
)
ORDER BY created_at DESC
LIMIT $3",
)
.bind(Self::DERIVATIVES_REV)
.bind(Self::MAX_DERIVATIVE_ATTEMPTS)
.bind(Self::BACKFILL_BATCH)
.fetch_all(&self.pool)
.await;
let rows = match rows {
Ok(r) => r,
Err(e) => {
tracing::warn!(error = ?e, "derivative backfill query failed");
return;
}
};
self.report_exhausted_derivatives().await;
if rows.is_empty() {
return;
}
tracing::info!("regenerating derivatives for {} upload(s)", rows.len());
// ONE task for the whole batch. The previous shape spawned a task per row, so a large
// backlog created thousands of live tasks that each held a pool handle and queued on
// the same two semaphore permits, competing with live uploads for the entire boot.
let worker = self.clone();
tokio::spawn(async move {
for (id, original_path, mime_type) in rows {
let _permit = worker.semaphore.acquire().await;
// Write-ahead, exactly as in the live path: if this row is the one that kills
// the process, this increment is the only thing that outlives the SIGKILL.
match Upload::begin_derivative_attempt(&worker.pool, id).await {
Ok(Some(n)) if n > Self::MAX_DERIVATIVE_ATTEMPTS => continue,
Ok(Some(_)) => {}
Ok(None) => continue,
Err(e) => {
tracing::warn!(error = ?e, %id, "could not record a backfill attempt; skipping");
continue;
}
}
let original = worker.media_path.join(&original_path);
match worker
.generate_image_derivatives(id, &original, &mime_type)
.await
{
Ok((preview_rel, display_rel)) => {
let _ = Upload::set_preview_path(&worker.pool, id, &preview_rel).await;
let _ = Upload::set_display_path(&worker.pool, id, &display_rel).await;
// Clears derivative_attempts too, so a row that failed transiently is
// not one boot closer to being abandoned.
let _ =
Upload::set_derivatives_rev(&worker.pool, id, Self::DERIVATIVES_REV)
.await;
tracing::info!("derivatives regenerated for upload {id}");
}
Err(e) => {
// Leave the existing derivatives and the original intact; this row is
// retried on the next start until its attempt budget runs out. The rev
// stays behind, which is the marker that it still needs doing.
tracing::warn!(error = ?e, %id, "derivative backfill failed; leaving as-is");
let _ =
Upload::record_derivative_failure(&worker.pool, id, &format!("{e:#}"))
.await;
}
}
}
});
}
/// Re-extract poster frames for videos that never got one.
///
/// A video interrupted by a restart is stranded: `startup_recovery` flips its
/// `compression_status` from `processing` to `failed` and nothing re-enqueues it, so
/// `thumbnail_path` stays NULL forever while the clip itself plays fine. The feed shows a
/// posterless tile for the rest of the event, and after
/// `FAILED_ORIGINAL_RETENTION_DAYS` the reclaim sweep is entitled to the original.
///
/// Shares `derivative_attempts` with the image backfill on purpose. Note the consequence,
/// which is intended rather than a bug to fix later: `extract_poster_frame` returning
/// `Ok(false)` is a NORMAL, permanent outcome for a sub-second clip (Live Photos,
/// mis-taps), and since the counter is write-ahead and only cleared by a real success,
/// those clips stop being re-ffmpeg'd on every boot once the budget is spent.
pub async fn backfill_video_posters(&self) {
let rows = sqlx::query_as::<_, (Uuid, String)>(
"SELECT id, original_path FROM upload
WHERE deleted_at IS NULL AND mime_type LIKE 'video/%'
AND thumbnail_path IS NULL
AND original_path <> ''
AND derivative_attempts < $1
ORDER BY created_at DESC
LIMIT $2",
)
.bind(Self::MAX_DERIVATIVE_ATTEMPTS)
.bind(Self::BACKFILL_BATCH)
.fetch_all(&self.pool)
.await;
let rows = match rows {
Ok(r) => r,
Err(e) => {
tracing::warn!(error = ?e, "video poster backfill query failed");
return;
}
};
if rows.is_empty() {
return;
}
tracing::info!("re-extracting posters for {} video(s)", rows.len());
let worker = self.clone();
tokio::spawn(async move {
for (id, original_path) in rows {
let _permit = worker.semaphore.acquire().await;
match Upload::begin_derivative_attempt(&worker.pool, id).await {
Ok(Some(n)) if n > Self::MAX_DERIVATIVE_ATTEMPTS => continue,
Ok(Some(_)) => {}
Ok(None) => continue,
Err(e) => {
tracing::warn!(error = ?e, %id, "could not record a poster attempt; skipping");
continue;
}
}
let original = worker.media_path.join(&original_path);
match worker.generate_video_thumbnail(id, &original).await {
Ok(Some(thumb_rel)) => {
if Upload::set_thumbnail_path(&worker.pool, id, &thumb_rel)
.await
.is_ok()
{
// Clears the attempt counter: a video that eventually succeeded
// must not carry a budget scar into a future pipeline revision.
let _ = Upload::set_derivatives_rev(
&worker.pool,
id,
Self::DERIVATIVES_REV,
)
.await;
tracing::info!("poster regenerated for upload {id}");
}
}
// No frame at all — normal for a very short clip. The tile stays
// posterless and the attempt is spent, which is what stops the retry.
Ok(None) => {
tracing::debug!(%id, "still no poster frame; leaving the tile as-is");
}
Err(e) => {
tracing::warn!(error = ?e, %id, "poster backfill failed; leaving as-is");
let _ =
Upload::record_derivative_failure(&worker.pool, id, &format!("{e:#}"))
.await;
}
}
}
});
}
/// Say out loud, once per boot, that some uploads have stopped being retried.
///
/// Without this the give-up is invisible: the loop stops (which is the point) but the
/// affected photos keep a stale or missing derivative forever with nothing to notice. The
/// originals are untouched, so this is recoverable once the cause is fixed — reset
/// `derivative_attempts` to 0 and restart.
async fn report_exhausted_derivatives(&self) {
let exhausted: Result<i64, _> = sqlx::query_scalar(
"SELECT count(*) FROM upload
WHERE deleted_at IS NULL AND mime_type LIKE 'image/%'
AND derivative_attempts >= $2
AND (
(display_path IS NULL AND preview_path IS NOT NULL)
OR derivatives_rev < $1
)",
)
.bind(Self::DERIVATIVES_REV)
.bind(Self::MAX_DERIVATIVE_ATTEMPTS)
.fetch_one(&self.pool)
.await;
if let Ok(count) = exhausted
&& count > 0
{
tracing::error!(
count,
"{count} upload(s) exhausted derivative regeneration and will no longer be \
retried; their originals are intact — see upload.derivative_last_error, fix \
the cause, then reset derivative_attempts to 0 and restart"
);
}
}
/// Extract the feed poster for a video. `Ok(None)` when the clip yields no frame — see
/// [`crate::services::video::extract_poster_frame`], which owns the seek order, the timeout and
/// the artifact check that this function used to be missing.
async fn generate_video_thumbnail(
&self,
upload_id: Uuid,
original: &Path,
) -> Result<Option<String>> {
let thumbs_dir = self.media_path.join("thumbnails"); let thumbs_dir = self.media_path.join("thumbnails");
tokio::fs::create_dir_all(&thumbs_dir).await?; tokio::fs::create_dir_all(&thumbs_dir).await?;
let thumb_filename = format!("{upload_id}.jpg"); let thumb_filename = format!("{upload_id}.jpg");
let thumb_path = thumbs_dir.join(&thumb_filename); let thumb_path = thumbs_dir.join(&thumb_filename);
// Hard timeout — a malformed video can hang `ffmpeg` indefinitely. Without a let produced =
// cap, the held compression-worker semaphore permit is never released and the crate::services::video::extract_poster_frame(original, &thumb_path, 800).await?;
// pool eventually deadlocks (no further uploads ever processed). 120s is well
// above the time to extract one frame from any sane input.
let mut child = tokio::process::Command::new("ffmpeg")
.args([
"-i",
original.to_str().unwrap_or_default(),
"-vframes",
"1",
"-ss",
"00:00:01",
"-vf",
"scale=800:-1",
"-y",
thumb_path.to_str().unwrap_or_default(),
])
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.kill_on_drop(true)
.spawn()
.context("failed to spawn ffmpeg")?;
let status = Ok(produced.then(|| format!("thumbnails/{thumb_filename}")))
match tokio::time::timeout(std::time::Duration::from_secs(120), child.wait()).await {
Ok(res) => res.context("ffmpeg wait failed")?,
Err(_) => {
let _ = child.kill().await;
anyhow::bail!("ffmpeg timeout after 120s");
} }
}
/// The blocking half of [`CompressionWorker::generate_image_derivatives`]: decode once, write
/// both derivatives, then optionally shrink a PNG original in place.
///
/// A free function rather than an inline closure so its memory behaviour is directly testable —
/// this is the code path that OOM-killed the container, and the fix is a scoping property that a
/// future edit could silently undo.
fn write_image_derivatives(
upload_id: Uuid,
original: &Path,
mime_type: &str,
preview_path: &Path,
display_path: &Path,
) -> Result<()> {
let preview_max = CompressionWorker::PREVIEW_MAX_EDGE;
let display_max = CompressionWorker::DISPLAY_MAX_EDGE;
// THE FULL-SIZE DECODE IS SCOPED TO THIS BLOCK ON PURPOSE, and the block yields the
// DISPLAY derivative rather than the original.
//
// `img` is up to 256 MiB (imaging::decode_limits max_alloc) and `resize` only BORROWS it,
// so it used to stay alive through both resizes AND the oxipng call below — which decodes
// the PNG a second time and holds a full-size buffer per filter trial. That measured
// ~1250 MiB of peak RSS for a 2.8 MiB input, inside a 1 GiB cgroup: the container was
// SIGKILLed, taking every SSE stream and every in-flight upload with it.
//
// A block rather than a bare `drop(img)` because a `drop` call is one careless edit away
// from being removed as redundant-looking — and note the `else` arm MOVES `img` out, which
// is what makes "the block's value is the only survivor" true in both arms.
let (display, width, height) = {
// Decompression-bomb limits + EXIF orientation, both in one place — see
// services::imaging for why neither may be skipped.
let img = crate::services::imaging::decode_oriented(original)?;
let (width, height) = (img.width(), img.height());
// Display: max 2048px for the diashow. Only DOWNSCALE — never upscale a smaller
// original (that adds bytes with no quality gain); re-encode it as JPEG as-is.
let display = if width > display_max || height > display_max {
img.resize(
display_max,
display_max,
image::imageops::FilterType::Lanczos3,
)
} else {
img
};
(display, width, height)
}; };
if !status.success() { display
// Best-effort: drain stderr for the log. .save_with_format(display_path, image::ImageFormat::Jpeg)
let mut stderr = Vec::new(); .context("failed to save display")?;
if let Some(mut handle) = child.stderr.take() {
use tokio::io::AsyncReadExt; // Preview: max 800px, derived from the DISPLAY, not from the original.
let _ = handle.read_to_end(&mut stderr).await; //
// Both derivatives used to resize the full-size decode independently, so a 8000x8000
// original paid for two full-size Lanczos passes and their intermediates — measured 520
// MiB peak even after the scoping fix above, which two concurrent workers cannot fit in a
// 1 GiB container. Chaining 8000 -> 2048 -> 800 makes the second pass operate on 2048px
// input, and the full-size buffer is already freed by the time it runs. Quality is not the
// trade-off here: a staged Lanczos3 downscale to 800px is visually indistinguishable from
// a single-step one (and is a standard technique for large ratios).
display
.resize(
preview_max,
preview_max,
image::imageops::FilterType::Lanczos3,
)
.save_with_format(preview_path, image::ImageFormat::Jpeg)
.context("failed to save preview")?;
drop(display);
let pixels = u64::from(width) * u64::from(height);
// If the original is PNG, try lossless compression in place — but only when its pixel count
// is inside the budget, and never for longer than OXIPNG_TIMEOUT. This is a best-effort size
// saving: declining it costs disk, while attempting it unbounded cost the whole container.
if mime_type == "image/png" {
if pixels <= CompressionWorker::OXIPNG_MAX_PIXELS {
let mut opts = oxipng::Options::from_preset(2);
opts.timeout = Some(CompressionWorker::OXIPNG_TIMEOUT);
let _ = oxipng::optimize(
&oxipng::InFile::Path(original.to_path_buf()),
&oxipng::OutFile::Path {
path: None,
preserve_attrs: true,
},
&opts,
);
} else {
tracing::info!(
%upload_id, pixels,
"skipping oxipng: above the pixel budget; the original is stored as uploaded"
);
} }
anyhow::bail!("ffmpeg failed: {}", String::from_utf8_lossy(&stderr));
} }
Ok(format!("thumbnails/{thumb_filename}")) Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
/// Peak resident set of THIS process, in bytes, from `/proc/self/status`.
fn peak_rss_bytes() -> u64 {
let status = std::fs::read_to_string("/proc/self/status").expect("procfs");
let line = status
.lines()
.find(|l| l.starts_with("VmHWM:"))
.expect("VmHWM");
let kb: u64 = line
.split_whitespace()
.nth(1)
.and_then(|v| v.parse().ok())
.expect("VmHWM value");
kb * 1024
}
/// Reset the kernel's peak-RSS watermark so the measurement covers only what follows.
/// Linux 4.0+; writing "5" to `clear_refs` resets `VmHWM` to the current RSS.
fn reset_peak_rss() {
let _ = std::fs::write("/proc/self/clear_refs", "5");
}
/// The pixel gate has to sit below what the axis limits allow, or it can never fire.
#[test]
fn the_oxipng_gate_is_reachable_within_the_decode_limits() {
const _: () = {
// imaging::decode_limits permits 12_000 x 12_000 = 144 MP. A gate above that would
// never skip anything.
assert!(CompressionWorker::OXIPNG_MAX_PIXELS < 12_000 * 12_000);
// ...and it must stay above a 48 MP camera, so real photos still get optimised.
assert!(CompressionWorker::OXIPNG_MAX_PIXELS >= 8_000_000);
};
}
/// The heavy-image gate has to classify the two cases the way the sizing assumed:
/// an ordinary phone photo must NOT serialise, and the giant must.
#[test]
fn the_heavy_gate_separates_a_phone_photo_from_a_giant() {
let dir = std::env::temp_dir().join(format!("es-heavy-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
// 12 MP, the shape of a default phone capture.
let ordinary = dir.join("ordinary.jpg");
image::RgbImage::new(4032, 3024).save(&ordinary).unwrap();
let ordinary_peak = crate::services::imaging::estimated_processing_peak_bytes(
&ordinary,
CompressionWorker::DISPLAY_MAX_EDGE,
)
.expect("header readable");
assert!(
ordinary_peak <= crate::services::imaging::HEAVY_IMAGE_BYTES,
"a 12 MP photo estimated at {} MiB would serialise the common path",
ordinary_peak / 1048576
);
// The 64 MP RGBA case that measured ~516 MiB peak.
let giant = dir.join("giant.png");
image::RgbaImage::new(8000, 8000).save(&giant).unwrap();
let giant_peak = crate::services::imaging::estimated_processing_peak_bytes(
&giant,
CompressionWorker::DISPLAY_MAX_EDGE,
)
.expect("header readable");
assert!(
giant_peak > crate::services::imaging::HEAVY_IMAGE_BYTES,
"an 8000x8000 RGBA original estimated at only {} MiB would be allowed to run \
concurrently with another one — 2x its real ~516 MiB peak does not fit in 1 GiB",
giant_peak / 1048576
);
// The estimate must also be in the right ballpark, not merely on the right side of the
// threshold: 244 MiB decode + 262 MiB f32 resize intermediate.
assert!(
(400..700).contains(&(giant_peak / 1048576)),
"estimate {} MiB is far from the measured ~516 MiB peak",
giant_peak / 1048576
);
let _ = std::fs::remove_dir_all(&dir);
}
/// The OOM that took the container down, measured rather than argued.
///
/// An 8000x8000 RGBA PNG passes admission: 256,000,000 bytes is just under the 256 MiB
/// `max_alloc`, and smooth content is a few MB on disk, far under any size cap. The old
/// code kept that ~244 MiB decode alive across an unbounded, multi-threaded oxipng run and
/// peaked at ~1250 MiB — inside a 1 GiB cgroup. Being SIGKILLed there is not a blip: the
/// row was already committed, so the boot backfill replayed the identical workload on every
/// restart.
///
/// `#[ignore]` because it allocates ~250 MiB and takes a few seconds. Run explicitly:
/// cargo test --release oom -- --ignored --nocapture --test-threads=1
/// It must run ALONE — `VmHWM` is per process, so a concurrent test would pollute it.
#[test]
#[ignore = "heavy: allocates ~250 MiB; run with --ignored --test-threads=1"]
fn a_large_png_stays_far_below_the_container_limit() {
const EDGE: u32 = 8_000;
let dir = std::env::temp_dir().join(format!("es-oom-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
let original = dir.join("big.png");
// Smooth gradient: ~244 MiB decoded, a couple of MB on disk. That gap is the whole
// point — file size tells you nothing about what a PNG costs to process.
{
let mut buf = image::RgbaImage::new(EDGE, EDGE);
for (x, y, px) in buf.enumerate_pixels_mut() {
*px = image::Rgba([(x >> 5) as u8, (y >> 5) as u8, ((x + y) >> 6) as u8, 255]);
}
buf.save(&original).unwrap();
}
// Everything above is fixture setup, not the code under test.
reset_peak_rss();
let before = peak_rss_bytes();
write_image_derivatives(
Uuid::new_v4(),
&original,
"image/png",
&dir.join("preview.jpg"),
&dir.join("display.jpg"),
)
.expect("derivatives");
let peak = peak_rss_bytes();
let on_disk = std::fs::metadata(&original).unwrap().len();
eprintln!(
"input {:.2} MiB on disk ({EDGE}x{EDGE}); peak RSS {:.0} MiB (was {:.0} MiB before)",
on_disk as f64 / 1048576.0,
peak as f64 / 1048576.0,
before as f64 / 1048576.0
);
assert!(dir.join("preview.jpg").exists() && dir.join("display.jpg").exists());
// The container gets 1 GiB and runs two of these concurrently. 600 MiB is a generous
// ceiling that the old code (~1250 MiB) could not have met.
assert!(
peak < 600 * 1024 * 1024,
"peak RSS {} MiB — the decode is being held across oxipng again, or the pixel \
gate stopped firing",
peak / 1048576
);
let _ = std::fs::remove_dir_all(&dir);
} }
} }

View File

@@ -148,3 +148,104 @@ pub async fn get_bool(cache: &ConfigCache, key: &str, default: bool) -> bool {
_ => default, _ => default,
} }
} }
#[cfg(test)]
mod seed_tests {
/// The value a fresh database actually ends up with for `key`, by replaying the migrations.
///
/// This exists because a `config::get_*` default is only a fallback for a MISSING key, and the
/// migrations seed nearly every key there is. So the literal in the handler is dead code on any
/// real install, and changing it changes nothing — which is exactly what happened to
/// `join_ip_rate_per_min`: it was raised 60 → 300 in `auth/handlers.rs` to stop one QR-code
/// burst from locking the venue out of `/join`, shipped, and did nothing at all, because
/// migration 017 seeds 60 and the seed wins. Nothing in the test suite could see it: the e2e
/// regression guard fires 12 concurrent joins, which is green at 60 and at 300 alike.
fn effective_seed(key: &str) -> Option<String> {
let dir = std::path::Path::new(env!("CARGO_MANIFEST_DIR")).join("migrations");
let mut files: Vec<_> = std::fs::read_dir(&dir)
.expect("migrations directory")
.filter_map(|e| e.ok().map(|e| e.path()))
.filter(|p| p.to_string_lossy().ends_with(".up.sql"))
.collect();
// Version order: migrations are applied in filename order and later ones override.
files.sort();
let mut value: Option<String> = None;
for path in files {
let sql = std::fs::read_to_string(&path).expect("readable migration");
for line in sql.lines() {
let line = line.trim();
if line.starts_with("--") {
continue;
}
// Seed form: ('key', 'value')
if let Some(rest) = line.strip_prefix(&format!("('{key}',"))
&& let Some(v) = rest.split('\'').nth(1)
{
value = Some(v.to_string());
}
// Update form: UPDATE config SET value = 'new' WHERE key = 'key' AND value = 'old'
if line.starts_with("UPDATE config SET value")
&& line.contains(&format!("key = '{key}'"))
&& let Some(new) = line.split('\'').nth(1)
{
let scoped_to = line
.rsplit_once("AND value = '")
.and_then(|(_, tail)| tail.split('\'').next().map(|s| s.to_string()));
// Only applies if the current value still matches the scope it was written for.
if scoped_to.is_none() || scoped_to.as_deref() == value.as_deref() {
value = Some(new.to_string());
}
}
}
}
value
}
/// The invariant, not the number: `join_ip:{ip}` is keyed on an address the WHOLE VENUE shares
/// behind NAT, and `/join` is the one screen with no auto-retry. A ceiling near the size of the
/// party is a ceiling on the party. Asserted against the effective seed rather than the code
/// default precisely because the code default is what silently did not apply.
#[test]
fn the_join_ceiling_a_real_install_gets_is_sized_for_a_whole_venue_arriving_at_once() {
let seeded = effective_seed("join_ip_rate_per_min")
.expect("join_ip_rate_per_min must be seeded by a migration");
let seeded: usize = seeded.parse().expect("numeric");
assert!(
seeded >= 300,
"a fresh database ends up with join_ip_rate_per_min = {seeded}. Every guest shares one \
NAT address, so this is the ceiling for the entire party scanning one QR code. Raise \
it with a value-scoped UPDATE migration (see 030) — changing the default in \
auth/handlers.rs does nothing, because the seed wins."
);
}
/// Pins the other half of the same trap: the seeded value must not exceed the ceiling the
/// handler clamps to, or an operator reading `GET /admin/config` sees a number that is not the
/// one being enforced.
#[test]
fn the_seeded_recover_name_ceiling_is_within_what_the_handler_will_honour() {
let seeded = effective_seed("recover_name_rate_per_15min")
.expect("recover_name_rate_per_15min must be seeded by a migration");
let seeded: usize = seeded.parse().expect("numeric");
assert!(
seeded <= crate::auth::handlers::RECOVER_NAME_CEILING_MAX,
"seeded recover_name_rate_per_15min = {seeded} exceeds RECOVER_NAME_CEILING_MAX = {}; \
the handler clamps at the point of use, so the advertised value would be a lie.",
crate::auth::handlers::RECOVER_NAME_CEILING_MAX
);
}
/// The parser itself, against a value migration 015 really does change. Without this a bug in
/// `effective_seed` makes both tests above vacuously green.
#[test]
fn the_seed_parser_follows_a_value_through_a_later_update_migration() {
assert_eq!(
effective_seed("upload_rate_per_hour").as_deref(),
Some("100"),
"005 seeds 10 and 015 raises it to 100; reading 10 here means the UPDATE form is not \
being applied, and every assertion built on this helper is worthless."
);
assert_eq!(effective_seed("no_such_key_anywhere"), None);
}
}

View File

@@ -76,6 +76,17 @@ impl Default for DiskCache {
} }
} }
/// UNCACHED free-space reading for the filesystem backing `path`.
///
/// Deliberately bypasses [`DiskCache`]. The cache exists for the quota poll, where a 15s-stale
/// number is fine because it is only ever advisory. The export preflight is the opposite case: it
/// decides whether to start writing a multi-GB archive, and the sibling export worker running
/// concurrently can move free space by tens of gigabytes well inside the TTL. A stale reading there
/// would authorise exactly the write that fills the disk.
pub fn free_bytes(path: &Path) -> Option<u64> {
read_disk_for_path(path).map(|d| d.free)
}
/// Resolve the filesystem backing `media_path` and read its total/free bytes. /// Resolve the filesystem backing `media_path` and read its total/free bytes.
/// ///
/// Snapshots the mount table via `sysinfo`, then delegates the selection to the pure /// Snapshots the mount table via `sysinfo`, then delegates the selection to the pure

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,388 @@
//! Shared image decoding.
//!
//! Exists so there is exactly ONE way to turn a file on disk into a `DynamicImage` in this
//! codebase. Two properties have to hold everywhere an image is decoded, and both were
//! previously re-derived per call site — which is how they drifted apart:
//!
//! - **EXIF orientation must be applied.** Phones do not rotate sensor data; they record how
//! the camera was held in a tag and store the pixels as shot. `image::open` and
//! `ImageReader::decode` both hand back the raw pixels and ignore that tag, and re-encoding
//! to JPEG writes no EXIF, so the derivative is permanently sideways while the untouched
//! original still renders upright. The compression worker was fixed; the export worker was
//! not, so every portrait photo came out sideways in the keepsake's HTML viewer.
//! - **Decode limits must be set.** The upload body cap bounds the file on disk, but a small
//! file can decode to enormous dimensions (a ~1 MB image expanding to 50k×50k px), OOM-ing
//! the box. `image::open` applies NO limits at all, so the export path was also decoding
//! arbitrary user-supplied images unbounded.
use anyhow::{Context, Result};
use image::{DynamicImage, ImageDecoder};
use std::path::Path;
/// Bounds for any decode of user-supplied image data. The per-axis cap covers any real phone
/// photo; `max_alloc` bounds the decoded buffer — but only because `decode_oriented` reserves
/// against it explicitly, see there.
///
/// Sized against the deployment: the app container is capped at 1 GiB and the compression
/// worker runs `compression_concurrency` decodes at once (default 2), so 256 MiB per decode
/// leaves headroom for the resize buffers and the runtime.
fn decode_limits() -> image::Limits {
let mut limits = image::Limits::default();
limits.max_image_width = Some(12_000);
limits.max_image_height = Some(12_000);
limits.max_alloc = Some(256 * 1024 * 1024);
limits
}
/// True when re-running the exact same work on the exact same bytes cannot possibly
/// succeed, so retrying only burns wall-clock and log noise.
///
/// Deliberately narrow. Only the `ImageError` variants that are a property of the *input*
/// count: the file will not shrink, gain codec support, or un-corrupt itself between
/// attempts. `IoError` is excluded on purpose — EMFILE under load, or a momentarily
/// unreadable file, is exactly the transient case the retry exists for. A FULL disk is the
/// one io error that must not be retried either, but for a different reason and with a
/// different remedy; see [`is_storage_full_error`].
pub fn is_permanent_image_error(err: &anyhow::Error) -> bool {
err.chain().any(|cause| {
matches!(
cause.downcast_ref::<image::ImageError>(),
Some(
image::ImageError::Limits(_)
| image::ImageError::Unsupported(_)
| image::ImageError::Decoding(_)
)
)
})
}
/// True when the failure is the media filesystem being out of space.
///
/// Deliberately separate from [`is_permanent_image_error`], which is about the *input*. ENOSPC
/// is about the *host*, and it is the one failure the retry loop actively makes worse: a disk
/// does not drain during six seconds of backoff, so all three attempts fail identically while
/// holding a compression permit that photos are queued behind.
///
/// The give-up path it fed was worse still. It refunded the guest's quota and soft-deleted the
/// row while deliberately RETAINING the original — so the bytes stayed on the full disk, the
/// photo vanished from the feed seconds after a `201 Created`, and the guest was handed back
/// the quota to upload it again into the same full disk. Each round shrank free space further.
pub fn is_storage_full_error(err: &anyhow::Error) -> bool {
fn is_full(io: &std::io::Error) -> bool {
// `StorageFull` is the portable classification; the raw ENOSPC catches the paths where
// the OS error was never mapped to a named kind.
io.kind() == std::io::ErrorKind::StorageFull || io.raw_os_error() == Some(28)
}
err.chain().any(|cause| {
// `image` wraps the io error in its own variant rather than exposing it as a source,
// so the plain downcast alone would miss every derivative-write failure.
cause.downcast_ref::<std::io::Error>().is_some_and(is_full)
|| matches!(
cause.downcast_ref::<image::ImageError>(),
Some(image::ImageError::IoError(io)) if is_full(io)
)
})
}
/// Build a decoder for `path` with the budget enforced, WITHOUT reading any pixels.
///
/// Single source of truth for "may this image be decoded at all": both the upload
/// admission check and the compression worker go through here, so they cannot disagree
/// about what is acceptable.
fn decoder_within_budget(path: &Path) -> Result<impl image::ImageDecoder> {
let mut reader = image::ImageReader::open(path)
.context("failed to open image")?
.with_guessed_format()
.context("failed to read image header")?;
let mut limits = decode_limits();
reader.limits(limits.clone());
// We need `into_decoder` rather than `decode()` to read the EXIF orientation tag before
// the pixels are consumed. But the two are NOT equivalent on safety: `decode()` performs
//
// limits.reserve(decoder.total_bytes())?;
//
// between building the decoder and reading the image, and `into_decoder()` skips it (the
// crate's own FIXME concedes `from_decoder` doesn't compensate). Nothing else enforces
// `max_alloc` — the JPEG decoder's `set_limits` only checks support and dimensions — so
// without the line below the budget is inert and the ONLY bound is the per-axis cap. That
// leaves 12000x12000 decodable at 412 MiB, and two concurrent at 824 MiB against a 1 GiB
// container. Re-add it, exactly as `decode()` does.
let mut decoder = reader.into_decoder().context("failed to decode image")?;
limits
.reserve(decoder.total_bytes())
.context("image too large to decode within the memory budget")?;
decoder
.set_limits(limits)
.context("image too large to decode within the memory budget")?;
Ok(decoder)
}
/// Rough peak heap an image will cost to turn into derivatives, read from the HEADER only —
/// no pixels are decoded. `None` when the header can't be read or the image is over budget
/// (the caller is about to fail on it anyway).
///
/// Two terms, and the second is the one that surprises:
///
/// - the decoded buffer, `width * height * channels`; and
/// - the resize intermediate. `image`'s Lanczos3 path accumulates in `f32`, so the buffer
/// between the horizontal and vertical passes is `new_width * old_height * 4 channels * 4
/// bytes` — 16 bytes per pixel-row-slot, not the 4 the output uses. For an 8000x8000
/// original that is 262 MiB on top of a 244 MiB decode, measured. It is bigger than the
/// decode for any tall image, which is why "the decode is bounded by max_alloc" was never
/// the whole story.
///
/// Used to decide whether an image is heavy enough to need exclusive use of the box's memory
/// headroom, NOT to reject anything.
pub fn estimated_processing_peak_bytes(path: &Path, display_edge: u32) -> Option<u64> {
let decoder = decoder_within_budget(path).ok()?;
let (width, height) = decoder.dimensions();
let decoded = decoder.total_bytes();
// Aspect-preserving fit into `display_edge`, matching DynamicImage::resize. No downscale
// means no intermediate at all.
let intermediate = if width > display_edge || height > display_edge {
let ratio = f64::from(display_edge) / f64::from(width.max(height));
let new_width = (f64::from(width) * ratio).round().max(1.0) as u64;
new_width * u64::from(height) * 16
} else {
0
};
Some(decoded.saturating_add(intermediate))
}
/// Megapixels an image would decode to, or `None` if its header can't be read. Used only
/// to put a concrete number in the message the guest sees.
pub fn megapixels(path: &Path) -> Option<f64> {
let reader = image::ImageReader::open(path)
.ok()?
.with_guessed_format()
.ok()?;
let (w, h) = reader.into_dimensions().ok()?;
Some(f64::from(w) * f64::from(h) / 1_000_000.0)
}
/// True when an image cannot be decoded specifically because it would exceed the memory
/// budget — read from the header, no pixels touched.
///
/// Called at upload admission so a guest who sends a 100 MP photo is told at the door, with
/// a reason they can act on, instead of the upload being accepted with a 201 and then
/// silently soft-deleted minutes later when the worker gives up on it.
///
/// Deliberately narrow: ONLY the budget. A corrupt, truncated or unsupported file also
/// fails to build a decoder, but rejecting those here would change a contract the
/// adversarial suite pins on purpose — acceptance follows the magic bytes, and a payload
/// with a valid JPEG header is accepted regardless of what follows it. Those go to the
/// compression worker as before, which handles them gracefully and (since the retry
/// classifier) no longer burns backoff on them.
pub fn exceeds_decode_budget(path: &Path) -> bool {
match decoder_within_budget(path) {
Ok(_) => false,
Err(e) => e.chain().any(|cause| {
matches!(
cause.downcast_ref::<image::ImageError>(),
Some(image::ImageError::Limits(_))
)
}),
}
}
/// Decode an image from disk with decompression-bomb limits applied and its EXIF
/// orientation baked into the pixels.
///
/// Blocking — call inside `spawn_blocking`.
pub fn decode_oriented(path: &Path) -> Result<DynamicImage> {
let mut decoder = decoder_within_budget(path)?;
// Cheap, and it happens BEFORE any pixels are read: an oversized image costs a header
// parse, not an allocation.
let orientation = decoder
.orientation()
.unwrap_or(image::metadata::Orientation::NoTransforms);
let mut img = DynamicImage::from_decoder(decoder).context("failed to decode image")?;
img.apply_orientation(orientation);
Ok(img)
}
/// Process-wide serialisation for memory-heavy image work.
///
/// The `app` container gets 1 GiB. A single 8000x8000 original measures ~516 MiB peak even with
/// the decode correctly scoped, so two overlapping giants is an OOM kill — and the kernel kills
/// the whole process, dropping every SSE stream and stranding every in-flight upload.
///
/// GLOBAL rather than a field on `CompressionWorker`, because the constraint is the container's
/// memory and there is more than one producer of this work. The export's own image path
/// (`services::export`) decodes and resizes every photo in the gallery — a thumbnail for each,
/// plus a 2000px re-encode for every original over 5 MB — and it ran in a bare `spawn_blocking`
/// with no permit at all. So "host taps Freigeben while the last phone photos are still
/// compressing" put an export decode and a heavy compression job in the same cgroup at the same
/// time, which is the scenario the permit exists to make impossible. Worse, it is self-repeating:
/// the OOM kill marks the export failed, and `recover_exports` re-spawns it on boot into the same
/// conditions.
///
/// Held across the blocking section and released on drop, including on error.
pub static HEAVY_IMAGE_PERMITS: std::sync::LazyLock<tokio::sync::Semaphore> =
std::sync::LazyLock::new(|| tokio::sync::Semaphore::new(1));
/// Estimated peak heap above which a job must take [`HEAVY_IMAGE_PERMITS`].
///
/// 150 MiB sits far above a normal phone photo (a 12 MP JPEG costs ~50 MiB all-in) so the common
/// path never serialises, and far below the point where two jobs stop fitting in the container.
pub const HEAVY_IMAGE_BYTES: u64 = 150 * 1024 * 1024;
#[cfg(test)]
mod tests {
use super::*;
/// Shared with the e2e suite rather than duplicating 568 KiB of binary: the same file
/// drives `02-upload/oversized-image` so both layers assert on one artefact.
const HUGE: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/huge-99mp.jpg"
);
#[test]
fn rejects_an_image_that_would_blow_the_allocation_budget() {
// 11000x9000 = 99 MP. Deliberately UNDER the 12000px per-axis cap, so the axis check
// cannot reject it — the allocation budget is the only thing that can, which is
// exactly what makes this a regression test rather than a restatement of the axis cap.
// 283 MiB decoded as RGB8 against a 256 MiB budget, from 568 KiB on disk.
//
// This failed before the guard was restored: `ImageReader::decode` performs
// `limits.reserve(decoder.total_bytes())`, and `into_decoder()` — which we need for
// the EXIF tag — skips it, so `max_alloc` was inert and this decoded happily.
// Map the Ok arm to its dimensions first: on failure `expect_err` Debug-prints the
// value, and Debug on a DynamicImage dumps every pixel — 283 MiB of output.
let err = decode_oriented(Path::new(HUGE))
.map(|img| (img.width(), img.height()))
.expect_err("a 99 MP image must be refused, not allocated");
let msg = format!("{err:#}");
assert!(
msg.to_lowercase().contains("limit") || msg.to_lowercase().contains("memory"),
"expected a limits error, got: {msg}"
);
}
#[test]
fn an_oversized_image_is_a_permanent_failure() {
// The retry loop must not burn 2s + 4s of backoff on this: the file will not shrink
// between attempts, so all three attempts reach the identical conclusion.
let err = decode_oriented(Path::new(HUGE))
.map(|img| (img.width(), img.height()))
.expect_err("fixture must exceed the budget");
assert!(
is_permanent_image_error(&err),
"a Limits error can never succeed on retry: {err:#}"
);
}
#[test]
fn a_plain_io_error_is_not_permanent() {
// The mirror that keeps the classifier honest. EMFILE under load, or a momentary
// unreadable file, is exactly what the retry exists for — misclassifying those as
// permanent would turn a transient blip back into the data loss round 1 fixed.
// (A FULL disk is its own case now; see the storage-full tests below.)
let err = decode_oriented(Path::new("/nonexistent/definitely-not-here.jpg"))
.map(|img| (img.width(), img.height()))
.expect_err("a missing file must error");
assert!(
!is_permanent_image_error(&err),
"an IO error must stay retryable: {err:#}"
);
}
#[test]
fn a_full_disk_is_recognised_through_both_wrappers() {
// The two shapes ENOSPC actually arrives in. A bare io::Error is what `tokio::fs` and
// `std::fs` produce; the `image` crate wraps its own in `ImageError::IoError`, which is
// NOT reachable via `source()` — so a chain walk that only downcast to io::Error would
// miss every derivative-write failure, i.e. the exact case this classifier exists for.
let bare = anyhow::Error::from(std::io::Error::from(std::io::ErrorKind::StorageFull))
.context("failed to write the preview");
assert!(is_storage_full_error(&bare), "bare io::Error: {bare:#}");
let wrapped = anyhow::Error::from(image::ImageError::IoError(std::io::Error::from(
std::io::ErrorKind::StorageFull,
)))
.context("failed to save the display derivative");
assert!(
is_storage_full_error(&wrapped),
"ImageError::IoError: {wrapped:#}"
);
}
#[test]
fn an_ordinary_io_error_is_not_a_full_disk() {
// Keeps the classifier from swallowing the general case: only ENOSPC may skip the retry
// and take the keep-the-row branch. Anything else must still be retried and, if it keeps
// failing, soft-deleted as before.
let missing = decode_oriented(Path::new("/nonexistent/definitely-not-here.jpg"))
.map(|img| (img.width(), img.height()))
.expect_err("a missing file must error");
assert!(
!is_storage_full_error(&missing),
"a missing file is not a full disk: {missing:#}"
);
}
#[test]
fn admission_rejects_only_the_over_budget_case() {
// Admission and processing must agree about SIZE — a photo accepted at the door and
// then rejected by the worker for being too big is the failure this pair prevents.
assert!(
exceeds_decode_budget(Path::new(HUGE)),
"admission must reject what the decoder rejects for size"
);
let ordinary = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/portrait-exif6.jpg"
);
assert!(
!exceeds_decode_budget(Path::new(ordinary)),
"admission must accept an ordinary photo"
);
}
#[test]
fn admission_does_not_reject_a_merely_undecodable_file() {
// The narrowing that keeps the adversarial contract intact: a payload with valid
// JPEG magic bytes and nothing behind them cannot be decoded, but acceptance follows
// the magic bytes by design (07-adversarial/file-upload-attacks). It is the worker's
// job to fail it, not admission's — admission is only the resource guard.
let dir = std::env::temp_dir().join("eventsnap-imaging-test");
std::fs::create_dir_all(&dir).expect("tmp dir");
let stub = dir.join("magic-only.jpg");
let mut bytes = vec![0u8; 1024];
bytes[..3].copy_from_slice(&[0xFF, 0xD8, 0xFF]);
std::fs::write(&stub, &bytes).expect("write stub");
assert!(
!exceeds_decode_budget(&stub),
"a corrupt file is not an over-budget file"
);
assert!(
decode_oriented(&stub)
.map(|i| (i.width(), i.height()))
.is_err(),
"...but it must still fail in the worker"
);
let _ = std::fs::remove_file(&stub);
}
#[test]
fn still_decodes_an_ordinary_photo_and_applies_orientation() {
// The guard must not have become a blanket refusal. This fixture is 40x20 stored with
// EXIF Orientation=6, so a correct decode returns it rotated to 20x40 portrait.
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/portrait-exif6.jpg"
);
let img = decode_oriented(Path::new(path)).expect("an ordinary photo must decode");
assert_eq!(
(img.width(), img.height()),
(20, 40),
"EXIF orientation must still be applied after restoring the guard"
);
}
}

View File

@@ -9,10 +9,13 @@
//! users staring at a spinner. Resetting them on startup recovers gracefully. //! users staring at a spinner. Resetting them on startup recovers gracefully.
//! //!
//! 2. **Periodic tasks** — pruning that should happen "every hour" rather than per //! 2. **Periodic tasks** — pruning that should happen "every hour" rather than per
//! request: expired sessions (otherwise the table grows unboundedly), and the //! request: expired sessions (otherwise the table grows unboundedly), the
//! rate-limiter's in-memory windows (so keys for IPs that left long ago don't //! rate-limiter's in-memory windows (so keys for IPs that left long ago don't
//! accumulate). //! accumulate), and the media of soft-deleted uploads — both the ones whose compression
//! permanently failed and the ones a guest or host deliberately removed — which are
//! retained for a recovery window and then reclaimed.
use std::path::PathBuf;
use std::time::Duration; use std::time::Duration;
use sqlx::PgPool; use sqlx::PgPool;
@@ -20,6 +23,42 @@ use sqlx::PgPool;
use crate::services::rate_limiter::RateLimiter; use crate::services::rate_limiter::RateLimiter;
use crate::services::sse_tickets::SseTicketStore; use crate::services::sse_tickets::SseTicketStore;
/// How long a permanently-failed upload's original is kept on disk before it is
/// reclaimed.
///
/// The compression worker stops deleting originals on failure — a transient error must
/// never destroy the guest's only copy of a photo they can't retake. But the row is
/// soft-deleted and the uploader's quota IS refunded, so without a sweep those bytes are
/// invisible, unowned, and free: a reproducible codec failure lets one guest accumulate
/// orphans at no personal cost, and because `active_uploaders` counts only users with
/// non-deleted uploads, dropping out of that count actually RAISES everyone's per-user
/// ceiling while the disk gets fuller.
///
/// Two weeks is comfortably longer than any single event, so an operator investigating a
/// failed upload still has the file, while the leak stays bounded.
const FAILED_ORIGINAL_RETENTION_DAYS: i64 = 14;
/// How long a DELIBERATELY deleted upload's files are kept before they are reclaimed.
///
/// The same leak, reached by the ordinary path rather than the exceptional one.
/// `soft_delete_in_event` stamps `deleted_at` and refunds `total_upload_bytes`, but nothing ever
/// removed the bytes — so the quota stopped bounding the disk. Upload 500 MB, delete, quota is back
/// to zero, upload another 500 MB: not an attack, just a guest curating their camera roll, which is
/// what people do. The host then sees guests hitting "Du hast dein Upload-Limit erreicht" while the
/// admin widget shows a disk full of files no upload row points at, and the quota message is
/// actively misleading because the space really is gone — just not to anyone the accounting can
/// name.
///
/// Much shorter than the failure window on purpose. Fourteen days outlives the whole event, so a
/// deliberate delete would never reclaim anything while it mattered. A day still gives an operator
/// a recovery window for a mis-tap.
///
/// NOTE what this does NOT do: within the window the bytes are still spent and still unaccounted,
/// so a guest deleting and re-uploading through an eight-hour event can outrun the sweep. Bounding
/// that would mean holding the quota until the file is actually reclaimed rather than refunding at
/// `deleted_at` — a deliberate trade, and the reason the low-disk warning exists.
const DELETED_UPLOAD_RETENTION_HOURS: i64 = 24;
/// Reset rows left in flight by a previous crashed instance. Run once on startup, /// Reset rows left in flight by a previous crashed instance. Run once on startup,
/// before the HTTP server starts taking requests, so users never observe the /// before the HTTP server starts taking requests, so users never observe the
/// half-state. /// half-state.
@@ -81,24 +120,270 @@ pub async fn startup_recovery(pool: &PgPool) {
} }
} }
/// How long a file in `originals/` may exist without a database row before it is treated as
/// abandoned.
///
/// This window is the ONLY thing making the sweep safe, because the upload handler renames the
/// temp file into its final path BEFORE committing the row: for a short moment a perfectly
/// healthy upload legitimately looks exactly like an orphan. Six hours is far beyond any live
/// request (a 576 MiB body over a bad venue uplink is minutes, and the request itself is bounded
/// by the reverse proxy) while still reclaiming the leak inside a single event.
///
/// DO NOT SHORTEN THIS to make a test faster — a value below the longest possible in-flight
/// upload deletes photos out from under the request that is committing them.
const ORPHAN_UPLOAD_RETENTION_HOURS: u64 = 6;
/// Spawns a background task that periodically: /// Spawns a background task that periodically:
/// - deletes session rows whose `expires_at` is more than a day in the past /// - deletes session rows whose `expires_at` is more than a day in the past
/// - prunes the in-memory rate-limiter HashMap of empty windows /// - prunes the in-memory rate-limiter HashMap of empty windows
/// - drops expired SSE tickets (30s TTL but the map keeps the slot until pruned) /// - drops expired SSE tickets (30s TTL but the map keeps the slot until pruned)
/// ///
/// Cadence is 1h — fine for both jobs at our scale. /// Cadence is 1h — fine for both jobs at our scale.
pub fn spawn_periodic_tasks(pool: PgPool, rate_limiter: RateLimiter, sse_tickets: SseTicketStore) { pub fn spawn_periodic_tasks(
pool: PgPool,
rate_limiter: RateLimiter,
sse_tickets: SseTicketStore,
media_path: PathBuf,
) {
// Supervised, because this one task carries EVERY piece of recurring hygiene in the app:
// session pruning, media reclamation, the orphan-temp sweep, and the rate-limiter and
// SSE-ticket maps. As a bare `tokio::spawn` with no retained handle, a single panic anywhere
// inside it stopped all five permanently and silently — no log line, no symptom until the
// disk or a HashMap grew into one. The supervisor re-spawns and, just as importantly, says
// so; it can never spin hot because the inner loop only returns by dying.
tokio::spawn(async move { tokio::spawn(async move {
loop {
let inner = tokio::spawn(periodic_loop(
pool.clone(),
rate_limiter.clone(),
sse_tickets.clone(),
media_path.clone(),
));
match inner.await {
Ok(()) => tracing::error!("periodic maintenance loop returned; restarting it"),
Err(e) => {
tracing::error!(error = ?e, "periodic maintenance task died; restarting it")
}
}
tokio::time::sleep(Duration::from_secs(60)).await;
}
});
}
/// The actual hygiene loop. Never returns in normal operation — see the supervisor above.
async fn periodic_loop(
pool: PgPool,
rate_limiter: RateLimiter,
sse_tickets: SseTicketStore,
media_path: PathBuf,
) {
// A crash left whatever the previous process was mid-upload behind, and the first periodic
// tick is an hour away — sweep once up front so a restart is also a cleanup.
sweep_orphan_upload_temps(&media_path).await;
let mut tick = tokio::time::interval(Duration::from_secs(3600)); let mut tick = tokio::time::interval(Duration::from_secs(3600));
// Fire the first tick immediately, then hourly. // Fire the first tick immediately, then hourly.
tick.tick().await; tick.tick().await;
loop { loop {
tick.tick().await; tick.tick().await;
cleanup_sessions(&pool).await; cleanup_sessions(&pool).await;
cleanup_deleted_media(&pool, &media_path).await;
sweep_orphan_upload_temps(&media_path).await;
// Runs AFTER the .tmp sweep, and covers the class that one structurally cannot see:
// an original that was renamed to its final name but whose transaction never
// committed. Those have no row, so `cleanup_deleted_media` (row-driven) can never
// find them, and `sweep_orphan_upload_temps` skips them because they no longer end
// in `.tmp` — they were permanently unowned, silently shrinking the free disk that
// `compute_storage_quota` divides among guests.
sweep_orphan_originals(&pool, &media_path).await;
rate_limiter.prune(); rate_limiter.prune();
sse_tickets.prune(); sse_tickets.prune();
} }
}); }
/// How long an upload's `.tmp` file must have been untouched before it is treated as abandoned.
///
/// This is an age on the MODIFICATION time, not on creation, and that is what makes an hour
/// safe rather than reckless: a live upload is being written to continuously, so its mtime keeps
/// advancing and it can never age into the sweep no matter how slow the connection. The clock
/// only starts once the writer stops — i.e. once the upload is genuinely dead.
const ORPHAN_TEMP_MAX_AGE: Duration = Duration::from_secs(3600);
/// Reclaim `.tmp` files left in the media tree by uploads that never finished.
///
/// `stream_field_to_file` removes its temp file on every error return, which covers everything
/// the handler can see. It cannot cover the case that actually happens at a party: the client
/// simply goes away — a phone sleeps, a guest walks out of range, the PWA is evicted mid-video —
/// and axum DROPS the handler future rather than returning an error, so no cleanup code runs at
/// all. The shutdown backstop force-exits in-flight handlers for the same net effect.
///
/// Nothing else reclaims these. `cleanup_deleted_media` only visits rows with `deleted_at`, and
/// an abandoned upload never got a row; `export::sweep_orphan_temps` is only ever pointed at the
/// exports volume. So before this, every abandonment stranded up to `max_video_size_mb` of
/// unowned bytes permanently — and worse than merely leaking, they were subtracted from what
/// everyone else could upload, because the per-user quota is computed from live free disk
/// (`compute_storage_quota`). On a 40 GB disk shared with `postgres_data`, an evening of flaky
/// venue wifi could take the event down.
async fn sweep_orphan_upload_temps(media_path: &std::path::Path) {
let originals = media_path.join("originals");
let mut event_dirs = match tokio::fs::read_dir(&originals).await {
Ok(d) => d,
// Absent before the first upload — not a problem worth logging every hour.
Err(e) if e.kind() == std::io::ErrorKind::NotFound => return,
Err(e) => {
tracing::warn!(error = ?e, path = %originals.display(), "orphan temp sweep: unreadable");
return;
}
};
let mut reclaimed = 0usize;
let mut bytes = 0u64;
while let Ok(Some(event_dir)) = event_dirs.next_entry().await {
let Ok(mut files) = tokio::fs::read_dir(event_dir.path()).await else {
continue;
};
while let Ok(Some(file)) = files.next_entry().await {
let path = file.path();
if path.extension().is_none_or(|e| e != "tmp") {
continue;
}
let Ok(meta) = file.metadata().await else {
continue;
};
// No mtime (or a clock that moved backwards) means we cannot show the file is
// abandoned, and deleting a live upload is far worse than leaking one temp file.
let abandoned = meta
.modified()
.ok()
.and_then(|m| m.elapsed().ok())
.is_some_and(|age| age >= ORPHAN_TEMP_MAX_AGE);
if !abandoned {
continue;
}
match tokio::fs::remove_file(&path).await {
Ok(()) => {
reclaimed += 1;
bytes += meta.len();
}
Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
tracing::warn!(error = ?e, path = %path.display(),
"orphan temp sweep: could not reclaim")
}
}
}
}
if reclaimed > 0 {
tracing::info!(
"orphan temp sweep: reclaimed {reclaimed} abandoned upload temp file(s), {} MiB",
bytes / (1024 * 1024)
);
}
}
/// Reclaim the media of soft-deleted uploads once they are past their retention window.
///
/// ONLY ever touches rows with `deleted_at IS NOT NULL`, so it can never reach a live upload. Two
/// classes, two windows, because the two deletes mean different things:
///
/// - a compression failure the guest didn't ask for and may want investigated —
/// [`FAILED_ORIGINAL_RETENTION_DAYS`];
/// - a deliberate removal by the guest or the host — [`DELETED_UPLOAD_RETENTION_HOURS`].
///
/// ALL FOUR paths are reclaimed, not just the original. The previous version cleared
/// `original_path` alone, which was right for its only case (a failed compression produces no
/// derivatives) but wrong the moment the sweep reaches a successfully processed upload: preview,
/// display and thumbnail are each a separate file on disk, none of them counted in
/// `original_size_bytes`, and nothing else ever removed them.
///
/// Every column is cleared in the same pass, which makes the sweep idempotent and stops a later run
/// re-reporting files that are already gone. The ROW is kept: it is the audit trail, it costs a few
/// hundred bytes, and `backfill_stale_derivatives` is guarded on `deleted_at IS NULL` so a nulled
/// `preview_path` can never make it regenerate what was just reclaimed.
async fn cleanup_deleted_media(pool: &PgPool, media_path: &std::path::Path) {
type Row = (
uuid::Uuid,
String,
Option<String>,
Option<String>,
Option<String>,
);
let rows = sqlx::query_as::<_, Row>(
"SELECT id, original_path, preview_path, display_path, thumbnail_path FROM upload
WHERE deleted_at IS NOT NULL
AND CASE WHEN compression_status = 'failed'
THEN deleted_at < NOW() - ($1 || ' days')::interval
ELSE deleted_at < NOW() - ($2 || ' hours')::interval
END
AND (original_path <> '' OR preview_path IS NOT NULL
OR display_path IS NOT NULL OR thumbnail_path IS NOT NULL)",
)
.bind(FAILED_ORIGINAL_RETENTION_DAYS.to_string())
.bind(DELETED_UPLOAD_RETENTION_HOURS.to_string())
.fetch_all(pool)
.await;
let rows = match rows {
Ok(r) => r,
Err(e) => {
tracing::warn!(error = ?e, "deleted-media sweep query failed");
return;
}
};
if rows.is_empty() {
return;
}
let mut reclaimed = 0usize;
for (id, original, preview, display, thumbnail) in rows {
let paths: Vec<String> = std::iter::once(original)
.filter(|p| !p.is_empty())
.chain([preview, display, thumbnail].into_iter().flatten())
.collect();
// All-or-nothing per row: the columns are only cleared once every file for that upload is
// gone. Clearing after a partial success would strand the survivors with nothing pointing
// at them — the same unowned-bytes state this sweep exists to drain.
let mut all_gone = true;
for rel in &paths {
let absolute = media_path.join(rel);
match tokio::fs::remove_file(&absolute).await {
Ok(()) => reclaimed += 1,
// Already gone (manual cleanup, restored backup) — still counts as reclaimed for
// the purpose of clearing the columns, or the row is re-selected every hour forever.
Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
tracing::warn!(error = ?e, %id, path = %absolute.display(),
"could not reclaim deleted media; leaving the row for the next sweep");
all_gone = false;
}
}
}
if !all_gone {
continue;
}
if let Err(e) = sqlx::query(
"UPDATE upload SET original_path = '', preview_path = NULL,
display_path = NULL, thumbnail_path = NULL
WHERE id = $1",
)
.bind(id)
.execute(pool)
.await
{
tracing::warn!(error = ?e, %id, "reclaimed the files but could not clear the paths");
}
}
if reclaimed > 0 {
tracing::info!(
"reclaimed {reclaimed} file(s) from soft-deleted uploads (deliberate deletes after \
{DELETED_UPLOAD_RETENTION_HOURS}h, compression failures after \
{FAILED_ORIGINAL_RETENTION_DAYS}d)"
);
}
} }
async fn cleanup_sessions(pool: &PgPool) { async fn cleanup_sessions(pool: &PgPool) {
@@ -113,3 +398,198 @@ async fn cleanup_sessions(pool: &PgPool) {
Err(e) => tracing::warn!("session cleanup failed: {e:#}"), Err(e) => tracing::warn!("session cleanup failed: {e:#}"),
} }
} }
/// Reclaim files in `originals/` that no upload row references.
///
/// The backstop behind [`TempFileGuard`](crate::handlers::upload). The guard covers the
/// process that is running; this covers the process that was killed — a SIGKILL, an OOM, or a
/// power cut leaves whatever bytes had been written with no `Drop` to reclaim them, and those
/// files are then permanently invisible: they have no row, so `cleanup_deleted_media` (which is
/// row-driven) can never see them, and they are not counted against any quota while still
/// consuming the free disk that `compute_storage_quota` divides among guests. On a single box
/// where all three volumes share a filesystem, that ends with Postgres unable to write WAL.
///
/// Two classes:
/// - `*.tmp` — an upload that never got as far as being renamed. Always safe past the window.
/// - everything else — a final-named original whose commit never happened.
async fn sweep_orphan_originals(pool: &PgPool, media_path: &std::path::Path) {
let originals = media_path.join("originals");
let cutoff = Duration::from_secs(ORPHAN_UPLOAD_RETENTION_HOURS * 3600);
// originals/{event_slug}/{uuid}.{ext} — one level of per-event directories.
let mut event_dirs = match tokio::fs::read_dir(&originals).await {
Ok(rd) => rd,
// Nothing uploaded yet; the directory is created lazily by the upload handler.
Err(_) => return,
};
let mut candidates: Vec<(String, std::path::PathBuf)> = Vec::new();
let mut temps_removed = 0u32;
while let Ok(Some(event_dir)) = event_dirs.next_entry().await {
if !event_dir
.file_type()
.await
.map(|t| t.is_dir())
.unwrap_or(false)
{
continue;
}
let slug = event_dir.file_name().to_string_lossy().to_string();
let Ok(mut files) = tokio::fs::read_dir(event_dir.path()).await else {
continue;
};
while let Ok(Some(entry)) = files.next_entry().await {
let Ok(meta) = entry.metadata().await else {
continue;
};
if !meta.is_file() {
continue;
}
// Too young to judge: an upload committing RIGHT NOW is indistinguishable from an
// orphan, because the rename precedes the commit.
//
// This sweep is also the backstop for the one case the upload handler's drop guard
// deliberately leaks: a client disconnect while `tx.commit()` is in flight disarms
// the guard first (so a COMMIT that Postgres applied anyway keeps its file), which
// means a COMMIT that did NOT apply leaves a final-named file with no row. The
// `NOT EXISTS` check below is what reclaims it. See `upload.rs`, the disarm site.
let recent = meta
.modified()
.ok()
.and_then(|m| m.elapsed().ok())
.is_none_or(|age| age < cutoff);
if recent {
continue;
}
let name = entry.file_name().to_string_lossy().to_string();
if name.ends_with(".tmp") {
// A `.tmp` never has a row by construction — no DB check needed.
if tokio::fs::remove_file(entry.path()).await.is_ok() {
temps_removed += 1;
}
continue;
}
candidates.push((format!("originals/{slug}/{name}"), entry.path()));
}
}
if temps_removed > 0 {
tracing::warn!(
"reclaimed {temps_removed} abandoned upload temp file(s) older than \
{ORPHAN_UPLOAD_RETENTION_HOURS}h"
);
}
if candidates.is_empty() {
return;
}
// One query per batch, not one per file: a backlog of thousands of orphans must not turn
// into thousands of round trips on an hourly timer.
let mut orphans_removed = 0u32;
for chunk in candidates.chunks(500) {
let paths: Vec<String> = chunk.iter().map(|(rel, _)| rel.clone()).collect();
// NO `deleted_at IS NULL` FILTER HERE. A soft-deleted row still points at its file
// during its retention window, and reclaiming that file is `cleanup_deleted_media`'s
// job — filtering here would race the two sweeps and destroy the exact files the
// recovery window exists to preserve.
let unreferenced: Result<Vec<(String,)>, _> = sqlx::query_as(
"SELECT p FROM unnest($1::text[]) AS p
WHERE NOT EXISTS (SELECT 1 FROM upload u WHERE u.original_path = p)",
)
.bind(&paths)
.fetch_all(pool)
.await;
let unreferenced = match unreferenced {
Ok(rows) => rows,
Err(e) => {
tracing::warn!(error = ?e, "orphan-original sweep query failed");
return;
}
};
for (rel,) in unreferenced {
if let Some((_, abs)) = chunk.iter().find(|(r, _)| *r == rel)
&& tokio::fs::remove_file(abs).await.is_ok()
{
orphans_removed += 1;
}
}
}
if orphans_removed > 0 {
tracing::warn!(
"reclaimed {orphans_removed} original(s) with no upload row, older than \
{ORPHAN_UPLOAD_RETENTION_HOURS}h"
);
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Build `<root>/originals/<event>/<name>` with `len` bytes, optionally back-dating its mtime
/// by `age`. Back-dating is the only way to test the sweep without sleeping through an hour.
fn temp_file(root: &std::path::Path, name: &str, len: usize, age: Option<Duration>) {
let dir = root.join("originals").join("wedding");
std::fs::create_dir_all(&dir).expect("create dir");
let path = dir.join(name);
let f = std::fs::File::create(&path).expect("create file");
std::io::Write::write_all(&mut &f, &vec![0u8; len]).expect("write");
if let Some(age) = age {
let when = std::time::SystemTime::now() - age;
f.set_modified(when).expect("set mtime");
}
}
fn exists(root: &std::path::Path, name: &str) -> bool {
root.join("originals").join("wedding").join(name).exists()
}
/// The two halves of the guarantee in one pass: an abandoned temp is reclaimed, and a temp
/// that is still being written to is NOT — the second matters more, because deleting a live
/// upload's temp file would corrupt a photo that was about to succeed.
#[tokio::test]
async fn the_sweep_reclaims_abandoned_temps_and_spares_live_ones() {
let root = std::env::temp_dir().join(format!("es-sweep-{}", uuid::Uuid::new_v4()));
// Abandoned: the writer died over an hour ago and nothing has touched it since.
temp_file(
&root,
"dead.tmp",
2048,
Some(ORPHAN_TEMP_MAX_AGE + Duration::from_secs(60)),
);
// Live: an upload in progress keeps advancing its mtime, so it always looks young —
// this is why the threshold is on modification time and not on creation time.
temp_file(&root, "inflight.tmp", 2048, None);
// A committed original. The sweep must only ever consider `.tmp`.
temp_file(&root, "keeper.jpg", 2048, Some(Duration::from_secs(86_400)));
sweep_orphan_upload_temps(&root).await;
assert!(
!exists(&root, "dead.tmp"),
"an abandoned temp must be reclaimed"
);
assert!(
exists(&root, "inflight.tmp"),
"a temp still being written to must survive — deleting it destroys a live upload"
);
assert!(
exists(&root, "keeper.jpg"),
"the sweep must never touch a committed original"
);
let _ = std::fs::remove_dir_all(&root);
}
/// Runs on every boot and every hour, so a media tree that does not exist yet (before the
/// first upload) must be a silent no-op rather than an error logged 24 times a day.
#[tokio::test]
async fn a_missing_media_tree_is_not_an_error() {
let root = std::env::temp_dir().join(format!("es-sweep-absent-{}", uuid::Uuid::new_v4()));
sweep_orphan_upload_temps(&root).await; // must simply return
}
}

View File

@@ -0,0 +1,112 @@
//! Cached sum of all media bytes the event is holding.
//!
//! The upload gate needs to know "how big would the keepsake be if we accept this file", because
//! the archive needs room for BOTH halves at once (`export::required_free_bytes` is
//! `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized). Asking that question
//! per upload has to be cheap, and it has to be cheap on the busiest write path in the app.
//!
//! `export::estimate_export_bytes` answers the same question exactly, but it aggregates
//! `original_size_bytes` across every upload row joined to `user` — fine once per release,
//! wasteful per upload and growing all evening. This sums `user.total_upload_bytes` instead:
//! one row per guest (~100), already maintained transactionally by the quota path, already
//! refunded on delete.
//!
//! The two differ slightly — this one counts uploads belonging to banned or hidden users, which
//! the export filters out. That skew is in the SAFE direction: it over-estimates the archive, so
//! the gate closes marginally early rather than marginally late. Never swap it for a cheaper
//! query that could under-estimate; an under-estimate authorises the very upload that makes the
//! keepsake unbuildable, which is the failure this exists to prevent.
use std::sync::{Arc, RwLock};
use std::time::{Duration, Instant};
use sqlx::PgPool;
/// How long a reading is trusted. Shorter than [`crate::services::disk`]'s TTL because this
/// number only ever grows and does so in the same request path that reads it — a stale value
/// under-counts the newest uploads, and under-counting is the direction that matters.
const TTL: Duration = Duration::from_secs(5);
/// Cheap-to-clone cache of the event's total media bytes. Lives in `AppState`.
#[derive(Clone)]
pub struct MediaTotalCache {
inner: Arc<RwLock<Option<(i64, Instant)>>>,
}
impl MediaTotalCache {
pub fn new() -> Self {
Self {
inner: Arc::new(RwLock::new(None)),
}
}
/// Drop the cached reading so the next `get()` re-queries.
///
/// Used by the e2e TRUNCATE endpoint for the same reason `DiskCache::invalidate` exists:
/// truncation removes every upload, and a surviving reading would make the next test's
/// gate compute against the previous test's data.
pub fn invalidate(&self) {
*self.inner.write().unwrap() = None;
}
/// Total bytes of media the event is holding, cached for [`TTL`].
///
/// Returns 0 when the query fails. That is a deliberate FAIL-OPEN, consistent with the
/// quota path and the export preflight: a database blip must not turn into "every upload
/// refused". The disk-space half of the gate still applies, so a failure here degrades the
/// check to the old flat-reserve behaviour rather than disabling it.
pub async fn get(&self, pool: &PgPool, event_slug: &str) -> i64 {
if let Some((bytes, at)) = *self.inner.read().unwrap()
&& at.elapsed() < TTL
{
return bytes;
}
// Scoped to THIS event (H12). The unscoped `SUM(total_upload_bytes) FROM "user"` summed
// every user row in the table, so reusing the install for a second event carried the first
// one's bytes into the second one's keepsake-headroom gate — closing uploads early with a
// message about "the event's storage" being full, counting media that belongs to a party
// that already happened (and whose files are never reclaimed either).
let queried = sqlx::query_scalar::<_, Option<i64>>(
"SELECT SUM(u.total_upload_bytes)::bigint FROM \"user\" u
JOIN event e ON e.id = u.event_id
WHERE e.slug = $1",
)
.bind(event_slug)
.fetch_one(pool)
.await;
let bytes = match queried {
Ok(v) => v.unwrap_or(0).max(0),
Err(e) => {
// FAIL OPEN, but do NOT cache the failure, and do NOT let it pass silently.
//
// Storing 0 here pinned the gate's view of the event at "empty" for the whole
// TTL. During that window `media_after` is just this upload, `keepsake_needs`
// collapses to ~2.2x one file, and the gate degrades to the flat 10 GB reserve —
// precisely the behaviour the two-halves design replaced, reappearing with no
// trace in the log. And the trigger correlates with the danger: with
// `max_connections = 10` and a 5s acquire timeout, this query fails exactly when
// a burst is in progress.
//
// Falling back to the LAST GOOD reading (however stale) is strictly better than
// 0: the total only ever grows, so a stale value under-counts slightly, while 0
// under-counts by everything.
let previous = self.inner.read().unwrap().map(|(b, _)| b);
tracing::warn!(
error = %e,
fallback_bytes = previous.unwrap_or(0),
"media total query failed; upload gate is running on a stale reading"
);
return previous.unwrap_or(0);
}
};
*self.inner.write().unwrap() = Some((bytes, Instant::now()));
bytes
}
}
impl Default for MediaTotalCache {
fn default() -> Self {
Self::new()
}
}

View File

@@ -1,7 +1,12 @@
pub mod audit;
pub mod compression; pub mod compression;
pub mod config; pub mod config;
pub mod disk; pub mod disk;
pub mod export; pub mod export;
pub mod imaging;
pub mod maintenance; pub mod maintenance;
pub mod media_total;
pub mod rate_limiter; pub mod rate_limiter;
pub mod sse_tickets; pub mod sse_tickets;
pub mod upload_admission;
pub mod video;

View File

@@ -7,7 +7,23 @@ use std::time::{Duration, Instant};
/// of recent requests and rejects new ones once the window is full. /// of recent requests and rejects new ones once the window is full.
#[derive(Clone)] #[derive(Clone)]
pub struct RateLimiter { pub struct RateLimiter {
windows: Arc<Mutex<HashMap<String, Vec<Instant>>>>, windows: Arc<Mutex<HashMap<String, Bucket>>>,
}
/// One key's recent hits, plus the window they were recorded under.
///
/// The `window` field is what makes pruning correct. `prune` used a single fixed 24 h ceiling for
/// every key, on the reasoning that 24 h is the longest window in use (export downloads) — but that
/// meant a `join:{ip}:{name}` key whose 60-SECOND window expired 23 hours ago was still retained.
/// Minting one costs a single 409 and no bcrypt, at 60/min per IP across three endpoints: roughly
/// 172,800 keys/day/IP, about 31 MB/day/IP inside a 1 GB container. The limiter became the
/// memory-exhaustion primitive it exists to prevent.
///
/// Storing the window per key makes the sweep drop each bucket as soon as ITS OWN window has
/// elapsed, which is also what the hot path already does on every check.
struct Bucket {
hits: Vec<Instant>,
window: Duration,
} }
impl RateLimiter { impl RateLimiter {
@@ -17,13 +33,14 @@ impl RateLimiter {
} }
} }
/// Returns `true` if the request is allowed, `false` if rate-limited.
pub fn check(&self, key: impl Into<String>, max: usize, window: Duration) -> bool {
self.check_with_retry(key, max, window).is_ok()
}
/// Returns `Ok(())` if allowed, `Err(retry_after_secs)` if rate-limited. /// Returns `Ok(())` if allowed, `Err(retry_after_secs)` if rate-limited.
/// `retry_after_secs` is how long until the oldest slot in the window expires. /// `retry_after_secs` is how long until the oldest slot in the window expires.
///
/// This is deliberately the ONLY entry point. There used to be a `check()` wrapper
/// returning a plain bool, and 7 of the 8 call sites used it and then hard-coded
/// `None` for the response's `Retry-After` — so a throttled client was told to back
/// off but never for how long. Forcing every caller through the `Result` makes the
/// retry delay impossible to discard by accident.
pub fn check_with_retry( pub fn check_with_retry(
&self, &self,
key: impl Into<String>, key: impl Into<String>,
@@ -33,20 +50,68 @@ impl RateLimiter {
let now = Instant::now(); let now = Instant::now();
let key = key.into(); let key = key.into();
let mut map = self.windows.lock().unwrap(); let mut map = self.windows.lock().unwrap();
let timestamps = map.entry(key).or_default(); let bucket = map.entry(key).or_insert_with(|| Bucket {
hits: Vec::new(),
window,
});
// A key's window can change under it when an admin edits the limit at runtime. Track the
// current one so `prune` expires the bucket on the window actually in force.
bucket.window = window;
let timestamps = &mut bucket.hits;
timestamps.retain(|&t| now.duration_since(t) < window); timestamps.retain(|&t| now.duration_since(t) < window);
if timestamps.len() < max { if timestamps.len() < max {
timestamps.push(now); timestamps.push(now);
Ok(()) Ok(())
} else { } else {
// The oldest timestamp expires at oldest + window; compute remaining seconds // The oldest timestamp expires at oldest + window; compute remaining seconds.
let oldest = timestamps[0]; //
// `first()`, not `[0]`: with `max == 0` the length check above is false even on an
// empty vec, so indexing would panic — WHILE HOLDING THIS MUTEX. That poisons it
// process-wide, so every subsequent `.lock().unwrap()` panics too: upload, feed,
// join, recover, social, export and the hourly maintenance task all die, and only
// a container restart brings them back. `max == 0` is not reachable through the
// admin API (every numeric spec has min = 1) but a direct DB edit would do it, and
// the blast radius does not justify the sharper syntax.
let Some(&oldest) = timestamps.first() else {
return Ok(());
};
let elapsed = now.duration_since(oldest); let elapsed = now.duration_since(oldest);
let remaining = window.saturating_sub(elapsed); let remaining = window.saturating_sub(elapsed);
Err(remaining.as_secs().max(1)) Err(remaining.as_secs().max(1))
} }
} }
/// Is `key` already at or above `max`, WITHOUT recording a hit?
///
/// Needed by limiters whose budget is spent by an outcome rather than by the request — the
/// per-IP failed-PIN ceiling charges only on a wrong PIN, so the gate at the top of the handler
/// has to be able to ask "is this IP shut out?" without itself consuming the budget it guards.
/// Using `check_with_retry` for that would charge every *successful* recovery too, and a venue
/// full of guests legitimately recovering their own devices would lock itself out.
///
/// Returns `Err(retry_after_secs)` when exhausted, mirroring `check_with_retry` so callers can
/// build the same 429.
pub fn peek(&self, key: &str, max: usize, window: Duration) -> Result<(), u64> {
let now = Instant::now();
let mut map = self.windows.lock().unwrap();
let Some(bucket) = map.get_mut(key) else {
return Ok(());
};
bucket
.hits
.retain(|&t| now.duration_since(t) < bucket.window);
if bucket.hits.len() < max {
return Ok(());
}
let Some(&oldest) = bucket.hits.first() else {
return Ok(());
};
Err(window
.saturating_sub(now.duration_since(oldest))
.as_secs()
.max(1))
}
/// Wipe every tracked window. Used by the test-mode truncate route so a previous /// Wipe every tracked window. Used by the test-mode truncate route so a previous
/// test's accumulated counters don't bleed into the next test's rate-limit checks. /// test's accumulated counters don't bleed into the next test's rate-limit checks.
pub fn clear(&self) { pub fn clear(&self) {
@@ -57,17 +122,23 @@ impl RateLimiter {
/// background task (see [`crate::services::maintenance`]) so that long-lived /// background task (see [`crate::services::maintenance`]) so that long-lived
/// processes don't accumulate one HashMap entry per IP that ever connected. /// processes don't accumulate one HashMap entry per IP that ever connected.
/// ///
/// Uses a conservative 24h ceiling — anything older than that is gone regardless /// Expires each bucket against ITS OWN window (see [`Bucket`]), not one global ceiling. The
/// of which endpoint's window it was tracked under (the longest window today is /// previous fixed 24 h ceiling retained per-minute keys for a full day — ~172,800 keys/day/IP
/// 24h for export downloads). If we ever add longer windows, raise this constant. /// at 60/min across three endpoints, each mintable with a single 409 and no bcrypt.
///
/// Holds the one global mutex for the length of the sweep, and that mutex is on the hot path of
/// upload, feed, join, recover, social and export — so the retain does the cheap thing per
/// bucket and nothing else. Correct pruning also keeps the map small enough that this stays
/// cheap, which the old ceiling actively undermined.
pub fn prune(&self) { pub fn prune(&self) {
let now = Instant::now(); let now = Instant::now();
let ceiling = Duration::from_secs(24 * 60 * 60);
let mut map = self.windows.lock().unwrap(); let mut map = self.windows.lock().unwrap();
let before = map.len(); let before = map.len();
map.retain(|_, ts| { map.retain(|_, bucket| {
ts.retain(|&t| now.duration_since(t) < ceiling); bucket
!ts.is_empty() .hits
.retain(|&t| now.duration_since(t) < bucket.window);
!bucket.hits.is_empty()
}); });
let dropped = before.saturating_sub(map.len()); let dropped = before.saturating_sub(map.len());
if dropped > 0 { if dropped > 0 {
@@ -84,6 +155,11 @@ impl RateLimiter {
/// appends is the real client. A client can prepend arbitrary spoofed values to /// appends is the real client. A client can prepend arbitrary spoofed values to
/// the left of XFF to dodge throttles — those are ignored here. This assumes /// the left of XFF to dodge throttles — those are ignored here. This assumes
/// exactly one trusted proxy (Caddy); revisit if that changes. /// exactly one trusted proxy (Caddy); revisit if that changes.
///
/// Pass the peer address as `fallback`, never a constant. Every caller used to pass
/// the literal `"unknown"`, so any request that arrived without XFF — i.e. anything
/// reaching the app directly rather than through Caddy — shared ONE bucket with every
/// other such request, turning the limiter into a self-inflicted global throttle.
pub fn client_ip(headers: &axum::http::HeaderMap, fallback: &str) -> String { pub fn client_ip(headers: &axum::http::HeaderMap, fallback: &str) -> String {
headers headers
.get("x-forwarded-for") .get("x-forwarded-for")
@@ -104,29 +180,35 @@ mod tests {
#[test] #[test]
fn allows_up_to_max_then_blocks() { fn allows_up_to_max_then_blocks() {
let rl = RateLimiter::new(); let rl = RateLimiter::new();
assert!(rl.check("k", 3, MIN)); assert!(rl.check_with_retry("k", 3, MIN).is_ok());
assert!(rl.check("k", 3, MIN)); assert!(rl.check_with_retry("k", 3, MIN).is_ok());
assert!(rl.check("k", 3, MIN)); assert!(rl.check_with_retry("k", 3, MIN).is_ok());
assert!(!rl.check("k", 3, MIN), "the 4th request must be blocked"); assert!(
rl.check_with_retry("k", 3, MIN).is_err(),
"the 4th request must be blocked"
);
} }
#[test] #[test]
fn keys_are_independent() { fn keys_are_independent() {
let rl = RateLimiter::new(); let rl = RateLimiter::new();
assert!(rl.check("a", 1, MIN)); assert!(rl.check_with_retry("a", 1, MIN).is_ok());
assert!(!rl.check("a", 1, MIN)); assert!(rl.check_with_retry("a", 1, MIN).is_err());
assert!(rl.check("b", 1, MIN), "a different key has its own window"); assert!(
rl.check_with_retry("b", 1, MIN).is_ok(),
"a different key has its own window"
);
} }
#[test] #[test]
fn window_slides_and_allows_again_after_expiry() { fn window_slides_and_allows_again_after_expiry() {
let rl = RateLimiter::new(); let rl = RateLimiter::new();
let w = Duration::from_millis(40); let w = Duration::from_millis(40);
assert!(rl.check("k", 1, w)); assert!(rl.check_with_retry("k", 1, w).is_ok());
assert!(!rl.check("k", 1, w)); assert!(rl.check_with_retry("k", 1, w).is_err());
std::thread::sleep(Duration::from_millis(55)); std::thread::sleep(Duration::from_millis(55));
assert!( assert!(
rl.check("k", 1, w), rl.check_with_retry("k", 1, w).is_ok(),
"the slot should expire once the window passes" "the slot should expire once the window passes"
); );
} }
@@ -191,10 +273,13 @@ mod tests {
#[test] #[test]
fn clear_resets_every_window() { fn clear_resets_every_window() {
let rl = RateLimiter::new(); let rl = RateLimiter::new();
assert!(rl.check("k", 1, MIN)); assert!(rl.check_with_retry("k", 1, MIN).is_ok());
assert!(!rl.check("k", 1, MIN)); assert!(rl.check_with_retry("k", 1, MIN).is_err());
rl.clear(); rl.clear();
assert!(rl.check("k", 1, MIN), "clear() must free the window"); assert!(
rl.check_with_retry("k", 1, MIN).is_ok(),
"clear() must free the window"
);
} }
/// `prune()` is a memory-leak guard: without it a long-lived process keeps one HashMap /// `prune()` is a memory-leak guard: without it a long-lived process keeps one HashMap
@@ -205,18 +290,23 @@ mod tests {
fn prune_drops_keys_whose_windows_have_fully_expired() { fn prune_drops_keys_whose_windows_have_fully_expired() {
let rl = RateLimiter::new(); let rl = RateLimiter::new();
// A key whose only timestamp is older than the 24h ceiling. We can't sleep for a day, // A key whose only timestamp is older than its own window. We can't sleep, so backdate
// so backdate the Instant directly. // the Instant directly. A ONE-MINUTE window here on purpose: the old prune applied a flat
// 24 h ceiling to every key, so this bucket — expired for over an hour of wall time —
// survived the sweep. That is the leak (H3), and pinning it needs a short-window key.
let ancient = Instant::now() let ancient = Instant::now()
.checked_sub(Duration::from_secs(25 * 60 * 60)) .checked_sub(Duration::from_secs(90 * 60))
.expect("backdating an Instant by 25h"); .expect("backdating an Instant by 90 minutes");
rl.windows rl.windows.lock().unwrap().insert(
.lock() "stale".to_string(),
.unwrap() Bucket {
.insert("stale".to_string(), vec![ancient]); hits: vec![ancient],
window: MIN,
},
);
// ...alongside a key that is still inside its window. // ...alongside a key that is still inside its window.
assert!(rl.check("live", 5, MIN)); assert!(rl.check_with_retry("live", 5, MIN).is_ok());
assert_eq!(rl.windows.lock().unwrap().len(), 2); assert_eq!(rl.windows.lock().unwrap().len(), 2);
rl.prune(); rl.prune();
@@ -239,13 +329,13 @@ mod tests {
// prune() dropped live keys, every background sweep would hand attackers a fresh // prune() dropped live keys, every background sweep would hand attackers a fresh
// budget. // budget.
let rl = RateLimiter::new(); let rl = RateLimiter::new();
assert!(rl.check("k", 1, MIN)); assert!(rl.check_with_retry("k", 1, MIN).is_ok());
assert!(!rl.check("k", 1, MIN)); assert!(rl.check_with_retry("k", 1, MIN).is_err());
rl.prune(); rl.prune();
assert!( assert!(
!rl.check("k", 1, MIN), rl.check_with_retry("k", 1, MIN).is_err(),
"prune() must not clear a window that is still active" "prune() must not clear a window that is still active"
); );
} }

View File

@@ -13,6 +13,88 @@ use rand::Rng;
/// stream open. Tickets are consumed on use and expire after `TTL`. /// stream open. Tickets are consumed on use and expire after `TTL`.
const TTL: Duration = Duration::from_secs(30); const TTL: Duration = Duration::from_secs(30);
/// Lifetime of a `Download` ticket, and it is deliberately far longer than [`TTL`].
///
/// A keepsake is up to ~1.4 GB over hotel or cellular wifi, so the download itself outlives a 30 s
/// window many times over — and a resumed transfer arrives minutes or hours after the ticket was
/// minted. A `Download` ticket is therefore a short-lived capability for ONE archive rather than a
/// single-shot nonce: [`SseTicketStore::redeem_download`] does not remove it, so a client may
/// resume with `Range` as many times as the transfer needs.
///
/// The abuse this does NOT open: the ticket is bound to a session (revoked with it), only mints at
/// `/export/ticket` where the 3/day limit is charged, and grants nothing but this event's own
/// keepsake — which every authenticated guest is entitled to download anyway. What it buys is that
/// one dropped connection no longer costs a guest a third of their daily allowance, at the
/// emotional payoff of the product.
const DOWNLOAD_TTL: Duration = Duration::from_secs(6 * 60 * 60);
/// The lifetime that applies to a given kind.
fn ttl_for(kind: TicketKind) -> Duration {
match kind {
TicketKind::Download(_) => DOWNLOAD_TTL,
TicketKind::Sse => TTL,
}
}
/// Ceiling on outstanding tickets across the whole process.
///
/// Not really about the bytes (~120 each) — about `issue` having had no bound of any kind.
/// Sized well above a real event: ~1000 concurrent clients each holding one live 30 s ticket.
const MAX_TICKETS: usize = 4096;
/// Live tickets one session may hold. Above 1 because two tabs sharing a token open their
/// EventSources concurrently; 4 absorbs that without letting a reconnect loop accumulate.
const MAX_TICKETS_PER_SESSION: usize = 4;
/// How many times one download ticket may be redeemed.
///
/// Making the ticket non-consuming is what lets a dropped transfer resume without spending another
/// of the guest's three daily downloads — but unbounded it also meant a single mint was an
/// unlimited download key for six hours, at BOTH archive endpoints, with the per-day limiter
/// (charged only at mint) never moving. On a 40 GB box serving ~1.4 GB archives that is the one
/// resource an ordinary guest could exhaust without doing anything obviously wrong.
///
/// 20 is far more than a resumed transfer needs (a browser retries a handful of times, not dozens)
/// and turns "unbounded until the ticket expires" into a bounded multiple. It does not make the
/// daily limit exact — that would mean charging per redemption, which would bill a client that
/// restarts from byte 0 instead of sending a `Range`, i.e. re-break the thing this exists to fix.
const MAX_DOWNLOAD_REDEMPTIONS: u32 = 20;
/// What a ticket may be redeemed for.
///
/// The store began life serving only SSE and stayed untyped when the export download started
/// reusing it, which silently made the two interchangeable. That is not a theoretical mixing
/// concern: `POST /stream/ticket` is rate-limited at 60/min per user and charges nothing, while
/// `POST /export/ticket` charges one of three PER-DAY downloads. An untyped ticket let any guest
/// mint at the cheap endpoint and redeem at the expensive one, so the daily export limit was
/// bypassable ~60×/minute — each redemption streaming the whole multi-GB keepsake, `no-store`,
/// off the same filesystem Postgres writes WAL to.
///
/// `consume` therefore requires the kind to MATCH. A ticket is only ever valid for the thing it
/// was minted for.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
pub enum TicketKind {
/// Opens the SSE stream (`GET /stream`). Cheap, high volume.
Sse,
/// Downloads ONE export archive. Expensive, rate-limited per day.
///
/// The archive is part of the ticket, not incidental to it. A bare `Download` ticket was
/// accepted by BOTH `/export/zip` and `/export/html` — they share one authenticator — so with
/// the redemption budget that makes a ticket resumable, a single mint authorised 20 transfers
/// spread across both archives. Three mints a day therefore bought 60 full downloads of a
/// ~1.4 GB keepsake, while the per-day limiter (charged only at mint) never moved.
Download(ExportKind),
}
/// Which archive a [`TicketKind::Download`] is good for.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
pub enum ExportKind {
/// `Gallery.<event>.<n>.zip` — the original media.
Zip,
/// `Memories.<event>.<n>.zip` — the offline HTML viewer.
Html,
}
#[derive(Clone)] #[derive(Clone)]
pub struct SseTicketStore { pub struct SseTicketStore {
inner: Arc<Mutex<HashMap<String, Entry>>>, inner: Arc<Mutex<HashMap<String, Entry>>>,
@@ -22,6 +104,10 @@ pub struct SseTicketStore {
struct Entry { struct Entry {
token_hash: String, token_hash: String,
issued_at: Instant, issued_at: Instant,
kind: TicketKind,
/// Times this ticket has been redeemed. Only meaningful for `Download` — see
/// [`MAX_DOWNLOAD_REDEMPTIONS`].
redemptions: u32,
} }
impl SseTicketStore { impl SseTicketStore {
@@ -39,26 +125,129 @@ impl SseTicketStore {
} }
/// Mint a new ticket bound to the caller's session (identified by token hash). /// Mint a new ticket bound to the caller's session (identified by token hash).
pub fn issue(&self, token_hash: String) -> String { ///
/// `None` when the store is at capacity — the caller should answer 503, not evict.
///
/// Three bounds, because `issue` had none: no size cap, no per-caller cap, and no rate
/// limit on the endpoint, while `prune` ran only hourly against a 30-second TTL. So any
/// authenticated session could loop the endpoint and grow the map for an hour.
pub fn issue(&self, token_hash: String, kind: TicketKind) -> Option<String> {
let ticket = random_ticket(); let ticket = random_ticket();
let mut map = self.inner.lock().unwrap(); let mut map = self.inner.lock().unwrap();
// Prune on issue rather than only hourly. This alone changes the bound from "tickets
// minted since the last maintenance tick" to "tickets live at once", which is what the
// 30 s TTL was always meant to express.
map.retain(|_, e| e.issued_at.elapsed() <= ttl_for(e.kind));
// Cap the caller's own outstanding tickets, evicting their oldest. NOT one-per-session:
// two tabs sharing a token open their EventSources concurrently, and having tab B
// invalidate tab A's unconsumed ticket looks exactly like a flaky SSE connection.
//
// SCOPED TO THE SAME KIND, which matters now that `Download` tickets live 6 h instead of
// being consumed on first use. A long-lived download ticket is ALWAYS the oldest entry for
// its session, so a kind-blind cap made it the first thing ordinary SSE churn threw away —
// and `/export` itself opens an SSE connection on the very session that just minted it.
// A couple of reconnects during the download (a wifi flap is enough; each attempt that
// returns early abandons an unconsumed ticket) evicted the ticket out from under a running
// transfer, so the next `Range` resume 401'd and the guest had to spend another of their
// three daily downloads. Two flaps and they were locked out of their own keepsake for a day.
//
// Per-kind, an SSE reconnect storm can still only evict SSE tickets, which is what the cap
// was written for; the guest's in-flight keepsake is no longer collateral.
let mut mine: Vec<(String, Instant)> = map
.iter()
.filter(|(_, e)| e.token_hash == token_hash && e.kind == kind)
.map(|(k, e)| (k.clone(), e.issued_at))
.collect();
if mine.len() >= MAX_TICKETS_PER_SESSION {
mine.sort_by_key(|(_, issued)| *issued);
for (key, _) in mine.iter().take(mine.len() - MAX_TICKETS_PER_SESSION + 1) {
map.remove(key);
}
}
// At capacity, REFUSE — never evict a stranger's ticket. Evicting would let one
// misbehaving client deny SSE to the whole venue, which is worse than failing the
// request that hit the ceiling.
if map.len() >= MAX_TICKETS {
tracing::warn!(
outstanding = map.len(),
"SSE ticket store at capacity; refusing to mint"
);
return None;
}
map.insert( map.insert(
ticket.clone(), ticket.clone(),
Entry { Entry {
token_hash, token_hash,
issued_at: Instant::now(), issued_at: Instant::now(),
kind,
redemptions: 0,
}, },
); );
ticket Some(ticket)
} }
/// Consume a ticket. Returns `Some(token_hash)` if the ticket exists and is /// Consume a ticket minted for `kind`. Returns `Some(token_hash)` if the ticket exists, is
/// not expired. Single-use: the ticket is removed regardless of whether it /// not expired, and was minted for this purpose. Single-use: the ticket is removed regardless
/// was still fresh, so a replay can't slip through after expiry. /// of whether it was still fresh, so a replay can't slip through after expiry.
pub fn consume(&self, ticket: &str) -> Option<String> { ///
/// A ticket of the WRONG kind is also removed. It was a valid ticket the caller legitimately
/// held, so this is not punitive — but leaving it would let a redemption loop probe the store
/// without ever spending anything, and the client has no legitimate reason to present a
/// ticket at the wrong endpoint.
/// Redeem a `Download` ticket WITHOUT consuming it.
///
/// Downloads must be resumable — see [`DOWNLOAD_TTL`]. A browser resumes by re-issuing the same
/// GET with a `Range` header, so a single-use ticket made `Accept-Ranges` a lie: the retry
/// authenticated against a ticket that the interrupted attempt had already spent, 401'd, and
/// the guest had to mint a new one, spending another of their three daily downloads. Two
/// dropped connections and they were locked out of their own keepsake for ~24 hours.
///
/// Still bound to a live session: the caller re-checks the session on every request, so
/// revoking a session (logout, "sign out everywhere", a host PIN reset) kills the download too.
pub fn redeem_download(&self, ticket: &str, want: ExportKind) -> Option<String> {
let mut map = self.inner.lock().unwrap();
let entry = map.get_mut(ticket)?;
// The archive must match the one this ticket was minted for. Both download routes share
// this authenticator, so without the payload check a ZIP ticket opened the HTML archive
// too and the redemption budget was spent across both.
if entry.kind != TicketKind::Download(want) {
tracing::warn!(
found = ?entry.kind,
?want,
"ticket presented for the wrong archive (or wrong kind); rejected"
);
return None;
}
if entry.issued_at.elapsed() > DOWNLOAD_TTL {
return None;
}
if entry.redemptions >= MAX_DOWNLOAD_REDEMPTIONS {
tracing::warn!(
redemptions = entry.redemptions,
"download ticket exceeded its redemption budget; refusing"
);
return None;
}
entry.redemptions += 1;
Some(entry.token_hash.clone())
}
pub fn consume(&self, ticket: &str, kind: TicketKind) -> Option<String> {
let mut map = self.inner.lock().unwrap(); let mut map = self.inner.lock().unwrap();
let entry = map.remove(ticket)?; let entry = map.remove(ticket)?;
if entry.issued_at.elapsed() > TTL { if entry.issued_at.elapsed() > ttl_for(entry.kind) {
return None;
}
if entry.kind != kind {
tracing::warn!(
expected = ?kind,
found = ?entry.kind,
"ticket presented at the wrong endpoint; rejected"
);
return None; return None;
} }
Some(entry.token_hash) Some(entry.token_hash)
@@ -68,7 +257,7 @@ impl SseTicketStore {
/// long-running process doesn't accumulate stale tickets. /// long-running process doesn't accumulate stale tickets.
pub fn prune(&self) { pub fn prune(&self) {
let mut map = self.inner.lock().unwrap(); let mut map = self.inner.lock().unwrap();
map.retain(|_, e| e.issued_at.elapsed() <= TTL); map.retain(|_, e| e.issued_at.elapsed() <= ttl_for(e.kind));
} }
} }
@@ -84,14 +273,66 @@ fn random_ticket() -> String {
mod tests { mod tests {
use super::*; use super::*;
/// `issue` now returns `Option`; in every test below the store is far from capacity, so an
/// `expect` here documents that refusing is exceptional rather than routine.
fn issue(store: &SseTicketStore, hash: &str) -> String {
store
.issue(hash.into(), TicketKind::Sse)
.expect("store has capacity")
}
/// The store is shared by two endpoints with wildly different costs: `/stream/ticket` is
/// 60/min per user and free, `/export/ticket` charges one of three PER-DAY downloads. While
/// entries were untyped, a ticket minted at the cheap endpoint opened the expensive one — so
/// the daily export limit could be bypassed ~60×/minute, each redemption streaming the whole
/// multi-GB keepsake off the disk Postgres writes WAL to.
///
/// Asserted in BOTH directions so this cannot be "fixed" by a check that only guards one.
#[test]
fn a_ticket_is_only_valid_for_the_purpose_it_was_minted_for() {
let store = SseTicketStore::new();
let sse = store.issue("h".into(), TicketKind::Sse).unwrap();
assert_eq!(
store.consume(&sse, TicketKind::Download(ExportKind::Zip)),
None,
"an SSE ticket must not open the export download"
);
let dl = store
.issue("h".into(), TicketKind::Download(ExportKind::Zip))
.unwrap();
assert_eq!(
store.consume(&dl, TicketKind::Sse),
None,
"a download ticket must not open the SSE stream"
);
// And the matching cases still work, so the guard is not simply rejecting everything.
let sse = store.issue("h".into(), TicketKind::Sse).unwrap();
assert_eq!(store.consume(&sse, TicketKind::Sse).as_deref(), Some("h"));
let dl = store
.issue("h".into(), TicketKind::Download(ExportKind::Zip))
.unwrap();
assert_eq!(
store
.consume(&dl, TicketKind::Download(ExportKind::Zip))
.as_deref(),
Some("h")
);
}
#[test] #[test]
fn issue_then_consume_returns_the_hash_exactly_once() { fn issue_then_consume_returns_the_hash_exactly_once() {
let store = SseTicketStore::new(); let store = SseTicketStore::new();
let ticket = store.issue("hash-1".into()); let ticket = issue(&store, "hash-1");
assert_eq!(store.consume(&ticket).as_deref(), Some("hash-1")); assert_eq!(
store.consume(&ticket, TicketKind::Sse).as_deref(),
Some("hash-1")
);
// Single-use: a replay of the same ticket is rejected. // Single-use: a replay of the same ticket is rejected.
assert_eq!( assert_eq!(
store.consume(&ticket), store.consume(&ticket, TicketKind::Sse),
None, None,
"a consumed ticket must not be reusable" "a consumed ticket must not be reusable"
); );
@@ -100,14 +341,14 @@ mod tests {
#[test] #[test]
fn unknown_ticket_consumes_to_none() { fn unknown_ticket_consumes_to_none() {
let store = SseTicketStore::new(); let store = SseTicketStore::new();
assert_eq!(store.consume("never-issued"), None); assert_eq!(store.consume("never-issued", TicketKind::Sse), None);
} }
#[test] #[test]
fn issued_tickets_are_unique_and_hex() { fn issued_tickets_are_unique_and_hex() {
let store = SseTicketStore::new(); let store = SseTicketStore::new();
let a = store.issue("h".into()); let a = issue(&store, "h");
let b = store.issue("h".into()); let b = issue(&store, "h");
assert_ne!(a, b, "each ticket must be unique"); assert_ne!(a, b, "each ticket must be unique");
assert_eq!(a.len(), 48, "24 random bytes → 48 hex chars"); assert_eq!(a.len(), 48, "24 random bytes → 48 hex chars");
assert!(a.chars().all(|c| c.is_ascii_hexdigit())); assert!(a.chars().all(|c| c.is_ascii_hexdigit()));
@@ -116,29 +357,210 @@ mod tests {
#[test] #[test]
fn fresh_ticket_survives_prune() { fn fresh_ticket_survives_prune() {
let store = SseTicketStore::new(); let store = SseTicketStore::new();
let ticket = store.issue("h".into()); let ticket = issue(&store, "h");
store.prune(); // not expired → kept store.prune(); // not expired → kept
assert_eq!(store.consume(&ticket).as_deref(), Some("h")); assert_eq!(
store.consume(&ticket, TicketKind::Sse).as_deref(),
Some("h")
);
}
/// Build an entry that is already past the TTL.
fn insert_stale(store: &SseTicketStore, key: &str, token_hash: &str) {
store.inner.lock().unwrap().insert(
key.to_string(),
Entry {
kind: TicketKind::Sse,
token_hash: token_hash.into(),
issued_at: Instant::now()
.checked_sub(TTL + Duration::from_secs(1))
.expect("host uptime should exceed the ticket TTL"),
redemptions: 0,
},
);
} }
#[test] #[test]
fn expired_ticket_consumes_to_none() { fn expired_ticket_consumes_to_none() {
// Construct an entry that is already past the TTL and confirm consume() rejects it.
let store = SseTicketStore::new(); let store = SseTicketStore::new();
let stale = "stale-ticket".to_string(); insert_stale(&store, "stale-ticket", "h");
store.inner.lock().unwrap().insert(
stale.clone(),
Entry {
token_hash: "h".into(),
issued_at: Instant::now()
.checked_sub(TTL + Duration::from_secs(1))
.expect("host uptime should exceed the ticket TTL"),
},
);
assert_eq!( assert_eq!(
store.consume(&stale), store.consume("stale-ticket", TicketKind::Sse),
None, None,
"an expired ticket must not authenticate" "an expired ticket must not authenticate"
); );
} }
/// The TTL is 30 s but `prune` only ran hourly, so the map was really bounded by "tickets
/// minted in the last hour" — which is unbounded for a client in a loop.
#[test]
fn issuing_prunes_expired_entries() {
let store = SseTicketStore::new();
insert_stale(&store, "stale-a", "someone-else");
insert_stale(&store, "stale-b", "someone-else");
issue(&store, "h");
assert_eq!(
store.inner.lock().unwrap().len(),
1,
"issue must reclaim expired slots, not merely add to them"
);
}
/// Two tabs sharing a token is normal, so the per-session cap must be above 1 — but a
/// reconnect loop must not accumulate. The caller's OWN oldest is what gets evicted.
#[test]
fn a_session_is_capped_and_evicts_only_its_own_oldest() {
let store = SseTicketStore::new();
let stranger = issue(&store, "other-session");
let mut mine: Vec<String> = Vec::new();
for _ in 0..MAX_TICKETS_PER_SESSION + 2 {
mine.push(issue(&store, "mine"));
}
let live = mine
.iter()
.filter(|t| store.inner.lock().unwrap().contains_key(*t))
.count();
assert_eq!(live, MAX_TICKETS_PER_SESSION, "one session, bounded");
assert!(
store
.inner
.lock()
.unwrap()
.contains_key(&mine[mine.len() - 1]),
"the newest ticket is the one the caller is about to use"
);
assert_eq!(
store.consume(&stranger, TicketKind::Sse).as_deref(),
Some("other-session"),
"another session's ticket must survive — evicting it would let one client deny \
SSE to the venue"
);
}
/// At capacity the store REFUSES rather than evicting a stranger. Refusing fails the one
/// request that hit the ceiling; evicting would break an unrelated client's live stream.
#[test]
fn at_capacity_the_store_refuses_instead_of_evicting() {
let store = SseTicketStore::new();
{
let mut map = store.inner.lock().unwrap();
for i in 0..MAX_TICKETS {
map.insert(
format!("filler-{i}"),
Entry {
kind: TicketKind::Sse,
token_hash: format!("session-{i}"),
issued_at: Instant::now(),
redemptions: 0,
},
);
}
}
assert_eq!(
store.issue("newcomer".into(), TicketKind::Sse),
None,
"a full store must refuse, so the caller can answer 503"
);
assert!(
store.inner.lock().unwrap().contains_key("filler-0"),
"no existing ticket may be sacrificed to make room"
);
}
/// A download ticket is deliberately non-consuming so a dropped 1.4 GB transfer can resume
/// without spending one of the guest's three daily downloads. Unbounded, though, that made a
/// single mint an unlimited download key for six hours while the per-day limiter — charged
/// only at mint — never moved. This pins the bound without breaking resumption.
#[test]
fn a_download_ticket_resumes_freely_but_not_forever() {
let store = SseTicketStore::new();
let ticket = store
.issue("session-a".into(), TicketKind::Download(ExportKind::Zip))
.expect("fresh store should issue");
// Every redemption inside the budget returns the session, so a resumed transfer works.
for i in 0..MAX_DOWNLOAD_REDEMPTIONS {
assert_eq!(
store.redeem_download(&ticket, ExportKind::Zip).as_deref(),
Some("session-a"),
"redemption {i} should still be honoured"
);
}
// Past it the ticket is spent: the guest re-mints (and is charged) rather than holding
// an open-ended key.
assert_eq!(store.redeem_download(&ticket, ExportKind::Zip), None);
}
#[test]
fn a_download_ticket_opens_only_the_archive_it_was_minted_for() {
// Both download routes share one authenticator, so without the archive in the ticket a
// single mint was good for BOTH. Combined with the resume budget that made one mint worth
// 2 x MAX_DOWNLOAD_REDEMPTIONS transfers of a multi-GB keepsake, while the per-day limit —
// charged only at mint — never moved.
let store = SseTicketStore::new();
let zip = store
.issue("s".into(), TicketKind::Download(ExportKind::Zip))
.unwrap();
assert_eq!(
store.redeem_download(&zip, ExportKind::Html),
None,
"a ZIP ticket must not open the HTML archive"
);
assert_eq!(
store.redeem_download(&zip, ExportKind::Zip).as_deref(),
Some("s"),
"...and the refusal above must be about the archive, not a spent ticket"
);
let html = store
.issue("s".into(), TicketKind::Download(ExportKind::Html))
.unwrap();
assert_eq!(store.redeem_download(&html, ExportKind::Zip), None);
assert_eq!(
store.redeem_download(&html, ExportKind::Html).as_deref(),
Some("s")
);
}
#[test]
fn sse_churn_cannot_evict_a_running_download() {
// A `Download` ticket lives 6 h, so it is ALWAYS the oldest entry for its session — and a
// kind-blind per-session cap therefore threw it away first. `/export` opens its own SSE
// connection on the same session, so a couple of reconnects during the transfer evicted
// the ticket out from under it: the next `Range` resume 401'd and the guest spent another
// of their three daily downloads. Two wifi flaps and they lost their keepsake for a day.
let store = SseTicketStore::new();
let download = store
.issue("one-session".into(), TicketKind::Download(ExportKind::Zip))
.unwrap();
// Far more SSE churn than the per-session cap, all on the same session.
for _ in 0..(MAX_TICKETS_PER_SESSION * 3) {
store.issue("one-session".into(), TicketKind::Sse).unwrap();
}
assert_eq!(
store.redeem_download(&download, ExportKind::Zip).as_deref(),
Some("one-session"),
"an in-flight keepsake download must survive an SSE reconnect storm"
);
}
/// The kind split is what stops a free SSE ticket from redeeming a rate-limited download.
#[test]
fn an_sse_ticket_is_never_redeemable_as_a_download() {
let store = SseTicketStore::new();
let sse = store
.issue("session-b".into(), TicketKind::Sse)
.expect("fresh store should issue");
assert_eq!(store.redeem_download(&sse, ExportKind::Zip), None);
// And it is still usable for what it IS, so the rejection above is about kind, not
// the ticket having been quietly spent.
assert_eq!(
store.consume(&sse, TicketKind::Sse).as_deref(),
Some("session-b")
);
}
} }

View File

@@ -0,0 +1,150 @@
//! Admission control for upload bodies, budgeted in BYTES rather than requests.
//!
//! ## Why this has to exist
//!
//! The keepsake headroom gate in `handlers::upload` cannot bound a burst, and the reason is
//! structural rather than a bug in the gate: the request body is streamed to a temp file during
//! multipart parsing, so the bytes are already on disk by the time any check runs. The gate can
//! only refuse to COMMIT them. Nothing upstream limited how many bodies stream at once — axum has
//! no such limit, the tower stack is just `TraceLayer`, and Caddy passes requests straight
//! through.
//!
//! So the failure mode is the ordinary one, not an attack: the ceremony ends, ~100 guests tap
//! "upload all", and ~100 bodies stream concurrently. At phone-video sizes that is 10-20 GB of
//! `.tmp` files on a 40 GB volume, none of it visible to the gate, and `DISK_RESERVE_BYTES` — the
//! 10 GB standing between the party and Postgres losing the volume it writes WAL to — is consumed
//! by transient files. The `.tmp` sweeper only reclaims files idle for an hour, correctly, which
//! means nothing reclaims a burst on this timescale.
//!
//! ## Why bytes and not a request count
//!
//! A flat "N concurrent uploads" limit has to be sized for the worst case (a 500 MB video), which
//! makes it absurdly restrictive for the common case (a 3 MB photo). Budgeting bytes lets one
//! 500 MB video and two hundred photos coexist under the same ceiling, and it means the ceiling is
//! stated in the unit the disk actually cares about.
//!
//! The reservation is the streaming CAP, not the real size — the real size is unknowable until the
//! body has been read, which is far too late. Reserving the cap is deliberately pessimistic; that
//! pessimism is the safety margin.
//!
//! ## Why a permit and not a counter
//!
//! `OwnedSemaphorePermit` releases on drop. Every path out of the upload handler — success, error,
//! a client vanishing mid-body, a panic — therefore returns the reservation without any explicit
//! bookkeeping. A hand-rolled `AtomicI64` would need a decrement on each of those paths, and the
//! one that gets missed is the one that leaks the budget until restart.
use std::sync::Arc;
use std::time::Duration;
use tokio::sync::{OwnedSemaphorePermit, Semaphore};
/// Total transient upload bytes allowed on disk at once, in MiB.
///
/// Sized against `DISK_RESERVE_BYTES` (10 GB): the reserve must survive a full burst with room to
/// spare, since Postgres is writing WAL to the same filesystem throughout. 4 GiB leaves ~6 GB of
/// the reserve untouched at the worst moment.
///
/// It is NOT a throughput limit. On 2 vCPU the box cannot usefully ingest more than this at once
/// anyway — compression, ffmpeg, Postgres and TLS all contend for the same two cores — so the
/// budget mostly converts "everything is slow and the disk fills" into "a few uploads wait".
const BUDGET_MIB: u32 = 4096;
/// How long an upload waits for room before being told to come back.
///
/// Long enough to absorb the burst (a photo holds its reservation for well under a second), short
/// enough that a guest is not left staring at a spinner. On timeout the handler answers 503 with
/// `Retry-After`, which the client queue already treats as transient and retries with backoff.
const WAIT: Duration = Duration::from_secs(20);
#[derive(Clone)]
pub struct UploadAdmission {
permits: Arc<Semaphore>,
}
impl UploadAdmission {
pub fn new() -> Self {
Self {
permits: Arc::new(Semaphore::new(BUDGET_MIB as usize)),
}
}
/// Reserve room for a body capped at `cap_bytes`. The returned permit must be held for as long
/// as the temp file exists.
///
/// `None` means the wait timed out and the caller should shed the request.
///
/// A cap larger than the whole budget is clamped rather than refused. Otherwise an operator
/// raising `max_video_size_mb` above the budget would make `acquire_many` unsatisfiable and
/// every video upload would hang until timeout — a config change silently disabling video for
/// the event. Clamped, such an upload simply gets the whole budget to itself, which is the
/// honest interpretation of "one file may fill the machine".
pub async fn reserve(&self, cap_bytes: usize) -> Option<OwnedSemaphorePermit> {
let mib = cap_bytes.div_ceil(1024 * 1024).max(1);
let want = u32::try_from(mib).unwrap_or(BUDGET_MIB).min(BUDGET_MIB);
match tokio::time::timeout(WAIT, self.permits.clone().acquire_many_owned(want)).await {
Ok(Ok(permit)) => Some(permit),
// The semaphore is never closed, so `Err` here is unreachable in practice; treat it
// the same as a timeout rather than panicking on the upload path.
Ok(Err(_)) => None,
Err(_) => {
tracing::warn!(
requested_mib = want,
"upload admission timed out; shedding to keep transient temp files bounded"
);
None
}
}
}
}
impl Default for UploadAdmission {
fn default() -> Self {
Self::new()
}
}
#[cfg(test)]
mod tests {
use super::*;
/// The budget must actually block once exhausted — otherwise this whole module is decoration.
#[tokio::test]
async fn a_full_budget_sheds_instead_of_admitting() {
let admission = UploadAdmission::new();
let whole = admission
.reserve(BUDGET_MIB as usize * 1024 * 1024)
.await
.expect("first reservation takes the whole budget");
// Nothing left: a second reservation must not be granted. Raced against a short timeout so
// the test does not sit for the full WAIT.
let blocked =
tokio::time::timeout(Duration::from_millis(150), admission.reserve(1024 * 1024)).await;
assert!(
blocked.is_err(),
"budget exhausted, yet a reservation was granted"
);
// ...and releasing the permit makes room again, so the budget is not a one-way latch.
drop(whole);
assert!(
admission.reserve(1024 * 1024).await.is_some(),
"budget did not recover after the permit was dropped"
);
}
/// A cap above the whole budget must be clamped, not left unsatisfiable. Unclamped,
/// `acquire_many` for more permits than exist never completes, so raising
/// `max_video_size_mb` past the budget would silently hang every video upload for 20s and
/// then shed it.
#[tokio::test]
async fn a_cap_larger_than_the_budget_is_clamped_rather_than_unsatisfiable() {
let admission = UploadAdmission::new();
let oversized = (BUDGET_MIB as usize + 4096) * 1024 * 1024;
assert!(
admission.reserve(oversized).await.is_some(),
"an over-budget cap must still be admittable on an idle server"
);
}
}

View File

@@ -0,0 +1,263 @@
//! Poster-frame extraction, shared by the compression worker and the HTML export.
//!
//! Both used to spawn `ffmpeg` themselves with the same broken invocation:
//!
//! ```text
//! ffmpeg -i <src> -vframes 1 -ss 00:00:01 -vf scale=… -y <out>
//! ```
//!
//! `-ss` AFTER `-i` is an output-side seek. Against a clip of a second or less ffmpeg exits **0 and
//! writes nothing** — and both call sites gated on the exit status, so neither noticed. The worker
//! then wrote `thumbnail_path` for a file that was never created (404 in the live feed) and the
//! export listed the entry in `data.json` while the ZIP writer skipped it (a broken image tile in
//! the keepsake). Every server-side signal stayed green. Phones produce such clips constantly:
//! mis-taps, Live Photos, boomerangs.
//!
//! This module exists for the same reason `imaging.rs` does — that one was created when compression
//! and export duplicated decode logic, and it paid off immediately when the `max_alloc` fix landed
//! in both workers at once. Same duplication, same fix.
use std::path::Path;
use std::time::Duration;
use anyhow::{Context, Result};
/// A malformed video can hang `ffmpeg` indefinitely. In the compression worker that never releases
/// the semaphore permit and the pool eventually deadlocks; in the export worker it strands the job
/// at `running` so the keepsake never completes. `export.rs` had NO timeout at all before this
/// module — sharing the spawn fixes that too.
/// 45s, not the 120s this started at. The timeout is not a budget for honest work — a poster
/// frame from a phone clip takes well under a second, and `-ss` before `-i` means even a 500 MB
/// file seeks rather than scans. It is purely the ceiling on how long a pathological input may
/// hold a compression permit that guests' photos are queued behind, so it should be as tight as
/// it can be without ever cutting off real work.
const FFMPEG_TIMEOUT: Duration = Duration::from_secs(45);
/// Seek positions to try, in order.
///
/// One second first: the opening frame of a real video is often black, a fade-in, or motion-blurred
/// as the camera settles, so it makes a poor poster. Zero second as the fallback, which is what
/// makes short clips work — and it is genuinely required, not defensive. Moving `-ss` before `-i`
/// (an input-side seek) is necessary but NOT sufficient: seeking to 1 s in a 1.000 s clip is still
/// past the last frame, and ffmpeg still exits 0 having written nothing. Verified against the real
/// production image.
const SEEK_POSITIONS: &[&str] = &["00:00:01", "0"];
/// Extract one poster frame from `src` into `dest`, scaled to `width` px wide.
///
/// `Ok(false)` means the video yielded no frame — a normal outcome for a very short or unusual
/// clip, NOT an error. Callers must degrade (no poster) rather than fail the upload: treating this
/// as an error would soft-delete every sub-second video, turning a cosmetic defect into data loss.
///
/// `Err` is reserved for something genuinely wrong — a hang we had to kill, or a failure to spawn.
pub async fn extract_poster_frame(src: &Path, dest: &Path, width: u32) -> Result<bool> {
for seek in SEEK_POSITIONS {
// A stale file from a previous attempt would be indistinguishable from a fresh success.
let _ = tokio::fs::remove_file(dest).await;
run_ffmpeg(src, dest, width, seek).await?;
// THE CHECK BOTH CALL SITES WERE MISSING: ask the filesystem, not the exit status.
// Non-empty, because a zero-byte file is not a poster either.
if tokio::fs::metadata(dest)
.await
.map(|m| m.is_file() && m.len() > 0)
.unwrap_or(false)
{
return Ok(true);
}
}
// Leave nothing behind for a caller to mistake for a result.
let _ = tokio::fs::remove_file(dest).await;
Ok(false)
}
/// Run one ffmpeg attempt. A non-zero exit is NOT an error here — the artifact check above is the
/// authority, and a corrupt input that fails at 1 s may still yield a frame at 0.
async fn run_ffmpeg(src: &Path, dest: &Path, width: u32, seek: &str) -> Result<()> {
let child = tokio::process::Command::new("ffmpeg")
.args([
// BEFORE -i: an input-side seek. See SEEK_POSITIONS.
"-ss",
seek,
"-i",
src.to_str().unwrap_or_default(),
"-vframes",
"1",
"-vf",
&format!("scale={width}:-1"),
"-y",
dest.to_str().unwrap_or_default(),
])
// ffmpeg writes the poster to `dest` itself; nothing here ever reads stdout, so
// giving it a pipe only created something that could fill.
.stdout(std::process::Stdio::null())
// stderr IS piped — it is the only diagnostic when a clip yields no frame — but it
// must be DRAINED, which is the whole point of `wait_with_output` below.
.stderr(std::process::Stdio::piped())
.kill_on_drop(true)
.spawn()
.context("failed to spawn ffmpeg")?;
// `wait_with_output`, NOT `wait`. ffmpeg is verbose on stderr (banner, stream info,
// per-frame progress) and `wait()` reads neither pipe — so once the ~64 KiB pipe buffer
// filled, ffmpeg blocked writing, `wait()` never returned, and the call burned the full
// timeout. That is not merely slow: the timeout is an `Err`, so after 2 seek positions x
// 3 compression attempts the caller soft-deletes a perfectly playable video for a
// poster-frame failure. `wait_with_output` polls the pipe and the exit status together.
//
// It also CONSUMES the child, so the explicit `child.kill()` that used to sit on the
// timeout arm cannot exist here — and is not needed: `kill_on_drop(true)` is set above,
// and dropping the future on timeout drops the child with it.
let out = match tokio::time::timeout(FFMPEG_TIMEOUT, child.wait_with_output()).await {
Ok(res) => res.context("ffmpeg wait failed")?,
Err(_) => anyhow::bail!("ffmpeg timed out after {}s", FFMPEG_TIMEOUT.as_secs()),
};
// A non-zero exit is not an error (see the doc comment) — the artifact check in
// `extract_poster_frame` is the authority. Log the tail so a systematically failing
// format is diagnosable without turning it into data loss.
if !out.status.success() {
tracing::debug!(
seek,
status = ?out.status,
stderr = %tail_lines(&out.stderr, 10),
"ffmpeg exited non-zero; the artifact check decides"
);
}
Ok(())
}
/// Last `n` lines of a child's stderr, lossily decoded.
///
/// Bounded on purpose: ffmpeg's stderr is unbounded, and the reason we now drain it is that
/// unbounded output used to be a hazard. Emitting all of it into a log line — into container
/// logs that are themselves size-capped — would just move the problem.
fn tail_lines(bytes: &[u8], n: usize) -> String {
let text = String::from_utf8_lossy(bytes);
let lines: Vec<&str> = text.lines().filter(|l| !l.trim().is_empty()).collect();
lines[lines.len().saturating_sub(n)..].join(" | ")
}
#[cfg(test)]
mod tests {
use super::*;
/// Is there a usable `ffmpeg` on PATH?
///
/// The poster-frame path shells out, and `extract_poster_frame` documents `Err` as meaning
/// "a hang or a SPAWN failure" — which is exactly what a missing binary produces. So on a
/// machine without ffmpeg the test below stops exercising the case it names (missing INPUT)
/// and instead reports a code defect that isn't there. The runtime image installs ffmpeg
/// (see backend/Dockerfile), so this only ever skips on a bare developer machine.
fn ffmpeg_available() -> bool {
std::process::Command::new("ffmpeg")
.arg("-version")
.stdout(std::process::Stdio::null())
.stderr(std::process::Stdio::null())
.status()
.is_ok()
}
/// The order is the whole fix. `-ss` must precede `-i`, and 0 must be tried after 1 s.
#[test]
fn the_fallback_seek_exists_and_comes_last() {
assert_eq!(
SEEK_POSITIONS,
&["00:00:01", "0"],
"1s first for a better poster, 0 as the fallback that makes short clips work"
);
}
/// A missing input yields no frame rather than an error: the caller must degrade to "no
/// poster", never fail the upload. `Err` is reserved for a hang or a spawn failure.
#[tokio::test]
async fn a_missing_source_yields_no_frame_rather_than_an_error() {
if !ffmpeg_available() {
eprintln!(
"SKIP a_missing_source_yields_no_frame_rather_than_an_error: no ffmpeg on PATH. \
A missing binary is a spawn failure, which this function returns Err for by \
design, so the missing-INPUT case cannot be exercised here. Install ffmpeg to \
run it (the runtime image already has it)."
);
return;
}
let dir = std::env::temp_dir().join(format!("es-video-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
let dest = dir.join("out.jpg");
let got = extract_poster_frame(Path::new("/nonexistent/clip.mp4"), &dest, 400).await;
match got {
Ok(false) => {}
other => panic!("expected Ok(false) for a missing input, got {other:?}"),
}
assert!(
!dest.exists(),
"a failed extraction must leave nothing a caller could mistake for a poster"
);
let _ = std::fs::remove_dir_all(&dir);
}
#[test]
fn the_stderr_tail_is_bounded_and_survives_invalid_utf8() {
let noisy: Vec<u8> = (0..500)
.map(|i| format!("line {i}\n"))
.collect::<String>()
.into_bytes();
let got = tail_lines(&noisy, 3);
assert_eq!(got, "line 497 | line 498 | line 499");
// ffmpeg emits filenames verbatim, so its stderr is not guaranteed to be UTF-8.
assert_eq!(tail_lines(&[b'o', b'k', 0xff], 5), "ok\u{fffd}");
assert_eq!(tail_lines(b"", 5), "");
}
/// A real extraction must finish in a small fraction of `FFMPEG_TIMEOUT`.
///
/// Wall-clock is the ONLY observable of the bug this guards: piping stderr and then
/// calling `wait()` (which drains nothing) blocks ffmpeg on a full pipe buffer until the
/// timeout fires, and the timeout is an `Err`, so the upload is soft-deleted. The
/// assertion is deliberately on elapsed time, not on the exit status.
///
/// Honest limitation: our fixture is quiet enough not to fill a 64 KiB pipe on its own,
/// so this catches a regression to `wait()` only in combination with a verbose input. It
/// is still worth pinning — a reverted drain plus any chatty clip is data loss.
#[tokio::test]
async fn a_real_clip_yields_a_poster_well_inside_the_timeout() {
if tokio::process::Command::new("ffmpeg")
.arg("-version")
.stdout(std::process::Stdio::null())
.stderr(std::process::Stdio::null())
.status()
.await
.is_err()
{
eprintln!("skipping: ffmpeg not on PATH");
return;
}
let src = Path::new("../e2e/fixtures/media/sample.mp4");
if !src.exists() {
eprintln!("skipping: {} missing", src.display());
return;
}
let dir = std::env::temp_dir().join(format!("es-video-ok-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
let dest = dir.join("poster.jpg");
let started = std::time::Instant::now();
let got = extract_poster_frame(src, &dest, 400).await;
let elapsed = started.elapsed();
assert!(matches!(got, Ok(true)), "expected a poster, got {got:?}");
assert!(dest.metadata().unwrap().len() > 0);
assert!(
elapsed < FFMPEG_TIMEOUT / 4,
"extraction took {elapsed:?}; a drained stderr finishes in well under \
{FFMPEG_TIMEOUT:?} — this is the pipe-deadlock regression guard"
);
let _ = std::fs::remove_dir_all(&dir);
}
}

View File

@@ -5,8 +5,10 @@ use crate::config::AppConfig;
use crate::services::compression::CompressionWorker; use crate::services::compression::CompressionWorker;
use crate::services::config::ConfigCache; use crate::services::config::ConfigCache;
use crate::services::disk::DiskCache; use crate::services::disk::DiskCache;
use crate::services::media_total::MediaTotalCache;
use crate::services::rate_limiter::RateLimiter; use crate::services::rate_limiter::RateLimiter;
use crate::services::sse_tickets::SseTicketStore; use crate::services::sse_tickets::SseTicketStore;
use crate::services::upload_admission::UploadAdmission;
#[derive(Clone, Debug)] #[derive(Clone, Debug)]
pub struct SseEvent { pub struct SseEvent {
@@ -38,6 +40,12 @@ pub struct AppState {
pub config_cache: ConfigCache, pub config_cache: ConfigCache,
/// Cached total/free bytes for the media filesystem (quota + admin stats). /// Cached total/free bytes for the media filesystem (quota + admin stats).
pub disk_cache: DiskCache, pub disk_cache: DiskCache,
/// Cached sum of all media bytes, for the upload gate's keepsake-headroom check.
pub media_total: MediaTotalCache,
/// Byte budget for upload bodies currently streaming to temp files. The headroom gate can
/// only refuse to COMMIT bytes that are already on disk; this is what bounds how many get
/// there at once.
pub upload_admission: UploadAdmission,
} }
impl AppState { impl AppState {
@@ -63,6 +71,8 @@ impl AppState {
sse_tickets: SseTicketStore::new(), sse_tickets: SseTicketStore::new(),
config_cache, config_cache,
disk_cache: DiskCache::new(), disk_cache: DiskCache::new(),
media_total: MediaTotalCache::new(),
upload_admission: UploadAdmission::new(),
} }
} }
} }

View File

@@ -1 +0,0 @@
export const env={}

File diff suppressed because one or more lines are too long

View File

@@ -1 +0,0 @@
import{u as o,n as t,o as c}from"./CcONa1Mr.js";function u(e){throw new Error("https://svelte.dev/e/lifecycle_outside_component")}function r(e){t===null&&u(),o(()=>{const n=c(e);if(typeof n=="function")return n})}export{r as o};

View File

@@ -1 +0,0 @@
import{f as l,g as o,p as u,i as n,j as d,k as m,h as p,e as _,m as v,l as k}from"./CcONa1Mr.js";class w{anchor;#t=new Map;#s=new Map;#e=new Map;#i=new Set;#f=!0;constructor(t,s=!0){this.anchor=t,this.#f=s}#a=t=>{if(this.#t.has(t)){var s=this.#t.get(t),e=this.#s.get(s);if(e)l(e),this.#i.delete(s);else{var f=this.#e.get(s);f&&(this.#s.set(s,f.effect),this.#e.delete(s),f.fragment.lastChild.remove(),this.anchor.before(f.fragment),e=f.effect)}for(const[i,a]of this.#t){if(this.#t.delete(i),i===t)break;const r=this.#e.get(a);r&&(o(r.effect),this.#e.delete(a))}for(const[i,a]of this.#s){if(i===s||this.#i.has(i))continue;const r=()=>{if(Array.from(this.#t.values()).includes(i)){var c=document.createDocumentFragment();v(a,c),c.append(n()),this.#e.set(i,{effect:a,fragment:c})}else o(a);this.#i.delete(i),this.#s.delete(i)};this.#f||!e?(this.#i.add(i),u(a,r,!1)):r()}}};#r=t=>{this.#t.delete(t);const s=Array.from(this.#t.values());for(const[e,f]of this.#e)s.includes(e)||(o(f.effect),this.#e.delete(e))};ensure(t,s){var e=m,f=k();if(s&&!this.#s.has(t)&&!this.#e.has(t))if(f){var i=document.createDocumentFragment(),a=n();i.append(a),this.#e.set(t,{effect:d(()=>s(a)),fragment:i})}else this.#s.set(t,d(()=>s(this.anchor)));if(this.#t.set(e,t),f){for(const[r,h]of this.#s)r===t?e.unskip_effect(h):e.skip_effect(h);for(const[r,h]of this.#e)r===t?e.unskip_effect(h.effect):e.skip_effect(h.effect);e.oncommit(this.#a),e.ondiscard(this.#r)}else p&&(this.anchor=_),this.#a(e)}}export{w as B};

File diff suppressed because one or more lines are too long

View File

@@ -1 +0,0 @@
import{b as c,h as o,a as l,E as b,r as p,s as v,c as g,d,e as m}from"./CcONa1Mr.js";import{B as y}from"./BRDva_z9.js";function k(f,h,_=!1){var n;o&&(n=m,l());var s=new y(f),u=_?b:0;function t(a,r){if(o){var e=p(n);if(a!==parseInt(e.substring(1))){var i=v();g(i),s.anchor=i,d(!1),s.ensure(a,r),d(!0);return}}s.ensure(a,r)}c(()=>{var a=!1;h((r,e=0)=>{a=!0,t(e,r)}),a||t(-1,null)},u)}export{k as i};

File diff suppressed because one or more lines are too long

View File

@@ -1 +0,0 @@
import{A as v,i as d,B as l,C as u,D as T,T as p,F as h,h as i,e as s,R as E,a as y,G as g,c as w,H as N}from"./CcONa1Mr.js";const A=globalThis?.window?.trustedTypes&&globalThis.window.trustedTypes.createPolicy("svelte-trusted-html",{createHTML:t=>t});function M(t){return A?.createHTML(t)??t}function x(t){var r=v("template");return r.innerHTML=M(t.replaceAll("<!>","<!---->")),r.content}function n(t,r){var e=l;e.nodes===null&&(e.nodes={start:t,end:r,a:null,t:null})}function b(t,r){var e=(r&p)!==0,f=(r&h)!==0,a,_=!t.startsWith("<!>");return()=>{if(i)return n(s,null),s;a===void 0&&(a=x(_?t:"<!>"+t),e||(a=u(a)));var o=f||T?document.importNode(a,!0):a.cloneNode(!0);if(e){var c=u(o),m=o.lastChild;n(c,m)}else n(o,o);return o}}function C(t=""){if(!i){var r=d(t+"");return n(r,r),r}var e=s;return e.nodeType!==g?(e.before(e=d()),w(e)):N(e),n(e,e),e}function O(){if(i)return n(s,null),s;var t=document.createDocumentFragment(),r=document.createComment(""),e=d();return t.append(r,e),n(r,e),t}function P(t,r){if(i){var e=l;((e.f&E)===0||e.nodes.end===null)&&(e.nodes.end=s),y();return}t!==null&&t.before(r)}const L="5";typeof window<"u"&&((window.__svelte??={}).v??=new Set).add(L);export{P as a,n as b,O as c,b as f,C as t};

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View File

@@ -1 +0,0 @@
import{l as o,a as r}from"../chunks/eAGLaJx1.js";export{o as load_css,r as start};

View File

@@ -1 +0,0 @@
import{c as s,a as c}from"../chunks/RsTAN2PN.js";import{b as l,E as p,t as i}from"../chunks/CcONa1Mr.js";import{B as m}from"../chunks/BRDva_z9.js";function u(n,r,...e){var o=new m(n);l(()=>{const t=r()??null;o.ensure(t,t&&(a=>t(a,...e)))},p)}const f=!0,_=!1,g=Object.freeze(Object.defineProperty({__proto__:null,prerender:f,ssr:_},Symbol.toStringTag,{value:"Module"}));function h(n,r){var e=s(),o=i(e);u(o,()=>r.children),c(n,e)}export{h as component,g as universal};

View File

@@ -1 +0,0 @@
import{a as i,f as h}from"../chunks/RsTAN2PN.js";import{q as g,t as v,v as d,w as l,x as s,y as a,z as x}from"../chunks/CcONa1Mr.js";import{s as o}from"../chunks/Bb9JxzU7.js";import{s as _,p}from"../chunks/eAGLaJx1.js";const $={get error(){return p.error},get status(){return p.status}};_.updated.check;const m=$;var k=h("<h1> </h1> <p> </p>",1);function z(c,f){g(f,!0);var t=k(),r=v(t),n=s(r,!0);a(r);var e=x(r,2),u=s(e,!0);a(e),d(()=>{o(n,m.status),o(u,m.error?.message)}),i(c,t),l()}export{z as component};

File diff suppressed because one or more lines are too long

View File

@@ -1 +0,0 @@
{"version":"1778876725548"}

File diff suppressed because one or more lines are too long

View File

@@ -255,3 +255,91 @@ pub async fn downloadable(pool: &PgPool, event_id: Uuid, export_type: &str) -> O
.expect("downloadable") .expect("downloadable")
.flatten() .flatten()
} }
/// Insert an upload of `size` bytes, optionally already soft-deleted.
pub async fn seed_upload(
pool: &PgPool,
event_id: Uuid,
user_id: Uuid,
size: i64,
deleted: bool,
) -> Uuid {
sqlx::query_scalar(
"INSERT INTO upload (event_id, user_id, original_path, mime_type,
original_size_bytes, deleted_at)
VALUES ($1, $2, 'originals/x.jpg', 'image/jpeg', $3,
CASE WHEN $4 THEN NOW() ELSE NULL END)
RETURNING id",
)
.bind(event_id)
.bind(user_id)
.bind(size)
.bind(deleted)
.fetch_one(pool)
.await
.expect("seed upload")
}
/// Flip the moderation flags a ban sets.
pub async fn set_user_moderation(pool: &PgPool, user_id: Uuid, banned: bool, hidden: bool) {
sqlx::query("UPDATE \"user\" SET is_banned = $2, uploads_hidden = $3 WHERE id = $1")
.bind(user_id)
.bind(banned)
.bind(hidden)
.execute(pool)
.await
.expect("set moderation");
}
/// SRC: `services/export.rs::query_uploads` — the visibility filter, verbatim, projected down to
/// `(id, original_size_bytes)`. This is the row set that ACTUALLY lands in the archives.
///
/// Production builds this WHERE from `export_visibility_where!()`, shared with
/// `estimate_export_bytes`. A copy here can pin the behaviour but CANNOT detect production moving
/// away from it — that is what sharing the fragment is for, not this.
pub async fn export_visible_uploads(pool: &PgPool, event_id: Uuid) -> Vec<(Uuid, i64)> {
sqlx::query_as(
"SELECT u.id, u.original_size_bytes
FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE
GROUP BY u.id, usr.display_name
ORDER BY u.created_at ASC",
)
.bind(event_id)
.fetch_all(pool)
.await
.expect("export_visible_uploads")
}
/// SRC: `services/export.rs::estimate_export_bytes` — verbatim. Same caveat as above: production
/// shares its WHERE with `query_uploads` via `export_visibility_where!()`, so these two copies
/// agreeing proves the behaviour, not the absence of drift.
pub async fn estimate_export_bytes(pool: &PgPool, event_id: Uuid) -> i64 {
let (bytes,): (i64,) = sqlx::query_as(
"SELECT COALESCE(SUM(u.original_size_bytes), 0)::bigint
FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE",
)
.bind(event_id)
.fetch_one(pool)
.await
.expect("estimate_export_bytes");
bytes
}
/// SRC: `services/export.rs::ensure_export_space` — the armed-job count, verbatim.
pub async fn armed_job_count(pool: &PgPool, event_id: Uuid) -> i64 {
let (n,): (i64,) = sqlx::query_as(
"SELECT COUNT(*) FROM export_job
WHERE event_id = $1 AND status IN ('pending', 'running')",
)
.bind(event_id)
.fetch_one(pool)
.await
.expect("armed_job_count");
n
}

View File

@@ -0,0 +1,163 @@
//! DB-backed tests for the export disk preflight.
//!
//! The keepsake used to be built with NO free-space check at all, and the failure that produced was
//! not "the export failed" but "the deliverable is stuck and the escape hatch needs the space that
//! isn't there":
//!
//! 1. A takedown bumps the epoch and re-arms both halves.
//! 2. The ZIP hits ENOSPC partway through a multi-GB write.
//! 3. The job row is now `failed` at the CURRENT epoch, so readiness
//! (`epoch = event.export_epoch AND status = 'done'`) is false and `GET /export/zip` 404s —
//! while the last good archive sits on disk, unreferenced and unreachable.
//! 4. `POST /host/export/rebuild` re-arms the same doomed write.
//!
//! Two changes close it: reclaim the superseded generation BEFORE building (so peak usage is one
//! generation, not two) and refuse up front with a number the host can act on.
//!
//! What these tests pin is the ESTIMATE — the part that decides. The arithmetic on top of it lives
//! in `services/export.rs`'s unit tests; the filesystem selection lives in `is_superseded_archive`.
//!
//! ON DRIFT, precisely, because it is easy to overclaim here. The hazard is that `query_uploads`
//! (which selects the rows the archives are built from) and `estimate_export_bytes` (which sizes
//! them) could disagree — and an estimate missing rows the archive writes UNDER-reserves, the one
//! direction that reintroduces the ENOSPC. **These tests cannot catch that**, and neither can any
//! test in this harness: both sides here are `SRC:`-marked hand-copies in `tests/common/mod.rs`,
//! so if production moved and the copies didn't, they would sit still and keep passing.
//!
//! That is fixed where it can be — the two queries now share one `export_visibility_where!()`
//! fragment in `services/export.rs`, so they cannot diverge by construction. What is left for
//! these tests is what the convention is genuinely good at: pinning the BEHAVIOUR, so a change
//! that deliberately alters the filter has to come here and say so.
mod common;
use common::*;
use sqlx::PgPool;
/// The estimate must equal the sum over EXACTLY the rows `query_uploads` returns — computed from
/// that row set, not from a restatement of its WHERE clause.
///
/// PINS: which uploads the preflight is allowed to count. Each excluded row below is excluded by a
/// DIFFERENT predicate, so a change that drops or weakens any one of them fails here and has to be
/// argued for. (It does not detect production drifting away from these copies — see the file
/// header; `export_visibility_where!()` is what makes that impossible.)
#[sqlx::test]
async fn the_estimate_sums_exactly_the_rows_the_archive_will_contain(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let visible = seed_user(&pool, event_id, "Anna").await;
let banned = seed_user(&pool, event_id, "Ben").await;
let hidden = seed_user(&pool, event_id, "Cara").await;
seed_upload(&pool, event_id, visible, 1_000, false).await;
seed_upload(&pool, event_id, visible, 2_500, false).await;
// Each of these is excluded from the archive by a DIFFERENT predicate.
seed_upload(&pool, event_id, visible, 9_000, true).await; // soft-deleted
seed_upload(&pool, event_id, banned, 9_000, false).await; // uploader banned
seed_upload(&pool, event_id, hidden, 9_000, false).await; // uploads hidden
set_user_moderation(&pool, banned, true, true).await;
set_user_moderation(&pool, hidden, false, true).await;
let rows = export_visible_uploads(&pool, event_id).await;
let expected: i64 = rows.iter().map(|(_, bytes)| bytes).sum();
assert_eq!(rows.len(), 2, "only Anna's two live uploads are archived");
assert_eq!(
estimate_export_bytes(&pool, event_id).await,
expected,
"the preflight must size the gallery the export will actually write"
);
assert_eq!(expected, 3_500);
}
/// An event with nothing to archive estimates zero rather than NULL.
///
/// PREVENTS: `SUM()` over no rows returning NULL and the decode blowing up — which would abort the
/// export with a type error instead of building an (entirely legitimate) empty keepsake.
#[sqlx::test]
async fn an_empty_gallery_estimates_zero_not_null(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
assert_eq!(estimate_export_bytes(&pool, event_id).await, 0);
// And with a user who has uploaded nothing.
seed_user(&pool, event_id, "Anna").await;
assert_eq!(estimate_export_bytes(&pool, event_id).await, 0);
}
/// A release arms both halves, so the preflight sees a count of 2 and reserves for the pair.
///
/// PREVENTS: the concurrency under-reservation. `spawn_export_jobs` starts the ZIP and HTML workers
/// at the same instant, and BOTH are gallery-sized (`Memories.zip` streams the original for every
/// video and every image at or under 5 MB, all `Compression::Stored`). A worker reserving only for
/// itself would see "it fits", its sibling would independently see the same, and together they
/// would ENOSPC — which is why `required_free_bytes` multiplies by this count.
#[sqlx::test]
async fn a_release_arms_both_halves_so_the_preflight_reserves_for_two(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user = seed_user(&pool, event_id, "Anna").await;
seed_upload(&pool, event_id, user, 1_000, false).await;
assert_eq!(
armed_job_count(&pool, event_id).await,
0,
"nothing is armed before the release"
);
let epoch = release_gallery(&pool, "wedding").await.expect("released");
assert_eq!(
armed_job_count(&pool, event_id).await,
2,
"a release arms zip AND html — both compete for the same disk"
);
// A worker that has claimed its half is still competing; `running` must keep counting.
assert!(claim_job(&pool, event_id, "zip", epoch).await);
assert_eq!(
armed_job_count(&pool, event_id).await,
2,
"claiming moves pending -> running, which must not drop out of the reservation"
);
// Only a FINISHED half stops competing.
assert!(finalize_job(&pool, event_id, "zip", epoch, "exports/Gallery.zip").await);
assert_eq!(
armed_job_count(&pool, event_id).await,
1,
"a done half no longer needs space reserved for it"
);
}
/// A ViewerOnly regeneration re-arms only the HTML half, so the preflight reserves for one.
///
/// PREVENTS: over-reservation refusing a rebuild that fits perfectly well. Moderating a comment
/// carries the finished ZIP forward untouched; demanding room for a second copy of it would fail
/// the one operation that needs no new gallery-sized write at all.
#[sqlx::test]
async fn a_viewer_only_regeneration_reserves_for_one_half(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user = seed_user(&pool, event_id, "Anna").await;
seed_upload(&pool, event_id, user, 1_000, false).await;
let epoch = release_gallery(&pool, "wedding").await.expect("released");
for t in ["zip", "html"] {
assert!(claim_job(&pool, event_id, t, epoch).await);
assert!(finalize_job(&pool, event_id, t, epoch, &format!("exports/{t}")).await);
}
assert_eq!(armed_job_count(&pool, event_id).await, 0);
// A moderated comment: bump the epoch, carry the ZIP forward, re-arm only the viewer.
let (_, _, next) = bump_epoch(&pool, "wedding").await.expect("bumped");
assert!(
carry_zip_forward(&pool, event_id, next).await,
"the finished ZIP is re-stamped, not rebuilt"
);
let mut conn = pool.acquire().await.expect("acquire");
enqueue_types_at_epoch(&mut conn, event_id, next, &["html"]).await;
assert_eq!(
armed_job_count(&pool, event_id).await,
1,
"only the viewer is being rebuilt, so only one archive's worth of space is needed"
);
}

View File

@@ -0,0 +1,352 @@
//! DB-backed tests for the deleted-media sweep (`services/maintenance.rs`).
//!
//! Context, in two halves.
//!
//! The compression worker deliberately no longer deletes an upload's original when its transcode
//! fails — a transient ENOSPC or a codec panic must never destroy the only copy of a photo a guest
//! cannot retake. But the row is soft-deleted and the uploader's quota IS refunded, so those bytes
//! become invisible, unowned and free.
//!
//! The SAME hole was reachable by the ordinary path, and that one is not an edge case at all:
//! `soft_delete_in_event` refunds `total_upload_bytes` on every guest or host delete and nothing
//! removed the files, so the quota stopped bounding the disk. Upload 500 MB, delete, quota back to
//! zero, upload another 500 MB — a guest curating their camera roll, which is what people do. The
//! sweep used to reach only `compression_status = 'failed'`, so it never touched this case; the
//! test below that now asserts an owner-deleted upload IS reclaimed is the one that used to assert
//! the opposite.
//!
//! Two windows, because the two deletes mean different things: 14 days for a failure an operator
//! may want to investigate, 24 hours for a removal someone asked for (14 days outlives the whole
//! event, so a deliberate delete would never reclaim anything while it mattered).
//!
//! The selection predicate is the whole safety argument — it must reach both leftovers and never a
//! live upload — so that is what these pin, following the same "reproduce the SQL verbatim" pattern
//! as `upload_concurrency.rs`. `#[sqlx::test]` gives each test a fresh, migrated database.
mod common;
use common::*;
use sqlx::PgPool;
use uuid::Uuid;
const FAILED_DAYS: i64 = 14;
const DELETED_HOURS: i64 = 24;
/// SRC: `services/maintenance.rs::cleanup_deleted_media` — the selection, verbatim.
async fn sweep_selects(pool: &PgPool, failed_days: i64, deleted_hours: i64) -> Vec<Uuid> {
type Row = (Uuid, String, Option<String>, Option<String>, Option<String>);
sqlx::query_as::<_, Row>(
"SELECT id, original_path, preview_path, display_path, thumbnail_path FROM upload
WHERE deleted_at IS NOT NULL
AND CASE WHEN compression_status = 'failed'
THEN deleted_at < NOW() - ($1 || ' days')::interval
ELSE deleted_at < NOW() - ($2 || ' hours')::interval
END
AND (original_path <> '' OR preview_path IS NOT NULL
OR display_path IS NOT NULL OR thumbnail_path IS NOT NULL)",
)
.bind(failed_days.to_string())
.bind(deleted_hours.to_string())
.fetch_all(pool)
.await
.expect("sweep query")
.into_iter()
.map(|(id, ..)| id)
.collect()
}
/// Seed an upload aged `deleted_hours_ago` (None = live), with optional derivative paths.
async fn seed_aged_upload(
pool: &PgPool,
event_id: Uuid,
user_id: Uuid,
status: &str,
deleted_hours_ago: Option<i64>,
original_path: &str,
derivatives: bool,
) -> Uuid {
sqlx::query_scalar(
"INSERT INTO upload (event_id, user_id, original_path, mime_type, original_size_bytes,
compression_status, deleted_at,
preview_path, display_path, thumbnail_path)
VALUES ($1, $2, $3, 'image/jpeg', 1000, $4,
CASE WHEN $5::bigint IS NULL THEN NULL
ELSE NOW() - ($5::text || ' hours')::interval END,
CASE WHEN $6 THEN 'previews/p.jpg' END,
CASE WHEN $6 THEN 'displays/d.jpg' END,
CASE WHEN $6 THEN 'thumbs/t.jpg' END)
RETURNING id",
)
.bind(event_id)
.bind(user_id)
.bind(original_path)
.bind(status)
.bind(deleted_hours_ago)
.bind(derivatives)
.fetch_one(pool)
.await
.expect("seed upload")
}
/// A live upload is untouchable no matter how the windows are configured.
///
/// PREVENTS: the catastrophic loosening. Everything else here is about reclaiming more; this is the
/// one assertion that must never bend.
#[sqlx::test]
async fn a_live_upload_is_never_selected(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-live").await;
let user_id = seed_user(&pool, event_id, "Sweeper").await;
for status in ["done", "failed", "processing", "pending"] {
let live = seed_aged_upload(
&pool,
event_id,
user_id,
status,
None,
"originals/e/live.jpg",
true,
)
.await;
assert!(
!sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
.await
.contains(&live),
"a non-deleted upload with status {status} must never be swept"
);
}
}
/// THE FIX. An upload a guest or host deliberately deleted is reclaimed once past 24 hours.
///
/// PREVENTS: the regression back to a sweep scoped to `compression_status = 'failed'`, which is
/// what let the quota stop bounding the disk. This assertion is the inverse of the one this file
/// used to make.
#[sqlx::test]
async fn a_deliberately_deleted_upload_is_reclaimed_after_a_day(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-deleted").await;
let user_id = seed_user(&pool, event_id, "Curator").await;
let deleted = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/owner.jpg",
true,
)
.await;
// Still inside the window — a mis-tap is recoverable for a day.
let recent = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(2),
"originals/e/recent.jpg",
true,
)
.await;
let selected = sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await;
assert!(
selected.contains(&deleted),
"a deliberate delete past the window must be reclaimed — this is the leak"
);
assert!(
!selected.contains(&recent),
"a delete inside the window keeps its recovery grace"
);
}
/// The two windows are independent: a failure is retained far longer than a deliberate delete.
///
/// PREVENTS: collapsing them into one. Applying 24h to failures would destroy the recovery window
/// the retained-original fix exists to provide; applying 14 days to deliberate deletes would mean
/// nothing is ever reclaimed during an event.
#[sqlx::test]
async fn the_two_retention_windows_do_not_bleed_into_each_other(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-windows").await;
let user_id = seed_user(&pool, event_id, "Windows").await;
// 48h old: past the deliberate window, nowhere near the failure window.
let failed_recent = seed_aged_upload(
&pool,
event_id,
user_id,
"failed",
Some(48),
"originals/e/f-recent.jpg",
false,
)
.await;
let deleted_same_age = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/d-same.jpg",
false,
)
.await;
// 30 days old: past both.
let failed_old = seed_aged_upload(
&pool,
event_id,
user_id,
"failed",
Some(30 * 24),
"originals/e/f-old.jpg",
false,
)
.await;
let selected = sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await;
assert!(
!selected.contains(&failed_recent),
"a 2-day-old compression failure is still inside its 14-day recovery window"
);
assert!(
selected.contains(&deleted_same_age),
"a deliberate delete of the same age is past its 24-hour window"
);
assert!(
selected.contains(&failed_old),
"a 30-day-old failure is past both windows"
);
}
/// Boundary behaviour on both windows.
#[sqlx::test]
async fn retention_windows_are_honoured_at_the_boundary(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-boundary").await;
let user_id = seed_user(&pool, event_id, "Boundary").await;
let cases = [
("failed", 13 * 24, false, "13 days"),
("failed", 15 * 24, true, "15 days"),
("done", 23, false, "23 hours"),
("done", 25, true, "25 hours"),
];
for (status, hours, expected, label) in cases {
let id = seed_aged_upload(
&pool,
event_id,
user_id,
status,
Some(hours),
"originals/e/b.jpg",
false,
)
.await;
assert_eq!(
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
.await
.contains(&id),
expected,
"a {status} upload deleted {label} ago: expected swept={expected}"
);
sqlx::query("DELETE FROM upload WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("clean up");
}
}
/// A row is re-selected until EVERY one of its paths is cleared.
///
/// PREVENTS: two failures at once. The sweep used to clear `original_path` alone, which was right
/// for its only case (a failed compression produces no derivatives) but leaves preview, display and
/// thumbnail on disk the moment it reaches a successfully processed upload — three files per
/// upload, none of them counted in `original_size_bytes`, that nothing else ever removes. And a row
/// whose paths are all cleared must stop coming back, or every hourly tick logs a phantom reclaim
/// forever.
#[sqlx::test]
async fn a_row_is_reselected_until_every_path_is_cleared(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-idempotent").await;
let user_id = seed_user(&pool, event_id, "Idem").await;
let id = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/once.jpg",
true,
)
.await;
assert_eq!(sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await, [id]);
// Clearing only the original is NOT enough — the derivatives are still on disk.
sqlx::query("UPDATE upload SET original_path = '' WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("clear original");
assert_eq!(
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await,
[id],
"derivatives left behind must keep the row selected"
);
sqlx::query(
"UPDATE upload SET preview_path = NULL, display_path = NULL, thumbnail_path = NULL
WHERE id = $1",
)
.bind(id)
.execute(&pool)
.await
.expect("clear derivatives");
assert!(
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
.await
.is_empty(),
"a fully swept row must not come back"
);
}
/// The derivative backfill must never resurrect what the sweep just reclaimed.
///
/// PREVENTS: an interaction, not a bug in either piece. The sweep nulls `preview_path`, and
/// `backfill_stale_derivatives` selects on `display_path IS NULL AND preview_path IS NOT NULL` —
/// close enough that a future edit to either could have the backfill re-decode an original that is
/// no longer on disk, on every boot. `deleted_at IS NULL` is what keeps them apart.
#[sqlx::test]
async fn the_backfill_ignores_swept_rows(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-backfill").await;
let user_id = seed_user(&pool, event_id, "Backfill").await;
seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/gone.jpg",
true,
)
.await;
// SRC: `services/compression.rs::backfill_stale_derivatives` — the selection, verbatim.
let backfilled: Vec<(Uuid, String, String)> = sqlx::query_as(
"SELECT id, original_path, mime_type FROM upload
WHERE deleted_at IS NULL AND mime_type LIKE 'image/%'
AND original_path IS NOT NULL
AND (
(display_path IS NULL AND preview_path IS NOT NULL)
OR derivatives_rev < $1
)",
)
.bind(1i16)
.fetch_all(&pool)
.await
.expect("backfill query");
assert!(
backfilled.is_empty(),
"a soft-deleted row must be invisible to the backfill, before or after sweeping"
);
}

Some files were not shown because too many files have changed in this diff Show More