Compare commits

...

60 Commits

Author SHA1 Message Date
fabi
4a9cf878ea Merge branch 'fix/reject-zero-quota-tolerance' 2026-07-29 21:37:39 +02:00
fabi
1e5953a566 Merge branch 'fix/keepsake-viewer-integrity' 2026-07-29 21:37:39 +02:00
fabi
50d1b5b06d fix(admin): reject quota_tolerance = 0 instead of silently blocking every upload
Zero is inside the documented 0–1 range and catastrophic. The per-user limit is
`free_disk * tolerance / active_uploaders`, so a tolerance of 0 makes every limit 0 and
refuses EVERY upload -- mid-event, with "Du hast dein Upload-Limit für dieses Event
erreicht", an error naming the wrong cause entirely. An admin reaching for an off-switch
wants `storage_quota_enabled`; the rejection now says so.

Rejecting the value rather than raising the floor. A floor of 0.01 was the obvious fix
and it is wrong: very small tolerances are legitimate -- they are how a large disk is
throttled down to a sensible per-guest ceiling, and how the quota specs steer it
(tolerance = target * active / free lands around 1e-5 on the 174 GB volume this suite
runs on). A floor would forbid real configurations, and would have broken the entire
storage-quota describe block, to prevent one typo. Verified: those four tests still pass.

Tests: the rejection, that the stored value is untouched (validation fully precedes any
write), and the mirror -- 0.00001 still round-trips -- so the guard can't quietly become
a floor later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 21:37:39 +02:00
fabi
bac30404e3 fix(export): stop a caption bricking the viewer, and ship readable archives
Two defects in the keepsake, both silent server-side and both only visible by
extracting the real artifact and trying to use it.

1. A CAPTION COULD BRICK THE VIEWER.

The viewer's data is inlined as `<script>window.__EXPORT_DATA__={…}</script>` -- it has
to be, since guests open index.html over file:// where fetching a sibling data.json is
blocked. The escape was `</` -> `<\/`. Against XSS that holds; I fired
`</script><img src=x onerror=…>` through a real Chromium parser and it round-trips
inert.

It does not stop the caption steering the HTML TOKENIZER. `<!--<script` with no later
`-->` drives the parser into script-data-double-escaped state, where the template's own
`</script>` only steps back to script-data-escaped instead of closing the element.
Everything after it -- including the viewer bundle -- is swallowed as script data.
Nothing executes and nothing leaks: `__EXPORT_DATA__` is simply never assigned and the
keepsake opens blank. A denial of the deliverable, not an XSS.

Reproduced in Chromium before changing anything, and the near-miss is worth recording:
`<!--<script>alert(1)</script>-->` comes back CLEAN, because the trailing `-->` returns
the parser to script-data state. A probe using the terminated form quietly repairs the
thing it is testing for.

Fix: escape every `<` as `<`, not just `</`. `<` never appears in JSON structural
syntax -- only inside string values -- so a global replace is sound, and one rule covers
`</script`, `<!--` and `<script` together. That is the point: the old escape was named
for the single case it handled. Only the INLINED copy is escaped; data.json is written
separately, in no HTML context, and stays literal.

2. EVERY ENTRY IN BOTH ARCHIVES WAS STORED MODE 0000.

`ZipEntryBuilder::new` leaves the external file attribute at zero and async_zip's host
compatibility defaults to Unix, so `unzip -Z` showed `?---------` on every line of both
Gallery.zip and Memories.zip. Windows Explorer ignores Unix modes, which is why this
survived; on Linux and macOS `unzip` faithfully applies what the archive asks for and
the guest gets a folder of photos none of which they can open.

Unconditional -- every keepsake ever produced, no hostile input required -- and
invisible server-side: the export succeeds, the ZIP is well-formed, the job writes
`done`, /export/status is green.

Found by accident. The browser test for defect 1 failed with ERR_ACCESS_DENIED on
file://, which looked exactly like a Playwright sandbox quirk; I twice "worked around"
it (fresh context, then a separately launched browser) before checking the extracted
files and finding mode 000. The workaround was suppressing a real bug. Both workarounds
are gone -- the ordinary `page` fixture loads the archive fine now.

Fix: all six ZipEntryBuilder sites route through one `keepsake_entry` helper stamping
`S_IFREG | 0644`, so the mode cannot be forgotten at a call site.

Tests: 3 unit (no `<` survives; the payload still decodes to the original value, because
this is a transport encoding and not a sanitiser; a clean payload is untouched) and 2
e2e that release for real, download the real archives, and check them from outside the
app -- one opening index.html over file:// in Chromium and asserting the viewer booted,
the captions came back verbatim and nothing executed; one asserting every stored mode
and every extracted file is readable. Both assertions verified to FAIL against the
pre-fix artifacts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 21:37:27 +02:00
fabi
5d0c7cd949 Merge branch 'refactor/share-export-visibility-filter' 2026-07-29 20:50:04 +02:00
fabi
a20b96d893 refactor(export): share the visibility filter between the row query and the estimate
`query_uploads` selects the rows the archives are built from; `estimate_export_bytes`
sizes them for the disk preflight. They stated the same WHERE clause separately, and
the direction of drift matters: an estimate that MISSES rows the archive writes
under-reserves, which is precisely the ENOSPC the preflight exists to prevent.

The integration test claimed to guard this and cannot. Both sides of
`the_estimate_sums_exactly_the_rows_the_archive_will_contain` are `SRC:`-marked
hand-copies in tests/common/mod.rs -- neither is production code -- so drift means
production moved while both copies sat still, and the test goes on passing. The
convention is sound for pinning behaviour; it is structurally incapable of detecting
divergence from the thing it copies.

So fix it where it can be fixed. One `export_visibility_where!()` fragment,
`concat!`-ed into both queries at compile time (still `&'static str`, no allocation),
with the `u`/`usr` alias contract stated. Divergence is now impossible by
construction rather than watched for.

The tests keep their value and lose the overclaim: the docstrings now say they pin
WHICH uploads may be counted -- each excluded row in the fixture is excluded by a
different predicate, so weakening any one of them still fails here -- and say plainly
that they do not detect drift, with a pointer to what does.

No behaviour change. The filters were verified identical before the hoist
(`u.event_id = $1 AND u.deleted_at IS NULL AND usr.uploads_hidden = FALSE AND
usr.is_banned = FALSE`); 99 backend tests still pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 20:50:04 +02:00
fabi
24ac862f81 Merge branch 'fix/video-poster-race' 2026-07-29 20:10:54 +02:00
fabi
a3d8ae72e3 fix(e2e): stop the video poster assertion racing the ffmpeg thumbnail
Pre-existing, and it fired for real during the full-suite run on a cold stack.

The lightbox binds `poster={upload.thumbnail_url ?? undefined}`, so the attribute is
absent until compression produces the thumbnail. This test asserted on it immediately
after seeding, never waiting for the worker -- unlike the Range test further down the
same file, which does poll. Against a warm stack the worker usually wins; against a
freshly rebuilt one (`stack:down -v`, cold ffmpeg) it doesn't.

That is the worst possible time for a false failure: the first run after a rebuild is
exactly when you are trying to establish whether a change broke something. Poll for
`compression_status = 'done'` before the poster assertion. The `src` assertion needs
no wait and keeps none.

Verified with --repeat-each=3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 20:10:54 +02:00
fabi
e1653cc54e Merge branch 'chore/db-memory-and-social-rate-limits' 2026-07-29 19:57:34 +02:00
fabi
35390800c7 chore: raise the db memory limit and rate-limit social writes
Two smaller operational items.

POSTGRES 512M -> 1G. DATABASE_MAX_CONNECTIONS is 30 for a ~100-guest event (feed
polling + SSE + uploads at once), and 30 backends plus Postgres 16's default
shared_buffers leaves very little headroom at 512M. An OOM here doesn't degrade one
feature -- every request path touches the database, so it takes the event down.
Memory is the cheaper knob than shrinking the pool back and reintroducing the
queueing it was raised to fix. .env.example now names the pairing explicitly, the way
it already does for COMPRESSION_WORKER_CONCURRENCY.

SOCIAL WRITES WERE UNTHROTTLED. toggle_like, add_comment and delete_comment were the
only mutating endpoints in the app with no limit at all -- upload, join, recover,
export and admin login all carry one. Asymmetric coverage rather than a deliberate
decision.

Low severity, and honestly so: a like fans an SSE broadcast to every client, but the
export regeneration a comment deletion triggers is contained (REGEN_DEBOUNCE 20s,
workers born with their epoch, superseded ones inert). So the ceiling is 120/min --
far above anything a real guest produces. This bounds a script, not an enthusiastic
double-tapper.

ONE bucket across all three actions: separate buckets would let a caller triple the
aggregate write rate by alternating between them. Keyed per USER, matching the feed
and upload limits -- at a venue every guest is behind one NAT, and an IP key is what
made the /join and /feed limits turn guests away in the first place.

Migration 020 seeds both keys, and both are wired into the admin allowlist, the
config UI and the e2e reseed -- the step two earlier per-area toggles missed, which
left switches that existed in code and could never be flipped.

Tests: 4 e2e, including that the shared bucket really is shared (the part most likely
to be lost in a refactor) and that one guest hitting the ceiling doesn't block
another behind the same IP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:57:34 +02:00
fabi
14ebe1e543 Merge branch 'docs/backup-restore-and-quota-tolerance' 2026-07-29 19:51:57 +02:00
fabi
a4a4e46c53 docs: add a restore procedure, fix the backup cadence, and correct quota_tolerance
Four things, all found by the same question: what does an operator standing at the
venue actually need?

A RESTORE PROCEDURE. There was none anywhere, and a backup you have never restored
isn't a backup. Two hazards worth writing down: media must be extracted preserving
ownership (the app runs as uid 100 / gid 101, and a root-owned restore makes every
upload fail with EACCES surfacing as a generic 500), and the app must be STOPPED
first, because migrations run on boot and a live pool will fight the restore.

Both the backup and the restore commands were run against the real stack before being
written down, which caught two that would have failed:

  - The plain `pg_dump` did not restore: `psql` aborted on `ERROR: schema
    "_sqlx_test" already exists`. pg_dump emits no DROPs without --clean --if-exists,
    so the documented dump could only ever be restored into an empty database. Fixed
    at the source (the dump is now self-cleaning) and verified end to end: 16 tables
    back, exit 0.
  - `--same-owner` does not exist in BusyBox tar, which is what `alpine` ships, so
    the extract aborted before unpacking anything. `--numeric-owner` plus the
    explicit chown, verified to land 100:101.

BACKUP CADENCE. "Weekly offsite" is the wrong shape when every irreplaceable byte is
created in one eight-hour window and nobody can retake a wedding. The backup that
matters runs that night, and again after the release so the keepsake is captured.
Also: take the DB dump and the media tarball back to back, or you get rows pointing
at files the dump doesn't know about.

quota_tolerance WAS DOCUMENTED AS SOMETHING IT ISN'T. .env.example called it "fraction
of disk that triggers the low-storage warning". It is the multiplier in
`floor(free_disk * tolerance / active_uploaders)` -- so an operator who wants "warn me
later" and sets 0.95 is actually authorising guests to fill 95% of the disk, moving
the fixed point from 43% to ~49% and eating the export headroom. The admin UI labelled
it "Toleranz (0-1)" with no explanation at all, which invites exactly that reading;
it is now "Speicher-Anteil für Gäste" with the formula in the hint. Wrong docs on a
tuning knob are worse than no docs.

SIZING. New section with the arithmetic: three volumes on one filesystem, the quota
fixed point at tolerance/(1+tolerance), and the fact the 80 GB baseline does not cover
the keepsake -- both archives are built concurrently and each is roughly a second copy
of every original. Provision ~3x expected media, or give exports its own volume.

Also ticks the low-disk alert off the roadmap, since it now exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:51:48 +02:00
fabi
e6e8a52d87 Merge branch 'feat/low-disk-warning' 2026-07-29 19:48:06 +02:00
fabi
43c2a0d09c feat(host): warn about low disk before it becomes unrecoverable
Storage visibility existed in exactly one place: a passive Speicherauslastung widget
on the ADMIN dashboard. A host who isn't the admin had no view of it, and nothing
warned anyone. README carried "Low-disk alert (< 10 GB free)" under Planned since v1.

Two things make this a safety net rather than a nice-to-have. postgres_data,
media_data and exports_data are all Docker named volumes on ONE filesystem, so
running out doesn't degrade a subsystem -- Postgres stops being able to write and the
whole event goes down. And the keepsake needs room for two gallery-sized archives,
which the export preflight can only ever refuse AFTER the release, when the event is
over and every remedy is harder.

So the threshold is not a fixed number alone. It fires on the 10 GB floor the README
always named, OR on "you could not build the keepsake right now" -- the trigger a
host can still act on, computed with the same arithmetic the preflight uses. Unknown
free space is NOT low: it fails open like the upload quota and the preflight do,
because a banner that cries wolf on an unreadable mount is a banner nobody reads.

Carried on GET /host/event, which the dashboard already fetches on load and on every
reload -- no new endpoint, no new poll. Rendered above everything else including the
PIN-reset queue, and it names the consequence (the event, not just the download)
rather than only the number.

Also fixes the host page's formatBytes, which topped out at MB: 30 GB free would have
rendered as "30720.0 MB", and a guest with 2 GB of uploads was already being shown
that way in the user list.

Tests: 5 unit on the threshold (including that plenty of free space is still low when
the keepsake wouldn't fit -- the case a fixed threshold misses entirely), 3 e2e.
The e2e drives it through `original_size_bytes` rather than a genuinely full disk:
the estimate is pure SQL over that column, so overstating one row moves the
accounting without touching a byte on disk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:48:06 +02:00
fabi
6818cabf91 Merge branch 'fix/reclaim-deleted-originals' 2026-07-29 19:42:03 +02:00
fabi
f777764839 fix(maintenance): reclaim the media of deliberately deleted uploads
The quota stopped bounding the disk. `soft_delete_in_event` stamps `deleted_at` and
refunds `total_upload_bytes`, but nothing ever removed the bytes, and the hourly
sweep reached only `compression_status = 'failed'`. Upload 500 MB, delete, quota back
to zero, upload another 500 MB. Not an attack -- a guest curating their camera roll,
which is what people do. The host then sees guests hitting "Du hast dein Upload-Limit
erreicht" while the admin widget shows a disk full of files no upload row points at,
and the quota message is actively misleading because the space really is gone, just
not to anyone the accounting can name.

Two retention windows, because the two deletes mean different things. A compression
failure keeps its 14 days: the guest didn't ask for it and may not be able to retake
the photo. A deliberate removal gets 24 hours -- 14 days outlives the whole event, so
a deliberate delete would never reclaim anything while it mattered, and a day still
covers a mis-tap.

Wider than reported: ALL FOUR paths are reclaimed, not just the original. Preview,
display and thumbnail are each a separate file, none counted in
`original_size_bytes`, and nothing ever removed them either. That was invisible while
the sweep only saw failed compressions (which produce no derivatives) and becomes
three leaked files per upload the moment it reaches a successful one. A row is
re-selected until every path is cleared, and the columns are cleared only once every
file for that upload is gone -- clearing after a partial success would strand the
survivors in exactly the unowned state this drains.

`backfill_stale_derivatives` selects on `display_path IS NULL AND preview_path IS NOT
NULL`, which is close enough to the post-sweep state to be worth pinning: it is
guarded on `deleted_at IS NULL`, so it cannot re-decode an original that is no longer
on disk. Covered.

Residual, deliberately: within the 24h window the bytes are still spent and still
unaccounted, so delete-and-re-upload through an eight-hour event can outrun the
sweep. Bounding that means holding the quota until the file is reclaimed rather than
refunding at `deleted_at`. The low-disk warning is the net under it.

Tests: 6 DB-backed, replacing 3. The one asserting an owner-deleted upload IS
reclaimed is the exact inverse of what this file used to assert.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:41:51 +02:00
fabi
aeb958f6ba Merge branch 'fix/export-disk-preflight' 2026-07-29 19:38:39 +02:00
fabi
281eb3bec7 fix(export): refuse an export that cannot fit, and stop peaking at two generations
Nothing in export.rs ever asked whether the keepsake would fit. Both archives write
their media `Compression::Stored`, so each is essentially a byte-for-byte second copy
of the originals -- Gallery.zip always, and Memories.zip for every video and every
image at or under 5 MB. On the documented CX33 (80 GB, all three volumes on one
filesystem) the upload quota's fixed point leaves ~40 GB free, and a release spawns
BOTH halves concurrently against it.

The failure is not "the export failed", it is "the deliverable is stuck":

  1. ENOSPC lands partway through a multi-GB write.
  2. The epoch has already moved, so the job row is `failed` at the CURRENT
     generation and readiness (epoch = event.export_epoch AND status = 'done') is
     false -- GET /export/zip 404s.
  3. The last good archive sits on disk, unreferenced and unreachable.
  4. POST /host/export/rebuild, the only escape, re-arms the same doomed write.

Three changes.

Reclaim before building. `prune_stale_export_files` ran only after the new archive
was written, renamed and finalised. That reads as durability but buys nothing: the
moment `invalidate_and_arm` bumps the epoch the old archive is ALREADY unreachable,
so keeping it reserves gigabytes for a download nobody can perform -- and for a
takedown it is content someone explicitly asked to have removed. Peak usage is now
one generation. Narrower than the post-finalize prune on purpose: final archives
only, never a `.tmp` or a `viewer_tmp_` dir, since a superseded worker can still be
streaming into those and at build START is far more likely to be alive.

Preflight the space. SUM(original_size_bytes) over exactly `query_uploads`'
visibility filter, +10% for ZIP overhead, multiplied by the number of armed jobs --
without that multiplier each of the two concurrent halves independently sees "it
fits" and together they don't. Runs AFTER claim_job, not before as reported: bailing
before the claim leaves the row `pending` with no worker and no error, the
spinner-forever state `mark_failed`'s status guard exists to prevent. Fails open when
the mount can't be read, exactly as the upload quota does.

Show the host the reason. /export/status returned {status, progress_pct} and nothing
else, so the host dashboard could only render "fehlgeschlagen" next to the retry
button. The message was written to the row and surfaced solely in the ADMIN job list
-- a different screen, possibly a different person. It now travels with the status,
and only on a failure, so a message left on a since-succeeded row can't appear beside
a green "ist bereit".

Tests: 10 unit (the u128 clamp caught a real bug in the first draft -- saturating_mul
then /100 turns an overflow into a number ~100x too small, the one direction that
authorises the write being guarded against; the carried-forward archive must survive
its own older epoch in the filename), 4 DB-backed (the estimate is asserted against
the row set the archive actually contains, not against a restatement of the WHERE
clause, so the two queries cannot drift), 3 e2e over the four-hop plumbing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:38:26 +02:00
fabi
8c93cbb045 Merge branch 'fix/narrow-admission-check' 2026-07-29 07:55:56 +02:00
fabi
ceb68939a7 fix(upload): narrow the admission check to the memory budget only
The admission check I just added rejected ANY image the decoder couldn't build —
corrupt, truncated, or unsupported, not only over-budget. That broke two
adversarial tests, and they were right to break.

07-adversarial/file-upload-attacks pins, deliberately, that acceptance follows the
MAGIC BYTES: a payload whose first three bytes are a JPEG header is accepted
regardless of what follows, because the security property under test is that the
client-declared Content-Type has no influence. Both failing cases upload 1024
bytes of JPEG magic followed by zeros. Rejecting those at admission is a
different, broader contract than the one asked for, and rewriting an adversarial
test to match new behaviour is precisely the thing that needs justifying rather
than doing quietly.

So admission now checks only what it was meant to: `exceeds_decode_budget`
returns true solely for `ImageError::Limits`. A corrupt file goes to the
compression worker exactly as before — which handles it gracefully and, since the
retry classifier in the previous commit, no longer burns backoff on it. The
resource guard is the part that had to move earlier; nothing else did.

Tests: the size agreement between admission and the worker is still asserted in
both directions, plus a new one writing a magic-bytes-only stub and asserting
admission accepts it WHILE the worker still rejects it — pinning the boundary
between the two checks so a future widening fails here rather than in the
adversarial suite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 07:55:56 +02:00
fabi
674ea87bbd Merge branch 'fix/no-retry-on-permanent-failure' 2026-07-29 07:29:21 +02:00
fabi
fae12bd7ec fix(upload): refuse undecodable images at the door, and stop retrying them
Two halves of the same complaint: an oversized photo was accepted with a 201 and
then silently soft-deleted minutes later, after the worker had burned six seconds
of backoff re-reaching a conclusion it could not change.

Admission. The compression budget now runs at upload time, against the header
only, so a guest is told immediately and told why:

  "Bild hat zu viele Bildpunkte (ca. 99 Megapixel) und kann nicht verarbeitet
   werden. Bitte verkleinere es und lade es erneut hoch."

instead of watching the photo vanish behind a vague "could not be processed" —
which arrived only if they happened to still be on the feed with that card
loaded. Nothing is stored, so there is no row to soft-delete and no orphan for
the sweep to reclaim.

Admission and the worker share ONE function (`decoder_within_budget`), so they
cannot drift apart and start disagreeing about what is acceptable — a photo
accepted at the door and rejected by the worker would be worse than either
behaviour alone. The worker keeps its own check: the backfill decodes files that
predate this check, and defence in depth is the whole reason the budget exists.

Retries. The loop retried every failure, including ones that are a property of
the input. An image over the budget, a corrupt file, an unsupported format: each
fails identically on all three attempts, so the only effect was 2s + 4s of sleep
and three near-identical warnings before the same outcome. `is_permanent_image_error`
classifies the `ImageError` variants that cannot change between attempts — Limits,
Unsupported, Decoding — and the loop gives up on those at once. `IoError` is
deliberately excluded: an ENOSPC while writing a derivative is exactly the
transient case the retry exists for, and misclassifying it would turn a blip back
into the data loss round 1 fixed. Measured: retry log lines went from 3 per
oversized upload to 0.

Tests: unit tests for both sides of the classifier (a Limits error is permanent, a
missing file is not) and for admission agreeing with the decoder on accept AND
reject. The e2e spec is rewritten for the new contract — 400 with an actionable
message, nothing stored, backend alive after a burst of four — plus a mirror
asserting an ordinary photo still uploads and processes, since a budget that
rejected everything would satisfy the other two.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 07:29:21 +02:00
fabi
2bd54d7f0b Merge branch 'chore/unshare-permission-settings' 2026-07-29 07:18:08 +02:00
fabi
528960d201 chore: take the Bash(*) permission change back out of the shared settings
`.claude/settings.json` is committed and applies to anyone who clones. Fabi's
local `allow: ["Bash(*)"]` plus deny list ended up in it, inside f0d69f1 — a
commit about the image decode guard, which has nothing to do with permissions.

That was my mistake, twice. The file was already modified when I started the
round: my `git status --short` check printed "(clean)" from an unconditional
`echo` rather than from the status output, so I read a dirty tree as clean. Then
`git add -A` swept it into an unrelated commit, and I reported afterwards that I
had left it untouched. Neither the check nor the claim was true.

Restores the shared file to its previous three narrow entries. The permission
setup itself is preserved, moved to `.claude/settings.local.json`, which
`.gitignore:34` covers precisely so per-user permissions stay per-user — the
existing 442 entries there are kept alongside it.

Not rewriting f0d69f1 to erase this: main is unpushed so it would be safe, but a
visible correction is worth more than a tidy history, and a rebase across the
merge commits carries more risk than the mistake does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 07:18:08 +02:00
fabi
6b7da8fb07 Merge branch 'fix/deploy-health-and-caddy-reload' 2026-07-28 22:34:01 +02:00
fabi
faf2e62a29 fix(deploy): route /health in production, and actually apply Caddyfile changes
Two defects in the update procedure I wrote last round, both of which make a
successful-looking deploy a lie.

1. The documented health check could never pass.

`curl -fsS https://DOMAIN/health` 404s against a perfectly healthy production
stack. The backend registers /health on its ROOT router, not under /api/v1, and
the production Caddyfile proxies only /api/* and /media/* — so /health fell
through to the SvelteKit catch-all, which has no such route and returns its 404
page. With -f, curl exits 22 and the `&& echo` never runs. My own gloss
("Anything other than ok means check the logs") then sent the operator chasing a
phantom outage.

e2e/Caddyfile.test has carried `reverse_proxy /health app:3000` since it was
written — precisely because the catch-all would otherwise swallow it. Production
never did. Per the fix-the-gap-not-the-doc call, production gets the same line,
and /health joins the no-store matcher so a cached response can't report the last
known state instead of the current one. Verified by running the production
Caddyfile against the real backend: /health -> 200 "ok", Cache-Control: no-store,
with /api/v1/event and / unaffected.

2. The sequence never reloaded Caddy, so a Caddyfile-only change was dropped.

`--build` only rebuilds services with a `build:` section, and caddy is a pinned
upstream image. Compose decides whether to recreate a container from its config
hash, which covers the mount SPECIFICATION but not the mounted file's CONTENTS —
so a git pull that changes ./Caddyfile produces no delta, Compose reports
`Running`, and Caddy serves its old config indefinitely. Exit code 0 throughout.

Round 1's iOS download fix (137c4ee) is exactly this shape: Caddyfile plus four
e2e files, so 100% of its production effect is in that one file. Following the
README to the letter deployed it, showed both image IDs changing, and left iOS
downloads broken.

Demonstrated rather than assumed — added a probe header to a Caddyfile, ran the
old sequence (`up -d --build`): header absent, change silently dropped. Ran the
new step 4 (`up -d --force-recreate caddy`): header served.

`--force-recreate` rather than `restart` or `caddy reload` because the bind mount
is resolved to an inode at container-create time and git pull replaces the file
rather than editing in place, so a restart can re-read the stale content — the
exact failure I hit in round 1 when `caddy reload` didn't pick up an edit.

Also rewrites the "db and caddy are untouched … so data volumes survive" sentence.
I wrote it as reassurance; "caddy is untouched" was the bug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:34:01 +02:00
fabi
6d5c488e14 Merge branch 'fix/decode-allocation-guard' 2026-07-28 22:31:09 +02:00
fabi
f0d69f1cda fix(imaging): restore the decode allocation guard I removed in round 1
This is a regression I introduced, not a pre-existing gap. Before 05948d8 the
compression worker used `ImageReader::decode()`, which does:

    let mut decoder = Self::make_decoder(format, self.inner, limits.clone())?;
    limits.reserve(decoder.total_bytes())?;   // enforces max_alloc
    decoder.set_limits(limits)?;

Reading the EXIF orientation tag needs `into_decoder()` instead, and that skips
the reserve entirely — the crate's own FIXME concedes `from_decoder` doesn't
compensate. Nothing else enforces `max_alloc`: the JPEG decoder's `set_limits`
only checks support and dimensions. So the 256 MiB budget has been inert since
that commit, and round 2 then propagated the weakened path into export.rs through
the shared helper, in a commit whose message claimed the helper "carries" the
decompression-bomb cap. It didn't, and the comment saying max_alloc "hard-caps
the decode allocation" was simply false.

What was left was only the per-axis cap, which permits 12000x12000 — 412 MiB
decoded, 824 MiB for the two concurrent decodes the worker runs by default,
against a 1 GiB container. Deploy-blocking right now because bumping
DERIVATIVES_REV makes the first boot after a deploy re-decode the entire gallery
two at a time: an OOM kill there restarts the container, which re-runs the
backfill. A boot loop, on the first deploy of these fixes.

Re-add the reserve exactly as `decode()` does it. Per the budget decision it stays
at 256 MiB (~89 MP for RGB8, above any mainstream phone's real output); two
concurrent decodes now peak at 512 MiB. Oversized images take the graceful path
from round 1 — original retained, quota refunded, upload-error toast — and fail
after the header parse but BEFORE any pixels are read, so they cost a header read
rather than an allocation. Measured peak during a concurrent oversized burst: 3.0
MiB.

Test parity is the other half, and the reason this was invisible: the e2e app
container had NO memory limit while production is capped at 1 GiB, so a decode
that would OOM-kill production simply succeeded in CI. Mirror the 1 GiB cap in
docker-compose.test.yml. That is the third divergence of this shape, after WebKit
missing from CI and /health existing only in Caddyfile.test.

Tests: a fixture that is 568 KiB on disk and 283 MiB decoded (11000x9000 = 99 MP,
deliberately UNDER the per-axis cap so the axis check cannot be what rejects it).
A unit test asserts the refusal — it fails against the old code, which decoded it
into an 11000x9000 buffer — with a companion asserting an ordinary photo still
decodes AND still gets its orientation applied, so the guard didn't become a
blanket refusal. An e2e test uploads it singly and as a concurrent pair, asserting
compression lands in 'failed' and the backend is still serving and still
processing afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:31:09 +02:00
fabi
64eccb8672 Merge branch 'chore/prettier' 2026-07-28 21:28:11 +02:00
fabi
eefa476765 chore: satisfy prettier in frontend and e2e
`checks.yml` runs `npm run format:check` for both projects and both were failing.

- frontend/src/lib/ui-store.ts is mine, unformatted since the round-1 upload-queue
  badge fix — the same miss as the rustfmt one: I gated on svelte-check and eslint
  but never on format:check.
- e2e/loadtest/* and e2e/shots.mjs have been unformatted since 7758270 and are
  unrelated to the audit work. Fixed here because they block the same gate and the
  fix is mechanical; no behaviour change in either project.

Still red and deliberately NOT fixed here: `npm run lint` in the frontend reports
`svelte/prefer-svelte-reactivity` on routes/diashow/+page.svelte:208 (a mutable
`Set` where the rule wants `SvelteSet`), pre-existing since 5009590. That one is a
real reactivity change in code I have no test coverage for, so it belongs in its
own change rather than smuggled into a formatting commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 21:28:11 +02:00
fabi
0932e2a470 Merge branch 'chore/rustfmt' 2026-07-28 21:26:28 +02:00
fabi
117c0c547f chore(backend): satisfy cargo fmt
`checks.yml` runs `cargo fmt --check`, and it has been failing since the round-1
audit fixes: I gated those on `cargo build` and `cargo clippy` but never ran fmt,
so three files drifted then and eight more this round. Pure formatting — no
behaviour change; clippy stays at zero and all 70 backend tests still pass.

Worth noting for next time: clippy passing is not evidence fmt does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 21:26:28 +02:00
fabi
c49bf875d9 Merge branch 'ci/webkit' 2026-07-28 20:57:08 +02:00
fabi
6e7c4565cd ci(e2e): run WebKit, so the iOS guarantees actually gate a PR
The workflow installed only Chromium and ran chromium-desktop + chromium-mobile.
iOS Safari is the app's stated primary user — a wedding guest opening a QR link —
and WebKit is the only engine in the matrix that reproduces two of its behaviours:

  - it enforces X-Frame-Options on the hidden download iframe, so a site-wide DENY
    makes the keepsake download silently do nothing. Blink hands attachments to
    the download manager before the frame check and never notices.
  - it abandons a <video> load unless its Range probe gets a 206.

Both of those shipped. Adding 06-export to the webkit project in the round-1 fix
bought nothing on a PR, because CI never ran that project at all — the regression
test written specifically to catch the blocker only ever executed locally.

Runs 71 tests (67 pass, 4 skip on the documented IndexedDB-blob harness
limitation) in ~1.5 minutes locally, using the exact command added here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:57:08 +02:00
fabi
1a7a531c90 Merge branch 'fix/recover-ip-ceiling' 2026-07-28 20:54:58 +02:00
fabi
6920e5bf7a fix(recover): cap name cycling, and stop bcrypt blocking the runtime
Round 1 gave /join a per-IP ceiling and left /recover with only its
`recover:{ip}:{name}` bucket. That key is right for the job it was written for —
stopping someone who knows a display name (they're listed on the feed) from
burning the victim's 3-strike PIN counter and locking them out on repeat. But the
name is ATTACKER-CHOSEN, so cycling names mints a fresh 5-attempt bucket every
time and the per-IP cost is unbounded.

What sits behind that limiter makes it worse than a normal flood: every call runs
a cost-12 bcrypt verify, including an UNCONDITIONAL throwaway verify for names
that don't exist — added deliberately to close a timing oracle. So an unknown name
is the single cheapest way to make the server do ~200ms of hashing.

Adds `recover_ip_rate_per_min` (default 30, migration 019), checked BEFORE the
per-name bucket so a name generator can't walk past it. 30/min is far above any
real recovery attempt while capping a flood. The per-name bucket is untouched and
remains the anti-guessing control.

The second half matters as much as the first: bcrypt was running inline on the
async runtime everywhere. At cost 12 that pins a tokio worker thread for ~200ms,
and there is only one per core — so a login flood stalled every other request on
the box, including the feed. There was no spawn_blocking anywhere in the auth
module, despite SECURITY-BACKLOG claiming bcrypt had been offloaded.

Route all of it through `verify_password` / `hash_password` on the blocking pool.
That covers /recover, /admin/login, the host PIN reset, and — the one most likely
to bite at a real event — the PIN hash minted on every single /join. Saturating
the blocking pool degrades logins; saturating the worker threads degrades
everything.

Tests: cycling distinct names from one IP now hits the ceiling with a Retry-After,
and — the assertion that keeps the fix honest — repeated wrong PINs against ONE
name are still throttled with the ceiling set generously high, so the ceiling
added protection rather than replacing it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:54:58 +02:00
fabi
58f718bdce Merge branch 'fix/export-exif-orientation' 2026-07-28 20:46:38 +02:00
fabi
3c984e2932 fix(export): apply EXIF orientation in the keepsake too
Round 1 fixed EXIF orientation in the compression worker, which corrected the live
app — feed preview and diashow display. The export worker was missed, and it does
not reuse those derivatives: it re-decodes the originals itself with `image::open`,
which ignores the orientation tag, then re-encodes to JPEG, which drops the tag —
so the viewer has no way to recover it.

The damage was oddly shaped, which is exactly why it reads as a viewer bug:

  Gallery.zip originals              correct  (byte-copied, EXIF intact)
  Memories viewer grid thumbnails    SIDEWAYS (always)
  Memories viewer full image >5 MB   SIDEWAYS (re-encoded at 2000px)
  Memories viewer full image ≤5 MB   correct  (streamed byte-for-byte)

So in the keepsake people actually keep, every portrait photo in the grid was on
its side, and clicking through silently "fixed" small photos but not large ones.

Rather than paste the decoder dance a third time, extract `services::imaging::
decode_oriented` and route both workers through it, so there is exactly one way to
turn a file on disk into a DynamicImage. It carries a second invariant that had
also drifted: `image::open` applies NO decode limits, so the export path was
decoding arbitrary user-supplied images unbounded — the decompression-bomb cap
existed only in the compression worker. Both now come as a pair, which is the
point of having one function.

Not done: switching export to consume the existing `display` derivative. It would
fix orientation and drop a redundant full-resolution decode per photo, but it
would also replace the pristine ≤5 MB originals in the keepsake with 2048px
re-encodes — a real quality regression in the one artefact people keep forever.

Test uploads the round-1 fixture (40x20 landscape tagged Orientation=6), runs a
real export, pulls the thumbnail out of Memories.zip and asserts it came back
portrait — with a sanity check that the source really is stored landscape, so the
test can't pass against a pipeline that does nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:46:38 +02:00
fabi
d6fdc13da9 Merge branch 'fix/compression-orphan-quota' 2026-07-28 20:40:15 +02:00
fabi
c14ccd2df1 fix(compression): reclaim failed originals instead of leaking them
Round 1 stopped the compression worker deleting an upload's original on failure —
a transient ENOSPC or a codec panic must never destroy the only copy of a photo a
guest cannot retake. But it left `Upload::soft_delete`'s quota refund in place, so
the bytes stayed on disk while the uploader was charged nothing for them.

That is worse than it first looks. The row is soft-deleted, so the file is
invisible and unowned; a guest hitting a reproducible codec failure can accumulate
orphans indefinitely at zero personal cost. And `active_uploaders` counts only
users with non-deleted uploads, so dropping out of that count RAISES everyone's
per-user ceiling — the leak loosens the very quota meant to contain it.

Keep the refund: the uploader didn't cause the failure and shouldn't silently lose
quota to it. Bound the leak instead, with an hourly sweep alongside the existing
session cleanup in `spawn_periodic_tasks`, reclaiming failed originals older than
14 days — comfortably longer than any single event, so an operator investigating a
failed upload still has the file.

The selection predicate is the entire safety argument, so it is deliberately
narrow: `compression_status = 'failed'` AND soft-deleted AND past the window AND
`original_path <> ''`. That is exactly the state the give-up path leaves behind,
and it cannot reach a live upload, an owner-deleted one, or a failure still inside
its recovery window. `original_path` is cleared after a successful reclaim, which
makes the sweep idempotent — otherwise a row whose file is already gone is
re-selected on every tick forever. The row itself is kept as the audit trail.

Tests reproduce the selection verbatim (same pattern as upload_concurrency) and
assert it against five near-misses that must survive, both sides of the retention
boundary, and the idempotence property.

Also fixes two comments in export.rs still claiming "the compression worker
hard-deletes an original when its transcode fails" — no longer true, and the
defensive handling they justify is now justified by this sweep and by ordinary
deletes instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:40:15 +02:00
fabi
9c8cc7c069 Merge branch 'fix/role-store-identity-reset' 2026-07-28 20:35:40 +02:00
fabi
1485df5469 fix(auth): bind the role store to the identity, not to the tab
The role store I added in the moderation work is a module-level singleton seeded
ONCE at import. `goto()` is a client-side navigation, so leaving and re-joining in
the same tab re-imports no module and re-runs no onMount — the previous user's
role simply stayed resident. Nothing reset it: not join, recover, admin login,
"Event verlassen", `clearAuth`, nor the api.ts 401 auto-clear.

So a host who left, followed by a guest joining on the same phone, left that guest
with `isStaff === true` and a "🚫 Beitrag entfernen" action on other people's
photos. The backend 403s the delete, so this was a false affordance rather than a
privilege escalation — but `/feed` never fetched `/me/context`, so unlike every
other route it never self-corrected either. It survived until a hard reload.

The mirror case was equally broken and easier to overlook: a guest who recovered
into a host account got NO host affordances.

`clearAuth` already had a hook registry for exactly this shape of problem, with a
comment explaining it exists to avoid circular imports. Add the missing mirror,
`onSetAuth`, fired by both `setAuth` and `setAdminAuth` after the new token is
resident, and have the role store register on both sides: clear to null on
logout, re-seed from the new token on login. That also gives
`syncRoleFromToken` — dead code with zero callers since I introduced it — its
intended purpose.

Seeding from the claim fixes the reported bug, but the claim is frozen for the
token's 30-day life, so a promotion or demotion still wouldn't reach the feed.
`/feed` now calls the existing `refreshEventState()` on mount, which fetches
`/me/context` and applies both the authoritative role and the lock/release state
in one request. The feed is the one route gating a destructive action on the role,
so it should not be the only route running on a stale claim.

Tests: 04-host/role-identity-reset drives the real flows. The first asserts the
host DOES see the action before asserting the newcomer does not — a negative
assertion alone would pass against a build that shipped no moderation at all. The
second covers the mirror, promoting a guest server-side while their resident token
still claims `role: guest`, so a fix that only cleared the role would fail it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:35:40 +02:00
fabi
81e5017f27 Merge branch 'fix/video-playback' 2026-07-28 20:29:16 +02:00
fabi
813a9fa500 fix(video): play the actual video, and answer Range requests
Every video in the app was unplayable. Two independent defects, either one
sufficient on its own, and nothing in the suite covered either — no test
anywhere played media or asserted a `<video>` src.

1. The lightbox handed `<video>` a JPEG.

`pickMediaUrl` is mime-agnostic, and compression only ever produces a THUMBNAIL
for a video (one `ffmpeg -vframes 1` frame) — no preview, no display. So in the
DEFAULT saver mode the element's src resolved to `/api/v1/upload/{id}/thumbnail`,
served as `image/jpeg` with `nosniff` so the browser can't even sniff its way
out. Chromium reports DEMUXER_ERROR_COULD_NOT_OPEN.

Fixed in the lightbox rather than in `pickMediaUrl`: FeedListCard shares that
helper and legitimately wants the thumbnail for its `<img>` poster, so a central
mime branch would break the feed. This mirrors the rule the diashow already
applies ("videos play the original file directly"). Added `preload="none"` so
saver-mode guests on cellular still fetch nothing until they press play — there
is no smaller video derivative to offer them — plus `playsinline`, without which
iOS hijacks playback into fullscreen.

2. `stream_media_file` ignored Range entirely.

It took no request headers, so it could not see `Range`; it always returned 200
with the whole body and never sent Accept-Ranges or Content-Range. iOS Safari
opens every `<video>` with a `Range: bytes=0-1` probe and abandons the load
without a 206 — so video failed on the app's primary platform even in `original`
mode, where the src was already correct.

Adds single-range support (`bytes=N-`, `bytes=N-M`, `bytes=-S`) with 206 +
Content-Range, 416 + `bytes */len` past EOF, and Accept-Ranges advertised on
every response. Anything it won't handle — multi-range, non-bytes units, garbage
— falls back to a full 200, which RFC 9110 explicitly permits and which is safer
than guessing. All four media routes share the helper, so seeking works
uniformly.

`get_original` now serves `inline` instead of `attachment`. An attachment
disposition is hostile to a `<video>` element, and this route is the only source
of playable video bytes; it also matches what the UI promises, since the action
is labelled "Original anzeigen" — view, not download. `no-store` is deliberately
kept so a takedown still revokes access promptly; ranges work fine under it, the
client just re-fetches.

Tests: 11 unit tests pin the parser (the iOS `bytes=0-1` probe, inclusive ends,
suffix ranges, clamping past EOF, 416 vs 200, malformed fallbacks). A new
03-feed/video-playback spec asserts the src is the original and not the
thumbnail, that the browser accepts the bytes as media (readyState > 0, no
MediaError), that no video bytes are delivered before play, and that Range
returns the correct 206 slices and a 416 past EOF — verified on both Chromium
and WebKit.

The "not downloaded before play" test asserts no *delivered body* rather than no
request: WebKit opens a connection for a preload="none" video and immediately
aborts it (GET, no Range, status 0, nothing transferred) while Chromium issues
nothing at all. The portable guarantee is that no response carrying bytes
completes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:29:16 +02:00
fabi
f03e392f8c Merge branch 'docs/upgrade-path' 2026-07-28 20:17:32 +02:00
fabi
3d94bbd6fb docs(deploy): document the update path — up -d alone ships nothing
The README only ever described a fresh install. There was no update section
anywhere, and `--build` appeared nowhere in the docs.

That matters because `app` and `frontend` are `build:` services with no published
image tag, and Compose has no source-change detection: if an image by that name
exists it is reused. So the natural `git pull && docker compose up -d` reports
"Container app-1 Running", rebuilds nothing, and exits 0. A deploy that shipped
none of the new code is indistinguishable from a successful one — which is how
eleven merged fixes can sit in the repo and never reach the box.

Verified both halves against a real stack rather than asserting them: with a
source change staged, `up -d` left the image ID untouched; `up -d --build`
produced a new image ID and a healthy /health.

Adds an "Updating an existing deployment" section covering backup-before-migrate,
pull, rebuild, health check, and an image-ID comparison to prove a build actually
happened. Also spells out the rollback trap: migrations run on boot and are not
undone by checking out an older commit, so rolling back code without restoring
the snapshot leaves the schema ahead of the binary and the app refusing to start.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 20:17:32 +02:00
fabi
96a22cfe27 Merge branch 'test/webkit-idb-blob-limitation' 2026-07-28 19:14:37 +02:00
fabi
c6e9350f78 test(e2e): document why WebKit can't run the client-queue upload tests
Three 02-upload tests have been failing on `webkit-iphone` on main — `02-upload`
was already in that project's testMatch, so this is pre-existing red, not
something the audit work introduced.

Root cause is the harness, not the app. Playwright's Linux WebKit build cannot
store Blobs in IndexedDB at all: `put()` fails with "UnknownError: Error
preparing Blob/File data to be stored in object store". Confirmed it is not
about how Playwright delivers files — a Blob constructed in-page with
`new Blob([bytes])` fails identically, while Chromium stores both that and a
`setInputFiles` File without complaint.

That breaks every test driving the composer (FAB → UploadSheet → /upload →
submit), because `handleSubmit` awaits `addToQueue`, which persists the file
before it can navigate. The symptom is a submit button stuck on "Wird
hochgeladen…" and a timeout waiting for /feed — which reads like an app hang and
cost real time to run down.

Skip those four (the three above plus the new rejection-visible) on WebKit only,
behind a named helper carrying the full explanation, so the next person gets the
answer instead of the investigation. Deliberately narrow: WebKit still runs every
API-driven upload test, all of 01-auth, 03-feed and 06-export — including the
keepsake download, which only WebKit can meaningfully verify. Chromium continues
to run all 14.

Worth being explicit, since these tests exist to protect iOS: real Safari
supports Blobs in IndexedDB, so this is NOT evidence that the offline upload
queue is broken on the platform. It does mean that guarantee is currently
unverifiable in CI and rests on Chromium coverage plus manual device testing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 19:14:37 +02:00
fabi
537a11b0a4 Merge branch 'docs/backup-and-test-defaults' 2026-07-28 07:49:52 +02:00
fabi
27e4004cc8 docs(backup): make the backup commands work; fix the e2e/prod divergences
Backup. Both documented commands failed on the shipped stack, and the sentence
explaining them was wrong too:

- `pg_dump $DATABASE_URL` — `DATABASE_URL` is only ever in the compose
  environment, never an operator's shell, and it points at `db:5432`, which is
  compose-internal DNS. The app image has no postgres client either.
- `> /media/backups/…` — `/media` is a named volume mounted inside the app
  container, not a host path, and nothing ever creates a `backups` subdirectory.
- `rsync /opt/eventsnap/media/` — that path does not exist anywhere.
- "a single path to back up" — false, and dangerously so: exports were moved to
  their own `exports_data` volume precisely so a keepsake (which contains every
  photo in the event) can't be served off the media tree. Backing up only
  `media_data` silently loses every generated keepsake.

Rewritten as three commands — db via `docker compose exec -T db pg_dump`, and one
`docker run … tar` per volume — all verified against the running stack. The
volume mounts use `/src`, not `/media`: I hit the footgun while testing this.
Docker pre-populates an EMPTY volume from the image's own directory and chowns it
to match, so `-v media_data:/media alpine` tars alpine's cdrom/floppy/usb, writes
them into the volume, and leaves it root-owned so the non-root app can no longer
write. Mounting where the image has nothing avoids all of it. Documented inline
so the next person doesn't rediscover it.

Also correct the architecture notes: `/media/*` no longer routes to the backend
(that static tree was removed as a gating bypass), and `exports_data` was missing
from the volume list — the one volume an operator most needs to know about.

e2e stack: add the `EXPORT_PATH` + `/exports` volume it was missing. The file
says "mirrors production layout"; without these, exports landed on the container's
writable layer at the default path, so export-leak and export-video wrote real
archives into ephemeral storage and the "exports live outside media" invariant
was never actually exercised.

Pre-existing red test, unrelated to the audit: all four 02-upload/quota tests
have been failing since 4464147 "stop /me/quota leaking raw disk to guests"
(2026-07-19), which post-dates the spec's last edit. `setLimitTo` calibrated
`quota_tolerance` from `free_disk_bytes` read through the GUEST's token — a field
that commit deliberately zeroes for non-staff. Dividing by it yields a NaN
tolerance, so every test in the block died in the helper. Read the calibration
inputs through a staff token and keep reading the ceiling back through the guest,
whose limit is the thing under test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 07:49:52 +02:00
fabi
c4e9b89af0 Merge branch 'fix/upload-pipeline-integrity' 2026-07-28 07:19:25 +02:00
fabi
05948d8268 fix(upload): stop destroying originals, apply EXIF orientation, surface rejections
Three defects in the same pipeline, each of which loses a photo or misrepresents
one.

1. A transient error destroyed the guest's only copy.

`process`'s error arm unconditionally `remove_file`d the original. Every failure
routed there: `create_dir_all`, both derivative `save_with_format` calls (disk
full is the canonical case, and it arrives exactly when many guests upload at
once), a panic inside the image codec, or a momentary DB-pool exhaustion. The
row is only SOFT-deleted, so the bytes were the sole unrecoverable part — and
they were the part we deleted. The author already knew this was wrong next door:
`backfill_missing_display` says it "must NEVER soft-delete an upload that already
has a working preview".

Retry up to 3 times with backoff (re-checking the e2e generation guard after each
sleep), and on final failure keep the refund + soft-delete but leave the original
on disk, logging its path. A failed upload is now recoverable instead of gone.

2. Every portrait photo was stored sideways.

Phones don't rotate sensor data — they record the camera orientation in EXIF and
store the pixels as shot. `decode()` returns those raw pixels and the JPEG
re-encode writes no EXIF, so the 800px preview, the 2048px diashow display and
the keepsake were all rotated 90°, while "Original anzeigen" rendered upright
because the original keeps its tag. That asymmetry is why it reads as a viewer
bug. There was no EXIF handling anywhere in the repo and no exif crate.

Read the tag via `into_decoder()` (which carries the decode Limits through, so
the decompression-bomb cap is untouched) and apply it. Missing/malformed tags
fall back to NoTransforms — most images have none.

Existing derivatives are already baked wrong, so migration 018 adds
`derivatives_rev` and `backfill_missing_display` becomes
`backfill_stale_derivatives`: it now also picks up anything below the current rev
and regenerates it once from the original, which still carries its EXIF. Videos
are marked current in the migration — ffmpeg already honours the rotation matrix.
Bump DERIVATIVES_REV for any future change that invalidates derivatives.

3. A rejected upload vanished without a word.

`UploadQueue.svelte` — 162 lines holding the ONLY renderer of an item's error
text, the only "Erneut" retry button and the only rate-limit countdown — was
never imported anywhere, so `retryItem`, `removeItem` and `clearCompleted` were
unreachable at runtime. On a terminal rejection the store purged the blob and
wrote a clear German reason into `entry.error` "so the UI shows a clear reason".
There was no such UI. And `uploadBadgeCount` counted only pending/uploading, so
the badge decremented exactly as if the upload had succeeded.

Mount the queue on /upload, toast the reason immediately (the flow sends the user
to /feed straight after staging, so the list alone would still miss them), and
count blocked/error in the badge so a failure can't read as success.

Tests: 02-upload/exif-orientation uploads a 40x20 fixture tagged Orientation=6
and asserts both derivatives come back PORTRAIT, with a sanity check that the
source really is stored landscape. 02-upload/rejection-visible bans the uploader
between staging and sending, then asserts the toast, the queue row with the
server's reason, and that the item is still counted.

Note: 02-upload/quota's 4 failures are pre-existing and unrelated — see the next
commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 07:19:25 +02:00
fabi
d4237ad2ad Merge branch 'feat/host-moderation-ui' 2026-07-27 22:33:33 +02:00
fabi
be6d56f278 feat(moderation): let a host remove a guest's photo or comment from the UI
`DELETE /host/upload/{id}` and `DELETE /host/comment/{id}` were complete on the
backend — transactional, SSE-broadcasting, audit-logged — and had zero frontend
callers. The feed context sheet offered "Löschen" only when
`target.user_id === myUserId`, so the only lever a host actually had against an
unwanted photo was banning the uploader.

That is both disproportionate and ineffective. A ban doesn't retract what was
already posted, and it makes things strictly worse for comments: the ban check
runs BEFORE the ownership check on the guest delete route, so banning an abusive
author leaves their comment on screen and permanently undeletable by them. With
no host affordance, nobody could remove it at all.

- feed: hosts/admins get "Beitrag entfernen" on other people's posts, routed to
  the host endpoint (the guest route 403s anything the caller doesn't own) with
  moderation-specific confirm copy. Own-post "Löschen" is unchanged.
- lightbox: same for comments, via /host/comment/{id}.
- Ban semantics are deliberately untouched (USER_JOURNEYS §10 — banned users keep
  read access and cannot write). The deadlock is broken by giving the host a way
  in, not by loosening the ban.

Live role (this had to come first). `getRole()` decodes the JWT claim, but the
token is never reissued — the backend slides the session row forward and treats
the DB row as authoritative. The claim is therefore frozen for the token's
lifetime: up to 30 days. A guest promoted at the party saw no Host-Dashboard and
no moderation actions until they signed out and back in, even though
`/me/context` had been returning their real role on every page load and 4 of its
6 call sites dropped the field on the floor.

Add `role-store.ts`: seeded from the claim so there's no flash of the wrong nav,
then corrected by every `/me/context` response. Point the ad-hoc `getRole()`
callers at it (account, upload, host, admin, and the new feed gate). The host and
admin dashboards now derive `myRole` reactively, so a demotion disables their
controls immediately instead of at next login.

Tests: 04-host/moderation-ui drives the real UI — host removes a guest photo and
it's gone from /feed server-side; a plain guest is offered nothing on someone
else's post (the mirror that keeps the first test honest); a promoted guest gains
the dashboard on reload while their token still carries `role: guest`; and a host
removes the comment of an already-banned guest, asserting first that the author's
own delete 403s so the deadlock is real.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 22:33:33 +02:00
fabi
0d8e83d392 Merge branch 'fix/rate-limit-shared-nat' 2026-07-27 21:56:49 +02:00
fabi
89057d605f fix(rate-limit): key the guest-facing limiters per user, not per IP
At a venue every guest is behind one NAT, so an IP-keyed limiter hands the
whole party a single bucket. On a fresh deploy 12 guests arriving together
meant 5 joined and 7 were turned away, with no Retry-After telling them when
to retry. `/feed` (60/min) and `/export` (3 per DAY — the fourth guest to
fetch their keepsake locked out until tomorrow) had the same defect.

`feed_delta` was already keyed per-user and its comment states the exact
rationale ("so one client can't starve others behind a shared NAT"); this
makes its siblings match.

- feed:{ip}   -> feed:{user_id}    (auth was already in scope)
- export:{ip} -> export:{user_id}  (resolved from the download ticket's
  session, which was previously looked up and discarded)
- join:{ip}: pre-auth, so there is no user to key on. Split in two — a loose
  per-IP ceiling that only bounds raw volume (new `join_ip_rate_per_min`,
  default 60, migration 017), plus the real 5/60s anti-spam bucket keyed
  per (ip, name), mirroring the existing `recover:{ip}:{name}`.

admin_login / recover / pin_reset_req stay IP-keyed on purpose and are now
commented as such: they guard credential guessing, where a per-user or
per-name key would just hand an attacker a fresh bucket per guess.

Retry-After: the machinery existed but 7 of 8 sites called `check()` and
hard-coded `None`, so a throttled client was told to back off but never for
how long. Delete the bool `check()` wrapper entirely so `check_with_retry`
is the only entry point and the delay cannot be discarded by accident. Also
surface it for the PIN lockout, where the deadline was already known.

Fix the "unknown" fallback while here: every client_ip() caller passed that
literal, so any request without X-Forwarded-For — anything reaching the app
directly rather than through Caddy — shared ONE global bucket. Serve with
connect-info and use the peer address.

Tests: the reseed forces every limiter toggle off before each test, which is
why this whole class was invisible. Add 01-auth/rate-limit-shared-nat, which
enables them and asserts 12 guests share an IP without collision, that one
guest hammering their own name IS still throttled (so the fix re-keys rather
than removes the limit), and that feed/export buckets are per-user. Retarget
the ddos join test at the new per-IP ceiling — it asserted the defect.

Also seed `admin_login_rate_enabled` (read by the handler, seeded by no
migration and no reseed) and register `join_ip_rate_per_min` in the admin
config allowlist. Unrelated pre-existing red test fixed: 01-auth/join
asserted a "Willkommen!" heading the wedding redesign removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:56:49 +02:00
fabi
688dc614d7 Merge branch 'fix/media-gating-percent-escape' 2026-07-27 21:21:18 +02:00
fabi
42416d76e2 fix(media): close the percent-escape bypass of the media gate
`/media/%70reviews/{id}.jpg` served a taken-down photo to anyone,
unauthenticated. Verified against the running stack: the literal path 404s,
the escaped one returned 200 with the full image. Same for displays,
thumbnails and originals, and any escaped byte in any position works.

Cause: the block was four `nest_service("/media/previews", 404)` route
matches sitting above a `ServeDir` on `/media`. axum matches on the RAW path
(matchit does no percent-decoding), while `ServeDir` percent-decodes when it
resolves the file. So `%70reviews` missed every blocker, fell through to the
ServeDir, and was decoded back to `previews/` on disk — reaching the bytes
with no soft-delete and no ban-hide check. That defeats a host takedown,
which is the entire point of the gate.

Remove the `/media` route tree outright instead of racing the decoder.
Nothing needs it: every media URL the backend emits is already a gated
`/api/v1/upload/{id}/{original,preview,display,thumbnail}` alias
(handlers::feed), the frontend contains zero `/media/` references, and the
`/media` in config.rs/disk.rs is the filesystem path while `media/` in
export.rs is a path inside the zip. `/media/**` now 404s regardless of
encoding. The route's own comment already said it "serves nothing" — it
wasn't a backstop, it was the vector.

Caddy keeps proxying /media/* deliberately: the app 404s it, and forwarding
means the e2e gating specs exercise the app's refusal exactly as production
would rather than being masked by the SvelteKit 404 page.

Extend the gating spec with the encoded variants — asserting only the literal
spelling is what let this sit undetected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:21:18 +02:00
fabi
cec69e804a Merge branch 'fix/ios-keepsake-download-webkit' 2026-07-27 21:10:41 +02:00
fabi
137c4ee8a1 fix(export): let the keepsake download through X-Frame-Options on iOS
The keepsake download navigates a hidden, same-origin iframe (deliberately:
a top-level navigation to a 404/429 would unload the PWA). Caddy stamped a
site-wide `X-Frame-Options: DENY` that also covered the proxied `/api/*`.

Blink hands a `Content-Disposition: attachment` response to the download
manager at the network layer, so Chromium never noticed. WebKit enforces XFO
on the frame navigation first and aborts the load — so on iOS Safari, the
app's primary platform, tapping Download did nothing at all, silently.

Carve the two export endpoints out to SAMEORIGIN, which still blocks
cross-origin framing. Implemented as two disjoint matchers rather than an
override: Caddy applies the FIRST `header` directive outermost, so it wins on
write and a later, more specific `header` is silently ignored (verified
against the running test stack).

Also close the test gap that let this ship:

- `06-export` ran on chromium-desktop only; add it to `webkit-iphone`, the
  only engine that enforces XFO on the download frame.
- No test in the suite ever clicked a download button — every archive
  assertion used Node `fetch`, which has no frame and no XFO enforcement.
  Add a spec that clicks it and awaits a real `download` event. Verified
  falsifiable: with the blanket DENY reinstated it fails and reports the
  WebKit refusal as the cause.
- Fix `ExportPage`'s card-scoped locators, which matched nothing: the cards
  carry `class="card p-5"` (a Tailwind `@apply` component class), never the
  `rounded-xl` the page object looked for. This had left the "shows enabled
  download buttons" test red on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:10:26 +02:00
77 changed files with 5085 additions and 383 deletions

View File

@@ -18,6 +18,10 @@ POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password
POSTGRES_DB=eventsnap
# Connection pool size. Default 10. For a busy event (~100 guests polling the feed
# + SSE + uploads at once) raise to ~30 so requests don't queue on a pool permit.
# PAIRED WITH THE DB CONTAINER'S MEMORY LIMIT: 30 backends plus Postgres 16's default
# shared_buffers is already snug in the 1G that docker-compose.yml allots the `db`
# service. If you raise this, raise `db.deploy.resources.limits.memory` with it — an
# OOM in Postgres doesn't degrade one feature, it takes the whole event down.
DATABASE_MAX_CONNECTIONS=30
# ── Authentication ────────────────────────────────────────────────────────────
@@ -54,8 +58,26 @@ EXPORT_PATH=/exports
# max image size 20 MB
# max video size 500 MB
# estimated guests 100
# quota tolerance 0.75 (fraction of disk that triggers the low-storage warning)
# quota tolerance 0.75 (see below — NOT a warning threshold)
# Adjust these in the admin UI before the event if needed.
#
# quota_tolerance is the MULTIPLIER IN THE PER-USER QUOTA FORMULA, not the point at
# which anything warns you:
#
# per_user_limit = floor(free_disk * quota_tolerance / active_uploaders)
#
# It is recomputed against LIVE free space on every upload, so it self-throttles: guests
# converge on a fixed point at tolerance/(1+tolerance) of the free space you started
# with — 43% at 0.75, i.e. ~30 GB of a fresh 70 GB.
#
# Raising it therefore AUTHORISES GUESTS TO FILL MORE OF THE DISK. Setting 0.95 in the
# belief that it means "warn me later" moves the fixed point to ~49% and eats the
# headroom the keepsake needs — and the keepsake needs a lot, because Gallery.zip and
# Memories.zip are each roughly a second copy of every original (both store media
# uncompressed). Budget for media + 2x media, or move exports to their own volume.
#
# 0.75 is the tested default. Lower it if the box is tight; raise it only if you have
# provisioned export headroom separately.
# ── Workers ───────────────────────────────────────────────────────────────────
# Number of parallel image/video compression workers. Default 2. This is the main

View File

@@ -7,7 +7,7 @@ on:
jobs:
e2e:
name: Playwright E2E (chromium-desktop)
name: Playwright E2E (chromium + webkit)
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
@@ -25,7 +25,7 @@ jobs:
- name: Install Playwright browsers
working-directory: ./e2e
run: npx playwright install --with-deps chromium
run: npx playwright install --with-deps chromium webkit
- name: Bring up the test stack
working-directory: ./e2e
@@ -54,6 +54,22 @@ jobs:
working-directory: ./e2e
run: npm run test:e2e -- --project=chromium-mobile
# iOS Safari is the app's stated primary user (a wedding guest opening a QR link), and
# WebKit is the ONLY engine here that reproduces two of its behaviours:
# - it enforces X-Frame-Options on the download iframe, so a site-wide `DENY` makes the
# keepsake download silently do nothing. Blink hands attachments to the download
# manager first and never notices. That shipped once already.
# - it abandons a <video> load without a 206 response to its Range probe.
# Both regressions are invisible to every Chromium project, so running WebKit is what
# actually gates them on a PR rather than on someone remembering to test locally.
#
# The project is scoped in playwright.config.ts to the journeys a guest walks
# (01-auth, 02-upload, 03-feed, 06-export); four IndexedDB-blob tests skip themselves
# there — see helpers/webkit.ts for why that is the harness and not the app.
- name: Run E2E tests (webkit / iOS)
working-directory: ./e2e
run: npm run test:e2e -- --project=webkit-iphone
- name: Upload Playwright report
if: failure()
uses: actions/upload-artifact@v4

View File

@@ -9,34 +9,63 @@
header {
Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Content-Type-Options "nosniff"
X-Frame-Options "DENY"
Referrer-Policy "strict-origin-when-cross-origin"
}
# X-Frame-Options: DENY everywhere EXCEPT the keepsake download endpoints, which
# are navigated in a HIDDEN, SAME-ORIGIN iframe so a 404/429 can't unload the PWA
# (see frontend/src/routes/export/+page.svelte). WebKit enforces XFO *before*
# honouring Content-Disposition, so a blanket DENY makes the download silently do
# nothing on iOS Safari — the app's primary platform. SAMEORIGIN still blocks
# cross-origin framing.
#
# Split into two disjoint matchers rather than an override: Caddy applies the
# FIRST header directive outermost, so it wins on write — a later, more specific
# `header` would be silently ignored.
@framable path /api/v1/export/zip /api/v1/export/html
@not_framable not path /api/v1/export/zip /api/v1/export/html
header @framable X-Frame-Options "SAMEORIGIN"
header @not_framable X-Frame-Options "DENY"
# SvelteKit frontend — static assets with long-lived cache (content-hashed filenames)
@hashed_assets path_regexp hashed /_app/immutable/.*\.[a-f0-9]{8,}\.(js|css|woff2)$
header @hashed_assets Cache-Control "public, max-age=31536000, immutable"
# Preview/thumbnail images. These are now served by the app through a
# visibility-checked alias (/api/v1/upload/{id}/{preview,thumbnail}) so moderation
# can revoke access; direct /media/previews|thumbnails is 404-blocked at the app.
# Preview/thumbnail images. These are served by the app through a visibility-checked
# alias (/api/v1/upload/{id}/{preview,thumbnail}) so moderation can revoke access;
# the app serves no /media route at all, so there is no direct path to the bytes.
# Privately cacheable for a short window (the app sets the same header; this is the
# edge carve-out from the blanket no-store below). Kept short so a moderated image
# stops being served to a direct-URL holder promptly.
@media_api path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail
header @media_api Cache-Control "private, max-age=300"
# API — never cache, EXCEPT the gated image routes above.
# API and health — never cache, EXCEPT the gated image routes above. A cached health
# response would report the last known state rather than the current one.
@api {
path /api/*
path /api/* /health
not path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail
}
header @api Cache-Control "no-store"
# Route API and media requests to the Rust backend
# Route API and media requests to the Rust backend.
#
# The app serves no /media route at all (see the note in backend/src/main.rs) — media
# bytes are reachable only through the visibility-checked /api/v1/upload aliases, so
# /media/* forwards to a plain 404. The proxy line is kept deliberately: it means the
# edge faithfully hands /media to the app, so if a future change ever re-introduces a
# static media route the e2e gating specs see it here exactly as production would,
# instead of being masked by the SvelteKit 404 page.
reverse_proxy /api/* app:3000
reverse_proxy /media/* app:3000
# The backend registers /health on its ROOT router, not under /api/v1, so it needs its
# own line — without it the catch-all below hands /health to SvelteKit, which has no
# such route and returns its 404 page. That made the documented post-deploy check
# (`curl -fsS https://DOMAIN/health`) fail 100% of the time on a perfectly healthy
# stack. e2e/Caddyfile.test has always carried this line; production never did.
reverse_proxy /health app:3000
# Everything else goes to SvelteKit frontend
reverse_proxy frontend:3001
}

View File

@@ -1133,16 +1133,19 @@ eventsnap/
### Backup Strategy
```bash
# Daily (e.g. as a separate Compose service or cron on the VPS)
pg_dump $DATABASE_URL | gzip > /media/backups/db_$(date +%Y-%m-%d).sql.gz
Three artefacts in three places: the database, the `media_data` volume
(originals + derivatives), and the **separate** `exports_data` volume. See
[README.md](README.md#backup) for the exact commands.
# Weekly: rsync /media volume to Hetzner Storage Box
rsync -az /opt/eventsnap/media/ \
user@u123456.your-storagebox.de:backup/eventsnap/
```
Everything runs through `docker compose` / `docker run`, because `DATABASE_URL`
and the `/media` and `/exports` paths only exist inside the compose network —
they are not host paths, and `DATABASE_URL` is never exported into an operator's
shell.
The `/media` volume contains originals, previews, thumbnails, generated exports, and DB backups — a single volume to back up.
Export archives are deliberately outside `MEDIA_PATH` (`EXPORT_PATH=/exports`): a
keepsake contains every photo in the event, and keeping it off the media tree is
what stops it being reachable except through the ticket-gated handler. A backup
of the media volume alone silently loses every generated keepsake.
---

231
README.md
View File

@@ -34,7 +34,6 @@ A guest scans the QR code on their way in, types their name, and is immediately
### Planned (v1.x)
- Individual file download button
- Low-disk alert (< 10 GB free)
- Event banner / cover image
- Chunked resumable upload for large videos
- Host-curated story highlights
@@ -115,6 +114,65 @@ Caddy automatically obtains a Let's Encrypt certificate on first start. The app
> docker compose -f docker-compose.yml -f docker-compose.dev.yml up
> ```
### Updating an existing deployment
> **`docker compose up -d` alone will NOT deploy your changes.** `app` and `frontend` are
> `build:` services with no published image tag, and Compose has no source-change detection:
> if an image with that name already exists it is reused. After a `git pull` the command
> reports `Container … Running`, changes nothing, and **exits 0** — so a deploy that shipped
> nothing looks exactly like a successful one. `--build` is what makes it real.
```bash
cd /path/to/eventsnap
# 1. Back up first — migrations run automatically on boot and are not reversible in place.
# (See "Backup" below; the database dump is the one that matters here.)
# 2. Fetch the new code.
git pull
# 3. Rebuild and restart the application services. --build is NOT optional.
docker compose up -d --build
# 4. Apply any Caddyfile change. Step 3 does NOT do this — see the warning below.
docker compose up -d --force-recreate caddy
# 5. Confirm the app came back up. Anything other than "ok" means check the logs.
curl -fsS https://DOMAIN/health && echo
# 6. Confirm a NEW image was actually built. Note the IMAGE ID before you start and
# compare — it must have changed. (Ignore the CREATED column; it reports the base
# layer's age, not this build's.) An unchanged ID means step 3 ran without --build
# and you are still serving the old code.
docker compose images app frontend
```
Migrations are applied by the backend on startup, so step 3 covers them. If `app` stays
unhealthy afterwards, `docker compose logs app` will name the failing migration — and note
that a migration applied by a *newer* build is not removed by checking out an older commit,
so rolling back code without restoring the database snapshot from step 1 leaves the schema
ahead of the binary and the app refusing to boot.
> **Why step 4 exists.** `--build` only rebuilds services that have a `build:` section, and
> `caddy` is a pinned upstream image. Compose decides whether to recreate a container from its
> *config hash*, which covers the mount **specification** (`./Caddyfile:/etc/caddy/Caddyfile:ro`)
> but **not the file's contents** — so a `git pull` that changes `./Caddyfile` produces no
> delta, Compose reports `Running`, and Caddy keeps serving its old config indefinitely. Exit
> code 0 throughout.
>
> That is not hypothetical: the fix that made the keepsake download work on iOS
> (`137c4ee`) touched the Caddyfile and four e2e files and nothing else, so **all** of its
> production effect lives in that one file. Without step 4 you deploy it, watch both image IDs
> change, and iOS downloads stay broken.
>
> `--force-recreate` rather than `restart` or `caddy reload`: the bind mount is resolved to an
> **inode** when the container is created, and `git pull` replaces the file instead of editing
> it in place, so the container can still be bound to the old, now-unlinked inode. A restart
> then re-reads the stale content. Recreating the container re-resolves the path.
`db` is never touched, and recreating `caddy` does not disturb the `caddy_data` volume, so the
TLS certificate and all data volumes survive.
### Generate required secrets
```bash
@@ -162,23 +220,176 @@ See [.env.example](.env.example) for the full list with descriptions and default
└────────┘
```
- `/api/*` and `/media/*` → Rust backend
- `/api/*` → Rust backend
- Everything else → SvelteKit frontend (`adapter-node`)
- Named volumes: `postgres_data`, `media_data`, `caddy_data`
- Named volumes: `postgres_data`, `media_data`, `exports_data`, `caddy_data`
Media is **not** served as static files. Every image goes through a
visibility-checked alias (`/api/v1/upload/{id}/{preview,display,thumbnail,original}`)
so a host takedown or a ban actually revokes access to the bytes.
---
## Sizing the disk
`postgres_data`, `media_data` and `exports_data` are all Docker named volumes under
`/var/lib/docker/volumes`, so **they share one filesystem**. Filling it does not
degrade one subsystem — Postgres stops being able to write and the whole event goes
down.
Uploads are self-limiting. `per_user_limit = free_disk × quota_tolerance ÷
active_uploaders` is recomputed against live free space on every upload, so guests
converge on a fixed point at `tolerance / (1 + tolerance)` of the free space you
started with — **43%** at the default 0.75. On an 80 GB box with ~70 GB free after
the OS and images, media settles at ~30 GB and stops.
**The keepsake is what the 80 GB baseline does not cover.** `Gallery.zip` and
`Memories.zip` are built concurrently and each is roughly a second copy of every
original: both write their media `Compression::Stored`, and `Memories.zip` streams the
untouched original for every video and for every image at or under 5 MB. So a release
wants room for **two more copies of the gallery** on top of the gallery itself.
| Stage | Used | Free (80 GB box) |
|---|---|---|
| Fresh box (OS + images) | ~10 GB | ~70 GB |
| Guests reach the quota fixed point | ~40 GB | ~40 GB |
| Host releases → both archives | ~100 GB | **ENOSPC** |
Two ways to size for it:
- **Provision ~3× your expected media** on one volume (media + two archives), or
- **give `exports_data` its own volume** so a full export cannot reach Postgres, and
size that one at ~2× expected media.
This is no longer silent. The export refuses up front with the two numbers rather than
hitting ENOSPC halfway through a multi-GB write, a rebuild reclaims the superseded
generation before it starts (so peak is one generation, not two), and the host
dashboard warns as soon as the keepsake would not fit — which is the only point at
which anyone can still do something about it.
---
## Backup
```bash
# Database snapshot
pg_dump $DATABASE_URL | gzip > /media/backups/db_$(date +%Y-%m-%d).sql.gz
There are **three** things to back up, and they live in three different places.
`DATABASE_URL` and the container paths (`/media`, `/exports`) are meaningful only
*inside* the compose network — they are not host paths, and `DATABASE_URL` is
never exported into an operator's shell — so every command below runs through
`docker compose` from the repo directory.
# Weekly offsite sync (Hetzner Storage Box or similar)
rsync -az /opt/eventsnap/media/ user@storagebox.example.com:backup/eventsnap/
```bash
# 1. Database snapshot. Runs pg_dump inside the db container (the app image has no
# postgres client), reading credentials from the compose environment.
# --clean --if-exists makes the dump SELF-CLEANING: without it the restore below
# aborts on the first "already exists" against a database that has ever booted,
# which is every database you would actually want to restore over.
mkdir -p ./backups
docker compose exec -T db \
sh -c 'pg_dump --clean --if-exists -U "$POSTGRES_USER" "$POSTGRES_DB"' \
| gzip > ./backups/db_$(date +%Y-%m-%d).sql.gz
# 2. Uploaded media (originals + derivatives) out of the named volume.
# NOTE the mountpoint is /src, not /media: if the volume is ever empty, Docker
# pre-populates a fresh mount from the image's own directory, and alpine ships a
# /media containing cdrom/floppy/usb. Mounting somewhere the image has nothing
# avoids silently tarring (and polluting the volume with) those.
docker run --rm \
-v eventsnap_media_data:/src:ro -v "$PWD/backups":/backup \
alpine tar czf /backup/media_$(date +%Y-%m-%d).tar.gz -C /src .
# 3. Export archives — a SEPARATE volume (see the security note below).
docker run --rm \
-v eventsnap_exports_data:/src:ro -v "$PWD/backups":/backup \
alpine tar czf /backup/exports_$(date +%Y-%m-%d).tar.gz -C /src .
# Offsite sync of the three artefacts above.
rsync -az ./backups/ user@storagebox.example.com:backup/eventsnap/
```
The `/media` volume holds originals, previews, thumbnails, exports, and DB backups — a single path to back up.
Volume names are prefixed with the compose project name — `eventsnap_` if you run
from a directory called `eventsnap`. Confirm yours with `docker volume ls`.
> **Exports are deliberately NOT under `/media`.** They live on their own
> `exports_data` volume (`EXPORT_PATH=/exports`) because a keepsake archive
> contains every photo in the event; keeping it outside the media tree is what
> stops it being reachable except through the ticket-gated download handler.
> Backing up only the media volume therefore loses every generated keepsake.
### When to run it
**A nightly cron is the wrong shape for this app.** Every irreplaceable byte is
created inside one eight-hour window, and nobody can retake a wedding. Run the three
commands above:
1. **The night of the event**, once uploads have stopped. This is the backup that
matters; everything else is a formality.
2. **After the host releases the gallery**, so the generated keepsake is captured too.
3. Weekly thereafter, until the event is archived and torn down.
Take the DB dump and the media tarball **back to back**, without uploads in flight
between them. Upload rows reference files by path — a database from 22:00 and a media
volume from 23:00 gives you rows pointing at files the dump doesn't know about, and
rows whose files aren't in the tarball. Locking uploads from the host dashboard first
(**Uploads sperren**) makes the pair genuinely consistent.
---
## Restore
An untested backup is not a backup. Run this once against a scratch host **before**
the event — it is roughly ten minutes, and it is the only way to find out that your
tarball is empty or your dump is truncated while that is still a small problem.
```bash
# 0. Stop the app FIRST. Migrations run on boot and a live pool will fight the
# restore — a booting app against a half-restored schema can leave the migration
# table and the schema disagreeing, which is its own recovery problem.
# Leave `db` running: the dump is restored through it.
docker compose stop app caddy
# 1. Database. The dump carries its own DROPs (step 1 of Backup), so this replaces
# rather than collides. A dump taken WITHOUT --clean --if-exists will abort here
# on the first "already exists" — restore that one into a fresh empty database
# instead.
gunzip -c ./backups/db_2026-07-29.sql.gz \
| docker compose exec -T db \
sh -c 'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" --set ON_ERROR_STOP=1'
# 2. Media. NOTE the `--numeric-owner` and the chown: the app runs as a
# NON-ROOT user (uid 100, gid 101 — `addgroup -S app && adduser -S app`), and a
# restore that lands root-owned files makes every upload fail with EACCES deep in
# the write path, surfacing to the guest as a generic 500 with nothing in the UI
# to suggest permissions. The explicit chown is what guarantees it — BusyBox tar
# (which is what `alpine` ships) has no --same-owner, and restores ownership only
# because it runs as root here.
docker run --rm \
-v eventsnap_media_data:/dst -v "$PWD/backups":/backup:ro \
alpine sh -c 'tar xzf /backup/media_2026-07-29.tar.gz -C /dst \
--numeric-owner && chown -R 100:101 /dst'
# 3. Exports. Same volume-name caveat, same ownership rules.
docker run --rm \
-v eventsnap_exports_data:/dst -v "$PWD/backups":/backup:ro \
alpine sh -c 'tar xzf /backup/exports_2026-07-29.tar.gz -C /dst \
--numeric-owner && chown -R 100:101 /dst'
# 4. Back up. Migrations run, then export recovery re-arms any keepsake whose file
# didn't come back with the volume.
docker compose up -d app caddy
docker compose logs -f app # watch for "migrations applied"
# 5. Verify — all three, not just the first.
curl -fsS https://DOMAIN/health && echo # → ok
# … then sign in as host and confirm the feed renders images (proves the media
# volume restored AND is readable by uid 100), and that the keepsake downloads.
```
If the media volume restored but images 404 while the feed lists them, the paths are
there and the bytes aren't — check `docker compose exec app ls -ln /media/originals`
and confirm both the files and the `100:101` ownership.
The restore is deliberately **not** automated. It is rare, destructive, and the one
operation where a script that half-works is worse than a checklist someone reads.
---
@@ -254,7 +465,7 @@ Open:
- [ ] SSE delta-fetch on foreground reconnect (scaffolded in [sse.ts](frontend/src/lib/sse.ts), not wired)
- [ ] Live diashow / slideshow mode — see [docs/CONCEPT_DIASHOW.md](docs/CONCEPT_DIASHOW.md)
- [ ] Individual file download button per post
- [ ] Low-disk alert (< 10 GB free)
- [x] Low-disk alert — host dashboard warns below 10 GB free, or whenever the keepsake would not fit
- [ ] Event banner / cover image
- [ ] Chunked resumable upload for files > 100 MB
- [ ] Shared Tailwind config between main app and export-viewer

View File

@@ -0,0 +1 @@
DELETE FROM config WHERE key IN ('join_ip_rate_per_min', 'admin_login_rate_enabled');

View File

@@ -0,0 +1,18 @@
-- Per-IP flood ceiling for /join, and the `admin_login_rate_enabled` toggle that
-- every prior migration forgot to seed.
--
-- Rationale: /join was throttled at 5 requests per 60s keyed on the client IP. At a
-- venue every guest is behind one NAT, so the whole party shared a single bucket —
-- 12 guests scanning the QR code within a few seconds meant 5 got in and 7 were
-- turned away. The handler now keys the real anti-spam bucket per (ip, name), the
-- same shape as `recover:{ip}:{name}`, and keeps only a loose per-IP ceiling to bound
-- raw volume. 60/min comfortably covers a whole wedding arriving at once while still
-- capping a flood from a single source.
--
-- `admin_login_rate_enabled` is read by auth::handlers::admin_login with a code
-- default of `true`, but no migration ever inserted it, so it was invisible to the
-- admin config UI and to the e2e reseed. Seed it explicitly.
INSERT INTO config (key, value) VALUES
('join_ip_rate_per_min', '60'),
('admin_login_rate_enabled', 'true')
ON CONFLICT (key) DO NOTHING;

View File

@@ -0,0 +1,2 @@
DROP INDEX IF EXISTS idx_upload_derivatives_rev;
ALTER TABLE upload DROP COLUMN IF EXISTS derivatives_rev;

View File

@@ -0,0 +1,18 @@
-- Track which revision of the derivative pipeline produced an upload's preview/display.
--
-- Rev 1 applies the EXIF orientation tag. Everything generated before it decoded the raw
-- sensor pixels and re-encoded to JPEG (which writes no EXIF), so every portrait phone photo
-- was stored sideways in the feed preview, the diashow display and the keepsake — while the
-- untouched original still rendered upright.
--
-- Existing rows default to 0 so the startup backfill can find and re-generate them exactly
-- once; bump the constant in services/compression.rs if the pipeline ever changes again.
ALTER TABLE upload ADD COLUMN IF NOT EXISTS derivatives_rev SMALLINT NOT NULL DEFAULT 0;
-- Only image derivatives are affected — video thumbnails are extracted by ffmpeg, which
-- already honours the rotation matrix. Mark them current so the backfill skips them.
UPDATE upload SET derivatives_rev = 1 WHERE mime_type NOT LIKE 'image/%';
CREATE INDEX IF NOT EXISTS idx_upload_derivatives_rev
ON upload (derivatives_rev)
WHERE deleted_at IS NULL;

View File

@@ -0,0 +1 @@
DELETE FROM config WHERE key = 'recover_ip_rate_per_min';

View File

@@ -0,0 +1,19 @@
-- Per-IP flood ceiling for /recover, mirroring the one migration 017 added for /join.
--
-- Rationale: /recover is keyed `recover:{ip}:{name}` at 5 per 15 minutes. That is the
-- right shape for its actual job — stopping someone who knows a display name (they are
-- visible on the feed) from burning the victim's 3-strike PIN counter and locking them
-- out repeatedly. But the name is attacker-chosen, so cycling names mints a fresh bucket
-- every time and the per-IP cost is unbounded.
--
-- Behind that limiter sits a cost-12 bcrypt verify, including an UNCONDITIONAL throwaway
-- verify for names that don't exist — deliberately, to close a timing oracle. So an
-- unknown name is the cheapest possible way to make the server do ~200ms of hashing.
-- Without a ceiling, one client can saturate the box's CPU with a name generator.
--
-- 30/min is far above any real recovery attempt (a guest tries their PIN a handful of
-- times) while capping a name-cycling flood. The per-(ip, name) bucket is unchanged and
-- remains the anti-guessing control.
INSERT INTO config (key, value) VALUES
('recover_ip_rate_per_min', '30')
ON CONFLICT (key) DO NOTHING;

View File

@@ -0,0 +1 @@
DELETE FROM config WHERE key IN ('social_rate_per_min', 'social_rate_enabled');

View File

@@ -0,0 +1,16 @@
-- Per-user rate limit for social writes (likes, comments, comment deletions).
--
-- These were the only writes in the app with no limit at all. Every other mutating
-- path -- upload, join, recover, export, admin login -- carries one; social.rs
-- carried none, so the coverage was asymmetric rather than deliberately open.
--
-- Severity is genuinely low for an invited-guest event, and the amplification worry
-- turned out to be contained: a like fans an SSE broadcast to ~100 clients, but the
-- export regeneration it could otherwise trigger is debounced (REGEN_DEBOUNCE 20s)
-- and superseded workers are inert. So this closes the gap for symmetry, not urgency,
-- and the ceiling is set high enough that no real guest will ever meet it -- a
-- double-tapping enthusiast at a wedding is not the thing being defended against.
INSERT INTO config (key, value) VALUES
('social_rate_per_min', '120'),
('social_rate_enabled', 'true')
ON CONFLICT (key) DO NOTHING;

View File

@@ -1,11 +1,12 @@
use std::time::Duration;
use axum::Json;
use axum::extract::State;
use axum::extract::{ConnectInfo, State};
use axum::http::{HeaderMap, StatusCode};
use chrono::Utc;
use rand::Rng;
use serde::{Deserialize, Serialize};
use std::net::SocketAddr;
use uuid::Uuid;
use crate::auth::jwt;
@@ -33,22 +34,32 @@ pub struct JoinResponse {
pub async fn join(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<JoinRequest>,
) -> Result<(StatusCode, Json<JoinResponse>), AppError> {
let ip = client_ip(&headers, "unknown");
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let join_rate_on = config::get_bool(&state.config_cache, "join_rate_enabled", true).await;
if rate_limits_on
&& join_rate_on
&& !state
.rate_limiter
.check(format!("join:{ip}"), 5, Duration::from_secs(60))
{
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None,
));
// Coarse per-IP flood ceiling. `/join` is pre-auth so there is no user to key on, and
// at a venue EVERY guest arrives from one public IP — a tight per-IP bucket meant the
// 6th person through the door was turned away by the 5 ahead of them. So the per-IP
// limit here only bounds raw volume; the real anti-spam bucket is per-name below.
// Cheap enough to run before validation, which keeps a flood of malformed bodies from
// being free.
if rate_limits_on && join_rate_on {
let ip_ceiling = config::get_usize(&state.config_cache, "join_ip_rate_per_min", 60).await;
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("join_ip:{ip}"),
ip_ceiling,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
let display_name = body.display_name.trim();
@@ -66,6 +77,23 @@ pub async fn join(
));
}
// Per-guest bucket, keyed like the `recover:{ip}:{name}` limiter below. This carries
// the original 5/60s anti-spam intent, but one guest retrying can no longer consume
// the allowance of everyone else sharing the venue's NAT.
if rate_limits_on && join_rate_on {
let name_key = display_name.to_lowercase();
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("join:{ip}:{name_key}"),
5,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
let event = Event::find_or_create(
&state.pool,
&state.config.event_slug,
@@ -83,7 +111,7 @@ pub async fn join(
// Generate a 4-digit PIN
let pin: String = format!("{:04}", rand::rng().random_range(0..10000u32));
let pin_hash = bcrypt::hash(&pin, 12).map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?;
let pin_hash = hash_password(pin.clone(), 12).await?;
// The pre-check above is racy: two simultaneous joins with the same name can both
// pass it, and the DB's unique index then rejects the loser. Map that unique
@@ -147,8 +175,31 @@ fn dummy_pin_hash() -> &'static str {
})
}
/// Run a bcrypt verify on the blocking pool.
///
/// bcrypt at cost 12 is ~200ms of deliberate CPU. Called inline on an async task it pins a
/// tokio WORKER thread for that whole time, and the runtime only has one per core — so a
/// flood of `/recover` or `/admin/login` attempts stalls every other request on the box,
/// including the feed. Offloading moves that cost to the blocking pool, which is sized for
/// exactly this and whose saturation degrades logins rather than the whole app.
async fn verify_password(candidate: String, hash: String) -> bool {
tokio::task::spawn_blocking(move || bcrypt::verify(&candidate, &hash).unwrap_or(false))
.await
.unwrap_or(false)
}
/// Hash a secret on the blocking pool. Same reasoning as [`verify_password`] — and this one
/// runs on the busiest auth path there is, since every guest who joins gets a PIN hashed.
pub async fn hash_password(secret: String, cost: u32) -> Result<String, AppError> {
tokio::task::spawn_blocking(move || bcrypt::hash(&secret, cost))
.await
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))
}
pub async fn recover(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<RecoverRequest>,
) -> Result<Json<RecoverResponse>, AppError> {
@@ -159,19 +210,39 @@ pub async fn recover(
// burn through 3 wrong PINs and lock the victim for 15 minutes — repeated
// every 15 minutes, indefinitely. 5 attempts per 15 minutes per (IP, name)
// softens that into a real cost.
let ip = client_ip(&headers, "unknown");
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let recover_rate_on = config::get_bool(&state.config_cache, "recover_rate_enabled", true).await;
if rate_limits_on && recover_rate_on {
// Coarse per-IP ceiling FIRST. The per-(ip, name) bucket below is the anti-guessing
// control, but the name is attacker-chosen, so cycling names mints a fresh bucket
// every time and leaves the per-IP cost unbounded. That matters more here than
// anywhere else: every call runs a cost-12 bcrypt verify, including an
// unconditional throwaway one for names that don't exist (see below), so an unknown
// name is the CHEAPEST way to make the server do ~200ms of hashing. Checked before
// the per-name bucket so a name generator can't walk past it.
let ip_ceiling =
config::get_usize(&state.config_cache, "recover_ip_rate_per_min", 30).await;
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("recover_ip:{ip}"),
ip_ceiling,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Versuche. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
let name_key = display_name.to_lowercase();
if !state.rate_limiter.check(
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("recover:{ip}:{name_key}"),
5,
Duration::from_secs(15 * 60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Versuche. Bitte warte kurz und versuche es erneut.".into(),
None,
Some(retry_after_secs),
));
}
}
@@ -188,7 +259,7 @@ pub async fn recover(
// PIN — so "no such name" and "wrong PIN" are indistinguishable by response or
// timing. Display names are already public on the feed, but this still closes
// the /recover enumeration + timing oracle.
let _ = bcrypt::verify(&body.pin, dummy_pin_hash());
let _ = verify_password(body.pin.clone(), dummy_pin_hash().to_string()).await;
return Err(AppError::Unauthorized("PIN ist falsch.".into()));
}
@@ -200,16 +271,19 @@ pub async fn recover(
// is effectively permanently fragile.
if let Some(locked_until) = user.pin_locked_until {
if Utc::now() < locked_until {
// The exact deadline is known, so surface it as Retry-After instead of
// making the client guess at the "15 Minuten" in the copy.
let retry_after_secs = (locked_until - Utc::now()).num_seconds().max(1) as u64;
return Err(AppError::TooManyRequests(
"Zu viele Versuche. Bitte warte 15 Minuten.".into(),
None,
Some(retry_after_secs),
));
}
// Lockout window expired — wipe the counter and the timestamp.
User::reset_pin_attempts(&state.pool, user.id).await?;
}
let pin_matches = bcrypt::verify(&body.pin, &user.recovery_pin_hash).unwrap_or(false);
let pin_matches = verify_password(body.pin.clone(), user.recovery_pin_hash.clone()).await;
if pin_matches {
// Reset failed attempts on success
@@ -274,6 +348,7 @@ pub struct AdminLoginResponse {
pub async fn admin_login(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<AdminLoginRequest>,
) -> Result<Json<AdminLoginResponse>, AppError> {
@@ -287,23 +362,31 @@ pub async fn admin_login(
// verify) but with no IP-level limit a determined attacker can still mount
// a long-running guess campaign. 5 attempts / minute / IP is plenty for
// honest typos.
let ip = client_ip(&headers, "unknown");
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let admin_rate_on =
config::get_bool(&state.config_cache, "admin_login_rate_enabled", true).await;
// Stays keyed by IP on purpose: this guards a single shared credential, so a per-user
// or per-name key would just hand an attacker a fresh bucket per guess.
if rate_limits_on
&& admin_rate_on
&& !state
.rate_limiter
.check(format!("admin_login:{ip}"), 5, Duration::from_secs(60))
&& let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("admin_login:{ip}"),
5,
Duration::from_secs(60),
)
{
return Err(AppError::TooManyRequests(
"Zu viele Anmeldeversuche. Bitte warte kurz und versuche es erneut.".into(),
None,
Some(retry_after_secs),
));
}
let valid = bcrypt::verify(&body.password, &state.config.admin_password_hash).unwrap_or(false);
let valid = verify_password(
body.password.clone(),
state.config.admin_password_hash.clone(),
)
.await;
if !valid {
tracing::warn!(ip = %ip, "admin_login: wrong password");
@@ -329,8 +412,7 @@ pub async fn admin_login(
let dummy_pin: String = (0..32)
.map(|_| rand::rng().random_range(b'a'..=b'z') as char)
.collect();
let dummy_hash =
bcrypt::hash(&dummy_pin, 4).map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?;
let dummy_hash = hash_password(dummy_pin.clone(), 4).await?;
let user = User::create(&state.pool, event.id, admin_name, &dummy_hash).await?;
sqlx::query("UPDATE \"user\" SET role = 'admin' WHERE id = $1")
.bind(user.id)
@@ -390,22 +472,23 @@ pub struct PinResetRequestBody {
/// feed already exposes.
pub async fn request_pin_reset(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<PinResetRequestBody>,
) -> Result<StatusCode, AppError> {
let display_name = body.display_name.trim();
let ip = client_ip(&headers, "unknown");
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
if rate_limits_on {
let name_key = display_name.to_lowercase();
if !state.rate_limiter.check(
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("pin_reset_req:{ip}:{name_key}"),
3,
Duration::from_secs(15 * 60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None,
Some(retry_after_secs),
));
}
}

View File

@@ -3,13 +3,13 @@ use std::time::Duration;
use axum::Json;
use axum::extract::{Query, State};
use axum::http::{HeaderMap, StatusCode};
use axum::http::StatusCode;
use serde::{Deserialize, Serialize};
use uuid::Uuid;
use crate::auth::middleware::RequireAdmin;
use crate::error::AppError;
use crate::services::config;
use crate::services::rate_limiter::client_ip;
use crate::state::AppState;
// ── DTOs ─────────────────────────────────────────────────────────────────────
@@ -120,6 +120,16 @@ pub async fn patch_config(
("upload_rate_per_hour", true, 1.0, 100_000.0),
("feed_rate_per_min", true, 1.0, 100_000.0),
("export_rate_per_day", true, 1.0, 100_000.0),
// Loose per-IP ceiling on /join. The real anti-spam bucket is per (ip, name); this
// only bounds raw volume from one source, so it must stay well above the size of a
// party arriving at once (see migration 017).
("join_ip_rate_per_min", true, 1.0, 100_000.0),
// Same shape for /recover: the per-(ip, name) bucket is the anti-guessing control,
// this only bounds a name-cycling flood in front of a cost-12 bcrypt (migration 019).
("recover_ip_rate_per_min", true, 1.0, 100_000.0),
// Aggregate ceiling on likes + comments + comment deletions, per user per minute.
// These were the only mutating endpoints with no limit at all (migration 020).
("social_rate_per_min", true, 1.0, 100_000.0),
("quota_tolerance", false, 0.0, 1.0),
("estimated_guest_count", true, 1.0, 1_000_000.0),
];
@@ -134,6 +144,7 @@ pub async fn patch_config(
// missing from this allowlist — so the switch existed in code and could never be flipped.
"admin_login_rate_enabled",
"recover_rate_enabled",
"social_rate_enabled",
"quota_enabled",
"storage_quota_enabled",
"upload_count_quota_enabled",
@@ -186,6 +197,23 @@ pub async fn patch_config(
"Wert für {key} liegt außerhalb des zulässigen Bereichs ({min}{max})."
)));
}
// Zero is in range and catastrophic. `quota_tolerance` is the multiplier in
// `free_disk * tolerance / active_uploaders`, so 0 makes every per-user limit 0 and
// refuses EVERY upload — mid-event, with "Du hast dein Upload-Limit für dieses Event
// erreicht", an error naming the wrong cause entirely. `storage_quota_enabled` is the
// intended off-switch.
//
// Rejecting the value rather than raising the floor: very small tolerances are
// legitimate (they are how a large disk is throttled down to a sensible per-guest
// ceiling, and how the e2e quota tests steer it — around 1e-5 on a 174 GB volume), so
// a floor of, say, 0.01 would forbid real configurations to prevent one typo.
if key_str == "quota_tolerance" && n == 0.0 {
return Err(AppError::BadRequest(
"quota_tolerance = 0 würde jeden Upload blockieren. Zum Abschalten der \
Speicher-Quote stattdessen „Speicher-Quote aktiv“ ausschalten."
.into(),
));
}
} else if BOOL_KEYS.contains(&key_str) {
match value.trim().to_ascii_lowercase().as_str() {
"true" | "false" | "1" | "0" | "yes" | "no" | "on" | "off" => {}
@@ -320,25 +348,26 @@ pub async fn export_ticket(
}
/// Validate a download ticket (single-use) and confirm its session still exists.
async fn authenticate_download_ticket(state: &AppState, ticket: &str) -> Result<(), AppError> {
/// Resolve a single-use download ticket to the user who minted it. The caller needs the
/// id to key the export rate limit per-user (see `enforce_export_rate`).
async fn authenticate_download_ticket(state: &AppState, ticket: &str) -> Result<Uuid, AppError> {
let token_hash = state
.sse_tickets
.consume(ticket)
.ok_or_else(|| AppError::Unauthorized("Ticket ungültig oder abgelaufen.".into()))?;
crate::models::session::Session::find_by_token_hash(&state.pool, &token_hash)
let session = crate::models::session::Session::find_by_token_hash(&state.pool, &token_hash)
.await
.map_err(|e| AppError::Internal(e.into()))?
.ok_or_else(|| AppError::Unauthorized("Sitzung nicht gefunden.".into()))?;
Ok(())
Ok(session.user_id)
}
pub async fn download_zip(
State(state): State<AppState>,
Query(q): Query<DownloadQuery>,
headers: HeaderMap,
) -> Result<axum::response::Response, AppError> {
authenticate_download_ticket(&state, &q.ticket).await?;
enforce_export_rate(&state, &headers).await?;
let user_id = authenticate_download_ticket(&state, &q.ticket).await?;
enforce_export_rate(&state, user_id).await?;
let path =
resolve_export_file(&state, "zip", "Der ZIP-Export ist noch nicht verfügbar.").await?;
@@ -389,10 +418,9 @@ async fn resolve_export_file(
pub async fn download_html(
State(state): State<AppState>,
Query(q): Query<DownloadQuery>,
headers: HeaderMap,
) -> Result<axum::response::Response, AppError> {
authenticate_download_ticket(&state, &q.ticket).await?;
enforce_export_rate(&state, &headers).await?;
let user_id = authenticate_download_ticket(&state, &q.ticket).await?;
enforce_export_rate(&state, user_id).await?;
let path =
resolve_export_file(&state, "html", "Der HTML-Export ist noch nicht verfügbar.").await?;
@@ -449,8 +477,13 @@ pub async fn export_status(
// worker superseded mid-run) is meaningless — surfacing its frozen `running`/77% would show a
// progress bar that never moves for a keepsake nobody is building. It reads as "locked" (no
// current job), which is exactly what it is.
let jobs: Vec<(String, String, i16)> = sqlx::query_as(
"SELECT j.type::text, j.status::text, j.progress_pct
// `error_message` is carried here, not just on the admin dashboard's job list. The host is the
// one who releases the keepsake and the one who owns the "Erneut versuchen" button, but this
// endpoint used to hand them a bare `failed` — so a fully actionable reason (notably the disk
// preflight's "needs X GB, Y GB free") was written to the row and then shown to nobody who
// could act on it. An admin-only diagnostic is not a diagnostic for the person on the spot.
let jobs: Vec<(String, String, i16, Option<String>)> = sqlx::query_as(
"SELECT j.type::text, j.status::text, j.progress_pct, j.error_message
FROM export_job j
JOIN event e ON e.id = j.event_id
WHERE e.id = $1 AND j.epoch = e.export_epoch",
@@ -461,9 +494,21 @@ pub async fn export_status(
let job_status = |type_name: &str| {
jobs.iter()
.find(|(t, _, _)| t == type_name)
.map(|(_, status, pct)| serde_json::json!({ "status": status, "progress_pct": pct }))
.unwrap_or_else(|| serde_json::json!({ "status": "locked", "progress_pct": 0 }))
.find(|(t, _, _, _)| t == type_name)
.map(|(_, status, pct, err)| {
serde_json::json!({
"status": status,
"progress_pct": pct,
// Only on a failure. A stale message left on a row that has since been re-armed
// would otherwise show an error next to a running progress bar.
"error_message": if status == "failed" { err.clone() } else { None },
})
})
.unwrap_or_else(|| {
serde_json::json!({
"status": "locked", "progress_pct": 0, "error_message": null,
})
})
};
Ok(Json(serde_json::json!({
@@ -476,21 +521,24 @@ pub async fn export_status(
/// Centralised guard for the export rate limit. Same pattern as upload/feed: master
/// switch + per-endpoint switch + numeric value, all stored in `config` and read on
/// each request.
async fn enforce_export_rate(state: &AppState, headers: &HeaderMap) -> Result<(), AppError> {
async fn enforce_export_rate(state: &AppState, user_id: Uuid) -> Result<(), AppError> {
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let export_rate_on = config::get_bool(&state.config_cache, "export_rate_enabled", true).await;
if !(rate_limits_on && export_rate_on) {
return Ok(());
}
let ip = client_ip(headers, "unknown");
let limit = config::get_usize(&state.config_cache, "export_rate_per_day", 3).await;
if !state
.rate_limiter
.check(format!("export:{ip}"), limit, Duration::from_secs(86400))
{
// Keyed per-user. This was the worst of the IP-keyed limiters: 3 downloads per DAY
// shared across every guest behind the venue's public IP, so the fourth person to
// fetch their keepsake was locked out until the next day.
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("export:{user_id}"),
limit,
Duration::from_secs(86400),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None,
Some(retry_after_secs),
));
}
Ok(())

View File

@@ -2,7 +2,6 @@ use std::time::Duration;
use axum::Json;
use axum::extract::{Query, State};
use axum::http::HeaderMap;
use chrono::{DateTime, Utc};
use serde::{Deserialize, Serialize};
use uuid::Uuid;
@@ -10,7 +9,6 @@ use uuid::Uuid;
use crate::auth::middleware::AuthUser;
use crate::error::AppError;
use crate::services::config;
use crate::services::rate_limiter::client_ip;
use crate::state::AppState;
#[derive(Deserialize)]
@@ -61,21 +59,23 @@ struct FeedRow {
pub async fn feed(
State(state): State<AppState>,
auth: AuthUser,
headers: HeaderMap,
Query(q): Query<FeedQuery>,
) -> Result<Json<FeedResponse>, AppError> {
let ip = client_ip(&headers, "unknown");
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let feed_rate_on = config::get_bool(&state.config_cache, "feed_rate_enabled", true).await;
if rate_limits_on && feed_rate_on {
let rate_limit = config::get_usize(&state.config_cache, "feed_rate_per_min", 60).await;
if !state
.rate_limiter
.check(format!("feed:{ip}"), rate_limit, Duration::from_secs(60))
{
// Keyed per-user, exactly like `feed_delta` below: at a venue every guest shares
// one public IP, so an IP key gave the whole party a single 60/min bucket and the
// fastest scroller starved everyone else.
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("feed:{}", auth.user_id),
rate_limit,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None,
Some(retry_after_secs),
));
}
}
@@ -225,14 +225,14 @@ pub async fn feed_delta(
let feed_rate_on = config::get_bool(&state.config_cache, "feed_rate_enabled", true).await;
if rate_limits_on && feed_rate_on {
let rate_limit = config::get_usize(&state.config_cache, "feed_rate_per_min", 60).await;
if !state.rate_limiter.check(
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("feed_delta:{}", auth.user_id),
rate_limit,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
None,
Some(retry_after_secs),
));
}
}

View File

@@ -35,6 +35,32 @@ pub struct EventStatus {
pub is_active: bool,
pub uploads_locked: bool,
pub export_released: bool,
/// Free space on the volume the keepsake is written to. `None` when the mount can't be
/// resolved — the UI hides the widget rather than rendering a confident zero.
pub disk_free_bytes: Option<u64>,
/// What a full keepsake build would need right now (both halves).
pub keepsake_required_bytes: u64,
/// Whether the host should be warned. See [`disk_is_low`].
pub disk_low: bool,
}
/// Absolute floor below which free space is worth surfacing regardless of gallery size — the
/// threshold the README has carried on the roadmap since v1.
const LOW_DISK_FLOOR_BYTES: u64 = 10_000_000_000;
/// Is free space low enough that the host needs to know?
///
/// Two triggers, because a fixed threshold answers the wrong question. `postgres_data`,
/// `media_data` and `exports_data` are all Docker named volumes on one filesystem, so a full disk
/// does not degrade one subsystem — it stops Postgres writing and takes the event down. That is
/// what the absolute floor is for.
///
/// The second trigger is the one that actually earns its place: the keepsake needs room for two
/// gallery-sized archives, and the only moment a host can do anything about that is BEFORE they
/// release. Warning at "you could not build the keepsake right now" turns a post-event dead end
/// into a decision someone can still make.
fn disk_is_low(free: u64, keepsake_required: u64) -> bool {
free < LOW_DISK_FLOOR_BYTES || free < keepsake_required
}
/// Count non-banned hosts/admins in the event OTHER than `excluding` — the operators
@@ -72,11 +98,29 @@ pub async fn get_event_status(
.await?
.ok_or_else(|| AppError::NotFound("Event nicht gefunden.".into()))?;
// Measured on the EXPORT volume, not the media one: that is where the cliff is, and it is a
// distinct mount point even when both are backed by the same filesystem. The cached reading is
// right here — this is advisory, polled on every dashboard load, and a 15s-stale number costs
// nothing (unlike the export preflight, which reads uncached because it is about to write).
let free = state
.disk_cache
.snapshot(&state.config.export_path)
.map(|d| d.free);
let keepsake_required_bytes =
crate::services::export::keepsake_space_required(&state.pool, event.id)
.await
.unwrap_or(0);
Ok(Json(EventStatus {
name: event.name,
is_active: event.is_active,
uploads_locked: event.uploads_locked_at.is_some(),
export_released: event.export_released_at.is_some(),
disk_free_bytes: free,
keepsake_required_bytes,
// Unknown free space is NOT low. Fails open, exactly as the upload quota and the export
// preflight do: a scary banner on an unreadable mount would train the host to ignore it.
disk_low: free.is_some_and(|f| disk_is_low(f, keepsake_required_bytes)),
}))
}
@@ -433,7 +477,7 @@ pub async fn reset_user_pin(
}
let pin: String = format!("{:04}", rand::rng().random_range(0..10000u32));
let pin_hash = bcrypt::hash(&pin, 12).map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?;
let pin_hash = crate::auth::handlers::hash_password(pin.clone(), 12).await?;
sqlx::query(
"UPDATE \"user\"
@@ -765,3 +809,45 @@ pub async fn release_gallery(
Ok(StatusCode::NO_CONTENT)
}
#[cfg(test)]
mod tests {
use super::{LOW_DISK_FLOOR_BYTES, disk_is_low};
const GB: u64 = 1_000_000_000;
#[test]
fn a_healthy_disk_with_room_for_the_keepsake_is_not_low() {
assert!(!disk_is_low(40 * GB, 25 * GB));
}
#[test]
fn the_absolute_floor_fires_even_when_the_gallery_is_tiny() {
// All three volumes share one filesystem, so running out doesn't degrade one subsystem —
// Postgres stops being able to write and the event goes down. A 1 GB gallery would clear
// the keepsake test comfortably; the floor is what catches this.
assert!(disk_is_low(5 * GB, GB));
assert!(disk_is_low(LOW_DISK_FLOOR_BYTES - 1, 0));
assert!(!disk_is_low(LOW_DISK_FLOOR_BYTES, 0));
}
#[test]
fn plenty_of_space_is_still_low_when_the_keepsake_would_not_fit() {
// THE case the fixed threshold misses, and the one that matters: 30 GB free is nowhere near
// any floor, but a 30 GB gallery needs room for TWO archives. The host can act on this
// before releasing; after releasing, they cannot.
assert!(disk_is_low(30 * GB, 66 * GB));
}
#[test]
fn the_keepsake_trigger_is_exact_at_the_boundary() {
assert!(!disk_is_low(66 * GB, 66 * GB), "exactly enough is enough");
assert!(disk_is_low(66 * GB - 1, 66 * GB));
}
#[test]
fn an_empty_gallery_needs_nothing_and_only_the_floor_applies() {
assert!(!disk_is_low(11 * GB, 0));
assert!(disk_is_low(9 * GB, 0));
}
}

View File

@@ -10,8 +10,40 @@ use crate::error::AppError;
use crate::models::comment::{Comment, CommentDto};
use crate::models::hashtag::{self, Hashtag};
use crate::models::upload::Upload;
use crate::services::config;
use crate::state::AppState;
/// Throttle a social write. Keyed PER USER, like the feed and upload limits and for the same
/// reason: at a venue every guest sits behind one NAT, so an IP key hands the whole party a
/// single bucket and the most active guest starves everyone else.
///
/// These were the only mutating endpoints in the app with no limit at all — the coverage was
/// asymmetric, not deliberately open. The ceiling is set well above anything a real guest
/// produces; this bounds a script, not an enthusiastic double-tapper.
async fn check_social_rate(state: &AppState, user_id: Uuid) -> Result<(), AppError> {
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let social_rate_on = config::get_bool(&state.config_cache, "social_rate_enabled", true).await;
if !(rate_limits_on && social_rate_on) {
return Ok(());
}
let rate_limit = config::get_usize(&state.config_cache, "social_rate_per_min", 120).await;
// ONE bucket across likes, comments and comment deletions. Separate buckets would let a
// caller triple the aggregate write rate just by alternating between them.
state
.rate_limiter
.check_with_retry(
format!("social:{user_id}"),
rate_limit,
std::time::Duration::from_secs(60),
)
.map_err(|retry_after_secs| {
AppError::TooManyRequests(
"Zu viele Aktionen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
)
})
}
#[derive(Serialize)]
pub struct LikeResponse {
/// The caller's like state *after* this toggle. The client sets `liked_by_me` from
@@ -35,6 +67,7 @@ pub async fn toggle_like(
if user.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into()));
}
check_social_rate(&state, auth.user_id).await?;
// Event-scope: the upload must belong to the caller's event (404 otherwise),
// matching the host handlers' find_by_id_and_event pattern.
@@ -141,6 +174,7 @@ pub async fn add_comment(
if user.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into()));
}
check_social_rate(&state, auth.user_id).await?;
// Event-scope: only comment on an upload that belongs to the caller's event.
Upload::find_by_id_and_event(&state.pool, upload_id, auth.event_id)
@@ -216,6 +250,7 @@ pub async fn delete_comment(
if auth.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into()));
}
check_social_rate(&state, auth.user_id).await?;
let comment = Comment::find_by_id(&state.pool, comment_id)
.await?
.ok_or_else(|| AppError::NotFound("Kommentar nicht gefunden.".into()))?;

View File

@@ -13,9 +13,10 @@ use crate::auth::middleware::RequireAdmin;
use crate::error::AppError;
use crate::state::AppState;
/// Truncates every event-scoped table, wipes media on disk, and reseeds the
/// `config` table from migration defaults. Requires an admin JWT — even with
/// `EVENTSNAP_TEST_MODE=1` it cannot be hit anonymously.
/// Truncates every event-scoped table, wipes media on disk, and reseeds the `config`
/// table: numeric values from the migration defaults, but every feature toggle forced
/// OFF (production seeds them ON — see the note at the reseed below). Requires an admin
/// JWT — even with `EVENTSNAP_TEST_MODE=1` it cannot be hit anonymously.
pub async fn truncate_all(
State(state): State<AppState>,
RequireAdmin(_auth): RequireAdmin,
@@ -40,8 +41,19 @@ pub async fn truncate_all(
.execute(&state.pool)
.await?;
// Reseed config mirrors migrations 005, 009 and 015. Kept in sync by hand
// because pulling SQL out of the migration files at runtime is fragile.
// Reseed config. The NUMERIC values mirror migrations 005/015/016/017/019; the BOOLEAN
// toggles deliberately do NOT — migration 009 seeds every one of them `true`
// (production), and this forces them `false` so the suite isn't fighting rate limits
// and quotas it isn't testing.
//
// Be aware of what that costs: this runs as an auto-fixture before EVERY test, so no
// test starts from production's config unless it explicitly turns a toggle back on
// (02-upload/rate-limit, 07-adversarial/ddos, 01-auth/rate-limit-nat, …). That blind
// spot is exactly why an entire class of per-IP limiter bugs went unnoticed: the
// limiters were simply off. When adding a limiter or quota, add a spec that enables it.
//
// Kept in sync by hand because pulling SQL out of the migration files at runtime is
// fragile — if you add a config key in a migration, add it here too.
sqlx::query(
r#"INSERT INTO config (key, value) VALUES
('max_image_size_mb', '20'),
@@ -49,6 +61,9 @@ pub async fn truncate_all(
('upload_rate_per_hour', '100'),
('feed_rate_per_min', '60'),
('export_rate_per_day', '3'),
('join_ip_rate_per_min', '60'),
('recover_ip_rate_per_min', '30'),
('social_rate_per_min', '120'),
('quota_tolerance', '0.75'),
('estimated_guest_count', '100'),
('compression_concurrency', '2'),
@@ -57,6 +72,8 @@ pub async fn truncate_all(
('feed_rate_enabled', 'false'),
('export_rate_enabled', 'false'),
('join_rate_enabled', 'false'),
('social_rate_enabled', 'false'),
('admin_login_rate_enabled', 'false'),
('quota_enabled', 'false'),
('storage_quota_enabled', 'false'),
('upload_count_quota_enabled', 'false'),

View File

@@ -233,6 +233,26 @@ pub async fn upload(
)));
}
// Images only: refuse anything the compression worker could never decode, reading just
// the header. Without this the upload is accepted with a 201 and then silently
// soft-deleted minutes later when the worker gives up — the guest sees the photo
// vanish with, at best, a vague "could not be processed". Rejecting here gives them a
// reason at the door that they can act on, and it uses the SAME budget the worker
// enforces, so admission and processing cannot disagree.
if mime.starts_with("image/") && crate::services::imaging::exceeds_decode_budget(&temp_abs) {
let mp = crate::services::imaging::megapixels(&temp_abs);
tracing::info!(
%mime, megapixels = ?mp,
"rejecting an image that exceeds the decode budget at admission"
);
let _ = tokio::fs::remove_file(&temp_abs).await;
let detail = mp.map_or(String::new(), |mp| format!(" (ca. {mp:.0} Megapixel)"));
return Err(AppError::BadRequest(format!(
"Bild hat zu viele Bildpunkte{detail} und kann nicht verarbeitet werden. \
Bitte verkleinere es und lade es erneut hoch."
)));
}
// Per-user storage quota — dynamic formula based on available disk space and the
// number of active uploaders. Gated by master + per-area toggles so the admin can
// disable it on trusted instances.
@@ -647,12 +667,95 @@ pub async fn compute_storage_quota(state: &AppState) -> QuotaEstimate {
}
}
/// Outcome of parsing a `Range` request header against a known file length.
#[derive(Debug, PartialEq, Eq)]
enum RangeSpec {
/// No `Range` header, or one we deliberately don't honour (multi-range, non-`bytes`
/// unit, malformed). RFC 9110 lets a server ignore a Range it can't process and reply
/// 200 with the full body, which is what every one of these cases does.
Full,
/// A single satisfiable range, resolved to inclusive absolute offsets.
Partial { start: u64, end: u64 },
/// Syntactically valid but starts beyond EOF — must be answered 416, not 200, or a
/// player can loop re-requesting it.
Unsatisfiable,
}
/// Parse a single-range `bytes=` header against `len`.
///
/// Deliberately supports only the three forms a media element actually sends —
/// `bytes=N-`, `bytes=N-M`, `bytes=-S` (suffix) — and treats everything else as `Full`.
/// Multi-range responses need `multipart/byteranges`, which no `<video>` requires.
fn parse_range(header: Option<&str>, len: u64) -> RangeSpec {
let Some(raw) = header else {
return RangeSpec::Full;
};
let Some(spec) = raw.trim().strip_prefix("bytes=") else {
return RangeSpec::Full;
};
// Multi-range → fall back to the whole body rather than lie about the content.
if spec.contains(',') {
return RangeSpec::Full;
}
let Some((from, to)) = spec.split_once('-') else {
return RangeSpec::Full;
};
let (from, to) = (from.trim(), to.trim());
// A zero-length file can satisfy no range at all.
if len == 0 {
return if from.is_empty() && to.is_empty() {
RangeSpec::Full
} else {
RangeSpec::Unsatisfiable
};
}
let (start, end) = if from.is_empty() {
// Suffix form: the last `to` bytes.
let Ok(suffix) = to.parse::<u64>() else {
return RangeSpec::Full;
};
if suffix == 0 {
return RangeSpec::Unsatisfiable;
}
(len.saturating_sub(suffix), len - 1)
} else {
let Ok(start) = from.parse::<u64>() else {
return RangeSpec::Full;
};
let end = if to.is_empty() {
len - 1
} else {
match to.parse::<u64>() {
// An end past EOF is clamped, not rejected (RFC 9110 §14.1.1).
Ok(end) => end.min(len - 1),
Err(_) => return RangeSpec::Full,
}
};
(start, end)
};
if start >= len || start > end {
RangeSpec::Unsatisfiable
} else {
RangeSpec::Partial { start, end }
}
}
/// Stream a media file from disk into an HTTP response with a fixed set of security
/// headers. Every media response (original, preview, thumbnail) goes through here so
/// they consistently carry `X-Content-Type-Options: nosniff` (defense-in-depth against
/// headers. Every media response (original, preview, display, thumbnail) goes through here
/// so they consistently carry `X-Content-Type-Options: nosniff` (defense-in-depth against
/// content-type confusion, even if the edge proxy is bypassed) plus an explicit
/// `Content-Disposition` and `Cache-Control`.
///
/// Honours a single `Range`. This is not an optimisation: iOS Safari opens every `<video>`
/// with a `Range: bytes=0-1` probe and abandons the load unless it gets a `206` with a
/// `Content-Range`. Without this, video is unplayable on the app's primary platform no
/// matter what `src` the element is given. `Accept-Ranges: bytes` is advertised on every
/// response so clients know seeking is available before they ask.
async fn stream_media_file(
req_headers: &axum::http::HeaderMap,
absolute: &std::path::Path,
content_type: String,
disposition: &str,
@@ -660,30 +763,60 @@ async fn stream_media_file(
) -> Result<axum::response::Response, AppError> {
use axum::body::Body;
use axum::http::{Response, StatusCode, header};
use tokio::io::{AsyncReadExt, AsyncSeekExt};
use tokio_util::io::ReaderStream;
if !absolute.exists() {
return Err(AppError::NotFound("Datei nicht gefunden.".into()));
}
let file = tokio::fs::File::open(absolute)
let mut file = tokio::fs::File::open(absolute)
.await
.map_err(|e| AppError::Internal(e.into()))?;
let metadata = file
let len = file
.metadata()
.await
.map_err(|e| AppError::Internal(e.into()))?;
let stream = ReaderStream::new(file);
.map_err(|e| AppError::Internal(e.into()))?
.len();
Response::builder()
.status(StatusCode::OK)
.header(header::CONTENT_TYPE, content_type)
.header(header::CONTENT_DISPOSITION, disposition)
.header(header::CONTENT_LENGTH, metadata.len())
.header(header::CACHE_CONTROL, cache_control)
.header(header::X_CONTENT_TYPE_OPTIONS, "nosniff")
.body(Body::from_stream(stream))
.map_err(|e| AppError::Internal(e.into()))
let range = parse_range(
req_headers.get(header::RANGE).and_then(|v| v.to_str().ok()),
len,
);
let base = |status: StatusCode| {
Response::builder()
.status(status)
.header(header::CONTENT_TYPE, content_type.clone())
.header(header::CONTENT_DISPOSITION, disposition)
.header(header::CACHE_CONTROL, cache_control)
.header(header::X_CONTENT_TYPE_OPTIONS, "nosniff")
.header(header::ACCEPT_RANGES, "bytes")
};
match range {
RangeSpec::Full => base(StatusCode::OK)
.header(header::CONTENT_LENGTH, len)
.body(Body::from_stream(ReaderStream::new(file)))
.map_err(|e| AppError::Internal(e.into())),
RangeSpec::Partial { start, end } => {
file.seek(std::io::SeekFrom::Start(start))
.await
.map_err(|e| AppError::Internal(e.into()))?;
let span = end - start + 1;
base(StatusCode::PARTIAL_CONTENT)
.header(header::CONTENT_LENGTH, span)
.header(header::CONTENT_RANGE, format!("bytes {start}-{end}/{len}"))
.body(Body::from_stream(ReaderStream::new(file.take(span))))
.map_err(|e| AppError::Internal(e.into()))
}
RangeSpec::Unsatisfiable => base(StatusCode::RANGE_NOT_SATISFIABLE)
.header(header::CONTENT_RANGE, format!("bytes */{len}"))
.body(Body::empty())
.map_err(|e| AppError::Internal(e.into())),
}
}
/// Streaming download of the original file behind an upload. Used by:
@@ -700,6 +833,7 @@ async fn stream_media_file(
/// [`get_preview`] / [`get_thumbnail`]).
pub async fn get_original(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
@@ -711,10 +845,22 @@ pub async fn get_original(
.file_name()
.and_then(|n| n.to_str())
.unwrap_or("original");
let disposition = format!("attachment; filename=\"{filename}\"");
// `inline`, not `attachment`. This route is the only source of playable video bytes
// (there is no video derivative), and an attachment disposition is hostile to a
// `<video>` element — Safari in particular. It also matches what the UI promises:
// the action is labelled "Original anzeigen", i.e. view, not download.
let disposition = format!("inline; filename=\"{filename}\"");
// Full-res original: force download, never cache at the edge.
stream_media_file(&absolute, media.mime_type, &disposition, "no-store").await
// Full-res original: never cache at the edge, so a takedown revokes access promptly.
// Range requests still work under no-store; the client simply re-fetches each range.
stream_media_file(
&headers,
&absolute,
media.mime_type,
&disposition,
"no-store",
)
.await
}
/// Streaming access to an upload's compressed **preview** image. Gated exactly like
@@ -729,6 +875,7 @@ pub async fn get_original(
/// `upload-deleted` / `user-hidden` SSE events, so this only bounds the raw-URL edge case.
pub async fn get_preview(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
@@ -739,6 +886,7 @@ pub async fn get_preview(
.ok_or_else(|| AppError::NotFound("Vorschau nicht verfügbar.".into()))?;
let absolute = state.config.media_path.join(&rel);
stream_media_file(
&headers,
&absolute,
"image/jpeg".to_string(),
"inline",
@@ -753,6 +901,7 @@ pub async fn get_preview(
/// back to the original.
pub async fn get_display(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
@@ -763,6 +912,7 @@ pub async fn get_display(
.ok_or_else(|| AppError::NotFound("Anzeige nicht verfügbar.".into()))?;
let absolute = state.config.media_path.join(&rel);
stream_media_file(
&headers,
&absolute,
"image/jpeg".to_string(),
"inline",
@@ -775,6 +925,7 @@ pub async fn get_display(
/// [`get_preview`].
pub async fn get_thumbnail(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
@@ -785,6 +936,7 @@ pub async fn get_thumbnail(
.ok_or_else(|| AppError::NotFound("Thumbnail nicht verfügbar.".into()))?;
let absolute = state.config.media_path.join(&rel);
stream_media_file(
&headers,
&absolute,
"image/jpeg".to_string(),
"inline",
@@ -795,7 +947,108 @@ pub async fn get_thumbnail(
#[cfg(test)]
mod tests {
use super::quota_limit_bytes;
use super::{RangeSpec, parse_range, quota_limit_bytes};
// `Range` handling exists because iOS Safari probes every `<video>` with
// `Range: bytes=0-1` and abandons the load without a 206. These pin the forms a
// media element actually sends, plus the edges that decide 206 vs 200 vs 416.
#[test]
fn no_range_header_is_a_full_response() {
assert_eq!(parse_range(None, 100), RangeSpec::Full);
}
#[test]
fn open_ended_range_runs_to_eof() {
assert_eq!(
parse_range(Some("bytes=10-"), 100),
RangeSpec::Partial { start: 10, end: 99 }
);
}
#[test]
fn closed_range_is_inclusive_on_both_ends() {
// The iOS probe. Two bytes, 0 and 1 — an exclusive end would return one.
assert_eq!(
parse_range(Some("bytes=0-1"), 100),
RangeSpec::Partial { start: 0, end: 1 }
);
}
#[test]
fn suffix_range_returns_the_last_n_bytes() {
assert_eq!(
parse_range(Some("bytes=-20"), 100),
RangeSpec::Partial { start: 80, end: 99 }
);
}
#[test]
fn suffix_longer_than_the_file_clamps_to_the_whole_file() {
assert_eq!(
parse_range(Some("bytes=-500"), 100),
RangeSpec::Partial { start: 0, end: 99 }
);
}
#[test]
fn end_past_eof_is_clamped_not_rejected() {
// RFC 9110 §14.1.1 — players routinely ask for more than is there.
assert_eq!(
parse_range(Some("bytes=90-999"), 100),
RangeSpec::Partial { start: 90, end: 99 }
);
}
#[test]
fn start_past_eof_is_416_not_a_full_body() {
// Answering 200 here makes a player re-request forever.
assert_eq!(
parse_range(Some("bytes=100-"), 100),
RangeSpec::Unsatisfiable
);
assert_eq!(
parse_range(Some("bytes=200-300"), 100),
RangeSpec::Unsatisfiable
);
}
#[test]
fn inverted_range_is_unsatisfiable() {
assert_eq!(
parse_range(Some("bytes=50-10"), 100),
RangeSpec::Unsatisfiable
);
}
#[test]
fn unsupported_or_malformed_forms_fall_back_to_the_full_body() {
// Ignoring a Range we can't process and sending 200 is explicitly allowed, and
// safer than guessing. Multi-range would need multipart/byteranges, which no
// <video> asks for.
for header in [
"bytes=0-10,20-30", // multi-range
"items=0-10", // non-bytes unit
"bytes=abc-def", // garbage
"bytes=", // empty spec
"nonsense", // no unit at all
] {
assert_eq!(parse_range(Some(header), 100), RangeSpec::Full, "{header}");
}
}
#[test]
fn last_byte_is_reachable() {
assert_eq!(
parse_range(Some("bytes=99-99"), 100),
RangeSpec::Partial { start: 99, end: 99 }
);
}
#[test]
fn empty_file_satisfies_no_range() {
assert_eq!(parse_range(Some("bytes=0-"), 0), RangeSpec::Unsatisfiable);
assert_eq!(parse_range(None, 0), RangeSpec::Full);
}
#[test]
fn divides_free_space_by_uploaders_with_tolerance() {

View File

@@ -2,7 +2,6 @@ use anyhow::Result;
use axum::Router;
use axum::extract::DefaultBodyLimit;
use axum::routing::{delete, get, patch, post};
use tower_http::services::ServeDir;
use tower_http::trace::TraceLayer;
use tracing_subscriber::{layer::SubscriberExt, util::SubscriberInitExt};
@@ -45,10 +44,12 @@ async fn main() -> Result<()> {
let state = AppState::new(pool.clone(), config.clone());
// Backfill the big-screen display derivative for image uploads processed before it
// existed (v0.17.x). Fire-and-forget behind the compression semaphore; each upload
// falls back to its original in the diashow until its display is generated.
state.compression.backfill_missing_display().await;
// Regenerate image derivatives an older pipeline produced: the big-screen display for
// uploads processed before it existed (v0.17.x), and anything predating the current
// DERIVATIVES_REV (rev 1 applies the EXIF orientation, without which every portrait
// phone photo is stored sideways). Fire-and-forget behind the compression semaphore;
// originals are never touched, so a failure just retries on the next start.
state.compression.backfill_stale_derivatives().await;
// Re-spawn exports for events that were released but whose keepsake never finished
// (crash mid-export). Needs the media/export paths + SSE sender, so it runs here
@@ -69,6 +70,7 @@ async fn main() -> Result<()> {
pool,
state.rate_limiter.clone(),
state.sse_tickets.clone(),
config.media_path.clone(),
);
// Ensure media directories exist
@@ -229,45 +231,39 @@ async fn main() -> Result<()> {
api
};
// Serve media files from disk
let media_service = ServeDir::new(&config.media_path);
// NOTE: media is deliberately NOT served over HTTP.
//
// Files live under `media_path` so the compression worker and the export job can read
// them off disk, but nothing may pull them straight from `/media/**` — that bypasses
// the visibility checks (soft-delete + ban-hide) that make a host takedown stick.
// Every legitimate fetch goes through `/api/v1/upload/{id}/{original,preview,display,
// thumbnail}`, which filter via `find_visible_media`; those are the only media URLs the
// backend ever emits (see `handlers::feed`).
//
// This used to be a `ServeDir` on `/media` with four `nest_service` blockers on the
// subtrees above it. That was bypassable: axum routes on the RAW path while `ServeDir`
// percent-decodes afterwards, so `/media/%70reviews/{id}.jpg` missed every blocker,
// fell through to the `ServeDir`, and was decoded back to `previews/` on disk — serving
// a taken-down photo to anyone, unauthenticated. Any single escaped byte worked, in all
// four subtrees. Deleting the route removes the vector outright rather than racing the
// decoder; `/media/**` now 404s regardless of encoding.
let router = Router::new()
.route("/health", get(|| async { "ok" }))
.merge(api)
// Block direct HTTP access to ALL media subtrees. The files live under
// `media_path` (so the compression worker and export can read them off disk) but
// must NOT be pullable straight from `/media/**` — that bypasses the visibility
// checks (soft-delete + ban-hide) in the gated handlers. Every legitimate fetch
// goes through `/api/v1/upload/{id}/{original,preview,thumbnail}`, which filter
// via `find_visible_media`. The more specific nests take precedence over the
// `/media` ServeDir below (which, with all three subtrees blocked, now serves
// nothing — kept as a backstop).
.nest_service(
"/media/originals",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service(
"/media/previews",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service(
"/media/displays",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service(
"/media/thumbnails",
get(|| async { axum::http::StatusCode::NOT_FOUND }),
)
.nest_service("/media", media_service)
.layer(TraceLayer::new_for_http())
.with_state(state);
let listener = tokio::net::TcpListener::bind(("0.0.0.0", config.app_port)).await?;
tracing::info!("listening on {}", listener.local_addr()?);
axum::serve(listener, router)
.with_graceful_shutdown(shutdown_signal())
.await?;
// `into_make_service_with_connect_info` is required by the pre-auth handlers, which
// extract `ConnectInfo<SocketAddr>` to use the peer address as the rate-limit key when
// X-Forwarded-For is absent. Without it those extractors fail at runtime.
axum::serve(
listener,
router.into_make_service_with_connect_info::<std::net::SocketAddr>(),
)
.with_graceful_shutdown(shutdown_signal())
.await?;
Ok(())
}

View File

@@ -148,6 +148,17 @@ impl Upload {
Ok(())
}
/// Stamp which revision of the derivative pipeline produced this row's preview/display,
/// so the startup backfill can find rows generated by an older one exactly once.
pub async fn set_derivatives_rev(pool: &PgPool, id: Uuid, rev: i16) -> Result<(), sqlx::Error> {
sqlx::query("UPDATE upload SET derivatives_rev = $2 WHERE id = $1")
.bind(id)
.bind(rev)
.execute(pool)
.await?;
Ok(())
}
pub async fn set_thumbnail_path(
pool: &PgPool,
id: Uuid,

View File

@@ -48,6 +48,16 @@ impl CompressionWorker {
self.generation.fetch_add(1, Ordering::SeqCst);
}
/// How many times `do_process` is attempted before an upload is given up on. The
/// give-up path is user-visible (the photo disappears), so transient infrastructure
/// errors must not reach it.
const MAX_PROCESS_ATTEMPTS: u32 = 3;
/// Revision of the image-derivative pipeline. Bump this whenever a change makes existing
/// previews/displays wrong, so `backfill_stale_derivatives` regenerates them once on the
/// next start. Rev 1 = EXIF orientation is applied.
const DERIVATIVES_REV: i16 = 1;
/// Spawn a background task to process an uploaded file.
pub fn process(&self, upload_id: Uuid, original_path: String, mime_type: String) {
let worker = self.clone();
@@ -60,10 +70,42 @@ impl CompressionWorker {
if worker.generation.load(Ordering::SeqCst) != born_at {
return;
}
match worker
.do_process(upload_id, &original_path, &mime_type)
.await
{
// Retry before giving up. Most failures here are transient and self-clearing —
// an ENOSPC spike while several guests upload at once, a momentary DB-pool
// exhaustion, a panic inside the image codec — and the give-up path is
// user-visible data loss, so it is worth a few seconds to avoid entering it.
//
// But only for failures that CAN clear. An image that exceeds the decode budget,
// is corrupt, or is in an unsupported format fails identically on every attempt,
// so retrying it just burns 2s + 4s of backoff and writes three near-identical
// warnings before reaching the same conclusion. Give up on those immediately.
let mut attempt = 1u32;
let outcome = loop {
match worker
.do_process(upload_id, &original_path, &mime_type)
.await
{
Ok(v) => break Ok(v),
Err(e)
if attempt < Self::MAX_PROCESS_ATTEMPTS
&& !crate::services::imaging::is_permanent_image_error(&e) =>
{
tracing::warn!(
error = ?e, %upload_id, attempt,
"compression attempt failed; retrying"
);
tokio::time::sleep(std::time::Duration::from_secs(2u64.pow(attempt))).await;
attempt += 1;
// The data may have been reset while we slept (e2e TRUNCATE).
if worker.generation.load(Ordering::SeqCst) != born_at {
return;
}
}
Err(e) => break Err(e),
}
};
match outcome {
Ok(_) => {
tracing::info!("compression completed for upload {upload_id}");
let _ = worker.sse_tx.send(SseEvent {
@@ -72,21 +114,31 @@ impl CompressionWorker {
});
}
Err(e) => {
tracing::error!("compression failed for upload {upload_id}: {e:#}");
// Auto-cleanup: a failed transcode would otherwise leave a
// permanently broken feed card, silently charge the uploader's
// quota, and orphan the original on disk. Refund + soft-delete
// (one tx, so v_feed excludes it), remove the orphan file, then
// tell the uploader (upload-error toast) and evict the card
// everywhere (upload-deleted, already handled by the feed).
tracing::error!(
"compression failed for upload {upload_id} after {attempt} attempt(s): {e:#}"
);
// Refund + soft-delete (one tx, so v_feed excludes it) so a failed
// transcode doesn't leave a permanently broken feed card or silently
// charge the uploader's quota. Then tell the uploader (upload-error
// toast) and evict the card everywhere (upload-deleted).
//
// The ORIGINAL IS DELIBERATELY KEPT. This path used to `remove_file` it
// unconditionally, which meant any transient error — a disk-full blip
// while saving a derivative, a pool hiccup, a panic in the image codec —
// irreversibly destroyed the guest's only copy of a photo they can never
// retake. The row is only soft-deleted, so keeping the bytes makes the
// upload fully recoverable; the file is orphaned rather than lost, and
// the path is logged so it can be found. `backfill_stale_derivatives`
// already refuses to destroy data on error for exactly this reason.
let _ = Upload::set_compression_status(&worker.pool, upload_id, "failed").await;
if let Err(del) = Upload::soft_delete(&worker.pool, upload_id).await {
tracing::warn!(error = ?del, %upload_id, "failed to soft-delete after compression failure");
}
let orphan = worker.media_path.join(&original_path);
if let Err(rm) = tokio::fs::remove_file(&orphan).await {
tracing::warn!(error = ?rm, path = %orphan.display(), "failed to remove orphaned original");
}
tracing::warn!(
%upload_id,
path = %worker.media_path.join(&original_path).display(),
"original retained for recovery after compression failure"
);
let _ = worker.sse_tx.send(SseEvent {
event_type: "upload-error".to_string(),
data: serde_json::json!({ "upload_id": upload_id, "error": e.to_string() })
@@ -117,6 +169,7 @@ impl CompressionWorker {
.await?;
Upload::set_preview_path(&self.pool, upload_id, &preview_rel).await?;
Upload::set_display_path(&self.pool, upload_id, &display_rel).await?;
Upload::set_derivatives_rev(&self.pool, upload_id, Self::DERIVATIVES_REV).await?;
tracing::info!("preview + display generated for upload {upload_id}");
} else if mime_type.starts_with("video/") {
let thumb_rel = self.generate_video_thumbnail(upload_id, &original).await?;
@@ -158,21 +211,9 @@ impl CompressionWorker {
// Run blocking image operations in a spawn_blocking task
tokio::task::spawn_blocking(move || -> Result<()> {
// Reject decompression bombs *before* fully decoding: the upload body
// cap bounds the file size on disk, but a small file can still decode to
// enormous dimensions (e.g. a ~1 MB image expanding to 50k×50k px →
// gigabytes), OOM-ing the box during decode/resize. 12000×12000 covers
// any real phone photo; max_alloc hard-caps the decode allocation.
let mut reader = image::ImageReader::open(&original)
.context("failed to open image")?
.with_guessed_format()
.context("failed to read image header")?;
let mut limits = image::Limits::default();
limits.max_image_width = Some(12_000);
limits.max_image_height = Some(12_000);
limits.max_alloc = Some(256 * 1024 * 1024);
reader.limits(limits);
let img = reader.decode().context("failed to decode image")?;
// Decompression-bomb limits + EXIF orientation, both in one place — see
// services::imaging for why neither may be skipped.
let img = crate::services::imaging::decode_oriented(&original)?;
// Preview: max 800px, preserving aspect ratio (data-saver feed).
img.resize(
@@ -221,33 +262,41 @@ impl CompressionWorker {
))
}
/// One-time backfill: existing image uploads processed before the display derivative
/// existed have a preview but no `display_path`. Regenerate both derivatives for them
/// (decode is cheap and idempotent) and set the path. Unlike the failure path in
/// `process`, a backfill error is logged and skipped — it must NEVER soft-delete an
/// upload that already has a working preview. Fire-and-forget from startup.
pub async fn backfill_missing_display(&self) {
/// Regenerate image derivatives that an older pipeline produced. Fire-and-forget from
/// startup; picks up two cases, both of which leave the ORIGINAL untouched:
///
/// - uploads processed before the `display` derivative existed (preview but no
/// `display_path`), and
/// - uploads whose derivatives predate `DERIVATIVES_REV` — currently rev 1, which applies
/// the EXIF orientation. Everything generated before it is stored sideways for any
/// portrait phone photo.
///
/// Unlike the failure path in `process`, a backfill error is logged and skipped — it must
/// NEVER destroy or soft-delete an upload that already has a working preview.
pub async fn backfill_stale_derivatives(&self) {
let rows = sqlx::query_as::<_, (Uuid, String, String)>(
"SELECT id, original_path, mime_type FROM upload
WHERE display_path IS NULL AND preview_path IS NOT NULL
AND deleted_at IS NULL AND mime_type LIKE 'image/%'",
WHERE deleted_at IS NULL AND mime_type LIKE 'image/%'
AND original_path IS NOT NULL
AND (
(display_path IS NULL AND preview_path IS NOT NULL)
OR derivatives_rev < $1
)",
)
.bind(Self::DERIVATIVES_REV)
.fetch_all(&self.pool)
.await;
let rows = match rows {
Ok(r) => r,
Err(e) => {
tracing::warn!(error = ?e, "display backfill query failed");
tracing::warn!(error = ?e, "derivative backfill query failed");
return;
}
};
if rows.is_empty() {
return;
}
tracing::info!(
"backfilling display derivative for {} upload(s)",
rows.len()
);
tracing::info!("regenerating derivatives for {} upload(s)", rows.len());
for (id, original_path, mime_type) in rows {
let worker = self.clone();
tokio::spawn(async move {
@@ -260,12 +309,16 @@ impl CompressionWorker {
Ok((preview_rel, display_rel)) => {
let _ = Upload::set_preview_path(&worker.pool, id, &preview_rel).await;
let _ = Upload::set_display_path(&worker.pool, id, &display_rel).await;
tracing::info!("display backfilled for upload {id}");
let _ =
Upload::set_derivatives_rev(&worker.pool, id, Self::DERIVATIVES_REV)
.await;
tracing::info!("derivatives regenerated for upload {id}");
}
Err(e) => {
// Leave the existing preview intact; the diashow falls back to the
// original for this upload until a later successful pass.
tracing::warn!(error = ?e, %id, "display backfill failed; leaving as-is");
// Leave the existing derivatives and the original intact; this row is
// simply retried on the next start. The rev stays behind, which is the
// marker that it still needs doing.
tracing::warn!(error = ?e, %id, "derivative backfill failed; leaving as-is");
}
}
});

View File

@@ -76,6 +76,17 @@ impl Default for DiskCache {
}
}
/// UNCACHED free-space reading for the filesystem backing `path`.
///
/// Deliberately bypasses [`DiskCache`]. The cache exists for the quota poll, where a 15s-stale
/// number is fine because it is only ever advisory. The export preflight is the opposite case: it
/// decides whether to start writing a multi-GB archive, and the sibling export worker running
/// concurrently can move free space by tens of gigabytes well inside the TTL. A stale reading there
/// would authorise exactly the write that fills the disk.
pub fn free_bytes(path: &Path) -> Option<u64> {
read_disk_for_path(path).map(|d| d.free)
}
/// Resolve the filesystem backing `media_path` and read its total/free bytes.
///
/// Snapshots the mount table via `sysinfo`, then delegates the selection to the pure

View File

@@ -20,6 +20,29 @@ use crate::state::SseEvent;
static VIEWER_DIR: Dir<'_> = include_dir!("$CARGO_MANIFEST_DIR/static/export-viewer");
// ── Shared visibility filter ─────────────────────────────────────────────────
/// The predicate that decides what lands in a keepsake, as ONE definition.
///
/// Two queries have to agree on it: [`query_uploads`], which selects the rows the archives are
/// built from, and [`estimate_export_bytes`], which sizes them for the disk preflight. They used
/// to state it separately, and the direction of drift matters — an estimate that misses rows the
/// archive writes UNDER-reserves, which is the exact ENOSPC the preflight exists to prevent.
///
/// A `SRC:`-marked copy in the integration tests cannot catch that: drift means production moved
/// and the copy didn't, so both sides of such a test sit still and it keeps passing. Sharing the
/// fragment removes the failure by construction instead, and leaves the test doing what it is
/// actually good at — pinning the behaviour.
///
/// CONTRACT: callers must alias `upload` as `u` and join `"user"` as `usr`, and bind the event id
/// as `$1`.
macro_rules! export_visibility_where {
() => {
"WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE"
};
}
// ── DB query rows ────────────────────────────────────────────────────────────
#[derive(sqlx::FromRow)]
@@ -476,6 +499,15 @@ async fn run_zip_export(
return Ok(());
}
// Reclaim BEFORE measuring: the superseded archive is already unreachable, and the space it
// holds is very often exactly the space this rebuild needs.
prune_superseded_archives(pool, export_path, "Gallery", event_id, epoch).await;
// AFTER the claim, not before. A preflight that bailed before claiming would leave the row
// `pending` with no worker and no error — the spinner-forever state `mark_failed`'s status
// guard was widened to prevent. Failing here goes through the caller's `mark_failed`, so the
// host gets the reason.
ensure_export_space(pool, event_id, export_path).await?;
// On error, mark THIS generation failed — a no-op if we've since been superseded (the
// caller in `spawn_export_jobs` does it, epoch-guarded). Temp artifacts are cleaned up
// here so a failing export can't leak them (which is what fills the disk in the first place).
@@ -548,7 +580,7 @@ async fn run_zip_export_inner(
};
let entry_name = format!("{folder}/{date}_{name_safe}_{}.{ext}", row.id);
let builder = ZipEntryBuilder::new(entry_name.into(), Compression::Stored);
let builder = keepsake_entry(entry_name, Compression::Stored);
// Open BEFORE writing the entry header. A missing source is skipped (the media file
// was deleted, or its processing failed) — but it must be skipped without aborting the
@@ -658,6 +690,11 @@ async fn run_html_export(
return Ok(());
}
// See run_zip_export: reclaim the superseded generation first, then refuse at the door rather
// than ENOSPC mid-write.
prune_superseded_archives(pool, export_path, "Memories", event_id, epoch).await;
ensure_export_space(pool, event_id, export_path).await?;
let res = run_html_export_inner(
epoch,
event_id,
@@ -720,10 +757,11 @@ async fn run_html_export_inner(
let src = media_path.join(&row.original_path);
// Stat ONCE, up front, and skip this upload if the source is gone. The old code probed with
// `exists()` here and then did `metadata(&src).await?` further down — a TOCTOU whose `?`
// aborted the ENTIRE keepsake if the file vanished in between. It genuinely can: the
// compression worker hard-deletes an original when its transcode fails, and it can still be
// running when the gallery is released. A missing source must degrade one entry, never the
// whole archive (which, once released, the host cannot rebuild without reopening uploads).
// aborted the ENTIRE keepsake if the file vanished in between. It can still happen: the
// compression worker no longer deletes originals on failure, but the hourly sweep reclaims
// them once past the retention window, and an owner or host delete can land mid-export. A
// missing source must degrade one entry, never the whole archive (which, once released, the
// host cannot rebuild without reopening uploads).
let src_meta = match tokio::fs::metadata(&src).await {
Ok(m) => m,
Err(e) => {
@@ -790,7 +828,12 @@ async fn run_html_export_inner(
let thumb_path_clone = thumb_path.clone();
let thumb_result = tokio::task::spawn_blocking(move || -> Result<()> {
let img = image::open(&src_clone).context("failed to open image for thumbnail")?;
// `decode_oriented`, not `image::open`: the latter ignores the EXIF
// orientation tag AND applies no decode limits. Using it here is why every
// portrait photo came out sideways in the keepsake's HTML viewer grid — the
// re-encode below drops the tag, so the viewer cannot recover it.
let img = crate::services::imaging::decode_oriented(&src_clone)
.context("failed to open image for thumbnail")?;
let resized = img.resize(400, 400, image::imageops::FilterType::Lanczos3);
resized
.save_with_format(&thumb_path_clone, image::ImageFormat::Jpeg)
@@ -811,8 +854,12 @@ async fn run_html_export_inner(
let full_path_clone = full_path.clone();
let compress_result = tokio::task::spawn_blocking(move || -> Result<()> {
let img =
image::open(&src_clone).context("failed to open image for compression")?;
// Same reason as the thumbnail above. This branch only runs for originals
// over 5 MB, which is why the viewer's full image looked correct for small
// photos and sideways for large ones — an inconsistency that reads as a
// viewer bug rather than an export one.
let img = crate::services::imaging::decode_oriented(&src_clone)
.context("failed to open image for compression")?;
let resized = img.resize(2000, 2000, image::imageops::FilterType::Lanczos3);
resized
.save_with_format(&full_path_clone, image::ImageFormat::Jpeg)
@@ -926,7 +973,7 @@ async fn run_html_export_inner(
// Write data.json
{
let builder = ZipEntryBuilder::new("data.json".into(), Compression::Deflate);
let builder = keepsake_entry("data.json".into(), Compression::Deflate);
let mut entry = zip.write_entry_stream(builder).await?;
let mut cursor = AllowStdIo::new(std::io::Cursor::new(data_json.as_bytes()));
fcopy(&mut cursor, &mut entry).await?;
@@ -935,7 +982,7 @@ async fn run_html_export_inner(
// Write README.txt
{
let builder = ZipEntryBuilder::new("README.txt".into(), Compression::Deflate);
let builder = keepsake_entry("README.txt".into(), Compression::Deflate);
let mut entry = zip.write_entry_stream(builder).await?;
let mut cursor = AllowStdIo::new(std::io::Cursor::new(README_TEXT.as_bytes()));
fcopy(&mut cursor, &mut entry).await?;
@@ -953,9 +1000,9 @@ async fn run_html_export_inner(
for (name, source) in &media_manifest {
let path = source.path();
// Open-first: a source that disappeared between the manifest being built and now (the
// compression worker deletes originals on transcode failure) must skip this entry, not
// fail the whole viewer. Opening collapses the check and the use into one operation.
// Open-first: a source that disappeared between the manifest being built and now (a
// delete, or the hourly sweep reclaiming a long-failed original) must skip this entry,
// not fail the whole viewer. Opening collapses the check and the use into one operation.
let src_file = match tokio::fs::File::open(path).await {
Ok(f) => f,
Err(e) => {
@@ -968,7 +1015,7 @@ async fn run_html_export_inner(
};
let entry_name = format!("media/{name}");
let builder = ZipEntryBuilder::new(entry_name.into(), Compression::Stored);
let builder = keepsake_entry(entry_name, Compression::Stored);
let mut zip_entry = zip.write_entry_stream(builder).await?;
let mut f = src_file.compat();
fcopy(&mut f, &mut zip_entry).await?;
@@ -1027,7 +1074,7 @@ async fn run_html_export_inner(
// ── DB helpers ───────────────────────────────────────────────────────────────
async fn query_uploads(pool: &PgPool, event_id: Uuid) -> Result<Vec<ExportUploadRow>> {
Ok(sqlx::query_as::<_, ExportUploadRow>(
Ok(sqlx::query_as::<_, ExportUploadRow>(concat!(
"SELECT u.id, u.original_path, u.mime_type, u.caption,
usr.display_name AS uploader_name,
COUNT(DISTINCT l.user_id) AS like_count,
@@ -1035,11 +1082,12 @@ async fn query_uploads(pool: &PgPool, event_id: Uuid) -> Result<Vec<ExportUpload
FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
LEFT JOIN \"like\" l ON l.upload_id = u.id
WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE
",
export_visibility_where!(),
"
GROUP BY u.id, usr.display_name
ORDER BY u.created_at ASC",
)
))
.bind(event_id)
.fetch_all(pool)
.await?)
@@ -1151,6 +1199,197 @@ fn parse_gen_seq(name: &str, prefix: &str, suffix: &str) -> Option<i64> {
.ok()
}
/// Filenames a live (current-epoch, `done`) job row still points at — OFF LIMITS to every prune,
/// regardless of the epoch encoded in the name.
///
/// A ViewerOnly regeneration carries the finished ZIP forward by re-stamping its row to the new
/// epoch WITHOUT renaming the file, so `Gallery.<event>.<older>.zip` is still the served archive and
/// deleting it by filename-epoch would 404 the download.
async fn protected_files(pool: &PgPool, event_id: Uuid) -> Vec<String> {
sqlx::query_scalar::<_, String>(
"SELECT split_part(j.file_path, '/', -1) FROM export_job j
JOIN event e ON e.id = j.event_id
WHERE j.event_id = $1 AND j.epoch = e.export_epoch
AND j.status = 'done' AND j.file_path IS NOT NULL",
)
.bind(event_id)
.fetch_all(pool)
.await
.unwrap_or_default()
}
/// Reclaim superseded FINAL archives BEFORE this generation starts writing its own.
///
/// Peak disk usage used to be two full generations, because the only prune ran after the new archive
/// was written, renamed and finalised. That ordering reads as durability ("don't delete the good
/// keepsake before the replacement is safe") but it buys nothing: readiness is derived from
/// `job.epoch = event.export_epoch AND status = 'done'`, so the moment `invalidate_and_arm` bumps
/// the epoch the old archive is ALREADY unreachable — `GET /export/zip` 404s whether the file is on
/// disk or not. Keeping it only reserves gigabytes for a download nobody can perform, and for
/// `Affects::Both` (a takedown) it is content someone has explicitly asked to have removed. So a
/// rebuild reclaims first and peaks at one generation.
///
/// Narrower than [`prune_stale_export_files`] on purpose: FINAL archives only. Those are inert — a
/// superseded worker either already renamed its file (and will delete it itself when its guarded
/// `finalize_job` fails) or never will. `.tmp` files and `viewer_tmp_` staging dirs are NOT touched
/// here: a superseded worker can still be streaming into those, and at build START it is far more
/// likely to be alive than at finalize time.
async fn prune_superseded_archives(
pool: &PgPool,
exports_dir: &Path,
prefix: &str,
event_id: Uuid,
keep_seq: i64,
) {
let protected = protected_files(pool, event_id).await;
let final_prefix = format!("{prefix}.{event_id}.");
let mut rd = match tokio::fs::read_dir(exports_dir).await {
Ok(rd) => rd,
Err(_) => return,
};
let mut reclaimed = 0u64;
while let Ok(Some(entry)) = rd.next_entry().await {
let name = entry.file_name();
let name = name.to_string_lossy();
if !is_superseded_archive(&name, &final_prefix, keep_seq, &protected) {
continue;
}
let len = entry.metadata().await.map(|m| m.len()).unwrap_or(0);
if tokio::fs::remove_file(entry.path()).await.is_ok() {
reclaimed += len;
}
}
if reclaimed > 0 {
tracing::info!(
"reclaimed {reclaimed} bytes of superseded {prefix} archives before rebuilding \
event {event_id} @ epoch {keep_seq}"
);
}
}
/// Is `name` a FINAL archive of a strictly-older generation that no live row points at?
///
/// Pure so the two dangerous cases can be pinned without a filesystem: the carried-forward archive
/// (protected despite an older epoch in its name) and the in-flight `.tmp` (never matched at all).
fn is_superseded_archive(
name: &str,
final_prefix: &str,
keep_seq: i64,
protected: &[String],
) -> bool {
if protected.iter().any(|p| p == name) {
return false;
}
parse_gen_seq(name, final_prefix, ".zip").is_some_and(|n| n < keep_seq)
}
/// Bytes the media in this event's keepsake will occupy, as an UPPER BOUND per archive.
///
/// Both archives write their media entries `Compression::Stored`, so an archive is essentially a
/// byte-for-byte second copy of the originals: `Gallery.zip` always, and `Memories.zip` for every
/// video ([`MediaSource::Original`]) and every image at or under 5 MB. Images over 5 MB are
/// downscaled to 2000px first, which only makes this estimate more conservative — the direction we
/// want, since being wrong low means ENOSPC halfway through.
///
/// Shares [`query_uploads`]' visibility filter via [`export_visibility_where`], so hidden/banned
/// uploads can't be counted here but skipped there (or the reverse, which under-reserves).
pub async fn estimate_export_bytes(pool: &PgPool, event_id: Uuid) -> Result<u64> {
let (bytes,): (i64,) = sqlx::query_as(concat!(
"SELECT COALESCE(SUM(u.original_size_bytes), 0)::bigint
FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
",
export_visibility_where!(),
))
.bind(event_id)
.fetch_one(pool)
.await
.context("estimating the export size")?;
Ok(bytes.max(0) as u64)
}
/// Headroom multiplier over the raw media sum: ZIP central directory, per-entry headers, the
/// embedded viewer, and the HTML export's temp staging.
const EXPORT_SIZE_OVERHEAD_PCT: u64 = 110;
/// Bytes this export must have available, given the raw media sum and how many jobs are competing.
///
/// `armed` is the multiplier that keeps two concurrent workers honest. `spawn_export_jobs` starts
/// the ZIP and the HTML halves at the same instant and both are gallery-sized, so a worker that
/// reserved only for itself would see "it fits", its sibling would independently see the same, and
/// together they would not fit — which is precisely the ENOSPC this preflight exists to prevent.
/// Computed in `u128` and clamped, NOT with `saturating_mul`: saturating first and then dividing by
/// 100 quietly turns an overflow into a number ~100x too small, which is the one direction that
/// matters here — an under-estimate authorises the very write the preflight exists to refuse.
fn required_free_bytes(media_bytes: u64, armed: i64) -> u64 {
let needed = media_bytes as u128 * EXPORT_SIZE_OVERHEAD_PCT as u128 / 100
* armed.max(1).min(i64::from(u32::MAX)) as u128;
needed.min(u64::MAX as u128) as u64
}
/// Free bytes a full keepsake build would need RIGHT NOW, both halves included.
///
/// The same arithmetic the preflight uses, exposed so the host dashboard can warn BEFORE the
/// release rather than reporting a failure after it. The preflight can only ever say "this didn't
/// fit"; at that point the gallery is full, the event is over, and the remedies (ask guests to stop
/// uploading, grow the volume) are all much harder. Hard-codes both halves because that is what a
/// release arms.
pub async fn keepsake_space_required(pool: &PgPool, event_id: Uuid) -> Result<u64> {
Ok(required_free_bytes(
estimate_export_bytes(pool, event_id).await?,
2,
))
}
/// Refuse to start an export that cannot fit, with a reason the host can act on.
///
/// Without this the failure mode is ENOSPC halfway through a multi-GB write, and the wreckage
/// outlives it: the epoch has already moved, so the job row is `failed` at the CURRENT generation
/// and `GET /export/zip` 404s, while the last good archive sits on disk unreferenced. The host's
/// only escape (`POST /host/export/rebuild`) needs the very space that isn't there. Failing at the
/// door instead leaves the disk untouched and puts a number in front of the operator.
///
/// Both halves are spawned concurrently and both are gallery-sized, so a worker must reserve for
/// its live sibling too — otherwise each independently sees "it fits", and together they don't.
async fn ensure_export_space(pool: &PgPool, event_id: Uuid, export_path: &Path) -> Result<()> {
let media_bytes = estimate_export_bytes(pool, event_id).await?;
// Every job armed at any epoch for this event that hasn't finished is competing for this disk.
let (armed,): (i64,) = sqlx::query_as(
"SELECT COUNT(*) FROM export_job
WHERE event_id = $1 AND status IN ('pending', 'running')",
)
.bind(event_id)
.fetch_one(pool)
.await
.context("counting armed export jobs")?;
let needed = required_free_bytes(media_bytes, armed);
// `None` = the mount couldn't be resolved. Fail OPEN, exactly as the upload quota does: refusing
// to build the keepsake because we can't read a number would be a worse failure than trying.
let Some(free) = crate::services::disk::free_bytes(export_path) else {
tracing::warn!("export preflight: disk snapshot unavailable; proceeding without the check");
return Ok(());
};
if free < needed {
let gb = |b: u64| b as f64 / 1_000_000_000.0;
tracing::error!(
needed,
free,
armed,
"export preflight: not enough free space to build the keepsake for event {event_id}"
);
anyhow::bail!(
"Nicht genug Speicherplatz für das Keepsake: benötigt ca. {:.1} GB, frei sind {:.1} GB. \
Bitte Speicher freigeben und das Keepsake anschließend neu erstellen.",
gb(needed),
gb(free)
);
}
Ok(())
}
/// Best-effort removal of stale per-generation export artifacts for one export type. Deletes
/// ONLY strictly-older generations (`n < keep_seq`) — never `keep_seq`'s own current file,
/// and never a NEWER generation that a concurrent re-release may already be producing (that
@@ -1167,20 +1406,7 @@ async fn prune_stale_export_files(
event_id: Uuid,
keep_seq: i64,
) {
// Files still referenced by a live (current-epoch) job row are OFF LIMITS regardless of the
// epoch in their name. A ViewerOnly regeneration carries the finished ZIP forward by re-stamping
// its row to the new epoch WITHOUT renaming the file — so `Gallery.<event>.<older>.zip` is still
// the served archive, and deleting it by filename-epoch would 404 the download.
let protected: Vec<String> = sqlx::query_scalar::<_, String>(
"SELECT split_part(j.file_path, '/', -1) FROM export_job j
JOIN event e ON e.id = j.event_id
WHERE j.event_id = $1 AND j.epoch = e.export_epoch
AND j.status = 'done' AND j.file_path IS NOT NULL",
)
.bind(event_id)
.fetch_all(pool)
.await
.unwrap_or_default();
let protected = protected_files(pool, event_id).await;
// EVERY shape is event-scoped. All events share one exports volume, so a name keyed only by
// generation would let event A's prune delete event B's live keepsake (and let two events
@@ -1326,6 +1552,36 @@ async fn maybe_broadcast_complete(
/// double-clicking `index.html` (file://), where browsers block a cross-origin
/// `fetch()` of a sibling `data.json` — so the data is inlined into the page.
/// (`data.json` is still written separately for the http-served case.)
/// Permissions stamped on every entry in both archives: `rw-r--r--`.
///
/// `ZipEntryBuilder::new` leaves the external file attribute at zero, and the host compatibility
/// defaults to Unix — so every entry was written with a stored mode of **0000**. Windows Explorer
/// ignores Unix modes and was fine, which is exactly why this survived: on Linux and macOS
/// `unzip` faithfully applies what the archive asks for, and the guest gets a directory of files
/// none of which they can open. `?---------` on every line of `unzip -Z`.
///
/// That is the keepsake — the artifact the whole event exists to produce — arriving unreadable,
/// after distribution, with no server-side symptom at all.
/// `S_IFREG | 0644`. The type bits are included because the mode is written whole into the high
/// half of the external file attribute: without them extractors see a file of type "unknown"
/// (`unzip -Z` renders `?rw-r--r--`), which works but is not what the archive means to say.
const KEEPSAKE_ENTRY_MODE: u16 = 0o100_644;
/// Build a ZIP entry for the keepsake. ALL entries in both archives go through here so the mode
/// can't be forgotten at one of the six call sites.
fn keepsake_entry(name: String, compression: Compression) -> ZipEntryBuilder {
ZipEntryBuilder::new(name.into(), compression).unix_permissions(KEEPSAKE_ENTRY_MODE)
}
/// Escape a JSON payload for inlining inside a `<script>` element.
///
/// See the call site in [`write_viewer_with_data`] for why this is every `<` and not just `</`.
/// Kept separate so the property that matters — no `<` survives, and the value still decodes to
/// the original — can be asserted without building a ZIP.
fn escape_json_for_script(data_json: &str) -> String {
data_json.replace('<', "\\u003c")
}
async fn write_viewer_with_data(
dir: &include_dir::Dir<'_>,
zip: &mut ZipFileWriter<tokio::fs::File>,
@@ -1337,8 +1593,31 @@ async fn write_viewer_with_data(
if path == "index.html" {
let html = std::str::from_utf8(file.contents())
.context("export-viewer index.html is not valid UTF-8")?;
// Escape `</` so a caption containing `</script>` can't break out of the tag.
let safe = data_json.replace("</", "<\\/");
// Escape EVERY `<`, not just `</`.
//
// `</` -> `<\/` stops the obvious break-out (`</script><img onerror=…>`) and is inert
// against XSS. It does not stop the caption steering the HTML TOKENIZER. A caption
// containing `<!--<script` with no later `-->` puts the parser into
// script-data-double-escaped state; from there the template's own `</script>` only
// steps back to script-data-escaped instead of closing the element, and the rest of the
// document — including the viewer bundle — is swallowed as script data. Nothing
// executes and nothing leaks; `window.__EXPORT_DATA__` is simply never assigned and the
// keepsake opens blank.
//
// That failure is silent and POST-DISTRIBUTION: the export succeeds, the ZIP is
// well-formed, the job writes `done`, /export/status is green, and the host hands out a
// file that only fails when a guest double-clicks index.html — in every copy, with no
// way to fix it after the fact. Reachable from any guest-authored caption or comment,
// since both are embedded in the viewer.
//
// `<` never appears in JSON structural syntax — only inside string values — so a global
// replace is sound, and `<` is valid in both JSON and a JS string literal. One
// rule covers `</script`, `<!--` and `<script` together, which is the point: the
// previous escape was named for the single case it did handle.
//
// NOTE this is deliberately only for the INLINED copy. `data.json` is written
// separately, in no HTML context, and must stay literal.
let safe = escape_json_for_script(data_json);
// Match the live app's colour theme: inject the same `:root:root{…}` override
// the app builds at runtime so a rose/sage/custom event exports a rose/sage
// keepsake (not the embedded default gold). The CSS is generated purely from
@@ -1352,13 +1631,13 @@ async fn write_viewer_with_data(
Some(idx) => format!("{}{}{}", &html[..idx], head_inject, &html[idx..]),
None => format!("{head_inject}{html}"),
};
let builder = ZipEntryBuilder::new(path.into(), Compression::Deflate);
let builder = keepsake_entry(path, Compression::Deflate);
let mut entry = zip.write_entry_stream(builder).await?;
let mut cursor = AllowStdIo::new(std::io::Cursor::new(injected.as_bytes()));
fcopy(&mut cursor, &mut entry).await?;
entry.close().await?;
} else {
let builder = ZipEntryBuilder::new(path.into(), Compression::Deflate);
let builder = keepsake_entry(path, Compression::Deflate);
let mut entry = zip.write_entry_stream(builder).await?;
let mut cursor = AllowStdIo::new(std::io::Cursor::new(file.contents()));
fcopy(&mut cursor, &mut entry).await?;
@@ -1494,3 +1773,187 @@ So geht's:\n\
Alles ist lokal auf deinem Gerät gespeichert.\n\
\n\
Viel Freude mit den Erinnerungen!\n";
// ── Tests ────────────────────────────────────────────────────────────────────
#[cfg(test)]
mod tests {
use super::*;
const EVT: &str = "11111111-1111-1111-1111-111111111111";
fn gallery_prefix() -> String {
format!("Gallery.{EVT}.")
}
#[test]
fn a_strictly_older_archive_is_reclaimed() {
// The whole point: at the start of a rebuild at epoch 5, generation 4's archive is dead
// weight — readiness is derived from `epoch = event.export_epoch`, so it is already
// unreachable — and its bytes are very often exactly the bytes the rebuild needs.
assert!(is_superseded_archive(
&format!("Gallery.{EVT}.4.zip"),
&gallery_prefix(),
5,
&[]
));
}
#[test]
fn our_own_and_newer_generations_are_never_touched() {
// `keep_seq` is OUR generation; a NEWER one belongs to a re-release that has already
// superseded us, and deleting it would let a lagging worker nuke a live keepsake.
for seq in [5, 6] {
assert!(
!is_superseded_archive(
&format!("Gallery.{EVT}.{seq}.zip"),
&gallery_prefix(),
5,
&[]
),
"generation {seq} must survive a prune keeping 5"
);
}
}
#[test]
fn a_carried_forward_archive_survives_despite_an_older_epoch_in_its_name() {
// THE dangerous case. A ViewerOnly regeneration (a moderated comment) re-stamps the
// finished ZIP's row to the new epoch WITHOUT renaming the file, so the SERVED archive
// legitimately carries an older generation number. Pruning it by filename-epoch would 404
// the photo download to change nothing in it. The protected set is what stops that, and
// moving the prune to build-start makes this case reachable far more often.
let carried = format!("Gallery.{EVT}.4.zip");
assert!(!is_superseded_archive(
&carried,
&gallery_prefix(),
5,
std::slice::from_ref(&carried)
));
}
#[test]
fn temps_and_staging_dirs_are_out_of_scope_for_the_early_prune() {
// A superseded worker may still be streaming into these, and at build START it is much
// more likely to be alive than at finalize time. Only inert FINAL archives are reclaimed
// here; `prune_stale_export_files` still handles the rest after we win.
for name in [
format!("Gallery.{EVT}.4.tmp"),
format!("viewer_tmp_{EVT}_4"),
format!("Memories.{EVT}.4.zip"),
] {
assert!(
!is_superseded_archive(&name, &gallery_prefix(), 5, &[]),
"{name} must not be reclaimed by the pre-build prune"
);
}
}
#[test]
fn another_events_archive_is_never_reclaimed() {
// All events share one exports volume, so the prefix carries the event id.
let other = "22222222-2222-2222-2222-222222222222";
assert!(!is_superseded_archive(
&format!("Gallery.{other}.4.zip"),
&gallery_prefix(),
5,
&[]
));
}
#[test]
fn unrelated_files_are_left_alone() {
for name in ["Gallery.zip", "notes.txt", "Gallery..4.zip"] {
assert!(!is_superseded_archive(name, &gallery_prefix(), 5, &[]));
}
}
/// Every `<` is escaped, whatever it is part of.
///
/// PREVENTS the regression to `</` -> `<\\/`, which is named for the one case it handles.
/// `<!--<script` with no later `-->` drives the HTML tokenizer into
/// script-data-double-escaped state, where the template's own `</script>` no longer closes
/// the element — the viewer bundle is swallowed as script data, `__EXPORT_DATA__` is never
/// assigned, and the keepsake opens blank in every copy the host has already handed out.
#[test]
fn no_left_angle_bracket_survives_inlining() {
for payload in [
r#"{"caption":"<!--<script"}"#,
r#"{"caption":"</script><img src=x onerror=alert(1)>"}"#,
r#"{"caption":"<!--"}"#,
r#"{"caption":"<script>"}"#,
r#"{"caption":"a < b"}"#,
] {
let escaped = escape_json_for_script(payload);
assert!(
!escaped.contains('<'),
"a surviving `<` can still steer the tokenizer: {escaped}"
);
}
}
/// The escape must not change what the viewer READS — it is a transport encoding, not a
/// sanitiser. A caption is guest-authored text that has to render back exactly.
#[test]
fn the_payload_still_decodes_to_the_original_value() {
// `<` appears only inside JSON string values, never in structural syntax, so a global
// replace is sound — this is the assertion that says so.
for caption in [
"<!--<script",
"</script><img src=x onerror=alert(1)>",
"a < b und c > d",
"ganz normale Bildunterschrift",
"Herz <3",
] {
let json = serde_json::json!({ "posts": [{ "caption": caption }] }).to_string();
let escaped = escape_json_for_script(&json);
let back: serde_json::Value =
serde_json::from_str(&escaped).expect("the escaped form must still be valid JSON");
assert_eq!(
back["posts"][0]["caption"].as_str(),
Some(caption),
"the caption must survive the round trip unchanged"
);
}
}
/// Nothing else in the document is touched.
#[test]
fn a_payload_with_no_angle_brackets_is_unchanged() {
let json = r#"{"posts":[{"caption":"schönes Foto"}]}"#;
assert_eq!(escape_json_for_script(json), json);
}
#[test]
fn a_lone_armed_job_reserves_for_one_archive() {
// A ViewerOnly regeneration re-arms only the HTML half — reserving for two would refuse
// a rebuild that fits perfectly well.
assert_eq!(required_free_bytes(1_000, 1), 1_100);
}
#[test]
fn two_concurrent_halves_reserve_for_both() {
// The bug this exists to prevent: each worker independently sees "it fits", and together
// they don't. Both halves are gallery-sized, so the reservation must be for the pair.
assert_eq!(required_free_bytes(1_000, 2), 2_200);
}
#[test]
fn a_zero_count_still_reserves_for_one() {
// Defensive: a racing status transition must never yield a zero requirement, which would
// wave through an export of any size onto a full disk.
assert_eq!(required_free_bytes(1_000, 0), 1_100);
}
#[test]
fn an_empty_gallery_needs_nothing() {
assert_eq!(required_free_bytes(0, 2), 0);
}
#[test]
fn a_pathological_size_saturates_instead_of_wrapping() {
// u64 overflow would wrap to a TINY requirement and authorise the exact write we're
// guarding against — the failure mode must be "refuse", never "wrap and allow".
assert_eq!(required_free_bytes(u64::MAX, 2), u64::MAX);
}
}

View File

@@ -0,0 +1,263 @@
//! Shared image decoding.
//!
//! Exists so there is exactly ONE way to turn a file on disk into a `DynamicImage` in this
//! codebase. Two properties have to hold everywhere an image is decoded, and both were
//! previously re-derived per call site — which is how they drifted apart:
//!
//! - **EXIF orientation must be applied.** Phones do not rotate sensor data; they record how
//! the camera was held in a tag and store the pixels as shot. `image::open` and
//! `ImageReader::decode` both hand back the raw pixels and ignore that tag, and re-encoding
//! to JPEG writes no EXIF, so the derivative is permanently sideways while the untouched
//! original still renders upright. The compression worker was fixed; the export worker was
//! not, so every portrait photo came out sideways in the keepsake's HTML viewer.
//! - **Decode limits must be set.** The upload body cap bounds the file on disk, but a small
//! file can decode to enormous dimensions (a ~1 MB image expanding to 50k×50k px), OOM-ing
//! the box. `image::open` applies NO limits at all, so the export path was also decoding
//! arbitrary user-supplied images unbounded.
use anyhow::{Context, Result};
use image::{DynamicImage, ImageDecoder};
use std::path::Path;
/// Bounds for any decode of user-supplied image data. The per-axis cap covers any real phone
/// photo; `max_alloc` bounds the decoded buffer — but only because `decode_oriented` reserves
/// against it explicitly, see there.
///
/// Sized against the deployment: the app container is capped at 1 GiB and the compression
/// worker runs `compression_concurrency` decodes at once (default 2), so 256 MiB per decode
/// leaves headroom for the resize buffers and the runtime.
fn decode_limits() -> image::Limits {
let mut limits = image::Limits::default();
limits.max_image_width = Some(12_000);
limits.max_image_height = Some(12_000);
limits.max_alloc = Some(256 * 1024 * 1024);
limits
}
/// True when re-running the exact same work on the exact same bytes cannot possibly
/// succeed, so retrying only burns wall-clock and log noise.
///
/// Deliberately narrow. Only the `ImageError` variants that are a property of the *input*
/// count: the file will not shrink, gain codec support, or un-corrupt itself between
/// attempts. `IoError` is excluded on purpose — an ENOSPC while writing a derivative, or
/// EMFILE under load, is exactly the transient case the retry exists for.
pub fn is_permanent_image_error(err: &anyhow::Error) -> bool {
err.chain().any(|cause| {
matches!(
cause.downcast_ref::<image::ImageError>(),
Some(
image::ImageError::Limits(_)
| image::ImageError::Unsupported(_)
| image::ImageError::Decoding(_)
)
)
})
}
/// Build a decoder for `path` with the budget enforced, WITHOUT reading any pixels.
///
/// Single source of truth for "may this image be decoded at all": both the upload
/// admission check and the compression worker go through here, so they cannot disagree
/// about what is acceptable.
fn decoder_within_budget(path: &Path) -> Result<impl image::ImageDecoder> {
let mut reader = image::ImageReader::open(path)
.context("failed to open image")?
.with_guessed_format()
.context("failed to read image header")?;
let mut limits = decode_limits();
reader.limits(limits.clone());
// We need `into_decoder` rather than `decode()` to read the EXIF orientation tag before
// the pixels are consumed. But the two are NOT equivalent on safety: `decode()` performs
//
// limits.reserve(decoder.total_bytes())?;
//
// between building the decoder and reading the image, and `into_decoder()` skips it (the
// crate's own FIXME concedes `from_decoder` doesn't compensate). Nothing else enforces
// `max_alloc` — the JPEG decoder's `set_limits` only checks support and dimensions — so
// without the line below the budget is inert and the ONLY bound is the per-axis cap. That
// leaves 12000x12000 decodable at 412 MiB, and two concurrent at 824 MiB against a 1 GiB
// container. Re-add it, exactly as `decode()` does.
let mut decoder = reader.into_decoder().context("failed to decode image")?;
limits
.reserve(decoder.total_bytes())
.context("image too large to decode within the memory budget")?;
decoder
.set_limits(limits)
.context("image too large to decode within the memory budget")?;
Ok(decoder)
}
/// Megapixels an image would decode to, or `None` if its header can't be read. Used only
/// to put a concrete number in the message the guest sees.
pub fn megapixels(path: &Path) -> Option<f64> {
let reader = image::ImageReader::open(path)
.ok()?
.with_guessed_format()
.ok()?;
let (w, h) = reader.into_dimensions().ok()?;
Some(f64::from(w) * f64::from(h) / 1_000_000.0)
}
/// True when an image cannot be decoded specifically because it would exceed the memory
/// budget — read from the header, no pixels touched.
///
/// Called at upload admission so a guest who sends a 100 MP photo is told at the door, with
/// a reason they can act on, instead of the upload being accepted with a 201 and then
/// silently soft-deleted minutes later when the worker gives up on it.
///
/// Deliberately narrow: ONLY the budget. A corrupt, truncated or unsupported file also
/// fails to build a decoder, but rejecting those here would change a contract the
/// adversarial suite pins on purpose — acceptance follows the magic bytes, and a payload
/// with a valid JPEG header is accepted regardless of what follows it. Those go to the
/// compression worker as before, which handles them gracefully and (since the retry
/// classifier) no longer burns backoff on them.
pub fn exceeds_decode_budget(path: &Path) -> bool {
match decoder_within_budget(path) {
Ok(_) => false,
Err(e) => e.chain().any(|cause| {
matches!(
cause.downcast_ref::<image::ImageError>(),
Some(image::ImageError::Limits(_))
)
}),
}
}
/// Decode an image from disk with decompression-bomb limits applied and its EXIF
/// orientation baked into the pixels.
///
/// Blocking — call inside `spawn_blocking`.
pub fn decode_oriented(path: &Path) -> Result<DynamicImage> {
let mut decoder = decoder_within_budget(path)?;
// Cheap, and it happens BEFORE any pixels are read: an oversized image costs a header
// parse, not an allocation.
let orientation = decoder
.orientation()
.unwrap_or(image::metadata::Orientation::NoTransforms);
let mut img = DynamicImage::from_decoder(decoder).context("failed to decode image")?;
img.apply_orientation(orientation);
Ok(img)
}
#[cfg(test)]
mod tests {
use super::*;
/// Shared with the e2e suite rather than duplicating 568 KiB of binary: the same file
/// drives `02-upload/oversized-image` so both layers assert on one artefact.
const HUGE: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/huge-99mp.jpg"
);
#[test]
fn rejects_an_image_that_would_blow_the_allocation_budget() {
// 11000x9000 = 99 MP. Deliberately UNDER the 12000px per-axis cap, so the axis check
// cannot reject it — the allocation budget is the only thing that can, which is
// exactly what makes this a regression test rather than a restatement of the axis cap.
// 283 MiB decoded as RGB8 against a 256 MiB budget, from 568 KiB on disk.
//
// This failed before the guard was restored: `ImageReader::decode` performs
// `limits.reserve(decoder.total_bytes())`, and `into_decoder()` — which we need for
// the EXIF tag — skips it, so `max_alloc` was inert and this decoded happily.
// Map the Ok arm to its dimensions first: on failure `expect_err` Debug-prints the
// value, and Debug on a DynamicImage dumps every pixel — 283 MiB of output.
let err = decode_oriented(Path::new(HUGE))
.map(|img| (img.width(), img.height()))
.expect_err("a 99 MP image must be refused, not allocated");
let msg = format!("{err:#}");
assert!(
msg.to_lowercase().contains("limit") || msg.to_lowercase().contains("memory"),
"expected a limits error, got: {msg}"
);
}
#[test]
fn an_oversized_image_is_a_permanent_failure() {
// The retry loop must not burn 2s + 4s of backoff on this: the file will not shrink
// between attempts, so all three attempts reach the identical conclusion.
let err = decode_oriented(Path::new(HUGE))
.map(|img| (img.width(), img.height()))
.expect_err("fixture must exceed the budget");
assert!(
is_permanent_image_error(&err),
"a Limits error can never succeed on retry: {err:#}"
);
}
#[test]
fn a_plain_io_error_is_not_permanent() {
// The mirror that keeps the classifier honest. ENOSPC while writing a derivative, or
// EMFILE under load, is exactly what the retry exists for — misclassifying those as
// permanent would turn a transient blip back into the data loss round 1 fixed.
let err = decode_oriented(Path::new("/nonexistent/definitely-not-here.jpg"))
.map(|img| (img.width(), img.height()))
.expect_err("a missing file must error");
assert!(
!is_permanent_image_error(&err),
"an IO error must stay retryable: {err:#}"
);
}
#[test]
fn admission_rejects_only_the_over_budget_case() {
// Admission and processing must agree about SIZE — a photo accepted at the door and
// then rejected by the worker for being too big is the failure this pair prevents.
assert!(
exceeds_decode_budget(Path::new(HUGE)),
"admission must reject what the decoder rejects for size"
);
let ordinary = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/portrait-exif6.jpg"
);
assert!(
!exceeds_decode_budget(Path::new(ordinary)),
"admission must accept an ordinary photo"
);
}
#[test]
fn admission_does_not_reject_a_merely_undecodable_file() {
// The narrowing that keeps the adversarial contract intact: a payload with valid
// JPEG magic bytes and nothing behind them cannot be decoded, but acceptance follows
// the magic bytes by design (07-adversarial/file-upload-attacks). It is the worker's
// job to fail it, not admission's — admission is only the resource guard.
let dir = std::env::temp_dir().join("eventsnap-imaging-test");
std::fs::create_dir_all(&dir).expect("tmp dir");
let stub = dir.join("magic-only.jpg");
let mut bytes = vec![0u8; 1024];
bytes[..3].copy_from_slice(&[0xFF, 0xD8, 0xFF]);
std::fs::write(&stub, &bytes).expect("write stub");
assert!(
!exceeds_decode_budget(&stub),
"a corrupt file is not an over-budget file"
);
assert!(
decode_oriented(&stub)
.map(|i| (i.width(), i.height()))
.is_err(),
"...but it must still fail in the worker"
);
let _ = std::fs::remove_file(&stub);
}
#[test]
fn still_decodes_an_ordinary_photo_and_applies_orientation() {
// The guard must not have become a blanket refusal. This fixture is 40x20 stored with
// EXIF Orientation=6, so a correct decode returns it rotated to 20x40 portrait.
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/portrait-exif6.jpg"
);
let img = decode_oriented(Path::new(path)).expect("an ordinary photo must decode");
assert_eq!(
(img.width(), img.height()),
(20, 40),
"EXIF orientation must still be applied after restoring the guard"
);
}
}

View File

@@ -9,10 +9,13 @@
//! users staring at a spinner. Resetting them on startup recovers gracefully.
//!
//! 2. **Periodic tasks** — pruning that should happen "every hour" rather than per
//! request: expired sessions (otherwise the table grows unboundedly), and the
//! request: expired sessions (otherwise the table grows unboundedly), the
//! rate-limiter's in-memory windows (so keys for IPs that left long ago don't
//! accumulate).
//! accumulate), and the media of soft-deleted uploads — both the ones whose compression
//! permanently failed and the ones a guest or host deliberately removed — which are
//! retained for a recovery window and then reclaimed.
use std::path::PathBuf;
use std::time::Duration;
use sqlx::PgPool;
@@ -20,6 +23,42 @@ use sqlx::PgPool;
use crate::services::rate_limiter::RateLimiter;
use crate::services::sse_tickets::SseTicketStore;
/// How long a permanently-failed upload's original is kept on disk before it is
/// reclaimed.
///
/// The compression worker stops deleting originals on failure — a transient error must
/// never destroy the guest's only copy of a photo they can't retake. But the row is
/// soft-deleted and the uploader's quota IS refunded, so without a sweep those bytes are
/// invisible, unowned, and free: a reproducible codec failure lets one guest accumulate
/// orphans at no personal cost, and because `active_uploaders` counts only users with
/// non-deleted uploads, dropping out of that count actually RAISES everyone's per-user
/// ceiling while the disk gets fuller.
///
/// Two weeks is comfortably longer than any single event, so an operator investigating a
/// failed upload still has the file, while the leak stays bounded.
const FAILED_ORIGINAL_RETENTION_DAYS: i64 = 14;
/// How long a DELIBERATELY deleted upload's files are kept before they are reclaimed.
///
/// The same leak, reached by the ordinary path rather than the exceptional one.
/// `soft_delete_in_event` stamps `deleted_at` and refunds `total_upload_bytes`, but nothing ever
/// removed the bytes — so the quota stopped bounding the disk. Upload 500 MB, delete, quota is back
/// to zero, upload another 500 MB: not an attack, just a guest curating their camera roll, which is
/// what people do. The host then sees guests hitting "Du hast dein Upload-Limit erreicht" while the
/// admin widget shows a disk full of files no upload row points at, and the quota message is
/// actively misleading because the space really is gone — just not to anyone the accounting can
/// name.
///
/// Much shorter than the failure window on purpose. Fourteen days outlives the whole event, so a
/// deliberate delete would never reclaim anything while it mattered. A day still gives an operator
/// a recovery window for a mis-tap.
///
/// NOTE what this does NOT do: within the window the bytes are still spent and still unaccounted,
/// so a guest deleting and re-uploading through an eight-hour event can outrun the sweep. Bounding
/// that would mean holding the quota until the file is actually reclaimed rather than refunding at
/// `deleted_at` — a deliberate trade, and the reason the low-disk warning exists.
const DELETED_UPLOAD_RETENTION_HOURS: i64 = 24;
/// Reset rows left in flight by a previous crashed instance. Run once on startup,
/// before the HTTP server starts taking requests, so users never observe the
/// half-state.
@@ -87,7 +126,12 @@ pub async fn startup_recovery(pool: &PgPool) {
/// - drops expired SSE tickets (30s TTL but the map keeps the slot until pruned)
///
/// Cadence is 1h — fine for both jobs at our scale.
pub fn spawn_periodic_tasks(pool: PgPool, rate_limiter: RateLimiter, sse_tickets: SseTicketStore) {
pub fn spawn_periodic_tasks(
pool: PgPool,
rate_limiter: RateLimiter,
sse_tickets: SseTicketStore,
media_path: PathBuf,
) {
tokio::spawn(async move {
let mut tick = tokio::time::interval(Duration::from_secs(3600));
// Fire the first tick immediately, then hourly.
@@ -95,12 +139,117 @@ pub fn spawn_periodic_tasks(pool: PgPool, rate_limiter: RateLimiter, sse_tickets
loop {
tick.tick().await;
cleanup_sessions(&pool).await;
cleanup_deleted_media(&pool, &media_path).await;
rate_limiter.prune();
sse_tickets.prune();
}
});
}
/// Reclaim the media of soft-deleted uploads once they are past their retention window.
///
/// ONLY ever touches rows with `deleted_at IS NOT NULL`, so it can never reach a live upload. Two
/// classes, two windows, because the two deletes mean different things:
///
/// - a compression failure the guest didn't ask for and may want investigated —
/// [`FAILED_ORIGINAL_RETENTION_DAYS`];
/// - a deliberate removal by the guest or the host — [`DELETED_UPLOAD_RETENTION_HOURS`].
///
/// ALL FOUR paths are reclaimed, not just the original. The previous version cleared
/// `original_path` alone, which was right for its only case (a failed compression produces no
/// derivatives) but wrong the moment the sweep reaches a successfully processed upload: preview,
/// display and thumbnail are each a separate file on disk, none of them counted in
/// `original_size_bytes`, and nothing else ever removed them.
///
/// Every column is cleared in the same pass, which makes the sweep idempotent and stops a later run
/// re-reporting files that are already gone. The ROW is kept: it is the audit trail, it costs a few
/// hundred bytes, and `backfill_stale_derivatives` is guarded on `deleted_at IS NULL` so a nulled
/// `preview_path` can never make it regenerate what was just reclaimed.
async fn cleanup_deleted_media(pool: &PgPool, media_path: &std::path::Path) {
type Row = (
uuid::Uuid,
String,
Option<String>,
Option<String>,
Option<String>,
);
let rows = sqlx::query_as::<_, Row>(
"SELECT id, original_path, preview_path, display_path, thumbnail_path FROM upload
WHERE deleted_at IS NOT NULL
AND CASE WHEN compression_status = 'failed'
THEN deleted_at < NOW() - ($1 || ' days')::interval
ELSE deleted_at < NOW() - ($2 || ' hours')::interval
END
AND (original_path <> '' OR preview_path IS NOT NULL
OR display_path IS NOT NULL OR thumbnail_path IS NOT NULL)",
)
.bind(FAILED_ORIGINAL_RETENTION_DAYS.to_string())
.bind(DELETED_UPLOAD_RETENTION_HOURS.to_string())
.fetch_all(pool)
.await;
let rows = match rows {
Ok(r) => r,
Err(e) => {
tracing::warn!(error = ?e, "deleted-media sweep query failed");
return;
}
};
if rows.is_empty() {
return;
}
let mut reclaimed = 0usize;
for (id, original, preview, display, thumbnail) in rows {
let paths: Vec<String> = std::iter::once(original)
.filter(|p| !p.is_empty())
.chain([preview, display, thumbnail].into_iter().flatten())
.collect();
// All-or-nothing per row: the columns are only cleared once every file for that upload is
// gone. Clearing after a partial success would strand the survivors with nothing pointing
// at them — the same unowned-bytes state this sweep exists to drain.
let mut all_gone = true;
for rel in &paths {
let absolute = media_path.join(rel);
match tokio::fs::remove_file(&absolute).await {
Ok(()) => reclaimed += 1,
// Already gone (manual cleanup, restored backup) — still counts as reclaimed for
// the purpose of clearing the columns, or the row is re-selected every hour forever.
Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
tracing::warn!(error = ?e, %id, path = %absolute.display(),
"could not reclaim deleted media; leaving the row for the next sweep");
all_gone = false;
}
}
}
if !all_gone {
continue;
}
if let Err(e) = sqlx::query(
"UPDATE upload SET original_path = '', preview_path = NULL,
display_path = NULL, thumbnail_path = NULL
WHERE id = $1",
)
.bind(id)
.execute(pool)
.await
{
tracing::warn!(error = ?e, %id, "reclaimed the files but could not clear the paths");
}
}
if reclaimed > 0 {
tracing::info!(
"reclaimed {reclaimed} file(s) from soft-deleted uploads (deliberate deletes after \
{DELETED_UPLOAD_RETENTION_HOURS}h, compression failures after \
{FAILED_ORIGINAL_RETENTION_DAYS}d)"
);
}
}
async fn cleanup_sessions(pool: &PgPool) {
match sqlx::query("DELETE FROM session WHERE expires_at < NOW() - INTERVAL '1 day'")
.execute(pool)

View File

@@ -2,6 +2,7 @@ pub mod compression;
pub mod config;
pub mod disk;
pub mod export;
pub mod imaging;
pub mod maintenance;
pub mod rate_limiter;
pub mod sse_tickets;

View File

@@ -17,13 +17,14 @@ impl RateLimiter {
}
}
/// Returns `true` if the request is allowed, `false` if rate-limited.
pub fn check(&self, key: impl Into<String>, max: usize, window: Duration) -> bool {
self.check_with_retry(key, max, window).is_ok()
}
/// Returns `Ok(())` if allowed, `Err(retry_after_secs)` if rate-limited.
/// `retry_after_secs` is how long until the oldest slot in the window expires.
///
/// This is deliberately the ONLY entry point. There used to be a `check()` wrapper
/// returning a plain bool, and 7 of the 8 call sites used it and then hard-coded
/// `None` for the response's `Retry-After` — so a throttled client was told to back
/// off but never for how long. Forcing every caller through the `Result` makes the
/// retry delay impossible to discard by accident.
pub fn check_with_retry(
&self,
key: impl Into<String>,
@@ -84,6 +85,11 @@ impl RateLimiter {
/// appends is the real client. A client can prepend arbitrary spoofed values to
/// the left of XFF to dodge throttles — those are ignored here. This assumes
/// exactly one trusted proxy (Caddy); revisit if that changes.
///
/// Pass the peer address as `fallback`, never a constant. Every caller used to pass
/// the literal `"unknown"`, so any request that arrived without XFF — i.e. anything
/// reaching the app directly rather than through Caddy — shared ONE bucket with every
/// other such request, turning the limiter into a self-inflicted global throttle.
pub fn client_ip(headers: &axum::http::HeaderMap, fallback: &str) -> String {
headers
.get("x-forwarded-for")
@@ -104,29 +110,35 @@ mod tests {
#[test]
fn allows_up_to_max_then_blocks() {
let rl = RateLimiter::new();
assert!(rl.check("k", 3, MIN));
assert!(rl.check("k", 3, MIN));
assert!(rl.check("k", 3, MIN));
assert!(!rl.check("k", 3, MIN), "the 4th request must be blocked");
assert!(rl.check_with_retry("k", 3, MIN).is_ok());
assert!(rl.check_with_retry("k", 3, MIN).is_ok());
assert!(rl.check_with_retry("k", 3, MIN).is_ok());
assert!(
rl.check_with_retry("k", 3, MIN).is_err(),
"the 4th request must be blocked"
);
}
#[test]
fn keys_are_independent() {
let rl = RateLimiter::new();
assert!(rl.check("a", 1, MIN));
assert!(!rl.check("a", 1, MIN));
assert!(rl.check("b", 1, MIN), "a different key has its own window");
assert!(rl.check_with_retry("a", 1, MIN).is_ok());
assert!(rl.check_with_retry("a", 1, MIN).is_err());
assert!(
rl.check_with_retry("b", 1, MIN).is_ok(),
"a different key has its own window"
);
}
#[test]
fn window_slides_and_allows_again_after_expiry() {
let rl = RateLimiter::new();
let w = Duration::from_millis(40);
assert!(rl.check("k", 1, w));
assert!(!rl.check("k", 1, w));
assert!(rl.check_with_retry("k", 1, w).is_ok());
assert!(rl.check_with_retry("k", 1, w).is_err());
std::thread::sleep(Duration::from_millis(55));
assert!(
rl.check("k", 1, w),
rl.check_with_retry("k", 1, w).is_ok(),
"the slot should expire once the window passes"
);
}
@@ -191,10 +203,13 @@ mod tests {
#[test]
fn clear_resets_every_window() {
let rl = RateLimiter::new();
assert!(rl.check("k", 1, MIN));
assert!(!rl.check("k", 1, MIN));
assert!(rl.check_with_retry("k", 1, MIN).is_ok());
assert!(rl.check_with_retry("k", 1, MIN).is_err());
rl.clear();
assert!(rl.check("k", 1, MIN), "clear() must free the window");
assert!(
rl.check_with_retry("k", 1, MIN).is_ok(),
"clear() must free the window"
);
}
/// `prune()` is a memory-leak guard: without it a long-lived process keeps one HashMap
@@ -216,7 +231,7 @@ mod tests {
.insert("stale".to_string(), vec![ancient]);
// ...alongside a key that is still inside its window.
assert!(rl.check("live", 5, MIN));
assert!(rl.check_with_retry("live", 5, MIN).is_ok());
assert_eq!(rl.windows.lock().unwrap().len(), 2);
rl.prune();
@@ -239,13 +254,13 @@ mod tests {
// prune() dropped live keys, every background sweep would hand attackers a fresh
// budget.
let rl = RateLimiter::new();
assert!(rl.check("k", 1, MIN));
assert!(!rl.check("k", 1, MIN));
assert!(rl.check_with_retry("k", 1, MIN).is_ok());
assert!(rl.check_with_retry("k", 1, MIN).is_err());
rl.prune();
assert!(
!rl.check("k", 1, MIN),
rl.check_with_retry("k", 1, MIN).is_err(),
"prune() must not clear a window that is still active"
);
}

View File

@@ -255,3 +255,91 @@ pub async fn downloadable(pool: &PgPool, event_id: Uuid, export_type: &str) -> O
.expect("downloadable")
.flatten()
}
/// Insert an upload of `size` bytes, optionally already soft-deleted.
pub async fn seed_upload(
pool: &PgPool,
event_id: Uuid,
user_id: Uuid,
size: i64,
deleted: bool,
) -> Uuid {
sqlx::query_scalar(
"INSERT INTO upload (event_id, user_id, original_path, mime_type,
original_size_bytes, deleted_at)
VALUES ($1, $2, 'originals/x.jpg', 'image/jpeg', $3,
CASE WHEN $4 THEN NOW() ELSE NULL END)
RETURNING id",
)
.bind(event_id)
.bind(user_id)
.bind(size)
.bind(deleted)
.fetch_one(pool)
.await
.expect("seed upload")
}
/// Flip the moderation flags a ban sets.
pub async fn set_user_moderation(pool: &PgPool, user_id: Uuid, banned: bool, hidden: bool) {
sqlx::query("UPDATE \"user\" SET is_banned = $2, uploads_hidden = $3 WHERE id = $1")
.bind(user_id)
.bind(banned)
.bind(hidden)
.execute(pool)
.await
.expect("set moderation");
}
/// SRC: `services/export.rs::query_uploads` — the visibility filter, verbatim, projected down to
/// `(id, original_size_bytes)`. This is the row set that ACTUALLY lands in the archives.
///
/// Production builds this WHERE from `export_visibility_where!()`, shared with
/// `estimate_export_bytes`. A copy here can pin the behaviour but CANNOT detect production moving
/// away from it — that is what sharing the fragment is for, not this.
pub async fn export_visible_uploads(pool: &PgPool, event_id: Uuid) -> Vec<(Uuid, i64)> {
sqlx::query_as(
"SELECT u.id, u.original_size_bytes
FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE
GROUP BY u.id, usr.display_name
ORDER BY u.created_at ASC",
)
.bind(event_id)
.fetch_all(pool)
.await
.expect("export_visible_uploads")
}
/// SRC: `services/export.rs::estimate_export_bytes` — verbatim. Same caveat as above: production
/// shares its WHERE with `query_uploads` via `export_visibility_where!()`, so these two copies
/// agreeing proves the behaviour, not the absence of drift.
pub async fn estimate_export_bytes(pool: &PgPool, event_id: Uuid) -> i64 {
let (bytes,): (i64,) = sqlx::query_as(
"SELECT COALESCE(SUM(u.original_size_bytes), 0)::bigint
FROM upload u
JOIN \"user\" usr ON usr.id = u.user_id
WHERE u.event_id = $1 AND u.deleted_at IS NULL
AND usr.uploads_hidden = FALSE AND usr.is_banned = FALSE",
)
.bind(event_id)
.fetch_one(pool)
.await
.expect("estimate_export_bytes");
bytes
}
/// SRC: `services/export.rs::ensure_export_space` — the armed-job count, verbatim.
pub async fn armed_job_count(pool: &PgPool, event_id: Uuid) -> i64 {
let (n,): (i64,) = sqlx::query_as(
"SELECT COUNT(*) FROM export_job
WHERE event_id = $1 AND status IN ('pending', 'running')",
)
.bind(event_id)
.fetch_one(pool)
.await
.expect("armed_job_count");
n
}

View File

@@ -0,0 +1,163 @@
//! DB-backed tests for the export disk preflight.
//!
//! The keepsake used to be built with NO free-space check at all, and the failure that produced was
//! not "the export failed" but "the deliverable is stuck and the escape hatch needs the space that
//! isn't there":
//!
//! 1. A takedown bumps the epoch and re-arms both halves.
//! 2. The ZIP hits ENOSPC partway through a multi-GB write.
//! 3. The job row is now `failed` at the CURRENT epoch, so readiness
//! (`epoch = event.export_epoch AND status = 'done'`) is false and `GET /export/zip` 404s —
//! while the last good archive sits on disk, unreferenced and unreachable.
//! 4. `POST /host/export/rebuild` re-arms the same doomed write.
//!
//! Two changes close it: reclaim the superseded generation BEFORE building (so peak usage is one
//! generation, not two) and refuse up front with a number the host can act on.
//!
//! What these tests pin is the ESTIMATE — the part that decides. The arithmetic on top of it lives
//! in `services/export.rs`'s unit tests; the filesystem selection lives in `is_superseded_archive`.
//!
//! ON DRIFT, precisely, because it is easy to overclaim here. The hazard is that `query_uploads`
//! (which selects the rows the archives are built from) and `estimate_export_bytes` (which sizes
//! them) could disagree — and an estimate missing rows the archive writes UNDER-reserves, the one
//! direction that reintroduces the ENOSPC. **These tests cannot catch that**, and neither can any
//! test in this harness: both sides here are `SRC:`-marked hand-copies in `tests/common/mod.rs`,
//! so if production moved and the copies didn't, they would sit still and keep passing.
//!
//! That is fixed where it can be — the two queries now share one `export_visibility_where!()`
//! fragment in `services/export.rs`, so they cannot diverge by construction. What is left for
//! these tests is what the convention is genuinely good at: pinning the BEHAVIOUR, so a change
//! that deliberately alters the filter has to come here and say so.
mod common;
use common::*;
use sqlx::PgPool;
/// The estimate must equal the sum over EXACTLY the rows `query_uploads` returns — computed from
/// that row set, not from a restatement of its WHERE clause.
///
/// PINS: which uploads the preflight is allowed to count. Each excluded row below is excluded by a
/// DIFFERENT predicate, so a change that drops or weakens any one of them fails here and has to be
/// argued for. (It does not detect production drifting away from these copies — see the file
/// header; `export_visibility_where!()` is what makes that impossible.)
#[sqlx::test]
async fn the_estimate_sums_exactly_the_rows_the_archive_will_contain(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let visible = seed_user(&pool, event_id, "Anna").await;
let banned = seed_user(&pool, event_id, "Ben").await;
let hidden = seed_user(&pool, event_id, "Cara").await;
seed_upload(&pool, event_id, visible, 1_000, false).await;
seed_upload(&pool, event_id, visible, 2_500, false).await;
// Each of these is excluded from the archive by a DIFFERENT predicate.
seed_upload(&pool, event_id, visible, 9_000, true).await; // soft-deleted
seed_upload(&pool, event_id, banned, 9_000, false).await; // uploader banned
seed_upload(&pool, event_id, hidden, 9_000, false).await; // uploads hidden
set_user_moderation(&pool, banned, true, true).await;
set_user_moderation(&pool, hidden, false, true).await;
let rows = export_visible_uploads(&pool, event_id).await;
let expected: i64 = rows.iter().map(|(_, bytes)| bytes).sum();
assert_eq!(rows.len(), 2, "only Anna's two live uploads are archived");
assert_eq!(
estimate_export_bytes(&pool, event_id).await,
expected,
"the preflight must size the gallery the export will actually write"
);
assert_eq!(expected, 3_500);
}
/// An event with nothing to archive estimates zero rather than NULL.
///
/// PREVENTS: `SUM()` over no rows returning NULL and the decode blowing up — which would abort the
/// export with a type error instead of building an (entirely legitimate) empty keepsake.
#[sqlx::test]
async fn an_empty_gallery_estimates_zero_not_null(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
assert_eq!(estimate_export_bytes(&pool, event_id).await, 0);
// And with a user who has uploaded nothing.
seed_user(&pool, event_id, "Anna").await;
assert_eq!(estimate_export_bytes(&pool, event_id).await, 0);
}
/// A release arms both halves, so the preflight sees a count of 2 and reserves for the pair.
///
/// PREVENTS: the concurrency under-reservation. `spawn_export_jobs` starts the ZIP and HTML workers
/// at the same instant, and BOTH are gallery-sized (`Memories.zip` streams the original for every
/// video and every image at or under 5 MB, all `Compression::Stored`). A worker reserving only for
/// itself would see "it fits", its sibling would independently see the same, and together they
/// would ENOSPC — which is why `required_free_bytes` multiplies by this count.
#[sqlx::test]
async fn a_release_arms_both_halves_so_the_preflight_reserves_for_two(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user = seed_user(&pool, event_id, "Anna").await;
seed_upload(&pool, event_id, user, 1_000, false).await;
assert_eq!(
armed_job_count(&pool, event_id).await,
0,
"nothing is armed before the release"
);
let epoch = release_gallery(&pool, "wedding").await.expect("released");
assert_eq!(
armed_job_count(&pool, event_id).await,
2,
"a release arms zip AND html — both compete for the same disk"
);
// A worker that has claimed its half is still competing; `running` must keep counting.
assert!(claim_job(&pool, event_id, "zip", epoch).await);
assert_eq!(
armed_job_count(&pool, event_id).await,
2,
"claiming moves pending -> running, which must not drop out of the reservation"
);
// Only a FINISHED half stops competing.
assert!(finalize_job(&pool, event_id, "zip", epoch, "exports/Gallery.zip").await);
assert_eq!(
armed_job_count(&pool, event_id).await,
1,
"a done half no longer needs space reserved for it"
);
}
/// A ViewerOnly regeneration re-arms only the HTML half, so the preflight reserves for one.
///
/// PREVENTS: over-reservation refusing a rebuild that fits perfectly well. Moderating a comment
/// carries the finished ZIP forward untouched; demanding room for a second copy of it would fail
/// the one operation that needs no new gallery-sized write at all.
#[sqlx::test]
async fn a_viewer_only_regeneration_reserves_for_one_half(pool: PgPool) {
let event_id = seed_event(&pool, "wedding").await;
let user = seed_user(&pool, event_id, "Anna").await;
seed_upload(&pool, event_id, user, 1_000, false).await;
let epoch = release_gallery(&pool, "wedding").await.expect("released");
for t in ["zip", "html"] {
assert!(claim_job(&pool, event_id, t, epoch).await);
assert!(finalize_job(&pool, event_id, t, epoch, &format!("exports/{t}")).await);
}
assert_eq!(armed_job_count(&pool, event_id).await, 0);
// A moderated comment: bump the epoch, carry the ZIP forward, re-arm only the viewer.
let (_, _, next) = bump_epoch(&pool, "wedding").await.expect("bumped");
assert!(
carry_zip_forward(&pool, event_id, next).await,
"the finished ZIP is re-stamped, not rebuilt"
);
let mut conn = pool.acquire().await.expect("acquire");
enqueue_types_at_epoch(&mut conn, event_id, next, &["html"]).await;
assert_eq!(
armed_job_count(&pool, event_id).await,
1,
"only the viewer is being rebuilt, so only one archive's worth of space is needed"
);
}

View File

@@ -0,0 +1,352 @@
//! DB-backed tests for the deleted-media sweep (`services/maintenance.rs`).
//!
//! Context, in two halves.
//!
//! The compression worker deliberately no longer deletes an upload's original when its transcode
//! fails — a transient ENOSPC or a codec panic must never destroy the only copy of a photo a guest
//! cannot retake. But the row is soft-deleted and the uploader's quota IS refunded, so those bytes
//! become invisible, unowned and free.
//!
//! The SAME hole was reachable by the ordinary path, and that one is not an edge case at all:
//! `soft_delete_in_event` refunds `total_upload_bytes` on every guest or host delete and nothing
//! removed the files, so the quota stopped bounding the disk. Upload 500 MB, delete, quota back to
//! zero, upload another 500 MB — a guest curating their camera roll, which is what people do. The
//! sweep used to reach only `compression_status = 'failed'`, so it never touched this case; the
//! test below that now asserts an owner-deleted upload IS reclaimed is the one that used to assert
//! the opposite.
//!
//! Two windows, because the two deletes mean different things: 14 days for a failure an operator
//! may want to investigate, 24 hours for a removal someone asked for (14 days outlives the whole
//! event, so a deliberate delete would never reclaim anything while it mattered).
//!
//! The selection predicate is the whole safety argument — it must reach both leftovers and never a
//! live upload — so that is what these pin, following the same "reproduce the SQL verbatim" pattern
//! as `upload_concurrency.rs`. `#[sqlx::test]` gives each test a fresh, migrated database.
mod common;
use common::*;
use sqlx::PgPool;
use uuid::Uuid;
const FAILED_DAYS: i64 = 14;
const DELETED_HOURS: i64 = 24;
/// SRC: `services/maintenance.rs::cleanup_deleted_media` — the selection, verbatim.
async fn sweep_selects(pool: &PgPool, failed_days: i64, deleted_hours: i64) -> Vec<Uuid> {
type Row = (Uuid, String, Option<String>, Option<String>, Option<String>);
sqlx::query_as::<_, Row>(
"SELECT id, original_path, preview_path, display_path, thumbnail_path FROM upload
WHERE deleted_at IS NOT NULL
AND CASE WHEN compression_status = 'failed'
THEN deleted_at < NOW() - ($1 || ' days')::interval
ELSE deleted_at < NOW() - ($2 || ' hours')::interval
END
AND (original_path <> '' OR preview_path IS NOT NULL
OR display_path IS NOT NULL OR thumbnail_path IS NOT NULL)",
)
.bind(failed_days.to_string())
.bind(deleted_hours.to_string())
.fetch_all(pool)
.await
.expect("sweep query")
.into_iter()
.map(|(id, ..)| id)
.collect()
}
/// Seed an upload aged `deleted_hours_ago` (None = live), with optional derivative paths.
async fn seed_aged_upload(
pool: &PgPool,
event_id: Uuid,
user_id: Uuid,
status: &str,
deleted_hours_ago: Option<i64>,
original_path: &str,
derivatives: bool,
) -> Uuid {
sqlx::query_scalar(
"INSERT INTO upload (event_id, user_id, original_path, mime_type, original_size_bytes,
compression_status, deleted_at,
preview_path, display_path, thumbnail_path)
VALUES ($1, $2, $3, 'image/jpeg', 1000, $4,
CASE WHEN $5::bigint IS NULL THEN NULL
ELSE NOW() - ($5::text || ' hours')::interval END,
CASE WHEN $6 THEN 'previews/p.jpg' END,
CASE WHEN $6 THEN 'displays/d.jpg' END,
CASE WHEN $6 THEN 'thumbs/t.jpg' END)
RETURNING id",
)
.bind(event_id)
.bind(user_id)
.bind(original_path)
.bind(status)
.bind(deleted_hours_ago)
.bind(derivatives)
.fetch_one(pool)
.await
.expect("seed upload")
}
/// A live upload is untouchable no matter how the windows are configured.
///
/// PREVENTS: the catastrophic loosening. Everything else here is about reclaiming more; this is the
/// one assertion that must never bend.
#[sqlx::test]
async fn a_live_upload_is_never_selected(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-live").await;
let user_id = seed_user(&pool, event_id, "Sweeper").await;
for status in ["done", "failed", "processing", "pending"] {
let live = seed_aged_upload(
&pool,
event_id,
user_id,
status,
None,
"originals/e/live.jpg",
true,
)
.await;
assert!(
!sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
.await
.contains(&live),
"a non-deleted upload with status {status} must never be swept"
);
}
}
/// THE FIX. An upload a guest or host deliberately deleted is reclaimed once past 24 hours.
///
/// PREVENTS: the regression back to a sweep scoped to `compression_status = 'failed'`, which is
/// what let the quota stop bounding the disk. This assertion is the inverse of the one this file
/// used to make.
#[sqlx::test]
async fn a_deliberately_deleted_upload_is_reclaimed_after_a_day(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-deleted").await;
let user_id = seed_user(&pool, event_id, "Curator").await;
let deleted = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/owner.jpg",
true,
)
.await;
// Still inside the window — a mis-tap is recoverable for a day.
let recent = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(2),
"originals/e/recent.jpg",
true,
)
.await;
let selected = sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await;
assert!(
selected.contains(&deleted),
"a deliberate delete past the window must be reclaimed — this is the leak"
);
assert!(
!selected.contains(&recent),
"a delete inside the window keeps its recovery grace"
);
}
/// The two windows are independent: a failure is retained far longer than a deliberate delete.
///
/// PREVENTS: collapsing them into one. Applying 24h to failures would destroy the recovery window
/// the retained-original fix exists to provide; applying 14 days to deliberate deletes would mean
/// nothing is ever reclaimed during an event.
#[sqlx::test]
async fn the_two_retention_windows_do_not_bleed_into_each_other(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-windows").await;
let user_id = seed_user(&pool, event_id, "Windows").await;
// 48h old: past the deliberate window, nowhere near the failure window.
let failed_recent = seed_aged_upload(
&pool,
event_id,
user_id,
"failed",
Some(48),
"originals/e/f-recent.jpg",
false,
)
.await;
let deleted_same_age = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/d-same.jpg",
false,
)
.await;
// 30 days old: past both.
let failed_old = seed_aged_upload(
&pool,
event_id,
user_id,
"failed",
Some(30 * 24),
"originals/e/f-old.jpg",
false,
)
.await;
let selected = sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await;
assert!(
!selected.contains(&failed_recent),
"a 2-day-old compression failure is still inside its 14-day recovery window"
);
assert!(
selected.contains(&deleted_same_age),
"a deliberate delete of the same age is past its 24-hour window"
);
assert!(
selected.contains(&failed_old),
"a 30-day-old failure is past both windows"
);
}
/// Boundary behaviour on both windows.
#[sqlx::test]
async fn retention_windows_are_honoured_at_the_boundary(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-boundary").await;
let user_id = seed_user(&pool, event_id, "Boundary").await;
let cases = [
("failed", 13 * 24, false, "13 days"),
("failed", 15 * 24, true, "15 days"),
("done", 23, false, "23 hours"),
("done", 25, true, "25 hours"),
];
for (status, hours, expected, label) in cases {
let id = seed_aged_upload(
&pool,
event_id,
user_id,
status,
Some(hours),
"originals/e/b.jpg",
false,
)
.await;
assert_eq!(
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
.await
.contains(&id),
expected,
"a {status} upload deleted {label} ago: expected swept={expected}"
);
sqlx::query("DELETE FROM upload WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("clean up");
}
}
/// A row is re-selected until EVERY one of its paths is cleared.
///
/// PREVENTS: two failures at once. The sweep used to clear `original_path` alone, which was right
/// for its only case (a failed compression produces no derivatives) but leaves preview, display and
/// thumbnail on disk the moment it reaches a successfully processed upload — three files per
/// upload, none of them counted in `original_size_bytes`, that nothing else ever removes. And a row
/// whose paths are all cleared must stop coming back, or every hourly tick logs a phantom reclaim
/// forever.
#[sqlx::test]
async fn a_row_is_reselected_until_every_path_is_cleared(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-idempotent").await;
let user_id = seed_user(&pool, event_id, "Idem").await;
let id = seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/once.jpg",
true,
)
.await;
assert_eq!(sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await, [id]);
// Clearing only the original is NOT enough — the derivatives are still on disk.
sqlx::query("UPDATE upload SET original_path = '' WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("clear original");
assert_eq!(
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS).await,
[id],
"derivatives left behind must keep the row selected"
);
sqlx::query(
"UPDATE upload SET preview_path = NULL, display_path = NULL, thumbnail_path = NULL
WHERE id = $1",
)
.bind(id)
.execute(&pool)
.await
.expect("clear derivatives");
assert!(
sweep_selects(&pool, FAILED_DAYS, DELETED_HOURS)
.await
.is_empty(),
"a fully swept row must not come back"
);
}
/// The derivative backfill must never resurrect what the sweep just reclaimed.
///
/// PREVENTS: an interaction, not a bug in either piece. The sweep nulls `preview_path`, and
/// `backfill_stale_derivatives` selects on `display_path IS NULL AND preview_path IS NOT NULL` —
/// close enough that a future edit to either could have the backfill re-decode an original that is
/// no longer on disk, on every boot. `deleted_at IS NULL` is what keeps them apart.
#[sqlx::test]
async fn the_backfill_ignores_swept_rows(pool: PgPool) {
let event_id = seed_event(&pool, "sweep-backfill").await;
let user_id = seed_user(&pool, event_id, "Backfill").await;
seed_aged_upload(
&pool,
event_id,
user_id,
"done",
Some(48),
"originals/e/gone.jpg",
true,
)
.await;
// SRC: `services/compression.rs::backfill_stale_derivatives` — the selection, verbatim.
let backfilled: Vec<(Uuid, String, String)> = sqlx::query_as(
"SELECT id, original_path, mime_type FROM upload
WHERE deleted_at IS NULL AND mime_type LIKE 'image/%'
AND original_path IS NOT NULL
AND (
(display_path IS NULL AND preview_path IS NOT NULL)
OR derivatives_rev < $1
)",
)
.bind(1i16)
.fetch_all(&pool)
.await
.expect("backfill query");
assert!(
backfilled.is_empty(),
"a soft-deleted row must be invisible to the backfill, before or after sweeping"
);
}

View File

@@ -17,7 +17,15 @@ services:
deploy:
resources:
limits:
memory: 512M
# 1G, not 512M. DATABASE_MAX_CONNECTIONS defaults to 30 for a ~100-guest event
# (feed polling + SSE + uploads at once), and 30 backends plus Postgres 16's
# default shared_buffers leaves very little headroom at 512M. An OOM here does
# not degrade one feature — it takes the event down, because every request
# path touches the database. Memory is the cheaper knob than shrinking the
# pool back and reintroducing the queueing it was raised to fix.
#
# Raising DATABASE_MAX_CONNECTIONS further means raising this too.
memory: 1G
app:
build:

View File

@@ -12,10 +12,17 @@
# Mirror prod's security headers (minus HSTS, which is HTTPS-only).
header {
X-Content-Type-Options "nosniff"
X-Frame-Options "DENY"
Referrer-Policy "strict-origin-when-cross-origin"
}
# Mirror prod's export carve-out: the keepsake download targets a hidden
# same-origin iframe, and WebKit enforces XFO before Content-Disposition.
# Two disjoint matchers, not an override — see the comment in ../Caddyfile.
@framable path /api/v1/export/zip /api/v1/export/html
@not_framable not path /api/v1/export/zip /api/v1/export/html
header @framable X-Frame-Options "SAMEORIGIN"
header @not_framable X-Frame-Options "DENY"
reverse_proxy /api/* app:3000
reverse_proxy /media/* app:3000
reverse_proxy /health app:3000

View File

@@ -42,11 +42,29 @@ services:
EVENT_NAME: E2E Test Event
APP_PORT: '3000'
MEDIA_PATH: /media
# Exports MUST live outside MEDIA_PATH — see the note on the volume below and
# config.rs::validate. Omitting this left exports on the container's writable
# layer at the /exports default, so the test stack diverged from the prod layout
# it claims to mirror, and export-leak/export-video wrote real archives into
# ephemeral storage.
EXPORT_PATH: /exports
SESSION_EXPIRY_DAYS: '30'
EVENTSNAP_TEST_MODE: '1' # ENABLES /admin/__truncate — never set in prod
RUST_LOG: eventsnap_backend=info,tower_http=warn
volumes:
- media_data:/media
# Separate volume, exactly as in production: a keepsake archive contains every
# photo in the event, so it is kept off the media tree.
- exports_data:/exports
# Mirror production's cap (docker-compose.yml). The test stack having NO memory limit is
# why an unbounded image decode was invisible here: a 99 MP upload that would OOM-kill the
# 1 GiB production container simply succeeded in CI. A test environment more generous than
# production cannot catch a resource bug — the same shape as WebKit being absent from CI
# and /health existing only in Caddyfile.test.
deploy:
resources:
limits:
memory: 1G
expose:
- '3000'
@@ -75,3 +93,4 @@ services:
volumes:
media_data:
exports_data:

View File

@@ -54,6 +54,27 @@ export const db = {
);
},
async compressionStatus(uploadId: string): Promise<string | null> {
return withClient(async (c) => {
const r = await c.query<{ compression_status: string }>(
`SELECT compression_status FROM upload WHERE id = $1`,
[uploadId]
);
return r.rows[0]?.compression_status ?? null;
});
},
/** Which revision of the derivative pipeline produced this row's preview/display. */
async derivativesRev(uploadId: string): Promise<number | null> {
return withClient(async (c) => {
const r = await c.query<{ derivatives_rev: number }>(
`SELECT derivatives_rev FROM upload WHERE id = $1`,
[uploadId]
);
return r.rows[0]?.derivatives_rev ?? null;
});
},
async countUploadsForUser(userId: string): Promise<number> {
return withClient(async (c) => {
const r = await c.query<{ count: string }>(
@@ -84,6 +105,20 @@ export const db = {
});
},
/**
* Overstate an upload's recorded size.
*
* The keepsake size estimate and the low-disk threshold are pure SQL over
* `original_size_bytes` — no file is read — so this is the lever for driving "the keepsake
* would not fit" without a genuinely full disk. The bytes on disk are unchanged; only the
* accounting the warning reads from moves.
*/
async setUploadSizeBytes(uploadId: string, bytes: number) {
await withClient((c) =>
c.query(`UPDATE upload SET original_size_bytes = $2 WHERE id = $1`, [uploadId, bytes])
);
},
async setExportReleased(slug: string, released: boolean) {
await withClient((c) =>
c.query(`UPDATE event SET export_released_at = $2 WHERE slug = $1`, [
@@ -121,7 +156,8 @@ export const db = {
async fakeExportJob(
eventSlug: string,
type: 'zip' | 'html',
status: 'pending' | 'running' | 'done'
status: 'pending' | 'running' | 'done' | 'failed',
errorMessage: string | null = null
) {
await withClient(async (c) => {
const ev = await c.query<{ id: string; export_epoch: string }>(
@@ -130,11 +166,12 @@ export const db = {
);
if (ev.rows.length === 0) throw new Error(`No event with slug ${eventSlug}`);
await c.query(
`INSERT INTO export_job (event_id, type, status, progress_pct, completed_at, epoch)
VALUES ($1, $2::export_type, $3::export_status, $4, $5, $6)
`INSERT INTO export_job (event_id, type, status, progress_pct, completed_at, epoch,
error_message)
VALUES ($1, $2::export_type, $3::export_status, $4, $5, $6, $7)
ON CONFLICT (event_id, type) DO UPDATE
SET status = EXCLUDED.status, progress_pct = EXCLUDED.progress_pct,
epoch = EXCLUDED.epoch`,
epoch = EXCLUDED.epoch, error_message = EXCLUDED.error_message`,
[
ev.rows[0].id,
type,
@@ -142,6 +179,7 @@ export const db = {
status === 'done' ? 100 : 0,
status === 'done' ? new Date() : null,
ev.rows[0].export_epoch,
errorMessage,
]
);
});

29
e2e/helpers/webkit.ts Normal file
View File

@@ -0,0 +1,29 @@
import { test } from '@playwright/test';
/**
* Skip a test that depends on persisting a Blob/File in IndexedDB when running on
* Playwright's WebKit.
*
* The client upload queue (`frontend/src/lib/upload-queue.ts`) stores the file itself in
* IndexedDB so a backgrounded or reloaded phone can resume the upload. Playwright's Linux
* WebKit build cannot store Blobs in IndexedDB at all — `put()` fails with
* "UnknownError: Error preparing Blob/File data to be stored in object store". Verified to
* be the harness, not the app: a Blob constructed in-page with `new Blob([bytes])` fails
* exactly the same way, while Chromium stores both that and a `setInputFiles` File fine.
* Real iOS Safari supports Blobs in IndexedDB, so this is NOT evidence of a bug on the
* platform these tests exist to protect.
*
* Any test that drives the composer (FAB → UploadSheet → /upload → submit) hits this,
* because `handleSubmit` awaits `addToQueue`, which throws before it can navigate.
*
* This is deliberately narrow. WebKit still runs every API-driven upload test, the whole of
* 01-auth, 03-feed and 06-export — including the keepsake download, which only WebKit can
* meaningfully verify. If Playwright's WebKit ever gains IndexedDB Blob support, delete this
* helper and the four call sites.
*/
export function skipIfNoIdbBlobs(browserName: string) {
test.skip(
browserName === 'webkit',
"Playwright's Linux WebKit cannot store Blobs in IndexedDB (harness limitation, not an iOS one) — the client upload queue can't be exercised there"
);
}

View File

@@ -9,12 +9,12 @@ acting as the showcase display.
**Validate the shipping config.** We run at the real production defaults —
compression concurrency (`COMPRESSION_WORKER_CONCURRENCY`, default **2**), DB
pool (default **10**), quotas **on** — and answer: *does the app survive the
event, and how far behind real-time does the diashow fall?*
pool (default **10**), quotas **on** — and answer: _does the app survive the
event, and how far behind real-time does the diashow fall?_
The headline metric is **pipeline latency**: time from an upload succeeding to
its preview being ready (`upload-processed` SSE event) — i.e. *how long until the
photo appears on the diashow*. A backlog that builds is fine; a backlog that
its preview being ready (`upload-processed` SSE event) — i.e. _how long until the
photo appears on the diashow_. A backlog that builds is fine; a backlog that
**never drains** is a fail for a live event.
## Methodology: what we change vs. shipping
@@ -86,13 +86,13 @@ compression backlog to drain** before reporting.
## What the flags mean
| Flag | Meaning |
|------|---------|
| `✗ 5xx` | server errored under load — hard fail |
| `✗ 507` | quota rejected uploads — disk/quota misconfig for the event size |
| Flag | Meaning |
| ------------------------- | ---------------------------------------------------------------------------- |
| `✗ 5xx` | server errored under load — hard fail |
| `✗ 507` | quota rejected uploads — disk/quota misconfig for the event size |
| `✗ backlog did not drain` | compression can't keep up even after uploads stop — diashow never catches up |
| `⚠ pipeline p95 > 60s` | photos take >1 min to appear on the diashow at peak |
| `⚠ SSE resyncs` | live consumers lagged the broadcast channel |
| `⚠ pipeline p95 > 60s` | photos take >1 min to appear on the diashow at peak |
| `⚠ SSE resyncs` | live consumers lagged the broadcast channel |
## Knobs
@@ -101,7 +101,7 @@ All via env (see header of `driver.mjs`): `LT_GUESTS`, `LT_IMAGES`,
`LT_TRUNCATE`, `LT_DRAIN_TIMEOUT_SEC`, `LT_KEEP_RATELIMITS`, `LT_BASE`,
`LT_APP_CONTAINER`, `LT_DB_CONTAINER`.
To later answer *"what config should I deploy?"*, re-run with a rebuilt stack
To later answer _"what config should I deploy?"_, re-run with a rebuilt stack
that sets `COMPRESSION_WORKER_CONCURRENCY` higher (boot-time env var in
`docker-compose.test.yml`) and compare the pipeline-latency / drain numbers.

View File

@@ -10,17 +10,40 @@ const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const j = (r) => r.json();
const adminLogin = () =>
fetch(`${BASE}/api/v1/admin/login`, { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ password: 'admin-test-pw' }) }).then(j).then((b) => b.jwt);
fetch(`${BASE}/api/v1/admin/login`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ password: 'admin-test-pw' }),
})
.then(j)
.then((b) => b.jwt);
const join = (name) =>
fetch(`${BASE}/api/v1/join`, { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ display_name: name }) }).then(j);
fetch(`${BASE}/api/v1/join`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ display_name: name }),
}).then(j);
async function truncate(admin) {
await fetch(`${BASE}/api/v1/admin/__truncate`, { method: 'POST', headers: { Authorization: `Bearer ${admin}` } });
await fetch(`${BASE}/api/v1/admin/__truncate`, {
method: 'POST',
headers: { Authorization: `Bearer ${admin}` },
});
}
async function upload(jwt) {
const form = new FormData();
form.append('file', new Blob([readFileSync('/tmp/eventsnap-loadtest/photos/photo_000.jpg')], { type: 'image/jpeg' }), 'live.jpg');
form.append(
'file',
new Blob([readFileSync('/tmp/eventsnap-loadtest/photos/photo_000.jpg')], {
type: 'image/jpeg',
}),
'live.jpg'
);
form.append('caption', 'LIVE-PROBE');
const r = await fetch(`${BASE}/api/v1/upload`, { method: 'POST', headers: { Authorization: `Bearer ${jwt}` }, body: form });
const r = await fetch(`${BASE}/api/v1/upload`, {
method: 'POST',
headers: { Authorization: `Bearer ${jwt}` },
body: form,
});
return (await r.json()).id;
}
@@ -35,7 +58,9 @@ const browser = await chromium.launch();
const ctx = await browser.newContext({ viewport: { width: 1280, height: 800 } });
const page = await ctx.newPage();
const streamReqs = [];
page.on('request', (req) => { if (req.url().includes('/stream')) streamReqs.push(req.url().replace(BASE, '')); });
page.on('request', (req) => {
if (req.url().includes('/stream')) streamReqs.push(req.url().replace(BASE, ''));
});
await page.goto(`${BASE}/join`, { waitUntil: 'domcontentloaded' });
await page.evaluate((g) => {
@@ -53,7 +78,9 @@ const emptyBefore = (await page.locator('text=Noch keine Beiträge').count()) >
const streamOpened = streamReqs.length > 0;
console.log(`\n1) direct /diashow on empty event:`);
console.log(` "Noch keine Beiträge" shown: ${emptyBefore} (expected: true)`);
console.log(` SSE stream opened: ${streamOpened} ${streamOpened ? '('+streamReqs.join(', ')+')' : ''} (expected: true — this is the fix)`);
console.log(
` SSE stream opened: ${streamOpened} ${streamOpened ? '(' + streamReqs.join(', ') + ')' : ''} (expected: true — this is the fix)`
);
// Now a guest uploads a photo while the display is open.
console.log(`\n2) guest uploads a photo (display already open)…`);
@@ -65,7 +92,11 @@ for (let i = 0; i < 15; i++) {
await sleep(1000);
const stillEmpty = (await page.locator('text=Noch keine Beiträge').count()) > 0;
const imgs = await page.locator('img').count();
if (!stillEmpty && imgs > 0) { appeared = true; console.log(` photo appeared live after ~${i + 1}s (img rendered, placeholder gone)`); break; }
if (!stillEmpty && imgs > 0) {
appeared = true;
console.log(` photo appeared live after ~${i + 1}s (img rendered, placeholder gone)`);
break;
}
}
if (!appeared) console.log(` ✗ photo did NOT appear within 15s`);

View File

@@ -132,9 +132,19 @@ const truncate = (adminJwt) =>
api('/admin/__truncate', { method: 'POST', token: adminJwt, expect: [204] });
const CAPTIONS = [
'Was für ein magischer Tag 💍', 'Der erste Tanz 🕺', 'Prost! 🥂', 'Die Torte 🍰',
'Feuerwerk 🎆', 'Beste Freunde 💕', 'Was für eine Stimmung! 🎉', 'Details 🌸',
'Sonnenuntergang 🌅', 'Tanzfläche brennt 🔥', null, null, null,
'Was für ein magischer Tag 💍',
'Der erste Tanz 🕺',
'Prost! 🥂',
'Die Torte 🍰',
'Feuerwerk 🎆',
'Beste Freunde 💕',
'Was für eine Stimmung! 🎉',
'Details 🌸',
'Sonnenuntergang 🌅',
'Tanzfläche brennt 🔥',
null,
null,
null,
];
const TAGS = ['hochzeit', 'liebe', 'party', 'tanzen', 'natur', 'feier', 'freunde', 'dessert'];
@@ -238,9 +248,12 @@ async function sampleResources() {
const out = { ts: now() };
try {
const { stdout } = await execFileAsync('docker', [
'stats', '--no-stream', '--format',
'stats',
'--no-stream',
'--format',
'{{.Name}}|{{.CPUPerc}}|{{.MemUsage}}',
cfg.appContainer, cfg.dbContainer,
cfg.appContainer,
cfg.dbContainer,
]);
out.docker = stdout.trim();
} catch (e) {
@@ -259,8 +272,15 @@ async function sampleResources() {
function psql(sql) {
return execFileAsync('docker', [
'exec', cfg.dbContainer, 'psql', '-U', 'eventsnap_test', '-d', 'eventsnap_test',
'-tAc', sql,
'exec',
cfg.dbContainer,
'psql',
'-U',
'eventsnap_test',
'-d',
'eventsnap_test',
'-tAc',
sql,
]);
}
@@ -292,8 +312,13 @@ function summarize(nums) {
const s = [...nums].sort((a, b) => a - b);
const sum = s.reduce((a, b) => a + b, 0);
return {
n: s.length, min: s[0], max: s[s.length - 1], mean: Math.round(sum / s.length),
p50: pct(s, 50), p95: pct(s, 95), p99: pct(s, 99),
n: s.length,
min: s[0],
max: s[s.length - 1],
mean: Math.round(sum / s.length),
p50: pct(s, 50),
p95: pct(s, 95),
p99: pct(s, 99),
};
}
@@ -490,7 +515,9 @@ async function main() {
await sleep(cfg.windowSec * 1000 + 500);
await Promise.all(burstPromises);
clearInterval(ticker);
console.log(`[run] all bursts issued. uploaded ${done}, ok ${uploads.filter((u) => u.status === 201).length}`);
console.log(
`[run] all bursts issued. uploaded ${done}, ok ${uploads.filter((u) => u.status === 201).length}`
);
// Drain: wait for compression backlog to clear (SSE processed ⊇ successful ids)
console.log('[drain] waiting for compression backlog to clear…');
@@ -589,21 +616,28 @@ async function main() {
console.log('\n' + '━'.repeat(72));
console.log('RESULTS');
console.log('━'.repeat(72));
console.log(`uploads: ${okUploads.length}/${uploads.length} ok — byStatus ${JSON.stringify(byStatus)}`);
console.log(
`uploads: ${okUploads.length}/${uploads.length} ok — byStatus ${JSON.stringify(byStatus)}`
);
console.log(`upload latency ms: ${JSON.stringify(uploadLatency)}`);
console.log(`pipeline latency ms (upload→preview ready): ${JSON.stringify(pipelineLatency)}`);
console.log(`backlog drain: ${(drainMs / 1000).toFixed(1)}s, cleared=${report.drain.cleared}`);
console.log(`final db compression status: ${JSON.stringify(finalCounts)}`);
console.log(`sse: ${sseClients.length} conns, ${totalReconnects} reconnects, ${totalResyncs} resyncs`);
console.log(
`sse: ${sseClients.length} conns, ${totalReconnects} reconnects, ${totalResyncs} resyncs`
);
console.log(`\nfull report → ${outPath}`);
// Heuristic pass/fail flags (validate shipping config)
const flags = [];
const err5xx = Object.entries(byStatus).filter(([s]) => +s >= 500).reduce((a, [, n]) => a + n, 0);
const err5xx = Object.entries(byStatus)
.filter(([s]) => +s >= 500)
.reduce((a, [, n]) => a + n, 0);
if (err5xx > 0) flags.push(`${err5xx} server errors (5xx)`);
if (byStatus['507']) flags.push(`${byStatus['507']} quota rejections (507)`);
if (byStatus['413']) flags.push(`${byStatus['413']} too-large (413)`);
if (byStatus['429']) flags.push(`${byStatus['429']} rate-limited (429) — unexpected with limits off`);
if (byStatus['429'])
flags.push(`${byStatus['429']} rate-limited (429) — unexpected with limits off`);
if (!report.drain.cleared) flags.push(`✗ backlog did NOT drain within ${cfg.drainTimeoutSec}s`);
if (totalResyncs > sseClients.length) flags.push(`${totalResyncs} SSE resyncs (consumer lag)`);
if (pipelineLatency && pipelineLatency.p95 > 60000)

View File

@@ -37,12 +37,19 @@ export class ExportPage {
/**
* The "Download" button inside the card whose heading is `heading`.
*
* Scoped to the card element (`div.rounded-xl`) rather than "any div containing the
* heading" — the latter also matches the page wrapper, which contains BOTH cards' buttons.
* Scoped to the card element rather than "any div containing the heading" — the latter
* also matches the page wrapper, which contains BOTH cards' buttons.
*
* The scope class is `div.card` (see the ZIP/HTML cards in
* frontend/src/routes/export/+page.svelte). It was previously `div.rounded-xl`, which
* matched NOTHING: `.card` is a Tailwind `@apply` component class
* (frontend/src/lib/styles/components.css) so the DOM only ever carries `class="card p-5"`
* — and the utility it applies is `rounded-2xl` anyway. Both card-scoped locators were
* therefore dead, which is why the "shows enabled download buttons" test was red.
*/
private cardButton(heading: string): Locator {
return this.page
.locator('div.rounded-xl')
.locator('div.card')
.filter({ has: this.page.getByRole('heading', { name: heading, exact: true }) })
.getByRole('button', { name: 'Download', exact: true });
}

View File

@@ -146,8 +146,18 @@ export default defineConfig({
// that on `@smoke` — which exists on exactly two specs — meant the entire iOS guarantee was
// one happy path and one join test. Every other UA here is a secondary browser and a smoke
// check is proportionate; WebKit is not. Give it the core journeys the guest actually walks:
// join/recover, upload, and browse the feed.
testMatch: ['**/__smoke/**', '**/01-auth/**', '**/02-upload/**', '**/03-feed/**'],
// join/recover, upload, browse the feed — and take the keepsake home.
//
// 06-export is here because WebKit is the ONLY engine that enforces X-Frame-Options on the
// hidden download iframe. Excluding it is what let a site-wide `XFO: DENY` ship a keepsake
// download that silently did nothing on iOS. See 06-export/download-iframe.spec.ts.
testMatch: [
'**/__smoke/**',
'**/01-auth/**',
'**/02-upload/**',
'**/03-feed/**',
'**/06-export/**',
],
},
{
name: 'firefox-android',

View File

@@ -4,24 +4,44 @@ import { chromium, devices } from '@playwright/test';
import { Client } from 'pg';
import { readFileSync, mkdirSync } from 'node:fs';
const BASE = 'http://localhost:3101';
const PHOTOS = '/tmp/eventsnap-shots/photos';
const OUT = '/tmp/eventsnap-shots';
mkdirSync(OUT, { recursive: true });
const api = (path, opts = {}) =>
fetch(`${BASE}/api/v1${path}`, opts).then(async (r) => ({ status: r.status, body: await r.text().then((t) => { try { return JSON.parse(t); } catch { return t; } }) }));
fetch(`${BASE}/api/v1${path}`, opts).then(async (r) => ({
status: r.status,
body: await r.text().then((t) => {
try {
return JSON.parse(t);
} catch {
return t;
}
}),
}));
const authHeaders = (jwt, json = true) => ({ Authorization: `Bearer ${jwt}`, ...(json ? { 'Content-Type': 'application/json' } : {}) });
const authHeaders = (jwt, json = true) => ({
Authorization: `Bearer ${jwt}`,
...(json ? { 'Content-Type': 'application/json' } : {}),
});
async function adminLogin() {
const r = await api('/admin/login', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ password: 'admin-test-pw' }) });
const r = await api('/admin/login', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ password: 'admin-test-pw' }),
});
return r.body.jwt;
}
async function joinGuest(name) {
const r = await api('/join', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ display_name: name }) });
if (r.status !== 201) throw new Error('join failed ' + name + ' ' + r.status + ' ' + JSON.stringify(r.body));
const r = await api('/join', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ display_name: name }),
});
if (r.status !== 201)
throw new Error('join failed ' + name + ' ' + r.status + ' ' + JSON.stringify(r.body));
return r.body; // {jwt,pin,user_id}
}
async function upload(jwt, file, caption, hashtags) {
@@ -29,7 +49,11 @@ async function upload(jwt, file, caption, hashtags) {
form.append('file', new Blob([readFileSync(file)], { type: 'image/jpeg' }), 'photo.jpg');
if (caption) form.append('caption', caption);
if (hashtags) form.append('hashtags', hashtags);
const r = await fetch(`${BASE}/api/v1/upload`, { method: 'POST', headers: authHeaders(jwt, false), body: form });
const r = await fetch(`${BASE}/api/v1/upload`, {
method: 'POST',
headers: authHeaders(jwt, false),
body: form,
});
if (r.status !== 201) throw new Error('upload failed ' + r.status + ' ' + (await r.text()));
return (await r.json()).id;
}
@@ -37,16 +61,34 @@ async function like(jwt, id) {
await fetch(`${BASE}/api/v1/upload/${id}/like`, { method: 'POST', headers: authHeaders(jwt) });
}
async function comment(jwt, id, body) {
await fetch(`${BASE}/api/v1/upload/${id}/comments`, { method: 'POST', headers: authHeaders(jwt), body: JSON.stringify({ body }) });
await fetch(`${BASE}/api/v1/upload/${id}/comments`, {
method: 'POST',
headers: authHeaders(jwt),
body: JSON.stringify({ body }),
});
}
const pg = () => new Client({ host: 'localhost', port: 55432, user: 'eventsnap_test', password: 'eventsnap_test', database: 'eventsnap_test' });
const pg = () =>
new Client({
host: 'localhost',
port: 55432,
user: 'eventsnap_test',
password: 'eventsnap_test',
database: 'eventsnap_test',
});
async function truncate(adminJwt) {
await fetch(`${BASE}/api/v1/admin/__truncate`, { method: 'POST', headers: authHeaders(adminJwt) });
await fetch(`${BASE}/api/v1/admin/__truncate`, {
method: 'POST',
headers: authHeaders(adminJwt),
});
}
async function patchConfig(adminJwt, patch) {
await fetch(`${BASE}/api/v1/admin/config`, { method: 'PATCH', headers: authHeaders(adminJwt), body: JSON.stringify(patch) });
await fetch(`${BASE}/api/v1/admin/config`, {
method: 'PATCH',
headers: authHeaders(adminJwt),
body: JSON.stringify(patch),
});
}
// ---- SEED ----
@@ -54,9 +96,27 @@ console.log('[seed] admin login + reset');
let admin = await adminLogin();
await truncate(admin);
admin = await adminLogin();
await patchConfig(admin, { rate_limits_enabled: 'false', upload_rate_enabled: 'false', feed_rate_enabled: 'false', export_rate_enabled: 'false', join_rate_enabled: 'false', quota_enabled: 'false', storage_quota_enabled: 'false', upload_count_quota_enabled: 'false' });
await patchConfig(admin, {
rate_limits_enabled: 'false',
upload_rate_enabled: 'false',
feed_rate_enabled: 'false',
export_rate_enabled: 'false',
join_rate_enabled: 'false',
quota_enabled: 'false',
storage_quota_enabled: 'false',
upload_count_quota_enabled: 'false',
});
const guests = ['Anna Bauer', 'Lukas Weber', 'Mia Schulz', 'Jonas Fischer', 'Emma Wagner', 'Ben Hoffmann', 'Sophie Klein', 'Paul Richter'];
const guests = [
'Anna Bauer',
'Lukas Weber',
'Mia Schulz',
'Jonas Fischer',
'Emma Wagner',
'Ben Hoffmann',
'Sophie Klein',
'Paul Richter',
];
const accounts = {};
for (const g of guests) accounts[g] = await joinGuest(g);
console.log('[seed] joined', guests.length, 'guests');
@@ -94,13 +154,23 @@ await comment(accounts['Ben Hoffmann'].jwt, ids[1], 'Was ein Abend!');
console.log('[seed] likes + comments done');
// Make one guest a host so /host renders populated
await fetch(`${BASE}/api/v1/host/users/${accounts['Anna Bauer'].user_id}/role`, { method: 'PATCH', headers: authHeaders(admin), body: JSON.stringify({ role: 'host' }) });
await fetch(`${BASE}/api/v1/host/users/${accounts['Anna Bauer'].user_id}/role`, {
method: 'PATCH',
headers: authHeaders(admin),
body: JSON.stringify({ role: 'host' }),
});
// Wait for compression to finish so previews render
const c = pg(); await c.connect();
const c = pg();
await c.connect();
for (let t = 0; t < 40; t++) {
const r = await c.query(`SELECT COUNT(*)::int AS n FROM upload WHERE compression_status <> 'done' AND deleted_at IS NULL`);
if (r.rows[0].n === 0) { console.log('[seed] compression done'); break; }
const r = await c.query(
`SELECT COUNT(*)::int AS n FROM upload WHERE compression_status <> 'done' AND deleted_at IS NULL`
);
if (r.rows[0].n === 0) {
console.log('[seed] compression done');
break;
}
await new Promise((res) => setTimeout(res, 500));
}
await c.end();
@@ -113,17 +183,20 @@ const shot = async (label, theme, who, route, prep) => {
const page = await ctx.newPage();
// Seed localStorage on the origin
await page.goto(`${BASE}/join`, { waitUntil: 'domcontentloaded' });
await page.evaluate(({ jwt, pin, uid, name, theme, mode }) => {
localStorage.setItem('eventsnap_theme', theme);
localStorage.setItem('eventsnap_data_mode', mode);
if (jwt) {
localStorage.setItem('eventsnap_jwt', jwt);
localStorage.setItem('eventsnap_pin', pin);
localStorage.setItem('eventsnap_user_id', uid);
localStorage.setItem('eventsnap_display_name', name);
}
localStorage.setItem('eventsnap_guide_seen', '1');
}, { jwt: who?.jwt, pin: who?.pin, uid: who?.user_id, name: who?.name, theme, mode: 'saver' });
await page.evaluate(
({ jwt, pin, uid, name, theme, mode }) => {
localStorage.setItem('eventsnap_theme', theme);
localStorage.setItem('eventsnap_data_mode', mode);
if (jwt) {
localStorage.setItem('eventsnap_jwt', jwt);
localStorage.setItem('eventsnap_pin', pin);
localStorage.setItem('eventsnap_user_id', uid);
localStorage.setItem('eventsnap_display_name', name);
}
localStorage.setItem('eventsnap_guide_seen', '1');
},
{ jwt: who?.jwt, pin: who?.pin, uid: who?.user_id, name: who?.name, theme, mode: 'saver' }
);
await page.goto(`${BASE}${route}`, { waitUntil: 'domcontentloaded' });
await page.waitForTimeout(1200);
if (prep) await prep(page);
@@ -141,10 +214,17 @@ for (const theme of ['light', 'dark']) {
await shot('01-join', theme, null, '/join');
await shot('02-feed-list', theme, hostWho, '/feed');
await shot('03-feed-grid', theme, hostWho, '/feed', async (p) => {
await p.getByLabel('Rasteransicht').click().catch(() => {});
await p
.getByLabel('Rasteransicht')
.click()
.catch(() => {});
});
await shot('04-lightbox', theme, hostWho, '/feed', async (p) => {
await p.locator('img').first().click().catch(() => {});
await p
.locator('img')
.first()
.click()
.catch(() => {});
await p.waitForTimeout(500);
});
await shot('05-account', theme, hostWho, '/account');

View File

@@ -13,7 +13,11 @@ test.describe('Auth — join flow', () => {
const join = new JoinPage(page);
await join.goto();
await expect(page.getByRole('heading', { name: 'Willkommen!' })).toBeVisible();
// The join form's landing state. There is no "Willkommen!" heading — the wedding
// redesign (f243bfe) split it into a "Willkommen bei" lead-in plus the event name as
// the <h1>, and this assertion was never updated, so it had been failing since.
// Anchor on the testid the markup provides rather than on copy.
await expect(page.getByTestId('join-event-name')).toBeVisible();
const { pin } = await join.joinAs('Alice');
expect(pin).toMatch(/^\d{4}$/);

View File

@@ -0,0 +1,206 @@
/**
* Regression guard — the door must not close on a venue behind one NAT.
*
* `/join` was throttled 5 per 60s keyed purely on the client IP. Every guest at a venue
* arrives from the same public IP (that is what a NAT is), so the whole party shared one
* bucket: 12 guests scanning the QR code within a few seconds meant 5 got in and 7 were
* turned away — with no Retry-After to tell them when to try again. `/feed` (60/min) and
* `/export` (3/DAY) had the identical defect.
*
* These ran green for the same structural reason every time: the e2e reseed forces every
* limiter toggle OFF before each test, so nothing here was ever exercised. Enable them
* explicitly, exactly as 02-upload/rate-limit does.
*/
import { test, expect } from '../../fixtures/test';
import { BASE } from '../../helpers/env';
test.describe('Rate limits — guests behind a shared NAT', () => {
test('a dozen guests can all join from one IP, and 429s carry Retry-After', async ({
api,
adminToken,
}) => {
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
join_rate_enabled: 'true',
});
// Twelve DISTINCT guests, same source IP — the arrival burst at a real party.
const names = Array.from({ length: 12 }, (_, i) => `NatGuest${i}`);
const results = await Promise.all(
names.map((display_name) =>
fetch(`${BASE}/api/v1/join`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ display_name }),
})
)
);
const rejected = results.filter((r) => r.status === 429);
expect(
rejected.length,
`all 12 guests must get in from one IP; ${rejected.length} were turned away`
).toBe(0);
expect(results.every((r) => r.status === 201)).toBe(true);
});
test('one guest retrying their own name is still throttled, and told for how long', async ({
api,
adminToken,
}) => {
// The per-name bucket must still bite — otherwise the NAT fix would have simply
// removed the anti-spam limit rather than re-keyed it.
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
join_rate_enabled: 'true',
});
const attempt = () =>
fetch(`${BASE}/api/v1/join`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ display_name: 'RepeatOffender' }),
});
// 5 per 60s for the same (ip, name): the first succeeds (201), the next four collide
// with the taken name (409), and the sixth exhausts the bucket.
const codes: number[] = [];
for (let i = 0; i < 6; i++) codes.push((await attempt()).status);
expect(codes[0], 'the first join should succeed').toBe(201);
expect(codes.at(-1), 'the 6th attempt on one name must be throttled').toBe(429);
const throttled = await attempt();
expect(throttled.status).toBe(429);
const retryAfter = throttled.headers.get('retry-after');
expect(retryAfter, '429 must tell the client when to come back').toBeTruthy();
expect(Number(retryAfter)).toBeGreaterThan(0);
expect(Number(retryAfter)).toBeLessThanOrEqual(60);
});
test('the feed limit is per-user, not per-IP', async ({ api, adminToken, guest }) => {
// Two guests, one IP. With a limit of 3/min an IP key would let the first guest's
// three reads starve the second entirely.
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
feed_rate_enabled: 'true',
feed_rate_per_min: '3',
});
const a = await guest('FeedHog');
const b = await guest('FeedVictim');
const read = (jwt: string) =>
fetch(`${BASE}/api/v1/feed`, { headers: { Authorization: `Bearer ${jwt}` } });
// Guest A burns their whole allowance.
for (let i = 0; i < 3; i++) expect((await read(a.jwt)).status).toBe(200);
expect((await read(a.jwt)).status, "A's own 4th read is throttled").toBe(429);
// Guest B must be entirely unaffected.
expect((await read(b.jwt)).status, 'B must not inherit As exhausted bucket').toBe(200);
});
test('the export limit is per-user — one guest cannot spend the whole venues quota', async ({
api,
adminToken,
guest,
host,
db,
}) => {
// The sharpest case: 3 downloads per DAY on an IP key meant the 4th guest to fetch
// their keepsake was locked out until tomorrow.
await db.setExportReleased('e2e-test-event', true);
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
export_rate_enabled: 'true',
export_rate_per_day: '1',
});
const mintAndFetch = async (jwt: string) => {
const res = await fetch(`${BASE}/api/v1/export/ticket`, {
method: 'POST',
headers: { Authorization: `Bearer ${jwt}` },
});
const { ticket } = await res.json();
return fetch(`${BASE}/api/v1/export/zip?ticket=${encodeURIComponent(ticket)}`);
};
const a = await guest('ExportFirst');
const b = await guest('ExportSecond');
// A spends their single daily allowance. The archive itself may not exist (404) —
// what matters is that the limiter admitted the request rather than 429ing it.
expect((await mintAndFetch(a.jwt)).status).not.toBe(429);
expect((await mintAndFetch(a.jwt)).status, 'As second download is throttled').toBe(429);
// B shares A's IP and must still get their keepsake.
expect((await mintAndFetch(b.jwt)).status, 'B must not be locked out by As download').not.toBe(
429
);
// And the host too, for good measure.
expect((await mintAndFetch(host.jwt)).status).not.toBe(429);
});
});
test.describe('Rate limits — /recover name cycling', () => {
test('cycling names from one IP hits the ceiling, while one name is still throttled', async ({
api,
adminToken,
}) => {
// /recover is keyed `recover:{ip}:{name}` — right for its job (stopping someone who
// knows a display name from burning the victim's 3-strike PIN counter), but the name is
// ATTACKER-CHOSEN, so cycling names minted a fresh bucket every time. Behind it sits a
// cost-12 bcrypt verify, including an unconditional throwaway one for unknown names, so
// a name generator was the cheapest way to make the server hash forever.
//
// Squeeze the ceiling so the flood is reproducible without firing 30+ requests.
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
recover_rate_enabled: 'true',
recover_ip_rate_per_min: '5',
});
const attempt = (name: string) =>
fetch(`${BASE}/api/v1/recover`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ display_name: name, pin: '0000' }),
});
// Every name is distinct, so the per-name bucket can never fire — only the ceiling can.
const codes: number[] = [];
for (let i = 0; i < 12; i++) codes.push((await attempt(`Unbekannt${i}_${Date.now()}`)).status);
expect(
codes.filter((c) => c === 429).length,
'name cycling must be capped by the per-IP ceiling'
).toBeGreaterThan(0);
const throttled = await attempt(`Unbekannt99_${Date.now()}`);
expect(throttled.status).toBe(429);
expect(Number(throttled.headers.get('retry-after'))).toBeGreaterThan(0);
});
test('the per-name bucket still protects a real account', async ({ api, adminToken, guest }) => {
// The ceiling must not have REPLACED the anti-guessing control. With a generous ceiling,
// repeated wrong PINs against ONE name must still be shut down by the per-name bucket.
const victim = await guest('PinVictim');
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
recover_rate_enabled: 'true',
recover_ip_rate_per_min: '1000',
});
const attempt = () =>
fetch(`${BASE}/api/v1/recover`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ display_name: victim.displayName, pin: '9999' }),
});
const codes: number[] = [];
for (let i = 0; i < 7; i++) codes.push((await attempt()).status);
expect(codes.at(-1), 'guessing one name must still be throttled').toBe(429);
});
});

View File

@@ -26,6 +26,7 @@
* large fixture files.
*/
import { test, expect } from '../../fixtures/test';
import { skipIfNoIdbBlobs } from '../../helpers/webkit';
import { FeedPage, UploadSheet } from '../../page-objects';
import { mkdtempSync, copyFileSync, rmSync } from 'node:fs';
import { join } from 'node:path';
@@ -100,7 +101,9 @@ test.describe('Upload — client queue under a burst', () => {
guest,
signIn,
db,
browserName,
}) => {
skipIfNoIdbBlobs(browserName);
const g = await guest('BurstSerial');
await signIn(page, g);
await warmUploadChunk(page);
@@ -163,7 +166,9 @@ test.describe('Upload — client queue under a burst', () => {
guest,
signIn,
db,
browserName,
}) => {
skipIfNoIdbBlobs(browserName);
const g = await guest('BurstResume');
await signIn(page, g);
await warmUploadChunk(page);

View File

@@ -0,0 +1,80 @@
/**
* Regression guard — EXIF orientation must be applied when generating derivatives.
*
* Phones do not rotate sensor data. They shoot in the sensor's native landscape and record
* how the camera was held in an EXIF `Orientation` tag. `image`'s `decode()` returns the raw
* pixels and ignores that tag, and the JPEG re-encode writes no EXIF at all — so every
* portrait photo was stored SIDEWAYS in the 800px feed preview, the 2048px diashow display
* and the keepsake, while "Original anzeigen" still rendered it upright (the original keeps
* its tag). That asymmetry is why it reads as a viewer bug instead of a pipeline one.
*
* The fixture is 40x20 landscape pixels tagged Orientation=6 ("rotate 90° CW to display"),
* so a correctly-processed derivative is PORTRAIT (20x40). Asserting on the aspect ratio
* rather than the bytes keeps this robust across encoder changes.
*/
import { test, expect } from '../../fixtures/test';
import { uploadRaw } from '../../helpers/upload-client';
import { BASE } from '../../helpers/env';
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
const EXIF_FIXTURE = join(process.cwd(), 'fixtures', 'media', 'portrait-exif6.jpg');
/**
* Read a baseline/progressive JPEG's pixel dimensions from its SOF marker.
* Avoids pulling an image dependency into the suite for one assertion.
*/
function jpegSize(buf: Buffer): { width: number; height: number } {
let i = 2; // skip SOI
while (i < buf.length) {
if (buf[i] !== 0xff) {
i++;
continue;
}
const marker = buf[i + 1];
// SOF0..SOF15, excluding DHT (c4), JPGA (c8) and DAC (cc)
if (marker >= 0xc0 && marker <= 0xcf && marker !== 0xc4 && marker !== 0xc8 && marker !== 0xcc) {
return { height: buf.readUInt16BE(i + 5), width: buf.readUInt16BE(i + 7) };
}
i += 2 + buf.readUInt16BE(i + 2);
}
throw new Error('no SOF marker found — not a JPEG?');
}
test.describe('Upload — EXIF orientation', () => {
test('a rotated photo is upright in the preview and the display derivative', async ({
guest,
db,
}) => {
const g = await guest('SidewaysShooter');
const res = await uploadRaw(g.jwt, readFileSync(EXIF_FIXTURE), {
filename: 'portrait-exif6.jpg',
contentType: 'image/jpeg',
caption: 'hochkant',
});
expect(res.status).toBe(201);
const { id } = (await res.json()) as { id: string };
// Sanity: the SOURCE really is stored landscape with the tag, otherwise this test
// could pass against a pipeline that does nothing.
const source = jpegSize(readFileSync(EXIF_FIXTURE));
expect(source.width).toBeGreaterThan(source.height);
await expect
.poll(() => db.compressionStatus(id), { timeout: 30_000, intervals: [250] })
.toBe('done');
for (const variant of ['preview', 'display'] as const) {
const r = await fetch(`${BASE}/api/v1/upload/${id}/${variant}`, {
headers: { Authorization: `Bearer ${g.jwt}` },
});
expect(r.status, `${variant} must be served`).toBe(200);
const { width, height } = jpegSize(Buffer.from(await r.arrayBuffer()));
expect(
height,
`${variant} must be portrait (${width}x${height}) — EXIF orientation was not applied`
).toBeGreaterThan(width);
}
});
});

View File

@@ -4,6 +4,7 @@
* IndexedDB queue resumption after refresh, and SSE `upload-processed`.
*/
import { test, expect } from '../../fixtures/test';
import { skipIfNoIdbBlobs } from '../../helpers/webkit';
import { FeedPage, UploadSheet } from '../../page-objects';
import { SseListener } from '../../helpers/sse-listener';
import { join } from 'node:path';
@@ -49,7 +50,9 @@ test.describe('Upload — gallery path', () => {
guest,
signIn,
db,
browserName,
}) => {
skipIfNoIdbBlobs(browserName);
// Previously fixme'd: the UI queue never fired a POST. Root cause was NOT a
// navigation/blob timing quirk but an IndexedDB upgrade bug — the v1→v2
// `upgrade` callback opened a *new* transaction, which throws during a

View File

@@ -0,0 +1,100 @@
/**
* Regression guard — an image that would blow the decode budget must be refused at the
* door, with a reason the guest can act on, and must never allocate.
*
* Two defects met here.
*
* 1. The budget was inert. `max_alloc = 256 MiB` was set, but reading the EXIF orientation
* tag requires `ImageReader::into_decoder()`, which skips the
* `limits.reserve(decoder.total_bytes())` that `decode()` performs — and nothing else
* enforces it (the JPEG decoder's `set_limits` only checks support and dimensions). The
* only real bound was the 12000px per-axis cap, leaving two concurrent decodes at
* 824 MiB against a 1 GiB container. This suite could not have caught it either, because
* the e2e app container had NO memory limit while production is capped at 1 GiB; that cap
* is now mirrored in docker-compose.test.yml so these assertions mean something.
*
* 2. Even with the budget restored, the upload was ACCEPTED with a 201 and then silently
* soft-deleted minutes later when the worker gave up — the photo simply vanished, with at
* best a vague "could not be processed". Admission now runs the same budget check against
* the header, so the guest is told immediately and told why.
*
* Fixture: 11000x9000 = 99 MP, 568 KiB on disk. Deliberately UNDER the per-axis cap, so the
* axis check cannot be what rejects it — 283 MiB decoded against a 256 MiB budget. A fixture
* at 13000px would pass this test against a build with no budget at all.
*/
import { test, expect } from '../../fixtures/test';
import { uploadRaw } from '../../helpers/upload-client';
import { BASE } from '../../helpers/env';
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
const HUGE = join(process.cwd(), 'fixtures', 'media', 'huge-99mp.jpg');
const SAMPLE = join(process.cwd(), 'fixtures', 'media', 'sample.jpg');
test.describe('Upload — an oversized image is refused, not allocated', () => {
test('a 99 MP upload is rejected at admission with a readable reason', async ({ guest, db }) => {
test.setTimeout(60_000);
const g = await guest('BombThrower');
const before = await db.countUploadsForUser(g.userId);
const res = await uploadRaw(g.jwt, readFileSync(HUGE), {
filename: 'huge.jpg',
contentType: 'image/jpeg',
caption: 'zu gross',
});
// 4xx, not 201-then-vanish. The queue classifies this as terminal, so the guest gets the
// message rather than watching the photo disappear.
expect(res.status, 'an undecodable image must be refused at the door').toBe(400);
const body = (await res.json()) as { message?: string };
expect(body.message ?? '', 'the reason must be actionable, not generic').toMatch(/bildpunkte/i);
expect(body.message ?? '', 'and should name the size so it is obvious why').toMatch(/99/);
// Nothing was stored — no row to soft-delete later, no orphaned file to sweep.
expect(await db.countUploadsForUser(g.userId)).toBe(before);
// The backend never allocated: it is still alive and still doing useful work.
expect((await fetch(`${BASE}/health`)).status).toBe(200);
const ok = await uploadRaw(g.jwt, readFileSync(SAMPLE), {
filename: 'after.jpg',
contentType: 'image/jpeg',
});
expect(ok.status).toBe(201);
const after = (await ok.json()) as { id: string };
await expect.poll(() => db.compressionStatus(after.id), { timeout: 30_000 }).toBe('done');
});
test('a burst of oversized uploads leaves the container alive', async ({ guest }) => {
// The concurrent case is the one that OOM'd: `compression_concurrency` is 2, so decodes
// overlapped. Four at once is comfortably past that, and must still cost only header
// reads.
test.setTimeout(60_000);
const g = await guest('BombThrower2');
const bytes = readFileSync(HUGE);
const results = await Promise.all(
Array.from({ length: 4 }, (_, i) =>
uploadRaw(g.jwt, bytes, { filename: `huge-${i}.jpg`, contentType: 'image/jpeg' })
)
);
expect(results.map((r) => r.status)).toEqual([400, 400, 400, 400]);
expect(
(await fetch(`${BASE}/health`)).status,
'concurrent oversized uploads must not kill the backend'
).toBe(200);
});
test('an ordinary photo is unaffected by the admission check', async ({ guest, db }) => {
// The mirror that keeps the check honest: a budget that rejected everything would pass
// both tests above.
const g = await guest('NormalShooter');
const res = await uploadRaw(g.jwt, readFileSync(SAMPLE), {
filename: 'normal.jpg',
contentType: 'image/jpeg',
});
expect(res.status).toBe(201);
const { id } = (await res.json()) as { id: string };
await expect.poll(() => db.compressionStatus(id), { timeout: 30_000 }).toBe('done');
});
});

View File

@@ -56,14 +56,21 @@ function upload(jwt: string, name: string) {
/**
* Pick a `quota_tolerance` that makes the per-user ceiling land on `targetBytes`.
* limit = floor(free_disk * tolerance / max(active, 1)) ⇒ tolerance = target * active / free.
*
* `staffJwt` reads the calibration inputs, `jwt` is the guest the limit is being aimed at.
* They must be different tokens: `free_disk_bytes` and `active_uploaders` are raw server
* telemetry and `/me/quota` zeroes both for non-staff (handlers/me.rs — "must never reach a
* guest"). Calibrating off the guest's own response divides by zero and yields a NaN
* tolerance, which is what silently broke this whole describe block.
*/
async function setLimitTo(
api: any,
adminToken: string,
staffJwt: string,
jwt: string,
targetBytes: number
): Promise<number> {
const q = await quotaOf(jwt);
const q = await quotaOf(staffJwt);
expect(
q.free_disk_bytes,
'the disk must be readable, else quota fails OPEN and proves nothing'
@@ -72,6 +79,7 @@ async function setLimitTo(
const tolerance = (targetBytes * active) / (q.free_disk_bytes as number);
await api.patchConfig(adminToken, { quota_tolerance: tolerance.toExponential(12) });
// Read back through the GUEST, whose ceiling is the one under test.
const after = await quotaOf(jwt);
expect(after.enabled).toBe(true);
return after.limit_bytes as number;
@@ -89,10 +97,11 @@ test.describe('Upload — storage quota enforcement', () => {
api,
adminToken,
guest,
host,
}) => {
const g = await guest('QuotaOver');
// Ceiling below one file: the very first upload must be refused.
const limit = await setLimitTo(api, adminToken, g.jwt, Math.floor(SIZE / 2));
const limit = await setLimitTo(api, adminToken, host.jwt, g.jwt, Math.floor(SIZE / 2));
expect(limit).toBeLessThan(SIZE);
const res = await upload(g.jwt, 'too-big.jpg');
@@ -111,9 +120,10 @@ test.describe('Upload — storage quota enforcement', () => {
api,
adminToken,
guest,
host,
}) => {
const g = await guest('QuotaUnder');
await setLimitTo(api, adminToken, g.jwt, SIZE * 4);
await setLimitTo(api, adminToken, host.jwt, g.jwt, SIZE * 4);
expect((await upload(g.jwt, 'fine.jpg')).status).toBe(201);
expect((await quotaOf(g.jwt)).used_bytes).toBe(SIZE);
@@ -123,11 +133,12 @@ test.describe('Upload — storage quota enforcement', () => {
api,
adminToken,
guest,
host,
}) => {
const g = await guest('QuotaRacer');
// Room for exactly ONE file.
const limit = await setLimitTo(api, adminToken, g.jwt, Math.floor(SIZE * 1.5));
const limit = await setLimitTo(api, adminToken, host.jwt, g.jwt, Math.floor(SIZE * 1.5));
expect(limit).toBeGreaterThanOrEqual(SIZE);
expect(limit).toBeLessThan(SIZE * 2);
@@ -170,10 +181,11 @@ test.describe('Upload — storage quota enforcement', () => {
api,
adminToken,
guest,
host,
}) => {
// Zero test hits before this — and it is the source of the "X von Y MB genutzt" widget.
const g = await guest('QuotaWidget');
await setLimitTo(api, adminToken, g.jwt, SIZE * 10);
await setLimitTo(api, adminToken, host.jwt, g.jwt, SIZE * 10);
const before = await quotaOf(g.jwt);
expect(before.enabled).toBe(true);

View File

@@ -0,0 +1,87 @@
/**
* Regression guard — a rejected upload must tell the user something.
*
* `UploadQueue.svelte` was 162 lines of complete, working UI — the only renderer of an
* item's error text, the only "Erneut" retry button, the only rate-limit countdown — and it
* was never imported anywhere, so `retryItem`, `removeItem` and `clearCompleted` were all
* unreachable at runtime. On a terminal rejection the store dropped the blob, wrote a clear
* German reason into `entry.error` with the comment "so the UI shows a clear reason", and
* there was no such UI.
*
* Meanwhile the FAB badge counted only pending/uploading, so a rejected photo decremented it
* exactly as if it had succeeded. Net effect: the photo silently vanished — no toast, no
* queue row, no error text, and it never appeared in the feed.
*/
import { test, expect } from '../../fixtures/test';
import { skipIfNoIdbBlobs } from '../../helpers/webkit';
import { FeedPage, UploadSheet } from '../../page-objects';
import { join } from 'node:path';
const SAMPLE_JPG = join(process.cwd(), 'fixtures', 'media', 'sample.jpg');
test.describe('Upload — a rejected upload is surfaced', () => {
test('a terminally rejected upload toasts, and stays visible in the queue', async ({
page,
api,
host,
guest,
signIn,
browserName,
}) => {
skipIfNoIdbBlobs(browserName);
const g = await guest('RejectedUploader');
await signIn(page, g);
const feed = new FeedPage(page);
const sheet = new UploadSheet(page);
await feed.openUploadSheet();
await sheet.stageFiles([SAMPLE_JPG]);
await sheet.captionInput.waitFor({ state: 'visible', timeout: 10_000 });
// Ban the uploader between staging and sending, so the POST comes back 403 — a
// terminal 4xx the server will keep rejecting, which is the path that purges the blob.
await api.banUser(host.jwt, g.userId);
await sheet.submit();
// 1. The user is told, wherever they are (the flow lands them on /feed).
const toast = page.getByRole('region', { name: 'Benachrichtigungen' });
await expect(toast).toContainText(/sample\.jpg/i, { timeout: 15_000 });
// 2. The queue row survives with its reason and is reachable on /upload — this is what
// the orphaned component made impossible.
await page.goto('/upload');
const queue = page.getByText('Upload-Warteschlange');
await expect(queue, 'the upload queue must be rendered somewhere').toBeVisible({
timeout: 10_000,
});
// Both the status chip ("Gesperrt") and the server's reason ("Du bist gesperrt.") must
// render — the reason is the part that had no UI at all before.
await expect(page.getByText('Gesperrt', { exact: true })).toBeVisible();
await expect(page.getByText('Du bist gesperrt.')).toBeVisible();
// 3. The badge must not read as success. It counted only pending/uploading before, so a
// rejected item dropped it to 0 — indistinguishable from a completed upload.
await expect
.poll(
() =>
page.evaluate(async () => {
return new Promise<number>((resolve, reject) => {
const req = indexedDB.open('eventsnap-uploads', 3);
req.onerror = () => reject(req.error);
req.onsuccess = () => {
const tx = req.result.transaction('queue', 'readonly');
const all = tx.objectStore('queue').getAll();
all.onsuccess = () =>
resolve(
all.result.filter((r: { status: string }) => r.status === 'blocked').length
);
all.onerror = () => reject(all.error);
};
});
}),
{ timeout: 10_000 }
)
.toBe(1);
});
});

View File

@@ -0,0 +1,136 @@
/**
* Regression guard — likes, comments and comment deletions are rate limited.
*
* These were the only mutating endpoints in the app with no limit at all. Every other write path
* -- upload, join, recover, export, admin login -- carried one; `social.rs` carried none, so the
* coverage was asymmetric rather than deliberately open.
*
* Severity is genuinely low for an invited-guest event, and the amplification worry is contained:
* a like does fan an SSE broadcast to every connected client, but the export regeneration a
* comment deletion triggers is debounced (REGEN_DEBOUNCE 20s) and superseded workers are inert. So
* this closes the gap for symmetry, and the ceiling is set well above anything a real guest
* produces -- it bounds a script, not an enthusiastic double-tapper.
*
* The bucket is shared across all three actions on purpose: separate buckets would let a caller
* triple the aggregate write rate just by alternating between them. That is what the second test
* pins, and it is the part most likely to be lost in a refactor.
*
* Keyed per USER, not per IP — at a venue every guest is behind one NAT, so an IP key would hand
* the whole party one bucket. Third test.
*/
import { test, expect } from '../../fixtures/test';
import { seedUpload } from '../../helpers/seed';
import { BASE } from '../../helpers/env';
const like = (jwt: string, uploadId: string) =>
fetch(`${BASE}/api/v1/upload/${uploadId}/like`, {
method: 'POST',
headers: { Authorization: `Bearer ${jwt}` },
});
const comment = (jwt: string, uploadId: string, body: string) =>
fetch(`${BASE}/api/v1/upload/${uploadId}/comments`, {
method: 'POST',
headers: { Authorization: `Bearer ${jwt}`, 'Content-Type': 'application/json' },
body: JSON.stringify({ body }),
});
test.describe('Social — rate limit', () => {
test('a burst of likes past the ceiling returns 429 with Retry-After', async ({
api,
adminToken,
guest,
}) => {
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
social_rate_enabled: 'true',
social_rate_per_min: '3',
});
const g = await guest('Tapper');
const uploadId = await seedUpload(g.jwt);
// Sequential, not parallel: a toggle flips state, so ordering matters for the assertion.
const statuses: number[] = [];
for (let i = 0; i < 5; i++) statuses.push((await like(g.jwt, uploadId)).status);
expect(statuses.slice(0, 3), 'the first three are within the ceiling').toEqual([200, 200, 200]);
expect(statuses.slice(3), 'everything past it is refused').toEqual([429, 429]);
const limited = await like(g.jwt, uploadId);
expect(limited.status).toBe(429);
expect(
limited.headers.get('retry-after'),
'a 429 without Retry-After tells the client nothing about when to come back'
).toBeTruthy();
});
test('likes and comments share one bucket', async ({ api, adminToken, guest }) => {
// THE assertion. Per-action buckets would let a caller triple the aggregate write rate by
// alternating, which defeats the point of having a ceiling at all.
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
social_rate_enabled: 'true',
social_rate_per_min: '2',
});
const g = await guest('Mixer');
const uploadId = await seedUpload(g.jwt);
expect((await like(g.jwt, uploadId)).status).toBe(200);
expect((await comment(g.jwt, uploadId, 'schön!')).status).toBe(201);
// Two writes spent, whichever endpoints they went to.
expect(
(await comment(g.jwt, uploadId, 'noch eins')).status,
'a comment must consume the same budget a like does'
).toBe(429);
expect((await like(g.jwt, uploadId)).status).toBe(429);
});
test('one guest hitting the ceiling does not block another', async ({
api,
adminToken,
guest,
}) => {
// Keyed per user, not per IP. Every request in this suite comes from one address, which is
// exactly the venue-NAT shape that made the /join and /feed limits turn guests away.
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
social_rate_enabled: 'true',
social_rate_per_min: '2',
});
const noisy = await guest('Noisy');
const quiet = await guest('Quiet');
const uploadId = await seedUpload(noisy.jwt);
for (let i = 0; i < 3; i++) await like(noisy.jwt, uploadId);
expect((await like(noisy.jwt, uploadId)).status).toBe(429);
expect(
(await like(quiet.jwt, uploadId)).status,
'a second guest behind the same IP must have their own budget'
).toBe(200);
});
test('flipping social_rate_enabled off bypasses the limit', async ({
api,
adminToken,
guest,
}) => {
// The toggle has to actually be honoured, or the admin switch is decorative — the failure
// mode two other per-area toggles already shipped with.
await api.patchConfig(adminToken, {
rate_limits_enabled: 'true',
social_rate_enabled: 'false',
social_rate_per_min: '2',
});
const g = await guest('Unlimited');
const uploadId = await seedUpload(g.jwt);
const statuses: number[] = [];
for (let i = 0; i < 6; i++) statuses.push((await like(g.jwt, uploadId)).status);
expect(statuses.every((s) => s === 200)).toBe(true);
});
});

View File

@@ -0,0 +1,181 @@
/**
* Regression guard — videos must actually play.
*
* Two independent defects made every video unplayable, and nothing in the suite covered
* either one (no test anywhere played media or asserted a `<video>` src):
*
* 1. The lightbox fed `<video>` the URL from `pickMediaUrl`, which is mime-agnostic.
* Compression only ever produces a *thumbnail* for a video — one ffmpeg frame — so in
* the DEFAULT saver mode the element's src was `/api/v1/upload/{id}/thumbnail`: a JPEG,
* served as image/jpeg with nosniff so the browser can't even sniff its way out.
* Chromium reported DEMUXER_ERROR_COULD_NOT_OPEN.
*
* 2. `stream_media_file` ignored `Range` entirely — always 200 with the whole body, never
* an Accept-Ranges or Content-Range. iOS Safari opens every `<video>` with a
* `Range: bytes=0-1` probe and abandons the load without a 206, so video failed on the
* app's primary platform even in `original` data mode.
*
* Fixing either alone still leaves video broken, so both are asserted here.
*/
import { test, expect } from '../../fixtures/test';
import { uploadRaw } from '../../helpers/upload-client';
import { BASE } from '../../helpers/env';
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
const SAMPLE_MP4 = join(process.cwd(), 'fixtures', 'media', 'sample.mp4');
async function seedVideo(jwt: string): Promise<string> {
const res = await uploadRaw(jwt, readFileSync(SAMPLE_MP4), {
filename: 'clip.mp4',
contentType: 'video/mp4',
caption: 'ein Video',
});
if (res.status !== 201) throw new Error(`video seed failed: ${res.status}`);
return ((await res.json()) as { id: string }).id;
}
test.describe('Video — the lightbox plays it', () => {
test('the <video> src is the original, not the thumbnail JPEG', async ({
page,
guest,
signIn,
db,
}) => {
const g = await guest('VideoWatcher');
const id = await seedVideo(g.jwt);
// The poster assertion below needs the ffmpeg thumbnail to EXIST — the lightbox binds
// `poster={upload.thumbnail_url ?? undefined}`, so the attribute is simply absent until
// compression finishes. Without this wait the test races the worker and fails against a
// cold stack (first run after `stack:down -v`, cold ffmpeg), which is exactly when a suite
// is least likely to be believed. The `src` assertion is unconditional; only the poster
// needs the wait.
await expect.poll(() => db.compressionStatus(id), { timeout: 60_000 }).toBe('done');
await signIn(page, g);
await page.goto('/feed');
const card = page.locator('article').first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Bild vergrößern' }).click();
const video = page.locator('video');
await expect(video).toBeVisible({ timeout: 10_000 });
// The bug in one assertion: this was `/thumbnail` in the default data mode.
await expect(video).toHaveAttribute('src', `/api/v1/upload/${id}/original`);
// The poster SHOULD still be the thumbnail — that's what it's for.
await expect(video).toHaveAttribute('poster', `/api/v1/upload/${id}/thumbnail`);
// And the browser must accept the bytes as media. preload="none" means nothing is
// fetched until we ask, so drive a load explicitly and wait for metadata.
const readyState = await video.evaluate(async (el: HTMLVideoElement) => {
el.preload = 'metadata';
el.load();
await new Promise<void>((resolve) => {
if (el.readyState > 0) return resolve();
el.addEventListener('loadedmetadata', () => resolve(), { once: true });
el.addEventListener('error', () => resolve(), { once: true });
setTimeout(resolve, 15_000);
});
return { ready: el.readyState, err: el.error?.message ?? null };
});
expect(readyState.err, `the browser rejected the media: ${readyState.err}`).toBeNull();
expect(readyState.ready, 'video metadata must load').toBeGreaterThan(0);
});
test('the video body is not downloaded before the user presses play', async ({
page,
guest,
signIn,
}) => {
// saver mode exists to protect a guest's mobile data, and there is no smaller video
// derivative to fall back to — so opening the lightbox must not pull the file down.
//
// Asserting "no request at all" would be wrong: WebKit opens a connection for a
// preload="none" <video> and immediately ABORTS it (observed: GET with no Range,
// response status 0, nothing transferred), whereas Chromium issues nothing. The
// portable guarantee — and the one that actually protects the data plan — is that no
// response carrying the body ever completes. Dropping preload="none" fails this.
const g = await guest('VideoThrifty');
const id = await seedVideo(g.jwt);
await signIn(page, g);
const delivered: string[] = [];
page.on('response', (r) => {
if (!r.url().includes(`/upload/${id}/original`)) return;
// status 0 = aborted before any bytes landed.
if (r.status() === 200 || r.status() === 206) {
delivered.push(`${r.status()} len=${r.headers()['content-length'] ?? '?'}`);
}
});
await page.goto('/feed');
const card = page.locator('article').first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Bild vergrößern' }).click();
await expect(page.locator('video')).toBeVisible({ timeout: 10_000 });
await page.waitForTimeout(2000);
expect(delivered, 'no video bytes may be delivered before play').toEqual([]);
});
});
test.describe('Media — HTTP Range', () => {
test('a range request is answered 206 with the right slice', async ({ guest }) => {
const g = await guest('RangeReader');
const id = await seedVideo(g.jwt);
const url = `${BASE}/api/v1/upload/${id}/original`;
const full = await fetch(url);
expect(full.status).toBe(200);
expect(full.headers.get('accept-ranges'), 'clients must be told seeking works').toBe('bytes');
const total = Number(full.headers.get('content-length'));
expect(total).toBeGreaterThan(0);
// The exact probe iOS Safari opens a <video> with. Two bytes, inclusive.
const probe = await fetch(url, { headers: { Range: 'bytes=0-1' } });
expect(probe.status, 'iOS abandons the load without a 206').toBe(206);
expect(probe.headers.get('content-range')).toBe(`bytes 0-1/${total}`);
expect(Number(probe.headers.get('content-length'))).toBe(2);
expect((await probe.arrayBuffer()).byteLength).toBe(2);
// A mid-file seek must return the matching slice, not the whole body.
const mid = await fetch(url, { headers: { Range: 'bytes=10-19' } });
expect(mid.status).toBe(206);
expect(mid.headers.get('content-range')).toBe(`bytes 10-19/${total}`);
const midBytes = Buffer.from(await mid.arrayBuffer());
expect(midBytes.byteLength).toBe(10);
expect(midBytes).toEqual(Buffer.from(await full.arrayBuffer()).subarray(10, 20));
// Open-ended range runs to EOF.
const tail = await fetch(url, { headers: { Range: `bytes=${total - 5}-` } });
expect(tail.status).toBe(206);
expect(Number(tail.headers.get('content-length'))).toBe(5);
// Past EOF must be 416 — answering 200 makes a player re-request forever.
const bad = await fetch(url, { headers: { Range: `bytes=${total + 10}-` } });
expect(bad.status).toBe(416);
expect(bad.headers.get('content-range')).toBe(`bytes */${total}`);
});
test('image derivatives are range-capable too', async ({ guest, db }) => {
// Same helper serves all four media routes, so the guarantee is uniform.
const g = await guest('RangeImages');
const res = await uploadRaw(
g.jwt,
readFileSync(join(process.cwd(), 'fixtures', 'media', 'sample.jpg')),
{ filename: 'r.jpg', contentType: 'image/jpeg' }
);
const { id } = (await res.json()) as { id: string };
await expect.poll(() => db.compressionStatus(id), { timeout: 30_000 }).toBe('done');
const preview = await fetch(`${BASE}/api/v1/upload/${id}/preview`, {
headers: { Range: 'bytes=0-9' },
});
expect(preview.status).toBe(206);
expect(Number(preview.headers.get('content-length'))).toBe(10);
});
});

View File

@@ -0,0 +1,97 @@
/**
* Regression guard — the host is warned about storage BEFORE it becomes unrecoverable.
*
* Storage visibility used to exist in exactly one place: a passive "Speicherauslastung" widget on
* the ADMIN dashboard. A host who isn't the admin had no view of it at all, and nothing anywhere
* warned anyone. README listed a low-disk alert under "Planned (v1.x)".
*
* Two things make that a safety net rather than a nice-to-have:
*
* - `postgres_data`, `media_data` and `exports_data` are all Docker named volumes on ONE
* filesystem. A full disk doesn't degrade a subsystem; Postgres stops being able to write and
* the whole event goes down.
* - The keepsake needs room for TWO gallery-sized archives (both write their media
* `Compression::Stored`; `Memories.zip` streams the original for every video and every image
* at or under 5 MB). The export preflight can refuse cleanly, but only AFTER the release —
* when the event is over, the gallery is full, and every remedy is harder.
*
* So the threshold is deliberately NOT a fixed number alone. It fires on an absolute floor (10 GB,
* the figure the README always carried) OR on "you could not build the keepsake right now", which
* is the trigger a host can still act on.
*
* These drive it through `original_size_bytes` rather than a genuinely full disk: the estimate is
* pure SQL over that column, so overstating one row moves the accounting the warning reads without
* touching a byte on disk.
*/
import { test, expect } from '../../fixtures/test';
import { seedUpload } from '../../helpers/seed';
import { BASE } from '../../helpers/env';
/** Comfortably larger than any disk this suite could run on. */
const ABSURD_BYTES = 500_000_000_000_000;
test.describe('Host — low-disk warning', () => {
test('a gallery too big to export warns the host, with the numbers', async ({
page,
host,
guest,
signIn,
db,
}) => {
const g = await guest('BigShooter');
const uploadId = await seedUpload(g.jwt);
await db.setUploadSizeBytes(uploadId, ABSURD_BYTES);
await signIn(page, host);
await page.goto('/host');
const warning = page.getByTestId('low-disk-warning');
await expect(warning, 'the host must be warned before releasing').toBeVisible({
timeout: 15_000,
});
// The actionable half: not just "low", but "the keepsake cannot be built".
await expect(warning).toContainText(/nicht.*erstellt werden/i);
// And the consequence that makes it urgent — the event, not just the download.
await expect(warning).toContainText(/gesamte Event/i);
});
test('the API reports the requirement and the verdict together', async ({ host, guest, db }) => {
const g = await guest('BigShooter2');
const uploadId = await seedUpload(g.jwt);
await db.setUploadSizeBytes(uploadId, ABSURD_BYTES);
const res = await fetch(`${BASE}/api/v1/host/event`, {
headers: { Authorization: `Bearer ${host.jwt}` },
});
expect(res.status).toBe(200);
const body = (await res.json()) as {
disk_low: boolean;
disk_free_bytes: number | null;
keepsake_required_bytes: number;
};
expect(body.disk_low).toBe(true);
expect(
body.keepsake_required_bytes,
'both halves are armed by a release, so the requirement covers two archives'
).toBeGreaterThan(ABSURD_BYTES);
expect(body.disk_free_bytes).not.toBeNull();
expect(body.keepsake_required_bytes).toBeGreaterThan(body.disk_free_bytes!);
});
test('an ordinary gallery shows no warning at all', async ({ page, host, guest, signIn }) => {
// The mirror that keeps the above honest. A warning that is always on is a warning nobody
// reads — and it would sit at the very top of the dashboard, above the PIN-reset queue.
const g = await guest('NormalShooter');
await seedUpload(g.jwt);
await signIn(page, host);
await page.goto('/host');
// Wait for the dashboard to actually be loaded before asserting on an absence.
await expect(page.getByRole('heading', { name: 'Host-Dashboard' })).toBeVisible({
timeout: 15_000,
});
await expect(page.getByTestId('low-disk-warning')).toHaveCount(0);
});
});

View File

@@ -0,0 +1,149 @@
/**
* Regression guard — a host must be able to remove a guest's content FROM THE UI.
*
* `DELETE /host/upload/{id}` and `DELETE /host/comment/{id}` were fully implemented,
* transactional, SSE-broadcasting, audit-logged — and had zero frontend callers. The feed's
* context sheet offered "Löschen" only when `target.user_id === myUserId`, so the only lever
* a host actually had against an unwanted photo was banning the uploader. That is both
* disproportionate and ineffective: a ban doesn't retract what was already posted, and
* because the ban check runs BEFORE the ownership check on the guest delete route, banning
* the author makes their abusive comment permanently undeletable by them too.
*
* The API side was already covered (04-host/moderation). What was missing is the wiring,
* so these tests drive the real UI.
*/
import { test, expect } from '../../fixtures/test';
import { seedUpload, seedComment } from '../../helpers/seed';
import { BASE } from '../../helpers/env';
test.describe('Host — moderation from the UI', () => {
test("a host removes a guest's photo via the feed context sheet", async ({
page,
host,
guest,
signIn,
}) => {
const g = await guest('PhotoOffender');
const uploadId = await seedUpload(g.jwt);
await signIn(page, host);
await page.goto('/feed');
const card = page.locator('article').filter({ hasText: g.displayName }).first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Mehr Aktionen' }).click();
const remove = page.getByRole('button', { name: /beitrag entfernen/i });
await expect(remove, 'a host must be offered a removal action on a guest post').toBeVisible();
await remove.click();
const sheet = page.getByTestId('confirm-sheet');
await expect(sheet).toBeVisible();
// Moderation copy, not "delete my post" copy.
await expect(sheet).toContainText(/beitrag entfernen/i);
await page.getByTestId('confirm-sheet-confirm').click();
await expect(card).not.toBeVisible({ timeout: 10_000 });
// And it is really gone server-side, not just dropped from the local list.
const res = await fetch(`${BASE}/api/v1/feed`, {
headers: { Authorization: `Bearer ${host.jwt}` },
});
const body = await res.json();
expect(body.uploads.some((u: { id: string }) => u.id === uploadId)).toBe(false);
});
test('a guest is NOT offered any delete action on someone elses post', async ({
page,
guest,
signIn,
}) => {
// The mirror that makes the test above meaningful: if this affordance rendered for
// everyone, the host test would still pass on a build that shipped moderation to guests.
const author = await guest('SomeAuthor');
await seedUpload(author.jwt);
const viewer = await guest('NosyViewer');
await signIn(page, viewer);
await page.goto('/feed');
const card = page.locator('article').filter({ hasText: author.displayName }).first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Mehr Aktionen' }).click();
await expect(page.getByRole('button', { name: /beitrag entfernen/i })).toHaveCount(0);
await expect(page.getByRole('button', { name: /^löschen$/i })).toHaveCount(0);
});
test('a promoted guest gets host powers without signing out and back in', async ({
page,
api,
adminToken,
guest,
signIn,
}) => {
// The JWT is never reissued — the backend slides the session row forward and treats the
// DB row as authoritative. So the token of a promoted guest still claims `role: guest`
// for up to 30 days. The UI read that frozen claim, which meant a guest promoted at the
// party saw no Host-Dashboard and no moderation actions until they signed out and back
// in — while `/me/context` had been handing the client the real role all along.
const g = await guest('LatePromotion');
await signIn(page, g);
await page.goto('/account');
await expect(page.getByRole('link', { name: /host-dashboard/i })).toHaveCount(0);
// Promote mid-session. The token in localStorage is deliberately NOT refreshed.
await api.setRole(adminToken, g.userId, 'host');
const claim = JSON.parse(Buffer.from(g.jwt.split('.')[1], 'base64').toString());
expect(
claim.role,
'the token must still carry the stale claim for this to prove anything'
).toBe('guest');
await page.reload();
await expect(
page.getByRole('link', { name: /host-dashboard/i }),
'the live role from /me/context must win over the frozen JWT claim'
).toBeVisible({ timeout: 10_000 });
});
test('a host can remove the comment of a guest they have already banned', async ({
page,
api,
host,
guest,
signIn,
}) => {
// The deadlock this closes. Ban first, exactly as a host would react to abuse: from then
// on the author gets 403 on their own delete, so if the host has no removal affordance
// the comment is stuck on screen forever.
// The photo belongs to an innocent third party — a ban hides the banned user's OWN
// uploads, so if the comment sat on their own photo the whole card would vanish and
// there would be nothing left to moderate.
const victim = await guest('PhotoOwner');
const uploadId = await seedUpload(victim.jwt);
const author = await guest('CommentOffender');
const commentId = await seedComment(author.jwt, uploadId, 'unangebrachter Kommentar');
await api.banUser(host.jwt, author.userId);
// Confirm the deadlock really exists — the author cannot retract it themselves.
const selfDelete = await fetch(`${BASE}/api/v1/comment/${commentId}`, {
method: 'DELETE',
headers: { Authorization: `Bearer ${author.jwt}` },
});
expect(selfDelete.status, 'a banned author is blocked from their own delete').toBe(403);
await signIn(page, host);
await page.goto('/feed');
const card = page.locator('article').filter({ hasText: victim.displayName }).first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Bild vergrößern' }).click();
const comment = page.getByText('unangebrachter Kommentar');
await expect(comment).toBeVisible({ timeout: 10_000 });
await page.getByRole('button', { name: 'Kommentar entfernen' }).first().click();
await expect(comment).toHaveCount(0, { timeout: 10_000 });
});
});

View File

@@ -0,0 +1,102 @@
/**
* Regression guard — the role must follow the identity, not the tab.
*
* The `role` store is a module-level singleton seeded ONCE at import. `goto()` is a
* client-side navigation, so leaving and re-joining in the same tab re-imports nothing and
* re-runs no `onMount` — the previous user's role simply stayed. A host who left and a
* guest who then joined kept `isStaff === true` and were offered "🚫 Beitrag entfernen" on
* other people's photos. The backend 403s the delete, so it was a false affordance rather
* than a privilege escalation, but `/feed` never fetched `/me/context`, so it never
* self-corrected either — it survived until a hard reload.
*
* The mirror case matters just as much and is easier to forget: a guest who recovers into a
* host account must GAIN the affordance without a reload.
*/
import { test, expect } from '../../fixtures/test';
import { seedUpload } from '../../helpers/seed';
import { JoinPage } from '../../page-objects';
const REMOVE = /beitrag entfernen/i;
test.describe('Role — follows the identity across a same-tab switch', () => {
test('a guest joining after a host leaves does NOT inherit host actions', async ({
page,
host,
guest,
signIn,
}) => {
// Someone else's photo — the only kind the removal action is offered on.
const author = await guest('RoleAuthor');
await seedUpload(author.jwt);
// 1. Host is signed in and DOES see the moderation action. Establishing this first is
// what makes the negative assertion below meaningful.
await signIn(page, host);
await page.goto('/feed');
const card = page.locator('article').filter({ hasText: author.displayName }).first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Mehr Aktionen' }).click();
await expect(page.getByRole('button', { name: REMOVE })).toBeVisible();
await page.keyboard.press('Escape');
// 2. Host leaves, in-app — no reload. This is the path "Event verlassen" takes.
await page.goto('/account');
await page.getByRole('button', { name: /event verlassen/i }).click();
const confirm = page.getByTestId('confirm-sheet-confirm');
if (await confirm.isVisible().catch(() => false)) await confirm.click();
await page.waitForURL('**/join', { timeout: 10_000 });
// 3. A brand-new guest joins in the same tab — the real flow, PIN modal and all.
const join = new JoinPage(page);
await join.joinAs(`Nachzuegler${Date.now() % 100000}`);
await join.continueToFeed();
await expect(page).toHaveURL(/\/feed$/, { timeout: 15_000 });
// 4. They must NOT be offered moderation on someone else's photo.
const card2 = page.locator('article').filter({ hasText: author.displayName }).first();
await expect(card2).toBeVisible({ timeout: 15_000 });
await card2.getByRole('button', { name: 'Mehr Aktionen' }).click();
await expect(
page.getByRole('button', { name: REMOVE }),
'a fresh guest must not inherit the previous users role'
).toHaveCount(0);
});
test('a guest who recovers into a host account GAINS host actions without a reload', async ({
page,
api,
adminToken,
guest,
signIn,
}) => {
// The mirror. If the fix only cleared the role it would pass the test above and still
// leave a real host with no moderation until they reloaded.
const author = await guest('RoleAuthor2');
await seedUpload(author.jwt);
const futureHost = await guest('WillBeHost');
await signIn(page, futureHost);
await page.goto('/feed');
const card = page.locator('article').filter({ hasText: author.displayName }).first();
await expect(card).toBeVisible({ timeout: 15_000 });
await card.getByRole('button', { name: 'Mehr Aktionen' }).click();
await expect(page.getByRole('button', { name: REMOVE })).toHaveCount(0);
await page.keyboard.press('Escape');
// Promote them server-side. Their resident JWT still claims `role: guest`.
await api.setRole(adminToken, futureHost.userId, 'host');
const claim = JSON.parse(Buffer.from(futureHost.jwt.split('.')[1], 'base64').toString());
expect(claim.role, 'the token must still be stale for this to prove anything').toBe('guest');
// A plain in-app navigation back to the feed must pick up the live role.
await page.goto('/account');
await page.goto('/feed');
const card2 = page.locator('article').filter({ hasText: author.displayName }).first();
await expect(card2).toBeVisible({ timeout: 15_000 });
await card2.getByRole('button', { name: 'Mehr Aktionen' }).click();
await expect(
page.getByRole('button', { name: REMOVE }),
'the live role from /me/context must reach the feed'
).toBeVisible({ timeout: 10_000 });
});
});

View File

@@ -68,6 +68,40 @@ test.describe('Admin — config API', () => {
expect(cfg.privacy_note).toBe(note);
await api.patchConfig(adminToken, { privacy_note: '' });
});
test('quota_tolerance = 0 is rejected, with a pointer to the real off-switch', async ({
api,
adminToken,
}) => {
// Zero is inside the documented 01 range and catastrophic: the per-user limit is
// `free_disk * tolerance / active_uploaders`, so 0 refuses EVERY upload — mid-event, with
// "Du hast dein Upload-Limit für dieses Event erreicht", which names the wrong cause
// entirely. `storage_quota_enabled` is what an admin reaching for an off-switch wants.
const res = await fetch(
(process.env.E2E_FRONTEND_URL ?? 'http://localhost:3101') + '/api/v1/admin/config',
{
method: 'PATCH',
headers: { Authorization: `Bearer ${adminToken}`, 'Content-Type': 'application/json' },
body: JSON.stringify({ quota_tolerance: '0' }),
}
);
expect(res.status).toBe(400);
expect(
(await res.text()).toLowerCase(),
'the error must name the switch the admin actually wanted'
).toContain('speicher-quote');
// The value is untouched — validation fully precedes any write.
expect((await api.getConfig(adminToken)).quota_tolerance).toBe('0.75');
});
test('a very small quota_tolerance is still accepted', async ({ api, adminToken }) => {
// The mirror. Rejecting 0 must not become a floor: small tolerances are how a large disk is
// throttled to a sensible per-guest ceiling, and how the quota specs steer it (~1e-5 on a
// 174 GB volume). A floor of 0.01 would forbid real configurations to prevent one typo.
await api.patchConfig(adminToken, { quota_tolerance: '0.00001' });
expect((await api.getConfig(adminToken)).quota_tolerance).toBe('0.00001');
await api.patchConfig(adminToken, { quota_tolerance: '0.75' });
});
});
test.describe('Admin — stats', () => {

View File

@@ -0,0 +1,103 @@
/**
* Regression guard — the keepsake extracts to files a guest can actually open.
*
* `ZipEntryBuilder::new` leaves the external file attribute at zero, and async_zip's host
* compatibility defaults to Unix — so every entry in BOTH archives was written with a stored mode
* of 0000. `unzip -Z` showed `?---------` on every line.
*
* Windows Explorer ignores Unix modes, which is exactly why this survived. On Linux and macOS,
* `unzip` faithfully applies what the archive asks for, and the guest gets a folder of photos none
* of which they can open — plus an index.html the browser refuses with ERR_ACCESS_DENIED.
*
* Unconditional: it affected every keepsake ever produced, no hostile input required. And it is
* invisible server-side — the export succeeds, the ZIP is well-formed, the job writes `done`,
* /export/status is green. The only way to see it is to extract the real artifact and try to read
* it, which is what this does.
*
* Found while chasing an unrelated ERR_ACCESS_DENIED that looked like a Playwright quirk.
*/
import { test, expect } from '../../fixtures/test';
import { execFileSync } from 'node:child_process';
import { mkdtempSync, writeFileSync, rmSync, readFileSync, statSync, readdirSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { seedUpload } from '../../helpers/seed';
import { BASE } from '../../helpers/env';
/** Walk every file under `dir`, ignoring the archives we dropped there ourselves. */
function walk(dir: string, skip: string[] = []): string[] {
return readdirSync(dir, { withFileTypes: true }).flatMap((e) => {
const p = join(dir, e.name);
if (e.isDirectory()) return walk(p, skip);
return skip.includes(e.name) ? [] : [p];
});
}
test.describe('Export — the archives extract to readable files', () => {
test('every entry in both keepsake archives is owner-readable', async ({ host, guest, db }) => {
test.setTimeout(120_000);
const bearer = { Authorization: `Bearer ${host.jwt}` };
const g = await guest('Archivist');
const id = await seedUpload(g.jwt, { caption: 'ein Foto' });
await expect.poll(() => db.compressionStatus(id), { timeout: 30_000 }).toBe('done');
expect(
(await fetch(`${BASE}/api/v1/host/gallery/release`, { method: 'POST', headers: bearer }))
.status
).toBe(204);
await expect
.poll(
async () => {
const res = await fetch(`${BASE}/api/v1/export/status`, { headers: bearer });
const s = await res.json();
return s.zip?.status === 'done' && s.html?.status === 'done';
},
{ timeout: 90_000, intervals: [500] }
)
.toBe(true);
for (const kind of ['zip', 'html'] as const) {
const ticketRes = await fetch(`${BASE}/api/v1/export/ticket`, {
method: 'POST',
headers: bearer,
});
const { ticket } = (await ticketRes.json()) as { ticket: string };
const dl = await fetch(`${BASE}/api/v1/export/${kind}?ticket=${encodeURIComponent(ticket)}`);
expect(dl.status, `downloading the ${kind} archive`).toBe(200);
const dir = mkdtempSync(join(tmpdir(), `eventsnap-perms-${kind}-`));
try {
const zipPath = join(dir, 'archive.zip');
writeFileSync(zipPath, Buffer.from(await dl.arrayBuffer()));
// The mode as STORED in the archive — this is what a guest's unzip will apply. Reading it
// from the central directory catches the defect even on a filesystem that would mask it.
const listing = execFileSync('unzip', ['-Z', zipPath], { encoding: 'utf8' });
const modes = listing
.split('\n')
.filter((l) => /^[?d-][rwx-]{9}\s/.test(l))
.map((l) => l.slice(0, 10));
expect(modes.length, `${kind}: no entries listed`).toBeGreaterThan(0);
for (const m of modes) {
expect(m, `${kind}: an entry is stored mode ${m} — the guest cannot open it`).toMatch(
/^.r[w-]-/
);
}
// And extraction really does produce readable files.
execFileSync('unzip', ['-qo', zipPath, '-d', dir]);
const files = walk(dir, ['archive.zip']);
expect(files.length, `${kind}: nothing extracted`).toBeGreaterThan(0);
for (const f of files) {
expect(statSync(f).mode & 0o400, `${f} is not owner-readable`).toBeTruthy();
// The assertion that matters to a guest: the bytes actually come out.
expect(() => readFileSync(f)).not.toThrow();
}
} finally {
rmSync(dir, { recursive: true, force: true });
}
}
});
});

View File

@@ -0,0 +1,103 @@
/**
* Regression guard — the keepsake download must actually download, in WebKit.
*
* The bug this exists to catch: `/export` streams the archive by pointing a HIDDEN,
* SAME-ORIGIN iframe at `/api/v1/export/zip` (deliberately — a top-level navigation to a
* 404/429 would unload the PWA). Caddy stamped a site-wide `X-Frame-Options: DENY` that
* also covered `/api/*`. Blink hands a `Content-Disposition: attachment` response to the
* download manager at the network layer, so Chromium never noticed; WebKit enforces XFO on
* the frame navigation FIRST and aborts the load. Result: on iOS Safari — the app's primary
* platform — tapping Download did nothing, silently, with no error anywhere.
*
* Why the old suite was structurally blind to it:
* - `06-export` ran on `chromium-desktop` only (webkit-iphone's testMatch excluded it),
* - and no test in the entire suite ever CLICKED a download button; every archive
* assertion used Node `fetch`, which has no frame and therefore no XFO enforcement.
*
* So this spec must keep both properties to be worth anything: a real click, in WebKit.
*/
import { test, expect } from '../../fixtures/test';
import { ExportPage } from '../../page-objects';
import { BASE } from '../../helpers/env';
const SLUG = 'e2e-test-event';
function post(path: string, jwt: string) {
return fetch(BASE + path, { method: 'POST', headers: { Authorization: `Bearer ${jwt}` } });
}
/** The real export job runs image processing; give it head-room over the tiny fixtures. */
async function releaseAndWait(jwt: string) {
expect((await post('/api/v1/host/gallery/release', jwt)).status).toBe(204);
await expect
.poll(
async () => {
const res = await fetch(BASE + '/api/v1/export/status', {
headers: { Authorization: `Bearer ${jwt}` },
});
const s = await res.json();
return s.released === true && s.zip?.status === 'done';
},
{ timeout: 60_000, intervals: [500] }
)
.toBe(true);
}
test.describe('Export — the download actually fires in the browser', () => {
test.slow();
test('clicking Download triggers a real download event', async ({ page, host, signIn, db }) => {
await db.setExportReleased(SLUG, false);
await releaseAndWait(host.jwt);
await signIn(page, host);
const exportPage = new ExportPage(page);
await exportPage.goto();
await expect(exportPage.zipDownloadButton).toBeEnabled({ timeout: 10_000 });
// Capture the frame-level refusal that XFO produces, so a failure reports the CAUSE
// rather than just a timeout. WebKit logs "Refused to display ... in a frame because it
// set 'X-Frame-Options'"; Chromium logs nothing here, which is the whole problem.
const refusals: string[] = [];
page.on('console', (m) => {
if (/X-Frame-Options|Refused to display/i.test(m.text())) refusals.push(m.text());
});
const downloadPromise = page.waitForEvent('download', { timeout: 30_000 });
await exportPage.zipDownloadButton.click();
const download = await downloadPromise.catch((err) => {
throw new Error(
`No download event fired after clicking the ZIP button.` +
(refusals.length
? ` The browser refused the iframe navigation: ${refusals.join(' | ')}`
: ' No X-Frame-Options refusal was logged; check the ticket/readiness path.') +
`\n${err}`
);
});
expect(download.suggestedFilename()).toMatch(/\.zip$/i);
// The stream must produce real bytes, not a zero-length placeholder.
const path = await download.path();
expect(path).toBeTruthy();
expect(refusals, 'no frame should have been refused').toEqual([]);
});
test('the export endpoints are framable same-origin; everything else stays DENY', async () => {
// Locks the Caddy carve-out itself, independently of any browser. Cheap, and it fails
// loudly at the exact layer that regressed if someone reinstates a blanket DENY.
for (const path of ['/api/v1/export/zip', '/api/v1/export/html']) {
const res = await fetch(BASE + path);
expect(res.headers.get('x-frame-options')?.toUpperCase(), `${path} must be framable`).toBe(
'SAMEORIGIN'
);
}
for (const path of ['/', '/api/v1/feed', '/api/v1/event']) {
const res = await fetch(BASE + path);
expect(res.headers.get('x-frame-options')?.toUpperCase(), `${path} must stay DENY`).toBe(
'DENY'
);
}
});
});

View File

@@ -0,0 +1,117 @@
/**
* Regression guard — the keepsake must not be sideways.
*
* Round 1 taught the compression worker to apply EXIF orientation, which fixed the live app
* (feed preview + diashow display). The export worker was missed: it does NOT reuse those
* derivatives — it re-decodes the originals itself with `image::open`, which ignores the
* orientation tag — and then re-encodes to JPEG, which drops the tag, so the viewer has no
* way to recover it.
*
* The resulting damage was oddly shaped, which is what made it read as a viewer bug:
* - Gallery.zip originals → correct (byte-copied, EXIF intact)
* - Memories viewer grid thumbnails → ALWAYS sideways
* - Memories viewer full image >5 MB → sideways (re-encoded at 2000px)
* - Memories viewer full image ≤5 MB → correct (streamed byte-for-byte)
*
* So clicking a small photo silently "fixed" it and a large one didn't. This pins the
* thumbnail, which is the path every photo takes.
*/
import { test, expect } from '../../fixtures/test';
import { uploadRaw } from '../../helpers/upload-client';
import { BASE } from '../../helpers/env';
import { execFileSync } from 'node:child_process';
import { mkdtempSync, readFileSync, writeFileSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
// 40x20 landscape pixels tagged Orientation=6 ("rotate 90° CW to display"), so anything
// that honours the tag emits a PORTRAIT derivative.
const EXIF_FIXTURE = join(process.cwd(), 'fixtures', 'media', 'portrait-exif6.jpg');
/** Pixel dimensions from a JPEG's SOF marker — avoids an image dep for one assertion. */
function jpegSize(buf: Buffer): { width: number; height: number } {
let i = 2;
while (i < buf.length) {
if (buf[i] !== 0xff) {
i++;
continue;
}
const marker = buf[i + 1];
if (marker >= 0xc0 && marker <= 0xcf && marker !== 0xc4 && marker !== 0xc8 && marker !== 0xcc) {
return { height: buf.readUInt16BE(i + 5), width: buf.readUInt16BE(i + 7) };
}
i += 2 + buf.readUInt16BE(i + 2);
}
throw new Error('no SOF marker found — not a JPEG?');
}
test.describe('Export — EXIF orientation in the keepsake', () => {
test('the Memories viewer thumbnail of a rotated photo is upright', async ({ host, db }) => {
test.setTimeout(90_000);
const bearer = { Authorization: `Bearer ${host.jwt}` };
const src = readFileSync(EXIF_FIXTURE);
// Sanity: the SOURCE really is stored landscape, or this test proves nothing.
const srcSize = jpegSize(src);
expect(srcSize.width).toBeGreaterThan(srcSize.height);
const up = await uploadRaw(host.jwt, src, {
filename: 'hochkant.jpg',
contentType: 'image/jpeg',
caption: 'hochkant',
});
expect(up.status).toBe(201);
const { id } = (await up.json()) as { id: string };
await expect.poll(() => db.compressionStatus(id), { timeout: 30_000 }).toBe('done');
const rel = await fetch(`${BASE}/api/v1/host/gallery/release`, {
method: 'POST',
headers: bearer,
});
expect(rel.status).toBe(204);
await expect
.poll(
async () => {
const res = await fetch(`${BASE}/api/v1/export/status`, { headers: bearer });
return (await res.json()).html?.status;
},
{ timeout: 60_000, intervals: [500] }
)
.toBe('done');
const ticketRes = await fetch(`${BASE}/api/v1/export/ticket`, {
method: 'POST',
headers: bearer,
});
const { ticket } = (await ticketRes.json()) as { ticket: string };
const dl = await fetch(`${BASE}/api/v1/export/html?ticket=${encodeURIComponent(ticket)}`);
expect(dl.status).toBe(200);
const dir = mkdtempSync(join(tmpdir(), 'eventsnap-exif-'));
try {
const zipPath = join(dir, 'Memories.zip');
writeFileSync(zipPath, Buffer.from(await dl.arrayBuffer()));
const entries = execFileSync('unzip', ['-Z1', zipPath], { encoding: 'utf8' })
.split('\n')
.filter(Boolean);
const thumbEntry = entries.find((e) => e.includes(`${id}_thumb`));
expect(thumbEntry, `no thumbnail for ${id} in Memories.zip`).toBeTruthy();
// `-p` streams the entry to stdout. Extracting to disk instead fails with EACCES:
// the archive preserves the container's file mode, which the test user can't read.
const thumb = execFileSync('unzip', ['-p', zipPath, thumbEntry!], {
maxBuffer: 64 * 1024 * 1024,
});
const { width, height } = jpegSize(thumb);
expect(
height,
`the keepsake grid thumbnail is ${width}x${height} — EXIF orientation was not applied`
).toBeGreaterThan(width);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
});

View File

@@ -0,0 +1,95 @@
/**
* Regression guard — when the keepsake fails to build, the HOST must be told why.
*
* `/export/status` reported `{status, progress_pct}` and nothing else, so the host dashboard could
* only ever render "Keepsake-Erstellung fehlgeschlagen." next to an "Erneut versuchen" button. The
* reason WAS being written — `mark_failed` stores it on the job row — but it surfaced solely in the
* admin dashboard's job list. The host is the person who releases the gallery, owns the retry
* button, and is standing at the venue; the admin may be someone else entirely, or the same person
* without the password to hand.
*
* That matters most for the failure this shipped alongside: the export disk preflight. Its message
* names the two numbers that decide what to do ("benötigt ca. X GB, frei sind Y GB"), and without
* it "Erneut versuchen" fails identically, forever, with no hint that the answer is free some space.
*
* These drive the real UI and the real endpoint — the plumbing is four hops (SQL → handler JSON →
* store type → Svelte branch) and any one of them dropping the field restores the silent version.
*/
import { test, expect } from '../../fixtures/test';
import { BASE } from '../../helpers/env';
const SLUG = 'e2e-test-event';
const DISK_REASON =
'Nicht genug Speicherplatz für das Keepsake: benötigt ca. 42.0 GB, frei sind 3.0 GB. ' +
'Bitte Speicher freigeben und das Keepsake anschließend neu erstellen.';
test.describe('Export — a failed keepsake explains itself to the host', () => {
test('the failure reason reaches /export/status', async ({ host, db }) => {
await db.setExportReleased(SLUG, true);
await db.fakeExportJob(SLUG, 'zip', 'failed', DISK_REASON);
await db.fakeExportJob(SLUG, 'html', 'failed', DISK_REASON);
const res = await fetch(`${BASE}/api/v1/export/status`, {
headers: { Authorization: `Bearer ${host.jwt}` },
});
expect(res.status).toBe(200);
const body = (await res.json()) as {
zip: { status: string; error_message: string | null };
html: { status: string; error_message: string | null };
};
expect(body.zip.status).toBe('failed');
expect(
body.zip.error_message,
'the reason must travel with the status, not live only in the admin job list'
).toBe(DISK_REASON);
expect(body.html.error_message).toBe(DISK_REASON);
});
test('a succeeding export carries no stale reason', async ({ host, db }) => {
// The mirror that keeps the above honest: a handler that returned `error_message`
// unconditionally would pass the first test while showing an error next to a green
// "Keepsake ist bereit." Rows keep their last message until they are re-armed, so this is a
// real state, not a hypothetical one.
await db.setExportReleased(SLUG, true);
await db.fakeExportJob(SLUG, 'zip', 'failed', DISK_REASON);
await db.fakeExportJob(SLUG, 'zip', 'done', DISK_REASON);
const res = await fetch(`${BASE}/api/v1/export/status`, {
headers: { Authorization: `Bearer ${host.jwt}` },
});
const body = (await res.json()) as { zip: { status: string; error_message: string | null } };
expect(body.zip.status).toBe('done');
expect(
body.zip.error_message,
'a message left on a row that has since succeeded must not be shown'
).toBeNull();
});
test('the host dashboard renders the reason under the failure', async ({
page,
host,
signIn,
db,
}) => {
await db.setExportReleased(SLUG, true);
await db.fakeExportJob(SLUG, 'zip', 'failed', DISK_REASON);
await db.fakeExportJob(SLUG, 'html', 'failed', DISK_REASON);
await signIn(page, host);
await page.goto('/host');
await expect(page.getByText(/Keepsake-Erstellung fehlgeschlagen/i)).toBeVisible({
timeout: 15_000,
});
// The actionable half — the numbers, not just the verdict.
await expect(
page.getByText(/Nicht genug Speicherplatz/i),
'the host must see WHY, next to the only button they have'
).toBeVisible();
await expect(page.getByText(/3\.0 GB/)).toBeVisible();
// And the retry button is still mounted — it is deliberately outside the status branches.
await expect(page.getByTestId('export-rebuild')).toBeVisible();
});
});

View File

@@ -0,0 +1,125 @@
/**
* Regression guard — a guest-authored caption cannot brick the offline keepsake.
*
* The viewer's data is inlined as `<script>window.__EXPORT_DATA__={…};</script>` (it must be:
* guests open index.html over file://, where a cross-origin fetch of a sibling data.json is
* blocked). Captions and comments are guest text and land in that payload.
*
* The escape used to be `</` → `<\/`. Against XSS that holds — `</script><img src=x onerror=…>`
* round-trips inert. It does NOT stop the caption steering the HTML TOKENIZER: `<!--<script` with
* no later `-->` drives the parser into script-data-double-escaped state, where the template's own
* `</script>` steps back to script-data-escaped instead of closing the element. Everything after —
* including the viewer bundle — is swallowed as script data. Nothing executes and nothing leaks;
* `__EXPORT_DATA__` is never assigned and the keepsake renders blank.
*
* What makes it worth a browser-level test rather than a unit test alone: the failure is SILENT and
* POST-DISTRIBUTION. The export succeeds, the ZIP is well-formed, the job writes `done`,
* /export/status is green, and the host hands out a file that only fails when a guest
* double-clicks it — in every copy, unfixably. It is not visible by reading the escape. It is only
* visible by running a real parser over the real artifact, which is what this does: release, pull
* the actual Memories.zip, extract index.html, open it over file:// in Chromium, and assert the
* viewer actually booted.
*
* The near-miss worth recording: `<!--<script>alert(1)</script>-->` comes back CLEAN, because the
* trailing `-->` returns the parser to script-data state. A probe using the terminated form
* quietly repairs the very thing it is testing for. Only the unterminated variant exposes it.
*/
import { test, expect } from '../../fixtures/test';
import { execFileSync } from 'node:child_process';
import { mkdtempSync, writeFileSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { seedUpload } from '../../helpers/seed';
import { BASE } from '../../helpers/env';
/** Unterminated on purpose — see the header. The terminated form self-repairs. */
const TOKENIZER_PAYLOAD = '<!--<script';
/** The classic break-out. Already handled, kept so the fix can never regress on it. */
const BREAKOUT_PAYLOAD = '</script><img src=x onerror=window.__XSS__=1>';
test.describe('Export — a caption cannot brick the keepsake viewer', () => {
test('the exported viewer boots with a tokenizer-hostile caption in it', async ({
page,
host,
guest,
db,
}) => {
test.setTimeout(120_000);
const bearer = { Authorization: `Bearer ${host.jwt}` };
const g = await guest('Trickster');
const a = await seedUpload(g.jwt, { caption: TOKENIZER_PAYLOAD });
const b = await seedUpload(g.jwt, { caption: BREAKOUT_PAYLOAD });
for (const id of [a, b]) {
await expect.poll(() => db.compressionStatus(id), { timeout: 30_000 }).toBe('done');
}
expect(
(await fetch(`${BASE}/api/v1/host/gallery/release`, { method: 'POST', headers: bearer }))
.status
).toBe(204);
await expect
.poll(
async () => {
const res = await fetch(`${BASE}/api/v1/export/status`, { headers: bearer });
return (await res.json()).html?.status;
},
{ timeout: 90_000, intervals: [500] }
)
.toBe('done');
const ticketRes = await fetch(`${BASE}/api/v1/export/ticket`, {
method: 'POST',
headers: bearer,
});
const { ticket } = (await ticketRes.json()) as { ticket: string };
const dl = await fetch(`${BASE}/api/v1/export/html?ticket=${encodeURIComponent(ticket)}`);
expect(dl.status).toBe(200);
const dir = mkdtempSync(join(tmpdir(), 'eventsnap-viewer-'));
try {
const zipPath = join(dir, 'Memories.zip');
writeFileSync(zipPath, Buffer.from(await dl.arrayBuffer()));
// Extract the WHOLE archive: index.html pulls in the viewer's own JS/CSS, and the point of
// this test is that those later resources are still reachable by the parser.
execFileSync('unzip', ['-qo', zipPath, '-d', dir]);
// file://, not http://. That is how a guest actually opens the keepsake, and it is the
// whole reason the data is inlined rather than fetched from a sibling data.json.
let xss = false;
page.on('dialog', (d) => {
xss = true;
void d.dismiss();
});
await page.goto('file://' + join(dir, 'index.html'));
// 1. The payload was assigned at all. This is the assertion that fails on the old escape —
// the second script block is never reached, so the global stays undefined.
const captions = await page.evaluate(() => {
const d = (window as unknown as { __EXPORT_DATA__?: { posts?: { caption?: string }[] } })
.__EXPORT_DATA__;
return d?.posts?.map((p) => p.caption ?? '') ?? null;
});
expect(captions, '__EXPORT_DATA__ was never assigned — the viewer is bricked').not.toBeNull();
// 2. The captions survived verbatim. The escape is a transport encoding, not a sanitiser:
// a guest's text has to come back exactly, or we have silently rewritten their words.
expect(captions).toContain(TOKENIZER_PAYLOAD);
expect(captions).toContain(BREAKOUT_PAYLOAD);
// 3. And nothing executed.
expect(
await page.evaluate(() => (window as unknown as { __XSS__?: number }).__XSS__ === 1),
'the caption must be inert, not merely non-fatal'
).toBe(false);
expect(xss).toBe(false);
// 4. The viewer actually rendered — the whole document parsed, not just the head. If the
// tokenizer had swallowed the bundle, the body would be empty of viewer output.
await expect(page.locator('body')).not.toBeEmpty();
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
});

View File

@@ -15,7 +15,15 @@ test.describe('Adversarial — small-scale abuse', () => {
await api.patchConfig(adminToken, { rate_limits_enabled: 'true', join_rate_enabled: 'true' });
});
test('20 parallel /join from one IP — rate limiter catches the excess', async () => {
test('a /join flood from one IP is caught by the per-IP ceiling', async ({ api, adminToken }) => {
// This used to assert that 20 joins from one IP produced 429s under a 5/min per-IP
// bucket. That "protection" was the bug: at a venue every guest shares one public IP,
// so it turned real arriving guests away (see 01-auth/rate-limit-shared-nat). The
// anti-spam bucket is now per (ip, name); what remains per-IP is a loose ceiling whose
// job is only to bound raw volume. Squeeze the ceiling so a flood is reproducible here
// without firing 60+ requests.
await api.patchConfig(adminToken, { join_ip_rate_per_min: '5' });
const requests = Array.from({ length: 20 }, (_, i) =>
fetch(`${BASE}/api/v1/join`, {
method: 'POST',
@@ -24,7 +32,7 @@ test.describe('Adversarial — small-scale abuse', () => {
})
);
const statuses = (await Promise.all(requests)).map((r) => r.status);
// 5/min limit → at least some should be 429.
// Ceiling of 5 → the excess must be shed.
expect(statuses.filter((s) => s === 429).length).toBeGreaterThan(0);
// Server stays up — at least one succeeded.
expect(statuses.some((s) => s === 201 || s === 409)).toBe(true);

View File

@@ -45,6 +45,23 @@ test.describe('Media gating — moderation revokes preview access (F2)', () => {
const direct = await fetch(`${BASE}/media/previews/${id}.jpg`);
expect(direct.status, 'direct /media/previews must be blocked').toBe(404);
// …and it must stay blocked under percent-encoding. The block used to be four
// `nest_service("/media/previews", 404)` route matches sitting above a `/media`
// ServeDir. axum routes on the RAW path while ServeDir percent-decodes afterwards, so
// ONE escaped byte (`%70` = `p`) missed every blocker, fell through to the ServeDir,
// and was decoded back to `previews/` on disk — serving the bytes unauthenticated.
// Asserting only the literal spelling is what let that sit here undetected.
for (const variant of [
`/media/%70reviews/${id}.jpg`, // p
`/media/p%72eviews/${id}.jpg`, // r — any position works
`/media/%64isplays/${id}.jpg`, // d
`/media/%74humbnails/${id}.jpg`, // t
`/media/%6Friginals/${id}.jpg`, // o
]) {
const res = await fetch(`${BASE}${variant}`, { redirect: 'manual' });
expect(res.status, `${variant} must not bypass the media block`).toBe(404);
}
// Host deletes the upload → the preview must stop being served.
const del = await fetch(`${BASE}/api/v1/host/upload/${id}`, {
method: 'DELETE',

View File

@@ -86,6 +86,7 @@ export function setAuth(
localStorage.setItem(USER_ID_KEY, userId);
if (displayName) localStorage.setItem(DISPLAY_NAME_KEY, displayName);
isAuthenticated.set(true);
fireSetAuthHooks();
}
/**
@@ -106,6 +107,7 @@ export function setAdminAuth(jwt: string, userId: string, displayName?: string):
sessionStorage.setItem(USER_ID_KEY, userId);
if (displayName) sessionStorage.setItem(DISPLAY_NAME_KEY, displayName);
isAuthenticated.set(true);
fireSetAuthHooks();
}
// Hook registry: cross-cutting stores (export-status, etc.) register a callback
@@ -118,6 +120,26 @@ export function onClearAuth(fn: () => void): void {
clearAuthHooks.push(fn);
}
// The mirror of `onClearAuth`, for stores that must be RE-SEEDED when a new identity
// arrives rather than merely cleared. Without it, anything derived from the token
// survives a logout→login in the same tab: `goto()` is a client-side navigation, so no
// module is re-imported and no `onMount` re-runs, and the previous user's value simply
// stays. Fires after the new token is resident, so hooks can read it.
const setAuthHooks: Array<() => void> = [];
export function onSetAuth(fn: () => void): void {
setAuthHooks.push(fn);
}
function fireSetAuthHooks(): void {
for (const fn of setAuthHooks) {
try {
fn();
} catch {
/* hook failure is non-fatal */
}
}
}
export function clearAuth(): void {
if (!browser) return;
// Clear from BOTH stores — a guest token lives in localStorage, an admin token in

View File

@@ -4,6 +4,7 @@
import { api } from '$lib/api';
import { onSseEvent } from '$lib/sse';
import { getUserId } from '$lib/auth';
import { isStaff } from '$lib/role-store';
import { dataMode, pickMediaUrl } from '$lib/data-mode-store';
import { doubletap } from '$lib/actions/doubletap';
import { focusTrap } from '$lib/actions/focus-trap';
@@ -40,7 +41,20 @@
let heartBurst = $state(false);
let burstTimer: ReturnType<typeof setTimeout> | null = null;
const mediaSrc = $derived(pickMediaUrl($dataMode, upload));
// Videos always play the ORIGINAL. `pickMediaUrl` is mime-agnostic and compression only
// ever produces a *thumbnail* for a video (one ffmpeg frame), so in the default saver
// mode it hands back `/thumbnail` — a JPEG, served as image/jpeg with nosniff. Feeding
// that to <video> is why every video failed with DEMUXER_ERROR_COULD_NOT_OPEN. There is
// no smaller video derivative to offer, so `preload="none"` keeps saver-mode users on
// cellular from fetching anything until they actually press play; the poster is the
// thumbnail, which is what they saw in the feed anyway.
// Fixed here rather than in `pickMediaUrl` because FeedListCard shares that helper and
// legitimately wants the thumbnail for its <img>. Same rule as the diashow.
const mediaSrc = $derived(
isVideo(upload.mime_type)
? `/api/v1/upload/${upload.id}/original`
: pickMediaUrl($dataMode, upload)
);
function triggerHeartBurst() {
heartBurst = true;
@@ -114,9 +128,15 @@
}
}
async function deleteComment(id: string) {
/**
* `asHost` routes to the moderation endpoint. The guest route only ever deletes the
* caller's OWN comment, and it refuses a banned author outright — so without this a
* host who banned an abusive guest was left with the abuse still on screen and no way
* to remove it, since the ban itself blocks the author's own delete.
*/
async function deleteComment(id: string, asHost: boolean) {
try {
await api.delete(`/comment/${id}`);
await api.delete(asHost ? `/host/comment/${id}` : `/comment/${id}`);
comments = comments.filter((c) => c.id !== id);
} catch (e) {
toastError(e);
@@ -170,6 +190,8 @@
<video
src={mediaSrc}
controls
preload="none"
playsinline
class="max-h-[60vh] w-full object-contain"
poster={upload.thumbnail_url ?? undefined}
></video>
@@ -246,11 +268,11 @@
{formatTime(comment.created_at)}
</div>
</div>
{#if comment.user_id === userId}
{#if comment.user_id === userId || $isStaff}
<button
onclick={() => deleteComment(comment.id)}
onclick={() => deleteComment(comment.id, comment.user_id !== userId)}
class="shrink-0 text-gray-400 hover:text-red-500 dark:text-gray-500 dark:hover:text-red-400"
aria-label="Löschen"
aria-label={comment.user_id === userId ? 'Löschen' : 'Kommentar entfernen'}
>
<svg
class="h-3.5 w-3.5"

View File

@@ -1,5 +1,6 @@
import { writable } from 'svelte/store';
import { api } from './api';
import { setRole } from './role-store';
import type { MeContextDto } from './types';
/**
@@ -31,6 +32,9 @@ export async function refreshEventState(): Promise<void> {
const seq = stateSeq;
try {
const ctx = await api.get<MeContextDto>('/me/context');
// The role is unaffected by the close/reopen race guarded below, so apply it
// unconditionally — this is one of the refreshes that used to drop it.
setRole(ctx.role);
// A close/reopen landed while this was in flight — its result is now authoritative;
// don't overwrite it with our possibly-stale snapshot.
if (seq !== stateSeq) return;

View File

@@ -11,6 +11,8 @@ import { getToken, onClearAuth } from './auth';
export interface ExportJob {
status: 'locked' | 'pending' | 'running' | 'done' | 'failed';
progress_pct: number;
/** Only populated when `status === 'failed'` — the reason, phrased for the host. */
error_message: string | null;
}
export interface ExportStatusSnapshot {

View File

@@ -0,0 +1,45 @@
import { derived, writable } from 'svelte/store';
import { getRole, onClearAuth, onSetAuth } from './auth';
export type Role = 'guest' | 'host' | 'admin';
/**
* The viewer's LIVE role.
*
* The JWT is never reissued — the backend slides the session row forward instead and
* deliberately ignores the token's own role claim (`auth/middleware.rs`: "the live user row
* is authoritative"). So `getRole()`, which decodes the claim, is frozen for the lifetime of
* the token: up to 30 days for a guest. A guest promoted to host saw no Host-Dashboard until
* they signed out and back in, even though `/me/context` had already told the client their
* real role on the very next page load — it was fetched and the `role` field dropped on the
* floor in 4 of its 6 call sites.
*
* This store is seeded from the claim (so there is no flash of the wrong nav on boot) and
* corrected by every `/me/context` response via `setRole`. Read this instead of calling
* `getRole()` ad hoc.
*/
export const role = writable<Role | null>(getRole());
/** True for host and admin — the "can moderate" predicate used across the UI. */
export const isStaff = derived(role, ($role) => $role === 'host' || $role === 'admin');
/** Apply the authoritative role from a `/me/context` response. */
export function setRole(next: Role | null): void {
role.set(next);
}
/** Re-seed from the token, e.g. straight after a login/join that minted a new one. */
export function syncRoleFromToken(): void {
role.set(getRole());
}
// Bind the store to the identity lifecycle. The store is a module-level singleton seeded
// ONCE at import, and `goto()` navigations re-import nothing — so without these hooks the
// previous user's role survives a logout→login in the same tab. A host who left and a guest
// who then joined kept `isStaff === true` and were offered "Beitrag entfernen" on other
// people's photos (the backend 403s it, so it was a false affordance rather than an
// escalation — but the feed never fetches /me/context, so it never self-corrected either).
// The mirror case is just as wrong: a guest recovering into a host account got no host
// affordances at all.
onClearAuth(() => role.set(null));
onSetAuth(syncRoleFromToken);

View File

@@ -7,8 +7,18 @@ export const showBottomNav = writable(true);
// Controls the UploadSheet overlay. FAB sets true; sheet sets false.
export const uploadSheetOpen = writable(false);
// Count of items currently pending or uploading — shown as FAB badge.
// Count of items still needing attention — shown as FAB badge. Includes the terminal
// states on purpose: counting only pending/uploading meant a rejected upload decremented
// the badge exactly as if it had succeeded, so the failure was indistinguishable from a
// completed upload. 'blocked' and 'error' stay counted until the user clears or retries them.
export const uploadBadgeCount = derived(
queueItems,
($items) => $items.filter((i) => i.status === 'pending' || i.status === 'uploading').length
($items) =>
$items.filter(
(i) =>
i.status === 'pending' ||
i.status === 'uploading' ||
i.status === 'error' ||
i.status === 'blocked'
).length
);

View File

@@ -3,6 +3,7 @@ import { writable, get } from 'svelte/store';
import { getToken, getUserId, clearAuth } from './auth';
import { onSseEvent } from './sse';
import { refreshQuota } from './quota-store';
import { toast } from './toast-store';
export interface QueueItem {
id: string;
@@ -629,6 +630,11 @@ async function uploadItem(id: string): Promise<void> {
entry.error = e.message;
await database.put(STORE_NAME, entry);
updateItemStatus(id, 'blocked', e.message);
// Tell the user NOW. The queue list only lives on /upload, and the flow sends
// them straight to /feed after staging a photo — so without this a rejected
// upload was silently swallowed: the blob is gone, the FAB badge drops exactly
// as if it had succeeded, and the photo simply never appears.
toast(`${entry.fileName}: ${e.message}`, 'error');
return;
}
const msg = e instanceof Error ? e.message : 'Upload fehlgeschlagen.';

View File

@@ -16,6 +16,7 @@
import { api } from '$lib/api';
import type { MeContextDto } from '$lib/types';
import { eventState, markClosed, markOpened, refreshEventState } from '$lib/event-state-store';
import { setRole } from '$lib/role-store';
import { loadEventConfig } from '$lib/event-config-store';
let { children } = $props();
@@ -48,6 +49,9 @@
try {
const ctx = await api.get<MeContextDto>('/me/context');
privacyNote.set(ctx.privacy_note);
// The live role — the JWT claim is frozen for the token's lifetime, so a
// promotion/demotion only reaches the UI through this. See role-store.ts.
setRole(ctx.role);
eventState.set({
uploadsLocked: ctx.uploads_locked,
galleryReleased: ctx.gallery_released

View File

@@ -1,6 +1,7 @@
<script lang="ts">
import { goto } from '$app/navigation';
import { getToken, getDisplayName, getExpiry, getRole, clearAuth, currentPin } from '$lib/auth';
import { getToken, getDisplayName, getExpiry, clearAuth, currentPin } from '$lib/auth';
import { role, setRole } from '$lib/role-store';
import { clearQueue } from '$lib/upload-queue';
import { api } from '$lib/api';
import { onMount, onDestroy } from 'svelte';
@@ -16,7 +17,6 @@
import type { MeContextDto } from '$lib/types';
let displayName = $state<string | null>(null);
let role = $state<'guest' | 'host' | 'admin' | null>(null);
let expiry = $state<Date | null>(null);
let pinCopied = $state(false);
let leaveConfirmOpen = $state(false);
@@ -35,13 +35,13 @@
return;
}
displayName = getDisplayName();
role = getRole();
expiry = getExpiry();
// Refresh server-driven state. Quota + privacy note may have changed since last visit.
try {
const ctx = await api.get<MeContextDto>('/me/context');
privacyNote.set(ctx.privacy_note);
setRole(ctx.role);
} catch {
// non-fatal
}
@@ -177,10 +177,10 @@
</p>
<span
class="mt-0.5 inline-block rounded-full px-2.5 py-0.5 text-xs font-semibold {roleColor(
role
$role
)}"
>
{roleLabel(role)}
{roleLabel($role)}
</span>
</div>
</div>
@@ -192,7 +192,7 @@
</div>
<!-- Dashboards section (host + admin only) -->
{#if role === 'host' || role === 'admin'}
{#if $role === 'host' || $role === 'admin'}
<div class="card overflow-hidden">
<div class="border-b border-gray-100 px-5 py-3 dark:border-gray-700">
<h2
@@ -230,7 +230,7 @@
<path stroke-linecap="round" stroke-linejoin="round" d="M8.25 4.5l7.5 7.5-7.5 7.5" />
</svg>
</a>
{#if role === 'admin'}
{#if $role === 'admin'}
<a
href="/admin"
class="flex items-center gap-3 border-t border-gray-100 px-5 py-4 transition hover:bg-gray-50 dark:border-gray-700 dark:hover:bg-gray-700/50"
@@ -433,7 +433,7 @@
<!-- Per-user quota widget — staff-only (host/admin); guests never see the
server-derived storage figures. -->
{#if (role === 'host' || role === 'admin') && $quotaStore.enabled && $quotaStore.limit_bytes != null}
{#if ($role === 'host' || $role === 'admin') && $quotaStore.enabled && $quotaStore.limit_bytes != null}
<div class="card p-5">
<h2
class="mb-2 text-xs font-semibold uppercase tracking-wide text-gray-500 dark:text-gray-400"

View File

@@ -1,6 +1,7 @@
<script lang="ts">
import { goto } from '$app/navigation';
import { getToken, getRole } from '$lib/auth';
import { getToken } from '$lib/auth';
import { role as myRoleStore } from '$lib/role-store';
import { api } from '$lib/api';
import { onMount } from 'svelte';
import { toast, toastError } from '$lib/toast-store';
@@ -83,9 +84,16 @@
{ key: 'feed_rate_enabled', label: 'Feed-Limit aktiv', kind: 'bool' },
{ key: 'export_rate_enabled', label: 'Export-Limit aktiv', kind: 'bool' },
{ key: 'join_rate_enabled', label: 'Join-Limit aktiv', kind: 'bool' },
{ key: 'social_rate_enabled', label: 'Interaktions-Limit aktiv', kind: 'bool' },
{ key: 'upload_rate_per_hour', label: 'Upload-Limit pro Stunde', kind: 'number' },
{ key: 'feed_rate_per_min', label: 'Feed-Anfragen pro Minute', kind: 'number' },
{ key: 'export_rate_per_day', label: 'Export-Downloads pro Tag', kind: 'number' }
{ key: 'export_rate_per_day', label: 'Export-Downloads pro Tag', kind: 'number' },
{
key: 'social_rate_per_min',
label: 'Interaktionen pro Minute',
kind: 'number',
hint: 'Likes, Kommentare und Kommentar-Löschungen zusammen, pro Gast. Bewusst hoch angesetzt — soll ein Skript bremsen, keinen begeisterten Gast.'
}
]
},
{
@@ -104,7 +112,20 @@
kind: 'bool',
hint: 'Reserviert für künftige Anzahl-Limits.'
},
{ key: 'quota_tolerance', label: 'Toleranz (01)', kind: 'number' },
{
key: 'quota_tolerance',
label: 'Speicher-Anteil für Gäste (01)',
kind: 'number',
// "Toleranz (01)" with no hint invited exactly the wrong reading — that a higher
// number means "warn me later". It is the multiplier in
// `floor(freier Speicher × Anteil / aktive Uploader)`, so raising it authorises
// guests to fill MORE of the disk, not less.
hint:
'Anteil des freien Speichers, den alle Gäste zusammen belegen dürfen: ' +
'Limit = freier Speicher × Anteil ÷ aktive Uploader. Kein Warnschwellenwert — ' +
'ein höherer Wert gibt MEHR Speicher frei. Das Keepsake braucht zusätzlich ' +
'etwa das Doppelte der Mediengröße; 0,75 ist der getestete Standard.'
},
{ key: 'estimated_guest_count', label: 'Geschätzte Gästezahl', kind: 'number' }
]
},
@@ -262,7 +283,8 @@
let pinResetSubmitting = $state(false);
let pinModal = $state<{ name: string; pin: string } | null>(null);
const myRole = getRole();
// Live role, not the frozen JWT claim: a demotion must disable these controls at once.
const myRole = $derived($myRoleStore);
// Generic confirm-then-run for irreversible / privilege-changing actions
// (promote, demote, unban, release gallery). Reuses the shared ConfirmSheet.

View File

@@ -1,6 +1,7 @@
<script lang="ts">
import { goto } from '$app/navigation';
import { getToken, getUserId } from '$lib/auth';
import { isStaff } from '$lib/role-store';
import { api } from '$lib/api';
import { connectSse, disconnectSse, onSseEvent } from '$lib/sse';
import { onMount, onDestroy } from 'svelte';
@@ -17,6 +18,7 @@
import { pullToRefresh } from '$lib/actions/pull-to-refresh';
import { vibrate } from '$lib/haptics';
import { filterUploads } from '$lib/feed-filter';
import { refreshEventState } from '$lib/event-state-store';
import type { FeedUpload, FeedResponse, HashtagCount, DeltaResponse } from '$lib/types';
let uploads = $state<FeedUpload[]>([]);
@@ -34,7 +36,9 @@
let sentinel: HTMLDivElement;
let feedObserver: IntersectionObserver | null = null;
let inPlaceRefreshTimer: ReturnType<typeof setTimeout> | null = null;
let pendingDeleteId = $state<string | null>(null);
// `asHost` picks the endpoint AND the copy: removing someone else's photo is a
// moderation action, not "delete my post", and it hits the host route.
let pendingDelete = $state<{ id: string; asHost: boolean } | null>(null);
// ─────────────────────────────────────────────────────────────────────────
// onMount A — DOM side-effects only (overscroll lock). Synchronous, returns
@@ -87,7 +91,19 @@
icon: '🗑',
tone: 'danger',
onClick: () => {
pendingDeleteId = target.id;
pendingDelete = { id: target.id, asHost: false };
}
});
} else if ($isStaff) {
// Moderation. Without this the only lever a host had against an unwanted photo
// was banning the uploader — which is both disproportionate and ineffective,
// since a ban does not retract what they already posted.
actions.unshift({
label: 'Beitrag entfernen',
icon: '🚫',
tone: 'danger',
onClick: () => {
pendingDelete = { id: target.id, asHost: true };
}
});
}
@@ -99,14 +115,18 @@
}
async function confirmDelete() {
const id = pendingDeleteId;
if (!id) return;
pendingDeleteId = null;
const pending = pendingDelete;
if (!pending) return;
pendingDelete = null;
try {
await api.delete(`/upload/${id}`);
uploads = uploads.filter((u) => u.id !== id);
if (selectedUpload?.id === id) selectedUpload = null;
void refreshQuota();
// The guest route rejects anything the caller doesn't own, so a host removing
// someone else's photo must go through the host route. That one also emits the
// `upload-deleted` SSE and writes an audit-log entry.
await api.delete(pending.asHost ? `/host/upload/${pending.id}` : `/upload/${pending.id}`);
uploads = uploads.filter((u) => u.id !== pending.id);
if (selectedUpload?.id === pending.id) selectedUpload = null;
// Only our own delete frees our quota; a host removal refunds the uploader.
if (!pending.asHost) void refreshQuota();
} catch (e) {
toastError(e);
}
@@ -214,6 +234,15 @@
}
}
// Pull the authoritative role/event state. The root layout does this on a full page
// load, but arriving here from /join or /recover is a client-side navigation, so its
// onMount never re-runs — and the feed is the one route that gates a destructive
// action (host "Beitrag entfernen") on the role. Without this the feed would run on
// whatever the JWT claim said, which is frozen for the token's 30-day lifetime and so
// misses a promotion or demotion entirely. Cheap, and it refreshes the lock/release
// state in the same request.
void refreshEventState();
await Promise.all([loadFeed(), loadHashtags()]);
connectSse();
@@ -969,13 +998,15 @@
<!-- Branded delete confirmation — replaces window.confirm() -->
<ConfirmSheet
open={pendingDeleteId !== null}
title="Beitrag löschen?"
message="Diese Aktion kann nicht rückgängig gemacht werden."
confirmLabel="Löschen"
open={pendingDelete !== null}
title={pendingDelete?.asHost ? 'Beitrag entfernen?' : 'Beitrag löschen?'}
message={pendingDelete?.asHost
? 'Der Beitrag verschwindet für alle Gäste. Diese Aktion kann nicht rückgängig gemacht werden.'
: 'Diese Aktion kann nicht rückgängig gemacht werden.'}
confirmLabel={pendingDelete?.asHost ? 'Entfernen' : 'Löschen'}
tone="danger"
onConfirm={confirmDelete}
onCancel={() => (pendingDeleteId = null)}
onCancel={() => (pendingDelete = null)}
/>
<!-- First-visit onboarding guide -->

View File

@@ -1,6 +1,7 @@
<script lang="ts">
import { goto } from '$app/navigation';
import { getToken, getRole, getUserId } from '$lib/auth';
import { getToken, getUserId } from '$lib/auth';
import { role as myRoleStore } from '$lib/role-store';
import { api } from '$lib/api';
import type { MeContextDto } from '$lib/types';
import { onMount, onDestroy } from 'svelte';
@@ -27,6 +28,9 @@
is_active: boolean;
uploads_locked: boolean;
export_released: boolean;
disk_free_bytes: number | null;
keepsake_required_bytes: number;
disk_low: boolean;
}
interface PinResetRequest {
@@ -39,6 +43,7 @@
interface ExportJob {
status: string;
progress_pct: number;
error_message: string | null;
}
interface ExportStatusDto {
released: boolean;
@@ -70,6 +75,11 @@
let exportProgress = $derived(
Math.min(exportInfo?.zip?.progress_pct ?? 0, exportInfo?.html?.progress_pct ?? 0)
);
// Either half can carry the reason, and a disk failure usually fails both with the same text —
// so take the first one present rather than rendering it twice.
let exportError = $derived(
exportInfo?.zip?.error_message ?? exportInfo?.html?.error_message ?? null
);
// SSE unsubscribers, torn down on destroy.
let sseOff: Array<() => void> = [];
@@ -97,7 +107,8 @@
let pinResetSubmitting = $state(false);
let pinModal = $state<{ name: string; pin: string } | null>(null);
const myRole = getRole();
// Live role, not the frozen JWT claim: a demotion must disable these controls at once.
const myRole = $derived($myRoleStore);
const myUserId = getUserId();
// Generic confirm-then-run for the irreversible / privilege-changing actions
@@ -387,7 +398,11 @@
function formatBytes(bytes: number): string {
if (bytes < 1024) return `${bytes} B`;
if (bytes < 1024 * 1024) return `${(bytes / 1024).toFixed(1)} KB`;
return `${(bytes / (1024 * 1024)).toFixed(1)} MB`;
// GB matters here now that this also renders free disk and keepsake size — the previous
// version topped out at MB, so 30 GB free read as "30720.0 MB" (and a guest with 2 GB of
// uploads was already being rendered the same way in the user list).
if (bytes < 1024 * 1024 * 1024) return `${(bytes / (1024 * 1024)).toFixed(1)} MB`;
return `${(bytes / (1024 * 1024 * 1024)).toFixed(1)} GB`;
}
</script>
@@ -499,6 +514,38 @@
{error}
</div>
{:else if event}
<!-- ── Speicherwarnung ─────────────────────────────────────────────
Above everything else on purpose. All three volumes (postgres_data, media_data,
exports_data) sit on one filesystem, so running out doesn't degrade a subsystem —
it stops Postgres writing and takes the event down. And the keepsake needs room
for TWO gallery-sized archives, which is only actionable BEFORE the release: the
export preflight can say "this didn't fit", but by then the event is over and the
remedies are all much harder.
Only the admin dashboard had any storage visibility at all, and a host is often
not the admin. `disk_low` fails closed to "not low" on an unreadable mount, so
this cannot cry wolf. -->
{#if event.disk_low && event.disk_free_bytes !== null}
<div
class="rounded-xl border border-red-300 bg-red-50 p-4 dark:border-red-800 dark:bg-red-950/30"
data-testid="low-disk-warning"
>
<h2 class="font-semibold text-red-900 dark:text-red-200">Speicherplatz wird knapp</h2>
<p class="mt-1 text-sm text-red-800 dark:text-red-300">
Noch <strong>{formatBytes(event.disk_free_bytes)}</strong> frei.
{#if event.keepsake_required_bytes > event.disk_free_bytes}
Für das Keepsake werden derzeit ca.
<strong>{formatBytes(event.keepsake_required_bytes)}</strong> benötigt — es kann
momentan <strong>nicht</strong> erstellt werden.
{/if}
</p>
<p class="mt-1.5 text-xs text-red-700 dark:text-red-400">
Schaffe Speicher frei oder vergrößere den Datenträger. Wenn der Datenträger vollläuft,
fällt das gesamte Event aus — nicht nur der Download.
</p>
</div>
{/if}
<!-- ── PIN-Reset-Anfragen ──────────────────────────────────────── -->
{#if pinResetRequests.length > 0}
<div
@@ -684,6 +731,13 @@
</p>
{:else}
<p class="text-red-700 dark:text-red-300">Keepsake-Erstellung fehlgeschlagen.</p>
<!-- The reason, not just the verdict. "Erneut versuchen" is the only control here,
and for the one failure that is actually common — not enough disk — retrying
without freeing space fails identically forever. The backend already wrote a
message naming the numbers; it just wasn't reaching this screen. -->
{#if exportError}
<p class="mt-1 text-red-700/80 dark:text-red-300/80">{exportError}</p>
{/if}
{/if}
<!-- ONE button, mounted in every state — deliberately OUTSIDE the branches above.

View File

@@ -1,6 +1,7 @@
<script lang="ts">
import { goto } from '$app/navigation';
import { getToken, getRole } from '$lib/auth';
import { getToken } from '$lib/auth';
import { isStaff } from '$lib/role-store';
import { addToQueue, loadQueue } from '$lib/upload-queue';
import { toast } from '$lib/toast-store';
import { showBottomNav } from '$lib/ui-store';
@@ -9,6 +10,7 @@
import { onMount, onDestroy } from 'svelte';
import { quotaStore, refreshQuota } from '$lib/quota-store';
import ConfirmSheet from '$lib/components/ConfirmSheet.svelte';
import UploadQueue from '$lib/components/UploadQueue.svelte';
import IconButton from '$lib/components/IconButton.svelte';
import { vibrate } from '$lib/haptics';
import type { PendingFile } from '$lib/pending-upload-store';
@@ -27,7 +29,6 @@
// The storage widget is staff-only (host/admin). Guests never see server-derived
// storage figures — the backend also zeroes the raw-disk fields for non-staff, so
// this is UI-consistency on top of an API guarantee, not the security boundary.
const isStaff = getRole() === 'host' || getRole() === 'admin';
// Quick-tag chips derived from caption as the user types
let captionTags = $derived.by(() => {
@@ -261,7 +262,7 @@
<!-- Per-user quota — staff-only (never shown to guests), and also hidden when
admin disabled enforcement. -->
{#if isStaff && $quotaStore.enabled && $quotaStore.limit_bytes != null}
{#if $isStaff && $quotaStore.enabled && $quotaStore.limit_bytes != null}
<div class="px-4 pt-3 text-xs text-gray-500 dark:text-gray-400">
<div class="flex items-center justify-between">
<span
@@ -344,5 +345,13 @@
: 'Hochladen'}
{/if}
</button>
<!--
The queue list. This component existed, complete with the per-item error text, the
"Erneut" retry button and the rate-limit countdown — and was never imported anywhere,
so none of it could be reached. A rejected upload wrote a clear reason into a store
nothing rendered.
-->
<UploadQueue />
</div>
</div>