Third round. The decode allocation guard is a regression from round 1: swapping
reader.decode() for into_decoder() silently dropped the max_alloc enforcement
while keeping the comment that claimed it held.
Squashed from 3 commits, original messages preserved below.
──────── fix(imaging): restore the decode allocation guard I removed in round 1
This is a regression I introduced, not a pre-existing gap. Before 05948d8 the
compression worker used `ImageReader::decode()`, which does:
let mut decoder = Self::make_decoder(format, self.inner, limits.clone())?;
limits.reserve(decoder.total_bytes())?; // enforces max_alloc
decoder.set_limits(limits)?;
Reading the EXIF orientation tag needs `into_decoder()` instead, and that skips
the reserve entirely — the crate's own FIXME concedes `from_decoder` doesn't
compensate. Nothing else enforces `max_alloc`: the JPEG decoder's `set_limits`
only checks support and dimensions. So the 256 MiB budget has been inert since
that commit, and round 2 then propagated the weakened path into export.rs through
the shared helper, in a commit whose message claimed the helper "carries" the
decompression-bomb cap. It didn't, and the comment saying max_alloc "hard-caps
the decode allocation" was simply false.
What was left was only the per-axis cap, which permits 12000x12000 — 412 MiB
decoded, 824 MiB for the two concurrent decodes the worker runs by default,
against a 1 GiB container. Deploy-blocking right now because bumping
DERIVATIVES_REV makes the first boot after a deploy re-decode the entire gallery
two at a time: an OOM kill there restarts the container, which re-runs the
backfill. A boot loop, on the first deploy of these fixes.
Re-add the reserve exactly as `decode()` does it. Per the budget decision it stays
at 256 MiB (~89 MP for RGB8, above any mainstream phone's real output); two
concurrent decodes now peak at 512 MiB. Oversized images take the graceful path
from round 1 — original retained, quota refunded, upload-error toast — and fail
after the header parse but BEFORE any pixels are read, so they cost a header read
rather than an allocation. Measured peak during a concurrent oversized burst: 3.0
MiB.
Test parity is the other half, and the reason this was invisible: the e2e app
container had NO memory limit while production is capped at 1 GiB, so a decode
that would OOM-kill production simply succeeded in CI. Mirror the 1 GiB cap in
docker-compose.test.yml. That is the third divergence of this shape, after WebKit
missing from CI and /health existing only in Caddyfile.test.
Tests: a fixture that is 568 KiB on disk and 283 MiB decoded (11000x9000 = 99 MP,
deliberately UNDER the per-axis cap so the axis check cannot be what rejects it).
A unit test asserts the refusal — it fails against the old code, which decoded it
into an 11000x9000 buffer — with a companion asserting an ordinary photo still
decodes AND still gets its orientation applied, so the guard didn't become a
blanket refusal. An e2e test uploads it singly and as a concurrent pair, asserting
compression lands in 'failed' and the backend is still serving and still
processing afterwards.
──────── fix(deploy): route /health in production, and actually apply Caddyfile changes
Two defects in the update procedure I wrote last round, both of which make a
successful-looking deploy a lie.
1. The documented health check could never pass.
`curl -fsS https://DOMAIN/health` 404s against a perfectly healthy production
stack. The backend registers /health on its ROOT router, not under /api/v1, and
the production Caddyfile proxies only /api/* and /media/* — so /health fell
through to the SvelteKit catch-all, which has no such route and returns its 404
page. With -f, curl exits 22 and the `&& echo` never runs. My own gloss
("Anything other than ok means check the logs") then sent the operator chasing a
phantom outage.
e2e/Caddyfile.test has carried `reverse_proxy /health app:3000` since it was
written — precisely because the catch-all would otherwise swallow it. Production
never did. Per the fix-the-gap-not-the-doc call, production gets the same line,
and /health joins the no-store matcher so a cached response can't report the last
known state instead of the current one. Verified by running the production
Caddyfile against the real backend: /health -> 200 "ok", Cache-Control: no-store,
with /api/v1/event and / unaffected.
2. The sequence never reloaded Caddy, so a Caddyfile-only change was dropped.
`--build` only rebuilds services with a `build:` section, and caddy is a pinned
upstream image. Compose decides whether to recreate a container from its config
hash, which covers the mount SPECIFICATION but not the mounted file's CONTENTS —
so a git pull that changes ./Caddyfile produces no delta, Compose reports
`Running`, and Caddy serves its old config indefinitely. Exit code 0 throughout.
Round 1's iOS download fix (137c4ee) is exactly this shape: Caddyfile plus four
e2e files, so 100% of its production effect is in that one file. Following the
README to the letter deployed it, showed both image IDs changing, and left iOS
downloads broken.
Demonstrated rather than assumed — added a probe header to a Caddyfile, ran the
old sequence (`up -d --build`): header absent, change silently dropped. Ran the
new step 4 (`up -d --force-recreate caddy`): header served.
`--force-recreate` rather than `restart` or `caddy reload` because the bind mount
is resolved to an inode at container-create time and git pull replaces the file
rather than editing in place, so a restart can re-read the stale content — the
exact failure I hit in round 1 when `caddy reload` didn't pick up an edit.
Also rewrites the "db and caddy are untouched … so data volumes survive" sentence.
I wrote it as reassurance; "caddy is untouched" was the bug.
──────── chore: take the Bash(*) permission change back out of the shared settings
`.claude/settings.json` is committed and applies to anyone who clones. Fabi's
local `allow: ["Bash(*)"]` plus deny list ended up in it, inside f0d69f1 — a
commit about the image decode guard, which has nothing to do with permissions.
That was my mistake, twice. The file was already modified when I started the
round: my `git status --short` check printed "(clean)" from an unconditional
`echo` rather than from the status output, so I read a dirty tree as clean. Then
`git add -A` swept it into an unrelated commit, and I reported afterwards that I
had left it untouched. Neither the check nor the claim was true.
Restores the shared file to its previous three narrow entries. The permission
setup itself is preserved, moved to `.claude/settings.local.json`, which
`.gitignore:34` covers precisely so per-user permissions stay per-user — the
existing 442 entries there are kept alongside it.
Not rewriting f0d69f1 to erase this: main is unpushed so it would be safe, but a
visible correction is worth more than a tidy history, and a rebase across the
merge commits carries more risk than the mistake does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
72 lines
3.6 KiB
Caddyfile
72 lines
3.6 KiB
Caddyfile
{$DOMAIN} {
|
|
# Compress everything EXCEPT the SSE stream — gzip buffering delays
|
|
# "real-time" likes/comments until the ~30s keep-alive tick.
|
|
@compressible not path /api/v1/stream
|
|
encode @compressible zstd gzip
|
|
|
|
# Site-wide security headers (defense-in-depth). HSTS is free since Caddy
|
|
# already terminates TLS. nosniff also covers all of /media/*.
|
|
header {
|
|
Strict-Transport-Security "max-age=31536000; includeSubDomains"
|
|
X-Content-Type-Options "nosniff"
|
|
Referrer-Policy "strict-origin-when-cross-origin"
|
|
}
|
|
|
|
# X-Frame-Options: DENY everywhere EXCEPT the keepsake download endpoints, which
|
|
# are navigated in a HIDDEN, SAME-ORIGIN iframe so a 404/429 can't unload the PWA
|
|
# (see frontend/src/routes/export/+page.svelte). WebKit enforces XFO *before*
|
|
# honouring Content-Disposition, so a blanket DENY makes the download silently do
|
|
# nothing on iOS Safari — the app's primary platform. SAMEORIGIN still blocks
|
|
# cross-origin framing.
|
|
#
|
|
# Split into two disjoint matchers rather than an override: Caddy applies the
|
|
# FIRST header directive outermost, so it wins on write — a later, more specific
|
|
# `header` would be silently ignored.
|
|
@framable path /api/v1/export/zip /api/v1/export/html
|
|
@not_framable not path /api/v1/export/zip /api/v1/export/html
|
|
header @framable X-Frame-Options "SAMEORIGIN"
|
|
header @not_framable X-Frame-Options "DENY"
|
|
|
|
# SvelteKit frontend — static assets with long-lived cache (content-hashed filenames)
|
|
@hashed_assets path_regexp hashed /_app/immutable/.*\.[a-f0-9]{8,}\.(js|css|woff2)$
|
|
header @hashed_assets Cache-Control "public, max-age=31536000, immutable"
|
|
|
|
# Preview/thumbnail images. These are served by the app through a visibility-checked
|
|
# alias (/api/v1/upload/{id}/{preview,thumbnail}) so moderation can revoke access;
|
|
# the app serves no /media route at all, so there is no direct path to the bytes.
|
|
# Privately cacheable for a short window (the app sets the same header; this is the
|
|
# edge carve-out from the blanket no-store below). Kept short so a moderated image
|
|
# stops being served to a direct-URL holder promptly.
|
|
@media_api path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail
|
|
header @media_api Cache-Control "private, max-age=300"
|
|
|
|
# API and health — never cache, EXCEPT the gated image routes above. A cached health
|
|
# response would report the last known state rather than the current one.
|
|
@api {
|
|
path /api/* /health
|
|
not path /api/v1/upload/*/preview /api/v1/upload/*/thumbnail
|
|
}
|
|
header @api Cache-Control "no-store"
|
|
|
|
# Route API and media requests to the Rust backend.
|
|
#
|
|
# The app serves no /media route at all (see the note in backend/src/main.rs) — media
|
|
# bytes are reachable only through the visibility-checked /api/v1/upload aliases, so
|
|
# /media/* forwards to a plain 404. The proxy line is kept deliberately: it means the
|
|
# edge faithfully hands /media to the app, so if a future change ever re-introduces a
|
|
# static media route the e2e gating specs see it here exactly as production would,
|
|
# instead of being masked by the SvelteKit 404 page.
|
|
reverse_proxy /api/* app:3000
|
|
reverse_proxy /media/* app:3000
|
|
|
|
# The backend registers /health on its ROOT router, not under /api/v1, so it needs its
|
|
# own line — without it the catch-all below hands /health to SvelteKit, which has no
|
|
# such route and returns its 404 page. That made the documented post-deploy check
|
|
# (`curl -fsS https://DOMAIN/health`) fail 100% of the time on a perfectly healthy
|
|
# stack. e2e/Caddyfile.test has always carried this line; production never did.
|
|
reverse_proxy /health app:3000
|
|
|
|
# Everything else goes to SvelteKit frontend
|
|
reverse_proxy frontend:3001
|
|
}
|