Third round. The decode allocation guard is a regression from round 1: swapping reader.decode() for into_decoder() silently dropped the max_alloc enforcement while keeping the comment that claimed it held. Squashed from 3 commits, original messages preserved below. ──────── fix(imaging): restore the decode allocation guard I removed in round 1 This is a regression I introduced, not a pre-existing gap. Before05948d8the compression worker used `ImageReader::decode()`, which does: let mut decoder = Self::make_decoder(format, self.inner, limits.clone())?; limits.reserve(decoder.total_bytes())?; // enforces max_alloc decoder.set_limits(limits)?; Reading the EXIF orientation tag needs `into_decoder()` instead, and that skips the reserve entirely — the crate's own FIXME concedes `from_decoder` doesn't compensate. Nothing else enforces `max_alloc`: the JPEG decoder's `set_limits` only checks support and dimensions. So the 256 MiB budget has been inert since that commit, and round 2 then propagated the weakened path into export.rs through the shared helper, in a commit whose message claimed the helper "carries" the decompression-bomb cap. It didn't, and the comment saying max_alloc "hard-caps the decode allocation" was simply false. What was left was only the per-axis cap, which permits 12000x12000 — 412 MiB decoded, 824 MiB for the two concurrent decodes the worker runs by default, against a 1 GiB container. Deploy-blocking right now because bumping DERIVATIVES_REV makes the first boot after a deploy re-decode the entire gallery two at a time: an OOM kill there restarts the container, which re-runs the backfill. A boot loop, on the first deploy of these fixes. Re-add the reserve exactly as `decode()` does it. Per the budget decision it stays at 256 MiB (~89 MP for RGB8, above any mainstream phone's real output); two concurrent decodes now peak at 512 MiB. Oversized images take the graceful path from round 1 — original retained, quota refunded, upload-error toast — and fail after the header parse but BEFORE any pixels are read, so they cost a header read rather than an allocation. Measured peak during a concurrent oversized burst: 3.0 MiB. Test parity is the other half, and the reason this was invisible: the e2e app container had NO memory limit while production is capped at 1 GiB, so a decode that would OOM-kill production simply succeeded in CI. Mirror the 1 GiB cap in docker-compose.test.yml. That is the third divergence of this shape, after WebKit missing from CI and /health existing only in Caddyfile.test. Tests: a fixture that is 568 KiB on disk and 283 MiB decoded (11000x9000 = 99 MP, deliberately UNDER the per-axis cap so the axis check cannot be what rejects it). A unit test asserts the refusal — it fails against the old code, which decoded it into an 11000x9000 buffer — with a companion asserting an ordinary photo still decodes AND still gets its orientation applied, so the guard didn't become a blanket refusal. An e2e test uploads it singly and as a concurrent pair, asserting compression lands in 'failed' and the backend is still serving and still processing afterwards. ──────── fix(deploy): route /health in production, and actually apply Caddyfile changes Two defects in the update procedure I wrote last round, both of which make a successful-looking deploy a lie. 1. The documented health check could never pass. `curl -fsS https://DOMAIN/health` 404s against a perfectly healthy production stack. The backend registers /health on its ROOT router, not under /api/v1, and the production Caddyfile proxies only /api/* and /media/* — so /health fell through to the SvelteKit catch-all, which has no such route and returns its 404 page. With -f, curl exits 22 and the `&& echo` never runs. My own gloss ("Anything other than ok means check the logs") then sent the operator chasing a phantom outage. e2e/Caddyfile.test has carried `reverse_proxy /health app:3000` since it was written — precisely because the catch-all would otherwise swallow it. Production never did. Per the fix-the-gap-not-the-doc call, production gets the same line, and /health joins the no-store matcher so a cached response can't report the last known state instead of the current one. Verified by running the production Caddyfile against the real backend: /health -> 200 "ok", Cache-Control: no-store, with /api/v1/event and / unaffected. 2. The sequence never reloaded Caddy, so a Caddyfile-only change was dropped. `--build` only rebuilds services with a `build:` section, and caddy is a pinned upstream image. Compose decides whether to recreate a container from its config hash, which covers the mount SPECIFICATION but not the mounted file's CONTENTS — so a git pull that changes ./Caddyfile produces no delta, Compose reports `Running`, and Caddy serves its old config indefinitely. Exit code 0 throughout. Round 1's iOS download fix (137c4ee) is exactly this shape: Caddyfile plus four e2e files, so 100% of its production effect is in that one file. Following the README to the letter deployed it, showed both image IDs changing, and left iOS downloads broken. Demonstrated rather than assumed — added a probe header to a Caddyfile, ran the old sequence (`up -d --build`): header absent, change silently dropped. Ran the new step 4 (`up -d --force-recreate caddy`): header served. `--force-recreate` rather than `restart` or `caddy reload` because the bind mount is resolved to an inode at container-create time and git pull replaces the file rather than editing in place, so a restart can re-read the stale content — the exact failure I hit in round 1 when `caddy reload` didn't pick up an edit. Also rewrites the "db and caddy are untouched … so data volumes survive" sentence. I wrote it as reassurance; "caddy is untouched" was the bug. ──────── chore: take the Bash(*) permission change back out of the shared settings `.claude/settings.json` is committed and applies to anyone who clones. Fabi's local `allow: ["Bash(*)"]` plus deny list ended up in it, insidef0d69f1— a commit about the image decode guard, which has nothing to do with permissions. That was my mistake, twice. The file was already modified when I started the round: my `git status --short` check printed "(clean)" from an unconditional `echo` rather than from the status output, so I read a dirty tree as clean. Then `git add -A` swept it into an unrelated commit, and I reported afterwards that I had left it untouched. Neither the check nor the claim was true. Restores the shared file to its previous three narrow entries. The permission setup itself is preserved, moved to `.claude/settings.local.json`, which `.gitignore:34` covers precisely so per-user permissions stay per-user — the existing 442 entries there are kept alongside it. Not rewritingf0d69f1to erase this: main is unpushed so it would be safe, but a visible correction is worth more than a tidy history, and a rebase across the merge commits carries more risk than the mistake does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
3.3 KiB
3.3 KiB