Files
EventSnap/backend/src/services/imaging.rs
MechaCat02 ef6d3a077a
Some checks failed
Checks / Backend — cargo test + clippy + fmt (push) Failing after 1m5s
Checks / Frontend — vitest + svelte-check (push) Failing after 5m55s
Checks / E2E — typecheck + lint (push) Failing after 49s
E2E / Playwright E2E (chromium + webkit) (push) Failing after 10m42s
E2E / Cross-UA smoke matrix (push) Failing after 7m57s
Audit / cargo audit (backend) (push) Failing after 10m15s
Audit / npm audit (frontend) (push) Successful in 53s
fix: close what nine adversarial reviews found, most of it mine
Nine focused reviews (export state machine, upload path, auth/abuse, client
queue, guest UI, database, deploy/ops, regression hunt, test honesty). Every
finding below was re-verified against the code before being acted on; several
plausible-sounding ones were checked and rejected.

## Data loss and denial of service

**One request could OOM-kill the app container.** `client_upload_id` was read
with `Field::text()` — axum builds its multipart reader with no SizeLimit, so
the only bound was the route's 576 MiB body limit, then decoded into a second
full String. `caption` and `hashtags` go through `read_text_field_bounded` for
exactly this reason; this field arrived later and missed it. Any guest, one
request, and every SSE stream drops and every in-flight temp file is stranded.

**Nothing bounded concurrent upload bodies.** The headroom gate can only refuse
to COMMIT — the body is already streamed to a temp file by the time it runs, and
neither axum, the tower stack nor Caddy limits how many stream at once. ~100
guests tapping "upload all" after the ceremony puts 10-20 GB of .tmp on a 40 GB
volume, invisible to the gate, eating the reserve that keeps Postgres able to
write WAL. New `UploadAdmission` budgets bytes (not requests, so one video and
two hundred photos coexist) via a permit that releases on drop, so every exit
path returns it.

**The export decode bypassed the memory permit the compression path takes.**
Same class of work — decode + resize every image in the gallery — in a bare
spawn_blocking. A release fired while the last photos were still compressing put
both in the same 1 GiB cgroup; the OOM kill marks the export failed and
`recover_exports` re-spawns it into the same conditions on the next boot. The
permit is now process-wide in `imaging`, because the constraint it expresses is
the container's memory, not one worker's.

**`MediaTotalCache` cached its own failure as 0.** For the whole TTL the gate
then saw an empty event and collapsed to the flat reserve — the behaviour the
two-halves design replaced — with no log line. And the trigger correlates with
the danger: with max_connections 10 the query fails exactly during a burst. Now
falls back to the last good reading and says so.

**V8's heap ceiling sat above the frontend container's entire budget** (measured:
259 MB inside a 256M limit), so GC could never intervene and the only
backpressure was SIGKILL under an arrival burst.

## Guest-visible

**The feed stopped being newest-first after the first reconcile.** It fetches
whole 100-item server pages while `uploads` grows in 20s, so everything in the
gap was absent from `present`, classified as new, and prepended — ~80 photos
from earlier in the evening above the newest ones. It also stalled infinite
scroll, since the cursor still pointed at item 20 and the observer only re-fires
on a change. The union is now sorted on the server's own (created_at, id) key,
which additionally places an SSE arrival correctly.

**A stale `loadMoreError` outlived every refresh and filter change**, leaving a
false error above a button that returns immediately on `!nextCursor`.

**A failed derivative toasted "Ein Upload konnte nicht verarbeitet werden."** for
a photo sitting right there on screen — the handler still assumed 1d9fb11's
pre-fix behaviour (row deleted, quota refunded, card evicted), none of which is
true any more. It was the last surviving route for the "your photo is gone"
signal that fix set out to remove.

## Enforcement that existed only in comments

`recover_name_rate_per_15min` is clamped at the point of use: the ordering
`3 x ceiling <= PIN_LOCK_THRESHOLD` is the whole control against one source
locking any guest whose name is on the feed, it was asserted in a comment, and
`patch_config` accepted 1..100_000. The test pinned the default constant rather
than the enforced bound; it now pins the bound.

## Tests that could not fail

- The gate test asserted only its own premise (`500MB x 100 > 35GB`) and never
  touched the gate. It now checks both controls against the same state and
  requires them to disagree in the right direction.
- `the_banner_always_fires_before_the_upload_gate_closes` reduced to
  `G < G + G/4` — true for any margin, including zero, so it could not detect
  the banner moving to exactly the gate. It now pins the gap.
- `disk_is_low`'s `free < LOW_DISK_FLOOR_BYTES` clause was unreachable (warn_at
  is always >= 12.5 GB against a 10 GB floor). Two tests were named after it and
  neither could fail if it were deleted. Clause and constant removed.
- The suspension test I added last commit hard-coded the credit cap instead of
  importing it, so changing STALL_TIMEOUT_MS would leave it passing against a
  system that no longer exists. Now imports MAX_SUSPEND_CREDIT_MS.

## Stale comments corrected

The prune doc still argued at length for the pre-build ordering that 1d9fb11
reversed — a reader trusting it would reopen the blocker 0506369 fixed.
DISK_RESERVE_BYTES claimed to equal the banner threshold that 0506369
deliberately offset by 25%. And host.rs kept its own duplicate 10 GB literal
instead of importing the constant.

154/154 backend, 59/59 vitest, clippy clean, svelte-check 0 errors, eslint
clean, both builds, compose + caddy validate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 14:51:58 +02:00

389 lines
19 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! Shared image decoding.
//!
//! Exists so there is exactly ONE way to turn a file on disk into a `DynamicImage` in this
//! codebase. Two properties have to hold everywhere an image is decoded, and both were
//! previously re-derived per call site — which is how they drifted apart:
//!
//! - **EXIF orientation must be applied.** Phones do not rotate sensor data; they record how
//! the camera was held in a tag and store the pixels as shot. `image::open` and
//! `ImageReader::decode` both hand back the raw pixels and ignore that tag, and re-encoding
//! to JPEG writes no EXIF, so the derivative is permanently sideways while the untouched
//! original still renders upright. The compression worker was fixed; the export worker was
//! not, so every portrait photo came out sideways in the keepsake's HTML viewer.
//! - **Decode limits must be set.** The upload body cap bounds the file on disk, but a small
//! file can decode to enormous dimensions (a ~1 MB image expanding to 50k×50k px), OOM-ing
//! the box. `image::open` applies NO limits at all, so the export path was also decoding
//! arbitrary user-supplied images unbounded.
use anyhow::{Context, Result};
use image::{DynamicImage, ImageDecoder};
use std::path::Path;
/// Bounds for any decode of user-supplied image data. The per-axis cap covers any real phone
/// photo; `max_alloc` bounds the decoded buffer — but only because `decode_oriented` reserves
/// against it explicitly, see there.
///
/// Sized against the deployment: the app container is capped at 1 GiB and the compression
/// worker runs `compression_concurrency` decodes at once (default 2), so 256 MiB per decode
/// leaves headroom for the resize buffers and the runtime.
fn decode_limits() -> image::Limits {
let mut limits = image::Limits::default();
limits.max_image_width = Some(12_000);
limits.max_image_height = Some(12_000);
limits.max_alloc = Some(256 * 1024 * 1024);
limits
}
/// True when re-running the exact same work on the exact same bytes cannot possibly
/// succeed, so retrying only burns wall-clock and log noise.
///
/// Deliberately narrow. Only the `ImageError` variants that are a property of the *input*
/// count: the file will not shrink, gain codec support, or un-corrupt itself between
/// attempts. `IoError` is excluded on purpose — EMFILE under load, or a momentarily
/// unreadable file, is exactly the transient case the retry exists for. A FULL disk is the
/// one io error that must not be retried either, but for a different reason and with a
/// different remedy; see [`is_storage_full_error`].
pub fn is_permanent_image_error(err: &anyhow::Error) -> bool {
err.chain().any(|cause| {
matches!(
cause.downcast_ref::<image::ImageError>(),
Some(
image::ImageError::Limits(_)
| image::ImageError::Unsupported(_)
| image::ImageError::Decoding(_)
)
)
})
}
/// True when the failure is the media filesystem being out of space.
///
/// Deliberately separate from [`is_permanent_image_error`], which is about the *input*. ENOSPC
/// is about the *host*, and it is the one failure the retry loop actively makes worse: a disk
/// does not drain during six seconds of backoff, so all three attempts fail identically while
/// holding a compression permit that photos are queued behind.
///
/// The give-up path it fed was worse still. It refunded the guest's quota and soft-deleted the
/// row while deliberately RETAINING the original — so the bytes stayed on the full disk, the
/// photo vanished from the feed seconds after a `201 Created`, and the guest was handed back
/// the quota to upload it again into the same full disk. Each round shrank free space further.
pub fn is_storage_full_error(err: &anyhow::Error) -> bool {
fn is_full(io: &std::io::Error) -> bool {
// `StorageFull` is the portable classification; the raw ENOSPC catches the paths where
// the OS error was never mapped to a named kind.
io.kind() == std::io::ErrorKind::StorageFull || io.raw_os_error() == Some(28)
}
err.chain().any(|cause| {
// `image` wraps the io error in its own variant rather than exposing it as a source,
// so the plain downcast alone would miss every derivative-write failure.
cause.downcast_ref::<std::io::Error>().is_some_and(is_full)
|| matches!(
cause.downcast_ref::<image::ImageError>(),
Some(image::ImageError::IoError(io)) if is_full(io)
)
})
}
/// Build a decoder for `path` with the budget enforced, WITHOUT reading any pixels.
///
/// Single source of truth for "may this image be decoded at all": both the upload
/// admission check and the compression worker go through here, so they cannot disagree
/// about what is acceptable.
fn decoder_within_budget(path: &Path) -> Result<impl image::ImageDecoder> {
let mut reader = image::ImageReader::open(path)
.context("failed to open image")?
.with_guessed_format()
.context("failed to read image header")?;
let mut limits = decode_limits();
reader.limits(limits.clone());
// We need `into_decoder` rather than `decode()` to read the EXIF orientation tag before
// the pixels are consumed. But the two are NOT equivalent on safety: `decode()` performs
//
// limits.reserve(decoder.total_bytes())?;
//
// between building the decoder and reading the image, and `into_decoder()` skips it (the
// crate's own FIXME concedes `from_decoder` doesn't compensate). Nothing else enforces
// `max_alloc` — the JPEG decoder's `set_limits` only checks support and dimensions — so
// without the line below the budget is inert and the ONLY bound is the per-axis cap. That
// leaves 12000x12000 decodable at 412 MiB, and two concurrent at 824 MiB against a 1 GiB
// container. Re-add it, exactly as `decode()` does.
let mut decoder = reader.into_decoder().context("failed to decode image")?;
limits
.reserve(decoder.total_bytes())
.context("image too large to decode within the memory budget")?;
decoder
.set_limits(limits)
.context("image too large to decode within the memory budget")?;
Ok(decoder)
}
/// Rough peak heap an image will cost to turn into derivatives, read from the HEADER only —
/// no pixels are decoded. `None` when the header can't be read or the image is over budget
/// (the caller is about to fail on it anyway).
///
/// Two terms, and the second is the one that surprises:
///
/// - the decoded buffer, `width * height * channels`; and
/// - the resize intermediate. `image`'s Lanczos3 path accumulates in `f32`, so the buffer
/// between the horizontal and vertical passes is `new_width * old_height * 4 channels * 4
/// bytes` — 16 bytes per pixel-row-slot, not the 4 the output uses. For an 8000x8000
/// original that is 262 MiB on top of a 244 MiB decode, measured. It is bigger than the
/// decode for any tall image, which is why "the decode is bounded by max_alloc" was never
/// the whole story.
///
/// Used to decide whether an image is heavy enough to need exclusive use of the box's memory
/// headroom, NOT to reject anything.
pub fn estimated_processing_peak_bytes(path: &Path, display_edge: u32) -> Option<u64> {
let decoder = decoder_within_budget(path).ok()?;
let (width, height) = decoder.dimensions();
let decoded = decoder.total_bytes();
// Aspect-preserving fit into `display_edge`, matching DynamicImage::resize. No downscale
// means no intermediate at all.
let intermediate = if width > display_edge || height > display_edge {
let ratio = f64::from(display_edge) / f64::from(width.max(height));
let new_width = (f64::from(width) * ratio).round().max(1.0) as u64;
new_width * u64::from(height) * 16
} else {
0
};
Some(decoded.saturating_add(intermediate))
}
/// Megapixels an image would decode to, or `None` if its header can't be read. Used only
/// to put a concrete number in the message the guest sees.
pub fn megapixels(path: &Path) -> Option<f64> {
let reader = image::ImageReader::open(path)
.ok()?
.with_guessed_format()
.ok()?;
let (w, h) = reader.into_dimensions().ok()?;
Some(f64::from(w) * f64::from(h) / 1_000_000.0)
}
/// True when an image cannot be decoded specifically because it would exceed the memory
/// budget — read from the header, no pixels touched.
///
/// Called at upload admission so a guest who sends a 100 MP photo is told at the door, with
/// a reason they can act on, instead of the upload being accepted with a 201 and then
/// silently soft-deleted minutes later when the worker gives up on it.
///
/// Deliberately narrow: ONLY the budget. A corrupt, truncated or unsupported file also
/// fails to build a decoder, but rejecting those here would change a contract the
/// adversarial suite pins on purpose — acceptance follows the magic bytes, and a payload
/// with a valid JPEG header is accepted regardless of what follows it. Those go to the
/// compression worker as before, which handles them gracefully and (since the retry
/// classifier) no longer burns backoff on them.
pub fn exceeds_decode_budget(path: &Path) -> bool {
match decoder_within_budget(path) {
Ok(_) => false,
Err(e) => e.chain().any(|cause| {
matches!(
cause.downcast_ref::<image::ImageError>(),
Some(image::ImageError::Limits(_))
)
}),
}
}
/// Decode an image from disk with decompression-bomb limits applied and its EXIF
/// orientation baked into the pixels.
///
/// Blocking — call inside `spawn_blocking`.
pub fn decode_oriented(path: &Path) -> Result<DynamicImage> {
let mut decoder = decoder_within_budget(path)?;
// Cheap, and it happens BEFORE any pixels are read: an oversized image costs a header
// parse, not an allocation.
let orientation = decoder
.orientation()
.unwrap_or(image::metadata::Orientation::NoTransforms);
let mut img = DynamicImage::from_decoder(decoder).context("failed to decode image")?;
img.apply_orientation(orientation);
Ok(img)
}
/// Process-wide serialisation for memory-heavy image work.
///
/// The `app` container gets 1 GiB. A single 8000x8000 original measures ~516 MiB peak even with
/// the decode correctly scoped, so two overlapping giants is an OOM kill — and the kernel kills
/// the whole process, dropping every SSE stream and stranding every in-flight upload.
///
/// GLOBAL rather than a field on `CompressionWorker`, because the constraint is the container's
/// memory and there is more than one producer of this work. The export's own image path
/// (`services::export`) decodes and resizes every photo in the gallery — a thumbnail for each,
/// plus a 2000px re-encode for every original over 5 MB — and it ran in a bare `spawn_blocking`
/// with no permit at all. So "host taps Freigeben while the last phone photos are still
/// compressing" put an export decode and a heavy compression job in the same cgroup at the same
/// time, which is the scenario the permit exists to make impossible. Worse, it is self-repeating:
/// the OOM kill marks the export failed, and `recover_exports` re-spawns it on boot into the same
/// conditions.
///
/// Held across the blocking section and released on drop, including on error.
pub static HEAVY_IMAGE_PERMITS: std::sync::LazyLock<tokio::sync::Semaphore> =
std::sync::LazyLock::new(|| tokio::sync::Semaphore::new(1));
/// Estimated peak heap above which a job must take [`HEAVY_IMAGE_PERMITS`].
///
/// 150 MiB sits far above a normal phone photo (a 12 MP JPEG costs ~50 MiB all-in) so the common
/// path never serialises, and far below the point where two jobs stop fitting in the container.
pub const HEAVY_IMAGE_BYTES: u64 = 150 * 1024 * 1024;
#[cfg(test)]
mod tests {
use super::*;
/// Shared with the e2e suite rather than duplicating 568 KiB of binary: the same file
/// drives `02-upload/oversized-image` so both layers assert on one artefact.
const HUGE: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/huge-99mp.jpg"
);
#[test]
fn rejects_an_image_that_would_blow_the_allocation_budget() {
// 11000x9000 = 99 MP. Deliberately UNDER the 12000px per-axis cap, so the axis check
// cannot reject it — the allocation budget is the only thing that can, which is
// exactly what makes this a regression test rather than a restatement of the axis cap.
// 283 MiB decoded as RGB8 against a 256 MiB budget, from 568 KiB on disk.
//
// This failed before the guard was restored: `ImageReader::decode` performs
// `limits.reserve(decoder.total_bytes())`, and `into_decoder()` — which we need for
// the EXIF tag — skips it, so `max_alloc` was inert and this decoded happily.
// Map the Ok arm to its dimensions first: on failure `expect_err` Debug-prints the
// value, and Debug on a DynamicImage dumps every pixel — 283 MiB of output.
let err = decode_oriented(Path::new(HUGE))
.map(|img| (img.width(), img.height()))
.expect_err("a 99 MP image must be refused, not allocated");
let msg = format!("{err:#}");
assert!(
msg.to_lowercase().contains("limit") || msg.to_lowercase().contains("memory"),
"expected a limits error, got: {msg}"
);
}
#[test]
fn an_oversized_image_is_a_permanent_failure() {
// The retry loop must not burn 2s + 4s of backoff on this: the file will not shrink
// between attempts, so all three attempts reach the identical conclusion.
let err = decode_oriented(Path::new(HUGE))
.map(|img| (img.width(), img.height()))
.expect_err("fixture must exceed the budget");
assert!(
is_permanent_image_error(&err),
"a Limits error can never succeed on retry: {err:#}"
);
}
#[test]
fn a_plain_io_error_is_not_permanent() {
// The mirror that keeps the classifier honest. EMFILE under load, or a momentary
// unreadable file, is exactly what the retry exists for — misclassifying those as
// permanent would turn a transient blip back into the data loss round 1 fixed.
// (A FULL disk is its own case now; see the storage-full tests below.)
let err = decode_oriented(Path::new("/nonexistent/definitely-not-here.jpg"))
.map(|img| (img.width(), img.height()))
.expect_err("a missing file must error");
assert!(
!is_permanent_image_error(&err),
"an IO error must stay retryable: {err:#}"
);
}
#[test]
fn a_full_disk_is_recognised_through_both_wrappers() {
// The two shapes ENOSPC actually arrives in. A bare io::Error is what `tokio::fs` and
// `std::fs` produce; the `image` crate wraps its own in `ImageError::IoError`, which is
// NOT reachable via `source()` — so a chain walk that only downcast to io::Error would
// miss every derivative-write failure, i.e. the exact case this classifier exists for.
let bare = anyhow::Error::from(std::io::Error::from(std::io::ErrorKind::StorageFull))
.context("failed to write the preview");
assert!(is_storage_full_error(&bare), "bare io::Error: {bare:#}");
let wrapped = anyhow::Error::from(image::ImageError::IoError(std::io::Error::from(
std::io::ErrorKind::StorageFull,
)))
.context("failed to save the display derivative");
assert!(
is_storage_full_error(&wrapped),
"ImageError::IoError: {wrapped:#}"
);
}
#[test]
fn an_ordinary_io_error_is_not_a_full_disk() {
// Keeps the classifier from swallowing the general case: only ENOSPC may skip the retry
// and take the keep-the-row branch. Anything else must still be retried and, if it keeps
// failing, soft-deleted as before.
let missing = decode_oriented(Path::new("/nonexistent/definitely-not-here.jpg"))
.map(|img| (img.width(), img.height()))
.expect_err("a missing file must error");
assert!(
!is_storage_full_error(&missing),
"a missing file is not a full disk: {missing:#}"
);
}
#[test]
fn admission_rejects_only_the_over_budget_case() {
// Admission and processing must agree about SIZE — a photo accepted at the door and
// then rejected by the worker for being too big is the failure this pair prevents.
assert!(
exceeds_decode_budget(Path::new(HUGE)),
"admission must reject what the decoder rejects for size"
);
let ordinary = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/portrait-exif6.jpg"
);
assert!(
!exceeds_decode_budget(Path::new(ordinary)),
"admission must accept an ordinary photo"
);
}
#[test]
fn admission_does_not_reject_a_merely_undecodable_file() {
// The narrowing that keeps the adversarial contract intact: a payload with valid
// JPEG magic bytes and nothing behind them cannot be decoded, but acceptance follows
// the magic bytes by design (07-adversarial/file-upload-attacks). It is the worker's
// job to fail it, not admission's — admission is only the resource guard.
let dir = std::env::temp_dir().join("eventsnap-imaging-test");
std::fs::create_dir_all(&dir).expect("tmp dir");
let stub = dir.join("magic-only.jpg");
let mut bytes = vec![0u8; 1024];
bytes[..3].copy_from_slice(&[0xFF, 0xD8, 0xFF]);
std::fs::write(&stub, &bytes).expect("write stub");
assert!(
!exceeds_decode_budget(&stub),
"a corrupt file is not an over-budget file"
);
assert!(
decode_oriented(&stub)
.map(|i| (i.width(), i.height()))
.is_err(),
"...but it must still fail in the worker"
);
let _ = std::fs::remove_file(&stub);
}
#[test]
fn still_decodes_an_ordinary_photo_and_applies_orientation() {
// The guard must not have become a blanket refusal. This fixture is 40x20 stored with
// EXIF Orientation=6, so a correct decode returns it rotated to 20x40 portrait.
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../e2e/fixtures/media/portrait-exif6.jpg"
);
let img = decode_oriented(Path::new(path)).expect("an ordinary photo must decode");
assert_eq!(
(img.width(), img.height()),
(20, 40),
"EXIF orientation must still be applied after restoring the guard"
);
}
}