fix: gate uploads on keepsake headroom, and close five unattended-event gaps

The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit,
the committed comments and README's sizing section all assumed. That correction
is what the first change is about; the rest are the remaining pre-event items.

THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING

`required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each
gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The
upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let
uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free.
Every upload in that band succeeded and the keepsake could then never be built: the
product's entire promise, failing silently at the end of the night with nobody there.

The gate now enforces the invariant that actually matters — never accept an upload
that would make the keepsake unbuildable — sharing `required_free_bytes` with the
preflight so the two cannot drift into disagreeing about the same question. Uploads
stop at ~8 GB of media on this disk, with a German message naming the cause.
Refusing the 1001st photo beats losing all 1000.

`media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s,
rather than `estimate_export_bytes`'s join across every upload. It counts hidden and
banned users' bytes, which the export excludes — skew in the SAFE direction, so the
gate closes marginally early rather than late. Fails open on a query error.

A test pins the gate against the preflight across the whole gallery-size range, and
a second asserts the per-user floor alone would over-commit the volume — i.e. that
the global gate is what must bind.

THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED

`Date.now()` advances while a backgrounded phone is frozen but `setInterval` does
not, so the first tick after a screen lock read the whole sleep as silence and
aborted — re-sending a video from byte zero and burning one of five PERMANENT
auto-attempts. The interval is now its own suspension detector: a tick that arrives
125s late for a 5s schedule credits that window back, because a period the watchdog
could not observe is not evidence of silence.

Chosen over a `visibilitychange` listener, which only covers causes that fire that
event — a throttled-but-visible tab, a closed lid and an occluded window all freeze
timers without one — and which would have needed module state, an SSR guard and a
teardown for strictly less coverage. `performance.now()` was rejected because Safari
pauses it across system sleep on some paths and Chrome does not.

The credit buys one fresh window, not immunity: a socket iOS reaped while
backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout`
(5-60 min) with the queue's `processing` latch held.

Two latent leaks found while in there: `xhr.abort()` on a request already in
readyState DONE emits no `abort` event, so `settle()` never ran and the interval
re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button
silently stopped working); and a synchronous throw from `xhr.send` — a blob whose
backing store the OS purged — leaked the same way. Both closed.

OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER

red/amber/green were never in the @theme block and fell through to Tailwind v4's
`oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot
parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color`
falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to
Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already
in this pipeline — so modern browsers render exactly what they render today. Verified
against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal
(avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0.

Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on
the host dashboard was rendering default violet on EVERY browser, off-brand.

The keepsake viewer only picks this up on a rebuild, so its committed artefact is
rebuilt here too — still single-file, still zero external references.

A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER

With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a
redeploy or a dead uplink left the guest on the boot spinner with no message, no
reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing
nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a
timeout rather than feature detection: a SyntaxError in the bundle is invisible to any
capability check. Plus a <noscript>, since there was nothing at all to see without JS.

EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL

tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info.
If guests spend the evening hitting 429s or 413s, the post-event logs said nothing.
Now one WARN per client error; 5xx excluded because Internal already logs its source
chain and the pool-exhaustion 503 logs at construction.

A DEAD FRONTEND SERVED A BLANK 502

`handle_errors 5xx` with an inline German page (the caddy service mounts only the
Caddyfile, so there is no volume to ship a static file through). Verified empirically
against this config, not from documentation: an upstream 404 through `reverse_proxy`
still arrives as untouched `application/json`, and only a dial failure renders the
page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a
real 404/429 arriving, and swallowing those would have been worse than the blank 502.

CONFIG CORRECTIONS FOR THE REAL HARDWARE

DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count.
Since migration 024 a feed page costs well under a millisecond, so connections are no
longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host.
COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already
serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was
sized against is unreachable; dropping to 1 would halve light-path throughput and push
more feed tiles onto full-size originals. README's sizing section rewritten for the
actual disk.

Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest,
svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate,
compose YAML parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-08-08 22:07:56 +02:00
parent 1d9fb11c7b
commit eb0e405562
16 changed files with 624 additions and 64 deletions

View File

@@ -3,7 +3,8 @@ import {
classifyUploadStatus,
isReversibleLock,
entryToQueueItem,
shouldAbortForStall
shouldAbortForStall,
suspendedSinceLastTick
} from './upload-queue';
/**
@@ -140,4 +141,55 @@ describe('shouldAbortForStall', () => {
expect(shouldAbortForStall(now - 91_000, now, true)).toBe(false);
expect(shouldAbortForStall(now - 121_000, now, true)).toBe(true);
});
});
/**
* The watchdog measures SILENCE via `Date.now()`, but a backgrounded phone freezes the
* interval while the clock keeps running. Without crediting the un-run window back, the first
* tick after a screen lock reads the whole sleep as a stall and aborts a healthy upload —
* re-sending from byte zero and spending one of five permanent auto-attempts. That is what
* every phone does between shots at a party.
*
* Detecting the freeze from the tick gap (rather than from `visibilitychange`) also covers the
* causes that fire no visibility event at all: a throttled-but-visible tab, a closed lid, an
* occluded window.
*/
describe('suspendedSinceLastTick', () => {
const now = 1_000_000;
it('credits nothing for a tick that arrived on schedule', () => {
expect(suspendedSinceLastTick(now - 5_000, now, 5_000)).toBe(0);
});
it('credits nothing for ordinary timer jitter or throttling', () => {
expect(suspendedSinceLastTick(now - 6_900, now, 5_000)).toBe(0);
});
it('credits the whole frozen window when the interval did not run', () => {
// Screen locked ~2 minutes: a 5s interval arriving 130s late.
expect(suspendedSinceLastTick(now - 130_000, now, 5_000)).toBe(125_000);
});
it('credits nothing when the clock jumps backwards (NTP correction)', () => {
expect(suspendedSinceLastTick(now + 60_000, now, 5_000)).toBe(0);
});
it('a suspension longer than the stall ceiling does not abort a healthy upload', () => {
// The bug, end to end: 3 minutes suspended, interval resumes, no bytes since.
const lastProgressAt = now - 180_000;
const credited = Math.min(
now,
lastProgressAt + suspendedSinceLastTick(now - 185_000, now, 5_000)
);
expect(shouldAbortForStall(credited, now, false)).toBe(false);
});
it('but a socket still silent 91s AFTER resume is aborted, never left to xhr.timeout', () => {
// iOS reaps backgrounded sockets without firing `error`. The credit buys one fresh
// window, not immunity — otherwise a dead upload would hang for 5-60 minutes holding
// the queue's `processing` latch.
const resumedAt = now - 91_000;
expect(shouldAbortForStall(resumedAt, now, false)).toBe(true);
});
});

View File

@@ -94,6 +94,45 @@ export function shouldAbortForStall(
return now - lastActivityAt > ceiling;
}
/**
* Normal timer jitter/throttling budget. A tick later than `interval + this` did not run
* because the page was suspended, not because it was merely late.
*/
const SUSPEND_TOLERANCE_MS = 2_000;
/**
* Wall-clock the watchdog interval FAILED to cover because the page was suspended.
*
* `Date.now()` keeps advancing while a backgrounded phone is frozen, but `setInterval` does
* not run. So the first tick after a screen lock saw the entire sleep as "no bytes moved" and
* aborted a connection that was very possibly healthy — restarting a 200 MB video from byte
* zero, burning one of five PERMANENT auto-attempts (`chargeAttempt`), and breaking the drain
* loop for every other queued photo. A phone in a pocket between shots is the common case at a
* party, not an edge case.
*
* The interval is its own suspension detector: a tick scheduled 5s out that arrives 130s late
* means the page was frozen for ~125s, and that is exactly the window the watchdog had no
* right to measure. Deliberately chosen over a `visibilitychange` listener, which only covers
* the causes that happen to fire that event — a throttled-but-visible tab, a closed laptop lid
* and an occluded window all freeze timers without one. It also needs no listener, no
* module-level state, no SSR guard and no teardown.
*
* `performance.now()` was rejected as the clock source: Safari pauses it across system sleep
* on some paths while Chrome does not, which is precisely the non-uniformity that makes it
* unusable as the sole signal.
*
* Returns 0 for a normal tick, and 0 if the clock jumps BACKWARDS (an NTP correction) — that
* fails open, and the wall-clock `xhr.timeout` still bounds the request.
*/
export function suspendedSinceLastTick(
lastTickAt: number,
now: number,
intervalMs: number = STALL_CHECK_INTERVAL_MS
): number {
const overshoot = now - lastTickAt - intervalMs;
return overshoot > SUSPEND_TOLERANCE_MS ? overshoot : 0;
}
/**
* Wall-clock cap for one attempt, scaled by file size assuming a floor of ~8 kB/s — a
* deliberately pessimistic rate, because killing a slow-but-progressing upload would lose
@@ -905,11 +944,25 @@ async function uploadItem(id: string): Promise<void> {
// connection that never errors and never completes. Only "no bytes moved" catches
// that without also punishing a healthy slow link.
let lastProgressAt = Date.now();
let lastTickAt = Date.now();
let bodySent = false;
let stalled = false;
const stallTimer = setInterval(() => {
if (!shouldAbortForStall(lastProgressAt, Date.now(), bodySent)) return;
// Never fire twice. `xhr.abort()` on a request already in `readyState === DONE`
// emits NO `abort` event, so `settle()` would never run: the interval would keep
// running forever, re-aborting every 5s, and `activeUploads` would keep a stale
// entry so the guest's ✕ button silently did nothing.
if (stalled) return;
const now = Date.now();
// Credit back the window the page was frozen. The watchdog measures SILENCE, and
// a period in which it could not observe anything is not evidence of silence.
// Clamped to `now` so a progress event delivered right at resume cannot push the
// timestamp into the future.
lastProgressAt = Math.min(now, lastProgressAt + suspendedSinceLastTick(lastTickAt, now));
lastTickAt = now;
if (!shouldAbortForStall(lastProgressAt, now, bodySent)) return;
stalled = true;
clearInterval(stallTimer);
xhr.abort();
}, STALL_CHECK_INTERVAL_MS);
const settle = (fn: () => void) => {
@@ -1018,7 +1071,16 @@ async function uploadItem(id: string): Promise<void> {
else reject(new NetworkError('Abgebrochen'));
})
);
xhr.send(formData);
// `send` can throw SYNCHRONOUSLY — most plausibly on a phone whose OS purged the
// backing store for the blob, leaving a neutered File. The executor would turn that
// into a rejection and `uploadItem` would recover, but `settle()` never runs: the
// stall interval leaks and `activeUploads` keeps a stale entry, so the ✕ button on
// that item stops working for the rest of the session.
try {
xhr.send(formData);
} catch {
settle(() => reject(new NetworkError('Netzwerkfehler')));
}
});
// Success — remove blob from IndexedDB, mark done