fix: gate uploads on keepsake headroom, and close five unattended-event gaps
The box is 2 vCPU / 4 GB / 40 GB, not the 4 vCPU / 8 GB / 80 GB that the audit, the committed comments and README's sizing section all assumed. That correction is what the first change is about; the rest are the remaining pre-event items. THE ARCHIVE COULD BECOME UNBUILDABLE WHILE UPLOADS KEPT SUCCEEDING `required_free_bytes` is `media × 1.1 × 2` — the ZIP and the HTML viewer are each gallery-sized — and the export preflight also wants DISK_RESERVE_BYTES on top. The upload gate, though, only refused below a FLAT 10 GB reserve. On 40 GB that let uploads run to ~25 GB of media while a release needed `2.2 × 25 + 10` = 65 GB free. Every upload in that band succeeded and the keepsake could then never be built: the product's entire promise, failing silently at the end of the night with nobody there. The gate now enforces the invariant that actually matters — never accept an upload that would make the keepsake unbuildable — sharing `required_free_bytes` with the preflight so the two cannot drift into disagreeing about the same question. Uploads stop at ~8 GB of media on this disk, with a German message naming the cause. Refusing the 1001st photo beats losing all 1000. `media_total.rs` backs it: SUM(user.total_upload_bytes) over ~100 rows, cached 5s, rather than `estimate_export_bytes`'s join across every upload. It counts hidden and banned users' bytes, which the export excludes — skew in the SAFE direction, so the gate closes marginally early rather than late. Fails open on a query error. A test pins the gate against the preflight across the whole gallery-size range, and a second asserts the per-user floor alone would over-commit the volume — i.e. that the global gate is what must bind. THE WATCHDOG ABORTED HEALTHY UPLOADS EVERY TIME A PHONE WAS POCKETED `Date.now()` advances while a backgrounded phone is frozen but `setInterval` does not, so the first tick after a screen lock read the whole sleep as silence and aborted — re-sending a video from byte zero and burning one of five PERMANENT auto-attempts. The interval is now its own suspension detector: a tick that arrives 125s late for a 5s schedule credits that window back, because a period the watchdog could not observe is not evidence of silence. Chosen over a `visibilitychange` listener, which only covers causes that fire that event — a throttled-but-visible tab, a closed lid and an occluded window all freeze timers without one — and which would have needed module state, an SSR guard and a teardown for strictly less coverage. `performance.now()` was rejected because Safari pauses it across system sleep on some paths and Chrome does not. The credit buys one fresh window, not immunity: a socket iOS reaped while backgrounded still aborts ~90s after resume rather than hanging for `xhr.timeout` (5-60 min) with the queue's `processing` latch held. Two latent leaks found while in there: `xhr.abort()` on a request already in readyState DONE emits no `abort` event, so `settle()` never ran and the interval re-aborted every 5s forever while `activeUploads` kept a stale entry (the ✕ button silently stopped working); and a synchronous throw from `xhr.send` — a blob whose backing store the OS purged — leaked the same way. Both closed. OKLCH MADE THE DELETE BUTTON INVISIBLE ON SAMSUNG'S DEFAULT BROWSER red/amber/green were never in the @theme block and fell through to Tailwind v4's `oklch()` defaults, which Safari <15.4, Chrome <111 and Samsung Internet <22 cannot parse: `var(--color-red-600)` is then invalid at computed-value time, `background-color` falls back to transparent, and `.btn-danger` renders white text on nothing. Pinned to Tailwind's own defaults gamut-mapped to sRGB by Lightning CSS — the converter already in this pipeline — so modern browsers render exactly what they render today. Verified against seven hex fallbacks it had already emitted for the /alpha forms. rose and teal (avatar chips) had the same leak. The app CSS goes from 40 oklch declarations to 0. Also fixes `--color-purple-950`, which was simply missing: `dark:bg-purple-950/50` on the host dashboard was rendering default violet on EVERY browser, off-brand. The keepsake viewer only picks this up on a rebuild, so its committed artefact is rebuilt here too — still single-file, still zero external references. A BRICKED BOOT LOOKED LIKE A SPINNER FOREVER With `ssr = false` the page is empty until the bundle mounts, so a chunk 404 after a redeploy or a dead uplink left the guest on the boot spinner with no message, no reload control, and in a standalone PWA no URL bar. A 15s timeout in the existing nonce'd IIFE (no CSP change) swaps in German copy and a reload button. Deliberately a timeout rather than feature detection: a SyntaxError in the bundle is invisible to any capability check. Plus a <noscript>, since there was nothing at all to see without JS. EVERY 4xx WAS INVISIBLE AT ANY LOG LEVEL tower_http counts 4xx as a success, so it logs at DEBUG while production runs at info. If guests spend the evening hitting 429s or 413s, the post-event logs said nothing. Now one WARN per client error; 5xx excluded because Internal already logs its source chain and the pool-exhaustion 503 logs at construction. A DEAD FRONTEND SERVED A BLANK 502 `handle_errors 5xx` with an inline German page (the caddy service mounts only the Caddyfile, so there is no volume to ship a static file through). Verified empirically against this config, not from documentation: an upstream 404 through `reverse_proxy` still arrives as untouched `application/json`, and only a dial failure renders the page. That mattered — the keepsake download navigates a hidden iframe and DEPENDS on a real 404/429 arriving, and swallowing those would have been worse than the blank 502. CONFIG CORRECTIONS FOR THE REAL HARDWARE DATABASE_MAX_CONNECTIONS 30 → 15: sized to 2 vCPU rather than to the guest count. Since migration 024 a feed page costs well under a millisecond, so connections are no longer spent waiting, and 30 backends crowd the db container's 1 GB on a 4 GB host. COMPRESSION_WORKER_CONCURRENCY stays at 2 — the merged heavy-image permit already serialises anything over 150 MiB, so the "two 48 MP photos" worst case that number was sized against is unreachable; dropping to 1 would halve light-path throughput and push more feed tiles onto full-size originals. README's sizing section rewritten for the actual disk. Verified: 149/149 backend tests against a live Postgres, clippy clean, 57/57 vitest, svelte-check 0 errors, eslint clean, vite build, export-viewer rebuild, caddy validate, compose YAML parse. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -46,6 +46,33 @@
|
||||
document.head.appendChild(s);
|
||||
}
|
||||
} catch (_) {}
|
||||
|
||||
// Boot backstop. With `ssr = false` the page is EMPTY until the bundle mounts, so
|
||||
// anything that stops it mounting leaves the guest on the spinner below forever:
|
||||
// a chunk 404 after a redeploy, a dead uplink, or untranspiled syntax on an old
|
||||
// phone. There is no message, no reload control, and in a standalone PWA no URL
|
||||
// bar to escape from — the app is simply bricked for that guest, all evening.
|
||||
//
|
||||
// A TIMEOUT, deliberately, not feature detection: a SyntaxError in the bundle is
|
||||
// invisible to any capability check, whereas "still not painted" catches every
|
||||
// cause at once. The root layout removes #app-boot on mount, so its continued
|
||||
// presence IS the failure signal and nothing needs to cancel this.
|
||||
//
|
||||
// 15s is well beyond a slow-but-healthy load on venue wifi (the bundle is ~130 kB
|
||||
// gzipped); erring long matters more than erring short, because a false positive
|
||||
// here would tell a guest something is broken while it is quietly working.
|
||||
setTimeout(function () {
|
||||
var boot = document.getElementById('app-boot');
|
||||
if (!boot) return; // app mounted — nothing to do
|
||||
boot.innerHTML =
|
||||
'<div style="max-width:20rem;text-align:center">' +
|
||||
'<p style="font-family:Georgia,serif;font-size:1.125rem;font-weight:600;margin:0 0 .5rem">Die App konnte nicht geladen werden</p>' +
|
||||
'<p style="margin:0 0 1.25rem;color:#545350;line-height:1.5">Bitte prüf deine Verbindung und lade die Seite neu.</p>' +
|
||||
'<button id="app-boot-reload" style="font:inherit;font-weight:600;cursor:pointer;border:0;border-radius:.5rem;padding:.625rem 1.25rem;background:#8a6a2b;color:#fff">Neu laden</button>' +
|
||||
'</div>';
|
||||
var btn = document.getElementById('app-boot-reload');
|
||||
if (btn) btn.addEventListener('click', function () { location.reload(); });
|
||||
}, 15000);
|
||||
})();
|
||||
</script>
|
||||
%sveltekit.head%
|
||||
@@ -65,6 +92,14 @@
|
||||
<span class="app-boot__spinner"></span>
|
||||
<span class="app-boot__label">EventSnap</span>
|
||||
</div>
|
||||
<!-- The app is client-rendered end to end, so with JS off there is nothing at all to
|
||||
show. Say so, rather than leaving a permanent spinner over a blank page. -->
|
||||
<noscript>
|
||||
<div class="app-boot__noscript">
|
||||
<p><strong>EventSnap braucht JavaScript</strong></p>
|
||||
<p>Bitte aktiviere JavaScript im Browser und lade die Seite neu.</p>
|
||||
</div>
|
||||
</noscript>
|
||||
<style>
|
||||
#app-boot {
|
||||
position: fixed;
|
||||
@@ -103,6 +138,26 @@
|
||||
transform: rotate(360deg);
|
||||
}
|
||||
}
|
||||
/* Sits above #app-boot, which is `position: fixed` and would otherwise cover it. */
|
||||
.app-boot__noscript {
|
||||
position: fixed;
|
||||
inset: 0;
|
||||
z-index: 1;
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
gap: 0.25rem;
|
||||
padding: 2rem;
|
||||
text-align: center;
|
||||
background: #faf9f7;
|
||||
font-family: system-ui, -apple-system, 'Segoe UI', Roboto, sans-serif;
|
||||
color: #545350;
|
||||
}
|
||||
html.dark .app-boot__noscript {
|
||||
background: #100f0f;
|
||||
color: #a6a4a1;
|
||||
}
|
||||
</style>
|
||||
</body>
|
||||
</html>
|
||||
|
||||
@@ -3,7 +3,8 @@ import {
|
||||
classifyUploadStatus,
|
||||
isReversibleLock,
|
||||
entryToQueueItem,
|
||||
shouldAbortForStall
|
||||
shouldAbortForStall,
|
||||
suspendedSinceLastTick
|
||||
} from './upload-queue';
|
||||
|
||||
/**
|
||||
@@ -140,4 +141,55 @@ describe('shouldAbortForStall', () => {
|
||||
expect(shouldAbortForStall(now - 91_000, now, true)).toBe(false);
|
||||
expect(shouldAbortForStall(now - 121_000, now, true)).toBe(true);
|
||||
});
|
||||
|
||||
});
|
||||
|
||||
/**
|
||||
* The watchdog measures SILENCE via `Date.now()`, but a backgrounded phone freezes the
|
||||
* interval while the clock keeps running. Without crediting the un-run window back, the first
|
||||
* tick after a screen lock reads the whole sleep as a stall and aborts a healthy upload —
|
||||
* re-sending from byte zero and spending one of five permanent auto-attempts. That is what
|
||||
* every phone does between shots at a party.
|
||||
*
|
||||
* Detecting the freeze from the tick gap (rather than from `visibilitychange`) also covers the
|
||||
* causes that fire no visibility event at all: a throttled-but-visible tab, a closed lid, an
|
||||
* occluded window.
|
||||
*/
|
||||
describe('suspendedSinceLastTick', () => {
|
||||
const now = 1_000_000;
|
||||
|
||||
it('credits nothing for a tick that arrived on schedule', () => {
|
||||
expect(suspendedSinceLastTick(now - 5_000, now, 5_000)).toBe(0);
|
||||
});
|
||||
|
||||
it('credits nothing for ordinary timer jitter or throttling', () => {
|
||||
expect(suspendedSinceLastTick(now - 6_900, now, 5_000)).toBe(0);
|
||||
});
|
||||
|
||||
it('credits the whole frozen window when the interval did not run', () => {
|
||||
// Screen locked ~2 minutes: a 5s interval arriving 130s late.
|
||||
expect(suspendedSinceLastTick(now - 130_000, now, 5_000)).toBe(125_000);
|
||||
});
|
||||
|
||||
it('credits nothing when the clock jumps backwards (NTP correction)', () => {
|
||||
expect(suspendedSinceLastTick(now + 60_000, now, 5_000)).toBe(0);
|
||||
});
|
||||
|
||||
it('a suspension longer than the stall ceiling does not abort a healthy upload', () => {
|
||||
// The bug, end to end: 3 minutes suspended, interval resumes, no bytes since.
|
||||
const lastProgressAt = now - 180_000;
|
||||
const credited = Math.min(
|
||||
now,
|
||||
lastProgressAt + suspendedSinceLastTick(now - 185_000, now, 5_000)
|
||||
);
|
||||
expect(shouldAbortForStall(credited, now, false)).toBe(false);
|
||||
});
|
||||
|
||||
it('but a socket still silent 91s AFTER resume is aborted, never left to xhr.timeout', () => {
|
||||
// iOS reaps backgrounded sockets without firing `error`. The credit buys one fresh
|
||||
// window, not immunity — otherwise a dead upload would hang for 5-60 minutes holding
|
||||
// the queue's `processing` latch.
|
||||
const resumedAt = now - 91_000;
|
||||
expect(shouldAbortForStall(resumedAt, now, false)).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -94,6 +94,45 @@ export function shouldAbortForStall(
|
||||
return now - lastActivityAt > ceiling;
|
||||
}
|
||||
|
||||
/**
|
||||
* Normal timer jitter/throttling budget. A tick later than `interval + this` did not run
|
||||
* because the page was suspended, not because it was merely late.
|
||||
*/
|
||||
const SUSPEND_TOLERANCE_MS = 2_000;
|
||||
|
||||
/**
|
||||
* Wall-clock the watchdog interval FAILED to cover because the page was suspended.
|
||||
*
|
||||
* `Date.now()` keeps advancing while a backgrounded phone is frozen, but `setInterval` does
|
||||
* not run. So the first tick after a screen lock saw the entire sleep as "no bytes moved" and
|
||||
* aborted a connection that was very possibly healthy — restarting a 200 MB video from byte
|
||||
* zero, burning one of five PERMANENT auto-attempts (`chargeAttempt`), and breaking the drain
|
||||
* loop for every other queued photo. A phone in a pocket between shots is the common case at a
|
||||
* party, not an edge case.
|
||||
*
|
||||
* The interval is its own suspension detector: a tick scheduled 5s out that arrives 130s late
|
||||
* means the page was frozen for ~125s, and that is exactly the window the watchdog had no
|
||||
* right to measure. Deliberately chosen over a `visibilitychange` listener, which only covers
|
||||
* the causes that happen to fire that event — a throttled-but-visible tab, a closed laptop lid
|
||||
* and an occluded window all freeze timers without one. It also needs no listener, no
|
||||
* module-level state, no SSR guard and no teardown.
|
||||
*
|
||||
* `performance.now()` was rejected as the clock source: Safari pauses it across system sleep
|
||||
* on some paths while Chrome does not, which is precisely the non-uniformity that makes it
|
||||
* unusable as the sole signal.
|
||||
*
|
||||
* Returns 0 for a normal tick, and 0 if the clock jumps BACKWARDS (an NTP correction) — that
|
||||
* fails open, and the wall-clock `xhr.timeout` still bounds the request.
|
||||
*/
|
||||
export function suspendedSinceLastTick(
|
||||
lastTickAt: number,
|
||||
now: number,
|
||||
intervalMs: number = STALL_CHECK_INTERVAL_MS
|
||||
): number {
|
||||
const overshoot = now - lastTickAt - intervalMs;
|
||||
return overshoot > SUSPEND_TOLERANCE_MS ? overshoot : 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Wall-clock cap for one attempt, scaled by file size assuming a floor of ~8 kB/s — a
|
||||
* deliberately pessimistic rate, because killing a slow-but-progressing upload would lose
|
||||
@@ -905,11 +944,25 @@ async function uploadItem(id: string): Promise<void> {
|
||||
// connection that never errors and never completes. Only "no bytes moved" catches
|
||||
// that without also punishing a healthy slow link.
|
||||
let lastProgressAt = Date.now();
|
||||
let lastTickAt = Date.now();
|
||||
let bodySent = false;
|
||||
let stalled = false;
|
||||
const stallTimer = setInterval(() => {
|
||||
if (!shouldAbortForStall(lastProgressAt, Date.now(), bodySent)) return;
|
||||
// Never fire twice. `xhr.abort()` on a request already in `readyState === DONE`
|
||||
// emits NO `abort` event, so `settle()` would never run: the interval would keep
|
||||
// running forever, re-aborting every 5s, and `activeUploads` would keep a stale
|
||||
// entry so the guest's ✕ button silently did nothing.
|
||||
if (stalled) return;
|
||||
const now = Date.now();
|
||||
// Credit back the window the page was frozen. The watchdog measures SILENCE, and
|
||||
// a period in which it could not observe anything is not evidence of silence.
|
||||
// Clamped to `now` so a progress event delivered right at resume cannot push the
|
||||
// timestamp into the future.
|
||||
lastProgressAt = Math.min(now, lastProgressAt + suspendedSinceLastTick(lastTickAt, now));
|
||||
lastTickAt = now;
|
||||
if (!shouldAbortForStall(lastProgressAt, now, bodySent)) return;
|
||||
stalled = true;
|
||||
clearInterval(stallTimer);
|
||||
xhr.abort();
|
||||
}, STALL_CHECK_INTERVAL_MS);
|
||||
const settle = (fn: () => void) => {
|
||||
@@ -1018,7 +1071,16 @@ async function uploadItem(id: string): Promise<void> {
|
||||
else reject(new NetworkError('Abgebrochen'));
|
||||
})
|
||||
);
|
||||
xhr.send(formData);
|
||||
// `send` can throw SYNCHRONOUSLY — most plausibly on a phone whose OS purged the
|
||||
// backing store for the blob, leaving a neutered File. The executor would turn that
|
||||
// into a rejection and `uploadItem` would recover, but `settle()` never runs: the
|
||||
// stall interval leaks and `activeUploads` keeps a stale entry, so the ✕ button on
|
||||
// that item stops working for the rest of the session.
|
||||
try {
|
||||
xhr.send(formData);
|
||||
} catch {
|
||||
settle(() => reject(new NetworkError('Netzwerkfehler')));
|
||||
}
|
||||
});
|
||||
|
||||
// Success — remove blob from IndexedDB, mark done
|
||||
|
||||
@@ -95,11 +95,81 @@
|
||||
--color-purple-700: #6f5523;
|
||||
--color-purple-800: #59441e;
|
||||
--color-purple-900: #493819;
|
||||
/* The ramp stopped at 900 while blue/primary both run to 950, so
|
||||
* `dark:bg-purple-950/50` (routes/host/+page.svelte) fell through to Tailwind's default
|
||||
* violet — off-brand on every browser, modern ones included. Mirrors `--color-blue-950`. */
|
||||
--color-purple-950: #29200d;
|
||||
--color-violet-500: #ab8433;
|
||||
--color-violet-600: #8a6a2b;
|
||||
--color-accent-500: #ab8433;
|
||||
--color-accent-600: #8a6a2b;
|
||||
|
||||
/* ── Semantic: red / amber / green ───────────────────────────────────────────
|
||||
* PINNED TO HEX rather than inherited. Tailwind v4 ships these families as
|
||||
* `oklch()`, which Safari <15.4, Chrome <111 and Samsung Internet <22 (the default
|
||||
* browser on Samsung phones) cannot parse. The declaration is accepted but
|
||||
* `var(--color-red-600)` is then invalid at computed-value time, so `background-color`
|
||||
* falls back to `initial` — transparent — and `.btn-danger` in the component layer
|
||||
* renders white text on nothing. An invisible delete-confirm button.
|
||||
*
|
||||
* Note this is NOT the `@apply` hard-fail described for `primary` above: these families
|
||||
* have Tailwind defaults, so an unpinned stop does not break the build, it silently
|
||||
* reverts to oklch. Full 50-950 ramps are therefore about closing that silent-leak class
|
||||
* permanently, not about compiling.
|
||||
*
|
||||
* Values are Tailwind 4.2.2's own defaults gamut-mapped to sRGB by Lightning CSS — the
|
||||
* same converter already in this build pipeline, which is why they match the hex
|
||||
* fallbacks it emits for the `/alpha` opacity forms (verified against `#bf000f`,
|
||||
* `#460809`, `#461901`, `#032e15`, `#82181a`, `#ffa3a3`, `#0d542b` in the shipped CSS).
|
||||
* Naive channel clipping gives different, wrong values for the out-of-gamut stops.
|
||||
* Modern browsers therefore render exactly what they render today. */
|
||||
--color-red-50: #fef2f2;
|
||||
--color-red-100: #ffe2e2;
|
||||
--color-red-200: #ffcaca;
|
||||
--color-red-300: #ffa3a3;
|
||||
--color-red-400: #ff6568;
|
||||
--color-red-500: #fb2c36;
|
||||
--color-red-600: #e40014;
|
||||
--color-red-700: #bf000f;
|
||||
--color-red-800: #9f0712;
|
||||
--color-red-900: #82181a;
|
||||
--color-red-950: #460809;
|
||||
|
||||
--color-amber-50: #fffbeb;
|
||||
--color-amber-100: #fef3c6;
|
||||
--color-amber-200: #fee685;
|
||||
--color-amber-300: #ffd236;
|
||||
--color-amber-400: #fcbb00;
|
||||
--color-amber-500: #f99c00;
|
||||
--color-amber-600: #dd7400;
|
||||
--color-amber-700: #b75000;
|
||||
--color-amber-800: #953d00;
|
||||
--color-amber-900: #7b3306;
|
||||
--color-amber-950: #461901;
|
||||
|
||||
--color-green-50: #f0fdf4;
|
||||
--color-green-100: #dcfce7;
|
||||
--color-green-200: #b9f8cf;
|
||||
--color-green-300: #7bf1a8;
|
||||
--color-green-400: #05df72;
|
||||
--color-green-500: #00c758;
|
||||
--color-green-600: #00a544;
|
||||
--color-green-700: #008138;
|
||||
--color-green-800: #016630;
|
||||
--color-green-900: #0d542b;
|
||||
--color-green-950: #032e15;
|
||||
|
||||
/* Avatar chips (lib/avatar.ts) — same oklch problem, same fix. Only the stops those
|
||||
* chips actually use; nothing else in the app references rose or teal. */
|
||||
--color-rose-100: #ffe4e6;
|
||||
--color-rose-200: #ffccd3;
|
||||
--color-rose-700: #c20039;
|
||||
--color-rose-900: #8b0836;
|
||||
--color-teal-100: #cbfbf1;
|
||||
--color-teal-200: #96f7e4;
|
||||
--color-teal-700: #00776e;
|
||||
--color-teal-900: #0b4f4a;
|
||||
|
||||
/* ── Neutrals: pearl → silver → graphite (remaps `gray-*`). Whisper-warm
|
||||
* pearl at the light end (research: "warm pearl, not stark white"), turning
|
||||
* neutral-cool through the mids/darks so structure reads as silver, not sand. */
|
||||
|
||||
Reference in New Issue
Block a user