Six independent operational defects, none of which needed a new feature to fix. Log rotation. Docker's json-file driver is unbounded by default, and those files land on the HOST filesystem — outside every deploy.resources.limits in the compose file, and on the same disk as postgres_data and media_data. A full disk stops Postgres writing WAL, which takes the event down. Capped at 10m x 3 per service. Log level. RUST_LOG was set in neither .env.example nor docker-compose.yml, so the code fallback WAS the production level — and it was `debug`, with tower_http=debug emitting a line per request and per response into that unrotated file. Now info, with tower_http=warn to state that those spans are diagnostics, not an access log. ffmpeg pipe deadlock. run_ffmpeg piped stdout and stderr and then called wait(), which drains neither. Once the ~64 KiB pipe buffer filled, ffmpeg blocked writing and wait() never returned — burning the full 120s timeout, twice per seek position, three times per compression attempt. And the timeout is an Err, so the end state was a soft-deleted upload: a guest's playable video destroyed by a poster-frame failure. Now stdout is null (nothing ever read it) and stderr is drained by wait_with_output, whose tail is logged on a non-zero exit. Note wait_with_output consumes the child, so the old kill-on-timeout is gone; kill_on_drop(true) already covers it. Readiness probe. /health never touched the pool, so the disk-full endgame above stayed green all the way down. Adds /health/ready (SELECT 1 under 2s) as a SECOND route — the compose healthcheck deliberately keeps pointing at /health, because caddy gates its startup on it and a DB-dependent probe would turn a Postgres blip into the reverse proxy refusing to start. api.ts request timeout. The abort timer was cleared in a finally around fetch(), which resolves on the response HEAD — leaving res.text() uncovered and no longer abortable. An upstream that sends headers then stalls the body hung the call forever. The timer now lives until the body is read, including the 204 path (which otherwise leaked a live 20s timer per no-content request). Upload XHR watchdog. The XHR had no timeout while processQueue held the isProcessing latch across it; on a half-open socket neither error nor abort ever fires, so the latch pinned and the queue wedged. Bounds SILENCE rather than total duration — a 500 MB video over a venue uplink legitimately runs 30+ minutes while making steady progress. Rejects as NetworkError, which is already the retryable branch, so a stalled upload now recovers like any network blip. IndexedDB failures. addToQueue called getDb() unguarded and handleSubmit had no catch, so a private-mode refusal or a QuotaExceededError on a large blob left a permanent "Wird hochgeladen…" spinner, no toast, and — for an in-app camera capture — the only copy of the photo gone. Now reported as 'failed', which keeps the staged files on screen and stays on the page. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
144 lines
5.8 KiB
TypeScript
144 lines
5.8 KiB
TypeScript
import { describe, it, expect } from 'vitest';
|
|
import {
|
|
classifyUploadStatus,
|
|
isReversibleLock,
|
|
entryToQueueItem,
|
|
shouldAbortForStall
|
|
} from './upload-queue';
|
|
|
|
/**
|
|
* Regression guard for the upload-queue retry policy (H2 + M1). The bug being locked out:
|
|
* every 4xx except 429 was classified terminal, and a terminal item has its blob PURGED
|
|
* from IndexedDB. A 401 (a sliding session that lapsed, or a host PIN-reset) is a 4xx, so a
|
|
* guest's queued photos were irrecoverably destroyed the moment the `online` auto-resume
|
|
* fired against a dead session. 401 must be `auth` (blob kept, re-auth), NEVER `terminal`.
|
|
*/
|
|
describe('classifyUploadStatus', () => {
|
|
it('2xx → success', () => {
|
|
expect(classifyUploadStatus(200)).toBe('success');
|
|
expect(classifyUploadStatus(201)).toBe('success');
|
|
expect(classifyUploadStatus(299)).toBe('success');
|
|
});
|
|
|
|
it('401 → auth, NOT terminal (never purge the blob on a dead session)', () => {
|
|
expect(classifyUploadStatus(401)).toBe('auth');
|
|
});
|
|
|
|
it('429 → rate_limit (back off, auto-resume)', () => {
|
|
expect(classifyUploadStatus(429)).toBe('rate_limit');
|
|
});
|
|
|
|
it('408 → transient (request timeout is retryable, not terminal)', () => {
|
|
expect(classifyUploadStatus(408)).toBe('transient');
|
|
});
|
|
|
|
it('genuinely permanent 4xx → terminal (locked / banned / released / quota)', () => {
|
|
expect(classifyUploadStatus(403)).toBe('terminal'); // banned / locked / released
|
|
expect(classifyUploadStatus(413)).toBe('terminal'); // quota exhausted
|
|
expect(classifyUploadStatus(400)).toBe('terminal');
|
|
expect(classifyUploadStatus(404)).toBe('terminal');
|
|
});
|
|
|
|
it('5xx / unexpected → transient (retryable, blob kept)', () => {
|
|
expect(classifyUploadStatus(500)).toBe('transient');
|
|
expect(classifyUploadStatus(502)).toBe('transient');
|
|
expect(classifyUploadStatus(503)).toBe('transient');
|
|
});
|
|
});
|
|
|
|
/**
|
|
* Regression guard for the reversible-lock discrimination inside the `terminal` bucket — the
|
|
* branch that decides whether a 4xx KEEPS the blob (event closed / gallery released: a host can
|
|
* reopen and the photo resumes) or PURGES it (permanent ban / quota). Getting this wrong either
|
|
* loses a photo the guest expected to survive a reopen, or lets a banned device retry forever.
|
|
*/
|
|
describe('isReversibleLock', () => {
|
|
it('an `uploads_locked` code is reversible at any status (event closed / released)', () => {
|
|
expect(isReversibleLock(403, 'uploads_locked')).toBe(true);
|
|
expect(isReversibleLock(409, 'uploads_locked')).toBe(true);
|
|
});
|
|
|
|
it('a `forbidden` 403 (banned) is PERMANENT — purge, never resume', () => {
|
|
expect(isReversibleLock(403, 'forbidden')).toBe(false);
|
|
});
|
|
|
|
it('an unidentifiable 403 (unparseable proxy/WAF/captive-portal body) is treated reversible', () => {
|
|
// Losing a photo is the worst outcome; 403 is the reversible-lock status here.
|
|
expect(isReversibleLock(403, undefined)).toBe(true);
|
|
expect(isReversibleLock(403, null)).toBe(true);
|
|
expect(isReversibleLock(403, '')).toBe(true);
|
|
});
|
|
|
|
it('a non-403 permanent 4xx (e.g. 413 quota) is NOT reversible unless explicitly locked', () => {
|
|
expect(isReversibleLock(413, undefined)).toBe(false);
|
|
expect(isReversibleLock(400, 'bad_request')).toBe(false);
|
|
expect(isReversibleLock(413, 'uploads_locked')).toBe(true); // explicit tag still wins
|
|
});
|
|
});
|
|
|
|
/**
|
|
* Regression guard for the queue-rehydration mapping. The bug this locks out: `loadQueue`
|
|
* rebuilt items from IndexedDB WITHOUT copying `lastModified`, so a reloaded item had
|
|
* `lastModified === undefined`. addToQueue's dedup keys on (name, size, lastModified), so
|
|
* re-selecting the same file after a reload would MISS the duplicate and queue it twice.
|
|
*/
|
|
describe('entryToQueueItem', () => {
|
|
const base = {
|
|
id: 'e1',
|
|
userId: 'u1',
|
|
fileName: 'photo.jpg',
|
|
fileSize: 1234,
|
|
lastModified: 1_700_000_000_000,
|
|
mimeType: 'image/jpeg',
|
|
status: 'pending' as const
|
|
};
|
|
|
|
it('carries lastModified across rehydration (dedup depends on it)', () => {
|
|
expect(entryToQueueItem(base).lastModified).toBe(1_700_000_000_000);
|
|
});
|
|
|
|
it('downgrades an interrupted `uploading` entry to `pending` so it resumes', () => {
|
|
expect(entryToQueueItem({ ...base, status: 'uploading' }).status).toBe('pending');
|
|
});
|
|
|
|
it('a `done` entry reports 100% progress; others start at 0', () => {
|
|
expect(entryToQueueItem({ ...base, status: 'done' }).progress).toBe(100);
|
|
expect(entryToQueueItem(base).progress).toBe(0);
|
|
});
|
|
|
|
it('defaults caption/hashtags to empty strings', () => {
|
|
const item = entryToQueueItem(base);
|
|
expect(item.caption).toBe('');
|
|
expect(item.hashtags).toBe('');
|
|
});
|
|
});
|
|
|
|
/**
|
|
* The upload XHR had no timeout of any kind while `processQueue` held the `isProcessing`
|
|
* latch across it. On a half-open socket neither `error` nor `abort` ever fires, so the
|
|
* latch was pinned forever and the whole queue wedged with no recovery but a reload.
|
|
*
|
|
* The policy that matters: bound SILENCE, not total duration. A 500 MB video over a venue
|
|
* uplink legitimately runs 30+ minutes while making steady progress, and a flat total cap
|
|
* would kill exactly the uploads worth keeping.
|
|
*/
|
|
describe('shouldAbortForStall', () => {
|
|
const now = 1_000_000;
|
|
|
|
it('lets a long upload run as long as progress keeps arriving', () => {
|
|
// Two hours in, but progress landed a second ago.
|
|
expect(shouldAbortForStall(now - 1_000, now, false)).toBe(false);
|
|
});
|
|
|
|
it('aborts once the body stalls past the no-progress ceiling', () => {
|
|
expect(shouldAbortForStall(now - 89_000, now, false)).toBe(false);
|
|
expect(shouldAbortForStall(now - 91_000, now, false)).toBe(true);
|
|
});
|
|
|
|
it('applies the wider ceiling once the body is sent and progress goes quiet', () => {
|
|
// Silence that would abort mid-body is normal while waiting for the response.
|
|
expect(shouldAbortForStall(now - 91_000, now, true)).toBe(false);
|
|
expect(shouldAbortForStall(now - 121_000, now, true)).toBe(true);
|
|
});
|
|
});
|