Files
EventSnap/e2e/docker-compose.sim.yml
MechaCat02 2dd563b3ee
Some checks failed
Audit / cargo audit (backend) (push) Failing after 9m31s
Audit / npm audit (frontend) (push) Successful in 1m2s
Checks / Backend — cargo test + clippy + fmt (push) Failing after 1m7s
Checks / Frontend — vitest + svelte-check (push) Failing after 5m37s
Checks / Keepsake viewer — builds, self-contained, committed artifact in sync (push) Failing after 5m13s
Checks / E2E — typecheck + lint (push) Failing after 36s
E2E / Playwright E2E (chromium + webkit) (push) Failing after 9m8s
E2E / Cross-UA smoke matrix (push) Failing after 4m17s
fix(upload): a truncated body no longer destroys the guest's photo, and the in-app camera can be switched off
Two event-day failures, both frontend-only.

TRUNCATED UPLOADS PURGED THE PHOTO. An iPhone guest uploading from the gallery
inside WhatsApp's browser got "Error parsing `multipart/form-data` request" and
the item went to "Gesperrt" with no retry. That message is AXUM's own multipart
rejection — a plain-text 400, no JSON envelope — which means the request body
never arrived intact. It is a transport failure, not a verdict on the file.

`classifyUploadStatus` maps every 4xx to `terminal`, and terminal PURGES the blob
from IndexedDB and offers no retry. So a webview hiccup deleted the only copy the
guest had, and told them the photo was rejected.

Every 400 the app itself raises carries `bad_request` in a JSON envelope (too
large, wrong type, caption too long, NUL byte), so an unparseable 400 is
distinguishable and is now a NetworkError: blob kept, retry offered. This is the
same rule the 403 branch already applies — "an unparseable body must NOT purge
the blob, losing a photo is the worst outcome" — extended to the status that was
actually hit. Retrying is safe because nothing was parsed, so nothing was stored
and no quota was charged, and `X-Client-Upload-Id` makes a duplicate impossible.

IN-APP CAMERA SWITCH. `PUBLIC_CAMERA_ENABLED=false` removes the "Kamera — Jetzt
aufnehmen" entry from the upload sheet. On some phones `getUserMedia` fails when
switching front/back ("Kamera konnte nicht gestartet werden") or when asked for
video, and those failures are per-device and undiagnosable mid-event; the switch
removes the broken path rather than leaving guests to find it. Nothing is lost:
the gallery picker reaches the phone's own camera app and handles video.

Read at RUNTIME via `$env/dynamic/public`, so flipping it is a compose variable
and `up -d frontend`, not a rebuild. Deliberately NOT routed through the
backend's event payload like `comments_enabled`: that flag describes the event,
this one describes what the client can do — the backend cannot tell a camera
upload from a gallery upload and has no stake in it. Keeping it off the app image
also means no backend release on the day of the event.

The onboarding step and the in-app-browser hint drop their camera wording when it
is off, so no text promises a button that is not there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 09:30:59 +02:00

157 lines
5.8 KiB
YAML

# EventSnap EVENT SIMULATION stack — models the real production box, not CI.
#
# Difference from docker-compose.test.yml (which is tuned for fast, unconstrained
# Playwright runs): this file reproduces the CX22 the event actually runs on —
# * 2 vCPU total, shared by all four services
# * 4 GB RAM, split by the same per-service limits production ships
# * a REAL 30 GB filesystem for media + exports (loopback ext4), so statvfs
# inside the container returns true numbers and the disk gate / 507 path is
# exercised for real rather than simulated.
#
# Why cpuset on every service: production's per-service `cpus` ceilings sum to
# 1.5 + 1.2 + 0.6 + 0.5 = 3.8 on a box that has 2. That oversubscription is the
# point — the ceilings only bind when something else is competing, and cpu_shares
# decides who wins. Reproducing that on a 12-core workstation requires confining
# every container to the SAME two cores; without cpuset each service would get its
# ceiling simultaneously and the contention under test would never happen.
#
# Bring up: docker compose -f docker-compose.sim.yml up -d --build
# Tear down: docker compose -f docker-compose.sim.yml down -v
#
# The 30 GB volume is created out-of-band (see e2e/loadtest/sim-disk.sh) because a
# loopback device must be attached by root; it is declared `external` here.
x-cpuset: &cpuset '0,1'
services:
db:
image: postgres:16-alpine
cpuset: *cpuset
environment:
POSTGRES_USER: eventsnap_test
POSTGRES_PASSWORD: eventsnap_test
POSTGRES_DB: eventsnap_test
healthcheck:
test: ['CMD-SHELL', 'pg_isready -U eventsnap_test -d eventsnap_test']
interval: 3s
timeout: 3s
retries: 30
ports:
- '55433:5432'
volumes:
# Named (not anonymous) so a `down -v` really wipes it and so its on-disk size
# can be measured against the 10 GB DISK_RESERVE that is meant to cover it.
- sim_pgdata:/var/lib/postgresql/data
# Production values, verbatim (docker-compose.yml db service).
deploy:
resources:
limits:
memory: 1G
cpus: '1.5'
reservations:
memory: 256M
cpu_shares: 2048
memswap_limit: 1152m
app:
# The SHIPPED release image, not a local build. Verified identical to HEAD:
# `git diff v0.17.5 HEAD -- backend/` is empty, and v0.17.5/v0.17.6 share one
# app digest on purpose (the v0.17.6 bump was frontend-only). Running the real
# artifact means the simulation tests what the event will actually run.
image: ${SIM_APP_IMAGE:-registry.mc02.dev/eventsnap/app:v0.17.5}
cpuset: *cpuset
depends_on:
db:
condition: service_healthy
environment:
DATABASE_URL: postgres://eventsnap_test:eventsnap_test@db:5432/eventsnap_test
JWT_SECRET: 00112233445566778899aabbccddeeff00112233445566778899aabbccddeeff
# bcrypt("admin-test-pw"), cost 4. $ doubled to escape compose interpolation.
ADMIN_PASSWORD_HASH: $$2b$$04$$XKJJkNX6BOi6y3S42DA5JOWwk4oxc8DHPL6.MrPfJI2vpnccZjP32
EVENT_SLUG: sim-wedding
EVENT_NAME: Hochzeit Simulation
APP_PORT: '3000'
# Both live on the SAME 30 GB filesystem, as they do on the real VPS — but in
# sibling directories, because config.rs::validate requires EXPORT_PATH outside
# MEDIA_PATH (a keepsake archive contains every photo in the event).
MEDIA_PATH: /disk/media
EXPORT_PATH: /disk/exports
SESSION_EXPIRY_DAYS: '30'
# Production pins these; the CI stack leaves them at code defaults. Sized to 2 vCPU.
DATABASE_MAX_CONNECTIONS: '15'
COMPRESSION_WORKER_CONCURRENCY: '2'
# Production disables comments for this event (docker-compose.yml, product decision).
COMMENTS_ENABLED: 'false'
# Mirrors the deployed setting so the harness can audit against the real gate.
KEEPSAKE_ENABLED: ${SIM_KEEPSAKE:-true}
# The ONE deviation from production: enables /admin/__truncate so the harness can
# reset between runs. Never set on the real box.
EVENTSNAP_TEST_MODE: '1'
RUST_LOG: eventsnap_backend=info,tower_http=warn
volumes:
- sim_disk:/disk
deploy:
resources:
limits:
memory: 1G
cpus: '1.2'
reservations:
memory: 256M
cpu_shares: 512
memswap_limit: 1152m
expose:
- '3000'
frontend:
# Shipped release image; `git diff v0.17.6 HEAD -- frontend/` is empty.
image: ${SIM_FE_IMAGE:-registry.mc02.dev/eventsnap/frontend:v0.17.6}
cpuset: *cpuset
depends_on:
- app
environment:
PORT: '3001'
HOST: '0.0.0.0'
ORIGIN: 'http://localhost:3102'
PUBLIC_CAMERA_ENABLED: ${SIM_CAMERA:-true}
deploy:
resources:
limits:
memory: 256M
cpus: '0.6'
cpu_shares: 256
memswap_limit: 320m
expose:
- '3001'
caddy:
image: caddy:2-alpine
cpuset: *cpuset
depends_on:
- app
- frontend
volumes:
- ./Caddyfile.sim:/etc/caddy/Caddyfile:ro
deploy:
resources:
limits:
memory: 256M
cpus: '0.5'
cpu_shares: 1024
memswap_limit: 320m
ports:
- '3102:3102'
volumes:
# 30 GB loopback ext4, created by e2e/loadtest/sim-disk.sh. Holds media AND exports,
# exactly as one VPS disk holds both.
sim_disk:
external: true
name: eventsnap_sim_media
# Postgres data stays on the host disk, NOT on the 30 GB volume. Scoping decision:
# the 30 GB budget under test is the one the APP manages and measures (statvfs on
# MEDIA_PATH drives the disk gate). Putting PG on the same volume would test
# filesystem exhaustion instead, and a full disk under Postgres risks ending the run
# for an infrastructural reason rather than an application one. Its growth is sampled
# separately and reported against DISK_RESERVE_BYTES (10 GB), which exists to cover it.
sim_pgdata: