* memswap_limit is now IN docker-compose.yml on all four services. Compose sets Memory but leaves MemorySwap unset, and Docker then permits swap equal to the memory limit — so following §5's "add 2 GB of swap" silently DOUBLED every ceiling, to ~5 GiB on a 3.82 GiB box. Nothing OOMs; instead Postgres's working set becomes swap-eligible on a shared-tenancy SSD, turning a bounded OOM-kill that restarts in seconds into unbounded latency with no signal but "everything is slow". The runbook told the operator to hand-add it, which also broke §0's own gate that docker-compose.yml must be unmodified. Verified rather than assumed: service-level memswap_limit does compose with deploy.resources.limits.memory (docker inspect → Memory=1073741824 MemorySwap=1207959552). * DATABASE_MAX_CONNECTIONS pinned in compose. It is the one env var that is now boot-FATAL when unparseable — the right call, but it means a stray quote or a trailing inline comment in .env crash-loops the app behind a live Caddy. MEDIA_PATH, EXPORT_PATH and APP_PORT are pinned for weaker reasons. * .env.example's quota narrative was sized for a CX33: "~30 GB of a fresh 70 GB" on a box with 40 GB. And on THIS box the fixed point never binds at all — ~210 MB/guest is below the 500 MiB floor, so everyone gets the floor and the per-user quota stops bounding aggregate growth. What actually stops uploads is the keepsake preflight at ~8 GB of media. That paragraph is what an operator reads when a guest is blocked, and it pointed at the wrong knob. * The emergency card gains the one disk symptom that can appear mid-event, where `df -h` — its only disk instruction — actively misleads: the gate fires ~10 GB + 2.2x media BEFORE the disk is full, so df shows ~20 GB free at the moment uploads are being refused. * Two code comments that now assert the opposite of the code: claim_job promised that "the update_progress liveness check bails such a worker out early" — it cannot, its predicate is on the job row, which a reopen does not touch, so a mid-export reopen grinds the whole gallery to completion on a 2-vCPU box during the live event. And prune_superseded_archives still argued "deleted bytes cannot be rolled back" as an invariant, after the reclaim path was changed to prune even when that will not close the shortfall. Both now describe what the code does. * Smaller corrections: runbook §3's "two 48 MP photos ≈ 800 MB" scenario is unreachable (compression.rs takes an exclusive heavy permit, so they serialise) and contradicted .env.example; "all four healthy" is wrong since caddy has no healthcheck; a README line reference pointed at a comment added by the same commit that broke it.
283 lines
15 KiB
YAML
283 lines
15 KiB
YAML
# Docker's default json-file driver has NO rotation at all, and every container writes to the
|
|
# same filesystem as postgres_data, media_data and exports_data. Filling that filesystem does
|
|
# not degrade one subsystem — Postgres stops being able to write and the whole event goes down
|
|
# (see README "Sizing the disk"). This caps logs at 30 MB per service, permanently.
|
|
#
|
|
# Paired with RUST_LOG in .env: without it the app falls back to `eventsnap_backend=debug,
|
|
# tower_http=debug` (main.rs), which is a debug line per HTTP request including every preview.
|
|
x-logging: &default-logging
|
|
driver: json-file
|
|
options:
|
|
max-size: "10m"
|
|
max-file: "3"
|
|
|
|
services:
|
|
db:
|
|
image: postgres:16-alpine
|
|
restart: unless-stopped
|
|
logging: *default-logging
|
|
env_file: .env
|
|
environment:
|
|
POSTGRES_USER: ${POSTGRES_USER}
|
|
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
|
|
POSTGRES_DB: ${POSTGRES_DB}
|
|
volumes:
|
|
- postgres_data:/var/lib/postgresql/data
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U ${POSTGRES_USER} -d ${POSTGRES_DB}"]
|
|
interval: 5s
|
|
timeout: 5s
|
|
retries: 10
|
|
deploy:
|
|
resources:
|
|
limits:
|
|
# 1G, not 512M. Postgres 16's default shared_buffers plus a pool of backends
|
|
# leaves very little headroom at 512M, and an OOM here does not degrade one
|
|
# feature — it takes the event down, because every request path touches the
|
|
# database.
|
|
#
|
|
# `.env.example` now sets DATABASE_MAX_CONNECTIONS=15, sized to the 2 vCPU this
|
|
# box has rather than to the guest count: since migration 024 replaced the feed
|
|
# view's GROUP BY with scalar subqueries, a feed page costs well under a
|
|
# millisecond, so connections are no longer spent waiting. Raising it back
|
|
# toward 30 means raising this limit with it.
|
|
memory: 1G
|
|
# CPU ceiling. Postgres is the one service that must never be starved: every request
|
|
# path touches it, so a CPU-bound image resize elsewhere degrades the whole event
|
|
# rather than one feature. 1.5 of 2 cores is a ceiling, not a reservation — it only
|
|
# binds when something else is competing.
|
|
cpus: '1.5'
|
|
reservations:
|
|
# Memory floor only. `reservations.cpus` USED TO BE HERE and did nothing: outside
|
|
# Swarm, `docker compose up` silently drops it — verified by inspecting a running
|
|
# container, where CpuShares, CpuQuota and CpusetCpus were all unset while
|
|
# `limits.cpus` and `reservations.memory` came through as NanoCpus and
|
|
# MemoryReservation. So the comment claiming it was "the piece that actually
|
|
# protects the database" described a guarantee the box never had.
|
|
#
|
|
# It matters on a CX22: the ceilings below sum to 1.2 + 0.6 + 0.5 = 2.3 on 2 vCPU,
|
|
# so the other services CAN oversubscribe the machine, and with every container on
|
|
# the default weight Postgres competes on equal footing with two image resizes and
|
|
# an ffmpeg poster. `cpu_shares` is the knob that survives the translation — see the
|
|
# weights on each service.
|
|
memory: 256M
|
|
# Relative CPU weight under contention (Docker default is 1024). Only consulted when the
|
|
# CPU is actually saturated, which is exactly the moment the database must not lose.
|
|
cpu_shares: 2048
|
|
# Caps memory+swap together, so the `memory` limit above stays the real ceiling.
|
|
#
|
|
# Compose sets `Memory` but leaves `MemorySwap` unset, and Docker then permits swap EQUAL to
|
|
# the memory limit — so following the runbook's "add 2 GB of swap" step silently DOUBLES every
|
|
# container ceiling, to 5 GiB of ceilings on a 3.82 GiB box. Nothing OOMs; instead Postgres's
|
|
# working set becomes swap-eligible on a shared-tenancy VPS SSD, turning a bounded OOM-kill
|
|
# (which restarts in seconds) into unbounded latency everywhere with no signal but "it's slow".
|
|
#
|
|
# The runbook used to tell the operator to add this BY HAND, which also broke its own §0 gate
|
|
# requiring docker-compose.yml to be unmodified. Shipped here instead. 1152m against a 1G limit
|
|
# leaves 128 MB of swap — enough to absorb a spike, not enough to hide one.
|
|
#
|
|
# Verified rather than assumed: service-level `memswap_limit` DOES compose with
|
|
# `deploy.resources.limits.memory` — `docker inspect` reports Memory=1073741824
|
|
# MemorySwap=1207959552.
|
|
memswap_limit: 1152m
|
|
|
|
app:
|
|
# Production PULLS a prebuilt image; it never compiles. A release build of this crate is
|
|
# fat-LTO over 427 dependencies (see backend/Cargo.toml [profile.release]) and peaks well
|
|
# above the RAM a 4 GB box has spare with the stack running — and a rollback would be a
|
|
# second build under pressure. Images are built on a workstation and pushed; see
|
|
# docker-compose.build.yml and DEPLOYMENT_RUNBOOK.md.
|
|
#
|
|
# There is deliberately NO `build:` key here: without one, a wrong tag fails instantly with
|
|
# "manifest unknown" instead of silently starting a 45-minute compile on the event server.
|
|
# The `:?` form fails loudly on an unset variable rather than resolving to an empty tag.
|
|
image: registry.mc02.dev/eventsnap/app:${EVENTSNAP_VERSION:?set EVENTSNAP_VERSION in .env}
|
|
restart: unless-stopped
|
|
logging: *default-logging
|
|
env_file: .env
|
|
environment:
|
|
# Default to info. Without this the code fallback in main.rs applies, which is
|
|
# `eventsnap_backend=debug,tower_http=debug` — a line per HTTP request AND per
|
|
# response, including every preview fetch, for a multi-day run. The x-logging cap
|
|
# above bounds the disk cost but not the CPU/IO one.
|
|
#
|
|
# Set here rather than only in `.env` because a stock deploy sets RUST_LOG nowhere,
|
|
# and this is the layer an operator will actually find when they need to raise it
|
|
# for a single event (`RUST_LOG=eventsnap_backend=debug docker compose up -d app`).
|
|
RUST_LOG: ${RUST_LOG:-info}
|
|
# Activates the production secret guard in config.rs — refuses to boot with
|
|
# placeholder JWT_SECRET / ADMIN_PASSWORD_HASH.
|
|
APP_ENV: production
|
|
# Pinned beside MEDIA_PATH for the same reason, and because nothing validates it:
|
|
# config.rs defaults it to /exports but never checks that it is a mount or that it
|
|
# differs from media_path. A stray EXPORT_PATH in .env builds the keepsake into the
|
|
# container's writable layer, where it passes every health check and disk preflight
|
|
# and then evaporates on the next `up -d`.
|
|
EXPORT_PATH: /exports
|
|
# The media volume is mounted at /media (below), so the app MUST write there.
|
|
# Pin it here rather than trusting .env: if MEDIA_PATH in .env points elsewhere
|
|
# (e.g. a host path used for running the backend natively) the container can't
|
|
# create it and every upload 500s with EACCES. `environment` overrides `env_file`,
|
|
# so this is authoritative for the container.
|
|
MEDIA_PATH: /media
|
|
# Third member of the MEDIA_PATH / EXPORT_PATH family, and the nastiest of the three because
|
|
# the app itself reports nothing wrong. The healthcheck below hardcodes 127.0.0.1:3000 and
|
|
# the Caddyfile hardcodes app:3000, while `.env.example` presents APP_PORT as an ordinary
|
|
# editable line. Change it there and the app boots and serves happily on the new port, the
|
|
# healthcheck fails forever, `app` never turns healthy — and because caddy is gated on
|
|
# `service_healthy`, CADDY NEVER STARTS AT ALL. Port 443 is dead for the whole event and the
|
|
# only diagnostic is `dependency failed to start`.
|
|
APP_PORT: "3000"
|
|
# Fourth member of the family, pinned for a reason the other three don't have: this one is
|
|
# boot-FATAL. `db.rs` rejects an unparseable value with `bail!` rather than falling back to
|
|
# the default (right call — an operator tuning a knob that silently never applied is worse),
|
|
# which means a stray quote, a trailing inline comment, or a smart-quote pasted into `.env`
|
|
# no longer degrades anything: it exits 1, and `restart: unless-stopped` crash-loops the app
|
|
# behind a live Caddy. `.trim()` covers whitespace and CRLF; it cannot cover those.
|
|
#
|
|
# Sized to the 2 vCPU this box has, not to the guest count — see `.env.example` and the
|
|
# `db` memory limit, which must be raised together with this.
|
|
DATABASE_MAX_CONNECTIONS: "15"
|
|
# Pinned for the same reason as MEDIA_PATH: `environment` beats `env_file`, so this cannot
|
|
# be lost by an operator who copies `.env.example` and edits only the secrets — which is
|
|
# the likely path, and `.env.example` ships the generic default of `true`.
|
|
#
|
|
# This is a product decision for this event, not a technical one: guests should be present
|
|
# at the party, not in a comment thread. Likes and captions stay on and are unaffected.
|
|
# Boot-time only, so changing it means `docker compose up -d`, not an admin toggle.
|
|
# To re-enable comments, delete this line and set COMMENTS_ENABLED in .env.
|
|
COMMENTS_ENABLED: "false"
|
|
depends_on:
|
|
db:
|
|
condition: service_healthy
|
|
# Longer than the app's own 10s shutdown backstop (main.rs SHUTDOWN_GRACE), because Docker's
|
|
# default stop timeout is ALSO 10s — so a redeploy raced the graceful drain and could SIGKILL
|
|
# the process at the exact moment it was finishing, truncating the in-flight upload the
|
|
# graceful shutdown exists to protect. The app always exits well before 20s.
|
|
stop_grace_period: 20s
|
|
volumes:
|
|
- media_data:/media
|
|
# Export archives live OUTSIDE /media so the public media ServeDir can't
|
|
# serve them — downloads go only through the ticket-gated handler.
|
|
- exports_data:/exports
|
|
expose:
|
|
- "3000"
|
|
healthcheck:
|
|
# Use 127.0.0.1, NOT localhost: the app binds IPv4 (0.0.0.0) but `localhost`
|
|
# resolves to ::1 (IPv6) first inside the container, so a localhost probe gets
|
|
# "connection refused" and the container never turns healthy — which would leave
|
|
# Caddy (gated on `condition: service_healthy` below) blocked forever on boot.
|
|
test: ["CMD-SHELL", "wget -q -O- http://127.0.0.1:3000/health || exit 1"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 20s
|
|
deploy:
|
|
resources:
|
|
limits:
|
|
# Bounds a runaway ffmpeg transcode (large uploads, 2 workers) so it can't
|
|
# OOM the single box and take down Postgres.
|
|
memory: 1G
|
|
# CPU ceiling for the two image workers + ffmpeg poster extraction. Bounded below
|
|
# 2.0 so the app can never take both cores on its own.
|
|
# COMPRESSION_WORKER_CONCURRENCY=2 is the memory bound; this is the CPU one.
|
|
cpus: '1.2'
|
|
# Half the default weight, and this is the ceiling's other half: the cap alone leaves
|
|
# 0.8 vCPU for db + frontend + caddy, which frontend and caddy can consume between them.
|
|
# Compression is throughput work with no guest waiting on it, so it yields to Postgres —
|
|
# which every request path, including the app's own, is blocked on.
|
|
cpu_shares: 512
|
|
# See the `db` service for why this is shipped rather than hand-added: without it, the
|
|
# runbook's swap step doubles this ceiling. 1152m against a 1G limit.
|
|
memswap_limit: 1152m
|
|
|
|
frontend:
|
|
# Pulled, not built — see the note on `app` above.
|
|
image: registry.mc02.dev/eventsnap/frontend:${EVENTSNAP_VERSION:?set EVENTSNAP_VERSION in .env}
|
|
restart: unless-stopped
|
|
logging: *default-logging
|
|
env_file: .env
|
|
environment:
|
|
# adapter-node behind Caddy TLS needs the public origin for CSRF checks on
|
|
# POST form actions — without it they fail only in production.
|
|
# `:?` for the same reason EVENTSNAP_VERSION uses it. A blank DOMAIN doesn't fail — it
|
|
# produces `https://` here and collapses the Caddyfile's site block below, so the stack
|
|
# comes up with no TLS and no site and the only symptom is a browser error.
|
|
ORIGIN: "https://${DOMAIN:?set DOMAIN in .env}"
|
|
# V8 sizes its old-space heap from the cgroup limit, but lands on ~101% of it (measured:
|
|
# heap_size_limit 259 MB inside a 256M container). So the JS heap ceiling sits ABOVE the
|
|
# container's entire budget — before base RSS (~60-90 MB), the C++ heap, or SSR response
|
|
# buffers, which are external memory V8 doesn't count at all. The practical effect is that
|
|
# V8 can never reach its own limit and run an emergency GC, so the only backpressure is a
|
|
# kernel SIGKILL: an arrival burst of ~100 guests SSR-rendering /join and /feed OOM-kills
|
|
# node, guests get the "Wir sind gleich zurück" page, it restarts, and the burst is still
|
|
# there. Setting the ceiling below the cgroup limit restores GC as the first line of defence.
|
|
NODE_OPTIONS: "--max-old-space-size=160"
|
|
depends_on:
|
|
- app
|
|
expose:
|
|
- "3001"
|
|
healthcheck:
|
|
# 127.0.0.1, not localhost — see the app healthcheck note above (IPv4 bind vs
|
|
# ::1 resolution would leave this container permanently unhealthy).
|
|
test: ["CMD-SHELL", "wget -q -O- http://127.0.0.1:3001/ >/dev/null 2>&1 || exit 1"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 15s
|
|
deploy:
|
|
resources:
|
|
limits:
|
|
memory: 256M
|
|
# Node SSR is bursty and not latency-critical for guests (the app is CSR after the
|
|
# first paint), so it yields first under contention.
|
|
cpus: '0.6'
|
|
# Lowest weight of the four, for the reason above: `ssr = false`, so this serves the shell
|
|
# and then guests talk to `app` directly. A slow shell delays a reload; a slow database
|
|
# breaks the event.
|
|
cpu_shares: 256
|
|
# See the `db` service for why this is shipped rather than hand-added: without it, the
|
|
# runbook's swap step doubles this ceiling. 320m against a 256M limit.
|
|
memswap_limit: 320m
|
|
|
|
caddy:
|
|
image: caddy:2-alpine
|
|
restart: unless-stopped
|
|
logging: *default-logging
|
|
environment:
|
|
# The Caddyfile's site address is `{$DOMAIN}`, read from THIS container's env.
|
|
# Without it, `{$DOMAIN}` expands to empty, the site block collapses, and Caddy
|
|
# serves nothing / fails to obtain a TLS cert. `env_file` alone wouldn't help —
|
|
# Caddy needs it in `environment`, and this keeps the Caddyfile the single source.
|
|
DOMAIN: ${DOMAIN:?set DOMAIN in .env}
|
|
ports:
|
|
- "80:80"
|
|
- "443:443"
|
|
volumes:
|
|
- ./Caddyfile:/etc/caddy/Caddyfile:ro
|
|
- caddy_data:/data
|
|
depends_on:
|
|
app:
|
|
condition: service_healthy
|
|
frontend:
|
|
condition: service_healthy
|
|
deploy:
|
|
resources:
|
|
limits:
|
|
memory: 256M
|
|
# TLS termination and static serving. Small but must stay responsive — a starved
|
|
# reverse proxy makes every service look down.
|
|
cpus: '0.5'
|
|
# Left at the Docker default (1024). Caddy is cheap but sits in front of everything, so
|
|
# it must not be the bottleneck; it is capped at 0.5 vCPU regardless.
|
|
cpu_shares: 1024
|
|
# See the `db` service for why this is shipped rather than hand-added: without it, the
|
|
# runbook's swap step doubles this ceiling. 320m against a 256M limit.
|
|
memswap_limit: 320m
|
|
|
|
volumes:
|
|
postgres_data:
|
|
media_data:
|
|
exports_data:
|
|
caddy_data:
|