fix(deploy): ship the swap ceilings, pin the last boot-fatal env var, and correct docs that misdirect
* memswap_limit is now IN docker-compose.yml on all four services. Compose sets Memory but leaves MemorySwap unset, and Docker then permits swap equal to the memory limit — so following §5's "add 2 GB of swap" silently DOUBLED every ceiling, to ~5 GiB on a 3.82 GiB box. Nothing OOMs; instead Postgres's working set becomes swap-eligible on a shared-tenancy SSD, turning a bounded OOM-kill that restarts in seconds into unbounded latency with no signal but "everything is slow". The runbook told the operator to hand-add it, which also broke §0's own gate that docker-compose.yml must be unmodified. Verified rather than assumed: service-level memswap_limit does compose with deploy.resources.limits.memory (docker inspect → Memory=1073741824 MemorySwap=1207959552). * DATABASE_MAX_CONNECTIONS pinned in compose. It is the one env var that is now boot-FATAL when unparseable — the right call, but it means a stray quote or a trailing inline comment in .env crash-loops the app behind a live Caddy. MEDIA_PATH, EXPORT_PATH and APP_PORT are pinned for weaker reasons. * .env.example's quota narrative was sized for a CX33: "~30 GB of a fresh 70 GB" on a box with 40 GB. And on THIS box the fixed point never binds at all — ~210 MB/guest is below the 500 MiB floor, so everyone gets the floor and the per-user quota stops bounding aggregate growth. What actually stops uploads is the keepsake preflight at ~8 GB of media. That paragraph is what an operator reads when a guest is blocked, and it pointed at the wrong knob. * The emergency card gains the one disk symptom that can appear mid-event, where `df -h` — its only disk instruction — actively misleads: the gate fires ~10 GB + 2.2x media BEFORE the disk is full, so df shows ~20 GB free at the moment uploads are being refused. * Two code comments that now assert the opposite of the code: claim_job promised that "the update_progress liveness check bails such a worker out early" — it cannot, its predicate is on the job row, which a reopen does not touch, so a mid-export reopen grinds the whole gallery to completion on a 2-vCPU box during the live event. And prune_superseded_archives still argued "deleted bytes cannot be rolled back" as an invariant, after the reclaim path was changed to prune even when that will not close the shortfall. Both now describe what the code does. * Smaller corrections: runbook §3's "two 48 MP photos ≈ 800 MB" scenario is unreachable (compression.rs takes an exclusive heavy permit, so they serialise) and contradicted .env.example; "all four healthy" is wrong since caddy has no healthcheck; a README line reference pointed at a comment added by the same commit that broke it.
This commit is contained in:
35
.env.example
35
.env.example
@@ -45,8 +45,12 @@ DATABASE_URL=postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/events
|
||||
POSTGRES_USER=eventsnap
|
||||
POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password
|
||||
POSTGRES_DB=eventsnap
|
||||
# Connection pool size. The code default is 10 (backend/src/db.rs) — set it explicitly,
|
||||
# because a `.env` written by hand from this file's secrets is otherwise silently on 10.
|
||||
# Connection pool size. The code default is 15 (DEFAULT_MAX_CONNECTIONS in backend/src/db.rs),
|
||||
# and docker-compose.yml pins this value in `app.environment` so an edit here cannot reach the
|
||||
# container. That pin is deliberate: since the value became boot-FATAL when unparseable — so an
|
||||
# operator tuning a knob that never took effect gets told, instead of silently staying on the
|
||||
# default — a stray quote or a trailing inline comment in `.env` would crash-loop the app behind
|
||||
# a live Caddy. Change the pin in compose, not this line.
|
||||
#
|
||||
# SIZE IT TO THE CORES, NOT TO THE GUESTS. The earlier advice here was ~30, reasoned from
|
||||
# "~100 guests polling the feed at once" back when a feed page cost ~449 ms and connections
|
||||
@@ -116,15 +120,26 @@ EXPORT_PATH=/exports
|
||||
# (upload::quota_limit_bytes). Earlier drafts of this file and the runbook both omitted
|
||||
# it and told operators it was inert; it is not.
|
||||
#
|
||||
# It is recomputed against LIVE free space on every upload, so it self-throttles: guests
|
||||
# converge on a fixed point at tolerance/(1+tolerance) of the free space you started
|
||||
# with — 43% at 0.75, i.e. ~30 GB of a fresh 70 GB.
|
||||
# It is recomputed against LIVE free space on every upload, so in principle it self-
|
||||
# throttles: guests converge on a fixed point at tolerance/(1+tolerance) of the free space
|
||||
# you started with — 43% at 0.75.
|
||||
#
|
||||
# Raising it therefore AUTHORISES GUESTS TO FILL MORE OF THE DISK. Setting 0.95 in the
|
||||
# belief that it means "warn me later" moves the fixed point to ~49% and eats the
|
||||
# headroom the keepsake needs — and the keepsake needs a lot, because Gallery.zip and
|
||||
# Memories.zip are each roughly a second copy of every original (both store media
|
||||
# uncompressed). Budget for media + 2x media, or move exports to their own volume.
|
||||
# ON THIS BOX THAT FIXED POINT NEVER BINDS, and it is worth knowing which knob actually
|
||||
# stops the disk filling. The arithmetic above used to be quoted as "~30 GB of a fresh
|
||||
# 70 GB", which is an 80 GB CX33; this deploys to a CX22 with 40 GB. At ~28 GB free and
|
||||
# estimated_guest_count = 100 flooring the divisor, the formula yields ~210 MB per guest —
|
||||
# BELOW the 500 MiB floor — so every guest is granted the floor and the per-user quota
|
||||
# stops bounding aggregate growth at all.
|
||||
#
|
||||
# What actually bounds it is the keepsake preflight in upload.rs: uploads are refused once
|
||||
# free < media x 1.1 x 2 + 10 GB, which on 40 GB lands at ~8 GB of media (README, "Sizing
|
||||
# the disk"). So if a guest reports being blocked, the number to look at is total media,
|
||||
# not this one.
|
||||
#
|
||||
# Raising this still AUTHORISES GUESTS TO FILL MORE OF THE DISK on a larger box, and it
|
||||
# still eats the headroom the keepsake needs — Gallery.zip and Memories.zip are each
|
||||
# roughly a second copy of every original (both store media uncompressed). Budget for
|
||||
# media + 2x media, or move exports to their own volume.
|
||||
#
|
||||
# 0.75 is the tested default. Lower it if the box is tight; raise it only if you have
|
||||
# provisioned export headroom separately.
|
||||
|
||||
Reference in New Issue
Block a user