fix(deploy): a permanent upload outage, a dead-on-arrival Caddy, and 11pm commands that don't run
* Caddy had no read_body, on the reasoning that "a slow body still has to actually send bytes". That is an argument about disk, and disk is not the scarce resource: upload_admission budgets concurrent bodies at 4096 MiB and reserves the DECLARED cap, so a video/* upload reserves 500 MiB. Eight connections that stall mid-body hold the whole budget, every other guest waits 20s and gets a 503, and it never recovers on its own — the permit is held until the handler returns. No attacker needed: eight guests starting real videos and walking out of AP range does it, and TCP will not reap those sockets for hours. 30m carries a 500 MB upload at ~2.2 Mbit/s, so it does not fail the uploads this product exists to collect. * APP_PORT is presented in .env.example as an ordinary editable line, while the healthcheck hardcodes 127.0.0.1:3000 and the Caddyfile hardcodes app:3000. Change it and the app boots and serves happily on the new port, the healthcheck fails forever, app never turns healthy — and because caddy is gated on service_healthy, CADDY NEVER STARTS. Port 443 dead for the whole event, sole diagnostic "dependency failed to start". Pinned in compose beside MEDIA_PATH and EXPORT_PATH, which are there for this reason. * Runbook §12's recovery commands do not run as written: unwrapped "$POSTGRES_USER" is expanded by the operator's shell, which does not have it, so psql answers `FATAL: role "" does not exist`. §9 documents that trap two hundred lines earlier and wraps its own calls in sh -c; §12 did not. This is the block you run with the app crash-looping behind a live Caddy. Its DELETE also hard-coded versions 21,22,23 as if to be copied verbatim, on a tree that now has 31 migrations — now explicitly an example, with the instruction to take the numbers from the actual boot error. * Migration counts corrected across the runbook and .env.example (22 -> 31, commit count 154 -> 196). All four were presented as literal command output the operator is invited to reproduce.
This commit is contained in:
@@ -537,9 +537,9 @@ explanation:
|
||||
|
||||
```
|
||||
$ git ls-tree --name-only v0.12.0 backend/migrations/ | wc -l
|
||||
12 # 6 migrations. HEAD has 22.
|
||||
12 # 6 migrations. HEAD has 31.
|
||||
$ git rev-list --count v0.12.0..HEAD
|
||||
154
|
||||
196
|
||||
```
|
||||
|
||||
`db.rs` runs `sqlx::migrate!()` with no `set_ignore_missing`, so an image built from a 6-migration
|
||||
@@ -591,7 +591,7 @@ covers both. It exists so that the rollback line in the emergency card is safe t
|
||||
catastrophic. **If you genuinely need to undo a code change during the event, you cannot; freeze
|
||||
early enough that you never have to.**
|
||||
|
||||
**Across a migration boundary — avoid by freezing.** If you must: all 22 migrations have paired
|
||||
**Across a migration boundary — avoid by freezing.** If you must: all 31 migrations have paired
|
||||
`.down.sql` files, but **none of them removes its own `_sqlx_migrations` row**, so that second step
|
||||
is mandatory and undocumented:
|
||||
|
||||
@@ -799,23 +799,29 @@ differently, the next boot aborts with *"migration 21 was previously applied but
|
||||
modified"*, `main` exits non-zero, and `restart: unless-stopped` makes it **permanent** — with
|
||||
Caddy still routing traffic to the dead container.
|
||||
|
||||
Main-line `021`–`023` are byte-identical to what shipped, so a box that only ever ran tagged
|
||||
Every main-line migration is byte-identical to what shipped, so a box that only ever ran tagged
|
||||
releases is unaffected. **Verify rather than assume** — run this against the server before any
|
||||
deploy:
|
||||
deploy.
|
||||
|
||||
Note the `sh -c` wrapping, for the same reason as §9: `POSTGRES_USER` and `POSTGRES_DB` live in
|
||||
`.env`, which Compose reads and **your shell does not**. Unwrapped, `-U "$POSTGRES_USER"` sends
|
||||
`-U ""` and psql answers `FATAL: role "" does not exist` — at 11pm, with the app crash-looping.
|
||||
|
||||
```bash
|
||||
docker compose exec -T db psql -U "$POSTGRES_USER" "$POSTGRES_DB" -c \
|
||||
"SELECT version, description, success FROM _sqlx_migrations ORDER BY version;"
|
||||
docker compose exec -T db sh -c \
|
||||
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "SELECT version, description, success FROM _sqlx_migrations ORDER BY version;"'
|
||||
```
|
||||
|
||||
If the app is already crash-looping on a renumbered migration, and **only** if you have confirmed
|
||||
the SQL in the new file is equivalent to what was actually applied:
|
||||
the SQL in the new file is equivalent to what was actually applied. Take the version numbers from
|
||||
the crash message and the query above — do **not** copy the ones below, which are an example:
|
||||
|
||||
```bash
|
||||
docker compose stop app
|
||||
docker compose exec -T db psql -U "$POSTGRES_USER" "$POSTGRES_DB" -c \
|
||||
"DELETE FROM _sqlx_migrations WHERE version IN (21,22,23);"
|
||||
docker compose start app # re-applies 021-023, then continues
|
||||
# Replace 21,22,23 with the versions the boot error actually named.
|
||||
docker compose exec -T db sh -c \
|
||||
'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "DELETE FROM _sqlx_migrations WHERE version IN (21,22,23);"'
|
||||
docker compose start app # re-applies exactly those, then continues
|
||||
```
|
||||
|
||||
This re-runs those migrations. They must be idempotent (`IF NOT EXISTS` / `IF EXISTS`) or this
|
||||
|
||||
Reference in New Issue
Block a user