after_connect puts lock_timeout = 5s on every pooled connection, and the
migrator inherited it. Migrations that take ACCESS EXCLUSIVE — 026's index
swap, 027's ADD COLUMN — then turn a short WAIT into a hard FAILURE.
The runbook installs an hourly pg_dump (§10.2) and tells the operator to back
up before deploying; pg_dump holds ACCESS SHARE on `upload` and `"user"` for
its whole run, and the runbook is full of psql snippets that do the same. Boot
into that window and the migration aborts, create_pool errors, main exits 1,
and `restart: unless-stopped` crash-loops the app behind a live Caddy. The
rollback is clean and a later retry succeeds, which is precisely what makes it
a baffling intermittent outage rather than an obvious one.
026's own comment reasons that "this runs at boot before the server accepts
requests, so the brief lock costs nothing" — true of the app's own sessions,
and it does not cover anything else on the database.
statement_timeout is dropped for the migrator too: a migration on a real table
can legitimately outlast the 15s a request is allowed.
Four migrations, all additive against a database that already has 001-025 applied.
026 narrows the client-upload idempotency index with `AND deleted_at IS NULL`. The old
index made a soft-deleted row keep its key forever, so a guest who deleted a photo and
re-sent the same one had the retry silently swallowed. The new indexed set is a strict
subset of the old, so it cannot fail on existing rows.
027 adds `client_join_id`, which lets a join retry after a lost response resume the same
account instead of 409ing on a name the caller itself owns. Every existing row gets NULL
and the partial index excludes NULLs, so it indexes nothing at creation.
028 brings the feed view's like/comment counts in line with what the feed actually renders:
a banned guest's rows were still counted, so a card showed "3 comments" above two.
`comment.rs` gets the matching `NOT u.is_banned` on the live read path — the export and
hashtag queries already filtered it, so the two views of one moderation action disagreed.
029 records host moderation actions, which were previously invisible after the fact.
Verified by applying 001-029 to a real Postgres against seeded data, including a
soft-deleted row holding a key and a banned user's like and comment. 026's down-migration
legitimately fails where a deleted and a live row share a key — that is inherent to the
direction, documented in the file, and sqlx never runs downs at boot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`/health` returned the literal string "ok" and touched nothing. Every request in this
app needs the database, so that answered a question nobody asked: the container
reported healthy while every request 500'd, and with no operator watching during the
event there was no signal at all. It now runs `SELECT 1` with a 2s timeout.
Verified live against the production stack: 200, stop Postgres, 503 "database timeout",
start Postgres, 200 again — with no app restart, because sqlx revalidates on acquire.
Deliberately NOT wired to automatic recovery. Compose's `restart: unless-stopped` does
not react to healthcheck state anyway, and an autoheal sidecar would be actively wrong
here: it would truncate every in-flight upload to "fix" an outage that, as the test
above shows, clears on its own. This is a diagnostic — including for the runbook's
event-day `curl`.
The pool had only `max_connections` set. Three additions:
- `acquire_timeout(5s)`. sqlx defaults to 30, so a DB blip parked every request AND all
~100 SSE session revalidations for half a minute before erroring — the app looked
hung rather than degraded, and the backlog outlived the blip.
- `min_connections(2)`, so the first request after the setup-to-guests-arriving gap
doesn't pay TCP + auth.
- `statement_timeout=15s` / `lock_timeout=5s` per connection. Without them a single
pathological query holds a pool slot indefinitely and no client-side timeout can take
it back, because the slot is only released when Postgres finishes.
Those two SETs are sent as two statements. `sqlx::query` uses the extended query
protocol, which permits exactly one per call — as `SET a; SET b` every new connection
failed, which surfaced as the pool never opening one and `create_pool` reporting a
connect timeout. Caught by booting against an empty database rather than a warm one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
README step 2 said to set "DOMAIN, JWT_SECRET, ADMIN_PASSWORD_HASH, EVENT_NAME, etc."
-- POSTGRES_PASSWORD was not in that list. .env.example said to set it and keep it in
sync with DATABASE_URL. The two documents disagreed, and both readings ended badly.
Branch A, README verbatim: the stack came up GREEN and healthy on
CHANGE_ME_use_a_strong_password -- a database credential published in the public repo.
The production secret guard covered JWT_SECRET and ADMIN_PASSWORD_HASH, and nothing
anywhere looked at the Postgres password.
Branch B, .env.example verbatim: a permanent restart loop, "password authentication
failed for user eventsnap".
The guard's own design made Branch B near-certain. It stops the APP on the first
`docker compose up -d` -- but not the `db` service in that same command, which
initialises its data directory and bakes in whatever password was in .env at that
moment. POSTGRES_PASSWORD is honoured ONLY at initdb. So the intended recovery -- see
the refusal, fix your secrets, boot again -- was exactly the sequence that broke it.
Nothing in the error named the cause, and the remedy (`down -v`) is both unguessable
and the one command you must never run once real data exists.
Three changes, which have to ship together: the guard alone would just move operators
out of Branch A and into Branch B.
- The guard now rejects a placeholder DATABASE_URL in production (the password rides
in that URL, which is what the app actually reads). Branch A can no longer boot.
- It reports EVERY unset secret in one message instead of returning on the first.
Fixing two secrets used to cost two boot cycles, on a stack where Caddy waits on the
unhealthy app throughout, and each avoidable cycle is another chance to reach for -v.
- A 28P01 handler in db.rs turns the unguessable failure into a self-explaining one:
it names the initdb semantics, gives `down -v` with an explicit "deletes db + media +
exports, no undo", and gives the ALTER ROLE alternative for when data already exists.
Docs: step 2 now names POSTGRES_PASSWORD and says every secret must be set BEFORE the
first up; the troubleshooting block covers the auth-failure loop and both remedies.
Also replaces htpasswd (apache2-utils -- not on a stock VPS) with
`docker run --rm caddy:2-alpine caddy hash-password`, an image the stack already pulls,
in README, .env.example and the guard's own message. Verified the output ($2a$14)
against the shipped $2y$12 example.
Verified on a real Postgres, not just in tests: Branch A refuses and names DATABASE_URL;
all three placeholders report in one boot; a volume initialised with one password and
connected to with another prints the diagnostic; an unrelated connect failure (dead
port) stays silent. 8 unit tests, including that non-prod ignores all of it so the e2e
stack is unaffected.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The backend had never been run through rustfmt. Doing it in one mechanical pass (134 files)
so no future functional diff is buried under formatting churn, then gating `cargo fmt
--check` in checks.yml so it stays clean.
Formatting only — no logic, SQL, or behaviour changed. Verified after the reformat:
cargo test 56 passed, clippy --all-targets -D warnings clean, cargo fmt --check clean.
This is the deferred cleanup noted when CI's Format step was first left out.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Follow-up to the comprehensive code review. Five batches:
1. Cross-event authorization: host_delete_upload, unban_user, and
host_delete_comment now scope by auth.event_id. Adds
Upload::find_by_id_and_event / soft_delete_in_event and a
Comment::soft_delete_in_event variant that joins through upload.
2. Token exposure: SSE auth no longer puts the JWT in the URL.
New /api/v1/stream/ticket endpoint mints a short-lived single-use
ticket bound to the session; the EventSource passes ?ticket=...
instead. Refuse to start in APP_ENV=production with the dev JWT
sentinel; warn loudly otherwise.
3. Account hardening: per-IP+name rate limit on /recover (mitigates
targeted lockout DoS), per-IP rate limit on /admin/login, random
32-char admin recovery PIN (replaces "0000"), structured tracing
events for wrong PIN, lockout, failed admin login, ban/unban/role
change/pin-reset/host-delete.
4. DoS / correctness: comment listing paginated (LIMIT 50 + ?before=
cursor), hashtag extraction whitelisted to ASCII alnum+underscore
(≤40 chars) with unit tests, display_name / caption / comment body
length validated in chars rather than bytes.
5. Cleanup: session-touch failures now logged, DATABASE_MAX_CONNECTIONS
env var (default 10).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>