Three deploy-readiness gaps, none of them in the export invariant itself. 1. The escape hatch was unreachable. `POST /host/export/rebuild` was added as the fix for "a failed export is terminal", but nothing in the frontend called it — so recovery required a terminal, the API docs and a valid JWT. The host page instead told the host to reopen uploads and re-release, which is the destructive act the endpoint exists to avoid (it unlocks the gallery to every guest and retracts the release). The failed state now offers "Erneut versuchen"; a ready keepsake can be rebuilt behind a confirm, since a rebuild makes it briefly undownloadable. 2. Migration 014 had never been run against anything. It is the repo's first destructive migration and its backfill has to carry the old notion of "downloadable" across exactly — get it wrong and a released event's keepsake silently 404s, on data that cannot be re-created. backend/scripts/rehearse-014.sh proves it on a throwaway postgres, against a real pg_dump or a synthetic DB covering every pre-014 state, asserting up, down and up-again all preserve downloadability. Verified non-vacuous: retiring nothing (ELSE -1 -> ELSE e.export_epoch) fails it, naming the three keepsakes it would resurrect. It also surfaced that the down migration is an inverse UP TO DRIFT: a ready flag that disagreed with its job row comes back FALSE. That heals corruption rather than restoring it, and costs nothing (no file_path to serve), so the assertion checks what actually matters — no genuinely downloadable keepsake loses its flag — and reports the normalisation instead of failing on it. 3. No advisory scanning. cargo audit + npm audit, in .github/workflows/ because Gitea Actions scans it too and e2e.yml already resolves actions/* from GitHub. The runner is not assumed to have cargo. RUSTSEC-2023-0071 (rsa/Marvin) is ignored SPECIFICALLY, not globally: it is unreachable here (HS256 + bcrypt, Postgres only) and unpatched, so failing on it would mean a permanently red pipeline everyone learns to ignore. Tests: e2e pins that a failed keepsake recovers via rebuild while the event stays released and uploads stay locked — so a future refactor implementing rebuild as an internal reopen+re-release fails — and that rebuild cannot publish an unreleased gallery. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
196 lines
9.7 KiB
Bash
Executable File
196 lines
9.7 KiB
Bash
Executable File
#!/usr/bin/env bash
|
||
#
|
||
# Dress rehearsal for migration 014 (export_epoch) — the repo's first DESTRUCTIVE migration.
|
||
#
|
||
# 014 DROPs event.export_zip_ready / export_html_ready and rewrites export_job.release_seq into an
|
||
# epoch. The dangerous part is not the DDL, it's the BACKFILL: it has to carry the old notion of
|
||
# "downloadable" across exactly. Get it wrong and a released event's keepsake silently 404s for
|
||
# every guest, on a dataset you cannot re-create.
|
||
#
|
||
# This script proves the backfill preserves downloadability, against a scratch copy of a REAL dump:
|
||
#
|
||
# ./rehearse-014.sh /path/to/prod-dump.sql # rehearse against production data
|
||
# ./rehearse-014.sh # synthetic: build a pre-014 DB covering every case
|
||
#
|
||
# It asserts:
|
||
# 1. up: {(event,type) downloadable BEFORE} == {(event,type) downloadable AFTER}
|
||
# 2. down: the old ready flags come back identical (so a rollback is actually a rollback)
|
||
# 3. up again: still identical (so a roll-forward after a rollback is safe)
|
||
#
|
||
# It NEVER touches your real database. Everything runs in a throwaway container.
|
||
#
|
||
# NOTE ON ROLLBACK (see the runbook header in 014_export_epoch.up.sql): sqlx runs migrations before
|
||
# serving and errors with VersionMissing on an unknown version, so the PREVIOUS image will not boot
|
||
# against the 014 schema — it crash-loops. Rolling back means running 014_export_epoch.down.sql BY
|
||
# HAND FIRST, then deploying the old image. Assertion 2 is what makes that safe. Rehearse it before
|
||
# you need it, not while the party is happening.
|
||
|
||
set -euo pipefail
|
||
|
||
DUMP="${1:-}"
|
||
MIG_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/../migrations" && pwd)"
|
||
CONTAINER="eventsnap-rehearse-014"
|
||
PGPASSWORD="rehearse"
|
||
DB="eventsnap_rehearsal"
|
||
|
||
cleanup() { docker rm -f "$CONTAINER" > /dev/null 2>&1 || true; }
|
||
trap cleanup EXIT
|
||
cleanup
|
||
|
||
echo "▸ Starting a throwaway postgres:16 (nothing here touches your real DB)…"
|
||
docker run --rm -d --name "$CONTAINER" \
|
||
-e POSTGRES_PASSWORD="$PGPASSWORD" -e POSTGRES_DB="$DB" \
|
||
postgres:16-alpine > /dev/null
|
||
|
||
psql() { docker exec -i -e PGPASSWORD="$PGPASSWORD" "$CONTAINER" psql -U postgres -d "$DB" -v ON_ERROR_STOP=1 "$@"; }
|
||
|
||
for _ in $(seq 1 30); do
|
||
if docker exec "$CONTAINER" pg_isready -U postgres > /dev/null 2>&1; then break; fi
|
||
sleep 1
|
||
done
|
||
|
||
if [[ -n "$DUMP" ]]; then
|
||
echo "▸ Restoring $DUMP …"
|
||
psql -q < "$DUMP"
|
||
else
|
||
echo "▸ No dump given — building a synthetic pre-014 DB (migrations 001→013)…"
|
||
for f in "$MIG_DIR"/0[0-1][0-9]_*.up.sql; do
|
||
[[ "$(basename "$f")" == 014_* ]] && continue
|
||
psql -q < "$f"
|
||
done
|
||
|
||
# Every state the backfill has to classify. If 014 mishandles ANY of these, assertion 1 fails.
|
||
# Every state the backfill has to classify. export_job is UNIQUE (event_id, type), so each event
|
||
# has at most one job per type — the cases are distinguished by event, not by piling up rows.
|
||
# A: released, both flags, both done → both stay downloadable
|
||
# B: released, ZIP flag only → ZIP downloadable, HTML not (a half-built keepsake)
|
||
# C: released, jobs done, NO flags → NOT downloadable (a superseded worker's finished job:
|
||
# the exact state the old ready-flag model used to leak)
|
||
# D: never released, jobs done → NOT downloadable (reopened, or built before release)
|
||
# E: released, ZIP flag set but the job is FAILED/RUNNING → NOT downloadable (flag/job disagree,
|
||
# which the old two-source model made representable at all)
|
||
echo "▸ Seeding: released+both, zip-only, done-but-stale, unreleased, flag/job-disagreement…"
|
||
psql -q <<'SQL'
|
||
INSERT INTO event (id, slug, name, export_released_at, export_zip_ready, export_html_ready) VALUES
|
||
('aaaaaaaa-0000-0000-0000-000000000001','ev-a','A', now(), TRUE, TRUE),
|
||
('aaaaaaaa-0000-0000-0000-000000000002','ev-b','B', now(), TRUE, FALSE),
|
||
('aaaaaaaa-0000-0000-0000-000000000003','ev-c','C', now(), FALSE, FALSE),
|
||
('aaaaaaaa-0000-0000-0000-000000000004','ev-d','D', NULL, FALSE, FALSE),
|
||
('aaaaaaaa-0000-0000-0000-000000000005','ev-e','E', now(), TRUE, TRUE);
|
||
|
||
INSERT INTO export_job (event_id, type, status, progress_pct, file_path, release_seq)
|
||
SELECT e.id, t::export_type, 'done'::export_status, 100, '/x/' || e.slug || '.' || t, 0
|
||
FROM event e CROSS JOIN (VALUES ('zip'),('html')) AS v(t);
|
||
|
||
-- E: the flags claim ready, the jobs say otherwise. Old model required BOTH, so neither is
|
||
-- downloadable — the migration must not "trust the flag" and resurrect them.
|
||
UPDATE export_job SET status = 'failed', file_path = NULL, progress_pct = 40
|
||
WHERE event_id = 'aaaaaaaa-0000-0000-0000-000000000005' AND type = 'zip';
|
||
UPDATE export_job SET status = 'running', file_path = NULL, progress_pct = 70
|
||
WHERE event_id = 'aaaaaaaa-0000-0000-0000-000000000005' AND type = 'html';
|
||
SQL
|
||
fi
|
||
|
||
echo "▸ Snapshotting what is downloadable under the OLD model…"
|
||
psql -q <<'SQL'
|
||
CREATE TABLE _rehearsal_before AS
|
||
SELECT j.event_id, j.type
|
||
FROM export_job j JOIN event e ON e.id = j.event_id
|
||
WHERE e.export_released_at IS NOT NULL
|
||
AND j.status = 'done'
|
||
AND ((j.type = 'zip' AND e.export_zip_ready) OR (j.type = 'html' AND e.export_html_ready));
|
||
|
||
-- For assertion 2: the exact flags a rollback has to reproduce.
|
||
CREATE TABLE _rehearsal_flags AS
|
||
SELECT id, export_zip_ready, export_html_ready FROM event;
|
||
SQL
|
||
|
||
before=$(psql -tAc "SELECT count(*) FROM _rehearsal_before")
|
||
echo " → $before (event,type) pair(s) downloadable before."
|
||
|
||
# The assertion. A symmetric difference of zero means the backfill carried the old notion of
|
||
# "downloadable" across EXACTLY: nothing gained (a stale keepsake resurrected), nothing lost (a
|
||
# guest's keepsake silently 404s).
|
||
assert_up_preserves() {
|
||
local label="$1" diff
|
||
diff=$(psql -tAc "
|
||
SELECT count(*) FROM (
|
||
(SELECT event_id, type FROM _rehearsal_before
|
||
EXCEPT SELECT event_id, type FROM export_current WHERE status = 'done')
|
||
UNION ALL
|
||
(SELECT event_id, type FROM export_current WHERE status = 'done'
|
||
EXCEPT SELECT event_id, type FROM _rehearsal_before)
|
||
) d")
|
||
if [[ "$diff" != "0" ]]; then
|
||
echo "✗ FAIL ($label): $diff (event,type) pair(s) changed downloadability across the migration."
|
||
psql -c "
|
||
(SELECT 'LOST — guests can no longer download this' AS problem, event_id, type FROM _rehearsal_before
|
||
EXCEPT ALL SELECT 'LOST — guests can no longer download this', event_id, type FROM export_current WHERE status='done')
|
||
UNION ALL
|
||
(SELECT 'GAINED — a stale keepsake became downloadable', event_id, type FROM export_current WHERE status='done'
|
||
EXCEPT ALL SELECT 'GAINED — a stale keepsake became downloadable', event_id, type FROM _rehearsal_before)"
|
||
exit 1
|
||
fi
|
||
echo "✓ $label: downloadability preserved exactly ($before pair(s))."
|
||
}
|
||
|
||
echo "▸ Applying 014 (up)…"
|
||
psql -q < "$MIG_DIR/014_export_epoch.up.sql"
|
||
assert_up_preserves "up"
|
||
|
||
echo "▸ Applying 014 (down) — the rollback path…"
|
||
psql -q < "$MIG_DIR/014_export_epoch.down.sql"
|
||
|
||
# The down migration is an inverse UP TO DRIFT, not a bit-exact inverse — deliberately.
|
||
#
|
||
# The old model let a ready flag disagree with its job row (flag TRUE over a running/failed job):
|
||
# the flag was a CACHED copy of a derivation, and keeping a cache in agreement with its source
|
||
# across concurrent workers is precisely what it kept failing at. The down migration re-derives the
|
||
# flags from the epoch state, so such a pair comes back FALSE. That HEALS the drift rather than
|
||
# faithfully restoring corruption, and it costs nothing: a flag over a non-done job had no file to
|
||
# serve (file_path is NULL until the worker finishes), so the old code 404'd on it anyway.
|
||
#
|
||
# What a rollback must NOT do is drop a flag for a keepsake that was genuinely downloadable
|
||
# (flag set AND the job actually done). That is the assertion.
|
||
lost=$(psql -tAc "
|
||
SELECT count(*) FROM _rehearsal_before b
|
||
WHERE NOT EXISTS (
|
||
SELECT 1 FROM event e
|
||
WHERE e.id = b.event_id
|
||
AND ((b.type = 'zip' AND e.export_zip_ready) OR (b.type = 'html' AND e.export_html_ready)))")
|
||
if [[ "$lost" != "0" ]]; then
|
||
echo "✗ FAIL (down): $lost downloadable keepsake(s) lost their ready flag — the rollback is LOSSY."
|
||
psql -c "
|
||
SELECT b.event_id, b.type, 'was downloadable, flag not restored' AS problem
|
||
FROM _rehearsal_before b
|
||
WHERE NOT EXISTS (
|
||
SELECT 1 FROM event e
|
||
WHERE e.id = b.event_id
|
||
AND ((b.type = 'zip' AND e.export_zip_ready) OR (b.type = 'html' AND e.export_html_ready)))"
|
||
exit 1
|
||
fi
|
||
echo "✓ down: every downloadable keepsake kept its flag — the old image sees a working state."
|
||
|
||
healed=$(psql -tAc "
|
||
SELECT count(*) FROM event e JOIN _rehearsal_flags f ON f.id = e.id
|
||
WHERE (e.export_zip_ready, e.export_html_ready) IS DISTINCT FROM (f.export_zip_ready, f.export_html_ready)")
|
||
if [[ "$healed" != "0" ]]; then
|
||
echo " ℹ $healed event(s) came back with a CLEARED flag that had drifted from its job row."
|
||
echo " Not a loss — those had no file to serve. The rollback normalises them; the old code"
|
||
echo " will rebuild on the next release. Listing them so the behaviour is not a surprise:"
|
||
psql -c "SELECT e.id, f.export_zip_ready AS was_zip, e.export_zip_ready AS now_zip,
|
||
f.export_html_ready AS was_html, e.export_html_ready AS now_html
|
||
FROM event e JOIN _rehearsal_flags f ON f.id = e.id
|
||
WHERE (e.export_zip_ready, e.export_html_ready)
|
||
IS DISTINCT FROM (f.export_zip_ready, f.export_html_ready)"
|
||
fi
|
||
|
||
echo "▸ Re-applying 014 (up) — roll forward after a rollback…"
|
||
psql -q < "$MIG_DIR/014_export_epoch.up.sql"
|
||
assert_up_preserves "up (after rollback)"
|
||
|
||
echo
|
||
echo "✓ Rehearsal passed. 014 is safe to deploy against this dataset."
|
||
[[ -z "$DUMP" ]] && echo " (synthetic data — re-run with a real pg_dump before you deploy for real)"
|
||
exit 0
|