Merge branch 'fix/postgres-password-deploy-trap'

This commit is contained in:
fabi
2026-07-30 19:28:34 +02:00
4 changed files with 265 additions and 32 deletions

View File

@@ -12,6 +12,14 @@ APP_ENV=production
# ── Database ────────────────────────────────────────────────────────────────── # ── Database ──────────────────────────────────────────────────────────────────
# Set a strong password and keep it in sync between DATABASE_URL and # Set a strong password and keep it in sync between DATABASE_URL and
# POSTGRES_PASSWORD. Generate one with: openssl rand -hex 24 # POSTGRES_PASSWORD. Generate one with: openssl rand -hex 24
#
# SET THIS BEFORE THE FIRST `docker compose up -d`. Postgres reads POSTGRES_PASSWORD
# only when it initialises its data directory, on that very first boot. Change it
# afterwards and the app authenticates with the new password against a volume still
# holding the old one — a permanent restart loop ("password authentication failed").
# The only ways out are restoring the old password or `docker compose down -v`, which
# deletes the database, the media and the exports. In production the app refuses to
# boot while this is still the placeholder below, so it cannot be missed by accident.
DATABASE_URL=postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap DATABASE_URL=postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap
POSTGRES_USER=eventsnap POSTGRES_USER=eventsnap
POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password POSTGRES_PASSWORD=CHANGE_ME_use_a_strong_password
@@ -30,7 +38,9 @@ JWT_SECRET=change_me_to_a_random_64_byte_hex_string
SESSION_EXPIRY_DAYS=30 SESSION_EXPIRY_DAYS=30
# Admin dashboard password (bcrypt hash). # Admin dashboard password (bcrypt hash).
# Generate with: htpasswd -bnBC 12 "" yourpassword | tr -d ':\n' # Generate with an image the stack already pulls (htpasswd needs apache2-utils, which
# a stock VPS does not have):
# docker run --rm caddy:2-alpine caddy hash-password --plaintext 'yourpassword'
# IMPORTANT: keep the SINGLE QUOTES. A bcrypt hash is full of `$` (e.g. $2b$12$…$…), # IMPORTANT: keep the SINGLE QUOTES. A bcrypt hash is full of `$` (e.g. $2b$12$…$…),
# and both Docker Compose's env_file interpolation and dotenvy's variable substitution # and both Docker Compose's env_file interpolation and dotenvy's variable substitution
# would otherwise eat the `$…` segments (reading them as unset vars) and corrupt the # would otherwise eat the `$…` segments (reading them as unset vars) and corrupt the

View File

@@ -97,17 +97,51 @@ eventsnap/
git clone https://git.mc02.dev/fabi/EventSnap.git eventsnap git clone https://git.mc02.dev/fabi/EventSnap.git eventsnap
cd eventsnap cd eventsnap
# 2. Configure environment # 2. Configure environment — set EVERY secret NOW, before step 3.
cp .env.example .env cp .env.example .env
nano .env # set DOMAIN, JWT_SECRET, ADMIN_PASSWORD_HASH, EVENT_NAME, etc. nano .env # DOMAIN, EVENT_NAME, EVENT_SLUG,
# JWT_SECRET, ADMIN_PASSWORD_HASH,
# POSTGRES_PASSWORD *and* the same password inside DATABASE_URL
# (see "Generate required secrets" below)
# 3. Start the stack # 3. Start the stack
docker compose up -d docker compose up -d
``` ```
> **Set every secret before step 3 — `POSTGRES_PASSWORD` especially.** Postgres reads it **only
> when it initialises its data directory**, which happens on the very first `docker compose up -d`.
> Changing it in `.env` afterwards does not change the stored password: the app then authenticates
> with the new one against a volume holding the old one, and you get a permanent restart loop with
> `password authentication failed for user "eventsnap"`. The only fixes are restoring the old
> password or `docker compose down -v`, which **deletes the database, the media and the exports**.
> Getting it right once, up front, costs nothing; getting it wrong costs the volume.
Caddy automatically obtains a Let's Encrypt certificate on first start. The app is live at `https://DOMAIN` within ~30 seconds. Caddy automatically obtains a Let's Encrypt certificate on first start. The app is live at `https://DOMAIN` within ~30 seconds.
> **If the site never comes up:** with `APP_ENV=production` the backend **refuses to boot** while `JWT_SECRET`/`ADMIN_PASSWORD_HASH` still hold the `.env.example` placeholders (this is deliberate — a publicly-known signing key is worse than downtime). Caddy then waits on the unhealthy `app` container and never serves. Check `docker compose logs app` — a "Refusing to start … placeholder …" line means you skipped step 2. Rotate the secrets (see below) and restart. > **If the site never comes up:** with `APP_ENV=production` the backend **refuses to boot** while
> `JWT_SECRET`, `ADMIN_PASSWORD_HASH` or the password inside `DATABASE_URL` still hold the
> `.env.example` placeholders (this is deliberate — a publicly-known signing key or database
> password is worse than downtime). Caddy then waits on the unhealthy `app` container and never
> serves. Check `docker compose logs app` — a "Refusing to start … placeholder …" line lists
> **every** unset secret at once, so one edit fixes them all.
>
> **If it comes up but keeps restarting with `password authentication failed for user
> "eventsnap"`:** `POSTGRES_PASSWORD` was changed after the database volume was created. Postgres
> applies that variable only at initialisation, so `.env` and the stored password have drifted
> apart permanently. `docker compose logs app` spells this out. Before the event, with nothing
> worth keeping:
>
> ```bash
> docker compose down -v && docker compose up -d # -v DELETES db + media + exports. No undo.
> ```
>
> **Once the event has real data, never do that.** Put the original password back into
> `DATABASE_URL`, or change the stored one instead:
>
> ```bash
> docker compose exec db psql -U "$POSTGRES_USER" -c \
> "ALTER ROLE eventsnap WITH PASSWORD 'the-password-now-in-your-.env';"
> ```
> **Production note:** `docker compose up -d` does **not** expose the database — Postgres is reachable only on the internal Docker network. For local development where you need host access to Postgres, opt into the dev overlay explicitly: > **Production note:** `docker compose up -d` does **not** expose the database — Postgres is reachable only on the internal Docker network. For local development where you need host access to Postgres, opt into the dev overlay explicitly:
> ```bash > ```bash
@@ -179,10 +213,19 @@ TLS certificate and all data volumes survive.
# JWT secret (64 random bytes) # JWT secret (64 random bytes)
openssl rand -hex 64 openssl rand -hex 64
# Admin password hash (bcrypt, cost 12) # Database password (goes in BOTH DATABASE_URL and POSTGRES_PASSWORD)
htpasswd -bnBC 12 "" yourpassword | tr -d ':\n' openssl rand -hex 24
# Admin password hash (bcrypt). Uses an image the stack already pulls, so it needs
# nothing installed on the host — `htpasswd` lives in apache2-utils, which a stock
# VPS does not have. Emits cost 14 rather than 12; that is fine (admin login is
# rate-limited and hashed off the async runtime), and any $2a/$2b/$2y hash verifies.
docker run --rm caddy:2-alpine caddy hash-password --plaintext 'yourpassword'
``` ```
Wrap the resulting hash in **single quotes** in `.env` — see the note there; a bcrypt
hash is full of `$`, and both Compose and dotenvy would otherwise eat those segments.
### Environment Variables ### Environment Variables
See [.env.example](.env.example) for the full list with descriptions and defaults. Key variables: See [.env.example](.env.example) for the full list with descriptions and defaults. Key variables:

View File

@@ -20,21 +20,53 @@ fn looks_placeholder(s: &str) -> bool {
/// Enforce secret hygiene. In production every guard is hard-fail: a booting app /// Enforce secret hygiene. In production every guard is hard-fail: a booting app
/// with a publicly-known signing key is worse than one that refuses to start. /// with a publicly-known signing key is worse than one that refuses to start.
/// Outside production the dev sentinel is tolerated (warned) so local dev is frictionless. /// Outside production the dev sentinel is tolerated (warned) so local dev is frictionless.
fn validate_secrets(is_prod: bool, jwt_secret: &str, admin_password_hash: &str) -> Result<()> { ///
/// EVERY failure is collected and reported together. Returning on the first one made fixing two
/// secrets cost two boot cycles — the operator rotates JWT_SECRET, restarts, and only then learns
/// about ADMIN_PASSWORD_HASH. Restarting this stack is not free (Caddy waits on the unhealthy app),
/// and each avoidable cycle is another chance to reach for `down -v`.
fn validate_secrets(
is_prod: bool,
jwt_secret: &str,
admin_password_hash: &str,
database_url: &str,
) -> Result<()> {
if is_prod { if is_prod {
let mut problems: Vec<&str> = Vec::new();
if looks_placeholder(jwt_secret) { if looks_placeholder(jwt_secret) {
return Err(anyhow!( problems.push(
"Refusing to start in production with a placeholder JWT_SECRET — \ "JWT_SECRET is still the .env.example placeholder — rotate it \
rotate it (openssl rand -hex 64)." (openssl rand -hex 64).",
)); );
} } else if jwt_secret.len() < 32 {
if jwt_secret.len() < 32 { problems.push("JWT_SECRET must be at least 32 characters.");
return Err(anyhow!("JWT_SECRET must be at least 32 characters."));
} }
if admin_password_hash.is_empty() || looks_placeholder(admin_password_hash) { if admin_password_hash.is_empty() || looks_placeholder(admin_password_hash) {
problems.push(
"ADMIN_PASSWORD_HASH is unset or still the .env.example placeholder — generate one \
(docker run --rm caddy:2-alpine caddy hash-password --plaintext '<password>').",
);
}
// The DATABASE_URL carries the Postgres password, so a placeholder here means the stack is
// running on `CHANGE_ME_use_a_strong_password` — a credential published in the repo. The
// app used to boot green on it, because this guard only ever covered the two secrets it
// was written for and nothing else looked at POSTGRES_PASSWORD at all.
//
// Read POSTGRES_PASSWORD's docs before changing this: it is applied ONLY at initdb, so the
// remedy is not "edit .env and restart" — see the 28P01 diagnostic in db.rs.
if looks_placeholder(database_url) {
problems.push(
"DATABASE_URL still carries the .env.example placeholder password — set a strong \
one (openssl rand -hex 24) in BOTH DATABASE_URL and POSTGRES_PASSWORD.",
);
}
if !problems.is_empty() {
return Err(anyhow!( return Err(anyhow!(
"Refusing to start in production without a real ADMIN_PASSWORD_HASH — \ "Refusing to start in production — {} secret(s) still unset or placeholder:\n - {}\n\
generate one (htpasswd -bnBC 12 '' <password> | tr -d ':\\n')." ALL secrets must be set BEFORE the first `docker compose up -d`: Postgres bakes \
POSTGRES_PASSWORD into its data directory on first boot and ignores later changes.",
problems.len(),
problems.join("\n - ")
)); ));
} }
} else if jwt_secret == DEV_JWT_SECRET_SENTINEL { } else if jwt_secret == DEV_JWT_SECRET_SENTINEL {
@@ -89,11 +121,12 @@ impl AppConfig {
let jwt_secret = std::env::var("JWT_SECRET").context("JWT_SECRET must be set")?; let jwt_secret = std::env::var("JWT_SECRET").context("JWT_SECRET must be set")?;
let admin_password_hash = std::env::var("ADMIN_PASSWORD_HASH").unwrap_or_default(); let admin_password_hash = std::env::var("ADMIN_PASSWORD_HASH").unwrap_or_default();
let database_url = std::env::var("DATABASE_URL").context("DATABASE_URL must be set")?;
validate_secrets(is_prod, &jwt_secret, &admin_password_hash)?; validate_secrets(is_prod, &jwt_secret, &admin_password_hash, &database_url)?;
Ok(Self { Ok(Self {
database_url: std::env::var("DATABASE_URL").context("DATABASE_URL must be set")?, database_url,
jwt_secret, jwt_secret,
session_expiry_days: std::env::var("SESSION_EXPIRY_DAYS") session_expiry_days: std::env::var("SESSION_EXPIRY_DAYS")
.unwrap_or_else(|_| "30".to_string()) .unwrap_or_else(|_| "30".to_string())
@@ -141,12 +174,18 @@ mod tests {
const REAL_SECRET: &str = "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2"; const REAL_SECRET: &str = "a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2";
const REAL_HASH: &str = "$2y$12$abcdefghijklmnopqrstuv.wxyzABCDEFGHIJKLMNOPQRSTUVWXYZ012"; const REAL_HASH: &str = "$2y$12$abcdefghijklmnopqrstuv.wxyzABCDEFGHIJKLMNOPQRSTUVWXYZ012";
const REAL_DB_URL: &str = "postgres://eventsnap:7f3a9c1e5b2d8a4f@db:5432/eventsnap";
#[test] #[test]
fn prod_rejects_shipped_placeholder_secret() { fn prod_rejects_shipped_placeholder_secret() {
// The exact string shipped in `.env` — >32 chars, so it must be caught by // The exact string shipped in `.env` — >32 chars, so it must be caught by
// the substring guard, not the length check. // the substring guard, not the length check.
let err = validate_secrets(true, "change_me_to_a_random_64_byte_hex_string", REAL_HASH); let err = validate_secrets(
true,
"change_me_to_a_random_64_byte_hex_string",
REAL_HASH,
REAL_DB_URL,
);
assert!( assert!(
err.is_err(), err.is_err(),
"placeholder JWT_SECRET must be rejected in prod" "placeholder JWT_SECRET must be rejected in prod"
@@ -155,29 +194,108 @@ mod tests {
#[test] #[test]
fn prod_rejects_dev_sentinel_and_short_secret() { fn prod_rejects_dev_sentinel_and_short_secret() {
assert!(validate_secrets(true, DEV_JWT_SECRET_SENTINEL, REAL_HASH).is_err()); assert!(validate_secrets(true, DEV_JWT_SECRET_SENTINEL, REAL_HASH, REAL_DB_URL).is_err());
assert!(validate_secrets(true, "tooshort", REAL_HASH).is_err()); assert!(validate_secrets(true, "tooshort", REAL_HASH, REAL_DB_URL).is_err());
} }
#[test] #[test]
fn prod_rejects_missing_or_placeholder_admin_hash() { fn prod_rejects_missing_or_placeholder_admin_hash() {
assert!(validate_secrets(true, REAL_SECRET, "").is_err()); assert!(validate_secrets(true, REAL_SECRET, "", REAL_DB_URL).is_err());
assert!(validate_secrets(true, REAL_SECRET, "$2y$12$placeholder_replace_me").is_err()); assert!(
validate_secrets(
true,
REAL_SECRET,
"$2y$12$placeholder_replace_me",
REAL_DB_URL
)
.is_err()
);
}
/// The stack used to come up GREEN on the database password published in the repo: this guard
/// covered the two secrets it was written for, and nothing anywhere looked at the Postgres
/// credential. README step 2 doesn't name POSTGRES_PASSWORD either, so following the
/// documented procedure verbatim shipped it.
#[test]
fn prod_rejects_the_shipped_placeholder_database_password() {
let shipped = "postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap";
let err = validate_secrets(true, REAL_SECRET, REAL_HASH, shipped).unwrap_err();
assert!(
err.to_string().contains("DATABASE_URL"),
"the refusal must name DATABASE_URL, not just fail: {err}"
);
// And it must point at the initdb trap, or the operator edits .env, restarts, and lands
// in a permanent auth-failure loop instead.
assert!(
err.to_string().contains("POSTGRES_PASSWORD"),
"the refusal must name POSTGRES_PASSWORD as the other half: {err}"
);
}
/// Every problem in ONE message. Reporting them one per boot made fixing two secrets cost two
/// restart cycles, on a stack where Caddy waits on the unhealthy app the whole time.
#[test]
fn prod_reports_every_placeholder_at_once() {
let err = validate_secrets(
true,
"change_me_to_a_random_64_byte_hex_string",
"$2y$12$placeholder_replace_me",
"postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap",
)
.unwrap_err()
.to_string();
for expected in ["JWT_SECRET", "ADMIN_PASSWORD_HASH", "DATABASE_URL"] {
assert!(err.contains(expected), "{expected} missing from: {err}");
}
assert!(
err.contains("3 secret(s)"),
"the count must match what is listed: {err}"
);
} }
#[test] #[test]
fn prod_accepts_real_secrets() { fn prod_accepts_real_secrets() {
assert!(validate_secrets(true, REAL_SECRET, REAL_HASH).is_ok()); assert!(validate_secrets(true, REAL_SECRET, REAL_HASH, REAL_DB_URL).is_ok());
}
/// A real password that happens to contain no placeholder substring must pass — including one
/// with URL-ish punctuation, so the guard can't be mistaken for a URL validator.
#[test]
fn prod_accepts_a_real_database_url_with_awkward_punctuation() {
assert!(
validate_secrets(
true,
REAL_SECRET,
REAL_HASH,
"postgres://eventsnap:aB3%24xY9-_.qW@db:5432/eventsnap"
)
.is_ok()
);
}
/// The e2e stack runs without APP_ENV=production, so none of this applies there — but assert
/// it, because a guard that tripped in e2e would be found the hard way.
#[test]
fn non_prod_ignores_a_placeholder_database_url() {
assert!(
validate_secrets(
false,
REAL_SECRET,
"",
"postgres://eventsnap:CHANGE_ME_use_a_strong_password@db:5432/eventsnap"
)
.is_ok()
);
} }
#[test] #[test]
fn non_prod_tolerates_dev_sentinel() { fn non_prod_tolerates_dev_sentinel() {
assert!(validate_secrets(false, DEV_JWT_SECRET_SENTINEL, "").is_ok()); assert!(validate_secrets(false, DEV_JWT_SECRET_SENTINEL, "", REAL_DB_URL).is_ok());
} }
#[test] #[test]
fn non_prod_still_rejects_short_non_sentinel_secret() { fn non_prod_still_rejects_short_non_sentinel_secret() {
assert!(validate_secrets(false, "tooshort", "").is_err()); assert!(validate_secrets(false, "tooshort", "", REAL_DB_URL).is_err());
} }
#[test] #[test]
@@ -185,9 +303,23 @@ mod tests {
// looks_placeholder lowercases before matching — an upper/mixed-case // looks_placeholder lowercases before matching — an upper/mixed-case
// placeholder must still be rejected in prod. // placeholder must still be rejected in prod.
assert!( assert!(
validate_secrets(true, "CHANGE_ME_TO_A_RANDOM_64_BYTE_HEX_STRING", REAL_HASH).is_err() validate_secrets(
true,
"CHANGE_ME_TO_A_RANDOM_64_BYTE_HEX_STRING",
REAL_HASH,
REAL_DB_URL
)
.is_err()
);
assert!(
validate_secrets(
true,
REAL_SECRET,
"$2Y$12$PLACEHOLDER_replace_me",
REAL_DB_URL
)
.is_err()
); );
assert!(validate_secrets(true, REAL_SECRET, "$2Y$12$PLACEHOLDER_replace_me").is_err());
} }
#[test] #[test]
@@ -197,7 +329,7 @@ mod tests {
const LEN_31: &str = "abcdefghijklmnopqrstuvwxyz01234"; const LEN_31: &str = "abcdefghijklmnopqrstuvwxyz01234";
assert_eq!(LEN_32.len(), 32); assert_eq!(LEN_32.len(), 32);
assert_eq!(LEN_31.len(), 31); assert_eq!(LEN_31.len(), 31);
assert!(validate_secrets(true, LEN_32, REAL_HASH).is_ok()); assert!(validate_secrets(true, LEN_32, REAL_HASH, REAL_DB_URL).is_ok());
assert!(validate_secrets(true, LEN_31, REAL_HASH).is_err()); assert!(validate_secrets(true, LEN_31, REAL_HASH, REAL_DB_URL).is_err());
} }
} }

View File

@@ -4,17 +4,65 @@ use sqlx::postgres::PgPoolOptions;
const DEFAULT_MAX_CONNECTIONS: u32 = 10; const DEFAULT_MAX_CONNECTIONS: u32 = 10;
/// SQLSTATE for `invalid_password`.
const PG_INVALID_PASSWORD: &str = "28P01";
/// Turn the one connect failure with an unguessable cause into a self-explaining one.
///
/// `POSTGRES_PASSWORD` is honoured ONLY when Postgres initialises its data directory. Change it in
/// `.env` afterwards and the app authenticates with the new password against a volume that still
/// holds the old one — a permanent restart loop whose only symptom is
/// `password authentication failed`.
///
/// The production secret guard makes that sequence NEARLY CERTAIN rather than rare: it stops the
/// app on the first `docker compose up -d`, but not the `db` service in that same command, which
/// initialises and bakes in whatever password was in `.env` at that moment. So the intended
/// recovery — see the refusal, fix your secrets, boot again — is exactly the sequence that breaks
/// it. Nothing in the error names the cause, and the remedy destroys data, so it is the last thing
/// an operator should guess at.
fn explain_auth_failure(err: &sqlx::Error) {
let is_auth_failure = match err {
sqlx::Error::Database(db) => db.code().as_deref() == Some(PG_INVALID_PASSWORD),
_ => false,
};
if !is_auth_failure {
return;
}
tracing::error!(
"Postgres rejected the credentials in DATABASE_URL (SQLSTATE {PG_INVALID_PASSWORD}).\n\
\n\
This almost always means POSTGRES_PASSWORD was changed AFTER the database volume was \
first created. Postgres applies that variable only when it initialises its data \
directory; editing .env and restarting does not change the stored password, so the two \
drift apart permanently.\n\
\n\
If the event has NOT started and you have no data worth keeping:\n\n \
docker compose down -v && docker compose up -d\n\n\
(-v DELETES the database, the uploaded media and the exports. There is no undo.)\n\
\n\
If you DO have data: restore the old password into DATABASE_URL instead, or change the \
stored one with ALTER ROLE inside the running db container. Never reach for -v to fix a \
login problem on a live event."
);
}
pub async fn create_pool(database_url: &str) -> Result<PgPool> { pub async fn create_pool(database_url: &str) -> Result<PgPool> {
let max_connections = std::env::var("DATABASE_MAX_CONNECTIONS") let max_connections = std::env::var("DATABASE_MAX_CONNECTIONS")
.ok() .ok()
.and_then(|s| s.parse::<u32>().ok()) .and_then(|s| s.parse::<u32>().ok())
.unwrap_or(DEFAULT_MAX_CONNECTIONS); .unwrap_or(DEFAULT_MAX_CONNECTIONS);
let pool = PgPoolOptions::new() let pool = match PgPoolOptions::new()
.max_connections(max_connections) .max_connections(max_connections)
.connect(database_url) .connect(database_url)
.await .await
.context("failed to connect to database")?; {
Ok(pool) => pool,
Err(e) => {
explain_auth_failure(&e);
return Err(e).context("failed to connect to database");
}
};
sqlx::migrate!() sqlx::migrate!()
.run(&pool) .run(&pool)