Files
EventSnap/backend/src/auth/handlers.rs
MechaCat02 05063694d2 fix: close eight regressions the audit pass found, five of them mine
Two adversarial reviews over 61119be, 1d9fb11 and eb0e405. The merge itself came
back clean — client_upload_id end to end, TempFileGuard's arm/retarget/disarm, the
supervised sweep wiring and v_feed's column parity were all verified sound. What
follows is what my own three commits broke.

BLOCKER — a post-release rebuild was permanently impossible, and it 404'd the keepsake

1d9fb11 deferred prune_superseded_archives to run only on success, so a failed
rebuild could no longer destroy the last good archive. It did not follow that
through: ensure_export_space runs BEFORE the prune, so at rebuild time the previous
generation is still on disk and counted against free. That halves the gallery a
rebuild can survive (~4.6 GB) relative to what the upload gate accepts (~7.8 GB) —
and it self-locks, because invalidate_and_arm bumps the epoch on COMMIT, which 404s
both download routes immediately, while the only code that could free the space now
runs only after a success that can never happen. A guest deleting their own photo is
enough to trigger it. Recovery needed `docker exec rm`.

Now two-phase: try to build while preserving the old generation; if that genuinely
does not fit, reclaim it and try once more. Strictly better than both the original
ordering and my change — the old archive is sacrificed only when it is the only way
to get a new one.

BLOCKER — the deferred prune could delete the last archive when a worker LOST the race

run_*_export_inner returned Ok(()) on the superseded/discard path, so `res.is_ok()`
fired the prune with the worker's own RETIRED epoch as keep_seq. At that moment the
winning generation is still `pending` with no file, so protected_files is empty and
the last good archive was deleted with no replacement. Exactly the invariant
deferring the prune was meant to establish. Returns Err(Superseded) now, which
abandon_if_superseded already swallows for the caller.

BLOCKER — the low-disk banner could never fire before the wall

eb0e405's gate refuses at `free < keepsake + DISK_RESERVE`, while disk_is_low warned
at `free < keepsake`. The two differ by the whole reserve, so the wall always came
first: every guest blocked from uploading while the host dashboard showed ~27 GB free
and no banner, with nobody on site. disk_is_low now shares the gate's expression plus
a 25% margin, and a test asserts the banner fires at the gate threshold across the
whole gallery-size range.

BLOCKER — I raised the unauthenticated bcrypt ceiling 24x on a 2 vCPU box

1d9fb11 moved admin_login's tight bucket after verify_password (correct — that is what
stops a guest locking the operator out) but replaced the incidental 5/min bound on
bcrypt with 120/min and nothing global. bcrypt is on spawn_blocking, but tokio's
blocking pool is 512 threads, so "off the runtime" is not "bounded": enough concurrent
verifies preempt both async workers and uploads, feed and SSE stall. Three
unauthenticated endpoints reach bcrypt and every guest shares one NAT IP, so per-IP
limits bound nothing globally. Adds a process-wide semaphore of `cores - 1` around both
verify and hash, and drops the ceiling to 30.

Also correcting my own claim: "a correct password is never throttled" was wrong. The
failure bucket cannot block it, but the CPU ceiling still can. The code comment said so;
the commit message did not.

BLOCKER — migration 025 could crash-loop the app on boot

Its UPDATE derives `Name (8hex)` with no guard against idx_user_event_name_ci. A guest
who had already joined as exactly that string makes the migration fail, which
propagates out of create_pool, exits main, and `restart: unless-stopped` turns it into
a permanent loop — a worse version of the lockout the migration exists to clean up.
Now skips colliding rows (create_admin_user already falls back to Admin-<8hex>, so the
cleanup is convenience, not load-bearing). Also `role = 'guest'` rather than
`<> 'admin'`, which was renaming legitimately promoted hosts named "Host".

DEGRADATION — the watchdog's suspension credit was unbounded

Background tabs are throttled to ~1 tick/min WITHOUT the network stack pausing, and the
tick gap cannot tell that from a freeze. Crediting every late tick grew the observed
silence by only one interval per real minute, so a dead socket took ~18 minutes to
detect while holding the queue's processing latch. Credit is now capped at one stall
window and REFILLS on real progress: an upload that is moving survives any number of
screen locks, while one that is silent and suspended is detected within ~3 minutes.

DEGRADATION — the 4xx log line was an unauthenticated log-injection vector

validate_display_name allowed newlines, several 4xx messages interpolate the name, and
%message wrote it unescaped. Two unauthenticated /join requests could forge arbitrary
lines in the only forensic record an unattended event has. Fixed at both ends: control
characters rejected at the door, and `detail = ?message` escapes on the way out (which
also stops colliding with tracing's reserved `message` field). 401/404 drop to DEBUG —
they carry no operator signal and were the cheapest lines for a scanner to use to roll
the 30 MB log window in minutes.

DEGRADATION — the quota floor was inverted exactly where it mattered

`computed.max(MIN.min(budget))`: `budget` is the whole disk's share, so below 500 MiB
the "floor" became the entire remaining budget and EVERY uploader was authorised all of
it — 400 MB free, 3 uploaders, 300 MB each. A test pinned that as correct under the name
`the_floor_never_exceeds_what_the_disk_can_back`. Both fixed.

Also replaces the headline gate test, which asserted its own precondition inside an `if`
on that precondition and could not fail. It now pins what actually binds the gate to the
preflight — that required_free_bytes charges for both halves — plus the ceiling band.

Verified: 151/151 backend tests against a live Postgres, clippy clean, 58/58 vitest,
svelte-check 0 errors, eslint clean, both builds, caddy validate, and the migration
collision reproduced against Postgres 16 before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 22:33:43 +02:00

804 lines
36 KiB
Rust

use std::time::Duration;
use axum::Json;
use axum::extract::{ConnectInfo, State};
use axum::http::{HeaderMap, StatusCode};
use chrono::Utc;
use rand::Rng;
use serde::{Deserialize, Serialize};
use std::net::SocketAddr;
use uuid::Uuid;
use crate::auth::jwt;
use crate::auth::middleware::AuthUser;
use crate::error::AppError;
use crate::models::event::Event;
use crate::models::session::Session;
use crate::models::user::{User, UserRole};
use crate::services::config;
use crate::services::rate_limiter::client_ip;
use crate::state::AppState;
/// Names a guest may not take.
///
/// Defence in depth only. The real fix for the admin-lockout defect is that `admin_login` now
/// resolves its user by ROLE rather than by name (see `User::find_admin_for_event`), which is
/// why homoglyph and zero-width bypasses of this list are not a concern: the name is no longer
/// load-bearing for anything. What this buys is that a guest cannot impersonate the host in the
/// feed's byline, and that "Admin" stays available for the admin row.
const RESERVED_DISPLAY_NAMES: &[&str] = &["admin", "administrator", "host", "eventsnap"];
fn is_reserved_display_name(name: &str) -> bool {
let name = name.trim().to_lowercase();
RESERVED_DISPLAY_NAMES.contains(&name.as_str())
}
/// Trim and bounds-check a display name.
///
/// Shared by `join`, `recover` and `request_pin_reset` so the length check happens BEFORE the
/// name is used to build a rate-limiter key. It was inline in `join` only, so on the other two
/// endpoints `format!("...:{ip}:{name_key}")` allocated from an unbounded, attacker-chosen
/// string and stored it in a HashMap pruned once an hour with a 24 h ceiling — turning the
/// limiter itself into the memory-exhaustion primitive it exists to prevent.
fn validate_display_name(raw: &str) -> Result<&str, AppError> {
let name = raw.trim();
let chars = name.chars().count();
if chars == 0 || chars > 50 {
return Err(AppError::BadRequest(
"Name muss zwischen 1 und 50 Zeichen lang sein.".into(),
));
}
// No control characters. NUL is the hard requirement — Postgres rejects 0x00 in TEXT with a
// 500, so catching it here turns an internal error into a clean 400 — but the rest matter
// too, and for reasons beyond tidiness:
//
// * Newlines make the name a LOG INJECTION vector. Several 4xx messages interpolate it
// ("Der Name \"X\" ist bereits vergeben.") and those are logged; a name carrying a
// newline plus a plausible timestamp prefix lets two unauthenticated requests forge
// entries in the only forensic record an unattended event has. `error.rs` escapes on the
// way out as well — this is the other half, and the half that keeps the forged text out
// of the database and out of the feed byline in the first place.
// * A bare CR or a bidi override renders as a name that is not what was typed, in the feed,
// the host dashboard's moderation list and the keepsake.
//
// Deliberately NOT a whitelist: guests have accents, emoji and non-Latin scripts in their
// names, and rejecting those would be worse than the problem.
if name.chars().any(|c| c.is_control()) {
return Err(AppError::BadRequest(
"Name enthält ungültige Zeichen.".into(),
));
}
Ok(name)
}
#[derive(Deserialize)]
pub struct JoinRequest {
pub display_name: String,
}
#[derive(Serialize)]
pub struct JoinResponse {
pub jwt: String,
pub pin: String,
pub user_id: Uuid,
pub is_new: bool,
}
pub async fn join(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<JoinRequest>,
) -> Result<(StatusCode, Json<JoinResponse>), AppError> {
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let join_rate_on = config::get_bool(&state.config_cache, "join_rate_enabled", true).await;
// Coarse per-IP flood ceiling. `/join` is pre-auth so there is no user to key on, and
// at a venue EVERY guest arrives from one public IP — a tight per-IP bucket meant the
// 6th person through the door was turned away by the 5 ahead of them. So the per-IP
// limit here only bounds raw volume; the real anti-spam bucket is per-name below.
// Cheap enough to run before validation, which keeps a flood of malformed bodies from
// being free.
if rate_limits_on && join_rate_on {
let ip_ceiling = config::get_usize(&state.config_cache, "join_ip_rate_per_min", 60).await;
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("join_ip:{ip}"),
ip_ceiling,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
let display_name = validate_display_name(&body.display_name)?;
if is_reserved_display_name(display_name) {
// 409, matching the name-taken response below, so the frontend's existing handling
// works unchanged. See RESERVED_DISPLAY_NAMES for why this exists.
return Err(AppError::Conflict(format!(
"Der Name \"{display_name}\" ist reserviert. Bitte wähle einen anderen."
)));
}
// Per-guest bucket, keyed like the `recover:{ip}:{name}` limiter below. This carries
// the original 5/60s anti-spam intent, but one guest retrying can no longer consume
// the allowance of everyone else sharing the venue's NAT.
if rate_limits_on && join_rate_on {
let name_key = display_name.to_lowercase();
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("join:{ip}:{name_key}"),
5,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
let event = Event::find_or_create(
&state.pool,
&state.config.event_slug,
&state.config.event_name,
)
.await?;
// Reject if a user with this name (case-insensitive) already exists
if User::name_taken(&state.pool, event.id, display_name).await? {
return Err(AppError::Conflict(format!(
"Der Name \"{}\" ist bereits vergeben.",
display_name
)));
}
// Generate a 4-digit PIN
let pin: String = format!("{:04}", rand::rng().random_range(0..10000u32));
let pin_hash = hash_password(pin.clone(), 12).await?;
// The pre-check above is racy: two simultaneous joins with the same name can both
// pass it, and the DB's unique index then rejects the loser. Map that unique
// violation to the same clean 409 the pre-check returns, not a generic 500.
let user = match User::create(&state.pool, event.id, display_name, &pin_hash).await {
Ok(u) => u,
Err(sqlx::Error::Database(db)) if db.is_unique_violation() => {
return Err(AppError::Conflict(format!(
"Der Name \"{}\" ist bereits vergeben.",
display_name
)));
}
Err(e) => return Err(e.into()),
};
let token = jwt::create_token(
user.id,
event.id,
user.role.clone(),
&state.config.jwt_secret,
state.config.session_expiry_days,
)
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?;
let token_hash = jwt::hash_token(&token);
let expires_at = Utc::now() + chrono::Duration::days(state.config.session_expiry_days);
Session::create(&state.pool, user.id, &token_hash, expires_at).await?;
Ok((
StatusCode::CREATED,
Json(JoinResponse {
jwt: token,
pin,
user_id: user.id,
is_new: true,
}),
))
}
/// Default for `recover_name_rate_per_15min` — wrong PINs allowed per (IP, name) per 15 min.
/// Mirrors migration 023; kept here so the invariant below can be asserted in a test.
const RECOVER_NAME_CEILING_DEFAULT: usize = 4;
/// Wrong PINs, from ALL sources, before the ACCOUNT itself is locked for 15 minutes.
///
/// This was 3, which sat BELOW the per-(IP, name) ceiling of 5 — and that ordering, not the
/// number, was the defect. Display names are public on the feed, so three requests from a single
/// IP locked any guest out of their own account, repeatable every 15 minutes, indefinitely. The
/// tier meant to protect a guest was the easiest way to attack them.
///
/// Raised deliberately far above the per-IP tier so the two do different jobs. The per-(IP, name)
/// bucket is what stops a guesser, and it costs the ATTACKER. This tier is the last line against
/// a DISTRIBUTED guesser, and it is the only one an attacker can turn on a victim — so reaching
/// it must require at least three distinct sources inside the decay window.
///
/// Brute-force cost is unchanged: 12 attempts per 15 minutes is 48/hour against one account, so
/// 10 000 four-digit PINs still take ~208 hours no matter how many IPs are used. An honest guest
/// fat-fingering a 4-digit PIN never comes close, and `increment_failed_pin` now decays the
/// streak after 15 minutes so yesterday's typos don't count toward today's.
const PIN_LOCK_THRESHOLD: i16 = 12;
#[derive(Deserialize)]
pub struct RecoverRequest {
pub display_name: String,
pub pin: String,
}
#[derive(Serialize)]
pub struct RecoverResponse {
pub jwt: String,
pub user_id: Uuid,
}
/// A real cost-12 bcrypt hash of a fixed dummy value, computed once and cached. Used
/// to run a constant-time-ish verify on the "unknown display name" branch of
/// [`recover`], so that branch costs the same as a real (user-exists) verify and can't
/// be told apart by timing.
fn dummy_pin_hash() -> &'static str {
static HASH: std::sync::OnceLock<String> = std::sync::OnceLock::new();
HASH.get_or_init(|| {
bcrypt::hash("eventsnap-dummy-not-a-real-pin", 12)
.expect("bcrypt hashing of a static dummy value cannot fail")
})
}
/// Run a bcrypt verify on the blocking pool.
///
/// bcrypt at cost 12 is ~200ms of deliberate CPU. Called inline on an async task it pins a
/// tokio WORKER thread for that whole time, and the runtime only has one per core — so a
/// flood of `/recover` or `/admin/login` attempts stalls every other request on the box,
/// including the feed. Offloading moves that cost to the blocking pool, which is sized for
/// exactly this and whose saturation degrades logins rather than the whole app.
/// Process-wide ceiling on CONCURRENT bcrypt work.
///
/// bcrypt is deliberately expensive — ~250 ms of a core at cost 12, and this deployment's own
/// runbook generates the admin hash at a higher cost than that. Every call is correctly on
/// `spawn_blocking`, but tokio's blocking pool defaults to 512 threads, so "off the async
/// runtime" is not the same as "bounded": enough concurrent hashes will preempt both async
/// worker threads a 2 vCPU box gets, and uploads, feed and SSE stall behind them.
///
/// Three unauthenticated endpoints reach bcrypt — `/join` (hash), `/recover` (verify, including
/// a deliberate throwaway verify for unknown names) and `/admin/login` (verify) — each with only
/// a per-IP bucket in front, and at a venue every guest shares one public IP. A per-IP limit
/// therefore bounds nothing globally.
///
/// `cores - 1` leaves a core for actually serving requests. Excess callers WAIT on the permit
/// rather than burning CPU, so a flood degrades to latency instead of an outage.
static BCRYPT_PERMITS: std::sync::LazyLock<tokio::sync::Semaphore> =
std::sync::LazyLock::new(|| {
let cores = std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(2);
tokio::sync::Semaphore::new(cores.saturating_sub(1).max(1))
});
async fn verify_password(candidate: String, hash: String) -> bool {
// `acquire()` only fails if the semaphore is closed, which never happens here.
let _permit = BCRYPT_PERMITS.acquire().await;
tokio::task::spawn_blocking(move || bcrypt::verify(&candidate, &hash).unwrap_or(false))
.await
.unwrap_or(false)
}
/// Hash a secret on the blocking pool. Same reasoning as [`verify_password`] — and this one
/// runs on the busiest auth path there is, since every guest who joins gets a PIN hashed.
pub async fn hash_password(secret: String, cost: u32) -> Result<String, AppError> {
// Same global ceiling as `verify_password` — `/join` hashes a PIN for every guest, and 100
// guests scanning the QR at once is the arrival burst this box has to survive.
let _permit = BCRYPT_PERMITS.acquire().await;
tokio::task::spawn_blocking(move || bcrypt::hash(&secret, cost))
.await
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))
}
pub async fn recover(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<RecoverRequest>,
) -> Result<Json<RecoverResponse>, AppError> {
// Validated BEFORE it is used as a rate-limiter key — see `validate_display_name`. The
// per-IP ceiling below is keyed only on the IP, so it is safe to run either side of this;
// the per-NAME bucket is not.
let display_name = validate_display_name(&body.display_name)?;
// Per-IP+name throttle BEFORE the per-user lockout counter. Without this an attacker who
// knows a display name (they're visible on the feed) can burn through the victim's wrong-PIN
// budget and lock them out, repeatedly. The ceiling here MUST stay below
// PIN_LOCK_THRESHOLD — see the constant for why that ordering is the whole control.
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let recover_rate_on = config::get_bool(&state.config_cache, "recover_rate_enabled", true).await;
if rate_limits_on && recover_rate_on {
// Coarse per-IP ceiling FIRST. The per-(ip, name) bucket below is the anti-guessing
// control, but the name is attacker-chosen, so cycling names mints a fresh bucket
// every time and leaves the per-IP cost unbounded. That matters more here than
// anywhere else: every call runs a cost-12 bcrypt verify, including an
// unconditional throwaway one for names that don't exist (see below), so an unknown
// name is the CHEAPEST way to make the server do ~200ms of hashing. Checked before
// the per-name bucket so a name generator can't walk past it.
let ip_ceiling =
config::get_usize(&state.config_cache, "recover_ip_rate_per_min", 30).await;
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("recover_ip:{ip}"),
ip_ceiling,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Versuche. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
let name_ceiling = config::get_usize(
&state.config_cache,
"recover_name_rate_per_15min",
RECOVER_NAME_CEILING_DEFAULT,
)
.await;
let name_key = display_name.to_lowercase();
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("recover:{ip}:{name_key}"),
name_ceiling,
Duration::from_secs(15 * 60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Versuche. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
let event = Event::find_by_slug(&state.pool, &state.config.event_slug)
.await?
.ok_or_else(|| AppError::NotFound("Event nicht gefunden.".into()))?;
let users = User::find_by_event_and_name(&state.pool, event.id, display_name).await?;
if users.is_empty() {
// No user with this name. Run a throwaway bcrypt verify so this branch takes
// the same time as the user-exists path, and return the SAME error as a wrong
// PIN — so "no such name" and "wrong PIN" are indistinguishable by response or
// timing. Display names are already public on the feed, but this still closes
// the /recover enumeration + timing oracle.
let _ = verify_password(body.pin.clone(), dummy_pin_hash().to_string()).await;
return Err(AppError::Unauthorized("PIN ist falsch.".into()));
}
for user in &users {
// Check PIN lockout. If the lockout has expired, also reset the failed-attempt
// counter so the user gets a fresh 3-strike window — otherwise the counter
// stays at 3+ and every subsequent wrong PIN immediately re-locks them, even
// after waiting out the cooldown. Without this reset, a once-locked account
// is effectively permanently fragile.
if let Some(locked_until) = user.pin_locked_until {
if Utc::now() < locked_until {
// The exact deadline is known, so surface it as Retry-After instead of
// making the client guess at the "15 Minuten" in the copy.
let retry_after_secs = (locked_until - Utc::now()).num_seconds().max(1) as u64;
return Err(AppError::TooManyRequests(
"Zu viele Versuche. Bitte warte 15 Minuten.".into(),
Some(retry_after_secs),
));
}
// Lockout window expired — wipe the counter and the timestamp.
User::reset_pin_attempts(&state.pool, user.id).await?;
}
let pin_matches = verify_password(body.pin.clone(), user.recovery_pin_hash.clone()).await;
if pin_matches {
// Reset failed attempts on success
User::reset_pin_attempts(&state.pool, user.id).await?;
let token = jwt::create_token(
user.id,
event.id,
user.role.clone(),
&state.config.jwt_secret,
state.config.session_expiry_days,
)
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?;
let token_hash = jwt::hash_token(&token);
let expires_at = Utc::now() + chrono::Duration::days(state.config.session_expiry_days);
Session::create(&state.pool, user.id, &token_hash, expires_at).await?;
return Ok(Json(RecoverResponse {
jwt: token,
user_id: user.id,
}));
}
// Wrong PIN — increment failure count
let attempts = User::increment_failed_pin(&state.pool, user.id).await?;
tracing::warn!(
user_id = %user.id,
event_id = %event.id,
ip = %ip,
attempts,
"recover: wrong PIN"
);
if attempts >= PIN_LOCK_THRESHOLD {
let lockout = Utc::now() + chrono::Duration::minutes(15);
User::lock_pin(&state.pool, user.id, lockout).await?;
tracing::warn!(
user_id = %user.id,
event_id = %event.id,
ip = %ip,
"recover: account locked for 15 minutes"
);
}
}
Err(AppError::Unauthorized("PIN ist falsch.".into()))
}
#[derive(Deserialize)]
pub struct AdminLoginRequest {
pub password: String,
}
#[derive(Serialize)]
pub struct AdminLoginResponse {
pub jwt: String,
/// The admin's user id + display name, so the client can populate a real identity
/// (own-post affordances, a name on the Account page) instead of a blank session.
pub user_id: Uuid,
pub display_name: String,
}
/// Requests per minute per IP that may reach `verify_password` at all.
///
/// Not a security control — the failure bucket below is. It bounds how deep a queue can form on
/// `BCRYPT_PERMITS`, which is what actually caps the CPU cost.
///
/// Still far above anything a person typing a password produces, but note the honest limitation:
/// unlike the failure bucket, this ceiling CAN refuse a correct password, and on venue NAT every
/// guest shares the operator's IP. It is a smaller number than it first was for exactly that
/// reason — the earlier 120 was chosen when this was the only bound on bcrypt, which made it both
/// too weak to cap CPU and too coarse to be safe for the operator.
const ADMIN_LOGIN_CPU_CEILING: usize = 30;
pub async fn admin_login(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<AdminLoginRequest>,
) -> Result<Json<AdminLoginResponse>, AppError> {
if state.config.admin_password_hash.is_empty() {
return Err(AppError::Forbidden(
"Admin-Login ist nicht konfiguriert.".into(),
));
}
// Throttling here is in two parts, and the ORDER is the whole point.
//
// A single tight IP-keyed bucket checked before the password was verified made this
// endpoint a denial-of-service against its own operator. Every guest at the venue shares
// one public IP behind NAT, `/admin/login` is a public linkable page, and the check ran
// BEFORE `verify_password` — so five requests a minute from any phone in the room kept the
// bucket permanently full and the admin, on that same IP, could never spend a slot.
// Successful logins consumed budget too, so a typo plus a retry on two devices did it by
// accident. And the escape hatch was circular: `admin_login_rate_enabled` can only be
// flipped through `PATCH /admin/config`, which needs the session being blocked.
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let admin_rate_on =
config::get_bool(&state.config_cache, "admin_login_rate_enabled", true).await;
// Part 1: a deliberately GENEROUS ceiling whose only job is to bound bcrypt CPU. Cost-12
// verification is ~250 ms of a core, so an unbounded endpoint is a CPU exhaustion vector
// regardless of whether anyone guesses right. No human typing a password reaches this.
if rate_limits_on
&& admin_rate_on
&& let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("admin_login_cpu:{ip}"),
ADMIN_LOGIN_CPU_CEILING,
Duration::from_secs(60),
)
{
return Err(AppError::TooManyRequests(
"Zu viele Anmeldeversuche. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
let valid = verify_password(
body.password.clone(),
state.config.admin_password_hash.clone(),
)
.await;
if !valid {
// Part 2: the tight bucket, charged ONLY on a wrong password. A correct password is
// never rate-limited, so no amount of guessing by anyone else can lock the operator
// out — which also dissolves the circular escape hatch above. Brute force is still
// bounded: every wrong guess costs a slot, and slots are per-IP.
if rate_limits_on
&& admin_rate_on
&& let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("admin_login_fail:{ip}"),
5,
Duration::from_secs(60),
)
{
tracing::warn!(ip = %ip, "admin_login: wrong password, failure bucket exhausted");
return Err(AppError::TooManyRequests(
"Zu viele Anmeldeversuche. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
tracing::warn!(ip = %ip, "admin_login: wrong password");
return Err(AppError::Unauthorized("Falsches Passwort.".into()));
}
let event = Event::find_or_create(
&state.pool,
&state.config.event_slug,
&state.config.event_name,
)
.await?;
// Find or create the admin user for this event — BY ROLE, never by name.
//
// The name lookup this replaces is what made admin login brickable. Migration 007 makes
// display_name unique per event case-insensitively and `join` had no reserved-name guard,
// so a guest joining as "admin" before the operator's first login made the lookup miss on
// role, the fallback `create("Admin")` violate that index, and `?` return a permanent 500 —
// taking out moderation, config and gallery release with no in-app recovery.
let admin_user = match User::find_admin_for_event(&state.pool, event.id).await? {
Some(u) => u,
None => create_admin_user(&state, event.id).await?,
};
tracing::info!(user_id = %admin_user.id, event_id = %event.id, ip = %ip, "admin_login: success");
let token = jwt::create_token(
admin_user.id,
event.id,
UserRole::Admin,
&state.config.jwt_secret,
1, // Admin sessions expire after 1 day
)
.map_err(|e| AppError::Internal(anyhow::anyhow!(e)))?;
let token_hash = jwt::hash_token(&token);
let expires_at = Utc::now() + chrono::Duration::days(1);
Session::create(&state.pool, admin_user.id, &token_hash, expires_at).await?;
Ok(Json(AdminLoginResponse {
jwt: token,
user_id: admin_user.id,
display_name: admin_user.display_name,
}))
}
/// Create this event's admin row on first successful admin login.
///
/// Prefers the name "Admin". If a legacy database has a guest squatting on it — the state
/// migration 023 renames away, but a row could also predate that or be created between
/// migrations — falls back to a suffixed name rather than failing the login.
///
/// PROMOTING THE SQUATTING ROW WOULD BE A SERIOUS MISTAKE, and is the obvious-looking fix, so
/// it is spelled out: that row carries a `recovery_pin_hash` the guest knows. Setting
/// `role = 'admin'` on it would hand them the admin dashboard through `/recover`, permanently,
/// via a path that needs no password. A separate row under an uglier name is worse UX and much
/// better security — and since the lookup is now by role, the fallback name never has to be
/// guessed again on a later login.
async fn create_admin_user(state: &AppState, event_id: Uuid) -> Result<User, AppError> {
// Admin authenticates via password, but the schema still requires a PIN hash. Generate a
// random unguessable one so the recovery path stays unusable as an escalation route even if
// the role flag were ever cleared.
let dummy_pin: String = (0..32)
.map(|_| rand::rng().random_range(b'a'..=b'z') as char)
.collect();
let dummy_hash = hash_password(dummy_pin, 4).await?;
match User::create_with_role(&state.pool, event_id, "Admin", &dummy_hash, UserRole::Admin).await
{
Ok(u) => Ok(u),
Err(sqlx::Error::Database(db)) if db.is_unique_violation() => {
let fallback = format!("Admin-{}", &Uuid::new_v4().to_string()[..8]);
tracing::warn!(
%event_id, %fallback,
"the name \"Admin\" is held by a non-admin user; creating the admin under a \
fallback name. Rename that guest to free it — do NOT promote their row, they \
know its recovery PIN."
);
Ok(User::create_with_role(
&state.pool,
event_id,
&fallback,
&dummy_hash,
UserRole::Admin,
)
.await?)
}
Err(e) => Err(e.into()),
}
}
pub async fn logout(State(state): State<AppState>, auth: AuthUser) -> Result<StatusCode, AppError> {
Session::delete_by_token_hash(&state.pool, &auth.token_hash).await?;
Ok(StatusCode::NO_CONTENT)
}
/// "Sign out everywhere" — revoke every session for the caller, not just the current one.
/// A single leaked/kept-alive device (a shared phone, a laptop recovered via PIN) can then
/// be cut from any of the user's devices.
pub async fn logout_all(
State(state): State<AppState>,
auth: AuthUser,
) -> Result<StatusCode, AppError> {
Session::delete_all_for_user(&state.pool, auth.user_id).await?;
Ok(StatusCode::NO_CONTENT)
}
#[derive(Deserialize)]
pub struct PinResetRequestBody {
pub display_name: String,
}
/// A guest who forgot their PIN asks a host to reset it in-app. Unauthenticated (they
/// can't log in without the PIN) and rate-limited. Always returns 204 regardless of
/// whether the name exists, so it can't enumerate display names beyond what the public
/// feed already exposes.
pub async fn request_pin_reset(
State(state): State<AppState>,
ConnectInfo(peer): ConnectInfo<SocketAddr>,
headers: HeaderMap,
Json(body): Json<PinResetRequestBody>,
) -> Result<StatusCode, AppError> {
let ip = client_ip(&headers, &peer.ip().to_string());
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
// Coarse per-IP ceiling FIRST, keyed only on the IP so its key is bounded by construction.
// /join got one of these in migration 017 and /recover in 019; this third unauthenticated
// endpoint was simply missed — migration 019's own comment describes exactly this attack.
// Without it, the per-name bucket below is no ceiling at all: the name is attacker-chosen,
// so cycling names mints a fresh bucket every request.
if rate_limits_on {
let ip_ceiling =
config::get_usize(&state.config_cache, "pin_reset_ip_rate_per_min", 30).await;
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("pin_reset_ip:{ip}"),
ip_ceiling,
Duration::from_secs(60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
// Validated BEFORE the per-name key is built, so an unbounded name can never be retained in
// the limiter map. NOTE the 204: this endpoint's contract is that it answers identically
// whether or not the name exists, so it cannot enumerate guests. A 400 here would be a new
// signal — it would distinguish a malformed name from a well-formed unknown one. Silence is
// the correct response, and matches what an empty name already did.
let Ok(display_name) = validate_display_name(&body.display_name) else {
return Ok(StatusCode::NO_CONTENT);
};
if rate_limits_on {
let name_key = display_name.to_lowercase();
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("pin_reset_req:{ip}:{name_key}"),
3,
Duration::from_secs(15 * 60),
) {
return Err(AppError::TooManyRequests(
"Zu viele Anfragen. Bitte warte kurz und versuche es erneut.".into(),
Some(retry_after_secs),
));
}
}
// Single statement so the existing-name and unknown-name paths do IDENTICAL work
// (same event+user index scans, an INSERT that matches 0 rows for an unknown name) —
// no timing oracle despite the always-204 contract. Admins recover via password, so
// they're excluded from the join and never get a reset request queued.
let _ = sqlx::query(
"INSERT INTO pin_reset_request (event_id, user_id)
SELECT e.id, u.id
FROM event e
JOIN \"user\" u
ON u.event_id = e.id
AND lower(u.display_name) = lower($2)
AND u.role <> 'admin'
WHERE e.slug = $1
ON CONFLICT (user_id) DO NOTHING",
)
.bind(&state.config.event_slug)
.bind(display_name)
.execute(&state.pool)
.await;
// Broadcast unconditionally (carries no per-name info) so an online host's request
// badge updates without revealing whether the name existed.
let _ = state
.sse_tx
.send(crate::state::SseEvent::new("pin-reset-requested", "{}"));
Ok(StatusCode::NO_CONTENT)
}
#[cfg(test)]
mod tests {
use super::*;
/// Control characters are rejected at the door. Newlines in particular: several 4xx messages
/// interpolate the display name and those are logged, so a name carrying a newline plus a
/// plausible prefix would let two unauthenticated requests forge lines in the event's only
/// forensic record. `error.rs` escapes on output too; this keeps it out of the database and
/// the feed byline in the first place.
#[test]
fn a_display_name_may_not_carry_control_characters() {
for bad in ["Anna\nERROR forged", "Anna\rX", "Anna\u{0}X", "A\u{7}B"] {
assert!(validate_display_name(bad).is_err(), "{bad:?} must be rejected");
}
// Real guests have accents, emoji and non-Latin names — never reject those.
for good in ["Anna", "Zo\u{eb}", "Jos\u{e9}", "\u{5c71}\u{7530}", "Anna \u{1f389}"] {
assert!(validate_display_name(good).is_ok(), "{good:?} must be allowed");
}
}
/// THE defect, stated as arithmetic: the account-lock threshold sat BELOW the per-(IP, name)
/// attempt ceiling, so a single IP could exhaust it and lock any guest whose display name is
/// visible on the feed — every 15 minutes, indefinitely. The tier meant to protect a guest
/// was the cheapest way to attack them.
///
/// The fix is the ORDERING, not either number on its own, so that is what this pins.
#[test]
fn one_ip_cannot_reach_the_account_lock() {
assert!(
PIN_LOCK_THRESHOLD as usize >= RECOVER_NAME_CEILING_DEFAULT * 3,
"locking a victim must require at least three distinct sources; \
threshold {PIN_LOCK_THRESHOLD} vs per-IP ceiling {RECOVER_NAME_CEILING_DEFAULT}"
);
}
/// Raising the threshold must not quietly weaken brute-force resistance. 4-digit PINs, and
/// the lockout window is 15 minutes, so an attacker gets PIN_LOCK_THRESHOLD tries per window.
#[test]
fn the_raised_threshold_still_makes_guessing_a_four_digit_pin_impractical() {
let attempts_per_hour = PIN_LOCK_THRESHOLD as u64 * 4; // four 15-minute windows
let hours_for_full_keyspace = 10_000 / attempts_per_hour;
assert!(
hours_for_full_keyspace >= 168,
"exhausting 10k PINs would take {hours_for_full_keyspace}h — under a week is too fast"
);
}
#[test]
fn reserved_names_are_matched_case_insensitively_and_trimmed() {
for name in ["admin", "Admin", "ADMIN", " Host ", "EventSnap"] {
assert!(is_reserved_display_name(name), "{name} must be reserved");
}
}
/// A substring match here would reject perfectly ordinary names, which is a worse outcome
/// than the impersonation the list guards against.
#[test]
fn names_that_merely_contain_a_reserved_word_are_allowed() {
for name in ["Administrata", "Hostess", "Adminah", "Ghost", "hosting"] {
assert!(!is_reserved_display_name(name), "{name} must be allowed");
}
}
#[test]
fn display_names_are_bounded_before_they_can_become_a_rate_limit_key() {
assert!(validate_display_name(" Lena ").is_ok());
assert_eq!(validate_display_name(" Lena ").unwrap(), "Lena");
// The case that made the limiter itself the exhaustion primitive.
assert!(validate_display_name(&"a".repeat(51)).is_err());
assert!(validate_display_name(&"a".repeat(2_000_000)).is_err());
assert!(validate_display_name(" ").is_err());
assert!(validate_display_name("bad\0name").is_err());
}
}