Follow-up to the comprehensive code review. Five batches: 1. Cross-event authorization: host_delete_upload, unban_user, and host_delete_comment now scope by auth.event_id. Adds Upload::find_by_id_and_event / soft_delete_in_event and a Comment::soft_delete_in_event variant that joins through upload. 2. Token exposure: SSE auth no longer puts the JWT in the URL. New /api/v1/stream/ticket endpoint mints a short-lived single-use ticket bound to the session; the EventSource passes ?ticket=... instead. Refuse to start in APP_ENV=production with the dev JWT sentinel; warn loudly otherwise. 3. Account hardening: per-IP+name rate limit on /recover (mitigates targeted lockout DoS), per-IP rate limit on /admin/login, random 32-char admin recovery PIN (replaces "0000"), structured tracing events for wrong PIN, lockout, failed admin login, ban/unban/role change/pin-reset/host-delete. 4. DoS / correctness: comment listing paginated (LIMIT 50 + ?before= cursor), hashtag extraction whitelisted to ASCII alnum+underscore (≤40 chars) with unit tests, display_name / caption / comment body length validated in chars rather than bytes. 5. Cleanup: session-touch failures now logged, DATABASE_MAX_CONNECTIONS env var (default 10). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
105 lines
3.8 KiB
Rust
105 lines
3.8 KiB
Rust
//! Startup recovery + periodic background hygiene.
|
|
//!
|
|
//! Two responsibilities:
|
|
//!
|
|
//! 1. **Startup sweep** — when the server boots, fix rows left in an "in-progress"
|
|
//! state by the previous (possibly crashed) instance. Compression and export jobs
|
|
//! each leave a status row when they begin; if the process is killed mid-run, that
|
|
//! row stays `'processing'` / `'running'` forever, blocking re-tries and leaving
|
|
//! users staring at a spinner. Resetting them on startup recovers gracefully.
|
|
//!
|
|
//! 2. **Periodic tasks** — pruning that should happen "every hour" rather than per
|
|
//! request: expired sessions (otherwise the table grows unboundedly), and the
|
|
//! rate-limiter's in-memory windows (so keys for IPs that left long ago don't
|
|
//! accumulate).
|
|
|
|
use std::time::Duration;
|
|
|
|
use sqlx::PgPool;
|
|
|
|
use crate::services::rate_limiter::RateLimiter;
|
|
use crate::services::sse_tickets::SseTicketStore;
|
|
|
|
/// Reset rows left in flight by a previous crashed instance. Run once on startup,
|
|
/// before the HTTP server starts taking requests, so users never observe the
|
|
/// half-state.
|
|
pub async fn startup_recovery(pool: &PgPool) {
|
|
// Uploads whose preview generation was interrupted. Marking them 'failed' is
|
|
// safer than re-queueing — the original file is still on disk, the user can
|
|
// delete + re-upload if they care, and we avoid double-processing risk.
|
|
match sqlx::query(
|
|
"UPDATE upload SET compression_status = 'failed'
|
|
WHERE compression_status = 'processing'",
|
|
)
|
|
.execute(pool)
|
|
.await
|
|
{
|
|
Ok(r) if r.rows_affected() > 0 => {
|
|
tracing::warn!(
|
|
"startup recovery: reset {} stuck upload(s) from 'processing' to 'failed'",
|
|
r.rows_affected()
|
|
);
|
|
}
|
|
Ok(_) => {}
|
|
Err(e) => tracing::error!("startup recovery: failed to sweep uploads: {e:#}"),
|
|
}
|
|
|
|
// Export jobs interrupted mid-run. Mark 'failed' so the host can re-trigger.
|
|
// The `UNIQUE(event_id, type)` constraint would otherwise block re-release.
|
|
match sqlx::query(
|
|
"UPDATE export_job
|
|
SET status = 'failed',
|
|
error_message = COALESCE(error_message, 'Server-Neustart während des Exports')
|
|
WHERE status = 'running'",
|
|
)
|
|
.execute(pool)
|
|
.await
|
|
{
|
|
Ok(r) if r.rows_affected() > 0 => {
|
|
tracing::warn!(
|
|
"startup recovery: reset {} stuck export job(s) from 'running' to 'failed'",
|
|
r.rows_affected()
|
|
);
|
|
}
|
|
Ok(_) => {}
|
|
Err(e) => tracing::error!("startup recovery: failed to sweep export_job: {e:#}"),
|
|
}
|
|
}
|
|
|
|
/// Spawns a background task that periodically:
|
|
/// - deletes session rows whose `expires_at` is more than a day in the past
|
|
/// - prunes the in-memory rate-limiter HashMap of empty windows
|
|
/// - drops expired SSE tickets (30s TTL but the map keeps the slot until pruned)
|
|
///
|
|
/// Cadence is 1h — fine for both jobs at our scale.
|
|
pub fn spawn_periodic_tasks(
|
|
pool: PgPool,
|
|
rate_limiter: RateLimiter,
|
|
sse_tickets: SseTicketStore,
|
|
) {
|
|
tokio::spawn(async move {
|
|
let mut tick = tokio::time::interval(Duration::from_secs(3600));
|
|
// Fire the first tick immediately, then hourly.
|
|
tick.tick().await;
|
|
loop {
|
|
tick.tick().await;
|
|
cleanup_sessions(&pool).await;
|
|
rate_limiter.prune();
|
|
sse_tickets.prune();
|
|
}
|
|
});
|
|
}
|
|
|
|
async fn cleanup_sessions(pool: &PgPool) {
|
|
match sqlx::query("DELETE FROM session WHERE expires_at < NOW() - INTERVAL '1 day'")
|
|
.execute(pool)
|
|
.await
|
|
{
|
|
Ok(r) if r.rows_affected() > 0 => {
|
|
tracing::info!("cleaned up {} expired session(s)", r.rows_affected());
|
|
}
|
|
Ok(_) => {}
|
|
Err(e) => tracing::warn!("session cleanup failed: {e:#}"),
|
|
}
|
|
}
|