Files
EventSnap/backend/src/handlers/upload.rs
MechaCat02 e52b2f1cd1 fix(upload): reclaim the bytes of uploads that never finish
Reclaim was a dozen explicit remove_file calls on the handler's return paths.
That covers every way the handler can FINISH and none of the ways it can simply
STOP. When a client disconnects mid-body — a phone leaving wifi, iOS killing a
backgrounded PWA, the user hitting back — axum drops the handler future at an
.await inside field.chunk() and no return path runs at all. The partial file then
survives forever: it has no upload row, so cleanup_deleted_media (row-driven)
can never see it, and no sweeper covered the originals directory. The triggers
are routine rather than adversarial, and the client keeps the blob and auto-
retries on every `online` event, so one large video over bad wifi leaves several
copies.

Those bytes are also invisible to the quota while still consuming the free disk
that compute_storage_quota divides among guests — so orphans silently shrink
every guest's ceiling while the admin widget under-reports. All three volumes
share one filesystem; the end state is Postgres unable to write WAL.

TempFileGuard is an RAII guard, because dropping the future is exactly what runs
Drop — it is the only construct that survives cancellation. Armed before the file
can exist, disarmed only after tx.commit() succeeds. The twelve explicit cleanups
are deleted so one owner holds the rule.

The subtler half is the rename. It happens BEFORE the commit, so between them the
file exists under its final name with no row pointing at it — an orphan that
looks legitimate. The guard is RETARGETED there rather than disarmed, and the
retarget sits on the same poll as the rename with no .await between, which is
what makes that window uncancellable.

sweep_orphan_originals is the backstop for the process that was killed, where no
Drop can run at all. Hourly, alongside the existing sweeps: .tmp files past the
window go unconditionally (a .tmp never has a row by construction), other files
are batched 500 at a time through a single NOT EXISTS query.

Two things that look like oversights and are not, both commented in place:
- the 6h window is what makes the sweep safe against the rename-before-commit
  ordering, since a committing upload is briefly indistinguishable from an
  orphan. It must not be shortened to speed up a test.
- the NOT EXISTS deliberately does NOT filter deleted_at IS NULL. A soft-deleted
  row still points at its file during its retention window, and reclaiming that
  is cleanup_deleted_media's job; filtering here would race the two sweeps and
  destroy the files the recovery window exists to preserve.

Verified against the real schema: given a live original, a soft-deleted one and a
true orphan, the query returns only the orphan.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 22:56:26 +02:00

1184 lines
47 KiB
Rust

use std::time::Duration;
use axum::Json;
use axum::extract::{Multipart, Path, State};
use axum::http::StatusCode;
use chrono::{DateTime, Utc};
use serde::Deserialize;
use uuid::Uuid;
use crate::auth::middleware::AuthUser;
use crate::error::AppError;
use crate::models::hashtag::{self, Hashtag};
use crate::models::upload::{Upload, UploadDto};
use crate::models::user::User;
use crate::services::config;
use crate::state::AppState;
const MAX_CAPTION_LENGTH: usize = 2000;
/// Owns the bytes an in-flight upload has written to disk, and deletes them unless the
/// request reaches the point where a database row takes ownership.
///
/// Reclaim used to be a dozen explicit `remove_file` calls on the handler's return paths.
/// That covers every way the handler can FINISH, and none of the ways it can simply STOP:
/// when a client disconnects mid-body — a phone leaving wifi, iOS killing a backgrounded
/// PWA, the user hitting back — axum drops the handler future at a `.await` inside
/// `field.chunk()`, and no return path runs at all. The partial file then survives forever:
/// it has no upload row, so `cleanup_deleted_media` (which is row-driven) can never see it,
/// and no sweeper existed for the originals directory. Those bytes are also invisible to the
/// quota while still consuming the free disk that `quota_limit_bytes` divides among guests.
///
/// A drop guard is the only construct that survives cancellation, because dropping the future
/// is exactly what runs it.
struct TempFileGuard {
/// `None` once disarmed — a row now owns these bytes.
path: Option<std::path::PathBuf>,
}
impl TempFileGuard {
fn new(path: std::path::PathBuf) -> Self {
Self { path: Some(path) }
}
/// Follow the bytes to their new location after a rename.
///
/// NOT `disarm`. Between the rename and the commit the file exists under its FINAL name
/// with still no row pointing at it, so that window needs guarding just as much as the
/// `.tmp` did — arguably more, since a leftover final-named original looks legitimate.
fn retarget(&mut self, path: std::path::PathBuf) {
self.path = Some(path);
}
/// Hand ownership to the committed row. Only correct after `tx.commit()` succeeds.
fn disarm(&mut self) {
self.path = None;
}
}
impl Drop for TempFileGuard {
fn drop(&mut self) {
let Some(path) = self.path.take() else {
return;
};
// std::fs, not tokio::fs: `Drop` cannot await, and a runtime-dependent unlink is not
// guaranteed a live runtime here (shutdown drops in-flight tasks).
match std::fs::remove_file(&path) {
Ok(()) => tracing::debug!(path = %path.display(), "reclaimed an abandoned upload"),
Err(e) if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
tracing::warn!(error = ?e, path = %path.display(), "failed to reclaim an abandoned upload")
}
}
}
}
/// Allowlist of accepted media types, keyed by the MIME that `infer` derives from
/// the file's magic bytes. The detected MIME (not the client-declared one) is what
/// we trust, store, and hand to the compression pipeline — so a text-based payload
/// (SVG/HTML/JS) can never be stored or served on-origin. Each entry maps to the
/// server-controlled file extension we persist the original under.
///
/// HEIC/HEIF are deliberately excluded: the preview pipeline (`image` crate, and
/// the bundled ffmpeg 6.1) cannot decode them, so accepting them would store files
/// that never get a thumbnail. iOS Safari already transcodes HEIC→JPEG when a photo
/// is selected via a file input, so this rejects only the rare HEIC-preserving
/// upload path — with a clear error rather than a silently broken post.
const ALLOWED_MEDIA: &[(&str, &str)] = &[
("image/jpeg", "jpg"),
("image/png", "png"),
("image/webp", "webp"),
("image/gif", "gif"),
("video/mp4", "mp4"),
("video/quicktime", "mov"),
("video/webm", "webm"),
];
pub async fn upload(
State(state): State<AppState>,
auth: AuthUser,
mut multipart: Multipart,
) -> Result<(StatusCode, Json<UploadDto>), AppError> {
// Rate limit: N uploads per hour per user. Gated by master + per-endpoint toggles.
let rate_limits_on = config::get_bool(&state.config_cache, "rate_limits_enabled", true).await;
let upload_rate_on = config::get_bool(&state.config_cache, "upload_rate_enabled", true).await;
if rate_limits_on && upload_rate_on {
let upload_rate =
config::get_i64(&state.config_cache, "upload_rate_per_hour", 100).await as usize;
if let Err(retry_after_secs) = state.rate_limiter.check_with_retry(
format!("upload:{}", auth.user_id),
upload_rate,
Duration::from_secs(3600),
) {
drain_multipart(multipart).await;
return Err(AppError::TooManyRequests(
"Du hast dein Upload-Limit für diese Stunde erreicht.".into(),
Some(retry_after_secs),
));
}
}
// Check if user is banned
let user = User::find_by_id(&state.pool, auth.user_id)
.await?
.ok_or_else(|| AppError::NotFound("Benutzer nicht gefunden.".into()))?;
if user.is_banned {
drain_multipart(multipart).await;
return Err(AppError::Forbidden("Du bist gesperrt.".into()));
}
// Check if uploads are locked
let event = crate::models::event::Event::find_by_slug(&state.pool, &state.config.event_slug)
.await?
.ok_or_else(|| AppError::NotFound("Event nicht gefunden.".into()))?;
if event.uploads_locked_at.is_some() {
drain_multipart(multipart).await;
// Reversible: a host can reopen the event, so the client keeps the queued blob and
// retries on `event-opened` rather than purging it (UploadsLocked, not Forbidden).
return Err(AppError::UploadsLocked("Uploads sind gesperrt.".into()));
}
// Belt-and-suspenders on top of the lock (release ⇒ lock): once the gallery is
// released the export has been snapshotted, so a late upload could never make it into
// the keepsake. Reject it explicitly rather than silently diverging the live feed.
// Also reversible (reopen clears `export_released_at`), so likewise UploadsLocked.
if event.export_released_at.is_some() {
drain_multipart(multipart).await;
return Err(AppError::UploadsLocked(
"Galerie wurde bereits freigegeben.".into(),
));
}
// Read config limits from DB
let max_image_mb: i64 = config::get_i64(&state.config_cache, "max_image_size_mb", 20).await;
let max_video_mb: i64 = config::get_i64(&state.config_cache, "max_video_size_mb", 500).await;
// The uploaded file is streamed straight to a temp file on disk (never buffered
// whole in memory — a 500 MB video used to cost 500 MB of RAM per concurrent
// upload). We only keep the first ≤512 bytes in memory for magic-byte sniffing.
// On success the temp file is renamed into place under its detected extension.
let upload_id = Uuid::new_v4();
let event_slug = &state.config.event_slug;
let originals_dir = state
.config
.media_path
.join(format!("originals/{event_slug}"));
let temp_abs = originals_dir.join(format!("{upload_id}.tmp"));
// Armed before anything can create the file, so there is no window in which bytes exist
// unowned. From here on, EVERY exit — return, `?`, panic, or the future being dropped
// mid-body by a client disconnect — reclaims them, and the explicit `remove_file` calls
// that used to be sprinkled over the return paths are gone. One owner, one rule.
let mut file_guard = TempFileGuard::new(temp_abs.clone());
let mut streamed: Option<(i64, Vec<u8>)> = None; // (size, head bytes for sniffing)
let mut caption: Option<String> = None;
let mut hashtags_csv: Option<String> = None;
// The multipart read is wrapped so the field loop can use `?` freely; reclaiming the temp
// file on failure is `file_guard`'s job, not this block's.
let parse_result: Result<(), AppError> = async {
while let Some(field) = multipart
.next_field()
.await
.map_err(|e| AppError::BadRequest(e.to_string()))?
{
let name = field.name().unwrap_or_default().to_string();
match name.as_str() {
"file" => {
// The client-declared Content-Type does NOT determine the stored
// MIME/extension — those come from the file's magic bytes below. The
// declared type only picks the streaming cap so an oversized body is
// aborted early; a mislabelled type only makes the cap *stricter*
// (safe), and the authoritative per-class check still runs on the
// detected type.
let declared = field.content_type().unwrap_or("").to_string();
let cap_bytes = if declared.starts_with("video/") {
(max_video_mb * 1024 * 1024) as usize
} else if declared.starts_with("image/") {
(max_image_mb * 1024 * 1024) as usize
} else {
(max_image_mb.max(max_video_mb) * 1024 * 1024) as usize
};
tokio::fs::create_dir_all(&originals_dir)
.await
.map_err(|e| AppError::Internal(e.into()))?;
streamed = Some(stream_field_to_file(field, &temp_abs, cap_bytes).await?);
}
"caption" => {
caption = Some(
field
.text()
.await
.map_err(|e| AppError::BadRequest(e.to_string()))?,
);
}
"hashtags" => {
hashtags_csv = Some(
field
.text()
.await
.map_err(|e| AppError::BadRequest(e.to_string()))?,
);
}
_ => {}
}
}
Ok(())
}
.await;
parse_result?;
// From here on the temp file may exist. Every exit reclaims it via `file_guard` — see
// TempFileGuard for why the explicit per-branch cleanup this replaced was not enough.
let (size, head) = match streamed {
Some(s) => s,
None => return Err(AppError::BadRequest("Keine Datei hochgeladen.".into())),
};
// Validate caption length. Counted in chars (code points) to match the
// "Zeichen" wording in the error message — `.len()` would be bytes and
// reject perfectly valid German/emoji captions early.
if let Some(ref cap) = caption
&& cap.chars().count() > MAX_CAPTION_LENGTH
{
return Err(AppError::BadRequest(format!(
"Beschreibung ist zu lang. Maximum: {} Zeichen.",
MAX_CAPTION_LENGTH
)));
}
// Determine the file type from its magic bytes and require it to be on the
// allowlist. `infer` returns None for text-based payloads (SVG/HTML/JS), so
// those are rejected outright — closing the stored-XSS vector. Both the MIME
// we persist and the on-disk extension come from the detected type, never from
// client-supplied values.
let kind = match infer::get(&head) {
Some(k) => k,
None => {
return Err(AppError::BadRequest(
"Dateityp nicht erkannt oder nicht unterstützt.".into(),
));
}
};
let (mime, ext) = match ALLOWED_MEDIA
.iter()
.find(|(allowed, _)| *allowed == kind.mime_type())
.map(|(m, e)| ((*m).to_string(), *e))
{
Some(v) => v,
None => {
return Err(AppError::BadRequest(format!(
"Dateityp wird nicht unterstützt: {}.",
kind.mime_type()
)));
}
};
// Validate file size against the authoritative per-detected-class limit.
let max_bytes = if mime.starts_with("video/") {
max_video_mb * 1024 * 1024
} else {
max_image_mb * 1024 * 1024
};
if size > max_bytes {
return Err(AppError::BadRequest(format!(
"Datei ist zu groß. Maximum: {} MB.",
max_bytes / (1024 * 1024)
)));
}
// Images only: refuse anything the compression worker could never decode, reading just
// the header. Without this the upload is accepted with a 201 and then silently
// soft-deleted minutes later when the worker gives up — the guest sees the photo
// vanish with, at best, a vague "could not be processed". Rejecting here gives them a
// reason at the door that they can act on, and it uses the SAME budget the worker
// enforces, so admission and processing cannot disagree.
if mime.starts_with("image/") && crate::services::imaging::exceeds_decode_budget(&temp_abs) {
let mp = crate::services::imaging::megapixels(&temp_abs);
tracing::info!(
%mime, megapixels = ?mp,
"rejecting an image that exceeds the decode budget at admission"
);
let detail = mp.map_or(String::new(), |mp| format!(" (ca. {mp:.0} Megapixel)"));
return Err(AppError::BadRequest(format!(
"Bild hat zu viele Bildpunkte{detail} und kann nicht verarbeitet werden. \
Bitte verkleinere es und lade es erneut hoch."
)));
}
// Per-user storage quota — dynamic formula based on available disk space and the
// number of active uploaders. Gated by master + per-area toggles so the admin can
// disable it on trusted instances.
let quota_on = config::get_bool(&state.config_cache, "quota_enabled", true).await;
let storage_quota_on =
config::get_bool(&state.config_cache, "storage_quota_enabled", true).await;
// When quota is enforced, this holds the byte ceiling so the increment UPDATE below can
// enforce it atomically (`WHERE total + size <= limit`). Without that guard, two
// concurrent uploads from the same user (e.g. phone + laptop) both pass this stale
// pre-check and both increment, blowing past the quota. The pre-check stays as a
// fast path that avoids the disk write when the user is already clearly over.
let mut quota_limit: Option<i64> = None;
if quota_on && storage_quota_on {
let estimate = compute_storage_quota(&state).await;
if let Some(limit) = estimate.limit_bytes {
quota_limit = Some(limit);
let prospective_total = user.total_upload_bytes.saturating_add(size);
if prospective_total > limit {
return Err(AppError::QuotaExceeded(
"Du hast dein Upload-Limit für dieses Event erreicht.".into(),
));
}
}
}
// All checks passed — atomically move the temp file to its final, extension-correct
// path (same directory, so the rename is cheap and atomic).
let relative_path = format!("originals/{event_slug}/{upload_id}.{ext}");
let absolute_path = state.config.media_path.join(&relative_path);
tokio::fs::rename(&temp_abs, &absolute_path)
.await
.map_err(|e| AppError::Internal(e.into()))?;
// THERE MUST BE NO `.await` BETWEEN THE RENAME AND THIS LINE. Both statements resolve on
// the same poll, so the future cannot be dropped between them and the guard is never
// pointing at a path that no longer holds the bytes. If the rename fails the guard still
// owns `temp_abs`, which is why retargeting comes after it rather than before.
file_guard.retarget(absolute_path.clone());
// Process hashtags from caption and explicit CSV
let mut tags: Vec<String> = Vec::new();
if let Some(ref cap) = caption {
tags.extend(hashtag::extract_hashtags(cap));
}
if let Some(ref csv) = hashtags_csv {
for tag in csv.split(',') {
let t = tag.trim().trim_start_matches('#').to_lowercase();
if !t.is_empty() {
tags.push(t);
}
}
}
tags.sort();
tags.dedup();
// Quota accounting, the upload row, and its hashtag links must be atomic: a
// crash between the bytes increment and the insert would permanently charge
// bytes with no row to reclaim them (silent quota erosion / spurious lockout).
let tx_result: Result<Upload, AppError> = async {
let mut tx = state.pool.begin().await?;
// RE-CHECK THE LOCK, UNDER A ROW LOCK, INSIDE THE COMMIT TX.
//
// The pre-flight check at the top of this handler ran BEFORE we streamed the body — which
// for a 500 MB video is minutes. Trusting it here is a TOCTOU that silently loses photos
// from the keepsake, and it is the real cause of the "stale keepsake" bug that survived
// three rounds of fixes inside the export state machine:
//
// 1. guest starts a big upload; the lock check passes (event open)
// 2. host releases the gallery → uploads lock, export workers snapshot the uploads table
// 3. this upload commits AFTER that snapshot → it shows up in the live feed but is
// MISSING from the downloaded keepsake, permanently (nothing ever regenerates it)
//
// `FOR SHARE` conflicts with the `UPDATE event` in `release_gallery`, which serializes us
// against it. Either we take the lock first — and release (hence the export snapshot) is
// strictly ordered after our commit, so the snapshot CONTAINS this upload — or release
// commits first and we observe the lock here and reject. Either way the keepsake is
// complete. `UploadsLocked` (not Forbidden) is reversible: the client keeps the blob and
// resumes it when the host reopens.
let (locked_at, released_at): (Option<DateTime<Utc>>, Option<DateTime<Utc>>) =
sqlx::query_as(
"SELECT uploads_locked_at, export_released_at FROM event WHERE id = $1 FOR SHARE",
)
.bind(auth.event_id)
.fetch_one(&mut *tx)
.await?;
if locked_at.is_some() || released_at.is_some() {
return Err(AppError::UploadsLocked("Uploads sind gesperrt.".into()));
}
// Increment the user's byte total. When a quota is in force, guard it atomically
// (`total + size <= limit`) so two concurrent uploads can't both slip past the
// stale pre-check — the loser's UPDATE matches 0 rows and we abort with the same
// terminal quota error (the tx rolls back on drop; the on-disk file is cleaned by
// the error path below).
let inc = if let Some(limit) = quota_limit {
sqlx::query(
"UPDATE \"user\" SET total_upload_bytes = total_upload_bytes + $2
WHERE id = $1 AND total_upload_bytes + $2 <= $3",
)
.bind(auth.user_id)
.bind(size)
.bind(limit)
.execute(&mut *tx)
.await?
} else {
sqlx::query(
"UPDATE \"user\" SET total_upload_bytes = total_upload_bytes + $2 WHERE id = $1",
)
.bind(auth.user_id)
.bind(size)
.execute(&mut *tx)
.await?
};
if inc.rows_affected() == 0 {
return Err(AppError::QuotaExceeded(
"Du hast dein Upload-Limit für dieses Event erreicht.".into(),
));
}
let upload = Upload::create(
&mut *tx,
auth.event_id,
auth.user_id,
&relative_path,
&mime,
size,
caption.as_deref(),
)
.await?;
for tag in &tags {
let h = Hashtag::upsert(&mut *tx, auth.event_id, tag).await?;
Hashtag::link_to_upload(&mut *tx, upload.id, h.id).await?;
}
tx.commit().await?;
Ok(upload)
}
.await;
// The committed row now references these bytes — hand ownership over. Anything other than
// a successful commit leaves the guard armed, so the file is reclaimed on the way out.
let upload = tx_result?;
file_guard.disarm();
// Spawn compression task
state
.compression
.process(upload.id, relative_path, mime.clone());
// Broadcast SSE event
let dto = UploadDto {
id: upload.id,
user_id: auth.user_id,
uploader_name: user.display_name,
preview_url: None,
thumbnail_url: None,
mime_type: mime,
caption,
hashtags: tags,
like_count: 0,
comment_count: 0,
liked_by_me: false,
created_at: upload.created_at,
};
let _ = state.sse_tx.send(crate::state::SseEvent::new(
"new-upload",
serde_json::to_string(&dto).unwrap_or_default(),
));
Ok((StatusCode::CREATED, Json(dto)))
}
#[derive(Deserialize)]
pub struct EditUploadRequest {
pub caption: Option<String>,
pub hashtags: Option<Vec<String>>,
}
pub async fn edit_upload(
State(state): State<AppState>,
auth: AuthUser,
Path(upload_id): Path<Uuid>,
Json(body): Json<EditUploadRequest>,
) -> Result<StatusCode, AppError> {
// Banned users keep read access but cannot mutate (USER_JOURNEYS §10).
if auth.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into()));
}
let upload = Upload::find_by_id_and_event(&state.pool, upload_id, auth.event_id)
.await?
.ok_or_else(|| AppError::NotFound("Upload nicht gefunden.".into()))?;
if upload.user_id != auth.user_id {
return Err(AppError::Forbidden("Nur eigene Uploads bearbeiten.".into()));
}
// Caption update + hashtag wipe-then-relink in one transaction, so a crash
// mid-relink can't leave the upload with its hashtags stripped.
//
// Editing is intentionally allowed while uploads are locked or the gallery is released — like
// comments and likes, the lock freezes *new uploads* only (USER_JOURNEYS §9.3). But a caption
// is embedded in the HTML viewer keepsake (the ZIP holds media only — see export.rs), so an
// edit AFTER release must regenerate the viewer, or the downloadable keepsake keeps showing the
// old caption forever while the live feed shows the new one. Same atomicity as delete_upload:
// the edit and its invalidation share one tx so a dropped handler can't leave them disagreeing.
// `Affects::ViewerOnly` carries the finished ZIP forward (the media didn't change); when the
// gallery isn't released, `invalidate_and_arm` returns None and this is a no-op.
let mut tx = state.pool.begin().await?;
if let Some(ref caption) = body.caption {
Upload::update_caption(&mut *tx, upload_id, Some(caption)).await?;
}
if let Some(ref hashtags) = body.hashtags {
Hashtag::unlink_all_from_upload(&mut *tx, upload_id).await?;
for tag in hashtags {
let h = Hashtag::upsert(&mut *tx, auth.event_id, tag).await?;
Hashtag::link_to_upload(&mut *tx, upload_id, h.id).await?;
}
}
let regen = crate::services::export::invalidate_and_arm(
&mut tx,
&state.config.event_slug,
crate::services::export::Affects::ViewerOnly,
)
.await?;
tx.commit().await?;
if let Some(r) = regen {
crate::handlers::host::start_regen(&state, r);
}
Ok(StatusCode::OK)
}
pub async fn delete_upload(
State(state): State<AppState>,
auth: AuthUser,
Path(upload_id): Path<Uuid>,
) -> Result<StatusCode, AppError> {
// Banned users keep read access but cannot mutate (USER_JOURNEYS §10).
if auth.is_banned {
return Err(AppError::Forbidden("Du bist gesperrt.".into()));
}
let upload = Upload::find_by_id_and_event(&state.pool, upload_id, auth.event_id)
.await?
.ok_or_else(|| AppError::NotFound("Upload nicht gefunden.".into()))?;
if upload.user_id != auth.user_id {
return Err(AppError::Forbidden("Nur eigene Uploads löschen.".into()));
}
// Atomic with the keepsake invalidation: a guest removing their own photo must have it removed
// from the downloadable archive too, and a half-applied delete would leave it there forever.
let mut tx = state.pool.begin().await?;
Upload::soft_delete_in_event(&mut tx, upload_id, auth.event_id).await?;
let regen = crate::services::export::invalidate_and_arm(
&mut tx,
&state.config.event_slug,
crate::services::export::Affects::Both,
)
.await?;
tx.commit().await?;
if let Some(r) = regen {
crate::handlers::host::start_regen(&state, r);
}
// Evict the card live on every other feed + the projector diashow — otherwise
// a self-deleted post lingers until each viewer manually reloads. Same event
// the host-delete path already emits and the frontend already handles.
let _ = state.sse_tx.send(crate::state::SseEvent::new(
"upload-deleted",
serde_json::json!({ "upload_id": upload_id }).to_string(),
));
Ok(StatusCode::NO_CONTENT)
}
/// Number of leading bytes retained in memory for magic-byte (`infer`) sniffing. Every
/// allowed type's signature sits well within this; 512 is comfortably generous.
const HEAD_SNIFF_BYTES: usize = 512;
/// Stream a multipart field straight to `dest`, aborting with a 400 the moment it
/// exceeds `max_bytes`. Only the first [`HEAD_SNIFF_BYTES`] bytes are kept in memory
/// (for type detection); the rest goes chunk-by-chunk to disk, so peak memory is a
/// single chunk rather than the whole file. Returns `(total_size, head_bytes)`. On any
/// error the partial temp file is removed so no stray `.tmp` is left behind.
async fn stream_field_to_file(
mut field: axum::extract::multipart::Field<'_>,
dest: &std::path::Path,
max_bytes: usize,
) -> Result<(i64, Vec<u8>), AppError> {
use tokio::io::AsyncWriteExt;
let mut file = tokio::fs::File::create(dest)
.await
.map_err(|e| AppError::Internal(e.into()))?;
let mut total: usize = 0;
let mut head: Vec<u8> = Vec::with_capacity(HEAD_SNIFF_BYTES);
loop {
let chunk = match field.chunk().await {
Ok(Some(c)) => c,
Ok(None) => break,
Err(e) => {
let _ = file.shutdown().await;
let _ = tokio::fs::remove_file(dest).await;
return Err(AppError::BadRequest(format!(
"Datei konnte nicht gelesen werden: {e}"
)));
}
};
total = total.saturating_add(chunk.len());
if total > max_bytes {
let _ = file.shutdown().await;
let _ = tokio::fs::remove_file(dest).await;
return Err(AppError::BadRequest(format!(
"Datei ist zu groß. Maximum: {} MB.",
max_bytes / (1024 * 1024)
)));
}
if head.len() < HEAD_SNIFF_BYTES {
let need = HEAD_SNIFF_BYTES - head.len();
head.extend_from_slice(&chunk[..need.min(chunk.len())]);
}
if let Err(e) = file.write_all(&chunk).await {
let _ = tokio::fs::remove_file(dest).await;
return Err(AppError::Internal(e.into()));
}
}
if let Err(e) = file.flush().await {
let _ = tokio::fs::remove_file(dest).await;
return Err(AppError::Internal(e.into()));
}
Ok((total as i64, head))
}
/// Drain a multipart body so the HTTP connection stays clean when returning an early error.
/// Without draining, the client may still be sending the body after we've sent our response,
/// which can corrupt the keep-alive connection for subsequent requests.
async fn drain_multipart(mut mp: Multipart) {
while let Ok(Some(mut field)) = mp.next_field().await {
while field.chunk().await.ok().flatten().is_some() {}
}
}
/// Snapshot of the dynamic per-user quota used both by the upload pre-check and the
/// `GET /me/quota` endpoint. `limit_bytes = None` means quota enforcement is currently
/// off (the frontend hides the widget in that case).
pub struct QuotaEstimate {
pub limit_bytes: Option<i64>,
pub active_uploaders: i64,
pub free_disk_bytes: i64,
/// The tolerance factor the limit above was computed with. Carried on the snapshot so the
/// number is self-describing; no caller reads it back today.
#[allow(dead_code)]
pub tolerance: f64,
}
/// Pure per-user quota formula: `floor((free_disk * tolerance) / max(active, 1))`.
/// Extracted from `compute_storage_quota` so it's unit-testable without a DB or disk.
fn quota_limit_bytes(free_disk: i64, tolerance: f64, active_uploaders: i64) -> i64 {
let active = active_uploaders.max(1);
((free_disk as f64 * tolerance) / active as f64).floor() as i64
}
/// Computes the per-user storage quota using
/// `floor((free_disk * tolerance) / max(active_uploaders, 1))`. Returns `limit_bytes =
/// None` whenever the storage quota is currently disabled — callers should skip the
/// check (upload handler) or hide the UI (quota endpoint).
pub async fn compute_storage_quota(state: &AppState) -> QuotaEstimate {
let quota_on = config::get_bool(&state.config_cache, "quota_enabled", true).await;
let storage_quota_on =
config::get_bool(&state.config_cache, "storage_quota_enabled", true).await;
let tolerance = config::get_f64(&state.config_cache, "quota_tolerance", 0.75).await;
let (active_count,): (i64,) =
sqlx::query_as("SELECT COUNT(DISTINCT user_id) FROM upload WHERE deleted_at IS NULL")
.fetch_one(&state.pool)
.await
.unwrap_or((0,));
let active = active_count.max(1);
// Cached disk reading. `None` means we couldn't resolve the media filesystem.
let disk = state.disk_cache.snapshot(&state.config.media_path);
let free_disk = disk.map(|d| d.free as i64).unwrap_or(0);
let limit_bytes = if quota_on && storage_quota_on {
match disk {
Some(d) => Some(quota_limit_bytes(d.free as i64, tolerance, active)),
// Fail OPEN, not closed: if the disk can't be read we don't know the real
// free space, and enforcing a 0-byte limit would reject every upload with a
// spurious "quota reached". Skip enforcement this round and warn instead.
None => {
tracing::warn!(
"disk snapshot unavailable; skipping storage-quota enforcement this round"
);
None
}
}
} else {
None
};
QuotaEstimate {
limit_bytes,
active_uploaders: active,
free_disk_bytes: free_disk,
tolerance,
}
}
/// Outcome of parsing a `Range` request header against a known file length.
#[derive(Debug, PartialEq, Eq)]
enum RangeSpec {
/// No `Range` header, or one we deliberately don't honour (multi-range, non-`bytes`
/// unit, malformed). RFC 9110 lets a server ignore a Range it can't process and reply
/// 200 with the full body, which is what every one of these cases does.
Full,
/// A single satisfiable range, resolved to inclusive absolute offsets.
Partial { start: u64, end: u64 },
/// Syntactically valid but starts beyond EOF — must be answered 416, not 200, or a
/// player can loop re-requesting it.
Unsatisfiable,
}
/// Parse a single-range `bytes=` header against `len`.
///
/// Deliberately supports only the three forms a media element actually sends —
/// `bytes=N-`, `bytes=N-M`, `bytes=-S` (suffix) — and treats everything else as `Full`.
/// Multi-range responses need `multipart/byteranges`, which no `<video>` requires.
fn parse_range(header: Option<&str>, len: u64) -> RangeSpec {
let Some(raw) = header else {
return RangeSpec::Full;
};
let Some(spec) = raw.trim().strip_prefix("bytes=") else {
return RangeSpec::Full;
};
// Multi-range → fall back to the whole body rather than lie about the content.
if spec.contains(',') {
return RangeSpec::Full;
}
let Some((from, to)) = spec.split_once('-') else {
return RangeSpec::Full;
};
let (from, to) = (from.trim(), to.trim());
// A zero-length file can satisfy no range at all.
if len == 0 {
return if from.is_empty() && to.is_empty() {
RangeSpec::Full
} else {
RangeSpec::Unsatisfiable
};
}
let (start, end) = if from.is_empty() {
// Suffix form: the last `to` bytes.
let Ok(suffix) = to.parse::<u64>() else {
return RangeSpec::Full;
};
if suffix == 0 {
return RangeSpec::Unsatisfiable;
}
(len.saturating_sub(suffix), len - 1)
} else {
let Ok(start) = from.parse::<u64>() else {
return RangeSpec::Full;
};
let end = if to.is_empty() {
len - 1
} else {
match to.parse::<u64>() {
// An end past EOF is clamped, not rejected (RFC 9110 §14.1.1).
Ok(end) => end.min(len - 1),
Err(_) => return RangeSpec::Full,
}
};
(start, end)
};
if start >= len || start > end {
RangeSpec::Unsatisfiable
} else {
RangeSpec::Partial { start, end }
}
}
/// Stream a media file from disk into an HTTP response with a fixed set of security
/// headers. Every media response (original, preview, display, thumbnail) goes through here
/// so they consistently carry `X-Content-Type-Options: nosniff` (defense-in-depth against
/// content-type confusion, even if the edge proxy is bypassed) plus an explicit
/// `Content-Disposition` and `Cache-Control`.
///
/// Honours a single `Range`. This is not an optimisation: iOS Safari opens every `<video>`
/// with a `Range: bytes=0-1` probe and abandons the load unless it gets a `206` with a
/// `Content-Range`. Without this, video is unplayable on the app's primary platform no
/// matter what `src` the element is given. `Accept-Ranges: bytes` is advertised on every
/// response so clients know seeking is available before they ask.
async fn stream_media_file(
req_headers: &axum::http::HeaderMap,
absolute: &std::path::Path,
content_type: String,
disposition: &str,
cache_control: &str,
) -> Result<axum::response::Response, AppError> {
use axum::body::Body;
use axum::http::{Response, StatusCode, header};
use tokio::io::{AsyncReadExt, AsyncSeekExt};
use tokio_util::io::ReaderStream;
if !absolute.exists() {
return Err(AppError::NotFound("Datei nicht gefunden.".into()));
}
let mut file = tokio::fs::File::open(absolute)
.await
.map_err(|e| AppError::Internal(e.into()))?;
let len = file
.metadata()
.await
.map_err(|e| AppError::Internal(e.into()))?
.len();
let range = parse_range(
req_headers.get(header::RANGE).and_then(|v| v.to_str().ok()),
len,
);
let base = |status: StatusCode| {
Response::builder()
.status(status)
.header(header::CONTENT_TYPE, content_type.clone())
.header(header::CONTENT_DISPOSITION, disposition)
.header(header::CACHE_CONTROL, cache_control)
.header(header::X_CONTENT_TYPE_OPTIONS, "nosniff")
.header(header::ACCEPT_RANGES, "bytes")
};
match range {
RangeSpec::Full => base(StatusCode::OK)
.header(header::CONTENT_LENGTH, len)
.body(Body::from_stream(ReaderStream::new(file)))
.map_err(|e| AppError::Internal(e.into())),
RangeSpec::Partial { start, end } => {
file.seek(std::io::SeekFrom::Start(start))
.await
.map_err(|e| AppError::Internal(e.into()))?;
let span = end - start + 1;
base(StatusCode::PARTIAL_CONTENT)
.header(header::CONTENT_LENGTH, span)
.header(header::CONTENT_RANGE, format!("bytes {start}-{end}/{len}"))
.body(Body::from_stream(ReaderStream::new(file.take(span))))
.map_err(|e| AppError::Internal(e.into()))
}
RangeSpec::Unsatisfiable => base(StatusCode::RANGE_NOT_SATISFIABLE)
.header(header::CONTENT_RANGE, format!("bytes */{len}"))
.body(Body::empty())
.map_err(|e| AppError::Internal(e.into())),
}
}
/// Streaming download of the original file behind an upload. Used by:
/// - the per-post "Original anzeigen" context action (`window.open`)
/// - `<img src>` / `<video src>` in the feed, lightbox, and diashow when the user is in
/// Data Mode = Original
///
/// **Auth model:** the route is intentionally unauthenticated so it works from
/// `<img src>` / `window.open`. The URL contains the upload's unguessable UUID. Unlike
/// raw `/media` files, this alias is the *only* way to fetch an original: direct
/// `/media/originals/**` access is blocked in the router, and this handler filters out
/// soft-deleted and ban-hidden uploads (via `find_visible_media`) so moderation actually
/// removes access to content. Preview and thumbnail variants are gated the same way (see
/// [`get_preview`] / [`get_thumbnail`]).
pub async fn get_original(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
.await?
.ok_or_else(|| AppError::NotFound("Upload nicht gefunden.".into()))?;
let absolute = state.config.media_path.join(&media.original_path);
let filename = absolute
.file_name()
.and_then(|n| n.to_str())
.unwrap_or("original");
// `inline`, not `attachment`. This route is the only source of playable video bytes
// (there is no video derivative), and an attachment disposition is hostile to a
// `<video>` element — Safari in particular. It also matches what the UI promises:
// the action is labelled "Original anzeigen", i.e. view, not download.
let disposition = format!("inline; filename=\"{filename}\"");
// Full-res original: never cache at the edge, so a takedown revokes access promptly.
// Range requests still work under no-store; the client simply re-fetches each range.
stream_media_file(
&headers,
&absolute,
media.mime_type,
&disposition,
"no-store",
)
.await
}
/// Streaming access to an upload's compressed **preview** image. Gated exactly like
/// [`get_original`]: `find_visible_media` drops soft-deleted / ban-hidden uploads, so a
/// deleted or moderated post's preview 404s here even though the file may still be on
/// disk. Direct `/media/previews/**` is blocked in the router, making this the only path
/// to a preview.
///
/// Served inline (it's an `<img src>`) and privately cacheable for a short window. The
/// window is intentionally short (5 min) so a moderated image stops being served to a
/// direct-URL holder promptly; the live feed already evicts the card instantly via the
/// `upload-deleted` / `user-hidden` SSE events, so this only bounds the raw-URL edge case.
pub async fn get_preview(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
.await?
.ok_or_else(|| AppError::NotFound("Upload nicht gefunden.".into()))?;
let rel = media
.preview_path
.ok_or_else(|| AppError::NotFound("Vorschau nicht verfügbar.".into()))?;
let absolute = state.config.media_path.join(&rel);
stream_media_file(
&headers,
&absolute,
"image/jpeg".to_string(),
"inline",
"private, max-age=300",
)
.await
}
/// Streaming access to an upload's big-screen **display** derivative (~2048px), used by the
/// diashow. Gated identically to [`get_preview`]. 404s when the derivative doesn't exist yet
/// (still compressing, or an old upload the backfill hasn't reached) — the diashow then falls
/// back to the original.
pub async fn get_display(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
.await?
.ok_or_else(|| AppError::NotFound("Upload nicht gefunden.".into()))?;
let rel = media
.display_path
.ok_or_else(|| AppError::NotFound("Anzeige nicht verfügbar.".into()))?;
let absolute = state.config.media_path.join(&rel);
stream_media_file(
&headers,
&absolute,
"image/jpeg".to_string(),
"inline",
"private, max-age=300",
)
.await
}
/// Streaming access to an upload's **thumbnail** (video poster). Gated identically to
/// [`get_preview`].
pub async fn get_thumbnail(
State(state): State<AppState>,
headers: axum::http::HeaderMap,
Path(upload_id): Path<Uuid>,
) -> Result<axum::response::Response, AppError> {
let media = Upload::find_visible_media(&state.pool, upload_id)
.await?
.ok_or_else(|| AppError::NotFound("Upload nicht gefunden.".into()))?;
let rel = media
.thumbnail_path
.ok_or_else(|| AppError::NotFound("Thumbnail nicht verfügbar.".into()))?;
let absolute = state.config.media_path.join(&rel);
stream_media_file(
&headers,
&absolute,
"image/jpeg".to_string(),
"inline",
"private, max-age=300",
)
.await
}
#[cfg(test)]
mod tests {
use super::{RangeSpec, parse_range, quota_limit_bytes};
// `Range` handling exists because iOS Safari probes every `<video>` with
// `Range: bytes=0-1` and abandons the load without a 206. These pin the forms a
// media element actually sends, plus the edges that decide 206 vs 200 vs 416.
#[test]
fn no_range_header_is_a_full_response() {
assert_eq!(parse_range(None, 100), RangeSpec::Full);
}
#[test]
fn open_ended_range_runs_to_eof() {
assert_eq!(
parse_range(Some("bytes=10-"), 100),
RangeSpec::Partial { start: 10, end: 99 }
);
}
#[test]
fn closed_range_is_inclusive_on_both_ends() {
// The iOS probe. Two bytes, 0 and 1 — an exclusive end would return one.
assert_eq!(
parse_range(Some("bytes=0-1"), 100),
RangeSpec::Partial { start: 0, end: 1 }
);
}
#[test]
fn suffix_range_returns_the_last_n_bytes() {
assert_eq!(
parse_range(Some("bytes=-20"), 100),
RangeSpec::Partial { start: 80, end: 99 }
);
}
#[test]
fn suffix_longer_than_the_file_clamps_to_the_whole_file() {
assert_eq!(
parse_range(Some("bytes=-500"), 100),
RangeSpec::Partial { start: 0, end: 99 }
);
}
#[test]
fn end_past_eof_is_clamped_not_rejected() {
// RFC 9110 §14.1.1 — players routinely ask for more than is there.
assert_eq!(
parse_range(Some("bytes=90-999"), 100),
RangeSpec::Partial { start: 90, end: 99 }
);
}
#[test]
fn start_past_eof_is_416_not_a_full_body() {
// Answering 200 here makes a player re-request forever.
assert_eq!(
parse_range(Some("bytes=100-"), 100),
RangeSpec::Unsatisfiable
);
assert_eq!(
parse_range(Some("bytes=200-300"), 100),
RangeSpec::Unsatisfiable
);
}
#[test]
fn inverted_range_is_unsatisfiable() {
assert_eq!(
parse_range(Some("bytes=50-10"), 100),
RangeSpec::Unsatisfiable
);
}
#[test]
fn unsupported_or_malformed_forms_fall_back_to_the_full_body() {
// Ignoring a Range we can't process and sending 200 is explicitly allowed, and
// safer than guessing. Multi-range would need multipart/byteranges, which no
// <video> asks for.
for header in [
"bytes=0-10,20-30", // multi-range
"items=0-10", // non-bytes unit
"bytes=abc-def", // garbage
"bytes=", // empty spec
"nonsense", // no unit at all
] {
assert_eq!(parse_range(Some(header), 100), RangeSpec::Full, "{header}");
}
}
#[test]
fn last_byte_is_reachable() {
assert_eq!(
parse_range(Some("bytes=99-99"), 100),
RangeSpec::Partial { start: 99, end: 99 }
);
}
#[test]
fn empty_file_satisfies_no_range() {
assert_eq!(parse_range(Some("bytes=0-"), 0), RangeSpec::Unsatisfiable);
assert_eq!(parse_range(None, 0), RangeSpec::Full);
}
#[test]
fn divides_free_space_by_uploaders_with_tolerance() {
// 1000 * 0.75 / 3 = 250
assert_eq!(quota_limit_bytes(1000, 0.75, 3), 250);
}
#[test]
fn floors_fractional_results() {
// 1000 * 0.75 / 7 = 107.14… → 107
assert_eq!(quota_limit_bytes(1000, 0.75, 7), 107);
}
#[test]
fn active_uploaders_below_one_is_clamped_to_one() {
// Guards against divide-by-zero when no one has uploaded yet.
assert_eq!(quota_limit_bytes(1000, 1.0, 0), 1000);
assert_eq!(quota_limit_bytes(1000, 1.0, -5), 1000);
}
#[test]
fn zero_free_disk_yields_zero() {
assert_eq!(quota_limit_bytes(0, 0.75, 3), 0);
}
#[test]
fn full_tolerance_is_identity_for_a_single_uploader() {
assert_eq!(quota_limit_bytes(500, 1.0, 1), 500);
}
/// The guard is the only reclaim mechanism that survives a client disconnect, so its three
/// states have to be exactly right — a wrong `disarm` leaks bytes forever, a wrong `Drop`
/// deletes a committed guest's photo.
mod temp_file_guard {
use super::super::TempFileGuard;
fn scratch(name: &str) -> std::path::PathBuf {
let dir = std::env::temp_dir().join(format!("es-guard-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
let p = dir.join(name);
std::fs::write(&p, b"bytes").unwrap();
p
}
#[test]
fn dropping_an_armed_guard_reclaims_the_file() {
let p = scratch("armed.tmp");
drop(TempFileGuard::new(p.clone()));
assert!(!p.exists(), "an abandoned upload must not survive the request");
}
#[test]
fn a_disarmed_guard_leaves_the_file_alone() {
let p = scratch("committed.jpg");
let mut g = TempFileGuard::new(p.clone());
g.disarm();
drop(g);
assert!(p.exists(), "a committed upload's bytes must never be deleted");
}
#[test]
fn retarget_follows_the_rename_and_forgets_the_old_path() {
let old = scratch("old.tmp");
let new = scratch("new.jpg");
let mut g = TempFileGuard::new(old.clone());
// The rename already moved the bytes; only the new path is at risk now.
std::fs::remove_file(&old).unwrap();
g.retarget(new.clone());
drop(g);
assert!(!new.exists(), "the final-named original is orphaned too until the row commits");
}
#[test]
fn a_guard_whose_file_is_already_gone_is_harmless() {
let p = scratch("vanished.tmp");
std::fs::remove_file(&p).unwrap();
drop(TempFileGuard::new(p)); // must not panic
}
}
}