Reproduced live, by accident, while smoke-testing on a machine with no ffmpeg: the clip uploaded fine, returned 201, and roughly six seconds later had `deleted_at` set and was gone from the feed. The `Ok(None)` "this clip yields no frame" case was already handled — that fix landed when sub-second clips were being destroyed. But the `?` on the call itself still routed every OTHER failure into the same give-up path, which soft-deletes: ffmpeg missing from the image, ffmpeg hanging on a truncated `.mov` and tripping the timeout, an ENOSPC on `thumbnails/`, or a DB blip in `set_thumbnail_path`. None of those says anything about the video, and `get_original` serves the file byte-for-byte, so a post that merely lacks a poster is fully watchable. No failure in the video branch may fail the upload. iPhone `.mov` is exactly the input most likely to trip it, and a wedding clip is not retakeable. ENOSPC gets its own classifier. It was the one failure the retry loop actively made worse: a disk does not drain during six seconds of backoff, so all three attempts failed identically while holding a compression permit that photos were queued behind — and the give-up path then refunded the quota and soft-deleted the row while deliberately KEEPING the original. That freed nothing, removed the photo seconds after a 201, and handed the guest the allowance to upload it again into the same full disk. Now: no retry, no refund, no delete. The row stays live and the photo is served from its original, and `backfill_stale_derivatives` regenerates the derivatives on the next start once there is room. `is_storage_full_error` has to look inside `ImageError::IoError` as well as at bare io errors, because `image` wraps rather than sources it and a plain chain walk would miss every derivative-write failure. FFMPEG_TIMEOUT drops 120s -> 45s. It was never a budget for honest work — a poster from a phone clip takes well under a second, and `-ss` before `-i` means even a 500 MB file seeks rather than scans. It is the ceiling on how long a pathological input holds a permit that guests' photos are waiting behind, so it should be as tight as it can be without cutting off real work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
171 lines
7.4 KiB
Rust
171 lines
7.4 KiB
Rust
//! Poster-frame extraction, shared by the compression worker and the HTML export.
|
|
//!
|
|
//! Both used to spawn `ffmpeg` themselves with the same broken invocation:
|
|
//!
|
|
//! ```text
|
|
//! ffmpeg -i <src> -vframes 1 -ss 00:00:01 -vf scale=… -y <out>
|
|
//! ```
|
|
//!
|
|
//! `-ss` AFTER `-i` is an output-side seek. Against a clip of a second or less ffmpeg exits **0 and
|
|
//! writes nothing** — and both call sites gated on the exit status, so neither noticed. The worker
|
|
//! then wrote `thumbnail_path` for a file that was never created (404 in the live feed) and the
|
|
//! export listed the entry in `data.json` while the ZIP writer skipped it (a broken image tile in
|
|
//! the keepsake). Every server-side signal stayed green. Phones produce such clips constantly:
|
|
//! mis-taps, Live Photos, boomerangs.
|
|
//!
|
|
//! This module exists for the same reason `imaging.rs` does — that one was created when compression
|
|
//! and export duplicated decode logic, and it paid off immediately when the `max_alloc` fix landed
|
|
//! in both workers at once. Same duplication, same fix.
|
|
|
|
use std::path::Path;
|
|
use std::time::Duration;
|
|
|
|
use anyhow::{Context, Result};
|
|
|
|
/// A malformed video can hang `ffmpeg` indefinitely. In the compression worker that never releases
|
|
/// the semaphore permit and the pool eventually deadlocks; in the export worker it strands the job
|
|
/// at `running` so the keepsake never completes. `export.rs` had NO timeout at all before this
|
|
/// module — sharing the spawn fixes that too.
|
|
/// 45s, not the 120s this started at. The timeout is not a budget for honest work — a poster
|
|
/// frame from a phone clip takes well under a second, and `-ss` before `-i` means even a 500 MB
|
|
/// file seeks rather than scans. It is purely the ceiling on how long a pathological input may
|
|
/// hold a compression permit that guests' photos are queued behind, so it should be as tight as
|
|
/// it can be without ever cutting off real work.
|
|
const FFMPEG_TIMEOUT: Duration = Duration::from_secs(45);
|
|
|
|
/// Seek positions to try, in order.
|
|
///
|
|
/// One second first: the opening frame of a real video is often black, a fade-in, or motion-blurred
|
|
/// as the camera settles, so it makes a poor poster. Zero second as the fallback, which is what
|
|
/// makes short clips work — and it is genuinely required, not defensive. Moving `-ss` before `-i`
|
|
/// (an input-side seek) is necessary but NOT sufficient: seeking to 1 s in a 1.000 s clip is still
|
|
/// past the last frame, and ffmpeg still exits 0 having written nothing. Verified against the real
|
|
/// production image.
|
|
const SEEK_POSITIONS: &[&str] = &["00:00:01", "0"];
|
|
|
|
/// Extract one poster frame from `src` into `dest`, scaled to `width` px wide.
|
|
///
|
|
/// `Ok(false)` means the video yielded no frame — a normal outcome for a very short or unusual
|
|
/// clip, NOT an error. Callers must degrade (no poster) rather than fail the upload: treating this
|
|
/// as an error would soft-delete every sub-second video, turning a cosmetic defect into data loss.
|
|
///
|
|
/// `Err` is reserved for something genuinely wrong — a hang we had to kill, or a failure to spawn.
|
|
pub async fn extract_poster_frame(src: &Path, dest: &Path, width: u32) -> Result<bool> {
|
|
for seek in SEEK_POSITIONS {
|
|
// A stale file from a previous attempt would be indistinguishable from a fresh success.
|
|
let _ = tokio::fs::remove_file(dest).await;
|
|
|
|
run_ffmpeg(src, dest, width, seek).await?;
|
|
|
|
// THE CHECK BOTH CALL SITES WERE MISSING: ask the filesystem, not the exit status.
|
|
// Non-empty, because a zero-byte file is not a poster either.
|
|
if tokio::fs::metadata(dest)
|
|
.await
|
|
.map(|m| m.is_file() && m.len() > 0)
|
|
.unwrap_or(false)
|
|
{
|
|
return Ok(true);
|
|
}
|
|
}
|
|
|
|
// Leave nothing behind for a caller to mistake for a result.
|
|
let _ = tokio::fs::remove_file(dest).await;
|
|
Ok(false)
|
|
}
|
|
|
|
/// Run one ffmpeg attempt. A non-zero exit is NOT an error here — the artifact check above is the
|
|
/// authority, and a corrupt input that fails at 1 s may still yield a frame at 0.
|
|
async fn run_ffmpeg(src: &Path, dest: &Path, width: u32, seek: &str) -> Result<()> {
|
|
let mut child = tokio::process::Command::new("ffmpeg")
|
|
.args([
|
|
// BEFORE -i: an input-side seek. See SEEK_POSITIONS.
|
|
"-ss",
|
|
seek,
|
|
"-i",
|
|
src.to_str().unwrap_or_default(),
|
|
"-vframes",
|
|
"1",
|
|
"-vf",
|
|
&format!("scale={width}:-1"),
|
|
"-y",
|
|
dest.to_str().unwrap_or_default(),
|
|
])
|
|
.stdout(std::process::Stdio::piped())
|
|
.stderr(std::process::Stdio::piped())
|
|
.kill_on_drop(true)
|
|
.spawn()
|
|
.context("failed to spawn ffmpeg")?;
|
|
|
|
match tokio::time::timeout(FFMPEG_TIMEOUT, child.wait()).await {
|
|
Ok(res) => {
|
|
res.context("ffmpeg wait failed")?;
|
|
}
|
|
Err(_) => {
|
|
let _ = child.kill().await;
|
|
anyhow::bail!("ffmpeg timed out after {}s", FFMPEG_TIMEOUT.as_secs());
|
|
}
|
|
}
|
|
Ok(())
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
/// Is there a usable `ffmpeg` on PATH?
|
|
///
|
|
/// The poster-frame path shells out, and `extract_poster_frame` documents `Err` as meaning
|
|
/// "a hang or a SPAWN failure" — which is exactly what a missing binary produces. So on a
|
|
/// machine without ffmpeg the test below stops exercising the case it names (missing INPUT)
|
|
/// and instead reports a code defect that isn't there. The runtime image installs ffmpeg
|
|
/// (see backend/Dockerfile), so this only ever skips on a bare developer machine.
|
|
fn ffmpeg_available() -> bool {
|
|
std::process::Command::new("ffmpeg")
|
|
.arg("-version")
|
|
.stdout(std::process::Stdio::null())
|
|
.stderr(std::process::Stdio::null())
|
|
.status()
|
|
.is_ok()
|
|
}
|
|
|
|
/// The order is the whole fix. `-ss` must precede `-i`, and 0 must be tried after 1 s.
|
|
#[test]
|
|
fn the_fallback_seek_exists_and_comes_last() {
|
|
assert_eq!(
|
|
SEEK_POSITIONS,
|
|
&["00:00:01", "0"],
|
|
"1s first for a better poster, 0 as the fallback that makes short clips work"
|
|
);
|
|
}
|
|
|
|
/// A missing input yields no frame rather than an error: the caller must degrade to "no
|
|
/// poster", never fail the upload. `Err` is reserved for a hang or a spawn failure.
|
|
#[tokio::test]
|
|
async fn a_missing_source_yields_no_frame_rather_than_an_error() {
|
|
if !ffmpeg_available() {
|
|
eprintln!(
|
|
"SKIP a_missing_source_yields_no_frame_rather_than_an_error: no ffmpeg on PATH. \
|
|
A missing binary is a spawn failure, which this function returns Err for by \
|
|
design, so the missing-INPUT case cannot be exercised here. Install ffmpeg to \
|
|
run it (the runtime image already has it)."
|
|
);
|
|
return;
|
|
}
|
|
let dir = std::env::temp_dir().join(format!("es-video-{}", std::process::id()));
|
|
std::fs::create_dir_all(&dir).unwrap();
|
|
let dest = dir.join("out.jpg");
|
|
|
|
let got = extract_poster_frame(Path::new("/nonexistent/clip.mp4"), &dest, 400).await;
|
|
|
|
match got {
|
|
Ok(false) => {}
|
|
other => panic!("expected Ok(false) for a missing input, got {other:?}"),
|
|
}
|
|
assert!(
|
|
!dest.exists(),
|
|
"a failed extraction must leave nothing a caller could mistake for a poster"
|
|
);
|
|
let _ = std::fs::remove_dir_all(&dir);
|
|
}
|
|
}
|