fix(files): make file writes transactional and close a quota bypass
Audit #6 and #8 for the last of the three stores, plus a bypass found on the way. **#6.** create/update/delete wrote the metadata row, then emitted best-effort, so an outbox failure left a committed file whose trigger never fired. `atomic_write::FilesWriter` commits the metadata row and the fan-out together. Files are the one store where the ordering is subtle, because the BYTES live on disk and cannot join a transaction: * create/update — blob first, then commit metadata + fan-out. A rollback unlinks the blob. (A crash at that exact point still orphans it; that hazard predates this change — the repo already wrote the blob and then inserted the row in a separate, failable statement — and the orphan is inert, referenced by nothing.) * delete — commit the metadata removal + fan-out FIRST, then unlink. The reverse order would destroy the bytes of a row that a rollback keeps, leaving a file that can never be read. **#8.** `GroupFilesService::create` read `total_bytes` on one connection and wrote on another, so concurrent uploads each saw the same pre-write total and together overshot the ceiling. This is the worst instance of the race in the codebase: the ceiling is DISK (10 GiB by default) and one file may be 100 MB, so a racing fleet overshoots by gigabytes. `PostgresGroupFilesWriter` takes the per-group advisory lock (on its own `files` key) across the check and the write. **The bypass.** `GroupFilesService::update` checked NO quota at all — so a 1-byte file could be updated to a 100 MB one without the ceiling ever being consulted, repeatedly, for unbounded disk. It now checks the projected total (the replaced file's bytes subtracted in SQL, so a same-size-or-smaller update near the cap still goes through). Also drive-by: `queue_e2e` asserted the ack the instant the marker appeared, but the marker is written DURING the handler and the ack happens after it returns — a zero-tolerance race. It polls now. (This does not fix the suite's flakiness, which reproduces on the pre-pass commit too.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -108,6 +108,19 @@ const THROW_HANDLER: &str = r#"
|
||||
throw "boom"
|
||||
"#;
|
||||
|
||||
/// Poll until the queue depth reaches `want` (or give up and return what we saw).
|
||||
async fn poll_queue_depth(pool: &PgPool, app_id: &str, queue: &str, want: i64) -> i64 {
|
||||
let mut depth = -1;
|
||||
for _ in 0..200 {
|
||||
depth = count_queue_messages(pool, app_id, queue).await;
|
||||
if depth == want {
|
||||
return depth;
|
||||
}
|
||||
tokio::time::sleep(Duration::from_millis(50)).await;
|
||||
}
|
||||
depth
|
||||
}
|
||||
|
||||
async fn poll_marker(pool: &PgPool, app_id: &str) -> Option<Value> {
|
||||
for _ in 0..200 {
|
||||
let row: Option<(Value,)> = sqlx::query_as(
|
||||
@@ -192,8 +205,10 @@ async fn queue_receive_acks_on_success() {
|
||||
assert_eq!(marker["queue"]["queue_name"], "jobs");
|
||||
assert_eq!(marker["queue"]["message"]["x"], 42);
|
||||
|
||||
// Ack deleted the row.
|
||||
assert_eq!(count_queue_messages(&pool, &app_id, "jobs").await, 0);
|
||||
// Ack deleted the row. The marker is written DURING the handler, but the ack
|
||||
// happens after it returns — so poll rather than assert on the instant the
|
||||
// marker appears (that window is real, and asserting into it is flaky).
|
||||
assert_eq!(poll_queue_depth(&pool, &app_id, "jobs", 0).await, 0);
|
||||
}
|
||||
|
||||
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
|
||||
|
||||
Reference in New Issue
Block a user