fix(files): make file writes transactional and close a quota bypass

Audit #6 and #8 for the last of the three stores, plus a bypass found on
the way.

**#6.** create/update/delete wrote the metadata row, then emitted
best-effort, so an outbox failure left a committed file whose trigger never
fired. `atomic_write::FilesWriter` commits the metadata row and the fan-out
together.

Files are the one store where the ordering is subtle, because the BYTES live
on disk and cannot join a transaction:

  * create/update — blob first, then commit metadata + fan-out. A rollback
    unlinks the blob. (A crash at that exact point still orphans it; that
    hazard predates this change — the repo already wrote the blob and then
    inserted the row in a separate, failable statement — and the orphan is
    inert, referenced by nothing.)
  * delete — commit the metadata removal + fan-out FIRST, then unlink. The
    reverse order would destroy the bytes of a row that a rollback keeps,
    leaving a file that can never be read.

**#8.** `GroupFilesService::create` read `total_bytes` on one connection and
wrote on another, so concurrent uploads each saw the same pre-write total and
together overshot the ceiling. This is the worst instance of the race in the
codebase: the ceiling is DISK (10 GiB by default) and one file may be 100 MB,
so a racing fleet overshoots by gigabytes. `PostgresGroupFilesWriter` takes
the per-group advisory lock (on its own `files` key) across the check and the
write.

**The bypass.** `GroupFilesService::update` checked NO quota at all — so a
1-byte file could be updated to a 100 MB one without the ceiling ever being
consulted, repeatedly, for unbounded disk. It now checks the projected total
(the replaced file's bytes subtracted in SQL, so a same-size-or-smaller
update near the cap still goes through).

Also drive-by: `queue_e2e` asserted the ack the instant the marker appeared,
but the marker is written DURING the handler and the ack happens after it
returns — a zero-tolerance race. It polls now. (This does not fix the
suite's flakiness, which reproduces on the pre-pass commit too.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-07-14 20:09:03 +02:00
parent 58bf0ab3ec
commit 80a0d31cd2
9 changed files with 1109 additions and 155 deletions

View File

@@ -247,12 +247,16 @@ pub async fn build_app(
// Kept for the v1.1.6 orphan sweeper (cleans stale `*.tmp.*` files).
let files_root = files_config.root.clone();
let files_repo = Arc::new(FsFilesRepo::new(pool.clone(), files_config));
let files: Arc<dyn FilesService> = Arc::new(FilesServiceImpl::new(
files_repo.clone(),
authz.clone(),
events.clone(),
files_max_size,
));
let files: Arc<dyn FilesService> = Arc::new(
FilesServiceImpl::new(
files_repo.clone(),
authz.clone(),
events.clone(),
files_max_size,
)
// Metadata row + trigger fan-out in one tx; a rollback unlinks the blob.
.with_atomic_writes(pool.clone(), files_root.clone()),
);
// §11.6 shared group collections (files): group-keyed blobs under
// `<root>/files/groups/<group_id>/...`, resolved from cx.app_id's chain.
// Shares the per-file size cap; fires `shared = true` triggers under the
@@ -265,7 +269,10 @@ pub async fn build_app(
authz.clone(),
files_max_size,
)
.with_events(events.clone()),
// Byte quota + metadata + shared fan-out in one advisory-locked tx: a
// fleet of concurrent uploads can no longer each see the same pre-write
// total and together overshoot the group's disk ceiling.
.with_atomic_writes(pool.clone(), files_root.clone()),
);
// v1.1.6 realtime: the in-process broadcaster is shared between the
// publish path (PubsubServiceImpl fans out to SSE subscribers after

View File

@@ -108,6 +108,19 @@ const THROW_HANDLER: &str = r#"
throw "boom"
"#;
/// Poll until the queue depth reaches `want` (or give up and return what we saw).
async fn poll_queue_depth(pool: &PgPool, app_id: &str, queue: &str, want: i64) -> i64 {
let mut depth = -1;
for _ in 0..200 {
depth = count_queue_messages(pool, app_id, queue).await;
if depth == want {
return depth;
}
tokio::time::sleep(Duration::from_millis(50)).await;
}
depth
}
async fn poll_marker(pool: &PgPool, app_id: &str) -> Option<Value> {
for _ in 0..200 {
let row: Option<(Value,)> = sqlx::query_as(
@@ -192,8 +205,10 @@ async fn queue_receive_acks_on_success() {
assert_eq!(marker["queue"]["queue_name"], "jobs");
assert_eq!(marker["queue"]["message"]["x"], 42);
// Ack deleted the row.
assert_eq!(count_queue_messages(&pool, &app_id, "jobs").await, 0);
// Ack deleted the row. The marker is written DURING the handler, but the ack
// happens after it returns — so poll rather than assert on the instant the
// marker appears (that window is real, and asserting into it is flaky).
assert_eq!(poll_queue_depth(&pool, &app_id, "jobs", 0).await, 0);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]