fix(outbox): reclaim stale claims — a crash mid-dispatch stranded events forever

Chasing the e2e flakiness turned up a production durability bug, not a test
bug.

The dispatcher claims an outbox row, executes it, then either deletes it
(success) or reschedules it (failure) — both of which clear the claim. If
the PROCESS DIES in between, neither runs. And `claim_due` only ever selects
`claimed_at IS NULL`. Nothing else in the codebase clears `outbox.claimed_at`
— grep it: there are exactly three writers, and those are two of them.

So a crash or restart mid-dispatch stranded every in-flight row PERMANENTLY.
Its trigger never fired and no retry could notice. The outbox is the
universal trigger path, so the loss covered kv / docs / files / cron /
pubsub / email / invoke_async / dead-letter alike. This is the same
durability class as the audit's #6 (lost trigger event), which was just
fixed at the WRITE end — this is the same hole at the READ end.

Every other claim-based store already had the safety net: `queue_messages`
and `group_queue_messages` have `reclaim_visibility_timeouts`,
`workflow_steps` has its own reclaim. The outbox was the one that didn't.

`OutboxRepo::reclaim_stale_claims(timeout)` returns rows whose claim is older
than `PICLOUD_OUTBOX_CLAIM_TIMEOUT_SEC` (default 600s), run from the
dispatcher's existing reclaim ticker. The default is deliberately generous —
a script may run for up to 300s (the `scripts.timeout_seconds` CHECK), so a
claim held past twice that is abandoned rather than slow; reclaiming a row a
LIVE dispatcher is still working on would double-execute it (survivable —
dispatch is at-least-once — but not worth courting).

A reclaim does NOT bump `attempt_count`: the handler never ran, so it must
not consume the row's retry budget, or repeated restarts would dead-letter an
event that executed zero times. (Same reasoning as the transient queue
`release` fixed earlier in this branch.)

This is also the root cause of the flaky e2e suites: each test drops its
dispatcher, and `claim_due` is not scoped per app or per dispatcher, so a
test's dispatcher could claim another test's row and strand it on teardown.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-07-14 20:38:29 +02:00
parent 966b209d2e
commit c859c9aab2
6 changed files with 236 additions and 0 deletions

View File

@@ -0,0 +1,153 @@
//! A dispatcher that dies mid-dispatch must not strand its claimed outbox rows.
//!
//! The dispatcher claims a row, executes it, then deletes it (success) or
//! reschedules it (failure) — both of which clear the claim. If the PROCESS DIES
//! in between, neither runs. And `claim_due` only ever selects
//! `claimed_at IS NULL`, so before `reclaim_stale_claims` that row was stranded
//! FOREVER: its trigger never fired, and no retry could notice. The outbox is
//! the universal trigger path (kv/docs/files/cron/pubsub/email/invoke_async/
//! dead-letter), so a crash or restart mid-dispatch silently lost all of it.
//!
//! Every other claim-based store — `queue_messages`, `group_queue_messages`,
//! `workflow_steps` — already had this reclaimer. The outbox was the one gap.
//!
//! Skips cleanly when `DATABASE_URL` is unset.
use picloud_manager_core::outbox_repo::{
NewOutboxRow, OutboxRepo, OutboxSourceKind, PostgresOutboxRepo,
};
use picloud_shared::AppId;
use sqlx::postgres::PgPoolOptions;
use sqlx::PgPool;
use uuid::Uuid;
async fn pool_or_skip() -> Option<PgPool> {
let Ok(url) = std::env::var("DATABASE_URL") else {
eprintln!("outbox_reclaim: DATABASE_URL unset — skipping");
return None;
};
let pool = PgPoolOptions::new()
.max_connections(3)
.connect(&url)
.await
.expect("connect");
sqlx::migrate!("./migrations")
.run(&pool)
.await
.expect("migrate");
Some(pool)
}
async fn mk_app(pool: &PgPool) -> Uuid {
let uniq = Uuid::new_v4().simple().to_string();
let (group,): (Uuid,) =
sqlx::query_as("INSERT INTO groups (slug, name) VALUES ($1, $1) RETURNING id")
.bind(format!("or-grp-{uniq}"))
.fetch_one(pool)
.await
.expect("group");
let (app,): (Uuid,) =
sqlx::query_as("INSERT INTO apps (slug, name, group_id) VALUES ($1, $1, $2) RETURNING id")
.bind(format!("or-app-{uniq}"))
.bind(group)
.fetch_one(pool)
.await
.expect("app");
app
}
/// Exactly `claim_due`'s predicate: a row is dispatchable iff its claim is clear.
async fn claimable(pool: &PgPool, id: Uuid) -> bool {
let (n,): (i64,) =
sqlx::query_as("SELECT COUNT(*) FROM outbox WHERE id = $1 AND claimed_at IS NULL")
.bind(id)
.fetch_one(pool)
.await
.expect("claimable");
n == 1
}
async fn attempt_count(pool: &PgPool, id: Uuid) -> i32 {
let (n,): (i32,) = sqlx::query_as("SELECT attempt_count FROM outbox WHERE id = $1")
.bind(id)
.fetch_one(pool)
.await
.expect("attempt_count");
n
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn a_stranded_claim_is_reclaimed_without_burning_the_retry_budget() {
let Some(pool) = pool_or_skip().await else {
return;
};
let repo = PostgresOutboxRepo::new(pool.clone());
let app = mk_app(&pool).await;
let id = repo
.insert(NewOutboxRow {
app_id: AppId::from(app),
source_kind: OutboxSourceKind::Kv,
trigger_id: None,
script_id: None,
reply_to: None,
payload: serde_json::json!({}),
origin_principal: None,
trigger_depth: 0,
root_execution_id: None,
})
.await
.expect("insert");
// A dispatcher claims it. (Asserted against `claim_due`'s own predicate
// rather than by calling it: the dev DB is shared, and a real `claim_due`
// would claim other tests' rows out from under them.)
sqlx::query("UPDATE outbox SET claimed_at = NOW(), claimed_by = 'dispatcher-a' WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("claim");
assert!(
!claimable(&pool, id).await,
"a claimed row is invisible to every other dispatcher — that is the trap"
);
// …and then the process dies. Nothing clears the claim, so before the
// reclaimer existed this row would sit here forever.
sqlx::query("UPDATE outbox SET claimed_at = NOW() - INTERVAL '20 minutes' WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("backdate");
let n = repo.reclaim_stale_claims(600).await.expect("reclaim");
assert!(n >= 1, "the stale claim must be reclaimed");
assert!(
claimable(&pool, id).await,
"the reclaimed row is dispatchable again — the trigger fires after all"
);
assert_eq!(
attempt_count(&pool, id).await,
0,
"the handler never ran, so a reclaim must NOT consume the retry budget — \
otherwise repeated restarts would dead-letter an event that executed zero times"
);
// A FRESH claim must not be reclaimed out from under a LIVE dispatcher.
sqlx::query("UPDATE outbox SET claimed_at = NOW(), claimed_by = 'dispatcher-b' WHERE id = $1")
.bind(id)
.execute(&pool)
.await
.expect("re-claim");
assert_eq!(
repo.reclaim_stale_claims(600).await.expect("reclaim again"),
0,
"a claim within the timeout belongs to a live dispatcher and must be left alone"
);
sqlx::query("DELETE FROM apps WHERE id = $1")
.bind(app)
.execute(&pool)
.await
.expect("cleanup");
}