Audit #6 and #8 for the last of the three stores, plus a bypass found on the way. **#6.** create/update/delete wrote the metadata row, then emitted best-effort, so an outbox failure left a committed file whose trigger never fired. `atomic_write::FilesWriter` commits the metadata row and the fan-out together. Files are the one store where the ordering is subtle, because the BYTES live on disk and cannot join a transaction: * create/update — blob first, then commit metadata + fan-out. A rollback unlinks the blob. (A crash at that exact point still orphans it; that hazard predates this change — the repo already wrote the blob and then inserted the row in a separate, failable statement — and the orphan is inert, referenced by nothing.) * delete — commit the metadata removal + fan-out FIRST, then unlink. The reverse order would destroy the bytes of a row that a rollback keeps, leaving a file that can never be read. **#8.** `GroupFilesService::create` read `total_bytes` on one connection and wrote on another, so concurrent uploads each saw the same pre-write total and together overshot the ceiling. This is the worst instance of the race in the codebase: the ceiling is DISK (10 GiB by default) and one file may be 100 MB, so a racing fleet overshoots by gigabytes. `PostgresGroupFilesWriter` takes the per-group advisory lock (on its own `files` key) across the check and the write. **The bypass.** `GroupFilesService::update` checked NO quota at all — so a 1-byte file could be updated to a 100 MB one without the ceiling ever being consulted, repeatedly, for unbounded disk. It now checks the projected total (the replaced file's bytes subtracted in SQL, so a same-size-or-smaller update near the cap still goes through). Also drive-by: `queue_e2e` asserted the ack the instant the marker appeared, but the marker is written DURING the handler and the ack happens after it returns — a zero-tolerance race. It polls now. (This does not fix the suite's flakiness, which reproduces on the pre-pass commit too.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
560 lines
17 KiB
Rust
560 lines
17 KiB
Rust
//! Low-level metadata (Postgres `group_files`) + blob bytes (filesystem) storage
|
|
//! for §11.6 group-shared FILES collections. A near-clone of [`crate::files_repo`]
|
|
//! keyed by the owning `group_id` instead of `app_id`.
|
|
//!
|
|
//! The security-sensitive disk mechanics — the atomic write+checksum protocol and
|
|
//! checksum-on-read — are **not** duplicated: they come from the owner-relative
|
|
//! free functions in [`crate::files_repo`] (`write_atomic_at` / `read_verify_at` /
|
|
//! `final_path_at`), called with this repo's owner sub-path. Group blobs shard at
|
|
//! `<root>/files/groups/<group_id>/<collection>/<id[0:2]>/<id>` — a `groups/` infix
|
|
//! disjoint from the per-app `files/<app_id>/...` subtree, so the orphan sweeper
|
|
//! ([crate::files_sweep]) covers both with one walk. Authorization, group
|
|
//! resolution, value validation, and content-type sanitization live one layer up
|
|
//! in `GroupFilesServiceImpl`.
|
|
|
|
use std::path::{Path, PathBuf};
|
|
|
|
use async_trait::async_trait;
|
|
use chrono::{DateTime, Utc};
|
|
use picloud_shared::{FileMeta, FileUpdate, FilesListPage, GroupId, NewFile};
|
|
use sqlx::PgPool;
|
|
use uuid::Uuid;
|
|
|
|
use crate::files_repo::{
|
|
decode_cursor, encode_cursor, final_path_at, read_verify_at, write_atomic_at, FileUpdated,
|
|
FilesRepoError,
|
|
};
|
|
|
|
const FILES_LIST_MAX_LIMIT: u32 = 1_000;
|
|
const FILES_LIST_DEFAULT_LIMIT: u32 = 100;
|
|
|
|
#[derive(Debug, thiserror::Error)]
|
|
pub enum GroupFilesRepoError {
|
|
#[error("database error: {0}")]
|
|
Db(#[from] sqlx::Error),
|
|
|
|
#[error("filesystem error: {0}")]
|
|
Io(String),
|
|
|
|
#[error("invalid collection name: {0}")]
|
|
InvalidCollection(String),
|
|
|
|
/// Bytes on disk no longer match the stored checksum (or are missing).
|
|
#[error("file content corrupted (checksum mismatch)")]
|
|
Corrupted,
|
|
|
|
#[error("invalid pagination cursor")]
|
|
InvalidCursor,
|
|
}
|
|
|
|
impl From<FilesRepoError> for GroupFilesRepoError {
|
|
fn from(e: FilesRepoError) -> Self {
|
|
match e {
|
|
FilesRepoError::Db(e) => Self::Db(e),
|
|
FilesRepoError::Io(s) => Self::Io(s),
|
|
FilesRepoError::InvalidCollection(c) => Self::InvalidCollection(c),
|
|
FilesRepoError::Corrupted => Self::Corrupted,
|
|
FilesRepoError::InvalidCursor => Self::InvalidCursor,
|
|
}
|
|
}
|
|
}
|
|
|
|
/// The owner-relative subdirectory of a **group**'s shared blobs under
|
|
/// `<root>/files/`: `groups/<group_id>`. A UUID app-dir can never equal the
|
|
/// literal `groups`, so app and group blob trees can't collide.
|
|
pub(crate) fn group_owner_dir(group_id: GroupId) -> PathBuf {
|
|
Path::new("groups").join(group_id.into_inner().to_string())
|
|
}
|
|
|
|
#[async_trait]
|
|
pub trait GroupFilesRepo: Send + Sync {
|
|
async fn create(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
new: NewFile,
|
|
) -> Result<FileMeta, GroupFilesRepoError>;
|
|
|
|
async fn head(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError>;
|
|
|
|
/// Reads + checksum-verifies the bytes. `Ok(None)` when no row exists;
|
|
/// `Err(Corrupted)` when the row exists but the bytes are missing/mismatched.
|
|
async fn get(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<Vec<u8>>, GroupFilesRepoError>;
|
|
|
|
/// `Ok(None)` when no row exists (the service maps that to `NotFound`).
|
|
async fn update(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
upd: FileUpdate,
|
|
) -> Result<Option<FileUpdated>, GroupFilesRepoError>;
|
|
|
|
async fn delete(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError>;
|
|
|
|
async fn list(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
cursor: Option<&str>,
|
|
limit: u32,
|
|
) -> Result<FilesListPage, GroupFilesRepoError>;
|
|
|
|
/// §11.6 quota: total stored bytes across the group's shared-files
|
|
/// collections. Default `Ok(0)` so non-Postgres impls skip the check.
|
|
async fn total_bytes(&self, group_id: GroupId) -> Result<u64, GroupFilesRepoError> {
|
|
let _ = group_id;
|
|
Ok(0)
|
|
}
|
|
}
|
|
|
|
/// Filesystem-bytes + Postgres-metadata repo for group-shared files.
|
|
pub struct FsGroupFilesRepo {
|
|
pool: PgPool,
|
|
root: PathBuf,
|
|
}
|
|
|
|
impl FsGroupFilesRepo {
|
|
#[must_use]
|
|
pub fn new(pool: PgPool, root: PathBuf) -> Self {
|
|
Self { pool, root }
|
|
}
|
|
|
|
/// Belt-and-suspenders path guard (the service validates at the SDK
|
|
/// boundary). Mirrors `FsFilesRepo::guard_collection`.
|
|
fn guard_collection(collection: &str) -> Result<(), GroupFilesRepoError> {
|
|
if collection.is_empty()
|
|
|| collection.contains('/')
|
|
|| collection.contains('\\')
|
|
|| collection.contains("..")
|
|
|| collection.contains('\0')
|
|
{
|
|
return Err(GroupFilesRepoError::InvalidCollection(
|
|
collection.to_string(),
|
|
));
|
|
}
|
|
Ok(())
|
|
}
|
|
|
|
fn write_atomic(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
bytes: &[u8],
|
|
) -> Result<String, GroupFilesRepoError> {
|
|
Ok(write_atomic_at(
|
|
&self.root,
|
|
&group_owner_dir(group_id),
|
|
collection,
|
|
id,
|
|
bytes,
|
|
)?)
|
|
}
|
|
}
|
|
|
|
#[async_trait]
|
|
impl GroupFilesRepo for FsGroupFilesRepo {
|
|
async fn create(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
new: NewFile,
|
|
) -> Result<FileMeta, GroupFilesRepoError> {
|
|
Self::guard_collection(collection)?;
|
|
let id = Uuid::new_v4();
|
|
let size = i64::try_from(new.data.len()).unwrap_or(i64::MAX);
|
|
let checksum = self.write_atomic(group_id, collection, id, &new.data)?;
|
|
|
|
let row: GroupFileRow = sqlx::query_as(
|
|
"INSERT INTO group_files \
|
|
(group_id, collection, id, name, content_type, size_bytes, checksum_sha256) \
|
|
VALUES ($1, $2, $3, $4, $5, $6, $7) \
|
|
RETURNING id, collection, name, content_type, size_bytes, \
|
|
checksum_sha256, created_at, updated_at",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.bind(&new.name)
|
|
.bind(&new.content_type)
|
|
.bind(size)
|
|
.bind(&checksum)
|
|
.fetch_one(&self.pool)
|
|
.await?;
|
|
Ok(row.into_meta())
|
|
}
|
|
|
|
async fn head(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError> {
|
|
Self::guard_collection(collection)?;
|
|
let row: Option<GroupFileRow> = sqlx::query_as(
|
|
"SELECT id, collection, name, content_type, size_bytes, \
|
|
checksum_sha256, created_at, updated_at \
|
|
FROM group_files WHERE group_id = $1 AND collection = $2 AND id = $3",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.fetch_optional(&self.pool)
|
|
.await?;
|
|
Ok(row.map(GroupFileRow::into_meta))
|
|
}
|
|
|
|
async fn get(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<Vec<u8>>, GroupFilesRepoError> {
|
|
Self::guard_collection(collection)?;
|
|
let row: Option<(String,)> = sqlx::query_as(
|
|
"SELECT checksum_sha256 FROM group_files \
|
|
WHERE group_id = $1 AND collection = $2 AND id = $3",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.fetch_optional(&self.pool)
|
|
.await?;
|
|
let Some((stored_checksum,)) = row else {
|
|
return Ok(None);
|
|
};
|
|
let bytes = read_verify_at(
|
|
&self.root,
|
|
&group_owner_dir(group_id),
|
|
collection,
|
|
id,
|
|
&stored_checksum,
|
|
)?;
|
|
Ok(Some(bytes))
|
|
}
|
|
|
|
async fn update(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
upd: FileUpdate,
|
|
) -> Result<Option<FileUpdated>, GroupFilesRepoError> {
|
|
Self::guard_collection(collection)?;
|
|
let Some(prev) = self.head(group_id, collection, id).await? else {
|
|
return Ok(None);
|
|
};
|
|
let size = i64::try_from(upd.data.len()).unwrap_or(i64::MAX);
|
|
let checksum = self.write_atomic(group_id, collection, id, &upd.data)?;
|
|
|
|
let row: GroupFileRow = sqlx::query_as(
|
|
"UPDATE group_files SET \
|
|
name = COALESCE($4, name), \
|
|
content_type = COALESCE($5, content_type), \
|
|
size_bytes = $6, \
|
|
checksum_sha256 = $7, \
|
|
updated_at = NOW() \
|
|
WHERE group_id = $1 AND collection = $2 AND id = $3 \
|
|
RETURNING id, collection, name, content_type, size_bytes, \
|
|
checksum_sha256, created_at, updated_at",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.bind(upd.name.as_deref())
|
|
.bind(upd.content_type.as_deref())
|
|
.bind(size)
|
|
.bind(&checksum)
|
|
.fetch_one(&self.pool)
|
|
.await?;
|
|
Ok(Some(FileUpdated {
|
|
new: row.into_meta(),
|
|
prev,
|
|
}))
|
|
}
|
|
|
|
async fn delete(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError> {
|
|
Self::guard_collection(collection)?;
|
|
let mut tx = self.pool.begin().await?;
|
|
let row: Option<GroupFileRow> = sqlx::query_as(
|
|
"SELECT id, collection, name, content_type, size_bytes, \
|
|
checksum_sha256, created_at, updated_at \
|
|
FROM group_files WHERE group_id = $1 AND collection = $2 AND id = $3 \
|
|
FOR UPDATE",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.fetch_optional(&mut *tx)
|
|
.await?;
|
|
|
|
let Some(row) = row else {
|
|
tx.rollback().await?;
|
|
return Ok(None);
|
|
};
|
|
|
|
sqlx::query("DELETE FROM group_files WHERE group_id = $1 AND collection = $2 AND id = $3")
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.execute(&mut *tx)
|
|
.await?;
|
|
tx.commit().await?;
|
|
|
|
// Row gone; unlink the bytes. A failure here leaves an orphan file
|
|
// (reclaimed by the sweep) — not fatal.
|
|
let path = final_path_at(&self.root, &group_owner_dir(group_id), collection, id);
|
|
if let Err(e) = std::fs::remove_file(&path) {
|
|
if e.kind() != std::io::ErrorKind::NotFound {
|
|
tracing::warn!(path = %path.display(), error = %e, "group files: unlink after delete failed (orphan)");
|
|
}
|
|
}
|
|
Ok(Some(row.into_meta()))
|
|
}
|
|
|
|
async fn list(
|
|
&self,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
cursor: Option<&str>,
|
|
limit: u32,
|
|
) -> Result<FilesListPage, GroupFilesRepoError> {
|
|
Self::guard_collection(collection)?;
|
|
let limit = if limit == 0 {
|
|
FILES_LIST_DEFAULT_LIMIT
|
|
} else {
|
|
limit.min(FILES_LIST_MAX_LIMIT)
|
|
};
|
|
let last_id = match cursor {
|
|
Some(c) => Some(decode_cursor(c)?),
|
|
None => None,
|
|
};
|
|
let take = i64::from(limit) + 1;
|
|
let rows: Vec<GroupFileRow> = sqlx::query_as(
|
|
"SELECT id, collection, name, content_type, size_bytes, \
|
|
checksum_sha256, created_at, updated_at \
|
|
FROM group_files \
|
|
WHERE group_id = $1 AND collection = $2 \
|
|
AND ($3::uuid IS NULL OR id > $3) \
|
|
ORDER BY id ASC \
|
|
LIMIT $4",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(last_id)
|
|
.bind(take)
|
|
.fetch_all(&self.pool)
|
|
.await?;
|
|
|
|
let mut files: Vec<FileMeta> = rows.into_iter().map(GroupFileRow::into_meta).collect();
|
|
let next_cursor = if files.len() > limit as usize {
|
|
files.truncate(limit as usize);
|
|
files.last().map(|m| encode_cursor(m.id))
|
|
} else {
|
|
None
|
|
};
|
|
Ok(FilesListPage { files, next_cursor })
|
|
}
|
|
|
|
async fn total_bytes(&self, group_id: GroupId) -> Result<u64, GroupFilesRepoError> {
|
|
let (n,): (i64,) = sqlx::query_as(
|
|
"SELECT COALESCE(SUM(size_bytes), 0)::bigint FROM group_files WHERE group_id = $1",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.fetch_one(&self.pool)
|
|
.await?;
|
|
Ok(u64::try_from(n).unwrap_or(0))
|
|
}
|
|
}
|
|
|
|
#[derive(sqlx::FromRow)]
|
|
pub(crate) struct GroupFileRow {
|
|
id: Uuid,
|
|
collection: String,
|
|
name: String,
|
|
content_type: String,
|
|
size_bytes: i64,
|
|
checksum_sha256: String,
|
|
created_at: DateTime<Utc>,
|
|
updated_at: DateTime<Utc>,
|
|
}
|
|
|
|
impl GroupFileRow {
|
|
pub(crate) fn into_meta(self) -> FileMeta {
|
|
FileMeta {
|
|
id: self.id,
|
|
collection: self.collection,
|
|
name: self.name,
|
|
content_type: self.content_type,
|
|
size: u64::try_from(self.size_bytes).unwrap_or(0),
|
|
checksum: self.checksum_sha256,
|
|
created_at: self.created_at,
|
|
updated_at: self.updated_at,
|
|
}
|
|
}
|
|
}
|
|
// ----------------------------------------------------------------------------
|
|
// Connection-scoped metadata operations (see `files_repo` for the rationale)
|
|
// ----------------------------------------------------------------------------
|
|
|
|
use crate::files_repo::{MetaFields, MetaPatch};
|
|
|
|
const GROUP_FILE_COLS: &str = "id, collection, name, content_type, size_bytes, \
|
|
checksum_sha256, created_at, updated_at";
|
|
|
|
pub(crate) async fn head_on<'c, E>(
|
|
exec: E,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError>
|
|
where
|
|
E: sqlx::PgExecutor<'c>,
|
|
{
|
|
let row: Option<GroupFileRow> = sqlx::query_as(&format!(
|
|
"SELECT {GROUP_FILE_COLS} FROM group_files \
|
|
WHERE group_id = $1 AND collection = $2 AND id = $3"
|
|
))
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.fetch_optional(exec)
|
|
.await?;
|
|
Ok(row.map(GroupFileRow::into_meta))
|
|
}
|
|
|
|
pub(crate) async fn insert_meta_on<'c, E>(
|
|
exec: E,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
meta: MetaFields<'_>,
|
|
) -> Result<FileMeta, GroupFilesRepoError>
|
|
where
|
|
E: sqlx::PgExecutor<'c>,
|
|
{
|
|
let row: GroupFileRow = sqlx::query_as(&format!(
|
|
"INSERT INTO group_files \
|
|
(group_id, collection, id, name, content_type, size_bytes, checksum_sha256) \
|
|
VALUES ($1, $2, $3, $4, $5, $6, $7) \
|
|
RETURNING {GROUP_FILE_COLS}"
|
|
))
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.bind(meta.name)
|
|
.bind(meta.content_type)
|
|
.bind(meta.size)
|
|
.bind(meta.checksum)
|
|
.fetch_one(exec)
|
|
.await?;
|
|
Ok(row.into_meta())
|
|
}
|
|
|
|
pub(crate) async fn update_meta_on<'c, E>(
|
|
exec: E,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
meta: MetaPatch<'_>,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError>
|
|
where
|
|
E: sqlx::PgExecutor<'c>,
|
|
{
|
|
let row: Option<GroupFileRow> = sqlx::query_as(&format!(
|
|
"UPDATE group_files SET \
|
|
name = COALESCE($4, name), \
|
|
content_type = COALESCE($5, content_type), \
|
|
size_bytes = $6, \
|
|
checksum_sha256 = $7, \
|
|
updated_at = NOW() \
|
|
WHERE group_id = $1 AND collection = $2 AND id = $3 \
|
|
RETURNING {GROUP_FILE_COLS}"
|
|
))
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.bind(meta.name)
|
|
.bind(meta.content_type)
|
|
.bind(meta.size)
|
|
.bind(meta.checksum)
|
|
.fetch_optional(exec)
|
|
.await?;
|
|
Ok(row.map(GroupFileRow::into_meta))
|
|
}
|
|
|
|
pub(crate) async fn delete_meta_on<'c, E>(
|
|
exec: E,
|
|
group_id: GroupId,
|
|
collection: &str,
|
|
id: Uuid,
|
|
) -> Result<Option<FileMeta>, GroupFilesRepoError>
|
|
where
|
|
E: sqlx::PgExecutor<'c>,
|
|
{
|
|
let row: Option<GroupFileRow> = sqlx::query_as(&format!(
|
|
"DELETE FROM group_files \
|
|
WHERE group_id = $1 AND collection = $2 AND id = $3 \
|
|
RETURNING {GROUP_FILE_COLS}"
|
|
))
|
|
.bind(group_id.into_inner())
|
|
.bind(collection)
|
|
.bind(id)
|
|
.fetch_optional(exec)
|
|
.await?;
|
|
Ok(row.map(GroupFileRow::into_meta))
|
|
}
|
|
|
|
/// §11.6 quota: the PROJECTED total stored bytes for the group AFTER this write
|
|
/// — the current SUM, minus the bytes of the file being replaced (`replacing =
|
|
/// Some(id)` on update; `None` on create), plus the incoming file's bytes.
|
|
///
|
|
/// The subtraction is what `GroupFilesService::update` was missing entirely: it
|
|
/// checked no quota at all, so a 1-byte file could be updated to a 100 MB one
|
|
/// without ever consulting the ceiling.
|
|
pub(crate) async fn projected_total_bytes_on<'c, E>(
|
|
exec: E,
|
|
group_id: GroupId,
|
|
replacing: Option<Uuid>,
|
|
incoming: i64,
|
|
) -> Result<u64, GroupFilesRepoError>
|
|
where
|
|
E: sqlx::PgExecutor<'c>,
|
|
{
|
|
let (n,): (i64,) = sqlx::query_as(
|
|
"SELECT ( \
|
|
COALESCE((SELECT SUM(size_bytes) FROM group_files WHERE group_id = $1), 0) \
|
|
- COALESCE((SELECT size_bytes FROM group_files \
|
|
WHERE group_id = $1 AND id = $2), 0) \
|
|
+ $3 \
|
|
)::BIGINT",
|
|
)
|
|
.bind(group_id.into_inner())
|
|
.bind(replacing)
|
|
.bind(incoming)
|
|
.fetch_one(exec)
|
|
.await?;
|
|
Ok(u64::try_from(n).unwrap_or(0))
|
|
}
|