Files
PiCloud/crates/shared/src/files.rs
MechaCat02 966b209d2e feat(quota): add per-app storage ceilings
An app's OWN store had no ceiling of any kind — only the group-SHARED
collections did, which is the inversion of what you'd expect. And
`authz::script_gate` returns `Ok(())` when there is no principal, so a
PUBLIC UNAUTHENTICATED route could reach `files::collection(..).create(..)`
and write blobs in a loop until the disk filled. `PICLOUD_FILES_MAX_FILE_SIZE_BYTES`
caps ONE file at 100 MB; nothing capped the count. Same for KV keys and docs
against Postgres.

Ceilings: `PICLOUD_APP_KV_MAX_ROWS` / `PICLOUD_APP_DOCS_MAX_ROWS` (100k) and
`PICLOUD_APP_FILES_MAX_TOTAL_BYTES` (10 GiB).

They are enforced DIFFERENTLY from the group ones, on purpose — porting the
group design as-is would have been a serious regression:

  * KV/docs check a ROW count, on INSERT only, with NO lock. `kv::set` is the
    hottest write path in the system; taking a per-app advisory lock on every
    set would serialize an app's entire data plane, and a `SUM(...)` scan would
    make write cost grow with the app's stored size. What the lock buys is
    small — unlocked overshoot is bounded by write concurrency (32) against a
    ceiling of 100k, ~0.03%. These are anti-DoS rails, not billing. An update
    adds no row and is bounded by the per-value cap, so it pays nothing at all.
  * No per-app BYTE ceiling for KV/docs: every value is already capped at
    `PICLOUD_KV_MAX_VALUE_BYTES`, so `max_rows x max_value_bytes` is ALREADY a
    finite bound. A second scan would buy nothing.
  * FILES are the exception and DO lock + sum: one blob may be 100 MB, so 32
    racing uploads could overshoot by gigabytes of real disk, and a file write
    is heavy enough that the lock and the SUM are lost in the noise. Checked on
    create AND update — an update that skipped the ceiling would be a free
    bypass (the same bug just fixed for group files: grow a 1-byte file to
    100 MB, repeat).

`group_quota` is renamed `quota`: it now serves both owners, and the module
doc is where the two enforcement strategies are contrasted.

Pinned by tests/atomic_write.rs: the key ceiling refuses a new key while
still allowing an update at the ceiling (rejecting updates would brick an app
the moment it filled up — worse than the DoS the rail exists to stop), and 10
concurrent uploads cannot exceed the disk ceiling, nor can an update grow past it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 20:32:32 +02:00

554 lines
19 KiB
Rust

//! `FilesService` — the v1.1.5 filesystem-backed blob store contract.
//!
//! Lives in `picloud-shared` (not `executor-core`) so the Rhai bridge,
//! the manager-core filesystem+Postgres impl, and any in-memory test
//! impl can all depend on the same trait without dragging
//! `executor-core` into a Postgres or filesystem dependency.
//!
//! Implementations MUST derive every storage `app_id` from `cx.app_id`
//! — never from a script-passed argument. That is the cross-app
//! isolation boundary; see `docs/sdk-shape.md`.
//!
//! `FilesService` is collection-scoped: scripts get a handle via
//! `files::collection(name)` and call
//! `create`/`head`/`get`/`update`/`delete`/`list` on it. The blob bytes
//! never travel through Postgres or through trigger payloads — the row
//! is metadata + a SHA-256 checksum; the bytes live on the filesystem.
use async_trait::async_trait;
use chrono::{DateTime, Utc};
use serde::{Deserialize, Serialize};
use thiserror::Error;
use uuid::Uuid;
use crate::SdkCallCx;
/// POSIX-portable filename cap (255 bytes).
pub const MAX_FILE_NAME_BYTES: usize = 255;
/// RFC 6838 puts a reasonable media-type ceiling around 127 chars.
pub const MAX_CONTENT_TYPE_BYTES: usize = 127;
/// Payload for `create` — a brand-new blob. The id is server-generated
/// (a UUID); scripts never supply it.
#[derive(Debug, Clone)]
pub struct NewFile {
pub name: String,
pub content_type: String,
pub data: Vec<u8>,
}
/// Payload for `update` — replacement bytes plus optional metadata. If
/// `name` / `content_type` are `None` the prior values are kept.
#[derive(Debug, Clone)]
pub struct FileUpdate {
pub data: Vec<u8>,
pub name: Option<String>,
pub content_type: Option<String>,
}
/// File metadata as scripts and triggers see it. Serialized into
/// `ServiceEvent.payload` (the blob bytes are NOT included — files are
/// too big to ship through trigger payloads), and surfaced to Rhai by
/// `head` / `list`.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct FileMeta {
pub id: Uuid,
pub collection: String,
pub name: String,
pub content_type: String,
pub size: u64,
/// Lowercase hex SHA-256 of the content.
pub checksum: String,
pub created_at: DateTime<Utc>,
pub updated_at: DateTime<Utc>,
}
/// One page of file metadata from `FilesService::list`. `next_cursor`
/// is `Some` when more pages exist, `None` when exhausted.
#[derive(Debug, Clone)]
pub struct FilesListPage {
pub files: Vec<FileMeta>,
pub next_cursor: Option<String>,
}
#[async_trait]
pub trait FilesService: Send + Sync {
/// Create a new blob; returns its server-generated id. Throws on a
/// missing required field, an over-limit blob, or an invalid
/// collection name.
async fn create(
&self,
cx: &SdkCallCx,
collection: &str,
new: NewFile,
) -> Result<Uuid, FilesError>;
/// Metadata only — no body read. `None` if the file is missing.
async fn head(
&self,
cx: &SdkCallCx,
collection: &str,
id: &str,
) -> Result<Option<FileMeta>, FilesError>;
/// Full content. `None` if missing. Verifies the stored checksum
/// against the bytes on disk and returns `FilesError::Corrupted`
/// when they diverge.
async fn get(
&self,
cx: &SdkCallCx,
collection: &str,
id: &str,
) -> Result<Option<Vec<u8>>, FilesError>;
/// Replace content (and optionally metadata). Throws `NotFound`
/// when the file doesn't exist.
async fn update(
&self,
cx: &SdkCallCx,
collection: &str,
id: &str,
upd: FileUpdate,
) -> Result<(), FilesError>;
/// Delete by id; returns whether the file was present.
async fn delete(&self, cx: &SdkCallCx, collection: &str, id: &str) -> Result<bool, FilesError>;
/// Cursor-paginated metadata listing (same shape as KV's list).
async fn list(
&self,
cx: &SdkCallCx,
collection: &str,
cursor: Option<&str>,
limit: u32,
) -> Result<FilesListPage, FilesError>;
}
/// Failure modes surfaced to the Rhai bridge. The bridge converts each
/// to a Rhai runtime error string; the discriminants exist so internal
/// callers (admin endpoints, tests) can react more precisely.
#[derive(Debug, Error)]
pub enum FilesError {
/// Empty collection name, or one containing a path separator / `..`
/// / NUL — rejected at the SDK boundary per `docs/sdk-shape.md`.
#[error("invalid collection name: {0}")]
InvalidCollection(String),
/// A required field on `create` was missing or empty. The string
/// names the field (`name` / `content_type` / `data`).
#[error("missing required field: {0}")]
MissingField(&'static str),
/// The app's stored blobs would exceed its total-bytes ceiling
/// (`PICLOUD_APP_FILES_MAX_TOTAL_BYTES`). Checked on the projected total, so
/// a same-size-or-smaller update near the cap still goes through.
#[error("files: app storage quota exceeded ({actual} bytes; limit {limit})")]
QuotaExceeded { limit: u64, actual: u64 },
/// Blob exceeds the per-file size cap (default 100 MB,
/// `PICLOUD_FILES_MAX_FILE_SIZE_BYTES`).
#[error("file too large: {size} bytes exceeds limit of {limit} bytes")]
TooLarge { size: usize, limit: usize },
/// Filename exceeds `MAX_FILE_NAME_BYTES`.
#[error("file name too long: {0} bytes exceeds 255")]
NameTooLong(usize),
/// Content-type exceeds `MAX_CONTENT_TYPE_BYTES`.
#[error("content_type too long: {0} bytes exceeds 127")]
ContentTypeTooLong(usize),
/// `update` on a non-existent file.
#[error("file not found")]
NotFound,
/// The bytes on disk no longer match the stored checksum — the
/// filesystem corrupted or a backup was misconfigured. The operator
/// decides what to do with the metadata-vs-bytes mismatch; the repo
/// does NOT auto-delete.
#[error("file content corrupted (checksum mismatch)")]
Corrupted,
/// Caller principal lacked the required capability. Only raised when
/// `cx.principal.is_some()` — scripts running with `principal: None`
/// (public HTTP) operate under script-as-gate semantics and skip
/// the capability check.
#[error("forbidden")]
Forbidden,
/// Anything else — Postgres unavailable, filesystem I/O error, etc.
#[error("files backend error: {0}")]
Backend(String),
}
impl NewFile {
/// Validate required fields + length caps at the SDK boundary.
///
/// Empty `data` is **accepted** as a valid stored state (v1.1.6
/// relaxed the v1.1.5 rejection — empty files are a legitimate use
/// case: sentinels, placeholders, zero-byte uploads. See HANDBACK
/// §7). `name` and `content_type` are still required.
///
/// # Errors
///
/// Returns the field-specific [`FilesError`] for the first failing
/// check.
pub fn validate(&self, max_size: usize) -> Result<(), FilesError> {
if self.name.trim().is_empty() {
return Err(FilesError::MissingField("name"));
}
if self.content_type.trim().is_empty() {
return Err(FilesError::MissingField("content_type"));
}
if self.name.len() > MAX_FILE_NAME_BYTES {
return Err(FilesError::NameTooLong(self.name.len()));
}
if self.content_type.len() > MAX_CONTENT_TYPE_BYTES {
return Err(FilesError::ContentTypeTooLong(self.content_type.len()));
}
if self.data.len() > max_size {
return Err(FilesError::TooLarge {
size: self.data.len(),
limit: max_size,
});
}
Ok(())
}
}
impl FileUpdate {
/// Validate the replacement bytes + any supplied metadata.
///
/// # Errors
///
/// Returns the field-specific [`FilesError`] for the first failing
/// check.
pub fn validate(&self, max_size: usize) -> Result<(), FilesError> {
// Empty replacement bytes are accepted (v1.1.6 relaxation —
// consistent with NewFile::validate; updating a file to zero
// bytes is as legitimate as creating one).
if let Some(name) = &self.name {
if name.trim().is_empty() {
return Err(FilesError::MissingField("name"));
}
if name.len() > MAX_FILE_NAME_BYTES {
return Err(FilesError::NameTooLong(name.len()));
}
}
if let Some(ct) = &self.content_type {
if ct.trim().is_empty() {
return Err(FilesError::MissingField("content_type"));
}
if ct.len() > MAX_CONTENT_TYPE_BYTES {
return Err(FilesError::ContentTypeTooLong(ct.len()));
}
}
if self.data.len() > max_size {
return Err(FilesError::TooLarge {
size: self.data.len(),
limit: max_size,
});
}
Ok(())
}
}
/// Content type used in place of any upload whose declared type would
/// render dangerously in a browser. See [`sanitize_stored_content_type`].
pub const SAFE_RENDER_FALLBACK: &str = "application/octet-stream";
/// Coerce a script-supplied `content_type` to a safe value for storage.
///
/// Audit 2026-06-11 C-2 — storage-side defense. The download response
/// also forces `Content-Disposition: attachment`, `X-Content-Type-Options:
/// nosniff`, and a restrictive CSP (see `manager-core::files_api::get_file`),
/// so this is belt-and-suspenders: even if a future code path serves a
/// file with `inline` again, the stored type can't be `text/html` /
/// `image/svg+xml` / `application/javascript` etc.
///
/// Allowlist (the audit's recommended set):
/// - `application/octet-stream`, `application/pdf`, `application/json`
/// - `text/plain`, `text/csv`
/// - `image/*` (except `image/svg+xml`)
/// - `audio/*`, `video/*`
///
/// Anything else returns `application/octet-stream`. Parameters
/// (`; charset=…`) are preserved when the base type is on the allowlist.
#[must_use]
pub fn sanitize_stored_content_type(ct: &str) -> String {
// Audit 2026-06-11 follow-up — reject any control character first.
// The download handler plumbs the (sanitized) content_type straight
// into a response header; an embedded CR/LF would otherwise survive
// the allowlist (e.g. "image/png\r\nX-Injected: 1" matches the
// `image/` prefix) and then panic `HeaderValue` construction on
// serve. Coercing to the opaque fallback closes the response-header
// injection / handler-panic vector for the stored value.
if ct.bytes().any(|b| b < 0x20 || b == 0x7f) {
return SAFE_RENDER_FALLBACK.to_string();
}
let base = ct
.split(';')
.next()
.unwrap_or("")
.trim()
.to_ascii_lowercase();
let in_allowlist = matches!(
base.as_str(),
"application/octet-stream"
| "application/pdf"
| "application/json"
| "text/plain"
| "text/csv"
) || (base.starts_with("image/") && !base.starts_with("image/svg"))
|| base.starts_with("audio/")
|| base.starts_with("video/");
if in_allowlist {
ct.to_string()
} else {
SAFE_RENDER_FALLBACK.to_string()
}
}
/// Fallback filename when a stored name sanitizes to nothing.
pub const SAFE_FILENAME_FALLBACK: &str = "download";
/// Coerce a stored filename into a value safe to embed in a
/// `Content-Disposition: attachment; filename="…"` header.
///
/// The stored `name` is script/upload-supplied and flows straight into a
/// response header. `HeaderValue` rejects control bytes (`< 0x20`, `0x7f`), so
/// a name containing e.g. a NUL or a raw byte would make the response builder
/// error — and the download handler `.expect()`s that build, turning an
/// attacker-influenced stored name into a per-request panic. We keep only
/// printable ASCII (`0x20..=0x7e`) minus the quoting metacharacters `"` and
/// `\`, replacing anything else with `_`. The result is always a valid
/// `HeaderValue`, so the handler can never panic on it. (Non-ASCII names lose
/// fidelity in this legacy `filename=` param; a future `filename*` UTF-8
/// parameter can restore it — this fix is about never panicking.)
#[must_use]
pub fn sanitize_stored_filename(name: &str) -> String {
// Keep printable ASCII (space included) minus the quoting metacharacters;
// DROP everything else (control bytes, non-ASCII) so the result is always a
// valid HeaderValue. A name that sanitizes to nothing falls back to a
// non-empty default.
let cleaned: String = name
.chars()
.filter(|&c| (c == ' ' || c.is_ascii_graphic()) && c != '"' && c != '\\')
.collect();
let trimmed = cleaned.trim();
if trimmed.is_empty() {
SAFE_FILENAME_FALLBACK.to_string()
} else {
trimmed.to_string()
}
}
/// Reject a collection name that is empty or could escape the per-app
/// files tree. UUID-shaped ids never produce traversal paths, but
/// collection names come from scripts so they're validated defensively
/// at both the SDK boundary and the repo.
///
/// # Errors
///
/// Returns [`FilesError::InvalidCollection`] when the name is empty or
/// contains `/`, `\`, `..`, or a NUL byte.
pub fn validate_collection(collection: &str) -> Result<(), FilesError> {
if collection.is_empty() {
return Err(FilesError::InvalidCollection("must not be empty".into()));
}
if collection.contains('/')
|| collection.contains('\\')
|| collection.contains("..")
|| collection.contains('\0')
{
return Err(FilesError::InvalidCollection(format!(
"collection {collection:?} must not contain '/', '\\', '..', or NUL"
)));
}
Ok(())
}
/// Stub used by the test harness so executor-core integration tests
/// (which don't touch files) can construct a `Services` bundle without
/// a filesystem or Postgres. Every call returns
/// `FilesError::Backend("...")` so accidental use surfaces clearly.
#[derive(Debug, Default, Clone, Copy)]
pub struct NoopFilesService;
#[async_trait]
impl FilesService for NoopFilesService {
async fn create(
&self,
_cx: &SdkCallCx,
_collection: &str,
_new: NewFile,
) -> Result<Uuid, FilesError> {
Err(FilesError::Backend("files is not wired in".into()))
}
async fn head(
&self,
_cx: &SdkCallCx,
_collection: &str,
_id: &str,
) -> Result<Option<FileMeta>, FilesError> {
Err(FilesError::Backend("files is not wired in".into()))
}
async fn get(
&self,
_cx: &SdkCallCx,
_collection: &str,
_id: &str,
) -> Result<Option<Vec<u8>>, FilesError> {
Err(FilesError::Backend("files is not wired in".into()))
}
async fn update(
&self,
_cx: &SdkCallCx,
_collection: &str,
_id: &str,
_upd: FileUpdate,
) -> Result<(), FilesError> {
Err(FilesError::Backend("files is not wired in".into()))
}
async fn delete(
&self,
_cx: &SdkCallCx,
_collection: &str,
_id: &str,
) -> Result<bool, FilesError> {
Err(FilesError::Backend("files is not wired in".into()))
}
async fn list(
&self,
_cx: &SdkCallCx,
_collection: &str,
_cursor: Option<&str>,
_limit: u32,
) -> Result<FilesListPage, FilesError> {
Err(FilesError::Backend("files is not wired in".into()))
}
}
#[cfg(test)]
mod content_type_sanitizer_tests {
use super::*;
#[test]
fn allowlist_passes_safe_types() {
for ct in [
"application/octet-stream",
"application/pdf",
"application/json",
"text/plain",
"text/csv",
"image/png",
"image/jpeg",
"image/webp",
"audio/mpeg",
"video/mp4",
] {
assert_eq!(sanitize_stored_content_type(ct), ct, "{ct} should pass");
}
}
#[test]
fn dangerous_render_types_are_coerced() {
for ct in [
"text/html",
"text/html; charset=utf-8",
"image/svg+xml",
"image/svg",
"application/xhtml+xml",
"application/javascript",
"text/javascript",
"application/ecmascript",
"application/x-shockwave-flash",
"text/xml",
"application/xml",
] {
assert_eq!(
sanitize_stored_content_type(ct),
SAFE_RENDER_FALLBACK,
"{ct} should coerce"
);
}
}
#[test]
fn case_and_whitespace_insensitive() {
assert_eq!(
sanitize_stored_content_type(" TEXT/HTML "),
SAFE_RENDER_FALLBACK
);
assert_eq!(
sanitize_stored_content_type("IMAGE/Svg+XML"),
SAFE_RENDER_FALLBACK
);
}
#[test]
fn parameters_are_preserved_for_safe_types() {
let ct = "text/plain; charset=utf-8";
assert_eq!(sanitize_stored_content_type(ct), ct);
}
#[test]
fn control_chars_are_coerced_even_under_an_allowlisted_prefix() {
// Audit 2026-06-11 follow-up — a CRLF in an otherwise
// allowlisted type must NOT pass through (it would inject a
// response header / panic HeaderValue on download).
for ct in [
"image/png\r\nX-Injected: 1",
"image/png\nX-Injected: 1",
"text/plain\r\nSet-Cookie: x=1",
"audio/mpeg\u{7f}",
"video/mp4\u{0}",
] {
assert_eq!(
sanitize_stored_content_type(ct),
SAFE_RENDER_FALLBACK,
"{ct:?} must coerce"
);
}
}
#[test]
fn stored_filename_is_header_safe() {
// Audit 2026-07-11 — a stored name with a control byte (NUL, 0x7f,
// CR/LF) or quote/backslash must sanitize to a value safe to embed in a
// `Content-Disposition` header, so the download handler can never panic
// building the response. The invariant HeaderValue requires: every byte
// is visible ASCII (0x20..=0x7e) and never `"` or `\` (quoting metachars).
for name in [
"ok.pdf",
"with space.txt",
"quote\".txt",
"back\\slash.txt",
"nul\u{0}byte.bin",
"del\u{7f}.bin",
"crlf\r\n.txt",
"café.pdf",
"\u{0}\u{7f}",
] {
let clean = sanitize_stored_filename(name);
assert!(!clean.is_empty(), "{name:?} produced an empty filename");
assert!(
clean
.bytes()
.all(|b| (0x20..=0x7e).contains(&b) && b != b'"' && b != b'\\'),
"{name:?} -> {clean:?} has a header-unsafe byte"
);
}
// An all-unsafe name falls back to a non-empty default.
assert_eq!(
sanitize_stored_filename("\u{0}\u{7f}\r\n"),
SAFE_FILENAME_FALLBACK
);
}
}