diff --git a/CLAUDE.md b/CLAUDE.md index 37104c6..117070a 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -10,7 +10,9 @@ Authoritative design: [serverless_cloud_blueprint.md](serverless_cloud_blueprint **v1.1.x — SDK foundation + services — is complete.** The SDK shape (handle pattern, `::` namespaces, `Services`/`SdkCallCx`; see [docs/sdk-shape.md](docs/sdk-shape.md), stdlib at [docs/stdlib-reference.md](docs/stdlib-reference.md)) fixed in v1.1.0, then KV, docs, modules, HTTP, cron, files, pub/sub, email, users, and durable queues + `invoke()` filled it in through **v1.1.9** — blueprint §12 has the table. Earlier groundwork: blueprint Phase 3 (admin auth, multi-app scoping, Phase 3.5 capability gating — `manager-core::authz::{can, require, Capability}`, migration `0006_users_authz.sql`). -**Current focus: v1.2 _Hierarchies_ — groups + the declarative project tool** ([docs/design/groups-and-project-tool.md](docs/design/groups-and-project-tool.md)). That doc's §11 uses its own **Phase 1–6 numbering, distinct from the blueprint product-phase numbering above — do not conflate them** (its "Phase 3" = group-inherited config, not admin auth). Implemented on `feat/groups-*` branches: §11 Phase 1 (declarative `pic plan`/`apply`/`prune` + env overlays), Phase 2 (single-parent groups tree + hierarchy-aware RBAC), Phase 3 (group-inherited, env-scoped `vars` + secrets resolved **live** via a recursive CTE — no materialized cache), Phase 4-lite (group-owned **endpoint** scripts: `scripts` polymorphic owner in `0050_group_scripts.sql`, `get_by_name_inherited`/`is_invocable_by_app` chain resolution, inherited `invoke()` + declarative route/trigger binding — all **live**, no body materialization), Phase 5 (the **declarative project tool maps onto the group tree**: the reconcile engine generalized to `ApplyOwner{App|Group}`, a `[group]` manifest kind, and a single atomic **tree apply** — `pic plan/apply --dir` reconciles a whole directory tree of `picloud.toml` nodes in one Postgres transaction, groups-before-apps so an app route can bind a group script created in the same tx; the bound token folds in each group's `structure_version`. Multi-repo attach-point ceiling + blast-radius and per-env approval gating are deferred; the single-owner ownership **claim** shipped as §7 M1 (see the tail); groups pre-exist), Phase 4b (group **modules** + the **lexical (sealed-by-default) import resolver**, §5.5: owner-polymorphic `ModuleScript`, origin-rooted `ModuleSource::resolve` walking the importing node's chain, `ExecRequest.script_owner` threaded from every dispatch + `invoke()` site, `_source`-driven lexical chaining in `PicloudModuleResolver` with the compiled-module cache re-keyed by `ScriptId`, group modules/imports allowed, single-node dangling-import `plan` check — an inherited group script's imports **seal to the group**, a leaf can't shadow them), §5.5 **extension points** (opt-in polymorphism — **§5.5 now complete**: marker table `0051_extension_points.sql` (owner-polymorphic, CASCADE — structurally a `secrets` name; default body = a co-located `kind=module` script), `ModuleSource::resolve_policy` with **nearest-declaration-kind-wins** — a concrete module resolves lexically, an EP marker resolves **dynamically against the inheriting app** (its override else the default body up-chain), `NoProvider` is a hard error; declarative-only authoring via the `[app]`/`[group]` manifest key `extension_points = [...]`, reconcile mirrors `secrets`, single-node no-provider `plan` check, read-only `pic extension-points ls` + `pull` round-trip — the app can **override** a group default, the deliberate inverse of the Phase 4b sealed import), §11.6 **group-level collections — KV + DOCS + FILES slices** (full cross-app shared read/write: a group declares a collection shared via the `[group]` manifest `collections = [...]` → owner-polymorphic marker `0052_group_collections.sql` with a `kind` discriminator + a per-kind group-keyed store: `0053_group_kv_entries.sql` (`kind='kv'`), `0054_group_docs.sql` (`kind='docs'`, the queryable-JSON store), and `0055_group_files.sql` (`kind='files'`, blob metadata in Postgres + bytes on disk under `/files/groups//...`, a `groups/` infix disjoint from the per-app `files//` subtree so the existing recursive orphan sweeper covers both with zero change) — no `app_id`, a shared row belongs to the group; CASCADE on group delete, an app delete leaves the data. Scripts use the **explicit** `kv::shared_collection("name")` / `docs::shared_collection("name")` / `files::shared_collection("name")` handles (`shared` alone is a Rhai reserved word); `GroupKv`/`GroupDocs`/`GroupFilesServiceImpl` resolve the owning group from `cx.app_id`'s ancestor chain **filtered by kind** (nearest-wins) — **that walk is the isolation boundary**, a foreign app gets `CollectionNotShared`; a `kv`, a `docs`, and a `files` collection of the same name are distinct stores. The docs slice reuses the `docs_filter` DSL — `build_find_query` generalized on its owner column (`docs`/`app_id` vs `group_docs`/`group_id`, both literals); the files slice likewise generalized the atomic-write + checksum-on-read path helpers on an owner-relative dir (one source for the security-sensitive disk mechanics). **Reads open** to any subtree script (anonymous incl. — the declaration is the grant), **writes require an authenticated editor+** on the owning group (`GroupKvRead/Write`, `GroupDocsRead/Write`, `GroupFilesRead/Write`, `script_gate_require_principal` fails closed on anon). Declarative authoring is the **string-or-table** form `collections = ["catalog", { name = "articles", kind = "docs" }, { name = "assets", kind = "files" }]` (bare string = kv); reconcile keys markers by `(name, kind)`; read-only `pic collections ls --group` shows a kind column. Topic shared collections shipped as D2 (storeless), queue as D3 — see below), and **§4.5 group TRIGGER templates** (live, event kinds — a `[group]` declares a `[[triggers.kv|docs|files|pubsub]]` template binding a group-owned handler; `triggers` gained a polymorphic owner `0056_group_triggers.sql` mirroring `0050`; the dispatcher's `list_matching_kv/docs/files` + the pubsub publish fan-out prepend `CHAIN_LEVELS_CTE` + `JOIN chain c ON (t.app_id = c.app_owner OR t.group_id = c.group_owner)` so a descendant app's event matches its own triggers **plus** ancestor-group templates in one query, the handler running under the firing `app_id` — **the chain walk is the isolation boundary**, a sibling-subtree app never matches; stateful kinds cron/queue/email need materialization — see M5 below; per-app opt-out deferred; read-only `pic triggers ls --group`), and **§4.5 group ROUTE templates** (live, inherited — a `[group]` declares a `[[routes]]` template binding a group-owned endpoint; `routes` gained a polymorphic owner `0057_group_routes.sql` mirroring `0056`. Unlike triggers (per-event SQL), routes serve from the in-memory `RouteTable`, so the HTTP hot path can't resolve inheritance per request — instead the table **rebuild** expands templates into each descendant app's slice via `RouteRepository::list_effective` (all-apps generalization of `CHAIN_LEVELS_CTE`: every app × its ancestor chain ⋈ routes), and `compile_effective_routes` applies **nearest-owner-wins shadowing** (an app's own identical binding shadows the inherited template; non-identical bindings coexist under the matcher's existing precedence — a route picks one winner, unlike a fanning trigger). Because the table is a cache, inheritance is rebuilt **full-live** through the single `rebuild_route_table` chokepoint on every edge that changes it: route CRUD, apply, **and tree mutations** — app create/delete (`apps_api`) + group reparent (`groups_api`) — so a new app under a group serves its templates instantly. Host-claim validation is skipped for a group template (descendants serve it on their own host claim; templates use `host_kind = any`). **The chain expansion is the isolation boundary** — a sibling-subtree app never inherits (pinned by `tests/group_route_templates.rs` + the `group_routes` journey); read-only `pic routes ls --group`. Deferred: multi-node snapshot propagation), and **§11 tail per-app opt-out (template suppression)** (a descendant declines an inherited group template: an `[app]` declares `[suppress]` with `triggers = [...]` (handler script names) + `routes = [...]` (paths) — **coarse by reference**, not a full definition, since template row-ids churn on re-apply but a reference is stable (re-apply NoOp, may decline several templates bound to the same script/path). App-only marker `0058_template_suppressions.sql` (`app_id NOT NULL` CASCADE, a `target_kind` discriminator), reconciled with the extension-point marker pattern (prunable → re-inherits). Consumed at the two resolution points: the trigger dispatch queries gain a correlated `NOT EXISTS` anti-join (gated to `t.group_id IS NOT NULL`), and `compile_effective_routes` drops an inherited (`depth > 0`) route at a suppressed path (loaded via `RouteRepository::list_route_suppressions`). **Inheritance-only** — the `group_id IS NOT NULL` / `depth > 0` gates mean an app can only decline what it inherits, never its own or a sibling's (pinned by `tests/template_suppression.rs` + the `suppress` journey); a dangling suppress is an apply-time warning; read-only `pic suppress ls --app`. **Trust-model consequence:** group templates are advisory-by-default (run *unless* a descendant declines) — a footgun for compliance hooks (audit/security triggers a tenant can opt out of)), and **§11 tail `sealed` (mandatory) group templates** (closes that footgun: a `[group]` marks a route/event-trigger template `sealed = true` and the two suppression filters skip it, so a descendant's `[suppress]` is ignored — it fires/serves on every descendant. Column `sealed BOOLEAN` on `triggers` + `routes` (`0059_sealed_templates.sql`); the trigger anti-join gains `AND t.sealed = FALSE` (a sealed row is never excluded → fires through), and `compile_effective_routes` gates its suppression `continue` on `!er.route.sealed`. `sealed` lives on the shared `Route` DTO + manager-core `Trigger` (both pure data) so the apply diff sees the current value — part of the route Update comparison + the trigger identity, so toggling it re-applies. Authored per-template (`sealed = true` on a `[[routes]]`/`[[triggers.kv|docs|files|pubsub]]`), **group-only** — `validate_bundle_for` rejects it on an app owner (an app resource is never inherited). Sealing only *strengthens* the guarantee (a sealed template can't be declined; it never grants new reach — the chain walk is still the isolation boundary). The dangling-suppress warning also flags a suppression matching only sealed templates ("… is sealed — the suppression has no effect"), and `pic triggers/routes ls --group` show a `sealed` column; pinned by `tests/sealed_templates.rs` + the `sealed` journey. Deferred: multi-node snapshot propagation), and **§11 tail M1 group-level suppression** (`template_suppressions` gained a polymorphic owner `0060_group_suppressions.sql` — a `[group]` declares `[suppress]` to decline a template it inherits from a higher ancestor for its **whole subtree**. Both filters generalized to the chain: the trigger anti-join joins the `chain` CTE (`ts.app_id = sc.app_owner OR ts.group_id = sc.group_owner`), and `list_route_suppressions` expands group suppressions across descendants via the all-apps `app_chain` CTE (`compile_effective_routes` unchanged). Still inheritance-only + `sealed` overrides; owner-polymorphic `suppression_repo::list_for_owner`/`insert`/`delete`; read-only `pic suppress ls --group`; the ineffective-suppress warning walks `GROUP_CHAIN_LEVELS_CTE` for a group node. Pinned by `tests/group_suppression.rs` + the `suppress` journey), and **§11.6 M2 shared-collection triggers** (a write to a group SHARED collection now fires a trigger — closes the "group trigger has no app to watch" gap. A group-owned trigger marked `shared = true` (`shared BOOLEAN` on `triggers`, `0061_shared_triggers.sql`) watches the group's shared collection; the per-app `list_matching_kv/docs/files` add `AND t.shared = FALSE`, new `list_matching_shared_{kv,docs,files}(owning_group,…)` select `shared = TRUE` triggers on the owning group — the `shared` flag is the namespace boundary. `ServiceEventEmitter` gained `emit_shared(cx, owning_group, event)`; `GroupKv/Docs/FilesServiceImpl` (via a `with_events` builder) emit on write, and the outbox emitter stamps the WRITER app_id so the handler runs under the writer. `shared` is authored on a group `[[triggers.kv|docs|files]]`, is part of the diff identity, and `validate_bundle_for` rejects it on an app owner / a non-collection kind / an undeclared collection. Read-only `shared` column in `pic triggers ls --group`; pinned by `tests/shared_triggers.rs` + the `shared_triggers` journey. Shared pubsub triggers shipped as D2 — see below), and **§11.6 M3 per-group quotas** (global env-var ceilings enforced in the group write path: `PICLOUD_GROUP_KV_MAX_ROWS`/`_DOCS_MAX_ROWS` (per-group row count, new-key only) + `_FILES_MAX_TOTAL_BYTES` (per-group total bytes); `group_quota` helpers + `count_rows`/`total_bytes` repo methods + `QuotaExceeded` errors), and **§11.6 M4 read-only operator admin API** (`group_blobs_api` mirrors the per-app kv/files admin surface for a group's shared collections: `GET /groups/{id}/{kv,docs,files}[/{collection}/{key|id}]`, authz `GroupKvRead`/`GroupDocsRead`/`GroupFilesRead`; the host hoists the group repos to share one instance; `pic kv ls/get --group`; reads-only, writes stay script-only), and **§4.5 M5 stateful group trigger templates via materialization** (cron + queue + email — the kinds that can't resolve live because each needs a per-app row: cron `last_fired_at`, the queue one-consumer advisory lock, the email sealed inbound secret. A group `[[triggers.cron|queue|email]]` template is **materialized** into an app-owned copy per descendant (`materialized_from` col, `0062_materialized_triggers.sql`) by `materialize::rematerialize_stateful_templates` — all-apps `app_chain` CTE ⋈ group-owned stateful templates, a precise create/delete diff preserving cron state — run full-live at the route-rebuild chokepoints (apply single+tree, app create/delete in `apps_api`, group reparent in `groups_api`; `AppsState`/`GroupsState` gained a `pool`). The scheduler + `list_active_queue_consumers` + `email_inbound_target` gained `AND t.app_id IS NOT NULL` so a group TEMPLATE is never dispatched directly (nor invocable via its own webhook URL) — only its per-app copies are; a queue copy is skipped-with-warning when the app already fills that queue's slot. **M5.5 email uses the shared-group-secret model:** the template resolves its `inbound_secret_ref` against the **group's own** secret store once at apply and seals it (`resolve_and_seal`/`insert_email_trigger_tx` generalized to a `SecretOwner`/`ScriptOwner`); email secrets are v0/no-AAD so the ciphertext is portable across rows — materialization copies the sealed bytes **verbatim** (no master key, no reseal, no new schema column, no threading into the CRUD hooks). All descendant webhooks share the one group HMAC secret; an unset group secret fails apply hard. Loading a group's own set-secret names into `CurrentState` (was hardcoded empty) lets the plan-time email-secret check resolve against the group. Pinned by `tests/stateful_templates.rs` + the `stateful_templates` journey. **Deferred:** the per-app (non-shared) email secret model), and **D1 the `materialized` column** (`pic triggers ls --app` now shows a read-only `materialized` column — a copy of an M5 group stateful template reads `true`, distinct from a hand-authored trigger; a derived `materialized` bool = `materialized_from IS NOT NULL` threaded row→domain→API→CLI mirroring `sealed`/`shared`), and **§11.6 D2 shared TOPICS + shared pub/sub triggers** (a group declares a storeless `topic` shared collection (`group_collections.kind` widened to `topic`+`queue` in `0064`); a `[[triggers.pubsub]] shared = true` handler watches it. `BundleTrigger::Pubsub` gained `shared` (part of the identity); `validate_bundle_for` requires the topic pattern's ROOT segment (`events.*` → `events`) be a declared `kind='topic'` collection, group-only. Scripts publish via the explicit `pubsub::shared_topic("events").publish("created", msg)` handle → `GroupPubsubServiceImpl` resolves the owning group (kind `topic`) from `cx.app_id`'s chain, requires editor+ (`GroupPubsubPublish`, fails closed on anon), and fans out via `PubsubRepo::fan_out_shared_publish` to `shared = true` pubsub triggers on that group, each outbox row stamped the WRITER app_id (M2 model). The per-app `fan_out_publish` gained `AND t.shared = FALSE` — a shared trigger never fires on a per-app publish and vice-versa (the `shared` flag is the namespace boundary); the owning-group chain walk is the isolation boundary. Pinned by `tests/shared_topics.rs` + the `shared_topics` journey. **Deferred:** shared-topic external SSE subscription), and **§11.6 D3 shared durable QUEUES** (a group declares a `queue` shared collection; any subtree app enqueues into ONE group-keyed store (`group_queue_messages`, `0065`, mirrors `queue_messages` keyed by `(group_id, collection)`; CASCADE on group delete) via `queue::shared_collection("name").enqueue(...)` → `GroupQueueServiceImpl` resolves the owning group (kind `queue`), requires editor+ (`GroupQueueEnqueue`, fails closed on anon). **Consumption is by COMPETING CONSUMERS:** a group `[[triggers.queue]] shared = true` consumer (shared threaded through `BundleTrigger::Queue` + identity) **materializes** a consumer copy per descendant app (`materialize` skips the M5 one-consumer-slot check for shared — each descendant intentionally gets a consumer), and all copies claim the SHARED store with `FOR UPDATE SKIP LOCKED` — each message delivered at-most-once across the subtree, scaling horizontally, each handler under its own `cx.app_id`. The dispatcher's queue arm gained a `shared_group` on `ActiveQueueConsumer` (recovered via a LEFT JOIN to the materialized copy's source template) and routes claim/ack/nack/terminal to the group store when set (`q_claim/q_ack/q_nack/q_terminal` helpers; a group claim normalizes to a `ClaimedMessage` under the consuming app); the reclaim task drains both stores. `validate_bundle_for` requires a shared queue on a group to name a declared `kind='queue'` collection; a shared queue on an app is rejected by the app-owner shared guard. Pinned by `tests/group_queue.rs` (competing-consumer at-most-once) + `stateful_templates.rs` + the `shared_queues` journey. **Deferred:** a group dead-letter store (an exhausted shared-queue message is dropped-with-warning, never silently lost)), and **§7 multi-repo ownership M1 — the single-owner claim** (the `owner_project` seam (0047) is now live behind a first-class `projects` table (`0066_projects.sql`, UUID pk + unique slug; `owner_project` FKs it `ON DELETE SET NULL` — un-claim, never cascade-destroy a tree). A `[project]` block (slug + optional name) in the repo's ROOT manifest declares identity (independent of the `[app]`/`[group]` XOR); the first apply with a new slug registers the project and **claims** each group node it touches. The claim runs inside the apply tx under the per-node advisory lock **before the diff**, so a conflict short-circuits with a **409** before any write. Pure policy `decide_group_claim`/`decide_app_owner` in `apply_service` (unit-tested, DB-free): unclaimed→claim, owner→no-op, foreign→conflict unless `--takeover` (which additionally requires `GroupAdmin` — ownership ⟂ RBAC, mapped to 403 vs the 409 conflict), no-project-into-a-claimed-subtree→conflict. **Apps carry no owner** — an app inherits ownership from its **nearest claimed ancestor group** (the ancestor walk, via `groups.ancestors` now carrying `owner_project` through its recursive CTE, is the isolation boundary); an unclaimed subtree stays open, so nothing changes until a repo first declares `[project]` (backward-compatible). `ProjectRepository` (read side) + tx free-fns `upsert_project_tx`/`read_group_owner_tx`/`write_group_owner_tx`; the claim deliberately does **not** bump `structure_version` (not a diff change → won't churn a pending bound plan). Wire: `project`/`takeover` on the apply request (both `#[serde(default)]` → the pre-M1 CLI stays compatible); the CLI surfaces the server's 409 message verbatim (covers `StateMoved` + `OwnershipConflict`). Visibility: `pic groups ls` `owner` column (server `list_with_owner` LEFT JOIN; the shared `Group` deserialize ignores the extra field so `pic groups tree`/dashboard are unaffected) + `pic apply --takeover`. Pinned by `apply_service` unit tests, `tests/projects_repo.rs`, and the `apply_ownership` journey. **M2 shipped — the attach-point ceiling:** a `[project] parent_group = ""` binds the repo under a pre-existing group; `check_within_attach` refuses (422 `OutsideAttachPoint`) any node not strictly within that subtree (a group node must be a *proper* descendant — you can't apply the attach point itself; an app node's group must be at-or-below it), resolved via `groups.ancestors` and enforced read-only before the claim in `apply_owner`/`apply_tree`; absent = instance root = no ceiling. **M3 shipped — plan preview + `pic projects ls`:** `pic plan` now carries the `[project]` and returns an `ownership` preview per node (`claim`/`owned`/`conflict`-owner-named/`unclaimed`, pure `preview_ownership`) plus, for a group node, the cross-repo **blast radius** (descendant apps owned by OTHER projects the change reaches, via `group_blast_radius` — a subtree CTE with per-group memoized nearest-claimed resolution); the attach ceiling is previewed at plan too. Read-only `pic projects ls` (`GET /api/v1/admin/projects`, `list_with_counts`) lists projects + owned-group counts. Pinned by the `apply_ownership` plan-preview case. **With M3 the §7 multi-repo ownership track (M1 claim · M2 attach ceiling · M3 preview + `pic projects ls`) is COMPLETE. Deferred (separate later work):** §6 structural-divergence detection + declarative group create/reparent (lifting "groups pre-exist"); per-env approval gating (`[project.environments]`)). Next: multi-node cluster mode. +**Current focus: v1.2 _Hierarchies_ — groups + the declarative project tool** ([docs/design/groups-and-project-tool.md](docs/design/groups-and-project-tool.md)). That doc's §11 uses its own **Phase 1–6 numbering, distinct from the blueprint product-phase numbering above — do not conflate them** (its "Phase 3" = group-inherited config, not admin auth). Implemented on `feat/groups-*` branches: §11 Phase 1 (declarative `pic plan`/`apply`/`prune` + env overlays), Phase 2 (single-parent groups tree + hierarchy-aware RBAC), Phase 3 (group-inherited, env-scoped `vars` + secrets resolved **live** via a recursive CTE — no materialized cache), Phase 4-lite (group-owned **endpoint** scripts: `scripts` polymorphic owner in `0050_group_scripts.sql`, `get_by_name_inherited`/`is_invocable_by_app` chain resolution, inherited `invoke()` + declarative route/trigger binding — all **live**, no body materialization), Phase 5 (the **declarative project tool maps onto the group tree**: the reconcile engine generalized to `ApplyOwner{App|Group}`, a `[group]` manifest kind, and a single atomic **tree apply** — `pic plan/apply --dir` reconciles a whole directory tree of `picloud.toml` nodes in one Postgres transaction, groups-before-apps so an app route can bind a group script created in the same tx; the bound token folds in each group's `structure_version`. The single-owner ownership **claim** shipped as §7 M1, the attach-point ceiling + blast-radius as §7 M2/M3, and per-env approval gating as §3 M3 — all server-authoritative (see the tail); declarative group **create/reparent** + structural-divergence detection shipped as §6 (`reconcile_group_structure_tx`/`reparent_group_tx`/`StructureMode`), so groups no longer need to pre-exist), Phase 4b (group **modules** + the **lexical (sealed-by-default) import resolver**, §5.5: owner-polymorphic `ModuleScript`, origin-rooted `ModuleSource::resolve` walking the importing node's chain, `ExecRequest.script_owner` threaded from every dispatch + `invoke()` site, `_source`-driven lexical chaining in `PicloudModuleResolver` with the compiled-module cache re-keyed by `ScriptId`, group modules/imports allowed, single-node dangling-import `plan` check — an inherited group script's imports **seal to the group**, a leaf can't shadow them), §5.5 **extension points** (opt-in polymorphism — **§5.5 now complete**: marker table `0051_extension_points.sql` (owner-polymorphic, CASCADE — structurally a `secrets` name; default body = a co-located `kind=module` script), `ModuleSource::resolve_policy` with **nearest-declaration-kind-wins** — a concrete module resolves lexically, an EP marker resolves **dynamically against the inheriting app** (its override else the default body up-chain), `NoProvider` is a hard error; declarative-only authoring via the `[app]`/`[group]` manifest key `extension_points = [...]`, reconcile mirrors `secrets`, single-node no-provider `plan` check, read-only `pic extension-points ls` + `pull` round-trip — the app can **override** a group default, the deliberate inverse of the Phase 4b sealed import), §11.6 **group-level collections — KV + DOCS + FILES slices** (full cross-app shared read/write: a group declares a collection shared via the `[group]` manifest `collections = [...]` → owner-polymorphic marker `0052_group_collections.sql` with a `kind` discriminator + a per-kind group-keyed store: `0053_group_kv_entries.sql` (`kind='kv'`), `0054_group_docs.sql` (`kind='docs'`, the queryable-JSON store), and `0055_group_files.sql` (`kind='files'`, blob metadata in Postgres + bytes on disk under `/files/groups//...`, a `groups/` infix disjoint from the per-app `files//` subtree so the existing recursive orphan sweeper covers both with zero change) — no `app_id`, a shared row belongs to the group; CASCADE on group delete, an app delete leaves the data. Scripts use the **explicit** `kv::shared_collection("name")` / `docs::shared_collection("name")` / `files::shared_collection("name")` handles (`shared` alone is a Rhai reserved word); `GroupKv`/`GroupDocs`/`GroupFilesServiceImpl` resolve the owning group from `cx.app_id`'s ancestor chain **filtered by kind** (nearest-wins) — **that walk is the isolation boundary**, a foreign app gets `CollectionNotShared`; a `kv`, a `docs`, and a `files` collection of the same name are distinct stores. The docs slice reuses the `docs_filter` DSL — `build_find_query` generalized on its owner column (`docs`/`app_id` vs `group_docs`/`group_id`, both literals); the files slice likewise generalized the atomic-write + checksum-on-read path helpers on an owner-relative dir (one source for the security-sensitive disk mechanics). **Reads open** to any subtree script (anonymous incl. — the declaration is the grant), **writes require an authenticated editor+** on the owning group (`GroupKvRead/Write`, `GroupDocsRead/Write`, `GroupFilesRead/Write`, `script_gate_require_principal` fails closed on anon). Declarative authoring is the **string-or-table** form `collections = ["catalog", { name = "articles", kind = "docs" }, { name = "assets", kind = "files" }]` (bare string = kv); reconcile keys markers by `(name, kind)`; read-only `pic collections ls --group` shows a kind column. Topic shared collections shipped as D2 (storeless), queue as D3 — see below), and **§4.5 group TRIGGER templates** (live, event kinds — a `[group]` declares a `[[triggers.kv|docs|files|pubsub]]` template binding a group-owned handler; `triggers` gained a polymorphic owner `0056_group_triggers.sql` mirroring `0050`; the dispatcher's `list_matching_kv/docs/files` + the pubsub publish fan-out prepend `CHAIN_LEVELS_CTE` + `JOIN chain c ON (t.app_id = c.app_owner OR t.group_id = c.group_owner)` so a descendant app's event matches its own triggers **plus** ancestor-group templates in one query, the handler running under the firing `app_id` — **the chain walk is the isolation boundary**, a sibling-subtree app never matches; stateful kinds cron/queue/email need materialization — see M5 below; per-app opt-out deferred; read-only `pic triggers ls --group`), and **§4.5 group ROUTE templates** (live, inherited — a `[group]` declares a `[[routes]]` template binding a group-owned endpoint; `routes` gained a polymorphic owner `0057_group_routes.sql` mirroring `0056`. Unlike triggers (per-event SQL), routes serve from the in-memory `RouteTable`, so the HTTP hot path can't resolve inheritance per request — instead the table **rebuild** expands templates into each descendant app's slice via `RouteRepository::list_effective` (all-apps generalization of `CHAIN_LEVELS_CTE`: every app × its ancestor chain ⋈ routes), and `compile_effective_routes` applies **nearest-owner-wins shadowing** (an app's own identical binding shadows the inherited template; non-identical bindings coexist under the matcher's existing precedence — a route picks one winner, unlike a fanning trigger). Because the table is a cache, inheritance is rebuilt **full-live** through the single `rebuild_route_table` chokepoint on every edge that changes it: route CRUD, apply, **and tree mutations** — app create/delete (`apps_api`) + group reparent (`groups_api`) — so a new app under a group serves its templates instantly. Host-claim validation is skipped for a group template (descendants serve it on their own host claim; templates use `host_kind = any`). **The chain expansion is the isolation boundary** — a sibling-subtree app never inherits (pinned by `tests/group_route_templates.rs` + the `group_routes` journey); read-only `pic routes ls --group`. Deferred: multi-node snapshot propagation), and **§11 tail per-app opt-out (template suppression)** (a descendant declines an inherited group template: an `[app]` declares `[suppress]` with `triggers = [...]` (handler script names) + `routes = [...]` (paths) — **coarse by reference**, not a full definition, since template row-ids churn on re-apply but a reference is stable (re-apply NoOp, may decline several templates bound to the same script/path). App-only marker `0058_template_suppressions.sql` (`app_id NOT NULL` CASCADE, a `target_kind` discriminator), reconciled with the extension-point marker pattern (prunable → re-inherits). Consumed at the two resolution points: the trigger dispatch queries gain a correlated `NOT EXISTS` anti-join (gated to `t.group_id IS NOT NULL`), and `compile_effective_routes` drops an inherited (`depth > 0`) route at a suppressed path (loaded via `RouteRepository::list_route_suppressions`). **Inheritance-only** — the `group_id IS NOT NULL` / `depth > 0` gates mean an app can only decline what it inherits, never its own or a sibling's (pinned by `tests/template_suppression.rs` + the `suppress` journey); a dangling suppress is an apply-time warning; read-only `pic suppress ls --app`. **Trust-model consequence:** group templates are advisory-by-default (run *unless* a descendant declines) — a footgun for compliance hooks (audit/security triggers a tenant can opt out of)), and **§11 tail `sealed` (mandatory) group templates** (closes that footgun: a `[group]` marks a route/event-trigger template `sealed = true` and the two suppression filters skip it, so a descendant's `[suppress]` is ignored — it fires/serves on every descendant. Column `sealed BOOLEAN` on `triggers` + `routes` (`0059_sealed_templates.sql`); the trigger anti-join gains `AND t.sealed = FALSE` (a sealed row is never excluded → fires through), and `compile_effective_routes` gates its suppression `continue` on `!er.route.sealed`. `sealed` lives on the shared `Route` DTO + manager-core `Trigger` (both pure data) so the apply diff sees the current value — part of the route Update comparison + the trigger identity, so toggling it re-applies. Authored per-template (`sealed = true` on a `[[routes]]`/`[[triggers.kv|docs|files|pubsub]]`), **group-only** — `validate_bundle_for` rejects it on an app owner (an app resource is never inherited). Sealing only *strengthens* the guarantee (a sealed template can't be declined; it never grants new reach — the chain walk is still the isolation boundary). The dangling-suppress warning also flags a suppression matching only sealed templates ("… is sealed — the suppression has no effect"), and `pic triggers/routes ls --group` show a `sealed` column; pinned by `tests/sealed_templates.rs` + the `sealed` journey. Deferred: multi-node snapshot propagation), and **§11 tail M1 group-level suppression** (`template_suppressions` gained a polymorphic owner `0060_group_suppressions.sql` — a `[group]` declares `[suppress]` to decline a template it inherits from a higher ancestor for its **whole subtree**. Both filters generalized to the chain: the trigger anti-join joins the `chain` CTE (`ts.app_id = sc.app_owner OR ts.group_id = sc.group_owner`), and `list_route_suppressions` expands group suppressions across descendants via the all-apps `app_chain` CTE (`compile_effective_routes` unchanged). Still inheritance-only + `sealed` overrides; owner-polymorphic `suppression_repo::list_for_owner`/`insert`/`delete`; read-only `pic suppress ls --group`; the ineffective-suppress warning walks `GROUP_CHAIN_LEVELS_CTE` for a group node. Pinned by `tests/group_suppression.rs` + the `suppress` journey), and **§11.6 M2 shared-collection triggers** (a write to a group SHARED collection now fires a trigger — closes the "group trigger has no app to watch" gap. A group-owned trigger marked `shared = true` (`shared BOOLEAN` on `triggers`, `0061_shared_triggers.sql`) watches the group's shared collection; the per-app `list_matching_kv/docs/files` add `AND t.shared = FALSE`, new `list_matching_shared_{kv,docs,files}(owning_group,…)` select `shared = TRUE` triggers on the owning group — the `shared` flag is the namespace boundary. `ServiceEventEmitter` gained `emit_shared(cx, owning_group, event)`; `GroupKv/Docs/FilesServiceImpl` (via a `with_events` builder) emit on write, and the outbox emitter stamps the WRITER app_id so the handler runs under the writer. `shared` is authored on a group `[[triggers.kv|docs|files]]`, is part of the diff identity, and `validate_bundle_for` rejects it on an app owner / a non-collection kind / an undeclared collection. Read-only `shared` column in `pic triggers ls --group`; pinned by `tests/shared_triggers.rs` + the `shared_triggers` journey. Shared pubsub triggers shipped as D2 — see below), and **§11.6 M3 per-group quotas** (global env-var ceilings enforced in the group write path: `PICLOUD_GROUP_KV_MAX_ROWS`/`_DOCS_MAX_ROWS` (per-group row count, new-key only), `PICLOUD_GROUP_{KV,DOCS}_MAX_TOTAL_BYTES` (per-group total stored bytes, projected-total check — Track A M4) + `_FILES_MAX_TOTAL_BYTES` (per-group total bytes); `group_quota` helpers + `count_rows`/`total_bytes` repo methods + `QuotaExceeded`/`TotalBytesQuotaExceeded` errors), and **§11.6 M4 read-only operator admin API** (`group_blobs_api` mirrors the per-app kv/files admin surface for a group's shared collections: `GET /groups/{id}/{kv,docs,files}[/{collection}/{key|id}]`, authz `GroupKvRead`/`GroupDocsRead`/`GroupFilesRead`; the host hoists the group repos to share one instance; `pic kv ls/get --group`; reads-only, writes stay script-only), and **§4.5 M5 stateful group trigger templates via materialization** (cron + queue + email — the kinds that can't resolve live because each needs a per-app row: cron `last_fired_at`, the queue one-consumer advisory lock, the email sealed inbound secret. A group `[[triggers.cron|queue|email]]` template is **materialized** into an app-owned copy per descendant (`materialized_from` col, `0062_materialized_triggers.sql` + the `0063_materialized_unique.sql` partial-unique index that makes rematerialization idempotent under concurrency) by `materialize::rematerialize_stateful_templates` — all-apps `app_chain` CTE ⋈ group-owned stateful templates, a precise create/delete diff preserving cron state — run full-live at the route-rebuild chokepoints (apply single+tree, app create/delete in `apps_api`, group reparent in `groups_api`; `AppsState`/`GroupsState` gained a `pool`). The scheduler + `list_active_queue_consumers` + `email_inbound_target` gained `AND t.app_id IS NOT NULL` so a group TEMPLATE is never dispatched directly (nor invocable via its own webhook URL) — only its per-app copies are; a queue copy is skipped-with-warning when the app already fills that queue's slot. **M5.5 email uses the shared-group-secret model:** the template resolves its `inbound_secret_ref` against the **group's own** secret store once at apply and seals it (`resolve_and_seal`/`insert_email_trigger_tx` generalized to a `SecretOwner`/`ScriptOwner`); email secrets are now **v1 AAD-bound to the SEALING OWNER** (Track A M3, migration `0069_email_secret_version` + `seal_email`/`open_email`; the group AAD is stable across rows, so materialization still copies the sealed bytes **verbatim** — the inbound path recovers the sealing group via `materialized_from` and opens under its AAD; a legacy v0 read path remains for pre-M3 rows). All descendant webhooks share the one group HMAC secret; an unset group secret fails apply hard. Loading a group's own set-secret names into `CurrentState` (was hardcoded empty) lets the plan-time email-secret check resolve against the group. Pinned by `tests/stateful_templates.rs` + the `stateful_templates` journey. **Deferred:** the per-app (non-shared) email secret model), and **D1 the `materialized` column** (`pic triggers ls --app` now shows a read-only `materialized` column — a copy of an M5 group stateful template reads `true`, distinct from a hand-authored trigger; a derived `materialized` bool = `materialized_from IS NOT NULL` threaded row→domain→API→CLI mirroring `sealed`/`shared`), and **§11.6 D2 shared TOPICS + shared pub/sub triggers** (a group declares a storeless `topic` shared collection (`group_collections.kind` widened to `topic`+`queue` in `0064`); a `[[triggers.pubsub]] shared = true` handler watches it. `BundleTrigger::Pubsub` gained `shared` (part of the identity); `validate_bundle_for` requires the topic pattern's ROOT segment (`events.*` → `events`) be a declared `kind='topic'` collection, group-only. Scripts publish via the explicit `pubsub::shared_topic("events").publish("created", msg)` handle → `GroupPubsubServiceImpl` resolves the owning group (kind `topic`) from `cx.app_id`'s chain, requires editor+ (`GroupPubsubPublish`, fails closed on anon), and fans out via `PubsubRepo::fan_out_shared_publish` to `shared = true` pubsub triggers on that group, each outbox row stamped the WRITER app_id (M2 model). The per-app `fan_out_publish` gained `AND t.shared = FALSE` — a shared trigger never fires on a per-app publish and vice-versa (the `shared` flag is the namespace boundary); the owning-group chain walk is the isolation boundary. Pinned by `tests/shared_topics.rs` + the `shared_topics` journey. **External SSE subscription shipped** (Track A M6): `GET /realtime/shared/topics/{topic}` streams a shared topic to external clients; `RealtimeAuthority::authorize_subscribe_shared` resolves the owning group from the subscriber app's chain (reads-open — the resolution IS the authorization, a foreign subtree 404s), and `GroupPubsubServiceImpl::with_realtime` bridges publish→broadcast after the durable fan-out. **Deferred:** multi-node broadcast propagation (cluster mode)), and **§11.6 D3 shared durable QUEUES** (a group declares a `queue` shared collection; any subtree app enqueues into ONE group-keyed store (`group_queue_messages`, `0065`, mirrors `queue_messages` keyed by `(group_id, collection)`; CASCADE on group delete) via `queue::shared_collection("name").enqueue(...)` → `GroupQueueServiceImpl` resolves the owning group (kind `queue`), requires editor+ (`GroupQueueEnqueue`, fails closed on anon). **Consumption is by COMPETING CONSUMERS:** a group `[[triggers.queue]] shared = true` consumer (shared threaded through `BundleTrigger::Queue` + identity) **materializes** a consumer copy per descendant app (`materialize` skips the M5 one-consumer-slot check for shared — each descendant intentionally gets a consumer), and all copies claim the SHARED store with `FOR UPDATE SKIP LOCKED` — each message delivered at-most-once across the subtree, scaling horizontally, each handler under its own `cx.app_id`. The dispatcher's queue arm gained a `shared_group` on `ActiveQueueConsumer` (recovered via a LEFT JOIN to the materialized copy's source template) and routes claim/ack/nack/terminal to the group store when set (`q_claim/q_ack/q_nack/q_terminal` helpers; a group claim normalizes to a `ClaimedMessage` under the consuming app); the reclaim task drains both stores. `validate_bundle_for` requires a shared queue on a group to name a declared `kind='queue'` collection; a shared queue on an app is rejected by the app-owner shared guard. Pinned by `tests/group_queue.rs` (competing-consumer at-most-once) + `stateful_templates.rs` + the `shared_queues` journey. **Group dead-letter store shipped** (Track A M2): an exhausted shared-queue message is preserved in `group_dead_letters` (`0068_group_dead_letters`) via `GroupQueueRepo::dead_letter` (atomic INSERT+DELETE) and is operator-visible at read-only `GET /api/v1/admin/groups/{id}/dead-letters` (`GroupKvRead`). **Deferred:** fan-out to a *shared* dead-letter trigger), and **§7 multi-repo ownership M1 — the single-owner claim** (the `owner_project` seam (0047) is now live behind a first-class `projects` table (`0066_projects.sql`, UUID pk + unique slug; `owner_project` FKs it `ON DELETE SET NULL` — un-claim, never cascade-destroy a tree). A `[project]` block (slug + optional name) in the repo's ROOT manifest declares identity (independent of the `[app]`/`[group]` XOR); the first apply with a new slug registers the project and **claims** each group node it touches. The claim runs inside the apply tx under the per-node advisory lock **before the diff**, so a conflict short-circuits with a **409** before any write. Pure policy `decide_group_claim`/`decide_app_owner` in `apply_service` (unit-tested, DB-free): unclaimed→claim, owner→no-op, foreign→conflict unless `--takeover` (which additionally requires `GroupAdmin` — ownership ⟂ RBAC, mapped to 403 vs the 409 conflict), no-project-into-a-claimed-subtree→conflict. **Apps carry no owner** — an app inherits ownership from its **nearest claimed ancestor group** (the ancestor walk, via `groups.ancestors` now carrying `owner_project` through its recursive CTE, is the isolation boundary); an unclaimed subtree stays open, so nothing changes until a repo first declares `[project]` (backward-compatible). `ProjectRepository` (read side) + tx free-fns `upsert_project_tx`/`read_group_owner_tx`/`write_group_owner_tx`; the claim deliberately does **not** bump `structure_version` (not a diff change → won't churn a pending bound plan). Wire: `project`/`takeover` on the apply request (both `#[serde(default)]` → the pre-M1 CLI stays compatible); the CLI surfaces the server's 409 message verbatim (covers `StateMoved` + `OwnershipConflict`). Visibility: `pic groups ls` `owner` column (server `list_with_owner` LEFT JOIN; the shared `Group` deserialize ignores the extra field so `pic groups tree`/dashboard are unaffected) + `pic apply --takeover`. Pinned by `apply_service` unit tests, `tests/projects_repo.rs`, and the `apply_ownership` journey. **M2 shipped — the attach-point ceiling:** a `[project] parent_group = ""` binds the repo under a pre-existing group; `check_within_attach` refuses (422 `OutsideAttachPoint`) any node not strictly within that subtree (a group node must be a *proper* descendant — you can't apply the attach point itself; an app node's group must be at-or-below it), resolved via `groups.ancestors` and enforced read-only before the claim in `apply_owner`/`apply_tree`; absent = instance root = no ceiling. **M3 shipped — plan preview + `pic projects ls`:** `pic plan` now carries the `[project]` and returns an `ownership` preview per node (`claim`/`owned`/`conflict`-owner-named/`unclaimed`, pure `preview_ownership`) plus, for a group node, the cross-repo **blast radius** (descendant apps owned by OTHER projects the change reaches, via `group_blast_radius` — a subtree CTE with per-group memoized nearest-claimed resolution); the attach ceiling is previewed at plan too. Read-only `pic projects ls` (`GET /api/v1/admin/projects`, `list_with_counts`) lists projects + owned-group counts. Pinned by the `apply_ownership` plan-preview case. **With M3 the §7 multi-repo ownership track (M1 claim · M2 attach ceiling · M3 preview + `pic projects ls`) is COMPLETE.** Also shipped: **§6 group-tree Tier 1** (declarative group create via dir-nesting · structural-divergence detection · declarative reparent — `reconcile_group_structure_tx`/`StructureMode`, 422 `StructuralDivergence`) and **§3 M3 the per-env approval gate**, now **server-authoritative** (migration `0067_project_environments`; the gate resolves the governing project from the target node's nearest-claimed ancestor — `governing_env_policy`/`_tree` — so omitting/spoofing `[project]` can't bypass it; an approved gated apply requires AppAdmin/GroupAdmin step-up + audit). + +**Track A (v1.2 deferred-gap closeout, migrations 0067–0069) shipped to local main:** M1 hermetic approval gate · M2 shared-queue dead-letter store · M3 email-secret AAD v0→v1 · M4 per-group KV/docs byte quotas · M5 `set_if` compare-and-swap for KV (per-app + shared + Rhai SDK) · M6 shared-topic external SSE. **Audit 2026-07-11 remediation** (migration 0070, admin-session absolute cap) also shipped. **With that, v1.2 _Hierarchies_ is complete.** Next: multi-node cluster mode (the deferred multi-node route/broadcast propagation lives there); the **Workflows** track (DAG workflows, nested workflows, service interceptors) is the other, not-yet-started half of v1.2. **Data-model invariant:** app-owned data-plane tables (KV, docs, files, …) start with `app_id UUID NOT NULL REFERENCES apps(id) ON DELETE CASCADE`; the group-inheritable tables — _config_ (`vars`, `secrets`) and now group-owned _code_ (`scripts`, `0050`) — instead carry a **polymorphic owner**: nullable `group_id` and `app_id` with an exactly-one CHECK and per-owner partial-unique indexes (config is `ON DELETE CASCADE`, scripts `RESTRICT` — code is not data). Inheritance resolves **live** down `apps.group_id → groups.parent_id` via `CHAIN_LEVELS_CTE` (no materialized view); nearest-owner-wins with an app's own row shadowing the inherited one (CoW). Every Rhai SDK call resolves its app from `cx.app_id`, never a script-passed arg, and a group script always runs under the *inheriting* app's `cx.app_id` (the cross-app isolation boundary). @@ -143,6 +145,6 @@ Environment variables consumed by the `picloud` binary: ## Out of MVP -Queue triggers, cron triggers, SMTP ingress, KV / docs / email / users / HTTP SDKs in scripts, interceptors, workflows, function-to-function `invoke()`, secrets, metrics dashboard. All deferred to v1.1+ per the blueprint. Don't pre-build for them — but don't make decisions that close the door on them either. +This section captured the original MVP cut. Most of it has since **shipped in v1.1.x**: queue triggers, cron triggers, inbound email (`email:receive`, HMAC-webhook model), KV / docs / email / users / HTTP SDKs, function-to-function `invoke()`, and secrets are all live (blueprint §12 Phase 4 table). **Still deferred:** the **Workflows** track — DAG workflows, nested workflows, and **service interceptors** (§9.4) — plus a metrics/observability dashboard, a raw SMTP-listener ingress, and **multi-node cluster mode**. Don't pre-build for them — but don't make decisions that close the door on them either. -**Pulled forward to Phase 3 (pre-v1.1):** admin auth, multi-app scoping. Cross-app data sharing (export/import) stays at v1.3+; the initial cut enforces strict isolation. See blueprint §11.5. +**Pulled forward to Phase 3 (pre-v1.1):** admin auth, multi-app scoping. The general cross-app **export/import** sharing model stays at v1.3+; note that v1.2 §11.6 shipped a narrower form — **group-owned shared collections** (KV/docs/files/topics/queues) let apps in one subtree share data through the owning group, with the ancestor-chain walk as the isolation boundary. See blueprint §11.5 + design-doc §11.6. diff --git a/crates/executor-core/src/sdk/kv.rs b/crates/executor-core/src/sdk/kv.rs index 45e2002..67e99bb 100644 --- a/crates/executor-core/src/sdk/kv.rs +++ b/crates/executor-core/src/sdk/kv.rs @@ -62,7 +62,7 @@ pub(super) fn register(engine: &mut RhaiEngine, services: &Services, cx: Arc **Status:** Active — design discussion captured 2026-06-18, revised 2026-06-20 with resolutions > grounded in precedent (GitLab, Kustomize, Helm, Terraform, Kubernetes Server-Side Apply, Pulumi) > and corrected against the codebase after three independent review passes (consistency, gaps, -> feasibility). **Scheduled as the *Hierarchies* track of v1.2** (blueprint §12 Phase 5). §11 Phases -> 1–3 (project tool, groups + hierarchy RBAC, group-inherited config) are **shipped**; Phases 4–6 -> remain. NB: §11's own Phase 1–6 numbering is local to this initiative and is **distinct from the +> feasibility). **Scheduled as the *Hierarchies* track of v1.2** (blueprint §12 Phase 5). **§11 Phases +> 1–6 have all shipped** — project tool, groups + hierarchy RBAC, group-inherited config, group +> scripts/modules + extension points, nested tree apply, §4.5 trigger/route templates, §11.6 shared +> collections, §7 multi-repo ownership, §3 M3 (now server-authoritative) approval gate, §6 group +> create/reparent — plus the Track A closeout and the 2026-07-11 audit remediation. **v1.2 Hierarchies +> is complete;** only multi-node cluster propagation and the separate Workflows track remain. NB: §11's +> own Phase 1–6 numbering is local to this initiative and is **distinct from the > blueprint product-phase numbering** — "Phase 3" here = group-inherited config, not the blueprint's > admin-auth phase. > @@ -792,9 +796,10 @@ apps owned by OTHER projects the change fans out to (`group_blast_radius`, a sub memoized nearest-claimed resolution). The attach ceiling is previewed at plan too (consistent with apply). Read-only `pic projects ls` lists every project + its owned-group count (`GET /api/v1/admin/projects`, `list_with_counts`). Pinned by the `apply_ownership` journey's plan-preview case. **With M3, the multi-repo -ownership track (§7 M1–M3) is complete. Deferred (separate later work):** §6 structural-divergence detection -(`--adopt-server-structure` / `--force-local-structure`); declarative group create/reparent (lifting "groups -pre-exist"); per-env approval gating (`[project.environments]`). +ownership track (§7 M1–M3) is complete.** The items once deferred here have all **since shipped**: §6 +structural-divergence detection + declarative group create/reparent (see the §6 M1/M2 blocks below) and +per-env approval gating (`[project.environments]`, now server-authoritative — see §3 M3). Only multi-node +cluster propagation remains deferred. **§6 group-tree M1 shipped — declarative group CREATE.** A `[group]` node's PARENT is now inferred from directory nesting (the nearest ancestor directory holding a `[group]`; the topmost group binds to the repo's @@ -1232,12 +1237,12 @@ Resolved items now live inline next to their topic. What genuinely remains: > **Status (Phase 5): ✅ shipped — single-repo nested tree apply, atomic.** A directory tree of > `picloud.toml` manifests (each declaring an `[app]` or `[group]` node) applies as ONE - > server-computed plan in ONE Postgres transaction. Per the §11.1 review, the **multi-repo - > single-owner / attach-point / takeover** layer (§7) is **deferred** for the solo-dev / single-repo - > start, as is **per-env approval-policy gating** (`[project.environments]`); env overlays - > (`picloud..toml`) already exist. Groups must **pre-exist** (`pic groups create`) — a manifest - > owns each node's *content*, not the tree *shape* (declarative group create/reparent is a later - > add). + > server-computed plan in ONE Postgres transaction. The **multi-repo single-owner / attach-point / + > takeover** layer (§7), **per-env approval-policy gating** (`[project.environments]`), and declarative + > **group create/reparent** were all deferred at Phase 5 and have **since shipped** (§7 M1–M3, §3 M3 + > server-authoritative, §6 M1/M2). Groups no longer need to pre-exist — a manifest can own the tree + > *shape* as well as each node's *content*. Env overlays (`picloud..toml`) already existed at + > Phase 5. > > Shipped surface: > - **Engine:** the reconcile engine generalized from app-only to an `ApplyOwner { App | Group }` @@ -1334,14 +1339,16 @@ Resolved items now live inline next to their topic. What genuinely remains: > UPDATE SKIP LOCKED` — at-most-once across the subtree, horizontally scaled, each handler under its own > `cx.app_id`. The dispatcher's queue arm routes claim/ack/nack/terminal to the group store when the > consumer's source template is a shared queue (`ActiveQueueConsumer.shared_group`). Pinned by - > `tests/group_queue.rs` + `stateful_templates.rs` + the `shared_queues` journey. **Deferred:** a group - > dead-letter store (an exhausted shared-queue message is dropped-with-warning). + > `tests/group_queue.rs` + `stateful_templates.rs` + the `shared_queues` journey. **Group dead-letter + > store shipped** (Track A M2, `group_dead_letters` `0068`): an exhausted shared-queue message is + > preserved (atomic `GroupQueueRepo::dead_letter`) and operator-visible at read-only + > `GET /api/v1/admin/groups/{id}/dead-letters`. Deferred: fan-out to a *shared* dead-letter trigger. > - > **Deferred (documented gaps):** shared-topic external SSE subscription; per-group total-size quotas + - > write-rate limits; - > CAS/`set_if`; an operator admin API for shared blobs (scripts use the SDK; `pic collections ls` shows - > the marker — matches KV/docs); app-declared collections. Multi-node tree-apply leans on the runtime - > backstop for no-op edges, as elsewhere. + > **Deferred (documented gaps):** per-group write-rate limits; app-declared (as opposed to group-owned) + > collections. The items once listed here — shared-topic external SSE subscription, per-group total-size + > byte quotas, CAS/`set_if`, and the read-only operator admin API for shared blobs — have all **shipped** + > (Track A M4–M6 and §11.6 M4). Multi-node tree-apply leans on the runtime backstop for no-op edges, as + > elsewhere. ### 11.1 Re-sequencing review (post-Phase-3) diff --git a/serverless_cloud_blueprint.md b/serverless_cloud_blueprint.md index 641e87a..b94d3db 100644 --- a/serverless_cloud_blueprint.md +++ b/serverless_cloud_blueprint.md @@ -1,7 +1,7 @@ # Project Blueprint: Lightweight Event-Based Serverless Cloud -**Status**: Phase 4 — Blueprint Complete -**Last Updated**: 2026-05-27 +**Status**: v1.1 shipped (SDK + services) · v1.2 *Hierarchies* track **complete**; v1.2 *Workflows* track + v1.3 cluster mode are next +**Last Updated**: 2026-07-12 (reconciled to shipped code — CLAUDE.md is the live source of truth) **Audience**: Solo developer (DIY self-hosted) --- @@ -48,55 +48,57 @@ A lightweight, self-hosted, event-driven compute platform that allows developers ### High-Level System Diagram +> **Note (shipped reality):** the original design spawned a Docker container per +> execution. That was replaced by an **embedded, in-process Rhai executor** (better +> latency, less infra — see §12 Phase 5 note). The MVP ships as ONE `picloud` binary +> that runs the Manager, Orchestrator, and Executor as in-process crates +> (`manager-core` / `orchestrator-core` / `executor-core`); the same crates split into +> separate per-node binaries in cluster mode (v1.3+). Caddy fronts everything; the +> dashboard is a SvelteKit static SPA under `/admin`. + ``` -┌─────────────────────────────────────────────────────────────────┐ -│ Self-Hosted Server │ -├─────────────────────────────────────────────────────────────────┤ -│ │ -│ ┌──────────────────────┐ ┌──────────────────────┐ │ -│ │ Web Dashboard │ │ Orchestrator API │ │ -│ │ (Alpine.js SPA) │ │ (Rust + Axum) │ │ -│ │ Port 3000 │ │ Port 8080 │ │ -│ └──────┬───────────────┘ └──────────┬───────────┘ │ -│ │ │ │ -│ │ Upload script │ HTTP requests │ -│ │ Manage scripts │ Script metadata │ -│ │ │ │ -│ └────────────────┬────────────────────┘ │ -│ │ │ -│ ┌───────▼────────┐ │ -│ │ PostgreSQL │ │ -│ │ (scripts, MD) │ │ -│ └────────────────┘ │ -│ │ │ -│ ┌────────────────┼────────────────┐ │ -│ │ │ │ │ -│ ┌────▼────┐ ┌────▼────┐ ┌────▼────┐ │ -│ │Container │ │Container │ │Container │ │ -│ │ Instance │ │ Instance │ │ Instance │ (on-demand) │ -│ │(Rhai Ex.)│ │(Rhai Ex.)│ │(Rhai Ex.)│ │ -│ └──────────┘ └──────────┘ └──────────┘ │ -│ │ │ │ │ -│ └─────────────────┼────────────────┘ │ -│ │ │ -│ ┌────────▼────────┐ │ -│ │ Docker Daemon │ │ -│ │ (container mgmt) │ │ -│ └─────────────────┘ │ -│ │ -└─────────────────────────────────────────────────────────────────┘ +┌──────────────────────────────────────────────────────────────────┐ +│ Self-Hosted Server (single node — MVP) │ +├──────────────────────────────────────────────────────────────────┤ +│ │ +│ HTTP/HTTPS ──▶ ┌────────────────────────┐ │ +│ │ Caddy 2 (reverse proxy, │ │ +│ │ auto-HTTPS) │ │ +│ └───────────┬────────────┘ │ +│ │ │ +│ ┌─────────────────▼───────────────────┐ │ +│ │ picloud — all-in-one binary │ │ +│ │ (Rust + Axum + Tokio) │ │ +│ │ │ │ +│ │ Manager Orchestrator Executor │ │ +│ │ (control (ingress + (embedded│ │ +│ │ plane) dispatch) Rhai + │ │ +│ │ sandbox)│ │ +│ │ in-process manager-core / │ │ +│ │ orchestrator-core / executor-core │ │ +│ │ crates (split per node in cluster │ │ +│ │ mode, v1.3+) │ │ +│ └──────┬────────────────────┬──────────┘ │ +│ │ │ │ +│ ┌───────▼────────┐ ┌───────▼─────────────┐ │ +│ │ PostgreSQL 15+ │ │ SvelteKit dashboard │ │ +│ │ control + data │ │ (/admin static SPA, │ │ +│ │ plane (JSONB) │ │ CodeMirror editor) │ │ +│ └────────────────┘ └─────────────────────┘ │ +│ │ +└──────────────────────────────────────────────────────────────────┘ ``` ### Data Flow: HTTP Request → Response -1. **HTTP Request** arrives at Orchestrator (`POST /api/execute/{script_id}`) -2. **Orchestrator** fetches script from PostgreSQL -3. **Docker daemon** spawns container from pre-built executor image -4. **Container startup** loads script into Rhai runtime + passes request context -5. **Rhai script** executes, processes request, returns JSON object -6. **Orchestrator** extracts `statusCode`, `headers`, `body` from response +1. **HTTP Request** arrives (via Caddy) at the Orchestrator (a user route, or `POST /api/v1/execute/{id}`) +2. **Orchestrator** resolves Host → app → route, fetches the script (cached) from PostgreSQL +3. **Orchestrator** dispatches to the local Executor through the `ExecutorClient` trait (in-process call in MVP; an HTTP client in cluster mode) +4. **Executor** runs the script in the **embedded Rhai engine** with sandbox limits, passing the request context +5. **Rhai script** executes, processes the request, returns a JSON object +6. **Orchestrator** extracts `statusCode`, `headers`, `body` from the response 7. **HTTP Response** sent to client -8. **Container** is destroyed (scale to zero) +8. No per-call container — the engine is embedded, so there is nothing to spin up or tear down --- @@ -150,7 +152,7 @@ rhai_executor --script $SCRIPT_PATH --request "$REQUEST_JSON" --- ### 3.3 Dashboard (Web UI) -**Framework**: Alpine.js (MVP), Svelte (v1.0+) +**Framework**: SvelteKit (static adapter) with a CodeMirror 6 script editor, served under `/admin`. *(The original MVP sketch below named Alpine.js; the dashboard shipped directly on SvelteKit — there was no Alpine phase.)* **Port**: 3000 (default) **Features (MVP):** @@ -160,15 +162,19 @@ rhai_executor --script $SCRIPT_PATH --request "$REQUEST_JSON" - List of deployed scripts - Simple "Deploy" / "Delete" actions -**Technology Stack:** -- HTML + CSS + Alpine.js -- Fetch API to call Orchestrator -- No build step (initially), just serve static files +**Technology Stack (shipped):** +- SvelteKit (static adapter) + CodeMirror 6 editor, `paths.base = '/admin'` +- Fetch API to call the Manager/Orchestrator; session/bearer-token auth +- Built to static assets (`npm run build`) and served behind Caddy --- ### 3.4 PostgreSQL Database -**Schema (MVP):** +**Schema (MVP sketch — NOT authoritative):** the block below is the original MVP shape. The **authoritative +schema is the migration set** in `crates/manager-core/migrations/` (through `0070` as of v1.2), which has +since added apps/domains, RBAC (`admin_users`/`app_members`/`api_keys`), the v1.1 data-plane services +(KV/docs/files/queues/…), and the v1.2 groups/collections/templates/projects tables. Treat this as +illustration only. ```sql CREATE TABLE scripts ( @@ -660,12 +666,12 @@ users.set_permissions(user_id, { | Layer | Technology | Rationale | |-------|-----------|-----------| | **Orchestrator** | Rust + Axum | Performance, safety, async-first; minimal overhead | -| **Dashboard** | Alpine.js + vanilla HTML/CSS | Zero dependencies, simple to deploy, fast enough for MVP | +| **Dashboard** | SvelteKit (static adapter) + CodeMirror 6 | Reactive SPA, static build served under `/admin` behind Caddy | | **Database** | PostgreSQL 15+ (`pgcrypto`) | Robust ACID database; JSONB carries data-plane values (v1.1+). See §8.1. | -| **Container Runtime** | Docker (Docker daemon) | Industry standard, simple CLI | -| **Executor Image** | Alpine Linux + Rhai | Minimal image size (~50-100MB), fast startup | +| **Reverse proxy** | Caddy 2 | Auto-HTTPS in prod; same Caddyfile shape for single-node and cluster | +| **Execution model** | Embedded, in-process Rhai engine (no per-call containers) | Lower latency + less infra than the original Docker-per-execution model | | **Scripting** | Rhai | Lightweight, embedded-friendly, safe by default | -| **Deployment** | Docker Compose (local) / systemd (production) | Simple multi-service orchestration | +| **Deployment** | Docker Compose (dev + single-node prod) | One `picloud` all-in-one binary + Postgres + Caddy | --- @@ -694,7 +700,14 @@ docker-compose up docker-compose -f docker-compose.prod.yml up -d ``` -### docker-compose.yml (MVP Template) +### docker-compose.yml (MVP Template — illustrative) + +> **Outdated shape.** This template predates two decisions that shipped: the executor is now **embedded +> in-process** (no Docker daemon / `docker.sock` mount, no per-call executor containers), and **Caddy** is +> the reverse proxy (not nginx). The current dev/single-node stack is: `postgres` + the all-in-one +> `picloud` binary + `caddy`. Kept below for historical shape; see the repo's `docker-compose.yml` and +> `caddy/` for the real thing. + ```yaml version: '3.8' services: @@ -738,7 +751,7 @@ volumes: **Purpose**: gate the admin API (`/api/v1/admin/*`) and dashboard (`/admin/*`) behind per-user authentication. Before this phase the surface was open — anyone reaching the bound port could create, edit, and delete scripts. -**Why per-user, not a shared secret**: shared admin passwords get shared between humans, leave no audit trail, and can't be revoked per-person. Per-user accounts solve all three. The initial cut deliberately stops there — no roles, no per-app permissions — because that scope is small enough to ship in a single phase without blocking Phase 3b. Roles + per-app permissions are queued for v1.3+. +**Why per-user, not a shared secret**: shared admin passwords get shared between humans, leave no audit trail, and can't be revoked per-person. Per-user accounts solve all three. The initial Phase-3a cut deliberately stopped there — no roles, no per-app permissions. Those **shipped shortly after in Phase 3.5** (`instance_role` on `admin_users`; `app_members` for per-app `app_admin`/`editor`/`viewer`; a unified `can(principal, capability)` gate), and **hierarchy-aware group RBAC** followed in v1.2 §11 Phase 2 (`effective_app_role`/`effective_group_role`). ### Naming: `admin_users` vs `users` @@ -863,7 +876,7 @@ Permission checks land in middleware that initially only enforces "authenticated **Deviations from the design below**: none of substance. Two operational notes: - The Hello-World seed lives in `crates/manager-core/seeds/hello.rhai` and is inserted by a Rust bootstrap step (`seed_hello_world_if_fresh`) rather than from the migration — keeps it testable and gives the dashboard editor real source to render. The migration always inserts the `default` app + `localhost` claim; the seed only fires when that app is otherwise empty. -- Per-app admin roles/permissions are deferred — every authenticated admin can act on every app. The middleware seam (`auth_middleware::require_admin`) is the place where role checks slot in later. +- Per-app admin roles/permissions **shipped in Phase 3.5** — `app_members` grants (`app_admin`/`editor`/`viewer`) enforced through the `can(principal, capability)` gate; hierarchy-aware group RBAC extended this in v1.2 §11 Phase 2. (Originally this slot deferred them; the `auth_middleware` seam is where the checks landed.) **Purpose**: PiCloud hosts multiple independent applications on one platform. Each app is the isolation boundary for scripts, routes, domains, and (later) data — App A cannot see or modify App B's resources except through HTTP calls between them. @@ -1229,7 +1242,7 @@ Three foundation pieces that must land before the v1.1 service expansion, becaus **3a. Admin auth** — ✓ shipped. See section 11.4. Per-user `admin_users` (not a shared secret), Argon2id passwords, env-var bootstrap of the first admin, session-token doubling as bearer token for API. No roles in this cut; schema is forward-compatible with later RBAC. -**3b. Multi-app scoping** — ✓ shipped. See section 11.5. `apps`, `app_domains`, `app_slug_history` tables; `app_id` columns on `scripts`, `routes`, `execution_logs`. Migration assigns existing data to a `default` app and always claims `localhost`; a Rust-side bootstrap inserts a `Hello World` script + `/hello` route when the default app is empty. Orchestrator dispatch is two-phase (Host → app → route trie). `/api/v1/execute/{id}/*` continues to work without a public domain claim. Dashboard is app-hierarchical (`/admin/apps`, `/admin/apps/{slug}/...`); API stays flat with new endpoints under `/api/v1/admin/apps/*` and a `?app=` filter on script listing. Per-app admin roles deferred. +**3b. Multi-app scoping** — ✓ shipped. See section 11.5. `apps`, `app_domains`, `app_slug_history` tables; `app_id` columns on `scripts`, `routes`, `execution_logs`. Migration assigns existing data to a `default` app and always claims `localhost`; a Rust-side bootstrap inserts a `Hello World` script + `/hello` route when the default app is empty. Orchestrator dispatch is two-phase (Host → app → route trie). `/api/v1/execute/{id}/*` continues to work without a public domain claim. Dashboard is app-hierarchical (`/admin/apps`, `/admin/apps/{slug}/...`); API stays flat with new endpoints under `/api/v1/admin/apps/*` and a `?app=` filter on script listing. Per-app admin roles shipped in Phase 3.5 (§11.6). **3c. Users, roles, and bearer-token auth (Phase 3.5)** — ✓ shipped. See section 11.6. Adds `instance_role` to `admin_users` (`owner`/`admin`/`member`), `app_members` for per-app `app_admin`/`editor`/`viewer` grants, and `api_keys` for `Authorization: Bearer pic_…` credentials. Unifies cookie-session and API-key paths behind a single `can(principal, capability)` gate; list endpoints filter by membership at SQL for `member` users. Dashboard surfaces, invites, MFA, service accounts, and the `picloud` CLI binary are deferred — schema room only. @@ -1237,7 +1250,7 @@ Three foundation pieces that must land before the v1.1 service expansion, becaus --- -### Phase 4: v1.1 (Expand Capabilities & Services) — Current focus +### Phase 4: v1.1 (Expand Capabilities & Services) — ✓ Shipped (through v1.1.9) Released in patch steps (v1.1.0 → v1.1.8), each landing one focused capability. The split lets each release ship behind tests + docs without long-lived branches. SDK shape (handle pattern, `::` namespace, error convention, `ExecutionGate`, `SdkCallCx`, `ServiceEventEmitter` — see §7.5 and [docs/sdk-shape.md](../docs/sdk-shape.md)) is fixed in v1.1.0; every subsequent release fills in the contents without re-litigating the shape. @@ -1257,9 +1270,10 @@ Released in patch steps (v1.1.0 → v1.1.8), each landing one focused capability ### Phase 5: v1.2 (Advanced Workflows & Hierarchies) -Two tracks under the same minor line. The **Hierarchies** track is in active development; the -**Workflows** track follows it (or runs in parallel — sequencing is an open call, see the design doc's -§11 re-sequencing note). +Two tracks under the same minor line. The **Hierarchies** track is **complete** (all §11 Phases 1–6 +shipped, plus the Track A closeout + the 2026-07-11 audit remediation); the **Workflows** track is the +remaining, not-yet-started half (sequencing was an open call — see the design doc's §11 re-sequencing +note). **Hierarchies — groups + the declarative project tool** (the detailed design and 6-phase breakdown live in [docs/design/groups-and-project-tool.md](../docs/design/groups-and-project-tool.md) §11; this @@ -1272,10 +1286,18 @@ absorbs the original one-line "Hierarchies" bullet into a full initiative): `0048_vars.sql`, `0049_group_secrets.sql`.)* - A **declarative, file-based project tool** (`pic plan`/`apply`/`prune`, env overlays, the three-state `enabled` lifecycle, bound-plan staleness check). *(§11 Phase 1 — shipped.)* -- Remaining: group-owned **scripts/modules** with the scope-aware import resolver (§11 Phase 4); the - project tool **mapping onto the group tree** (nested manifests, attach points, ownership — §11 Phase - 5). Group-level shared *data* (collections/topics) stays in v1.3 (the cross-app-sharing problem - below). +- **Group-owned scripts/modules** with the scope-aware, sealed-by-default import resolver + opt-in + extension points. *(§11 Phase 4/4b — shipped: `0050_group_scripts.sql`, `0051_extension_points.sql`.)* +- The **project tool mapping onto the group tree** — nested manifests, single atomic tree apply, §7 + multi-repo ownership (claim/attach-ceiling/takeover), §3 M3 server-authoritative per-env approval + gate, and §6 declarative group create/reparent. *(§11 Phase 5 + §6/§7 — shipped: `0066_projects.sql`, + `0067_project_environments.sql`.)* +- **Group-level shared *data*** — shared collections (KV/docs/files/topics/queues), §4.5 group + trigger/route templates, suppression + sealed templates, per-group quotas. Pulled forward from v1.3 + into v1.2 §11.6 (the ancestor-chain walk is the isolation boundary; the general cross-app + export/import model remains v1.3+). *(shipped: migrations `0052`–`0065`, `0068`.)* +- **Only remaining in this track:** multi-node cluster propagation (route-table + broadcast snapshot), + which lives in the Phase 6 / v1.3+ scaling work below. **Workflows:** - Function workflows (DAG execution, conditional branching, error handling) @@ -1291,7 +1313,7 @@ absorbs the original one-line "Hierarchies" bullet into a full initiative): - Cross-app data sharing (explicit export/import model — see section 11.5) - Script versioning + rollback (keep N historical versions in a side table; rollback endpoint) - Rate limiting on endpoints -- Auth (richer model: API keys, OAuth, etc.) +- Auth (richer model: OAuth / SSO — `Authorization: Bearer pic_…` **API keys already shipped in Phase 3.5**) - Metrics + monitoring dashboard - Distributed tracing (OpenTelemetry) - Webhooks for execution events @@ -1301,12 +1323,21 @@ absorbs the original one-line "Hierarchies" bullet into a full initiative): ## 7. Complete Rhai SDK Reference (MVP → v1.1+) +> **Notation note (shipped surface):** the tables below list capabilities with an +> illustrative `service.method(...)` shorthand. The **actual shipped SDK uses the +> collection-scoped handle pattern with `::` namespaces** — e.g. `kv::collection("x").get(key)` +> and `docs::collection("x").create(data)`, not `kv.get("x", key)`. Group-shared variants +> use the explicit `kv::shared_collection("x")` / `docs::shared_collection("x")` / +> `pubsub::shared_topic("x")` / `queue::shared_collection("x")` handles. See +> [docs/sdk-shape.md](docs/sdk-shape.md) and [docs/stdlib-reference.md](docs/stdlib-reference.md) +> for the authoritative signatures. + ### Storage & Data | Component | Methods | Availability | |-----------|---------|--------------| -| **KV Store** | `kv.get(collection, key)`, `kv.set(collection, key, value, ttl?)`, `kv.delete(collection, key)`, `kv.has(collection, key)` | v1.1 | -| **Documents** | `docs.create(collection, data, schema?)`, `docs.find(collection, id)`, `docs.update(collection, id, data, schema?)`, `docs.delete(collection, id)`, `docs.list(collection, opts?)`, `docs.query(collection, filter?)` | v1.1 | -| **S3** | `s3.get(key)`, `s3.put(key, data)`, `s3.delete(key)`, `s3.list(prefix?)` | v1.1 | +| **KV Store** | `kv::collection("x").get/set/delete/has(...)`, `set_if(...)` (CAS) | v1.1 | +| **Documents** | `docs::collection("x").create/find/update/delete/list/query(...)` | v1.1 | +| **S3** | `s3.get(key)`, `s3.put(key, data)`, `s3.delete(key)`, `s3.list(prefix?)` | v1.3+ (not shipped — see Phase 6) | | **Users** | `users.create(data)`, `users.get(id)`, `users.find_by_email(email)`, `users.search(query, limit, offset)`, `users.list(filters)`, `users.update(id, data)`, `users.authenticate(email, password)`, `users.update_password(id, old, new)`, `users.lock/unlock(id)`, `users.delete(id)`, `users.send_invite(email)`, `users.send_password_reset(email)`, `users.send_login_link(email)`, `users.has_role/permission(id, role/perm)`, `users.add/remove_role(id, role)` | v1.1 | ### Communication @@ -1524,6 +1555,13 @@ email.send({ ## 9. v1.2+ Future Vision: Workflows & Hierarchies +> **Status note:** the **Hierarchies** half of this vision has **shipped** in v1.2 (groups, inherited +> config/scripts/modules, shared collections, templates, multi-repo ownership, approval gates — see +> [docs/design/groups-and-project-tool.md](docs/design/groups-and-project-tool.md) §11 and §12 Phase 5 +> above). What remains genuinely future here is the **Workflows** track below — DAG workflows, nested +> workflows, and **service interceptors** (§9.4) — plus advanced `docs.query()` filters. Those are +> unbuilt and the designs below are still speculative. + ### 9.1 Function Workflows (DAG Execution) **Concept**: Chain multiple functions together in a directed acyclic graph (DAG).