diff --git a/docs/design/groups-and-project-tool.md b/docs/design/groups-and-project-tool.md new file mode 100644 index 0000000..dd19cbf --- /dev/null +++ b/docs/design/groups-and-project-tool.md @@ -0,0 +1,849 @@ +# Groups, Projects & the Declarative Project Tool + +> **Status:** Draft — design discussion captured 2026-06-18, revised 2026-06-20 with resolutions +> grounded in precedent (GitLab, Kustomize, Helm, Terraform, Kubernetes Server-Side Apply, Pulumi) +> and corrected against the codebase after three independent review passes (consistency, gaps, +> feasibility). Not yet scheduled into a phase. +> +> **Scope:** turns the imperative `pic` CLI into a declarative, file-based project tool, and +> introduces a server-side **groups** hierarchy so apps can share and inherit config/scripts +> without duplication. +> +> **Blueprint impact:** this reverses the §11.5 *snapshot-copy, not live-link* stance — top-down +> hierarchical inheritance (group → app) is now in scope, implemented as a *materialized, auto- +> invalidated* view (§5.1). Fold the outcome back into the blueprint before it drifts. + +--- + +## 1. Motivation + +Today `pic` is **imperative**: every command (`pic deploy`, `pic routes create`, `pic triggers +create-cron`, …) maps 1:1 to an admin API call. There is no memory of desired state, no project +file, and no way to express "these N apps are the staging/prod/tenant variants of one thing." + +Two gaps follow: + +1. **No declarative project layer.** Developers re-run commands by hand; nothing reconciles a repo + to an instance. +2. **No server-side notion of a project/group.** If we model environments as separate apps, the + *shared* base (config, scripts) is duplicated across every deployed app. The server sees N + unrelated apps. + +This document designs both halves: a **declarative manifest + project tool**, and a **groups +hierarchy** on the server that the project tool projects onto the filesystem. + +`manager-core` is the single writer to Postgres. We lean on that for two concrete things — an +**atomic desired-state write** (one DB transaction per apply, §4.2) and a **server-computed diff** +(§4.2) — *not* for continuous reconciliation, which we deliberately decline (drift handling is +detect-and-surface, §4.2). This is a narrower and more honest claim than "we can reconcile live +state": most comparable tools cannot do an all-or-nothing apply because their effects are +un-undoable cloud resources; ours are Postgres rows, so we can — within the boundary described in +§4.2. + +--- + +## 2. Vocabulary + +| Term | Meaning | +|---|---| +| **Group** | Server-side hierarchy node. Single-parent tree. Owns shared **code/config** definitions and members/roles. Nestable. | +| **App** | A deployable unit. The data-isolation boundary (`app_id`). Lives under a group. | +| **Environment** | A deployment variant (staging/production/…). **An environment *is* an app.** Also a *scope dimension* on config (§3). | +| **Tenant** | Modeled like an environment — a scope dimension / overlay axis, **not** a second parent (§5.4). | +| **Project** | *Not* a server entity. A CLI view over a repo-managed **subtree** of groups. | +| **Definition** | Entities that can be group-owned: **scripts/modules, vars, secret-refs only**. Routes, triggers, collections, topics are **app-scoped** (§3, §5.1). | +| **Data** | KV/docs/files collections + pub/sub topics. **Always app-owned** (`app_id`). The isolation boundary never moves. | +| **Manifest** | TOML file describing desired state for one node (group base) + per-env overlays. | +| **Effective view** | The materialized, per-app resolved set of definitions, served to the runtime and recomputed on any ancestor change (§5.1). | +| **Attach point** | Where a local working tree's root binds into the server group tree. Its ceiling — you cannot apply above it. | + +Relationship: **the base + per-env-overlay "project" is the degenerate one-group subtree** of the +general groups model. Nesting generalizes it to multi-group subtrees. Nothing about the +single-project model is lost. + +--- + +## 3. Configuration resolution + +This is the single rule the whole design hangs off; everything else (env vars, secrets, `enabled`, +overrides) resolves through it. **All three precedent systems converge on the same shape, so we +adopt it wholesale:** + +> ⚠️ **Net-new, not an extension (verified against the codebase).** PiCloud has **no env-agnostic +> config / `vars` layer today** — this resolution engine and the `vars` table are greenfield. Only +> `secrets` exists, and as a value-bearing table, not a ref. Read §3 as *new* infrastructure, not a +> modification of something present. + +> **Resolution = sparse, per-field merge; environment is a *pre-filter*, not a precedence tier; +> proximity wins across levels.** + +To resolve a key for an app in environment `E`: + +1. **Filter by environment first, per level.** A value scoped `@E` is eligible; a value scoped `*` + (env-agnostic) is the fallback. At a *single* level, `@E` beats `*`. Environment scope decides + *eligibility*, it is **not** a competing precedence rank. + *Evidence:* GitLab CI/CD variables use `environment_scope` exactly this way — a `production`-scoped + row simply isn't visible to a `staging` job ([GitLab CI/CD variables](https://docs.gitlab.com/ci/variables/)). +2. **Then nearest level wins.** Walk up `acme → team-a → blog → app`; the closest level that defines + the (filtered) key wins. GitLab states this verbatim: *"if the same variable name exists in a + group and its subgroups, the job uses the value from the closest subgroup."* +3. **Merge granularity:** maps/vars deep-merge **per key** (set `title`, still inherit `region`); + entities (scripts) replace **by identity** (name); deletion is an explicit tombstone. + *Evidence:* Kustomize strategic-merge patches are sparse and deep-merge maps but **replace lists + unless keyed** ([Kustomize patches](https://kubectl.docs.kubernetes.io/references/kustomize/kustomization/patches/)); + Helm deep-merges nested maps and uses `null` to delete a defaulted key + ([Helm values](https://helm.sh/docs/chart_template_guide/values_files/)). + +This one rule resolves several issues at once: it makes **group config environment-scopable** (a +group sets `db_url@production` and `db_url@staging`; descendants inherit the right one — no per-leaf +duplication), it defines **merge granularity** (maps per-key, entities per-identity), and it makes +**`enabled` just another sparse field** (§4.3). + +### 3.1 The env-scope manifest syntax + +```toml +# group manifest (e.g. team-a/picloud.toml) +[vars] +region = "eu" # env-agnostic (scope *) + +[vars.production] # scope @production +db_url = "postgres://prod/..." +[vars.staging] +db_url = "postgres://staging/..." + +[secrets] +names = ["stripe_key"] # name-only; values pushed per env via CLI (§4.6) +``` + +### 3.2 The one deliberately-novel precedence call + +With `db_url@production` at the root **and** a plain `db_url` at the leaf, the **leaf's env-agnostic +value wins** for a production app — proximity beats farther-level env-specificity. The research was +explicit that **no major system lets env-specificity override proximity across levels**; GitLab +structurally avoids the question by treating env as a per-level filter (above). So **proximity-first +is the evidence-backed default**, and we document it as a chosen rule: *inner scope shadows outer, +like lexical scoping.* To avoid this becoming a debugging nightmare at depth, `pic config +--effective --explain` must show **why** a value resolved (which level/scope it came from). + +> **Residual risk (carried, not solved):** multi-level + environment scopes is genuinely novel +> territory — GitLab is two-level (group→project). Proximity-first at arbitrary depth is consistent +> and defensible but not battle-tested. The `--explain` tooling is a hard requirement, not a nicety. + +--- + +## 4. Manifest & apply + +### 4.1 Manifest & layering + +- **TOML, data not code.** PiCloud runs untrusted scripts; the deployment descriptor stays inert, + diffable, server-validatable, and dashboard-renderable. Turing-completeness belongs in the Rhai + script, never the manifest. +- **Base + overlays.** A node's `picloud.toml` holds the shared base; `picloud..toml` overlays + add per-env vars/secrets/slug. Shared config is written once; overlays only carry deltas (the + Kustomize base+overlay model — overlays are sparse patches, not replacements). +- **Routes are top-level**, referencing scripts by name (one place to see all routing). +- **`pull` is first-class** — see §4.6 for its inheritance/masking semantics. + +### 4.2 The apply engine + +The IaC research reshaped this section materially; the relevant precedents are cited inline. + +**Plan/apply split with a bound plan artifact.** `pic plan` computes a reviewable diff +(`create / update / delete / replace`), the server persists it as a row, and `apply` executes +*exactly that stored plan* — **detecting and refusing if state moved underneath it** (state +version/serial check, cheap given the single writer). *Evidence:* Terraform's saved plan +*"reliably perform[s] an exact set of pre-approved changes, even if the configuration or state has +changed in the minutes since"* ([Terraform run](https://developer.hashicorp.com/terraform/cloud-docs/run)). +The "state moved" check covers **both** the per-node **content** version **and** the **tree-structure** +version (§6): a reparent or a new app created under the subtree between `plan` and `apply` changes the +true blast radius, so the bound plan must refuse if *either* counter moved. Both are **net-new** — the +existing `scripts.version` column is an unconditional write counter, **not** a compare-and-set guard, +and must not be mistaken for one. + +**Atomicity — the corrected position.** No major IaC tool does transactional apply, *because their +effects are un-undoable cloud resources* (you cannot un-create a VM). **That constraint does not +bind us:** a PiCloud manifest apply is **almost entirely Postgres row writes** (script source, +route/trigger defs, config). Therefore: + +- **The desired-state write is a single DB transaction** — all-or-nothing across the subtree. This + is achievable *because* our substrate is one transactional store (unlike Terraform's un-undoable + cloud resources). It removes the downstream-breakage hazard of an earlier draft (which proposed + per-node, stop-on-error apply — **now superseded**): there is no intermediate state where a parent + committed and a child did not. +- **Propagation is forward-convergent, not transactional.** The post-commit effective-view refresh + + cache invalidation (§5.1) live *outside* the transaction. So "atomic" means + **atomic-for-desired-state, eventually-consistent-for-effect**, with a bounded window where the DB + says new and a cache says old. + +> **Invariant to enforce — and it does NOT hold today (verified).** The single-transaction model +> requires apply to write **only Postgres**, but the *current* admin write paths violate that: route +> and domain writes prime in-process caches **synchronously, interleaved with the write** +> (`route_admin.rs:234`, `apps_api.rs:396`), and files fsync to disk. So moving to "commit the +> transaction, *then* refresh views" is a deliberate **restructuring of manager-core**, not a +> description of the status quo. Additionally, **domains** are a non-DB effect (Caddy/filesystem, out +> of band) — they must be **excluded from the transactional core** and handled by the convergence +> model below. Treat "apply writes only Postgres" as an invariant to *establish and guard*, not one +> we already have. + +**Idempotent, convergent recovery (for the non-transactional propagation and any future side +effects).** Operations are idempotent upserts keyed by stable identity; "re-apply converges" is the +recovery story; "rollback" means revert the declaration and re-apply forward. Per-node status +(applied/pending/failed/drifted) is recorded so a partial propagation failure is observable and +re-runnable. *Evidence:* every surveyed tool (Terraform, ArgoCD, Flux, Pulumi) chose idempotent +forward-only convergence over distributed rollback. + +**Drift model — upgraded from pure last-write-wins.** An earlier draft chose model "(c)": +last-write-wins, just log. That is *exactly* the pre-SSA `kubectl apply` footgun — KEP-555 lists the +scenario verbatim: *"User does an apply, then `kubectl edit`, then applies again: surprise!"* +Kubernetes Server-Side Apply exists to convert that silent clobber into an explicit conflict +([K8s SSA](https://kubernetes.io/docs/reference/using-api/server-side-apply/)). Concrete risk for us: +CI runs `pic apply` and silently re-enables a route an operator killed in an emergency. Resolution: + +- **Keep (c)'s simplicity for ordinary fields** (last-write-wins, change reported). +- **For the security-relevant subset only** — `enabled` and a secret's *reference/existence in + desired state* — track a "changed out-of-band" bit. If that subset was changed outside the manifest + (e.g. an operator disabled a route in the dashboard), `pic plan` surfaces it as a **labeled + conflict** and `apply` **refuses without `--force`**. This is a scoped version of SSA's field + ownership; we deliberately avoid full per-field managers (object bloat, complexity). +- **Secret *values* are explicitly out of scope of the conflict/state-version machinery.** Values + are never in the manifest or the plan (§4.6), so a routine `pic secret set` between `plan` and + `apply` is the *supported* workflow and does **not** trip the state-version refusal — the + state-version check covers manifest-managed *desired state* (definitions, including secret + *references* and `enabled`), not value rotation. +- Out-of-band changes render as a **distinct labeled diff** in `plan`, separate from intended + changes. *Evidence:* Terraform/Pulumi both make read-only drift detection default and label drift + distinctly ([Pulumi drift](https://www.pulumi.com/docs/iac/operations/stack-management/drift/)). + +> **Residual risk:** the conflict bit reintroduces some of the complexity (c) was chosen to avoid. +> The cruder fallback, if even that is too much: keep pure (c) but add a non-revertible +> **operational lock** flag an operator sets in an emergency that `apply` won't touch. Either closes +> the silent-revert hole; neither is free. + +**Gating high-stakes applies: trigger ≠ authorization.** Any CI trigger may `plan`, but applying to +a confirm-required env is a separate, default-off, explicitly-authorized step. A blanket `--yes` +covers ordinary confirms; **confirm-required envs require an explicit per-env `--approve `**, so +CI must opt in per environment. "Override a gate" is its own audited capability (maps onto +`manager-core::authz::can`). *Evidence:* Terraform Cloud parks runs in *Needs Confirmation* and +separates the *apply runs* permission from *manage policy overrides* +([TFC run states](https://developer.hashicorp.com/terraform/cloud-docs/run/states)). + +**Concurrency.** Start with a **coarse per-instance (or per-root-group) apply lock** — one apply at a +time — which is trivially correct for the single-node MVP and makes "last-commit-wins" hold. Refine +*later* to **per-blast-radius advisory locks** (lock the affected apps in `app_id` order so +overlapping radii serialize while disjoint ones proceed; queue triggers already use this advisory-lock +primitive) only if contention appears. Don't build the fine-grained version speculatively. + +**Blast radius — defined and bounded.** The blast radius is **the set of descendant apps whose +materialized effective view would actually change** — a *diff*, not "all descendants." Per changed +definition at node N it is `subtree(N)` minus apps that override that key nearer. `pic plan` +enumerates it for small radii; for large ones it **summarizes (count + sample)** and a threshold +triggers extra confirmation (`"this changes 4,213 apps — confirm"`). Because only scripts/vars/secret- +refs inherit (§5.1), blast radius applies to *those* changes; an app-scoped change (a route or +trigger) has blast radius = that single app. Root-level changes are accepted as expensive, rare, and +high-confirmation; the computation is bounded by `subtree(N)` size, and the confirmed radius is +re-validated at apply against the tree-structure version (above) so it can't go stale. + +**In-flight executions.** An apply that disables or replaces a script affects **new invocations +immediately** (via the `enabled` re-check + view invalidation; pending trigger outbox rows are dropped +at fire-time, §4.3 — so no separate outbox purge is needed). **In-flight *running* executions run to +completion**, and note (verified) the executor has **no external-cancel path today**: executions are +`spawn_blocking` Rhai calls interruptible only by their operation budget or a pre-set wall-clock +deadline self-checked in `engine.on_progress`. A true must-stop-now **kill-switch** is therefore a +*net-new* capability — buildable by having that same `on_progress` hook also poll a per-execution +cancel flag — **gated by an admin capability (`authz::can`) and audited**, scheduled as a later item, +not phase 1. Killing mid-run also risks partial side effects, so it stays the explicit extreme, never +default apply behavior. + +**Apply flow (end to end):** + +``` +CLI builds bundle (manifests + script sources) for the subtree + → server computes plan + blast radius (incl. descendant apps in OTHER repos) + → server persists plan artifact; checks state version + → dev reviews; confirm/approve per env policy + → server applies desired state in ONE DB transaction (all-or-nothing) + → server refreshes effective views + bumps generation + invalidates caches (convergent) + → server returns change report; CLI logs created / updated / re-enabled / pruned / conflicts +``` + +### 4.3 The three-state `enabled` lifecycle + +`enabled` is a real platform feature (DB + server runtime + UI badge/toggle), not CLI sugar, and — by +§3 — it is **just another sparse, proximity-resolved field**. + +| State | Meaning | Pruned? | +|---|---|---| +| declared, `enabled = true` (or omitted) | deployed, active | kept | +| declared, `enabled = false` | deployed but **inert** (route short-circuits, trigger doesn't fire, script not invocable) | **kept** — still desired state | +| absent from merged manifest | stale | **deleted** by prune / `--prune` | + +- Default `true`; last-write-wins on merge; a base `enabled = false` is inherited until an overlay + (env *or* a nearer group/app) explicitly sets `enabled = true`. Overriding a *different* field does + **not** implicitly re-enable. +- **Across the group axis (resolved):** a descendant disables an inherited (group-owned) script via a + **sparse override** — an override row that sets only `enabled = false`, inheriting the source. A + descendant can re-enable a parent-disabled entity because nearest-level wins. This is the same + sparse-field mechanism as everything else (Kustomize patches are sparse — you specify only what + changes). +- **Base = the superset across envs.** Per-env you may toggle `enabled` in *either* direction + (disable a base-active entity, or re-enable a base-disabled one — proximity wins, §3), but you + **cannot *remove* a base entity** in one env; true removal requires the entity not be in the base. + An entity unique to one env goes in that overlay; an entity present everywhere but off in staging + goes in base with `enabled = false` in the staging overlay. +- **Disabled = invisible.** External callers hitting a disabled route get **404** (indistinguishable + from absent — no info leak). +- **Schema note:** the `triggers` table **already** has `enabled` (+ `dispatch_mode`, retry columns) + and it is **honored at match/schedule time** (`trigger_repo.rs`, `cron_scheduler.rs` — verified). + New work is `enabled` on **scripts** and **routes** only, plus runtime honoring in the matcher / + invoker. +- **Outbox fire-time gap (verified):** the dispatcher does **not** re-check `enabled` on an + already-enqueued outbox row (`dispatcher.rs:699`), so disabling a trigger stops *new* matches while + a pending item still fires. **Fix:** add an `enabled` re-check (trigger *and* script) at fire time + in `resolve_trigger`, so a pending outbox row for a now-disabled trigger/script is dropped when it + comes up — closing the gap cheaply, with no cache or kill-switch involvement. This is the home for + the §5.1 security-disable guarantee on the trigger path. +- **Provenance caveat (accepted):** a single boolean carries no "manual vs manifest" provenance, so a + disabled entity looks the same however it got there — which is *why* the §4.2 conflict bit on the + `enabled`/secrets subset exists, to stop the next apply silently reverting an operational disable. + +### 4.4 Identity & naming + +- **Kebab everywhere.** One canonical identifier regex: `^[a-z0-9][a-z0-9-]{0,62}$` for project + names, env names, script names/slugs, trigger names — unified with the existing app-slug rule. +- **Scripts:** unique `name`/`slug` per app = merge/upsert key. +- **Routes:** identity = the triple **` `**, e.g. + `ANY *.beta.example.com /hello/:name`. `dispatch_mode`, `host_param_name`, etc. are *attributes* + overridable without changing identity. The CLI infers `host_kind`/`path_kind` from the pattern + syntax (`*`, `{name}`, `:name`, exact), with an explicit `kind` key as override (mirrors the UI). + Needs a normalization rule (default method `ANY`, default host `*`, case-folding) so manifest ↔ + server match exactly. +- **Slugs are instance-global, derived from the path.** Two identifiers coexist: + - **path** = `acme/team-a/blog` — hierarchical, group-scoped, display/organization/RBAC. + - **slug** = flat, instance-global, the deployment key. + The derived default is **`{flattened-path}-{env}`** (e.g. `team-a-blog-staging`) — unique by + construction in the multi-group world, unlike a bare `{leaf}-{env}` which collides whenever two + groups reuse a leaf name. On >63-char overflow, truncate + short hash suffix. Explicit override + allowed; the path never *is* the slug, it only *seeds* it. + +### 4.5 Triggers + +Triggers are **app-scoped, not group-inherited** (§5.1 explains why). Two senses of "app" are in play +and must not be conflated: a trigger is declared once in the **leaf (logical app)** base — an +authoring convenience — and **materializes into per-`app_id` rows**, one per env-app, where `app_id` +is the server-app isolation boundary (§5.2). It is never group-owned. + +Two distinct constraints: + +- **`name`** (explicit, kebab) = the merge/identity key + upsert target. Unique per app. + *(Triggers have no name column today — new.)* +- **Semantic uniqueness** = a *post-merge validation*: no two triggers may share their kind-specific + semantic key. Checked after merging, so overriding a base trigger per-env (reuse `name`, change a + field) is fine; two differently-named triggers with identical effect is an error. + +| Kind | Semantic key | Note | +|---|---|---| +| kv / docs / files | `(script, collection_glob, ops)` | **canonicalize `ops`** (sort + dedupe) | +| cron | `(script, schedule, timezone)` | exact TEXT match | +| pubsub | `(script, topic_pattern)` | | +| dead-letter | `(script, source_filter, trigger_id_filter, script_id_filter)` | NULL = literal value (not wildcard) for equality | +| email | `(script)` | no filters — one per script | +| **queue** | `(queue_name)` — **not** script-scoped | already server-enforced (advisory lock: one consumer per `(app_id, queue_name)`) | + +- **Matching vs identity:** `[]`/`NULL`/globs mean **"any/wildcard" at dispatch time** (unchanged + runtime matching) but are compared as **literal structural values for dedup**. We dedup on + **structural identity, never on overlap/subsumption** — `ops = []` and `ops = ["insert"]` are *not* + duplicates. Overlapping triggers coexist (multiple triggers firing on one event is already how + PiCloud works). +- **Name backfill** for existing nameless rows: `{kind}-{entity}-{n}`, where `{entity}` is the kind's + identity token (collection / queue / topic; for cron/email/dead-letter fall back to `{kind}-{n}`), + sanitized to the kebab regex, with `-{n}` (discovery order) guaranteeing per-app uniqueness. + +> **Ergonomic debt (accepted, watch it):** because triggers don't inherit, 100 tenant apps each +> needing the same 5 triggers = 500 declarations. The fix is group trigger/route **templates** that +> fan out per descendant (a *template/instantiation* mechanism, not inheritance) — deferred, but it +> bites early if tenant cardinality is high. Pressure-test against real tenant counts before +> committing to the narrow-inheritance choice (§5.1). + +### 4.6 Secrets & `pull` + +- **Name-only in the manifest; value pushed via CLI** (`pic secret set`, reads stdin). Unanimous + across every comparable tool — never commit secret values. +- **Env-scoped** like any var (`[secrets]` names declared once; values set per env). +- **Warn (don't block) if a referenced secret is not set.** This requires an app dev to see that a + group secret **exists / is set** (a boolean) without reading its value — an accepted, explicit + authz boundary. GitLab surfaces inherited masked variable *keys* the same way. +- **Email's `inbound_secret` is a reference**, not inline — same rule; the server already encrypts it + at rest. +- **`pull` exports own-rows only** (this node's overrides), **never effective/inherited state.** A + separate read-only **`pic config --effective`** shows the inherited result with **masked secrets + rendered as `` / `***`, never plaintext.** This makes pull-under-masking safe *by + construction* — you cannot pull a secret you cannot read, and you cannot accidentally duplicate + inherited config into a leaf. *Evidence:* GitLab shows inherited group variables in a project + read-only and separate from project-own variables. +- Flat pull for a new project; "smart" delta-pull (own-vs-effective diff) is server-computed since an + app dev's checkout lacks ancestor manifests. + +### 4.7 Apply-time warnings + +- Enabled route/trigger pointing at a **disabled script**. +- An **endpoint** script deployed with **no route and no trigger** (unreachable). Modules are exempt. +- Sandbox override exceeding the admin **ceiling**. +- Referenced secret not set. +- An out-of-band change to the `enabled`/secrets subset (surfaced as a conflict, §4.2). + +--- + +## 5. Groups & inheritance + +### 5.1 What inherits, and the runtime model + +- **GitLab-like, nested, single-parent tree.** Single parent keeps inheritance acyclic and + resolution deterministic (no diamond precedence). +- **Inherit code/config — narrowly. Inherit data — no.** Group-inheritable = **scripts/modules, + vars, secret-refs only.** Routes, triggers, collections, topics, files are **app-scoped.** + - *Why narrow:* routes and triggers are **bindings**, not pure code/config. A group-owned trigger + has no app data to watch (triggers fire on app-owned collections/topics/queues); a group-owned + route has no host (routing is Host→app first). Inheriting them is incoherent. There are two + distinct sharing mechanisms and an earlier draft conflated them: the **leaf base+overlay** shares + routes/triggers across *one app's environments*; **group inheritance** shares *code/config across + many apps*. *Evidence:* Serverless/SAM keep `events:` per-function and share logic via *layers* — + bindings local, code shared. + - *Why data stays app-owned:* a group script executes in the *inheriting app's* context, so its + `cx.app_id` still scopes data to that app. Group-level *collections/topics* would break `app_id` + as the isolation boundary — that is the v1.3 cross-app data-sharing problem and stays **out**. +- **Runtime model: a materialized effective view + versioned cache (not per-request live-resolve).** + An earlier draft said "live-resolve," which is underspecified and would fight the existing cache. + The real model: + - manager-core **resolves-at-write into a materialized per-app effective view** (§3 rule applied: + sparse merge, env filter, proximity, CoW, `enabled`). + - The orchestrator/executor serve from that view, **keyed by `app_id` + a generation/version**. The + app_id-keyed *shape* survives (today's route cache is already an `app_id`→routes map), **but the + substance is net-new (verified):** there is no generation counter anywhere today, the route cache + is rebuilt whole-table on every write rather than per-app, and script *bodies* are live-resolved + per request. Read "serve from it" as *extend the keying*, not *reuse the mechanism* — and note + that fanning out today's full-rescan invalidation to thousands of descendants would be a + regression, so per-app incremental invalidation is part of the build. + - On any write to a node, manager-core (single writer, knows the tree) **recomputes descendants' + views + bumps the generation + invalidates caches.** + - The materialized view is a **derived cache, not a second source of truth** — canonical config + still lives once at the owning node, so this does **not** reintroduce duplication. This dissolves + the long-running "snapshot vs. live" tension: no duplication *and* propagation *and* a fast hot + path. + +> **Residual risk (relocated, not solved):** cache invalidation is now a **correctness/security +> requirement** — disabling a script for a security reason must stop it running *everywhere* within +> bounded time. This is the same class as the existing PrincipalCache revocation lag. Require +> **synchronous invalidation for the security-relevant subset** (`enabled=false`, secret rotation) +> and accept bounded eventual staleness elsewhere; the hard SLA at fan-out to thousands of +> descendants is genuinely unsolved. + +### 5.2 Schema impact + +- Inheritable definition kinds get a polymorphic owner (`owner_kind ∈ {group, app}` + `owner_id`): + **scripts** (modify the existing table — and note its `app_id` FK is `ON DELETE RESTRICT`, not + CASCADE, so a pruning apply needs explicit ordering), plus **`vars` and `secret-refs`, which are + net-new tables** (they do not exist today — §3). So this is *one table modified + two invented*, + not "three tables touched." **Routes, triggers, collections, topics, files stay strictly + `app_id`-owned**, so the runtime isolation boundary stays fixed. (The CLAUDE.md `app_id NOT NULL … + CASCADE` rule is itself not universal today — scripts already use RESTRICT.) +- New: a `groups` table (single-parent, `parent_id`), group `membership`/roles, an `owner_project` + column on group nodes (§7), and the materialized effective-view store keyed by `app_id` + + generation (§5.1). Env-scoped values carry an `environment_scope` column (`*` or a specific env). + +### 5.3 RBAC + +- **Hierarchy-aware capabilities.** `authz::can(principal, cap, on=node)` resolves by walking + ancestors and taking the highest effective role. Instance → group(s) → app. +- **Inherited membership** (GitLab-style): a group admin is implicitly admin of every subgroup/app + beneath it. +- **Masked group secrets:** a group secret is *used by* an app at runtime but *not human-readable* by + the app's developers. Two orthogonal gates: **runtime resolution** (engine injects plaintext) vs + **human-read authz** (admin API returns a value only to a principal with rights at the *owning + group*). An app-scoped admin call never returns group secrets; runtime injection bypasses the human + gate. App devs may see a group secret **exists** (for the unset-warning, §4.6) but not its value. + **An app can run with config its own developers cannot see.** + + > **Crypto caveat (verified):** secrets today are AES-GCM sealed with AAD `secret:{app_id}:{name}`, + > and the decrypt path hard-codes `cx.app_id` (migration 0042). A **group-owned** secret is *not + > expressible* under this scheme — there is no group identity in the AAD. It needs a new AAD + > identity (e.g. `secret:group:{group_id}:{name}`) **and** an owner-aware decrypt path that resolves + > whether the inherited secret is group- or app-owned. This is the single hardest correctness detail + > of group secrets and gates phasing step 3. + +### 5.4 Tenants & single-parent + +Single-parent forbids an app combining two *sibling* groups' configs — which seems to threaten the +multi-tenant use case (shared-platform base + per-tenant overlay). It does not, because **the +shared base is an *ancestor*, not a sibling:** model tenants as leaf apps under a `tenants/` +subgroup that inherits the platform base up the chain (tenant leaf → `tenants` group → platform +group). The "two parents" intuition is satisfied by the *chain*. Better still, **a tenant is a scope +dimension like environment** (§3) — the same env-scope machinery generalizes to a `tenant` scope, so +multi-tenant needs no new hierarchy primitive. *Evidence:* GitLab is single-parent and serves +multi-tenant teams via subgroups; "combine two siblings" is handled by promoting shared config to a +common ancestor. + +--- + +### 5.5 Module & import resolution under inheritance + +A group-owned script may `import` modules, and inheritance makes "which module?" ambiguous (an +inherited script importing a module a leaf has shadowed could bind to app-dev code it never saw — a +trust inversion). The rule: + +- **Lexical by default.** An inherited script's imports resolve against the module set visible at the + **script's own defining node** (walking up from there), **not** the inheriting app's effective + view. A leaf cannot shadow a module an inherited group script depends on, and a group script + behaves identically across every app that inherits it — preserving both **determinism** and the + **trust boundary** (a security-authored shared script is tamper-proof from below). Ordinary lexical + scoping; "sealed by default", like non-`open` classes in Kotlin / `final` in Java. +- **"Defining node" of a CoW-overridden script = the node that authored *that body*.** An + inherited-unchanged script's defining node is the ancestor that owns it (imports sealed to the + ancestor's modules). A script a leaf *overrides* (CoW, §3) has the **leaf** as its defining node, so + the override's imports resolve from the leaf — which is *not* a trust inversion, because the app + owner wrote that body and is trusted within their own app. Consequence (intended): the same module + name can resolve to **two different bodies** in one app, depending on the importing script's defining + node. +- **Explicit extension points for opt-in polymorphism.** A module is marked an extension point in the + **manifest** (an `[extension_points]` declaration — *not* in Rhai source, keeping code inert and + giving the `plan` checker something to read). Such a module is one descendants are *expected* to + provide or override; **only** these resolve against the inheriting app's effective view. Controlled + template-method customization (a shared `render` whose `theme` module each tenant supplies) without + the blanket trust inversion of dynamic resolution. When a name is declared at more than one level — + or is concrete on one path and an extension point on another — **the nearest declaration's kind + wins** (proximity, §3); its default body (if any) is the inherited fallback. +- **Apply-time checks:** a dangling import (inherited script → missing module) is a `plan` error; an + extension point with no provider in a given app is an error for that app (a hard failure, joining + §4.7). + +> **Residual (verified):** executor-core's `PicloudModuleResolver` is app-scoped today and ignores the +> importing script's origin (`module_resolver.rs` passes `_source` unused). Rhai *does* expose that +> origin, so the lexical-vs-dynamic split is expressible — but it requires re-keying the resolver cache +> by owner identity and adding per-import policy (sealed vs. extension point), i.e. a real +> resolver+cache redesign, not a parameter tweak. Lands with phasing step 4. + +### 5.6 Tree lifecycle: delete, reparent, rename + +Structural mutations of the group tree, which the rest of the design depends on staying acyclic and +non-orphaning: + +- **Delete = RESTRICT, never implicit CASCADE.** Deleting a non-empty group is refused — an implicit + cascade would destroy descendant apps and their isolated data. CLI `--recursive` expands a delete + into *ordered, explicit, confirmed* child deletions; the DB FK stays RESTRICT. Corollary: a + referenced ancestor **cannot vanish while it has descendants**, so cross-repo read-only references + (§7) can't be orphaned by deletion — the RESTRICT protects them automatically. **App *data* is + destroyed only on explicit opt-in:** an app delete refuses unless `--purge-data`, which then removes + its KV/docs rows *and* its files blob tree under `PICLOUD_FILES_ROOT//` — a non-DB, + non-undoable effect run outside the transaction and logged. So `--recursive` group delete requires + `--purge-data` to touch any descendant app's data; without it, a non-empty app blocks the delete. +- **Reparent / rename: the slug is frozen at creation.** The path only *seeds* the derived slug + (§4.4); a move or rename updates the **display path** but never rewrites the **instance-global + slug** — the deployment key stays stable, external references don't break. After a move the slug no + longer mirrors the path (cosmetic, accepted). +- **Reparent recomputes descendant effective views** (it changes the resolution chain — the same + fan-out invalidation as a node write, §5.1) and is **doubly capability-gated**: group-admin at + *both* the source and destination parent (you remove from one ancestor's domain and add to + another's). Because it changes the resolution chain, a reparent is **validated like a plan and + refused (unless forced)** if the recompute would orphan a sparse `enabled`-override (now shadowing + nothing) or leave an extension point with no provider (§5.5) — a structural move must not silently + produce a state `apply` would have rejected. +- **Cycle guard, under the apply lock.** Reparent runs an **ancestor-walk check in manager-core** + (walk from the destination up to root; reject if it reaches the node being moved). A Postgres + `CHECK` can't express this; the guard is what guarantees §9's "resolution always terminates." + Single-parent + this guard = acyclic. **All structural mutations (reparent/rename/delete) take the + same coarse apply-lock (§4.2)**, so the ancestor-walk + `parent_id` write run serialized — two + concurrent reparents can't race into a cycle, and a reparent's view-recompute can't collide with an + overlapping apply on the materialized-view store. + +## 6. The CLI ↔ server projection + +- **Directories = groups** (the hierarchy axis). Single-parent falls out of the filesystem for free. +- **Overlay files = environments/apps** (the deployment-variant axis) — *not* subdirectories, + because envs share scripts/structure and only diverge on vars/secrets/slug; files structurally + prevent per-env script drift. +- **`scripts/` at every level cascades** up the tree; nearer overrides farther by name (CoW). +- **A leaf group = one logical app; its environments = the actual server apps.** Multiple distinct + apps = multiple sibling leaf groups. **Intermediate groups may also bear apps** (a dir may have + both subdirs and overlay files) — allowed, no special-casing. +- **Environment registry lives at the app-bearing node**, but **confirm-policy is inheritable** (set + "production always confirms" once at root; it flows down via the §3 mechanism). Leaves may declare + independent environment sets; a tree-wide `pic apply --env production` simply skips a leaf that has + no `production`. +- **Attach point:** the local root manifest declares where it binds into the server tree + (`parent_group = "acme"` or instance root). Ancestors above it are inherited/referenced but **not + present locally** — which makes the sparse checkout enforce the RBAC masking for free; effective + config needs a server round-trip (`pic config --effective`). +- **Stable IDs in gitignored `.picloud/`** (group IDs, instance URL, token ref) so a directory + rename/move maps to a server **reparent**, not delete+create. +- **Local/server structural divergence is detected, not silently fought.** Alongside the per-node + content version (§4.2), each node carries a **per-subtree structure version** (covering its own + parentage/subtree — **not one global counter**, so a structural edit in one team's subtree never + force-refuses an unrelated repo's plan). `pic plan` compares the local parent-by-ID (from + `.picloud/`) against the server's; on a structural mismatch (someone reparented server-side, or a + dir moved locally) it **refuses**, requiring an explicit `--adopt-server-structure` or + `--force-local-structure` (Terraform's detect-and-refuse on stale state). The content-version check + alone would miss a pure structural move; reparent/rename/delete each bump the affected subtree's + structure version. +- **Mono-repo** = attach at instance root. **Per-team repo** = attach at a subgroup, contain only its + slice. Same model, different attach depth. + +--- + +## 7. Ownership of shared nodes + +**Single-owner-per-node, ceilinged by the attach point.** + +1. Each group node is owned by exactly one **project-root** (the repo that *manages* it — contains + its manifest as something it applies, not merely references). The server records `owner_project`. +2. **Your attach point is your ceiling** — you cannot apply to anything above your local root. +3. **First apply claims; transfer is explicit and capability-gated.** A second repo applying to an + owned node is rejected (`owned by project X; use --takeover`); takeover needs group-admin + capability. This stops one team silently clobbering org-wide config — whose blast radius (via the + effective-view fan-out) is the whole subtree. +4. **Ownership ⟂ RBAC.** Ownership = *which manifest is authoritative*; RBAC = *whether this + principal may*. The owner still needs group-admin capability to apply. +5. A node with **no project claim is UI/API-owned** — the dashboard is its source of truth and no + manifest fights it. Every node is either manifest-owned (one repo) or UI-owned. + +**Corollary:** don't co-own a node — split config downward. Shared config lives *higher* (owned by a +platform/shared repo attaching at root); team-specific bits go into subgroups each team owns. + +--- + +## 8. Diagrams + +### 8.1 Server ownership & containment + +Only scripts/vars/secret-refs are group-ownable (polymorphic owner); routes/triggers/data are always +app-owned via `app_id`. + +```mermaid +graph TD + INST([Instance]) + INST --> RG["Group: acme (root)"] + RG --> SGA["Group: team-a"] + RG --> SGB["Group: team-b"] + SGA --> LGA["Group: blog (leaf)"] + SGA --> LGB["Group: shop (leaf)"] + LGA --> APP1["App: blog-staging"] + LGA --> APP2["App: blog-production"] + + RG -.->|"owns (shared)"| D0["scripts, vars, secret-refs"] + SGA -.->|owns| D1["team-a scripts, vars"] + APP2 -.->|"owns / overrides"| D2["app scripts, vars, secrets"] + + APP1 ==>|app_id| C1[("routes, triggers, KV/Docs/Files, Topics")] + APP2 ==>|app_id| C2[("routes, triggers, KV/Docs/Files, Topics")] +``` + +### 8.2 Entity identity & cardinalities + +```mermaid +erDiagram + GROUP ||--o{ GROUP : "parent-of (single parent)" + GROUP ||--o{ APP : contains + GROUP ||--o{ DEFINITION : "owns (scripts/vars/secret-refs)" + APP ||--o{ DEFINITION : "owns (override)" + APP ||--o{ APPSCOPED : "owns (routes/triggers/data, app_id)" + GROUP ||--o{ MEMBERSHIP : "role (inherited down)" + APP ||--o{ MEMBERSHIP : "role" +``` + +### 8.3 Config resolution (sparse merge, env filter, proximity-first) + +Effective value for one app+env, resolved by §3. Env-scope filters per level; nearest level wins; +maps merge per key. + +```mermaid +graph TB + I["Instance defaults"] --> RG["Group acme
secret stripe_key (ref)
script auth.rhai
db_url@production"] + RG --> SG["Group team-a
var region = eu"] + SG --> LG["Group blog
script render.rhai
var title = Blog"] + LG --> EV["Env overlay: production
var title = Blog PROD (override)
secret stripe_key = prod value"] + EV --> EFF[["Materialized effective view (app=blog-production):
auth.rhai (acme), render.rhai (blog)
title = Blog PROD (leaf overlay)
region = eu (team-a)
db_url = prod (acme @production filter)
stripe_key = prod"]] +``` + +### 8.4 Filesystem ↔ server mapping + +```mermaid +graph LR + subgraph FS["Git working tree"] + direction TB + R["acme/
picloud.toml
scripts/"] + R --> TA["team-a/
picloud.toml
scripts/"] + TA --> BL["blog/
picloud.toml (base)
picloud.staging.toml
picloud.production.toml
scripts/render.rhai"] + LINK[".picloud/ (gitignored)
group IDs, instance URL, token ref"] + end + subgraph SRV["PiCloud server"] + direction TB + G0["Group acme"] --> G1["Group team-a"] + G1 --> G2["Group blog"] + G2 --> A1["App blog-staging"] + G2 --> A2["App blog-production"] + end + R -.->|defines| G0 + TA -.->|defines| G1 + BL -.->|"base defines"| G2 + BL -.->|"staging.toml"| A1 + BL -.->|"production.toml"| A2 + LINK -.->|"stable IDs"| SRV +``` + +### 8.5 Multi-repo subtree views & single-owner ownership + +```mermaid +graph TD + subgraph Server["Server group tree (authoritative)"] + acme["acme (root)"] + acme --> ta["team-a"] + acme --> tb["team-b"] + ta --> blog["blog"] + tb --> shop["shop"] + end + subgraph PR["Repo: platform"] + pr["manages acme"] + end + subgraph AR["Repo: team-a"] + ar["attaches at acme
manages team-a + blog"] + end + pr ==>|owns| acme + ar ==>|owns| ta + ar ==>|owns| blog + ar -.->|"reference, read-only"| acme +``` + +### 8.6 Apply pipeline (bound plan → DB-atomic write → convergent propagation) + +```mermaid +sequenceDiagram + actor Dev + participant CLI as pic CLI + participant Mgr as manager-core (single writer) + participant DB as Postgres + participant View as effective views + caches + Dev->>CLI: pic plan --env production + CLI->>CLI: read manifests + .rhai sources, build subtree bundle + CLI->>Mgr: send bundle (whole subtree) + Mgr->>DB: read state + version of affected nodes + Mgr->>Mgr: diff + blast radius + persist plan artifact + Mgr-->>CLI: PLAN (changes, N descendant apps incl. other repos, conflicts) + CLI-->>Dev: show plan; confirm/approve per env policy + Dev->>CLI: pic apply (executes stored plan) + CLI->>Mgr: apply (plan id) + Mgr->>DB: refuse if content or structure version moved; else ONE transaction (all-or-nothing) + Mgr->>View: recompute effective views + bump generation + invalidate + Mgr-->>CLI: change report (created/updated/re-enabled/pruned/conflicts) + CLI-->>Dev: log changes +``` + +### 8.7 RBAC: masked group secrets + +```mermaid +graph TD + GA["Group admin (team-a)"] -->|"sets + can read"| GS["Group secret: stripe_key
owned by team-a, encrypted at rest"] + AD["App developer (blog)"] -- "cannot read value (sees exists)" --x GS + AD -->|"can edit"| AS["App script source + app vars"] + GS ==>|"runtime injects plaintext"| EX["Executor: running app script"] + AS --> EX + GATE["Two gates:
human-read authz vs runtime resolution"] + GATE -.-> GS + GATE -.-> EX +``` + +### 8.8 The three-state `enabled` lifecycle + +```mermaid +stateDiagram-v2 + [*] --> Active: declared (enabled=true/omitted) + Active --> Disabled: set enabled=false (manifest or UI toggle) + Disabled --> Active: set enabled=true (manifest or UI toggle) + Active --> Pruned: removed from manifest + prune/--prune + Disabled --> Pruned: removed from manifest + prune/--prune + Pruned --> [*] + note right of Disabled + Still desired state, NOT pruned. + Route 404s, trigger inert, script not invocable. + Out-of-band toggle on this field is conflict-guarded (4.2). + end note +``` + +--- + +## 9. Adoption & backfill + +Groups land onto a live instance with existing flat apps, so a migration is a prerequisite, not an +afterthought: + +- Create a **root group** (and/or a per-owner personal namespace, GitLab-style) and **reparent every + existing app** under it. Every app must have a parent from day one so resolution always terminates. +- Existing apps have no group-owned definitions, so their effective view = their own rows — the + materialized-view store can be backfilled trivially (identity resolution). +- The trigger `name` backfill (§4.5) runs in the same migration window. +- Existing app slugs are already instance-global, so no slug rewrite is needed; the path is new + metadata layered on top. + +--- + +## 10. Open questions & residual risks + +Resolved items now live inline next to their topic. What genuinely remains: + +- **Effective-view invalidation SLA (§5.1)** — the security-staleness guarantee at fan-out to many + descendants is unsolved; synchronous-for-security + eventual-elsewhere is the proposed shape, not a + proven one. Highest-risk open item. +- **Conflict bit vs. operational lock (§4.2)** — *decided:* phase 1 ships the `enabled` / secret- + reference **conflict bit**; the operational-lock flag is the documented fallback if the bit proves + too heavy. (Was listed as undecided; resolved here to match §11 phase 1.) +- **Multi-level env-scope precedence (§3.2)** — *decided default:* proximity-first with env as a + per-level filter. The open part is only *validation at depth*, which is why `pic config --effective + --explain` is a **phase-3 hard requirement** (when multi-level resolution first ships), not a + precondition to adopting the rule. +- **Inherited-membership revocation lag (§5.3)** — revoking a group admin must drop implicit admin on + every descendant app, but §5.1's synchronous-invalidation subset covers only `enabled`/secrets, not + **role revocation** — leaving an unbounded window. New residual risk; should get the same + synchronous-for-security bound, lands with phase 3's authz. +- **External execution cancel (§4.2 kill-switch)** — the executor has **no external-cancel path + today** (`spawn_blocking` + self-checked deadline, verified); the kill-switch is a net-new capability + (a cancel flag polled in `on_progress`), deferred past phase 1. Until it exists, the strongest stop + is op-budget/deadline + the dispatcher fire-time `enabled` re-check (§4.3) for the trigger path. +- **Narrow-inheritance vs. trigger/route templates (§4.5, §5.1)** — the per-app binding tax bites + early at high tenant cardinality. Decide whether templates are truly deferrable for your target. +- **`pull --factor`** — auto-extract a shared base by diffing two pulled envs (later nicety). + +--- + +## 11. Suggested phasing + +1. **Declarative project tool, single-app (no groups yet).** `init`, `pull`/`config --effective`, + manifest parse/validate, `plan` (bound artifact), `apply` (**atomic desired-state write — requires + the manager-core post-commit-refresh restructuring of §4.2, domains/files excluded from the + transactional core**), `prune`, secrets push, link state, env-scoped config. Adds `enabled` to + scripts/routes + the three-state runtime + the dispatcher fire-time `enabled` re-check (§4.3) + + trigger `name` column/backfill + the `enabled`/secrets conflict bit + the net-new content + + tree-structure version counters + a coarse per-instance apply lock; in-flight executions finish (no + kill). +2. **Groups as pure org/RBAC/UI container.** Nested groups (single-parent, `parent_id`, **delete = + RESTRICT**, reparent/rename with the **ancestor-walk cycle guard** + **slug-freeze** + + **tree-structure version**, §5.6), inherited membership, hierarchy-aware `can`, UI grouping, the §9 + backfill. No shared resources yet — cheap, no data-plane schema change. +3. **Group-inherited config** (vars, secret-refs, env-scoped). The net-new `vars`/`secret-refs` + tables + polymorphic owner; the group-secret AAD scheme (§5.3 caveat); masked group secrets; the + effective-view resolver + materialization + invalidation; **`config --effective --explain`** (hard + requirement, since multi-level resolution first ships here). +4. **Group-inherited scripts/modules.** CoW overrides; the **scope-aware module/import resolver + + extension points** (§5.5); cache-invalidation fan-out hardening; versioning/pinning if needed. +5. **Project tool maps onto groups.** Nested manifests, attach point, single-owner, server-computed + tree plan, per-env approval gating. +6. **(Much later) group-level collections/topics** — the v1.3 cross-app data-sharing problem, with a + real shared-scope authz model. Optionally, trigger/route **templates** (§4.5) if cardinality + demands. + +--- + +## 12. Contracts still to draft + +- The **apply bundle / plan artifact / change-report** wire contract (what the CLI ships, what the + server persists and returns), including the conflict and blast-radius shapes. +- The **effective-view resolver** (the read primitive) — the §3 rule made executable, plus the + materialization + invalidation protocol (§5.1). +- The **full manifest schema** spelling every block (scripts, routes, the 8 trigger kinds, storage + config, env-scoped vars, secret-refs, domains, `[project.environments]` + confirm policy).