Files
PiCloud/docs/design/groups-and-project-tool.md
MechaCat02 c600177fd6 docs(design): groups & declarative project-tool spec
The architecture spec for the full vision — a declarative project tool
(picloud.toml + pull/plan/apply/prune) and a server-side nested-groups
hierarchy with live-resolve inheritance. Reviewed across several passes
for consistency, gaps, and contradictions; resolutions are folded in
beside the topics they settle.

This commit lands the spec only. The first slice of implementation (the
single-app reconcile foundation) follows; groups, env-scoping, and the
`enabled` toggle are later milestones called out in the doc.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 21:52:01 +02:00

850 lines
52 KiB
Markdown

# Groups, Projects & the Declarative Project Tool
> **Status:** Draft — design discussion captured 2026-06-18, revised 2026-06-20 with resolutions
> grounded in precedent (GitLab, Kustomize, Helm, Terraform, Kubernetes Server-Side Apply, Pulumi)
> and corrected against the codebase after three independent review passes (consistency, gaps,
> feasibility). Not yet scheduled into a phase.
>
> **Scope:** turns the imperative `pic` CLI into a declarative, file-based project tool, and
> introduces a server-side **groups** hierarchy so apps can share and inherit config/scripts
> without duplication.
>
> **Blueprint impact:** this reverses the §11.5 *snapshot-copy, not live-link* stance — top-down
> hierarchical inheritance (group → app) is now in scope, implemented as a *materialized, auto-
> invalidated* view (§5.1). Fold the outcome back into the blueprint before it drifts.
---
## 1. Motivation
Today `pic` is **imperative**: every command (`pic deploy`, `pic routes create`, `pic triggers
create-cron`, …) maps 1:1 to an admin API call. There is no memory of desired state, no project
file, and no way to express "these N apps are the staging/prod/tenant variants of one thing."
Two gaps follow:
1. **No declarative project layer.** Developers re-run commands by hand; nothing reconciles a repo
to an instance.
2. **No server-side notion of a project/group.** If we model environments as separate apps, the
*shared* base (config, scripts) is duplicated across every deployed app. The server sees N
unrelated apps.
This document designs both halves: a **declarative manifest + project tool**, and a **groups
hierarchy** on the server that the project tool projects onto the filesystem.
`manager-core` is the single writer to Postgres. We lean on that for two concrete things — an
**atomic desired-state write** (one DB transaction per apply, §4.2) and a **server-computed diff**
(§4.2) — *not* for continuous reconciliation, which we deliberately decline (drift handling is
detect-and-surface, §4.2). This is a narrower and more honest claim than "we can reconcile live
state": most comparable tools cannot do an all-or-nothing apply because their effects are
un-undoable cloud resources; ours are Postgres rows, so we can — within the boundary described in
§4.2.
---
## 2. Vocabulary
| Term | Meaning |
|---|---|
| **Group** | Server-side hierarchy node. Single-parent tree. Owns shared **code/config** definitions and members/roles. Nestable. |
| **App** | A deployable unit. The data-isolation boundary (`app_id`). Lives under a group. |
| **Environment** | A deployment variant (staging/production/…). **An environment *is* an app.** Also a *scope dimension* on config (§3). |
| **Tenant** | Modeled like an environment — a scope dimension / overlay axis, **not** a second parent (§5.4). |
| **Project** | *Not* a server entity. A CLI view over a repo-managed **subtree** of groups. |
| **Definition** | Entities that can be group-owned: **scripts/modules, vars, secret-refs only**. Routes, triggers, collections, topics are **app-scoped** (§3, §5.1). |
| **Data** | KV/docs/files collections + pub/sub topics. **Always app-owned** (`app_id`). The isolation boundary never moves. |
| **Manifest** | TOML file describing desired state for one node (group base) + per-env overlays. |
| **Effective view** | The materialized, per-app resolved set of definitions, served to the runtime and recomputed on any ancestor change (§5.1). |
| **Attach point** | Where a local working tree's root binds into the server group tree. Its ceiling — you cannot apply above it. |
Relationship: **the base + per-env-overlay "project" is the degenerate one-group subtree** of the
general groups model. Nesting generalizes it to multi-group subtrees. Nothing about the
single-project model is lost.
---
## 3. Configuration resolution
This is the single rule the whole design hangs off; everything else (env vars, secrets, `enabled`,
overrides) resolves through it. **All three precedent systems converge on the same shape, so we
adopt it wholesale:**
> ⚠️ **Net-new, not an extension (verified against the codebase).** PiCloud has **no env-agnostic
> config / `vars` layer today** — this resolution engine and the `vars` table are greenfield. Only
> `secrets` exists, and as a value-bearing table, not a ref. Read §3 as *new* infrastructure, not a
> modification of something present.
> **Resolution = sparse, per-field merge; environment is a *pre-filter*, not a precedence tier;
> proximity wins across levels.**
To resolve a key for an app in environment `E`:
1. **Filter by environment first, per level.** A value scoped `@E` is eligible; a value scoped `*`
(env-agnostic) is the fallback. At a *single* level, `@E` beats `*`. Environment scope decides
*eligibility*, it is **not** a competing precedence rank.
*Evidence:* GitLab CI/CD variables use `environment_scope` exactly this way — a `production`-scoped
row simply isn't visible to a `staging` job ([GitLab CI/CD variables](https://docs.gitlab.com/ci/variables/)).
2. **Then nearest level wins.** Walk up `acme → team-a → blog → app`; the closest level that defines
the (filtered) key wins. GitLab states this verbatim: *"if the same variable name exists in a
group and its subgroups, the job uses the value from the closest subgroup."*
3. **Merge granularity:** maps/vars deep-merge **per key** (set `title`, still inherit `region`);
entities (scripts) replace **by identity** (name); deletion is an explicit tombstone.
*Evidence:* Kustomize strategic-merge patches are sparse and deep-merge maps but **replace lists
unless keyed** ([Kustomize patches](https://kubectl.docs.kubernetes.io/references/kustomize/kustomization/patches/));
Helm deep-merges nested maps and uses `null` to delete a defaulted key
([Helm values](https://helm.sh/docs/chart_template_guide/values_files/)).
This one rule resolves several issues at once: it makes **group config environment-scopable** (a
group sets `db_url@production` and `db_url@staging`; descendants inherit the right one — no per-leaf
duplication), it defines **merge granularity** (maps per-key, entities per-identity), and it makes
**`enabled` just another sparse field** (§4.3).
### 3.1 The env-scope manifest syntax
```toml
# group manifest (e.g. team-a/picloud.toml)
[vars]
region = "eu" # env-agnostic (scope *)
[vars.production] # scope @production
db_url = "postgres://prod/..."
[vars.staging]
db_url = "postgres://staging/..."
[secrets]
names = ["stripe_key"] # name-only; values pushed per env via CLI (§4.6)
```
### 3.2 The one deliberately-novel precedence call
With `db_url@production` at the root **and** a plain `db_url` at the leaf, the **leaf's env-agnostic
value wins** for a production app — proximity beats farther-level env-specificity. The research was
explicit that **no major system lets env-specificity override proximity across levels**; GitLab
structurally avoids the question by treating env as a per-level filter (above). So **proximity-first
is the evidence-backed default**, and we document it as a chosen rule: *inner scope shadows outer,
like lexical scoping.* To avoid this becoming a debugging nightmare at depth, `pic config
--effective --explain` must show **why** a value resolved (which level/scope it came from).
> **Residual risk (carried, not solved):** multi-level + environment scopes is genuinely novel
> territory — GitLab is two-level (group→project). Proximity-first at arbitrary depth is consistent
> and defensible but not battle-tested. The `--explain` tooling is a hard requirement, not a nicety.
---
## 4. Manifest & apply
### 4.1 Manifest & layering
- **TOML, data not code.** PiCloud runs untrusted scripts; the deployment descriptor stays inert,
diffable, server-validatable, and dashboard-renderable. Turing-completeness belongs in the Rhai
script, never the manifest.
- **Base + overlays.** A node's `picloud.toml` holds the shared base; `picloud.<env>.toml` overlays
add per-env vars/secrets/slug. Shared config is written once; overlays only carry deltas (the
Kustomize base+overlay model — overlays are sparse patches, not replacements).
- **Routes are top-level**, referencing scripts by name (one place to see all routing).
- **`pull` is first-class** — see §4.6 for its inheritance/masking semantics.
### 4.2 The apply engine
The IaC research reshaped this section materially; the relevant precedents are cited inline.
**Plan/apply split with a bound plan artifact.** `pic plan` computes a reviewable diff
(`create / update / delete / replace`), the server persists it as a row, and `apply` executes
*exactly that stored plan***detecting and refusing if state moved underneath it** (state
version/serial check, cheap given the single writer). *Evidence:* Terraform's saved plan
*"reliably perform[s] an exact set of pre-approved changes, even if the configuration or state has
changed in the minutes since"* ([Terraform run](https://developer.hashicorp.com/terraform/cloud-docs/run)).
The "state moved" check covers **both** the per-node **content** version **and** the **tree-structure**
version (§6): a reparent or a new app created under the subtree between `plan` and `apply` changes the
true blast radius, so the bound plan must refuse if *either* counter moved. Both are **net-new** — the
existing `scripts.version` column is an unconditional write counter, **not** a compare-and-set guard,
and must not be mistaken for one.
**Atomicity — the corrected position.** No major IaC tool does transactional apply, *because their
effects are un-undoable cloud resources* (you cannot un-create a VM). **That constraint does not
bind us:** a PiCloud manifest apply is **almost entirely Postgres row writes** (script source,
route/trigger defs, config). Therefore:
- **The desired-state write is a single DB transaction** — all-or-nothing across the subtree. This
is achievable *because* our substrate is one transactional store (unlike Terraform's un-undoable
cloud resources). It removes the downstream-breakage hazard of an earlier draft (which proposed
per-node, stop-on-error apply — **now superseded**): there is no intermediate state where a parent
committed and a child did not.
- **Propagation is forward-convergent, not transactional.** The post-commit effective-view refresh +
cache invalidation (§5.1) live *outside* the transaction. So "atomic" means
**atomic-for-desired-state, eventually-consistent-for-effect**, with a bounded window where the DB
says new and a cache says old.
> **Invariant to enforce — and it does NOT hold today (verified).** The single-transaction model
> requires apply to write **only Postgres**, but the *current* admin write paths violate that: route
> and domain writes prime in-process caches **synchronously, interleaved with the write**
> (`route_admin.rs:234`, `apps_api.rs:396`), and files fsync to disk. So moving to "commit the
> transaction, *then* refresh views" is a deliberate **restructuring of manager-core**, not a
> description of the status quo. Additionally, **domains** are a non-DB effect (Caddy/filesystem, out
> of band) — they must be **excluded from the transactional core** and handled by the convergence
> model below. Treat "apply writes only Postgres" as an invariant to *establish and guard*, not one
> we already have.
**Idempotent, convergent recovery (for the non-transactional propagation and any future side
effects).** Operations are idempotent upserts keyed by stable identity; "re-apply converges" is the
recovery story; "rollback" means revert the declaration and re-apply forward. Per-node status
(applied/pending/failed/drifted) is recorded so a partial propagation failure is observable and
re-runnable. *Evidence:* every surveyed tool (Terraform, ArgoCD, Flux, Pulumi) chose idempotent
forward-only convergence over distributed rollback.
**Drift model — upgraded from pure last-write-wins.** An earlier draft chose model "(c)":
last-write-wins, just log. That is *exactly* the pre-SSA `kubectl apply` footgun — KEP-555 lists the
scenario verbatim: *"User does an apply, then `kubectl edit`, then applies again: surprise!"*
Kubernetes Server-Side Apply exists to convert that silent clobber into an explicit conflict
([K8s SSA](https://kubernetes.io/docs/reference/using-api/server-side-apply/)). Concrete risk for us:
CI runs `pic apply` and silently re-enables a route an operator killed in an emergency. Resolution:
- **Keep (c)'s simplicity for ordinary fields** (last-write-wins, change reported).
- **For the security-relevant subset only** — `enabled` and a secret's *reference/existence in
desired state* — track a "changed out-of-band" bit. If that subset was changed outside the manifest
(e.g. an operator disabled a route in the dashboard), `pic plan` surfaces it as a **labeled
conflict** and `apply` **refuses without `--force`**. This is a scoped version of SSA's field
ownership; we deliberately avoid full per-field managers (object bloat, complexity).
- **Secret *values* are explicitly out of scope of the conflict/state-version machinery.** Values
are never in the manifest or the plan (§4.6), so a routine `pic secret set` between `plan` and
`apply` is the *supported* workflow and does **not** trip the state-version refusal — the
state-version check covers manifest-managed *desired state* (definitions, including secret
*references* and `enabled`), not value rotation.
- Out-of-band changes render as a **distinct labeled diff** in `plan`, separate from intended
changes. *Evidence:* Terraform/Pulumi both make read-only drift detection default and label drift
distinctly ([Pulumi drift](https://www.pulumi.com/docs/iac/operations/stack-management/drift/)).
> **Residual risk:** the conflict bit reintroduces some of the complexity (c) was chosen to avoid.
> The cruder fallback, if even that is too much: keep pure (c) but add a non-revertible
> **operational lock** flag an operator sets in an emergency that `apply` won't touch. Either closes
> the silent-revert hole; neither is free.
**Gating high-stakes applies: trigger ≠ authorization.** Any CI trigger may `plan`, but applying to
a confirm-required env is a separate, default-off, explicitly-authorized step. A blanket `--yes`
covers ordinary confirms; **confirm-required envs require an explicit per-env `--approve <env>`**, so
CI must opt in per environment. "Override a gate" is its own audited capability (maps onto
`manager-core::authz::can`). *Evidence:* Terraform Cloud parks runs in *Needs Confirmation* and
separates the *apply runs* permission from *manage policy overrides*
([TFC run states](https://developer.hashicorp.com/terraform/cloud-docs/run/states)).
**Concurrency.** Start with a **coarse per-instance (or per-root-group) apply lock** — one apply at a
time — which is trivially correct for the single-node MVP and makes "last-commit-wins" hold. Refine
*later* to **per-blast-radius advisory locks** (lock the affected apps in `app_id` order so
overlapping radii serialize while disjoint ones proceed; queue triggers already use this advisory-lock
primitive) only if contention appears. Don't build the fine-grained version speculatively.
**Blast radius — defined and bounded.** The blast radius is **the set of descendant apps whose
materialized effective view would actually change** — a *diff*, not "all descendants." Per changed
definition at node N it is `subtree(N)` minus apps that override that key nearer. `pic plan`
enumerates it for small radii; for large ones it **summarizes (count + sample)** and a threshold
triggers extra confirmation (`"this changes 4,213 apps — confirm"`). Because only scripts/vars/secret-
refs inherit (§5.1), blast radius applies to *those* changes; an app-scoped change (a route or
trigger) has blast radius = that single app. Root-level changes are accepted as expensive, rare, and
high-confirmation; the computation is bounded by `subtree(N)` size, and the confirmed radius is
re-validated at apply against the tree-structure version (above) so it can't go stale.
**In-flight executions.** An apply that disables or replaces a script affects **new invocations
immediately** (via the `enabled` re-check + view invalidation; pending trigger outbox rows are dropped
at fire-time, §4.3 — so no separate outbox purge is needed). **In-flight *running* executions run to
completion**, and note (verified) the executor has **no external-cancel path today**: executions are
`spawn_blocking` Rhai calls interruptible only by their operation budget or a pre-set wall-clock
deadline self-checked in `engine.on_progress`. A true must-stop-now **kill-switch** is therefore a
*net-new* capability — buildable by having that same `on_progress` hook also poll a per-execution
cancel flag — **gated by an admin capability (`authz::can`) and audited**, scheduled as a later item,
not phase 1. Killing mid-run also risks partial side effects, so it stays the explicit extreme, never
default apply behavior.
**Apply flow (end to end):**
```
CLI builds bundle (manifests + script sources) for the subtree
→ server computes plan + blast radius (incl. descendant apps in OTHER repos)
→ server persists plan artifact; checks state version
→ dev reviews; confirm/approve per env policy
→ server applies desired state in ONE DB transaction (all-or-nothing)
→ server refreshes effective views + bumps generation + invalidates caches (convergent)
→ server returns change report; CLI logs created / updated / re-enabled / pruned / conflicts
```
### 4.3 The three-state `enabled` lifecycle
`enabled` is a real platform feature (DB + server runtime + UI badge/toggle), not CLI sugar, and — by
§3 — it is **just another sparse, proximity-resolved field**.
| State | Meaning | Pruned? |
|---|---|---|
| declared, `enabled = true` (or omitted) | deployed, active | kept |
| declared, `enabled = false` | deployed but **inert** (route short-circuits, trigger doesn't fire, script not invocable) | **kept** — still desired state |
| absent from merged manifest | stale | **deleted** by prune / `--prune` |
- Default `true`; last-write-wins on merge; a base `enabled = false` is inherited until an overlay
(env *or* a nearer group/app) explicitly sets `enabled = true`. Overriding a *different* field does
**not** implicitly re-enable.
- **Across the group axis (resolved):** a descendant disables an inherited (group-owned) script via a
**sparse override** — an override row that sets only `enabled = false`, inheriting the source. A
descendant can re-enable a parent-disabled entity because nearest-level wins. This is the same
sparse-field mechanism as everything else (Kustomize patches are sparse — you specify only what
changes).
- **Base = the superset across envs.** Per-env you may toggle `enabled` in *either* direction
(disable a base-active entity, or re-enable a base-disabled one — proximity wins, §3), but you
**cannot *remove* a base entity** in one env; true removal requires the entity not be in the base.
An entity unique to one env goes in that overlay; an entity present everywhere but off in staging
goes in base with `enabled = false` in the staging overlay.
- **Disabled = invisible.** External callers hitting a disabled route get **404** (indistinguishable
from absent — no info leak).
- **Schema note:** the `triggers` table **already** has `enabled` (+ `dispatch_mode`, retry columns)
and it is **honored at match/schedule time** (`trigger_repo.rs`, `cron_scheduler.rs` — verified).
New work is `enabled` on **scripts** and **routes** only, plus runtime honoring in the matcher /
invoker.
- **Outbox fire-time gap (verified):** the dispatcher does **not** re-check `enabled` on an
already-enqueued outbox row (`dispatcher.rs:699`), so disabling a trigger stops *new* matches while
a pending item still fires. **Fix:** add an `enabled` re-check (trigger *and* script) at fire time
in `resolve_trigger`, so a pending outbox row for a now-disabled trigger/script is dropped when it
comes up — closing the gap cheaply, with no cache or kill-switch involvement. This is the home for
the §5.1 security-disable guarantee on the trigger path.
- **Provenance caveat (accepted):** a single boolean carries no "manual vs manifest" provenance, so a
disabled entity looks the same however it got there — which is *why* the §4.2 conflict bit on the
`enabled`/secrets subset exists, to stop the next apply silently reverting an operational disable.
### 4.4 Identity & naming
- **Kebab everywhere.** One canonical identifier regex: `^[a-z0-9][a-z0-9-]{0,62}$` for project
names, env names, script names/slugs, trigger names — unified with the existing app-slug rule.
- **Scripts:** unique `name`/`slug` per app = merge/upsert key.
- **Routes:** identity = the triple **`<method> <host> <path>`**, e.g.
`ANY *.beta.example.com /hello/:name`. `dispatch_mode`, `host_param_name`, etc. are *attributes*
overridable without changing identity. The CLI infers `host_kind`/`path_kind` from the pattern
syntax (`*`, `{name}`, `:name`, exact), with an explicit `kind` key as override (mirrors the UI).
Needs a normalization rule (default method `ANY`, default host `*`, case-folding) so manifest ↔
server match exactly.
- **Slugs are instance-global, derived from the path.** Two identifiers coexist:
- **path** = `acme/team-a/blog` — hierarchical, group-scoped, display/organization/RBAC.
- **slug** = flat, instance-global, the deployment key.
The derived default is **`{flattened-path}-{env}`** (e.g. `team-a-blog-staging`) — unique by
construction in the multi-group world, unlike a bare `{leaf}-{env}` which collides whenever two
groups reuse a leaf name. On >63-char overflow, truncate + short hash suffix. Explicit override
allowed; the path never *is* the slug, it only *seeds* it.
### 4.5 Triggers
Triggers are **app-scoped, not group-inherited** (§5.1 explains why). Two senses of "app" are in play
and must not be conflated: a trigger is declared once in the **leaf (logical app)** base — an
authoring convenience — and **materializes into per-`app_id` rows**, one per env-app, where `app_id`
is the server-app isolation boundary (§5.2). It is never group-owned.
Two distinct constraints:
- **`name`** (explicit, kebab) = the merge/identity key + upsert target. Unique per app.
*(Triggers have no name column today — new.)*
- **Semantic uniqueness** = a *post-merge validation*: no two triggers may share their kind-specific
semantic key. Checked after merging, so overriding a base trigger per-env (reuse `name`, change a
field) is fine; two differently-named triggers with identical effect is an error.
| Kind | Semantic key | Note |
|---|---|---|
| kv / docs / files | `(script, collection_glob, ops)` | **canonicalize `ops`** (sort + dedupe) |
| cron | `(script, schedule, timezone)` | exact TEXT match |
| pubsub | `(script, topic_pattern)` | |
| dead-letter | `(script, source_filter, trigger_id_filter, script_id_filter)` | NULL = literal value (not wildcard) for equality |
| email | `(script)` | no filters — one per script |
| **queue** | `(queue_name)`**not** script-scoped | already server-enforced (advisory lock: one consumer per `(app_id, queue_name)`) |
- **Matching vs identity:** `[]`/`NULL`/globs mean **"any/wildcard" at dispatch time** (unchanged
runtime matching) but are compared as **literal structural values for dedup**. We dedup on
**structural identity, never on overlap/subsumption**`ops = []` and `ops = ["insert"]` are *not*
duplicates. Overlapping triggers coexist (multiple triggers firing on one event is already how
PiCloud works).
- **Name backfill** for existing nameless rows: `{kind}-{entity}-{n}`, where `{entity}` is the kind's
identity token (collection / queue / topic; for cron/email/dead-letter fall back to `{kind}-{n}`),
sanitized to the kebab regex, with `-{n}` (discovery order) guaranteeing per-app uniqueness.
> **Ergonomic debt (accepted, watch it):** because triggers don't inherit, 100 tenant apps each
> needing the same 5 triggers = 500 declarations. The fix is group trigger/route **templates** that
> fan out per descendant (a *template/instantiation* mechanism, not inheritance) — deferred, but it
> bites early if tenant cardinality is high. Pressure-test against real tenant counts before
> committing to the narrow-inheritance choice (§5.1).
### 4.6 Secrets & `pull`
- **Name-only in the manifest; value pushed via CLI** (`pic secret set`, reads stdin). Unanimous
across every comparable tool — never commit secret values.
- **Env-scoped** like any var (`[secrets]` names declared once; values set per env).
- **Warn (don't block) if a referenced secret is not set.** This requires an app dev to see that a
group secret **exists / is set** (a boolean) without reading its value — an accepted, explicit
authz boundary. GitLab surfaces inherited masked variable *keys* the same way.
- **Email's `inbound_secret` is a reference**, not inline — same rule; the server already encrypts it
at rest.
- **`pull` exports own-rows only** (this node's overrides), **never effective/inherited state.** A
separate read-only **`pic config --effective`** shows the inherited result with **masked secrets
rendered as `<set>` / `***`, never plaintext.** This makes pull-under-masking safe *by
construction* — you cannot pull a secret you cannot read, and you cannot accidentally duplicate
inherited config into a leaf. *Evidence:* GitLab shows inherited group variables in a project
read-only and separate from project-own variables.
- Flat pull for a new project; "smart" delta-pull (own-vs-effective diff) is server-computed since an
app dev's checkout lacks ancestor manifests.
### 4.7 Apply-time warnings
- Enabled route/trigger pointing at a **disabled script**.
- An **endpoint** script deployed with **no route and no trigger** (unreachable). Modules are exempt.
- Sandbox override exceeding the admin **ceiling**.
- Referenced secret not set.
- An out-of-band change to the `enabled`/secrets subset (surfaced as a conflict, §4.2).
---
## 5. Groups & inheritance
### 5.1 What inherits, and the runtime model
- **GitLab-like, nested, single-parent tree.** Single parent keeps inheritance acyclic and
resolution deterministic (no diamond precedence).
- **Inherit code/config — narrowly. Inherit data — no.** Group-inheritable = **scripts/modules,
vars, secret-refs only.** Routes, triggers, collections, topics, files are **app-scoped.**
- *Why narrow:* routes and triggers are **bindings**, not pure code/config. A group-owned trigger
has no app data to watch (triggers fire on app-owned collections/topics/queues); a group-owned
route has no host (routing is Host→app first). Inheriting them is incoherent. There are two
distinct sharing mechanisms and an earlier draft conflated them: the **leaf base+overlay** shares
routes/triggers across *one app's environments*; **group inheritance** shares *code/config across
many apps*. *Evidence:* Serverless/SAM keep `events:` per-function and share logic via *layers*
bindings local, code shared.
- *Why data stays app-owned:* a group script executes in the *inheriting app's* context, so its
`cx.app_id` still scopes data to that app. Group-level *collections/topics* would break `app_id`
as the isolation boundary — that is the v1.3 cross-app data-sharing problem and stays **out**.
- **Runtime model: a materialized effective view + versioned cache (not per-request live-resolve).**
An earlier draft said "live-resolve," which is underspecified and would fight the existing cache.
The real model:
- manager-core **resolves-at-write into a materialized per-app effective view** (§3 rule applied:
sparse merge, env filter, proximity, CoW, `enabled`).
- The orchestrator/executor serve from that view, **keyed by `app_id` + a generation/version**. The
app_id-keyed *shape* survives (today's route cache is already an `app_id`→routes map), **but the
substance is net-new (verified):** there is no generation counter anywhere today, the route cache
is rebuilt whole-table on every write rather than per-app, and script *bodies* are live-resolved
per request. Read "serve from it" as *extend the keying*, not *reuse the mechanism* — and note
that fanning out today's full-rescan invalidation to thousands of descendants would be a
regression, so per-app incremental invalidation is part of the build.
- On any write to a node, manager-core (single writer, knows the tree) **recomputes descendants'
views + bumps the generation + invalidates caches.**
- The materialized view is a **derived cache, not a second source of truth** — canonical config
still lives once at the owning node, so this does **not** reintroduce duplication. This dissolves
the long-running "snapshot vs. live" tension: no duplication *and* propagation *and* a fast hot
path.
> **Residual risk (relocated, not solved):** cache invalidation is now a **correctness/security
> requirement** — disabling a script for a security reason must stop it running *everywhere* within
> bounded time. This is the same class as the existing PrincipalCache revocation lag. Require
> **synchronous invalidation for the security-relevant subset** (`enabled=false`, secret rotation)
> and accept bounded eventual staleness elsewhere; the hard SLA at fan-out to thousands of
> descendants is genuinely unsolved.
### 5.2 Schema impact
- Inheritable definition kinds get a polymorphic owner (`owner_kind ∈ {group, app}` + `owner_id`):
**scripts** (modify the existing table — and note its `app_id` FK is `ON DELETE RESTRICT`, not
CASCADE, so a pruning apply needs explicit ordering), plus **`vars` and `secret-refs`, which are
net-new tables** (they do not exist today — §3). So this is *one table modified + two invented*,
not "three tables touched." **Routes, triggers, collections, topics, files stay strictly
`app_id`-owned**, so the runtime isolation boundary stays fixed. (The CLAUDE.md `app_id NOT NULL …
CASCADE` rule is itself not universal today — scripts already use RESTRICT.)
- New: a `groups` table (single-parent, `parent_id`), group `membership`/roles, an `owner_project`
column on group nodes (§7), and the materialized effective-view store keyed by `app_id` +
generation (§5.1). Env-scoped values carry an `environment_scope` column (`*` or a specific env).
### 5.3 RBAC
- **Hierarchy-aware capabilities.** `authz::can(principal, cap, on=node)` resolves by walking
ancestors and taking the highest effective role. Instance → group(s) → app.
- **Inherited membership** (GitLab-style): a group admin is implicitly admin of every subgroup/app
beneath it.
- **Masked group secrets:** a group secret is *used by* an app at runtime but *not human-readable* by
the app's developers. Two orthogonal gates: **runtime resolution** (engine injects plaintext) vs
**human-read authz** (admin API returns a value only to a principal with rights at the *owning
group*). An app-scoped admin call never returns group secrets; runtime injection bypasses the human
gate. App devs may see a group secret **exists** (for the unset-warning, §4.6) but not its value.
**An app can run with config its own developers cannot see.**
> **Crypto caveat (verified):** secrets today are AES-GCM sealed with AAD `secret:{app_id}:{name}`,
> and the decrypt path hard-codes `cx.app_id` (migration 0042). A **group-owned** secret is *not
> expressible* under this scheme — there is no group identity in the AAD. It needs a new AAD
> identity (e.g. `secret:group:{group_id}:{name}`) **and** an owner-aware decrypt path that resolves
> whether the inherited secret is group- or app-owned. This is the single hardest correctness detail
> of group secrets and gates phasing step 3.
### 5.4 Tenants & single-parent
Single-parent forbids an app combining two *sibling* groups' configs — which seems to threaten the
multi-tenant use case (shared-platform base + per-tenant overlay). It does not, because **the
shared base is an *ancestor*, not a sibling:** model tenants as leaf apps under a `tenants/`
subgroup that inherits the platform base up the chain (tenant leaf → `tenants` group → platform
group). The "two parents" intuition is satisfied by the *chain*. Better still, **a tenant is a scope
dimension like environment** (§3) — the same env-scope machinery generalizes to a `tenant` scope, so
multi-tenant needs no new hierarchy primitive. *Evidence:* GitLab is single-parent and serves
multi-tenant teams via subgroups; "combine two siblings" is handled by promoting shared config to a
common ancestor.
---
### 5.5 Module & import resolution under inheritance
A group-owned script may `import` modules, and inheritance makes "which module?" ambiguous (an
inherited script importing a module a leaf has shadowed could bind to app-dev code it never saw — a
trust inversion). The rule:
- **Lexical by default.** An inherited script's imports resolve against the module set visible at the
**script's own defining node** (walking up from there), **not** the inheriting app's effective
view. A leaf cannot shadow a module an inherited group script depends on, and a group script
behaves identically across every app that inherits it — preserving both **determinism** and the
**trust boundary** (a security-authored shared script is tamper-proof from below). Ordinary lexical
scoping; "sealed by default", like non-`open` classes in Kotlin / `final` in Java.
- **"Defining node" of a CoW-overridden script = the node that authored *that body*.** An
inherited-unchanged script's defining node is the ancestor that owns it (imports sealed to the
ancestor's modules). A script a leaf *overrides* (CoW, §3) has the **leaf** as its defining node, so
the override's imports resolve from the leaf — which is *not* a trust inversion, because the app
owner wrote that body and is trusted within their own app. Consequence (intended): the same module
name can resolve to **two different bodies** in one app, depending on the importing script's defining
node.
- **Explicit extension points for opt-in polymorphism.** A module is marked an extension point in the
**manifest** (an `[extension_points]` declaration — *not* in Rhai source, keeping code inert and
giving the `plan` checker something to read). Such a module is one descendants are *expected* to
provide or override; **only** these resolve against the inheriting app's effective view. Controlled
template-method customization (a shared `render` whose `theme` module each tenant supplies) without
the blanket trust inversion of dynamic resolution. When a name is declared at more than one level —
or is concrete on one path and an extension point on another — **the nearest declaration's kind
wins** (proximity, §3); its default body (if any) is the inherited fallback.
- **Apply-time checks:** a dangling import (inherited script → missing module) is a `plan` error; an
extension point with no provider in a given app is an error for that app (a hard failure, joining
§4.7).
> **Residual (verified):** executor-core's `PicloudModuleResolver` is app-scoped today and ignores the
> importing script's origin (`module_resolver.rs` passes `_source` unused). Rhai *does* expose that
> origin, so the lexical-vs-dynamic split is expressible — but it requires re-keying the resolver cache
> by owner identity and adding per-import policy (sealed vs. extension point), i.e. a real
> resolver+cache redesign, not a parameter tweak. Lands with phasing step 4.
### 5.6 Tree lifecycle: delete, reparent, rename
Structural mutations of the group tree, which the rest of the design depends on staying acyclic and
non-orphaning:
- **Delete = RESTRICT, never implicit CASCADE.** Deleting a non-empty group is refused — an implicit
cascade would destroy descendant apps and their isolated data. CLI `--recursive` expands a delete
into *ordered, explicit, confirmed* child deletions; the DB FK stays RESTRICT. Corollary: a
referenced ancestor **cannot vanish while it has descendants**, so cross-repo read-only references
(§7) can't be orphaned by deletion — the RESTRICT protects them automatically. **App *data* is
destroyed only on explicit opt-in:** an app delete refuses unless `--purge-data`, which then removes
its KV/docs rows *and* its files blob tree under `PICLOUD_FILES_ROOT/<app_id>/` — a non-DB,
non-undoable effect run outside the transaction and logged. So `--recursive` group delete requires
`--purge-data` to touch any descendant app's data; without it, a non-empty app blocks the delete.
- **Reparent / rename: the slug is frozen at creation.** The path only *seeds* the derived slug
(§4.4); a move or rename updates the **display path** but never rewrites the **instance-global
slug** — the deployment key stays stable, external references don't break. After a move the slug no
longer mirrors the path (cosmetic, accepted).
- **Reparent recomputes descendant effective views** (it changes the resolution chain — the same
fan-out invalidation as a node write, §5.1) and is **doubly capability-gated**: group-admin at
*both* the source and destination parent (you remove from one ancestor's domain and add to
another's). Because it changes the resolution chain, a reparent is **validated like a plan and
refused (unless forced)** if the recompute would orphan a sparse `enabled`-override (now shadowing
nothing) or leave an extension point with no provider (§5.5) — a structural move must not silently
produce a state `apply` would have rejected.
- **Cycle guard, under the apply lock.** Reparent runs an **ancestor-walk check in manager-core**
(walk from the destination up to root; reject if it reaches the node being moved). A Postgres
`CHECK` can't express this; the guard is what guarantees §9's "resolution always terminates."
Single-parent + this guard = acyclic. **All structural mutations (reparent/rename/delete) take the
same coarse apply-lock (§4.2)**, so the ancestor-walk + `parent_id` write run serialized — two
concurrent reparents can't race into a cycle, and a reparent's view-recompute can't collide with an
overlapping apply on the materialized-view store.
## 6. The CLI ↔ server projection
- **Directories = groups** (the hierarchy axis). Single-parent falls out of the filesystem for free.
- **Overlay files = environments/apps** (the deployment-variant axis) — *not* subdirectories,
because envs share scripts/structure and only diverge on vars/secrets/slug; files structurally
prevent per-env script drift.
- **`scripts/` at every level cascades** up the tree; nearer overrides farther by name (CoW).
- **A leaf group = one logical app; its environments = the actual server apps.** Multiple distinct
apps = multiple sibling leaf groups. **Intermediate groups may also bear apps** (a dir may have
both subdirs and overlay files) — allowed, no special-casing.
- **Environment registry lives at the app-bearing node**, but **confirm-policy is inheritable** (set
"production always confirms" once at root; it flows down via the §3 mechanism). Leaves may declare
independent environment sets; a tree-wide `pic apply --env production` simply skips a leaf that has
no `production`.
- **Attach point:** the local root manifest declares where it binds into the server tree
(`parent_group = "acme"` or instance root). Ancestors above it are inherited/referenced but **not
present locally** — which makes the sparse checkout enforce the RBAC masking for free; effective
config needs a server round-trip (`pic config --effective`).
- **Stable IDs in gitignored `.picloud/`** (group IDs, instance URL, token ref) so a directory
rename/move maps to a server **reparent**, not delete+create.
- **Local/server structural divergence is detected, not silently fought.** Alongside the per-node
content version (§4.2), each node carries a **per-subtree structure version** (covering its own
parentage/subtree — **not one global counter**, so a structural edit in one team's subtree never
force-refuses an unrelated repo's plan). `pic plan` compares the local parent-by-ID (from
`.picloud/`) against the server's; on a structural mismatch (someone reparented server-side, or a
dir moved locally) it **refuses**, requiring an explicit `--adopt-server-structure` or
`--force-local-structure` (Terraform's detect-and-refuse on stale state). The content-version check
alone would miss a pure structural move; reparent/rename/delete each bump the affected subtree's
structure version.
- **Mono-repo** = attach at instance root. **Per-team repo** = attach at a subgroup, contain only its
slice. Same model, different attach depth.
---
## 7. Ownership of shared nodes
**Single-owner-per-node, ceilinged by the attach point.**
1. Each group node is owned by exactly one **project-root** (the repo that *manages* it — contains
its manifest as something it applies, not merely references). The server records `owner_project`.
2. **Your attach point is your ceiling** — you cannot apply to anything above your local root.
3. **First apply claims; transfer is explicit and capability-gated.** A second repo applying to an
owned node is rejected (`owned by project X; use --takeover`); takeover needs group-admin
capability. This stops one team silently clobbering org-wide config — whose blast radius (via the
effective-view fan-out) is the whole subtree.
4. **Ownership ⟂ RBAC.** Ownership = *which manifest is authoritative*; RBAC = *whether this
principal may*. The owner still needs group-admin capability to apply.
5. A node with **no project claim is UI/API-owned** — the dashboard is its source of truth and no
manifest fights it. Every node is either manifest-owned (one repo) or UI-owned.
**Corollary:** don't co-own a node — split config downward. Shared config lives *higher* (owned by a
platform/shared repo attaching at root); team-specific bits go into subgroups each team owns.
---
## 8. Diagrams
### 8.1 Server ownership & containment
Only scripts/vars/secret-refs are group-ownable (polymorphic owner); routes/triggers/data are always
app-owned via `app_id`.
```mermaid
graph TD
INST([Instance])
INST --> RG["Group: acme (root)"]
RG --> SGA["Group: team-a"]
RG --> SGB["Group: team-b"]
SGA --> LGA["Group: blog (leaf)"]
SGA --> LGB["Group: shop (leaf)"]
LGA --> APP1["App: blog-staging"]
LGA --> APP2["App: blog-production"]
RG -.->|"owns (shared)"| D0["scripts, vars, secret-refs"]
SGA -.->|owns| D1["team-a scripts, vars"]
APP2 -.->|"owns / overrides"| D2["app scripts, vars, secrets"]
APP1 ==>|app_id| C1[("routes, triggers, KV/Docs/Files, Topics")]
APP2 ==>|app_id| C2[("routes, triggers, KV/Docs/Files, Topics")]
```
### 8.2 Entity identity & cardinalities
```mermaid
erDiagram
GROUP ||--o{ GROUP : "parent-of (single parent)"
GROUP ||--o{ APP : contains
GROUP ||--o{ DEFINITION : "owns (scripts/vars/secret-refs)"
APP ||--o{ DEFINITION : "owns (override)"
APP ||--o{ APPSCOPED : "owns (routes/triggers/data, app_id)"
GROUP ||--o{ MEMBERSHIP : "role (inherited down)"
APP ||--o{ MEMBERSHIP : "role"
```
### 8.3 Config resolution (sparse merge, env filter, proximity-first)
Effective value for one app+env, resolved by §3. Env-scope filters per level; nearest level wins;
maps merge per key.
```mermaid
graph TB
I["Instance defaults"] --> RG["Group acme<br/>secret stripe_key (ref)<br/>script auth.rhai<br/>db_url@production"]
RG --> SG["Group team-a<br/>var region = eu"]
SG --> LG["Group blog<br/>script render.rhai<br/>var title = Blog"]
LG --> EV["Env overlay: production<br/>var title = Blog PROD (override)<br/>secret stripe_key = prod value"]
EV --> EFF[["Materialized effective view (app=blog-production):<br/>auth.rhai (acme), render.rhai (blog)<br/>title = Blog PROD (leaf overlay)<br/>region = eu (team-a)<br/>db_url = prod (acme @production filter)<br/>stripe_key = prod"]]
```
### 8.4 Filesystem ↔ server mapping
```mermaid
graph LR
subgraph FS["Git working tree"]
direction TB
R["acme/<br/>picloud.toml<br/>scripts/"]
R --> TA["team-a/<br/>picloud.toml<br/>scripts/"]
TA --> BL["blog/<br/>picloud.toml (base)<br/>picloud.staging.toml<br/>picloud.production.toml<br/>scripts/render.rhai"]
LINK[".picloud/ (gitignored)<br/>group IDs, instance URL, token ref"]
end
subgraph SRV["PiCloud server"]
direction TB
G0["Group acme"] --> G1["Group team-a"]
G1 --> G2["Group blog"]
G2 --> A1["App blog-staging"]
G2 --> A2["App blog-production"]
end
R -.->|defines| G0
TA -.->|defines| G1
BL -.->|"base defines"| G2
BL -.->|"staging.toml"| A1
BL -.->|"production.toml"| A2
LINK -.->|"stable IDs"| SRV
```
### 8.5 Multi-repo subtree views & single-owner ownership
```mermaid
graph TD
subgraph Server["Server group tree (authoritative)"]
acme["acme (root)"]
acme --> ta["team-a"]
acme --> tb["team-b"]
ta --> blog["blog"]
tb --> shop["shop"]
end
subgraph PR["Repo: platform"]
pr["manages acme"]
end
subgraph AR["Repo: team-a"]
ar["attaches at acme<br/>manages team-a + blog"]
end
pr ==>|owns| acme
ar ==>|owns| ta
ar ==>|owns| blog
ar -.->|"reference, read-only"| acme
```
### 8.6 Apply pipeline (bound plan → DB-atomic write → convergent propagation)
```mermaid
sequenceDiagram
actor Dev
participant CLI as pic CLI
participant Mgr as manager-core (single writer)
participant DB as Postgres
participant View as effective views + caches
Dev->>CLI: pic plan --env production
CLI->>CLI: read manifests + .rhai sources, build subtree bundle
CLI->>Mgr: send bundle (whole subtree)
Mgr->>DB: read state + version of affected nodes
Mgr->>Mgr: diff + blast radius + persist plan artifact
Mgr-->>CLI: PLAN (changes, N descendant apps incl. other repos, conflicts)
CLI-->>Dev: show plan; confirm/approve per env policy
Dev->>CLI: pic apply (executes stored plan)
CLI->>Mgr: apply (plan id)
Mgr->>DB: refuse if content or structure version moved; else ONE transaction (all-or-nothing)
Mgr->>View: recompute effective views + bump generation + invalidate
Mgr-->>CLI: change report (created/updated/re-enabled/pruned/conflicts)
CLI-->>Dev: log changes
```
### 8.7 RBAC: masked group secrets
```mermaid
graph TD
GA["Group admin (team-a)"] -->|"sets + can read"| GS["Group secret: stripe_key<br/>owned by team-a, encrypted at rest"]
AD["App developer (blog)"] -- "cannot read value (sees exists)" --x GS
AD -->|"can edit"| AS["App script source + app vars"]
GS ==>|"runtime injects plaintext"| EX["Executor: running app script"]
AS --> EX
GATE["Two gates:<br/>human-read authz vs runtime resolution"]
GATE -.-> GS
GATE -.-> EX
```
### 8.8 The three-state `enabled` lifecycle
```mermaid
stateDiagram-v2
[*] --> Active: declared (enabled=true/omitted)
Active --> Disabled: set enabled=false (manifest or UI toggle)
Disabled --> Active: set enabled=true (manifest or UI toggle)
Active --> Pruned: removed from manifest + prune/--prune
Disabled --> Pruned: removed from manifest + prune/--prune
Pruned --> [*]
note right of Disabled
Still desired state, NOT pruned.
Route 404s, trigger inert, script not invocable.
Out-of-band toggle on this field is conflict-guarded (4.2).
end note
```
---
## 9. Adoption & backfill
Groups land onto a live instance with existing flat apps, so a migration is a prerequisite, not an
afterthought:
- Create a **root group** (and/or a per-owner personal namespace, GitLab-style) and **reparent every
existing app** under it. Every app must have a parent from day one so resolution always terminates.
- Existing apps have no group-owned definitions, so their effective view = their own rows — the
materialized-view store can be backfilled trivially (identity resolution).
- The trigger `name` backfill (§4.5) runs in the same migration window.
- Existing app slugs are already instance-global, so no slug rewrite is needed; the path is new
metadata layered on top.
---
## 10. Open questions & residual risks
Resolved items now live inline next to their topic. What genuinely remains:
- **Effective-view invalidation SLA (§5.1)** — the security-staleness guarantee at fan-out to many
descendants is unsolved; synchronous-for-security + eventual-elsewhere is the proposed shape, not a
proven one. Highest-risk open item.
- **Conflict bit vs. operational lock (§4.2)** — *decided:* phase 1 ships the `enabled` / secret-
reference **conflict bit**; the operational-lock flag is the documented fallback if the bit proves
too heavy. (Was listed as undecided; resolved here to match §11 phase 1.)
- **Multi-level env-scope precedence (§3.2)** — *decided default:* proximity-first with env as a
per-level filter. The open part is only *validation at depth*, which is why `pic config --effective
--explain` is a **phase-3 hard requirement** (when multi-level resolution first ships), not a
precondition to adopting the rule.
- **Inherited-membership revocation lag (§5.3)** — revoking a group admin must drop implicit admin on
every descendant app, but §5.1's synchronous-invalidation subset covers only `enabled`/secrets, not
**role revocation** — leaving an unbounded window. New residual risk; should get the same
synchronous-for-security bound, lands with phase 3's authz.
- **External execution cancel (§4.2 kill-switch)** — the executor has **no external-cancel path
today** (`spawn_blocking` + self-checked deadline, verified); the kill-switch is a net-new capability
(a cancel flag polled in `on_progress`), deferred past phase 1. Until it exists, the strongest stop
is op-budget/deadline + the dispatcher fire-time `enabled` re-check (§4.3) for the trigger path.
- **Narrow-inheritance vs. trigger/route templates (§4.5, §5.1)** — the per-app binding tax bites
early at high tenant cardinality. Decide whether templates are truly deferrable for your target.
- **`pull --factor`** — auto-extract a shared base by diffing two pulled envs (later nicety).
---
## 11. Suggested phasing
1. **Declarative project tool, single-app (no groups yet).** `init`, `pull`/`config --effective`,
manifest parse/validate, `plan` (bound artifact), `apply` (**atomic desired-state write — requires
the manager-core post-commit-refresh restructuring of §4.2, domains/files excluded from the
transactional core**), `prune`, secrets push, link state, env-scoped config. Adds `enabled` to
scripts/routes + the three-state runtime + the dispatcher fire-time `enabled` re-check (§4.3) +
trigger `name` column/backfill + the `enabled`/secrets conflict bit + the net-new content +
tree-structure version counters + a coarse per-instance apply lock; in-flight executions finish (no
kill).
2. **Groups as pure org/RBAC/UI container.** Nested groups (single-parent, `parent_id`, **delete =
RESTRICT**, reparent/rename with the **ancestor-walk cycle guard** + **slug-freeze** +
**tree-structure version**, §5.6), inherited membership, hierarchy-aware `can`, UI grouping, the §9
backfill. No shared resources yet — cheap, no data-plane schema change.
3. **Group-inherited config** (vars, secret-refs, env-scoped). The net-new `vars`/`secret-refs`
tables + polymorphic owner; the group-secret AAD scheme (§5.3 caveat); masked group secrets; the
effective-view resolver + materialization + invalidation; **`config --effective --explain`** (hard
requirement, since multi-level resolution first ships here).
4. **Group-inherited scripts/modules.** CoW overrides; the **scope-aware module/import resolver +
extension points** (§5.5); cache-invalidation fan-out hardening; versioning/pinning if needed.
5. **Project tool maps onto groups.** Nested manifests, attach point, single-owner, server-computed
tree plan, per-env approval gating.
6. **(Much later) group-level collections/topics** — the v1.3 cross-app data-sharing problem, with a
real shared-scope authz model. Optionally, trigger/route **templates** (§4.5) if cardinality
demands.
---
## 12. Contracts still to draft
- The **apply bundle / plan artifact / change-report** wire contract (what the CLI ships, what the
server persists and returns), including the conflict and blast-radius shapes.
- The **effective-view resolver** (the read primitive) — the §3 rule made executable, plus the
materialization + invalidation protocol (§5.1).
- The **full manifest schema** spelling every block (scripts, routes, the 8 trigger kinds, storage
config, env-scoped vars, secret-refs, domains, `[project.environments]` + confirm policy).