Files
PiCloud/docs/design/groups-and-project-tool.md
MechaCat02 b46667c309 docs: reconcile the groups/project-tool initiative into the roadmap
The groups/project-tool work had no home in the blueprint's version
scheme, two "Phase N" numbering schemes overlapped, and CLAUDE.md's
"current focus" was stale (claimed v1.1.0 SDK foundation; we're at v1.1.9
with the SDK line complete). Reconcile all three:

* blueprint §12 Phase 5 (v1.2, already titled "Advanced Workflows &
  Hierarchies") — split into a Hierarchies track (the groups + project
  tool initiative, with §11 Phases 1–3 marked shipped) and a Workflows
  track, with sequencing flagged as an open call. This gives the work an
  explicit version home: the Hierarchies half of v1.2.
* design doc header — Status Draft → Active, "scheduled as the Hierarchies
  track of v1.2", and an explicit warning that §11's local Phase 1–6
  numbering is distinct from the blueprint product-phase numbering.
* design doc §11.1 — a post-Phase-3 re-sequencing review: resolve Phase 4
  scripts LIVE (drop the fan-out-invalidation sub-project, per the Phase 3
  result); consider a "Phase 5-lite" (declarative apply of vars/secrets +
  app scripts) before the hard group-scripts work; and force the two
  deferred decisions (trigger/route templates vs tenant cardinality;
  multi-repo ownership vs a single-repo start).
* CLAUDE.md — refresh "current focus" to v1.2 Hierarchies, and correct the
  now-broken "every v1.1+ table is app_id NOT NULL" invariant: the new
  group-inheritable config tables (vars, secrets) use a polymorphic owner
  (group_id XOR app_id), not app_id NOT NULL.

Doc-only; no code change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 07:12:25 +02:00

64 KiB
Raw Permalink Blame History

Groups, Projects & the Declarative Project Tool

Status: Active — design discussion captured 2026-06-18, revised 2026-06-20 with resolutions grounded in precedent (GitLab, Kustomize, Helm, Terraform, Kubernetes Server-Side Apply, Pulumi) and corrected against the codebase after three independent review passes (consistency, gaps, feasibility). Scheduled as the Hierarchies track of v1.2 (blueprint §12 Phase 5). §11 Phases 13 (project tool, groups + hierarchy RBAC, group-inherited config) are shipped; Phases 46 remain. NB: §11's own Phase 16 numbering is local to this initiative and is distinct from the blueprint product-phase numbering — "Phase 3" here = group-inherited config, not the blueprint's admin-auth phase.

Scope: turns the imperative pic CLI into a declarative, file-based project tool, and introduces a server-side groups hierarchy so apps can share and inherit config/scripts without duplication.

Blueprint impact: this reverses the §11.5 snapshot-copy, not live-link stance — top-down hierarchical inheritance (group → app) is now in scope. Originally designed as a materialized, auto-invalidated view (§5.1); Phase 3 shipped it as live, resolve-on-read inheritance (one recursive CTE per call, no cache), with materialization deferred to a future optimization (see the Phase 3 status note in §11.3). Fold the outcome back into the blueprint before it drifts.


1. Motivation

Today pic is imperative: every command (pic deploy, pic routes create, pic triggers create-cron, …) maps 1:1 to an admin API call. There is no memory of desired state, no project file, and no way to express "these N apps are the staging/prod/tenant variants of one thing."

Two gaps follow:

  1. No declarative project layer. Developers re-run commands by hand; nothing reconciles a repo to an instance.
  2. No server-side notion of a project/group. If we model environments as separate apps, the shared base (config, scripts) is duplicated across every deployed app. The server sees N unrelated apps.

This document designs both halves: a declarative manifest + project tool, and a groups hierarchy on the server that the project tool projects onto the filesystem.

manager-core is the single writer to Postgres. We lean on that for two concrete things — an atomic desired-state write (one DB transaction per apply, §4.2) and a server-computed diff (§4.2) — not for continuous reconciliation, which we deliberately decline (drift handling is detect-and-surface, §4.2). This is a narrower and more honest claim than "we can reconcile live state": most comparable tools cannot do an all-or-nothing apply because their effects are un-undoable cloud resources; ours are Postgres rows, so we can — within the boundary described in §4.2.


2. Vocabulary

Term Meaning
Group Server-side hierarchy node. Single-parent tree. Owns shared code/config definitions and members/roles. Nestable.
App A deployable unit. The data-isolation boundary (app_id). Lives under a group.
Environment A deployment variant (staging/production/…). An environment is an app. Also a scope dimension on config (§3).
Tenant Modeled like an environment — a scope dimension / overlay axis, not a second parent (§5.4).
Project Not a server entity. A CLI view over a repo-managed subtree of groups.
Definition Entities that can be group-owned: scripts/modules, vars, secret-refs only. Routes, triggers, collections, topics are app-scoped (§3, §5.1).
Data KV/docs/files collections + pub/sub topics. Always app-owned (app_id). The isolation boundary never moves.
Manifest TOML file describing desired state for one node (group base) + per-env overlays.
Effective view The materialized, per-app resolved set of definitions, served to the runtime and recomputed on any ancestor change (§5.1).
Attach point Where a local working tree's root binds into the server group tree. Its ceiling — you cannot apply above it.

Relationship: the base + per-env-overlay "project" is the degenerate one-group subtree of the general groups model. Nesting generalizes it to multi-group subtrees. Nothing about the single-project model is lost.


3. Configuration resolution

This is the single rule the whole design hangs off; everything else (env vars, secrets, enabled, overrides) resolves through it. All three precedent systems converge on the same shape, so we adopt it wholesale:

⚠️ Net-new, not an extension (verified against the codebase). PiCloud has no env-agnostic config / vars layer today — this resolution engine and the vars table are greenfield. Only secrets exists, and as a value-bearing table, not a ref. Read §3 as new infrastructure, not a modification of something present.

Resolution = sparse, per-field merge; environment is a pre-filter, not a precedence tier; proximity wins across levels.

To resolve a key for an app in environment E:

  1. Filter by environment first, per level. A value scoped @E is eligible; a value scoped * (env-agnostic) is the fallback. At a single level, @E beats *. Environment scope decides eligibility, it is not a competing precedence rank. Evidence: GitLab CI/CD variables use environment_scope exactly this way — a production-scoped row simply isn't visible to a staging job (GitLab CI/CD variables).
  2. Then nearest level wins. Walk up acme → team-a → blog → app; the closest level that defines the (filtered) key wins. GitLab states this verbatim: "if the same variable name exists in a group and its subgroups, the job uses the value from the closest subgroup."
  3. Merge granularity: maps/vars deep-merge per key (set title, still inherit region); entities (scripts) replace by identity (name); deletion is an explicit tombstone. Evidence: Kustomize strategic-merge patches are sparse and deep-merge maps but replace lists unless keyed (Kustomize patches); Helm deep-merges nested maps and uses null to delete a defaulted key (Helm values).

This one rule resolves several issues at once: it makes group config environment-scopable (a group sets db_url@production and db_url@staging; descendants inherit the right one — no per-leaf duplication), it defines merge granularity (maps per-key, entities per-identity), and it makes enabled just another sparse field (§4.3).

3.1 The env-scope manifest syntax

# group manifest (e.g. team-a/picloud.toml)
[vars]
region = "eu"                 # env-agnostic (scope *)

[vars.production]             # scope @production
db_url = "postgres://prod/..."
[vars.staging]
db_url = "postgres://staging/..."

[secrets]
names = ["stripe_key"]        # name-only; values pushed per env via CLI (§4.6)

3.2 The one deliberately-novel precedence call

With db_url@production at the root and a plain db_url at the leaf, the leaf's env-agnostic value wins for a production app — proximity beats farther-level env-specificity. The research was explicit that no major system lets env-specificity override proximity across levels; GitLab structurally avoids the question by treating env as a per-level filter (above). So proximity-first is the evidence-backed default, and we document it as a chosen rule: inner scope shadows outer, like lexical scoping. To avoid this becoming a debugging nightmare at depth, pic config --effective --explain must show why a value resolved (which level/scope it came from).

Residual risk (carried, not solved): multi-level + environment scopes is genuinely novel territory — GitLab is two-level (group→project). Proximity-first at arbitrary depth is consistent and defensible but not battle-tested. The --explain tooling is a hard requirement, not a nicety.


4. Manifest & apply

4.1 Manifest & layering

  • TOML, data not code. PiCloud runs untrusted scripts; the deployment descriptor stays inert, diffable, server-validatable, and dashboard-renderable. Turing-completeness belongs in the Rhai script, never the manifest.
  • Base + overlays. A node's picloud.toml holds the shared base; picloud.<env>.toml overlays add per-env vars/secrets/slug. Shared config is written once; overlays only carry deltas (the Kustomize base+overlay model — overlays are sparse patches, not replacements).
  • Routes are top-level, referencing scripts by name (one place to see all routing).
  • pull is first-class — see §4.6 for its inheritance/masking semantics.

4.2 The apply engine

The IaC research reshaped this section materially; the relevant precedents are cited inline.

Plan/apply split with a bound plan artifact. pic plan computes a reviewable diff (create / update / delete / replace), the server persists it as a row, and apply executes exactly that stored plandetecting and refusing if state moved underneath it (state version/serial check, cheap given the single writer). Evidence: Terraform's saved plan "reliably perform[s] an exact set of pre-approved changes, even if the configuration or state has changed in the minutes since" (Terraform run). The "state moved" check covers both the per-node content version and the tree-structure version (§6): a reparent or a new app created under the subtree between plan and apply changes the true blast radius, so the bound plan must refuse if either counter moved. Both are net-new — the existing scripts.version column is an unconditional write counter, not a compare-and-set guard, and must not be mistaken for one.

Atomicity — the corrected position. No major IaC tool does transactional apply, because their effects are un-undoable cloud resources (you cannot un-create a VM). That constraint does not bind us: a PiCloud manifest apply is almost entirely Postgres row writes (script source, route/trigger defs, config). Therefore:

  • The desired-state write is a single DB transaction — all-or-nothing across the subtree. This is achievable because our substrate is one transactional store (unlike Terraform's un-undoable cloud resources). It removes the downstream-breakage hazard of an earlier draft (which proposed per-node, stop-on-error apply — now superseded): there is no intermediate state where a parent committed and a child did not.
  • Propagation is forward-convergent, not transactional. The post-commit effective-view refresh + cache invalidation (§5.1) live outside the transaction. So "atomic" means atomic-for-desired-state, eventually-consistent-for-effect, with a bounded window where the DB says new and a cache says old.

Invariant to enforce — and it does NOT hold today (verified). The single-transaction model requires apply to write only Postgres, but the current admin write paths violate that: route and domain writes prime in-process caches synchronously, interleaved with the write (route_admin.rs:234, apps_api.rs:396), and files fsync to disk. So moving to "commit the transaction, then refresh views" is a deliberate restructuring of manager-core, not a description of the status quo. Additionally, domains are a non-DB effect (Caddy/filesystem, out of band) — they must be excluded from the transactional core and handled by the convergence model below. Treat "apply writes only Postgres" as an invariant to establish and guard, not one we already have.

Idempotent, convergent recovery (for the non-transactional propagation and any future side effects). Operations are idempotent upserts keyed by stable identity; "re-apply converges" is the recovery story; "rollback" means revert the declaration and re-apply forward. Per-node status (applied/pending/failed/drifted) is recorded so a partial propagation failure is observable and re-runnable. Evidence: every surveyed tool (Terraform, ArgoCD, Flux, Pulumi) chose idempotent forward-only convergence over distributed rollback.

Drift model — upgraded from pure last-write-wins. An earlier draft chose model "(c)": last-write-wins, just log. That is exactly the pre-SSA kubectl apply footgun — KEP-555 lists the scenario verbatim: "User does an apply, then kubectl edit, then applies again: surprise!" Kubernetes Server-Side Apply exists to convert that silent clobber into an explicit conflict (K8s SSA). Concrete risk for us: CI runs pic apply and silently re-enables a route an operator killed in an emergency. Resolution:

  • Keep (c)'s simplicity for ordinary fields (last-write-wins, change reported).
  • For the security-relevant subset onlyenabled and a secret's reference/existence in desired state — track a "changed out-of-band" bit. If that subset was changed outside the manifest (e.g. an operator disabled a route in the dashboard), pic plan surfaces it as a labeled conflict and apply refuses without --force. This is a scoped version of SSA's field ownership; we deliberately avoid full per-field managers (object bloat, complexity).
  • Secret values are explicitly out of scope of the conflict/state-version machinery. Values are never in the manifest or the plan (§4.6), so a routine pic secret set between plan and apply is the supported workflow and does not trip the state-version refusal — the state-version check covers manifest-managed desired state (definitions, including secret references and enabled), not value rotation.
  • Out-of-band changes render as a distinct labeled diff in plan, separate from intended changes. Evidence: Terraform/Pulumi both make read-only drift detection default and label drift distinctly (Pulumi drift).

Residual risk: the conflict bit reintroduces some of the complexity (c) was chosen to avoid. The cruder fallback, if even that is too much: keep pure (c) but add a non-revertible operational lock flag an operator sets in an emergency that apply won't touch. Either closes the silent-revert hole; neither is free.

Gating high-stakes applies: trigger ≠ authorization. Any CI trigger may plan, but applying to a confirm-required env is a separate, default-off, explicitly-authorized step. A blanket --yes covers ordinary confirms; confirm-required envs require an explicit per-env --approve <env>, so CI must opt in per environment. "Override a gate" is its own audited capability (maps onto manager-core::authz::can). Evidence: Terraform Cloud parks runs in Needs Confirmation and separates the apply runs permission from manage policy overrides (TFC run states).

Concurrency. Start with a coarse per-instance (or per-root-group) apply lock — one apply at a time — which is trivially correct for the single-node MVP and makes "last-commit-wins" hold. Refine later to per-blast-radius advisory locks (lock the affected apps in app_id order so overlapping radii serialize while disjoint ones proceed; queue triggers already use this advisory-lock primitive) only if contention appears. Don't build the fine-grained version speculatively.

Blast radius — defined and bounded. The blast radius is the set of descendant apps whose materialized effective view would actually change — a diff, not "all descendants." Per changed definition at node N it is subtree(N) minus apps that override that key nearer. pic plan enumerates it for small radii; for large ones it summarizes (count + sample) and a threshold triggers extra confirmation ("this changes 4,213 apps — confirm"). Because only scripts/vars/secret- refs inherit (§5.1), blast radius applies to those changes; an app-scoped change (a route or trigger) has blast radius = that single app. Root-level changes are accepted as expensive, rare, and high-confirmation; the computation is bounded by subtree(N) size, and the confirmed radius is re-validated at apply against the tree-structure version (above) so it can't go stale.

In-flight executions. An apply that disables or replaces a script affects new invocations immediately (via the enabled re-check + view invalidation; pending trigger outbox rows are dropped at fire-time, §4.3 — so no separate outbox purge is needed). In-flight running executions run to completion, and note (verified) the executor has no external-cancel path today: executions are spawn_blocking Rhai calls interruptible only by their operation budget or a pre-set wall-clock deadline self-checked in engine.on_progress. A true must-stop-now kill-switch is therefore a net-new capability — buildable by having that same on_progress hook also poll a per-execution cancel flag — gated by an admin capability (authz::can) and audited, scheduled as a later item, not phase 1. Killing mid-run also risks partial side effects, so it stays the explicit extreme, never default apply behavior.

Apply flow (end to end):

CLI builds bundle (manifests + script sources) for the subtree
  → server computes plan + blast radius (incl. descendant apps in OTHER repos)
  → server persists plan artifact; checks state version
  → dev reviews; confirm/approve per env policy
  → server applies desired state in ONE DB transaction (all-or-nothing)
  → server refreshes effective views + bumps generation + invalidates caches (convergent)
  → server returns change report; CLI logs created / updated / re-enabled / pruned / conflicts

4.3 The three-state enabled lifecycle

enabled is a real platform feature (DB + server runtime + UI badge/toggle), not CLI sugar, and — by §3 — it is just another sparse, proximity-resolved field.

State Meaning Pruned?
declared, enabled = true (or omitted) deployed, active kept
declared, enabled = false deployed but inert (route short-circuits, trigger doesn't fire, script not invocable) kept — still desired state
absent from merged manifest stale deleted by prune / --prune
  • Default true; last-write-wins on merge; a base enabled = false is inherited until an overlay (env or a nearer group/app) explicitly sets enabled = true. Overriding a different field does not implicitly re-enable.
  • Across the group axis (resolved): a descendant disables an inherited (group-owned) script via a sparse override — an override row that sets only enabled = false, inheriting the source. A descendant can re-enable a parent-disabled entity because nearest-level wins. This is the same sparse-field mechanism as everything else (Kustomize patches are sparse — you specify only what changes).
  • Base = the superset across envs. Per-env you may toggle enabled in either direction (disable a base-active entity, or re-enable a base-disabled one — proximity wins, §3), but you cannot remove a base entity in one env; true removal requires the entity not be in the base. An entity unique to one env goes in that overlay; an entity present everywhere but off in staging goes in base with enabled = false in the staging overlay.
  • Disabled = invisible. External callers hitting a disabled route get 404 (indistinguishable from absent — no info leak).
  • Schema note: the triggers table already has enabled (+ dispatch_mode, retry columns) and it is honored at match/schedule time (trigger_repo.rs, cron_scheduler.rs — verified). New work is enabled on scripts and routes only, plus runtime honoring in the matcher / invoker.
  • Fire-time re-check (shipped, test-backed): the dispatcher re-checks enabled at fire time on every async path, so a pending event for a now-disabled trigger/script is dropped (not executed) when it comes up — the §5.1 security-disable guarantee on the trigger path. Outbox arm: resolve_trigger sets active = trigger.enabled && script.enabled and the async-HTTP/invoke builders set active = script.enabled; dispatch_one drops the row via a single unified if !resolved.active gate. Queue arm: dispatch_one_queue runs a separate path (it claims from queue_messages, not the outbox), so in addition to list_active_queue_consumers filtering s.enabled/t.enabled at list time, it re-checks the freshly-read script.enabled before executing and releases the claim (nack) when disabled — closing the per-tick TOCTOU window (a script disabled after the list snapshot but before its in-flight message runs). Both arms are regression-locked: dispatcher::tests::outbox_enabled_gate::disabled_http_outbox_row_is_dropped_without_executing covers the unified !resolved.active outbox gate, and dispatcher::tests::queue_enabled_gate::disabled_queue_consumer_releases_claim_without_executing covers the queue arm. Same-app backstop: every async build path (resolve_trigger, build_http_request, build_invoke_request, dispatch_one_queue) rejects a row whose script's app_id ≠ the row's app_id, so a hand-edited/restored row can't run one app's script under another's SdkCallCx — the cross-app isolation boundary.
  • Provenance caveat (accepted): a single boolean carries no "manual vs manifest" provenance, so a disabled entity looks the same however it got there — which is why the §4.2 conflict bit on the enabled/secrets subset exists, to stop the next apply silently reverting an operational disable.

4.4 Identity & naming

  • Kebab everywhere. One canonical identifier regex: ^[a-z0-9][a-z0-9-]{0,62}$ for project names, env names, script names/slugs, trigger names — unified with the existing app-slug rule.
  • Scripts: unique name/slug per app = merge/upsert key.
  • Routes: identity = the triple <method> <host> <path>, e.g. ANY *.beta.example.com /hello/:name. dispatch_mode, host_param_name, etc. are attributes overridable without changing identity. The CLI infers host_kind/path_kind from the pattern syntax (*, {name}, :name, exact), with an explicit kind key as override (mirrors the UI). Needs a normalization rule (default method ANY, default host *, case-folding) so manifest ↔ server match exactly.
  • Slugs are instance-global, derived from the path. Two identifiers coexist:
    • path = acme/team-a/blog — hierarchical, group-scoped, display/organization/RBAC.
    • slug = flat, instance-global, the deployment key. The derived default is {flattened-path}-{env} (e.g. team-a-blog-staging) — unique by construction in the multi-group world, unlike a bare {leaf}-{env} which collides whenever two groups reuse a leaf name. On >63-char overflow, truncate + short hash suffix. Explicit override allowed; the path never is the slug, it only seeds it.

4.5 Triggers

Triggers are app-scoped, not group-inherited (§5.1 explains why). Two senses of "app" are in play and must not be conflated: a trigger is declared once in the leaf (logical app) base — an authoring convenience — and materializes into per-app_id rows, one per env-app, where app_id is the server-app isolation boundary (§5.2). It is never group-owned.

Two distinct constraints:

  • name (explicit, kebab) = the merge/identity key + upsert target. Unique per app. (Triggers have no name column today — new.)
  • Semantic uniqueness = a post-merge validation: no two triggers may share their kind-specific semantic key. Checked after merging, so overriding a base trigger per-env (reuse name, change a field) is fine; two differently-named triggers with identical effect is an error.
Kind Semantic key Note
kv / docs / files (script, collection_glob, ops) canonicalize ops (sort + dedupe)
cron (script, schedule, timezone) exact TEXT match
pubsub (script, topic_pattern)
dead-letter (script, source_filter, trigger_id_filter, script_id_filter) NULL = literal value (not wildcard) for equality
email (script) no filters — one per script
queue (queue_name)not script-scoped already server-enforced (advisory lock: one consumer per (app_id, queue_name))
  • Matching vs identity: []/NULL/globs mean "any/wildcard" at dispatch time (unchanged runtime matching) but are compared as literal structural values for dedup. We dedup on structural identity, never on overlap/subsumptionops = [] and ops = ["insert"] are not duplicates. Overlapping triggers coexist (multiple triggers firing on one event is already how PiCloud works).
  • Name backfill for existing nameless rows: {kind}-{entity}-{n}, where {entity} is the kind's identity token (collection / queue / topic; for cron/email/dead-letter fall back to {kind}-{n}), sanitized to the kebab regex, with -{n} (discovery order) guaranteeing per-app uniqueness.

Ergonomic debt (accepted, watch it): because triggers don't inherit, 100 tenant apps each needing the same 5 triggers = 500 declarations. The fix is group trigger/route templates that fan out per descendant (a template/instantiation mechanism, not inheritance) — deferred, but it bites early if tenant cardinality is high. Pressure-test against real tenant counts before committing to the narrow-inheritance choice (§5.1).

4.6 Secrets & pull

  • Name-only in the manifest; value pushed via CLI (pic secret set, reads stdin). Unanimous across every comparable tool — never commit secret values.
  • Env-scoped like any var ([secrets] names declared once; values set per env).
  • Warn (don't block) if a referenced secret is not set. This requires an app dev to see that a group secret exists / is set (a boolean) without reading its value — an accepted, explicit authz boundary. GitLab surfaces inherited masked variable keys the same way.
  • Email's inbound_secret is a reference, not inline — same rule; the server already encrypts it at rest.
  • pull exports own-rows only (this node's overrides), never effective/inherited state. A separate read-only pic config --effective shows the inherited result with masked secrets rendered as <set> / ***, never plaintext. This makes pull-under-masking safe by construction — you cannot pull a secret you cannot read, and you cannot accidentally duplicate inherited config into a leaf. Evidence: GitLab shows inherited group variables in a project read-only and separate from project-own variables.
  • Flat pull for a new project; "smart" delta-pull (own-vs-effective diff) is server-computed since an app dev's checkout lacks ancestor manifests.

4.7 Apply-time warnings

  • Enabled route/trigger pointing at a disabled script.
  • An endpoint script deployed with no route and no trigger (unreachable). Modules are exempt.
  • Sandbox override exceeding the admin ceiling.
  • Referenced secret not set.
  • An out-of-band change to the enabled/secrets subset (surfaced as a conflict, §4.2).

5. Groups & inheritance

5.1 What inherits, and the runtime model

  • GitLab-like, nested, single-parent tree. Single parent keeps inheritance acyclic and resolution deterministic (no diamond precedence).

  • Inherit code/config — narrowly. Inherit data — no. Group-inheritable = scripts/modules, vars, secret-refs only. Routes, triggers, collections, topics, files are app-scoped.

    • Why narrow: routes and triggers are bindings, not pure code/config. A group-owned trigger has no app data to watch (triggers fire on app-owned collections/topics/queues); a group-owned route has no host (routing is Host→app first). Inheriting them is incoherent. There are two distinct sharing mechanisms and an earlier draft conflated them: the leaf base+overlay shares routes/triggers across one app's environments; group inheritance shares code/config across many apps. Evidence: Serverless/SAM keep events: per-function and share logic via layers — bindings local, code shared.
    • Why data stays app-owned: a group script executes in the inheriting app's context, so its cx.app_id still scopes data to that app. Group-level collections/topics would break app_id as the isolation boundary — that is the v1.3 cross-app data-sharing problem and stays out.
  • Runtime model: a materialized effective view + versioned cache (not per-request live-resolve). An earlier draft said "live-resolve," which is underspecified and would fight the existing cache. The real model:

    ⚠️ Superseded for config by what Phase 3 actually shipped (§11.3): live resolution. For vars + secrets, PiCloud resolves on-demand via one recursive CTE per call — no materialized view, generation counter, or invalidation fan-out. There was no pre-existing config cache to "fight" (config is net-new, §3), and the org tree is small, so resolve-on-read is cheap and sidesteps the invalidation SLA below. The materialized model in this bullet stands only as the deferred plan for script-body inheritance (Phase 4) — where a hot per-request path may justify a cache — and as a future optimization for config if a deep tree ever demands it. Read the rest of §5.1 as "the materialization design, if/when we need it," not as Phase 3's runtime.

    • manager-core resolves-at-write into a materialized per-app effective view (§3 rule applied: sparse merge, env filter, proximity, CoW, enabled).
    • The orchestrator/executor serve from that view, keyed by app_id + a generation/version. The app_id-keyed shape survives (today's route cache is already an app_id→routes map), but the substance is net-new (verified): there is no generation counter anywhere today, the route cache is rebuilt whole-table on every write rather than per-app, and script bodies are live-resolved per request. Read "serve from it" as extend the keying, not reuse the mechanism — and note that fanning out today's full-rescan invalidation to thousands of descendants would be a regression, so per-app incremental invalidation is part of the build.
    • On any write to a node, manager-core (single writer, knows the tree) recomputes descendants' views + bumps the generation + invalidates caches.
    • The materialized view is a derived cache, not a second source of truth — canonical config still lives once at the owning node, so this does not reintroduce duplication. This dissolves the long-running "snapshot vs. live" tension: no duplication and propagation and a fast hot path.

Residual risk (relocated, not solved): cache invalidation is now a correctness/security requirement — disabling a script for a security reason must stop it running everywhere within bounded time. This is the same class as the existing PrincipalCache revocation lag. Require synchronous invalidation for the security-relevant subset (enabled=false, secret rotation) and accept bounded eventual staleness elsewhere; the hard SLA at fan-out to thousands of descendants is genuinely unsolved.

5.2 Schema impact

  • Inheritable definition kinds get a polymorphic owner (owner_kind ∈ {group, app} + owner_id): scripts (modify the existing table — and note its app_id FK is ON DELETE RESTRICT, not CASCADE, so a pruning apply needs explicit ordering), plus vars and secret-refs, which are net-new tables (they do not exist today — §3). So this is one table modified + two invented, not "three tables touched." Routes, triggers, collections, topics, files stay strictly app_id-owned, so the runtime isolation boundary stays fixed. (The CLAUDE.md app_id NOT NULL … CASCADE rule is itself not universal today — scripts already use RESTRICT.)
  • New: a groups table (single-parent, parent_id), group membership/roles, an owner_project column on group nodes (§7), and the materialized effective-view store keyed by app_id + generation (§5.1). Env-scoped values carry an environment_scope column (* or a specific env).

5.3 RBAC

  • Hierarchy-aware capabilities. authz::can(principal, cap, on=node) resolves by walking ancestors and taking the highest effective role. Instance → group(s) → app.

  • Inherited membership (GitLab-style): a group admin is implicitly admin of every subgroup/app beneath it.

  • Masked group secrets: a group secret is used by an app at runtime but not human-readable by the app's developers. Two orthogonal gates: runtime resolution (engine injects plaintext) vs human-read authz (admin API returns a value only to a principal with rights at the owning group). An app-scoped admin call never returns group secrets; runtime injection bypasses the human gate. App devs may see a group secret exists (for the unset-warning, §4.6) but not its value. An app can run with config its own developers cannot see.

    Crypto caveat (verified): secrets today are AES-GCM sealed with AAD secret:{app_id}:{name}, and the decrypt path hard-codes cx.app_id (migration 0042). A group-owned secret is not expressible under this scheme — there is no group identity in the AAD. It needs a new AAD identity (e.g. secret:group:{group_id}:{name}) and an owner-aware decrypt path that resolves whether the inherited secret is group- or app-owned. This is the single hardest correctness detail of group secrets and gates phasing step 3.

5.4 Tenants & single-parent

Single-parent forbids an app combining two sibling groups' configs — which seems to threaten the multi-tenant use case (shared-platform base + per-tenant overlay). It does not, because the shared base is an ancestor, not a sibling: model tenants as leaf apps under a tenants/ subgroup that inherits the platform base up the chain (tenant leaf → tenants group → platform group). The "two parents" intuition is satisfied by the chain. Better still, a tenant is a scope dimension like environment (§3) — the same env-scope machinery generalizes to a tenant scope, so multi-tenant needs no new hierarchy primitive. Evidence: GitLab is single-parent and serves multi-tenant teams via subgroups; "combine two siblings" is handled by promoting shared config to a common ancestor.


5.5 Module & import resolution under inheritance

A group-owned script may import modules, and inheritance makes "which module?" ambiguous (an inherited script importing a module a leaf has shadowed could bind to app-dev code it never saw — a trust inversion). The rule:

  • Lexical by default. An inherited script's imports resolve against the module set visible at the script's own defining node (walking up from there), not the inheriting app's effective view. A leaf cannot shadow a module an inherited group script depends on, and a group script behaves identically across every app that inherits it — preserving both determinism and the trust boundary (a security-authored shared script is tamper-proof from below). Ordinary lexical scoping; "sealed by default", like non-open classes in Kotlin / final in Java.
  • "Defining node" of a CoW-overridden script = the node that authored that body. An inherited-unchanged script's defining node is the ancestor that owns it (imports sealed to the ancestor's modules). A script a leaf overrides (CoW, §3) has the leaf as its defining node, so the override's imports resolve from the leaf — which is not a trust inversion, because the app owner wrote that body and is trusted within their own app. Consequence (intended): the same module name can resolve to two different bodies in one app, depending on the importing script's defining node.
  • Explicit extension points for opt-in polymorphism. A module is marked an extension point in the manifest (an [extension_points] declaration — not in Rhai source, keeping code inert and giving the plan checker something to read). Such a module is one descendants are expected to provide or override; only these resolve against the inheriting app's effective view. Controlled template-method customization (a shared render whose theme module each tenant supplies) without the blanket trust inversion of dynamic resolution. When a name is declared at more than one level — or is concrete on one path and an extension point on another — the nearest declaration's kind wins (proximity, §3); its default body (if any) is the inherited fallback.
  • Apply-time checks: a dangling import (inherited script → missing module) is a plan error; an extension point with no provider in a given app is an error for that app (a hard failure, joining §4.7).

Residual (verified): executor-core's PicloudModuleResolver is app-scoped today and ignores the importing script's origin (module_resolver.rs passes _source unused). Rhai does expose that origin, so the lexical-vs-dynamic split is expressible — but it requires re-keying the resolver cache by owner identity and adding per-import policy (sealed vs. extension point), i.e. a real resolver+cache redesign, not a parameter tweak. Lands with phasing step 4.

5.6 Tree lifecycle: delete, reparent, rename

Structural mutations of the group tree, which the rest of the design depends on staying acyclic and non-orphaning:

  • Delete = RESTRICT, never implicit CASCADE. Deleting a non-empty group is refused — an implicit cascade would destroy descendant apps and their isolated data. CLI --recursive expands a delete into ordered, explicit, confirmed child deletions; the DB FK stays RESTRICT. Corollary: a referenced ancestor cannot vanish while it has descendants, so cross-repo read-only references (§7) can't be orphaned by deletion — the RESTRICT protects them automatically. App data is destroyed only on explicit opt-in: an app delete refuses unless --purge-data, which then removes its KV/docs rows and its files blob tree under PICLOUD_FILES_ROOT/<app_id>/ — a non-DB, non-undoable effect run outside the transaction and logged. So --recursive group delete requires --purge-data to touch any descendant app's data; without it, a non-empty app blocks the delete.
  • Reparent / rename: the slug is frozen at creation. The path only seeds the derived slug (§4.4); a move or rename updates the display path but never rewrites the instance-global slug — the deployment key stays stable, external references don't break. After a move the slug no longer mirrors the path (cosmetic, accepted).
  • Reparent recomputes descendant effective views (it changes the resolution chain — the same fan-out invalidation as a node write, §5.1) and is doubly capability-gated: group-admin at both the source and destination parent (you remove from one ancestor's domain and add to another's). Because it changes the resolution chain, a reparent is validated like a plan and refused (unless forced) if the recompute would orphan a sparse enabled-override (now shadowing nothing) or leave an extension point with no provider (§5.5) — a structural move must not silently produce a state apply would have rejected.
  • Cycle guard, under the apply lock. Reparent runs an ancestor-walk check in manager-core (walk from the destination up to root; reject if it reaches the node being moved). A Postgres CHECK can't express this; the guard is what guarantees §9's "resolution always terminates." Single-parent + this guard = acyclic. All structural mutations (reparent/rename/delete) take the same coarse apply-lock (§4.2), so the ancestor-walk + parent_id write run serialized — two concurrent reparents can't race into a cycle, and a reparent's view-recompute can't collide with an overlapping apply on the materialized-view store.

6. The CLI ↔ server projection

  • Directories = groups (the hierarchy axis). Single-parent falls out of the filesystem for free.
  • Overlay files = environments/apps (the deployment-variant axis) — not subdirectories, because envs share scripts/structure and only diverge on vars/secrets/slug; files structurally prevent per-env script drift.
  • scripts/ at every level cascades up the tree; nearer overrides farther by name (CoW).
  • A leaf group = one logical app; its environments = the actual server apps. Multiple distinct apps = multiple sibling leaf groups. Intermediate groups may also bear apps (a dir may have both subdirs and overlay files) — allowed, no special-casing.
  • Environment registry lives at the app-bearing node, but confirm-policy is inheritable (set "production always confirms" once at root; it flows down via the §3 mechanism). Leaves may declare independent environment sets; a tree-wide pic apply --env production simply skips a leaf that has no production.
  • Attach point: the local root manifest declares where it binds into the server tree (parent_group = "acme" or instance root). Ancestors above it are inherited/referenced but not present locally — which makes the sparse checkout enforce the RBAC masking for free; effective config needs a server round-trip (pic config --effective).
  • Stable IDs in gitignored .picloud/ (group IDs, instance URL, token ref) so a directory rename/move maps to a server reparent, not delete+create.
  • Local/server structural divergence is detected, not silently fought. Alongside the per-node content version (§4.2), each node carries a per-subtree structure version (covering its own parentage/subtree — not one global counter, so a structural edit in one team's subtree never force-refuses an unrelated repo's plan). pic plan compares the local parent-by-ID (from .picloud/) against the server's; on a structural mismatch (someone reparented server-side, or a dir moved locally) it refuses, requiring an explicit --adopt-server-structure or --force-local-structure (Terraform's detect-and-refuse on stale state). The content-version check alone would miss a pure structural move; reparent/rename/delete each bump the affected subtree's structure version.
  • Mono-repo = attach at instance root. Per-team repo = attach at a subgroup, contain only its slice. Same model, different attach depth.

7. Ownership of shared nodes

Single-owner-per-node, ceilinged by the attach point.

  1. Each group node is owned by exactly one project-root (the repo that manages it — contains its manifest as something it applies, not merely references). The server records owner_project.
  2. Your attach point is your ceiling — you cannot apply to anything above your local root.
  3. First apply claims; transfer is explicit and capability-gated. A second repo applying to an owned node is rejected (owned by project X; use --takeover); takeover needs group-admin capability. This stops one team silently clobbering org-wide config — whose blast radius (via the effective-view fan-out) is the whole subtree.
  4. Ownership ⟂ RBAC. Ownership = which manifest is authoritative; RBAC = whether this principal may. The owner still needs group-admin capability to apply.
  5. A node with no project claim is UI/API-owned — the dashboard is its source of truth and no manifest fights it. Every node is either manifest-owned (one repo) or UI-owned.

Corollary: don't co-own a node — split config downward. Shared config lives higher (owned by a platform/shared repo attaching at root); team-specific bits go into subgroups each team owns.


8. Diagrams

8.1 Server ownership & containment

Only scripts/vars/secret-refs are group-ownable (polymorphic owner); routes/triggers/data are always app-owned via app_id.

graph TD
    INST([Instance])
    INST --> RG["Group: acme (root)"]
    RG --> SGA["Group: team-a"]
    RG --> SGB["Group: team-b"]
    SGA --> LGA["Group: blog (leaf)"]
    SGA --> LGB["Group: shop (leaf)"]
    LGA --> APP1["App: blog-staging"]
    LGA --> APP2["App: blog-production"]

    RG -.->|"owns (shared)"| D0["scripts, vars, secret-refs"]
    SGA -.->|owns| D1["team-a scripts, vars"]
    APP2 -.->|"owns / overrides"| D2["app scripts, vars, secrets"]

    APP1 ==>|app_id| C1[("routes, triggers, KV/Docs/Files, Topics")]
    APP2 ==>|app_id| C2[("routes, triggers, KV/Docs/Files, Topics")]

8.2 Entity identity & cardinalities

erDiagram
    GROUP ||--o{ GROUP : "parent-of (single parent)"
    GROUP ||--o{ APP : contains
    GROUP ||--o{ DEFINITION : "owns (scripts/vars/secret-refs)"
    APP ||--o{ DEFINITION : "owns (override)"
    APP ||--o{ APPSCOPED : "owns (routes/triggers/data, app_id)"
    GROUP ||--o{ MEMBERSHIP : "role (inherited down)"
    APP ||--o{ MEMBERSHIP : "role"

8.3 Config resolution (sparse merge, env filter, proximity-first)

Effective value for one app+env, resolved by §3. Env-scope filters per level; nearest level wins; maps merge per key.

graph TB
    I["Instance defaults"] --> RG["Group acme<br/>secret stripe_key (ref)<br/>script auth.rhai<br/>db_url@production"]
    RG --> SG["Group team-a<br/>var region = eu"]
    SG --> LG["Group blog<br/>script render.rhai<br/>var title = Blog"]
    LG --> EV["Env overlay: production<br/>var title = Blog PROD (override)<br/>secret stripe_key = prod value"]
    EV --> EFF[["Materialized effective view (app=blog-production):<br/>auth.rhai (acme), render.rhai (blog)<br/>title = Blog PROD (leaf overlay)<br/>region = eu (team-a)<br/>db_url = prod (acme @production filter)<br/>stripe_key = prod"]]

8.4 Filesystem ↔ server mapping

graph LR
    subgraph FS["Git working tree"]
      direction TB
      R["acme/<br/>picloud.toml<br/>scripts/"]
      R --> TA["team-a/<br/>picloud.toml<br/>scripts/"]
      TA --> BL["blog/<br/>picloud.toml (base)<br/>picloud.staging.toml<br/>picloud.production.toml<br/>scripts/render.rhai"]
      LINK[".picloud/ (gitignored)<br/>group IDs, instance URL, token ref"]
    end
    subgraph SRV["PiCloud server"]
      direction TB
      G0["Group acme"] --> G1["Group team-a"]
      G1 --> G2["Group blog"]
      G2 --> A1["App blog-staging"]
      G2 --> A2["App blog-production"]
    end
    R -.->|defines| G0
    TA -.->|defines| G1
    BL -.->|"base defines"| G2
    BL -.->|"staging.toml"| A1
    BL -.->|"production.toml"| A2
    LINK -.->|"stable IDs"| SRV

8.5 Multi-repo subtree views & single-owner ownership

graph TD
    subgraph Server["Server group tree (authoritative)"]
      acme["acme (root)"]
      acme --> ta["team-a"]
      acme --> tb["team-b"]
      ta --> blog["blog"]
      tb --> shop["shop"]
    end
    subgraph PR["Repo: platform"]
      pr["manages acme"]
    end
    subgraph AR["Repo: team-a"]
      ar["attaches at acme<br/>manages team-a + blog"]
    end
    pr ==>|owns| acme
    ar ==>|owns| ta
    ar ==>|owns| blog
    ar -.->|"reference, read-only"| acme

8.6 Apply pipeline (bound plan → DB-atomic write → convergent propagation)

sequenceDiagram
    actor Dev
    participant CLI as pic CLI
    participant Mgr as manager-core (single writer)
    participant DB as Postgres
    participant View as effective views + caches
    Dev->>CLI: pic plan --env production
    CLI->>CLI: read manifests + .rhai sources, build subtree bundle
    CLI->>Mgr: send bundle (whole subtree)
    Mgr->>DB: read state + version of affected nodes
    Mgr->>Mgr: diff + blast radius + persist plan artifact
    Mgr-->>CLI: PLAN (changes, N descendant apps incl. other repos, conflicts)
    CLI-->>Dev: show plan; confirm/approve per env policy
    Dev->>CLI: pic apply (executes stored plan)
    CLI->>Mgr: apply (plan id)
    Mgr->>DB: refuse if content or structure version moved; else ONE transaction (all-or-nothing)
    Mgr->>View: recompute effective views + bump generation + invalidate
    Mgr-->>CLI: change report (created/updated/re-enabled/pruned/conflicts)
    CLI-->>Dev: log changes

8.7 RBAC: masked group secrets

graph TD
    GA["Group admin (team-a)"] -->|"sets + can read"| GS["Group secret: stripe_key<br/>owned by team-a, encrypted at rest"]
    AD["App developer (blog)"] -- "cannot read value (sees exists)" --x GS
    AD -->|"can edit"| AS["App script source + app vars"]
    GS ==>|"runtime injects plaintext"| EX["Executor: running app script"]
    AS --> EX
    GATE["Two gates:<br/>human-read authz  vs  runtime resolution"]
    GATE -.-> GS
    GATE -.-> EX

8.8 The three-state enabled lifecycle

stateDiagram-v2
    [*] --> Active: declared (enabled=true/omitted)
    Active --> Disabled: set enabled=false (manifest or UI toggle)
    Disabled --> Active: set enabled=true (manifest or UI toggle)
    Active --> Pruned: removed from manifest + prune/--prune
    Disabled --> Pruned: removed from manifest + prune/--prune
    Pruned --> [*]
    note right of Disabled
      Still desired state, NOT pruned.
      Route 404s, trigger inert, script not invocable.
      Out-of-band toggle on this field is conflict-guarded (4.2).
    end note

9. Adoption & backfill

Groups land onto a live instance with existing flat apps, so a migration is a prerequisite, not an afterthought:

  • Create a root group (and/or a per-owner personal namespace, GitLab-style) and reparent every existing app under it. Every app must have a parent from day one so resolution always terminates.
  • Existing apps have no group-owned definitions, so their effective view = their own rows — the materialized-view store can be backfilled trivially (identity resolution).
  • The trigger name backfill (§4.5) runs in the same migration window.
  • Existing app slugs are already instance-global, so no slug rewrite is needed; the path is new metadata layered on top.

10. Open questions & residual risks

Resolved items now live inline next to their topic. What genuinely remains:

  • Effective-view invalidation SLA (§5.1) — the security-staleness guarantee at fan-out to many descendants is unsolved; synchronous-for-security + eventual-elsewhere is the proposed shape, not a proven one. (Was the highest-risk open item.) Moot for config as of Phase 3: vars/secrets ship with live resolution (no cache → no staleness window; a disable/rotation is visible on the next call). The SLA only re-arises if script-body materialization is built (Phase 4) or if config is ever materialized for performance; until then there is nothing to invalidate.
  • Conflict bit vs. operational lock (§4.2)decided: phase 1 ships the enabled / secret- reference conflict bit; the operational-lock flag is the documented fallback if the bit proves too heavy. (Was listed as undecided; resolved here to match §11 phase 1.)
  • Multi-level env-scope precedence (§3.2)decided default: proximity-first with env as a per-level filter. The open part is only validation at depth, which is why pic config --effective --explain is a phase-3 hard requirement (when multi-level resolution first ships), not a precondition to adopting the rule.
  • Inherited-membership revocation lag (§5.3) — revoking a group admin must drop implicit admin on every descendant app, but §5.1's synchronous-invalidation subset covers only enabled/secrets, not role revocation — leaving an unbounded window. New residual risk; should get the same synchronous-for-security bound, lands with phase 3's authz. Resolved — never materialized: Phase 2 RBAC and Phase 3 config both resolve roles/config live (no inherited-role or effective-view cache), so a revoked group grant is gone on the next request. The lag only exists if a cache is introduced.
  • External execution cancel (§4.2 kill-switch) — the executor has no external-cancel path today (spawn_blocking + self-checked deadline, verified); the kill-switch is a net-new capability (a cancel flag polled in on_progress), deferred past phase 1. Until it exists, the strongest stop is op-budget/deadline + the dispatcher fire-time enabled re-check (§4.3) for the trigger path.
  • Narrow-inheritance vs. trigger/route templates (§4.5, §5.1) — the per-app binding tax bites early at high tenant cardinality. Decide whether templates are truly deferrable for your target.
  • pull --factor — auto-extract a shared base by diffing two pulled envs (later nicety).

11. Suggested phasing

Status (Phase 1): shipped. init, pull, config --effective, manifest parse/validate, plan + apply (single-transaction desired-state write, post-commit route refresh, domains/files outside the tx), prune, secrets push, .picloud/ link state, env-scoped overlays (picloud.<env>.toml base+overlay), the three-state enabled lifecycle on scripts/routes (runtime 404 + dispatcher fire-time re-check), the trigger name column + backfill, a coarse per-app apply lock, and in-flight-finish-no-kill all landed.

The bound-plan staleness check is the content-fingerprint form (a state_token hash of the live state; apply 409s if it moved) rather than a persisted plan-artifact + serial — it delivers the §4.2 guarantee without a migration or interactive-write-path changes, and the token covers enabled/secret names so the enabled/secrets conflict is enforced on the plan→apply path.

Deferred, with rationale: the tree-structure version counter is a no-op seam until groups exist (Phase 2); vars + multi-level/--explain resolution and the richer per-env overlay are Phase 3; trigger upsert-by-name (in-place Update; the diff still Create/Deletes by semantic identity) and the dashboard enabled toggle are tracked follow-ups beyond the §11 bullets.

  1. Declarative project tool, single-app (no groups yet). init, pull/config --effective, manifest parse/validate, plan (bound artifact), apply (atomic desired-state write — requires the manager-core post-commit-refresh restructuring of §4.2, domains/files excluded from the transactional core), prune, secrets push, link state, env-scoped config. Adds enabled to scripts/routes + the three-state runtime + the dispatcher fire-time enabled re-check (§4.3) + trigger name column/backfill + the enabled/secrets conflict bit + the net-new content + tree-structure version counters + a coarse per-instance apply lock; in-flight executions finish (no kill).

  2. Groups as pure org/RBAC/UI container. Nested groups (single-parent, parent_id, delete = RESTRICT, reparent/rename with the ancestor-walk cycle guard + slug-freeze + tree-structure version, §5.6), inherited membership, hierarchy-aware can, UI grouping, the §9 backfill. No shared resources yet — cheap, no data-plane schema change.

    Status (Phase 2): shipped. Migration 0047_groups.sql adds groups (self-FK parent_id ON DELETE RESTRICT, instance-global frozen slug, structure_version, inert owner_project §7 seam) + group_members (same app_admin|editor|viewer literals as app_members) + apps.group_id (NOT NULL, RESTRICT) with the §9 single-root backfill. The tree lives in group_repo (reparent: ancestor-walk cycle guard under a coarse instance-wide structural advisory lock, slug-freeze, structure_version bump; delete=RESTRICT) and group_members_repo. Hierarchy-aware RBAC (§5.3): AuthzRepo::effective_app_role / effective_group_role resolve the highest role across the app's own membership + every ancestor group via one depth-bounded recursive CTE; authz::can's Member path folds inherited group roles, so a group_admin on any ancestor is implicitly app_admin beneath it — resolved live per request (revocation is immediate). New caps InstanceCreateGroup / Group{Read,Write,Admin} (no app_id ⇒ bound API keys can't manage groups). Surfaced via groups_api, pic groups …, and a dashboard group tree + detail/members. structure_version is bumped but not yet a consumer beyond the reparent path (CLI structural-drift detection is Phase 5).

    Deferred to Phase 3+ (unchanged): group-owned scripts/vars/secret-refs, the materialized effective-view resolver, masked group secrets + the group-secret AAD scheme, module/import resolution under inheritance, config --effective --explain, and personal-namespace root groups (the backfill seeds a single instance root). The §10 inherited-membership revocation-lag risk does not bite in Phase 2 because resolution is live (no inherited-role cache); it returns only when Phase 3's materialized views land.

  3. Group-inherited config (vars, secret-refs, env-scoped). The net-new vars/secret-refs tables + polymorphic owner; the group-secret AAD scheme (§5.3 caveat); masked group secrets; the effective-view resolver + materialization + invalidation; config --effective --explain (hard requirement, since multi-level resolution first ships here).

    Status (Phase 3): shipped — with one deliberate architecture change: LIVE resolution, not materialization. §5.1 / this bullet / §12 prescribed a materialized per-app effective view with a generation counter and fan-out invalidation. We shipped live, on-demand resolution instead: one depth-bounded recursive CTE (config_resolver::CHAIN_LEVELS_CTE, walking apps.group_id → groups.parent_id → root) runs per vars::get/secrets::get/config/effective, and the §3 merge/proximity/tombstone semantics run in a pure, unit-tested resolve() pass on the fetched rows. No materialized store, no generation, no invalidation protocol. Rationale: the group tree is small org structure, resolve-on-read is cheap, and it sidesteps §10's unsolved fan-out-invalidation SLA entirely. Consequence: §10's "effective-view invalidation SLA" and "inherited-membership revocation lag" are moot for config — with no cache, a security disable or secret rotation propagates immediately (same posture as Phase 2's live RBAC). Materialization (and a per-execution memo keyed by cx.execution_id) is re-scoped to a future optimization, to revisit only if a deep tree + hot path ever makes live resolution measurably costly.

    Shipped surface:

    • Schema: 0048_vars.sql (vars polymorphic owner [group_id XOR app_id], environment_scope, is_tombstone, partial-unique indexes) + apps.environment TEXT NULL (the explicit realization of "an environment is an app", §2 — set at app-create, read by the resolver, never entering SdkCallCx). 0049_group_secrets.sql reshapes the live secrets table to the same polymorphic-owner + environment_scope contract (PK → two partial-unique indexes); existing app rows backfill to app_id + scope '*' and decrypt byte-for-byte (the v1 AAD excludes the scope).
    • §3 resolver: env-filter in SQL (scope IN ('*', app_env)), then proximity-first + per-key deep-merge / replace / tombstone in Rust. A same-level @E value suppresses the same-level * (env is eligibility, not a merge tier — §3.1), incl. the map-vs-map case.
    • Vars: VarsService + vars::get/vars::all SDK + App/GroupVars{Read,Write} caps + {apps,groups}/{id}/vars admin CRUD + pic vars.
    • Group secrets: owner-discriminated AAD (secret:group:{group_id}:{name}; app AAD unchanged); inherited secrets::get (resolves app-own + ancestor-group, decrypts under the resolved owner's AAD, anchored to cx.app_id); the two orthogonal gates (§5.3): runtime injection (no human-read check) vs. the group-gated value endpoint GET /groups/{id}/secrets/{name}/value (GroupSecretsRead = group_admin). GroupSecretsWrite sits at script:write scope (matches its editor role tier).
    • Effective view (masked): GET /apps/{id}/config/effective → resolved vars (with provenance) + secrets masked to {status, owner, scope} (value never returned). pic config --effective [--explain] renders it.

    Deferred (unchanged): group-owned scripts/modules (Phase 4), the manifest [vars]/[secrets] apply reconciliation (Phase 5 — Phase 3 ships CRUD API + CLI), and materialization (above).

  4. Group-inherited scripts/modules. CoW overrides; the scope-aware module/import resolver + extension points (§5.5); cache-invalidation fan-out hardening; versioning/pinning if needed.

  5. Project tool maps onto groups. Nested manifests, attach point, single-owner, server-computed tree plan, per-env approval gating.

  6. (Much later) group-level collections/topics — the v1.3 cross-app data-sharing problem, with a real shared-scope authz model. Optionally, trigger/route templates (§4.5) if cardinality demands.

11.1 Re-sequencing review (post-Phase-3)

A review after Phase 3 shipped surfaced three adjustments to the 4→6 plan; carried here as decisions to take, not yet taken:

  • Phase 4 should resolve scripts live, not materialize. §11.4 carries "cache-invalidation fan-out hardening" as a deliverable, a holdover from the materialized-view model. Phase 3 proved live resolution (recursive CTE per call) is sufficient and sidesteps the §10 invalidation SLA, and script bodies are already live-resolved per request today (§5.1). So Phase 4 should resolve group-owned scripts live too and drop the fan-out-invalidation sub-project unless a measured hot-path cost appears. This removes the scariest item from Phase 4; the real remaining cost is the module/import resolver redesign (§5.5), which is unaffected by this and stays the gating risk.
  • Consider a "Phase 5-lite" before Phase 4. The declarative project tool already applies app scripts (Phase 1) and Phase 3 shipped the group vars/secrets CRUD, so a project tool that declares group vars/secrets + app scripts needs nothing from Phase 4 (group-owned scripts). The highest-leverage day-to-day workflow (declarative config apply) can ship ahead of the hard module-resolver work. Phase 4 then adds group-owned scripts to an already-working project tool.
  • Two deferred decisions to force, not default:
    • Trigger/route templates (§4.5): bundled into Phase 6 as "much later," but the per-app binding tax bites in Phase 5 at high tenant cardinality (100 tenants × 5 triggers = 500 declarations). Decide against actual target tenant counts before defaulting it late.
    • Multi-repo ownership (§7): the single-owner / attach-point / takeover machinery is substantial and only earns its keep with a second managing repo. For a solo-dev / single-repo start, a "one repo owns the whole subtree" model covers the near term; confirm the multi-repo need exists before building the ownership layer.

12. Contracts still to draft

  • The apply bundle / plan artifact / change-report wire contract (what the CLI ships, what the server persists and returns), including the conflict and blast-radius shapes.
  • The effective-view resolver (the read primitive) — the §3 rule made executable. (Phase 3: shipped as a live recursive-CTE resolver, config_resolver/secrets_repo::resolve; the materialization + invalidation protocol of §5.1 was not built — see the Phase 3 status note in §11.3. The remaining contract work here is only for script-body inheritance, Phase 4.)
  • The full manifest schema spelling every block (scripts, routes, the 8 trigger kinds, storage config, env-scoped vars, secret-refs, domains, [project.environments] + confirm policy).