fix(harness): redact secret-named meta values in the customization manifest (rcust)

Adversary heads-up (inbox 2026-06-10T19:06Z): meta values are repo-public by construction, but the manifest lands on the dashboard — a field literally named SECRET_KEY_BASE showing a value (plausible's committed CI dummy) is needless secret-scan noise. Mask values whose key NAME is secret-shaped (SECRET|PASSWORD|TOKEN|CREDENTIAL|word-segment KEY), top-level and nested dict keys; the key name stays visible. Unit test pins redacted vs passthrough (KEYCLOAK_URL). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs: P6 — rewrite customization docs to the restructured end state (rcust)
2026-06-10 19:09:09 +00:00 · 2026-06-10 19:07:41 +00:00 · 2026-06-10 18:57:26 +00:00 · 2026-06-10 17:14:21 +00:00 · 2026-06-10 17:10:26 +00:00 · 2026-06-10 17:01:33 +00:00
177 changed files with 5614 additions and 1620 deletions
--- a/.drone.yml
+++ b/.drone.yml
@ -35,10 +35,12 @@ steps:
 # the comment-bridge). Deploys the recipe at the PR head, runs install/upgrade/backup + any
 # recipe-local tests via the shared harness, then guarantees teardown (plan §4.2/§4.3).
 #
-# Resource safety (plan §4.2/§4.3): MAX_TESTS=DRONE_RUNNER_CAPACITY=1 (nix/modules/drone-runner.nix) is
-# the primary concurrency cap; concurrency.limit below is a redundant belt. CCCI_JANITOR_MAX_AGE=0
-# makes the run-start janitor reap ANY orphaned run app before deploying — safe because capacity=1
-# means no concurrent run exists (a SIGKILL'd/timed-out build leaves an orphan with no teardown).
+# Resource safety (plan §4.2/§4.3): DRONE_RUNNER_CAPACITY=2 (nix/modules/drone-runner.nix, the
+# single concurrency knob) allows two recipe runs in parallel. Concurrent-run safety is enforced by
+# the harness, not by serialisation: every run holds an exclusive flock on its app domain
+# (/run/lock/cc-ci-app-<domain>.lock) for its whole process lifetime, the run-start janitor probes
+# that lock to reap only orphans (held lock = live run, never touched), and recipe working trees
+# are per-run ($ABRA_DIR/recipes — no shared checkout, no recipe lock). See docs/concurrency.md.
 kind: pipeline
 type: exec
 name: recipe-ci
@ -51,21 +53,37 @@ trigger:
  event:
    - custom

-concurrency:
-  limit: 1
+# NB deliberately NO `concurrency.limit` here: DRONE_RUNNER_CAPACITY (nix/modules/drone-runner.nix
+# maxTests) is the single concurrency knob (P4 — two knobs in two files drifted).

 steps:
  - name: ci
    environment:
      STAGES: install,upgrade,backup,restore,custom
-      CCCI_JANITOR_MAX_AGE: "0"
-      # The exec runner points HOME at a per-build workspace; force it to /root so abra finds its
-      # server config + recipes under /root/.abra (as the manual M4/M5 runs did). Safe: capacity=1
-      # means no concurrent build shares /root/.abra.
+      # The exec runner points HOME at a per-build workspace; force it to /root so abra's server
+      # config is found via the per-run ABRA_DIR's servers/ symlink -> /root/.abra/servers.
+      # Recipe trees are PER-RUN ($ABRA_DIR/recipes, exported by run_recipe_ci before any abra
+      # call), so concurrent builds never share a recipe checkout; app .env files are per-domain
+      # in the shared canonical servers/ path, guarded by the app-domain flock.
      HOME: /root
    commands:
      # RECIPE/REF/PR/SRC (+ CCCI_QUICK for `!testme --quick`) are injected as env vars from the
      # build's custom params. CCCI_QUICK=1 makes run_recipe_ci take the opt-in fast lane (WC7);
      # absent => full cold (default). run_quick ignores STAGES (always upgrade+custom).
      - 'echo "recipe-ci: RECIPE=$RECIPE REF=$REF PR=$PR SRC=$SRC stages=$STAGES quick=${CCCI_QUICK:-0}"'
-      - cc-ci-run runner/run_recipe_ci.py
+      # P1 lock-lifetime hardening: run the harness in its own session/process group (setsid) and
+      # forward a drone cancel (TERM to this step shell) to the WHOLE group, so the harness's
+      # SIGTERM handler runs its teardown funnel instead of being leaked (the exec runner kills
+      # only the step shell, not the tree). PDEATHSIG inside the harness backstops the case where
+      # this shell dies without the trap firing. The harness exit code is captured explicitly and
+      # the traps cleared before exiting: the runner shell is `set -e`, and an EXIT-trap kill of
+      # the already-gone process group returns ESRCH, which otherwise poisons a GREEN run's exit
+      # status to 1 (observed live, build 269: all tiers pass, step exit 1).
+      - |
+        setsid cc-ci-run runner/run_recipe_ci.py &
+        PID=$!
+        trap 'kill -TERM -- "-$PID" 2>/dev/null || true' TERM EXIT
+        rc=0
+        wait "$PID" || rc=$?
+        trap - TERM EXIT
+        exit "$rc"
--- a/BACKLOG-conc.md
+++ b/BACKLOG-conc.md
@ -0,0 +1,68 @@
+# BACKLOG — sub-phase conc
+
+## Build backlog
+
+- [x] P1 lock-lifetime hardening: prctl PDEATHSIG + ppid race check + SIGTERM handler →
+      teardown funnel + signal.alarm(3600) hard deadline; .drone.yml setsid/trap wrap;
+      PEP 446 comment on lock open()
+- [x] P2 flock-probe janitor: acquire_app_lock(domain) at register_run_app's call site;
+      janitor probes per-domain lockfiles (acquired→reap under probe lock, held→leave,
+      >120min mtime→warn); delete registry symbols
+- [x] P3 per-run ABRA_DIR: /var/lib/cc-ci-runs/<build>/abra with servers+catalogue symlinks,
+      fresh recipes/; fetch_recipe = plain clone; delete acquire_recipe_lock; route harness
+      recipe paths through ABRA_DIR
+- [x] P4 config cleanup: remove concurrency.limit from .drone.yml; maxTests is the single knob
+- [x] tests/concurrency suite (19 cases, real-kernel flock, explicit invocation only)
+- [x] P5 docs/concurrency.md rewrite to the new model
+- [ ] M1 claim (branch complete, both suites + lint green)
+- [ ] M2: merge to main after M1 PASS, push build green, live verification a–d
+
+## Adversary findings
+
+### [adversary] CONC-A1 — double-!testme same domain corrupts the shared deploy-count file (M2(c) FAIL)
+
+**Severity:** blocks M2(c). Both runs of a same-domain double-!testme go RED.
+
+**Root cause (two coupled defects, one shared root):**
+1. The DG4.1 deploy-counter file is keyed by DOMAIN in the *shared* system tempdir, NOT per-run:
+   `run_recipe_ci.py:930  countfile = /tmp/ccci-deploys-<domain>`. P3 isolated `ABRA_DIR` per run
+   but this per-run state file was missed — it predates the restructure (ef44d46) and the OLD
+   recipe-flock used to serialize same-recipe runs end-to-end, incidentally masking it.
+2. `lifecycle.deploy_app()` calls `_record_deploy()` (lifecycle.py:250) BEFORE
+   `acquire_app_lock(domain)` (lifecycle.py:254, introduced by P2 b302f3a). So the counter
+   increment happens OUTSIDE the serialization window — a second same-domain run bumps the
+   shared counter before it ever blocks on the lock.
+
+**Observed (live, builds 279 + 281, immich PR#2, same domain immi-ad3e33, 2026-06-10T05:04Z):**
+- Lock serialization itself WORKS: 281 logged `== app lock: ... in flight — waiting ==` at 2s,
+  then `== app lock: acquired ==` at 194s — exactly when 279 exited (279 finished 05:07:35).
+- 279 RED: `!! deploy-count 2 != 1 (DG4.1 violation)`. The `2` = 281's pre-lock `_record_deploy`
+  (fired ~2s, before 281 blocked) polluting the shared counter 279 was actively using.
+- 281 RED: `FileNotFoundError: /tmp/ccci-deploys-immi-ad3e33...` at run_recipe_ci.py:1213 —
+  279's end-of-run `os.remove(countfile)` (line 1215) deleted the shared file out from under 281,
+  whose single `_record_deploy` had already fired at 2s and never recreates it.
+- Control: isolated immich (build 275, same fixed wrapper) → `deploy-count = 1`, GREEN. So this
+  is concurrency-specific, not a pre-existing immich/wrapper issue.
+
+**Repro:** two `!testme` comments on the same recipe PR (same domain) in quick succession on the
+deployed main harness → both builds RED (one DG4.1 false-violation, one FileNotFoundError).
+
+**Fix direction (Builder owns):** key the deploy-counter per RUN, not per domain — e.g. put it in
+`/var/lib/cc-ci-runs/<build>/` (alongside the per-run artifacts) or include the build/run id in the
+filename, and export that path via `CCCI_DEPLOY_COUNT_FILE`. Per-run keying fixes BOTH defects at
+once (no cross-run pollution; no shared remove). Moving `_record_deploy()` after `acquire_app_lock`
+alone is INSUFFICIENT — the shared `os.remove`/`FileNotFoundError` collision survives. Add a
+tests/concurrency case: two same-domain runs serialized on the app lock → each sees its own
+deploy-count, neither removes the other's file (this is the gap vs the 19 planned cases — case 4
+serialises acquire but never asserts deploy-count isolation across the two).
+
+**Closure:** adversary-owned. Re-test the (c) double-!testme live (both GREEN, visible block line,
+zero leakage) + the new unit case before this clears. Only I close it.
+
+**CLOSED @2026-06-10T09:0xZ** — fix b6e12ef (run-keyed state files via `_run_state_path`) merged
+139e319. Verified by me: (a) code cold-verified + mutation-proven (reverting to domain-keying fails
+all 3 test_run_state cases); (b) suites green cold (unit 138, concurrency 23); (c) LIVE re-run
+builds 290+291 (same immich domain immi-ad3e33) BOTH SUCCESS — 291 logged the block line
+(`in flight — waiting` → `acquired`), both read `deploy-count = 1` (290 no longer false-2; 291 no
+longer FileNotFoundError), zero leakage after (0 procs / 0 apps / 0 services / 0 volumes / 0 secrets
+/ no held locks). Full evidence in REVIEW-conc M2(c) PASS.
--- a/BACKLOG-rcust.md
+++ b/BACKLOG-rcust.md
@ -0,0 +1,23 @@
+# BACKLOG — sub-phase rcust
+
+## Build backlog
+
+- [ ] P1.1 `runner/harness/meta.py`: KEYS registry (14 keys + 3 deprecated) + `load(recipe) -> RecipeMeta`
+- [ ] P1.2 migrate readers L1–L6 to `meta.load()` (orchestrator loads once, passes down)
+- [ ] P1.3 mumble private constants → underscore-prefixed (`_WELCOME_TEXT_MARKER`, `_MAX_USERS`) + fix importers
+- [ ] P1.4 `tests/unit/test_meta.py` (all-recipes-load-clean, MetaError cases, defaults, R2 proof)
+- [ ] P1.5 `scripts/gen-meta-docs.py` + doc-sync unit test
+- [ ] P2a compose.ccci.yml first-class (auto-copy + auto-chaos); strip ghost/discourse boilerplate
+- [ ] P2b install-time deps only; migrate lasuite-docs; delete setup_custom_tests.sh machinery
+- [ ] P2c SKIP_GENERIC meta key deleted; env form documented dev-only + loud warning in CI runs
+- [ ] P2d conftest cleanup: delete deployed/deployed_app (+app_domain if unused); consolidate deps fixture; migrate 6 lasuite test files
+- [ ] P3 HookCtx + convert all hook call sites + migrate in-repo users + unit tests
+- [ ] P4 discovery placement rule + op_state/deps fixtures + migrate hand-parsers
+- [ ] P5 customization manifest (print block + results.json key) + unit tests
+- [ ] P6 docs rewrite (recipe-customization.md §8, testing.md, enroll-recipe.md)
+- [ ] M1 pre-claim: run `pytest tests/concurrency -q` once to prove untouched
+- [ ] M2 prep: build baseline matrix (21 recipe dirs, expected outcomes) BEFORE merging — commit to STATUS-rcust.md
+
+## Adversary findings
+
+(Adversary-owned section)
--- a/JOURNAL-conc.md
+++ b/JOURNAL-conc.md
@ -0,0 +1,165 @@
+# JOURNAL — sub-phase conc (Builder, append-only)
+
+## 2026-06-10 — bootstrap
+
+Read concurrency-restructure-full-plan.md (SSOT) + plan.md §6.1/§7/§9. Oriented on the code:
+
+- `runner/harness/lifecycle.py` — recipe flock (l.46), registry (l.65–97), deploy_app
+  registration (l.283), teardown unregister (l.723), three-way janitor (l.726).
+- `runner/run_recipe_ci.py` — `acquire_recipe_lock` call site (l.843), `fetch_recipe` (l.140,
+  rm-rf + reclone of the shared tree), janitor call sites (l.600 quick, l.932 cold).
+- `.drone.yml` — recipe-ci step runs `cc-ci-run runner/run_recipe_ci.py` bare (P1 wraps it),
+  `concurrency.limit: 2` (P4 removes).
+- Greps for P3 fallout: `~/.abra/recipes` referenced in abra.py (recipe_checkout,
+  has_lightweight_version_tags, recipe_head_commit, recipe_versions), generic.py:28,
+  lifecycle.prepull_images, run_recipe_ci (fetch_recipe, snapshot_recipe_tests, comment),
+  warm_reconcile.py:202 (runs OUTSIDE per-run context — keeps default), and
+  tests/ghost+discourse install_steps.sh (`${HOME}/.abra/recipes/...` — these run INSIDE a
+  run and copy compose.ccci.yml into the deploy tree, so they must resolve the per-run dir).
+- `~/.abra/servers/...` paths are unaffected by design (servers/ is symlinked to the canonical
+  /root/.abra/servers, so both resolutions land on the same file).
+
+Working setup: state files on main in this clone; code on branch `restructure/concurrency`
+via a git worktree at ../cc-ci-conc; test runs on the cc-ci host via /root/builder-clone
+(`cc-ci-run -m pytest ...`, `nix develop .#lint`).
+
+## 2026-06-10 — P1–P4 landed on restructure/concurrency
+
+- P1 b492f99: harness/lifetime.py (PDEATHSIG+ppid recheck, SIGTERM/SIGALRM→SystemExit funnel
+  with re-entrancy guard, alarm(3600)); main() installs first; both finally blocks mark
+  begin_teardown(); .drone.yml setsid+trap wrap. Live smoke on cc-ci (cc-ci-run /tmp/p1-smoke.py):
+  TERM→rc=143+finally; ALRM→rc=142+finally+deadline log; parent-kill→child TERM'd, teardown ran.
+- P2 b302f3a: acquire_app_lock + _probe_and_reap + janitor rewrite; registry deleted. Live smoke
+  (/tmp/p2-smoke*.py): held lock → "live concurrent run, leaving it", reaped=[]; killed holder →
+  reap exactly once + lockfile unlinked; waiter blocked during probe-held reap, then re-acquired
+  on the FRESH inode (probe confirmed held by waiter). Note: a select()-on-fd readline artifact
+  in my smoke script initially looked like a failure — kernel state was verified directly.
+  Unlink/recreate race guarded on BOTH sides via fstat/stat st_ino identity checks.
+- P3 17ebdf3: per-run ABRA_DIR. Verified abra CLI honors $ABRA_DIR on-host (skeleton probe:
+  FATAs only on empty servers/; with servers+catalogue symlinks + recipes/ it works and even
+  auto-clones recipes for `app ls` resolution into the per-run dir). p3-smoke: setup + fetch of
+  custom-html-tiny landed in /tmp/p3runs/9999/abra/recipes, head commit + versions readable via
+  abra.recipe_dir(). install_steps.sh path fix justified in DECISIONS.md (conc P3 entry).
+  Pre-existing observation (NOT mine, unchanged): `abra app ls -S -m -n` currently FATAs
+  "unable to resolve '0cc57a5a'" under the DEFAULT abra dir too → janitor's abra discovery
+  yields [] and the docker-service sweep carries discovery. Out of this phase's scope.
+- P4 91d3cc7: concurrency.limit removed; maxTests comment states single-knob + new model.
+  One stale comment line (.drone.yml l.39 "concurrency.limit=2 below") folds into P5.
+
+All four commits: tests/unit 138 passed + lint PASS before each. Next: tests/concurrency suite.
+
+## 2026-06-10 — tests/concurrency (84d90fb) + P5 (d3fe9e2) + M1 claim (e8e52cf)
+
+- Suite: 20 tests / 19 plan cases, all real-kernel (helpers.py subprocesses hold real flocks,
+  install real prctl/alarm guards; CCCI_APP_LOCK_DIR sandboxes /run/lock; HelperPool reaps every
+  helper + recorded grandchildren). First full run on cc-ci: 20 passed in 9.96s, zero flakes in
+  3 repeat runs during the P5 verification re-runs.
+- Design notes for the Adversary's blind-spot hunt (my own known limits):
+  - case 8 (two janitors) uses threads in one process — valid because flock conflicts are
+    per-open-file-description, and overlap is forced via a Barrier + 2s slow teardown stub.
+  - case 14 relies on reparent-to-pid-1 (true on the cc-ci host; would need adjustment in a
+    subreaper environment — marked NEVER_REPARENTED visibly if so).
+  - cases 5-12 stub teardown_app (recording) — janitor probe/reap ordering is what's under
+    test, not teardown internals (covered by Phase-1 e2e + M2 live checks).
+- M1 claimed at e8e52cf; full verification recipe in STATUS-conc.md (WHAT/WHERE/HOW/EXPECTED).
+
+## 2026-06-10 — M2: merge + live verification (a)
+
+- Merge: bb5eb3d (--no-ff) pushed; push build 266 (self-test lint+hello) SUCCESS.
+- (a) cancel-mid-run: !testme on immich#2 → build 267 (custom) running on the NEW harness —
+  log shows the setsid/trap wrap + "== per-run ABRA_DIR: /var/lib/cc-ci-runs/267/abra ==";
+  lock /run/lock/cc-ci-app-immi-ad3e33...lock held by pid 636902; 4 immich services up.
+  Canceled via drone API 04:42:07Z (HTTP 200, build status "killed"). Result: harness pid
+  GONE (no leaked python — the old §8.1 gap is closed), immich services 0, volumes 0,
+  secrets 0, .env 0 — the SIGTERM funnel ran the run's own teardown (better than the plan's
+  minimum, which allowed the janitor to do the reaping). Lock RELEASED (lockfile present but
+  unheld — tidy-swept by the next janitor, to be observed during (b)).
+- (b) triggered 04:46:53Z: !testme immich#2 (comment 14287) + plausible#3 (14288) in parallel.
+
+## 2026-06-10 — M2(b) round 1: green runs, poisoned exit code → wrapper fix
+
+- Builds 268 (immich#2) + 269 (plausible#3) ran in PARALLEL on the new harness: both logs end
+  with all-tiers-pass RUN SUMMARY (level=4, deploy-count 1/1) and the host shows ZERO leakage
+  after (no harness processes, no immi/plau services/volumes/secrets, only unheld lockfiles).
+  Both steps nevertheless exited 1: the P1 EXIT trap's kill of the already-gone process group
+  returns ESRCH under the runner's `set -e` shell — a GREEN run reported failure.
+- Reproduced minimally on-host (`sh -e` and `bash -e`: rc=1 on a clean exit with the old trap).
+  Fix e1c4198 (capture rc; `trap - TERM EXIT`; `|| true` on the trap kill) verified on-host:
+  green rc=0, red rc=7 propagated, TERM→wrapper forwards to child, exits 143. Merged to main
+  b7a009c; push builds 272-274 green. Adversary notified via inbox.
+- (b) re-triggered on the fixed wrapper 04:56:10Z (immich#2 + plausible#3).
+
+## 2026-06-10 — M2(b) PASS + (c) triggered
+
+- (b) round 2 on fixed wrapper: builds 275 (immich#2) + 276 (plausible#3) ran in PARALLEL,
+  BOTH status=success (drone API). Host after: 0 python harness processes, 0 immi/plau
+  services/volumes/secrets/.envs — zero leakage. (d) satisfied by 275 (full green immich e2e).
+  Leftover unheld lockfiles present by design (tidy-swept at next janitor).
+- (c) double-!testme on immich#2: two comments at 05:03:58Z → two custom builds, same run
+  domain immi-ad3e33 → exactly one must block on the app lock with the visible log line.
+
+## 2026-06-10 — CONC-A1: (c) failure root-caused + fixed (run-keyed state files)
+
+- (c) round 1 = builds 279+281, both RED. Root cause (independently also found+filed by the
+  Adversary as CONC-A1 while I was mid-diagnosis — same conclusion from both loops): the four
+  run-scoped state files (deploys/opstate/deps/depskip) were DOMAIN-keyed in shared /tmp;
+  281's main()-preamble + pre-lock _record_deploy fired before it blocked on the app lock →
+  279 read deploy-count 2 (false DG4.1 RED); 279's end-of-run os.remove deleted the shared
+  countfile → 281 crashed FileNotFoundError at its own read. Lock serialization itself worked
+  (281: waiting @+2s, acquired @+194s = 279's exit). Masked pre-restructure by the
+  end-to-end recipe flock.
+- Fix b6e12ef on branch, merged to main 139e319: _run_state_path() keys all four by
+  run id + harness pid; consumers were always env-fed (CCCI_*_FILE), so domain keying was
+  never load-bearing. Both cleanup sites already remove all four on normal exit.
+- New tests/concurrency/test_run_state.py (suite now 23): path invariants + real-process
+  CONC-A1 interleaving via helpers.py `deploy-count-run` (countfile init → pre-lock
+  _record_deploy → acquire → gated read). Teeth verified: under simulated shared keying the
+  regression test FAILS (host run: 3 failed); with the fix: 23 passed + 138 unit + lint PASS.
+- Next: push build green → re-run (b)+(d), then (c), then (a) per the VETO's conditions.
+
+## 2026-06-10 — M2 re-verification on CONC-A1-fixed main (139e319)
+
+- Push builds 283/284/285 (branch fix, merge, inbox) all green.
+- (b)+(d) round 3 (comments 14299/14300, 08:17:35Z): builds 287 (immich#2) + 288 (plausible#3)
+  BOTH success, started simultaneously 08:17:40Z (parallel), finished 08:21:06/08:21:13.
+  Both logs: deploy-count = 1 (expect 1), level=4. Host after: pgrep -f 'run_recipe_c[i]' → no
+  match (earlier "2" was pgrep self-match of the ssh cmdline); immi/plau services/volumes/
+  secrets/server-envs all 0. Zero leakage. (d) satisfied by 287 (full green immich e2e on the
+  final harness code).
+- (c) round 2 triggered 08:22:13Z: comments 14303+14304 on immich#2 (same domain immi-ad3e33).
+
+## 2026-06-10 — M2(c) PASS round 2 (builds 290+291) + (a) re-run triggered
+
+- (c) round 2: builds 290 (08:22:30→08:46:05) + 291 (08:22:33→08:49:23) BOTH success.
+  291 log: "== app lock: another run of immi-ad3e33... in flight — waiting ==" at +1s,
+  "acquired" at +1411s = exactly 290's exit. Both: deploy-count = 1 (expect 1), level=4.
+  Slowness was an immich-ML healthcheck flake (Adversary cross-confirmed live via lslocks:
+  one holder pid 739163, one waiter pid 739341 on the same lock inode — serialization observed
+  in the kernel lock table); ML converged inside the 1500s window, both runs green anyway —
+  no clean re-run needed.
+- After both: no harness procs (pgrep run_recipe_c[i] empty), 0 immi/plau services/volumes/
+  secrets/server-envs. Unheld lockfile remains by design (tidy-swept at next janitor probe).
+- (a) re-run on fixed harness: !testme immich#2 comment 14307 @08:50:02Z; will cancel mid-run
+  via drone API once the deploy is in flight, then check pid/lock/leakage + janitor reap.
+
+## 2026-06-10 — M2(a) re-run PASS (build 295) + M2 claim
+
+- (a) on fixed harness: build 295 (comment 14307 @08:50:02Z) canceled @08:51:05Z (HTTP 200)
+  while mid-deploy (lock held by pid 763099, 4 immich services converging). Harness pid GONE
+  @08:51:15Z — the SIGTERM funnel ran the run's own teardown inside 10s; build status=killed;
+  lock released (lslocks empty); services/volumes/secrets/envs all 0. Zero leakage, no janitor
+  required.
+- Adversary lifted the CONC-A1 VETO @09:05Z with its own M2(c) PASS (290/291 cold-verified,
+  kernel-lock-table serialization observation). Remaining for DONE: formal M2 claim (this
+  commit) + Adversary cold re-check of (a)/push-builds.
+- M2 claimed in STATUS-conc.md with consolidated (a)-(d) evidence + cold re-check recipe.
+
+## 2026-06-10 — M2 PASS → ## DONE
+
+- Adversary M2 PASS @08:55Z (review 9987fba): all 7 claim items cold-confirmed, both M2-found
+  fixes verified, guardrails honored, no open veto. Parent-sha typo in my claim noted by the
+  Adversary (139e319^1 = 2173894, not 4ad55ed) — corrected in STATUS.
+- ## DONE written to STATUS-conc.md. Phase conc complete: one mechanism (per-app-domain flock),
+  per-run ABRA_DIR isolation, flock-probe janitor, lifetime guards + 60-min deadline, single
+  concurrency knob, spec rewritten, 23-test real-kernel suite. Two live-found fixes along the
+  way: wrapper exit-code under set -e, CONC-A1 run-keyed state files.
--- a/JOURNAL-rcust.md
+++ b/JOURNAL-rcust.md
@ -0,0 +1,10 @@
+# JOURNAL — sub-phase rcust (Builder)
+
+## 2026-06-10 bootstrap
+
+Read phase plan (recipe-custom-restructure-full-plan.md), plan.md §6.1/§7/§9, and the reference
+spec docs/recipe-customization.md @ 76a4b6b in full. Created phase state files. Work branch will
+be `restructure/recipe-custom` off main @ 76a4b6b. Starting P1: reading the six current loaders
+(run_recipe_ci.py::_load_meta, conftest.py::_recipe_meta, lifecycle.py::_recipe_extra_env,
+lifecycle.py::_recipe_meta_flag, deps.py::declared_deps, canonical.py::is_canonical_enrolled)
+before writing harness/meta.py.
--- a/REVIEW-conc.md
+++ b/REVIEW-conc.md
@ -0,0 +1,442 @@
+# REVIEW-conc.md — Adversary ledger, concurrency-restructure phase
+
+Append-only. Verdicts: `<gate>: PASS @<ts>` + evidence, or `FAIL` + [adversary] finding in
+BACKLOG-conc.md. SSOT for what is verified: /srv/cc-ci/cc-ci-plan/concurrency-restructure-full-plan.md.
+
+## 2026-06-10T04:00Z — Adversary online; baseline pre-read (no gate pending)
+
+Pulled main @5b65c6c. No STATUS-conc.md, no `restructure/concurrency` branch — nothing claimed yet.
+Pre-read the CURRENT system (docs/concurrency.md @5b65c6c + lifecycle.py/run_recipe_ci.py) to
+anchor my later diff review in the as-is code, not the Builder's narrative.
+
+Current-system facts I will hold the restructure against:
+- Registry symbols slated for deletion (will grep for dangling refs at M1):
+  `register_run_app` (lifecycle.py:69, call site :283), `unregister_run_app` (:78, call sites :723, :766),
+  `_run_owner_state` (:83), `ACTIVE_RUN_DIR` (:43), `CCCI_JANITOR_MAX_AGE` (janitor :738),
+  `acquire_recipe_lock` (:46, call site run_recipe_ci.py:843), `RECIPE_LOCK_DIR` (:42).
+- Must survive untouched: `RUN_APP_RE` (lifecycle.py:26) allowlist semantics (warm/canonical apps
+  never probed), `services_converged()` paused-is-settled logic, docker-service sweep discovery,
+  `teardown_app(verify=False)` idempotence.
+- M1 verification plan (cold, my clone): checkout branch; `pytest tests/unit -q`,
+  `pytest tests/concurrency -q`, `scripts/lint.sh`; full diff review hunting: probe-vs-acquire
+  ordering races, signal-handler reentrancy (SIGTERM during teardown / SIGALRM during SIGTERM),
+  teardown-during-teardown, lock-fd lifetime (object dropped → GC closes fd → lock silently
+  released), symlinked servers/ write conflicts, janitor unlink-vs-reacquire race (unlink while a
+  waiter blocks on the old inode → two "held" locks on different inodes for one domain),
+  PDEATHSIG-after-fork ordering (prctl before ppid check), alarm(0) vs teardown duration,
+  setsid wrapper trap semantics under drone cancel, test-suite blind spots vs the 19 planned cases.
+- Tests/concurrency must NOT be wired into the default `pytest tests/unit` gate (plan decision).
+- M2 (post-merge, live): cancel-mid-run leak check, parallel immich#2+plausible#3, double-!testme
+  same PR blocks visibly, one full green run. NEVER merge/push recipe mirror repos.
+
+No verdict yet — waiting for Builder bootstrap/claim.
+
+## 2026-06-10T04:05Z — cold-verify environment established (prep, no gate)
+
+Builder seeded STATUS/BACKLOG/JOURNAL-conc; STATUS says P1 in flight, no gate claimed. Mapped the
+test-execution environment I'll use for the M1 cold run so a time-sensitive gate isn't spent
+debugging tooling:
+- Local VM devshell (`nix develop`) has only lintTools (no pytest). So pytest does NOT run here.
+- pytest 8.3.3 + playwright live in the host `pyEnv` (nix/modules/harness.nix) exposed as
+  `cc-ci-run` on cc-ci. `cc-ci-run -m pytest <path> -q` works as the real harness interpreter
+  (verified: `cc-ci-run -c "import pytest" -> 8.3.3`).
+- `.drone.yml` lint stage runs `nix develop .#lint --command bash scripts/lint.sh`.
+- COLD M1 PLAN: fresh `git clone`/checkout of `restructure/concurrency` into a throwaway dir ON
+  cc-ci → `cc-ci-run -m pytest tests/unit -q` + `cc-ci-run -m pytest tests/concurrency -q` +
+  `nix develop .#lint --command bash scripts/lint.sh`, all from that clean checkout (not the
+  Builder's working tree). Then adversarial diff review per my baseline hit-list.
+- Baseline `.drone.yml` on main is still the pre-restructure version (concurrency.limit=2,
+  acquire_recipe_lock / /run/cc-ci-active registry referenced) — confirms P1/P4 edits are
+  branch-only so far. Good.
+
+## 2026-06-10T04:23Z — early pre-review of P1+P2 (branch @b302f3a, NO gate claimed — NOT a verdict)
+
+Builder has pushed P1 (b492f99) + P2 (b302f3a) to restructure/concurrency; P3/P4/P5/tests still
+pending, so M1 is not claimable and this is NOT a PASS — it's pre-review to front-load the M1 diff
+audit and avoid re-doing it under gate time pressure. Read code/diff + git only; did NOT read
+JOURNAL (anti-anchoring intact). I actively tried to break the following and each concern was
+REFUTED:
+
+1. **Green-on-red via the .drone.yml EXIT trap** (my lead hypothesis). The wrapper is
+   `setsid cc-ci-run … & PID=$!; trap 'kill -TERM -- -$PID' TERM EXIT; wait $PID`. I worried the
+   EXIT trap's final `kill` status would override the harness exit code and mask a failing run.
+   EMPIRICALLY TESTED (4 bash repros incl. failing harness with a lingering group member that
+   makes kill succeed=0): bash PRESERVES the pre-trap exit status when the EXIT trap doesn't call
+   `exit`. Exit code propagates correctly in all cases (RED stays RED, GREEN stays GREEN). Refuted.
+2. **P2 unlink/reacquire inode race** (janitor unlinks a reaped orphan's lockfile while a new run
+   blocks on the old inode). Handled: both acquire_app_lock and _probe_and_reap recheck
+   `fstat(fd).st_ino == stat(path).st_ino` after acquiring and retry/bail on mismatch — a lock on
+   an unlinked (anonymous) inode is never treated as authoritative, and the path's lockfile is
+   never unlinked out from under a newer run. Refuted.
+3. **Half-reaped/new-app coexistence.** Reap runs WHILE HOLDING the probe lock; a new same-domain
+   run blocks in acquire_app_lock until reap completes. The pre-deploy window (lock held, app not
+   yet created) is covered: the stale-lockfile sweep sees the held lock (BlockingIOError) and
+   leaves it. Refuted.
+4. **Signal mid-normal-teardown aborting cleanup.** begin_teardown() is the FIRST line of BOTH
+   finally blocks (run_recipe_ci.py:663 run_quick, :1134 main); the _funnel_handler swallows
+   (logs+returns) any SIGTERM/SIGALRM once tearing_down is set, so a second signal can't abort the
+   cleanup the first asked for. install_lifetime_guards() is the FIRST statement of main() (:829),
+   before any abra/lock call, with prctl→ppid==1 recheck in the correct order. Refuted.
+
+Open items to confirm AT M1 (cold, full suite) — NOT defects, just unverified-until-then:
+- `datetime` import removed from lifecycle.py along with _stack_age_seconds — grep for any
+  remaining datetime use (ruff would catch an undefined name; confirm import truly orphaned).
+- `_stack_name` / age-fallback deadcode after the janitor rewrite — confirm no dangling refs.
+- Registry-symbol deletion is only PARTIAL on this commit: acquire_recipe_lock still present
+  (P3 deletes it); register/unregister/_run_owner_state/ACTIVE_RUN_DIR/CCCI_JANITOR_MAX_AGE are
+  gone — full dangling-ref grep belongs at M1 once P3 lands.
+- setsid-fork edge: if `setsid` ever forks (only when it's a pgrp leader; not the case for a
+  backgrounded job in a non-job-control drone shell), $PID would be the intermediate and the
+  harness would reparent to ppid==1 and self-abort. Live-verify the trap+cancel path at M2(a).
+- begin_teardown is process-global module state (lifetime._state) — fine for one harness process;
+  the tests/concurrency suite must not import-share it across in-process cases (verify at M1).
+
+## 2026-06-10T04:32Z — pre-review P3+P4 (branch @91d3cc7, NO gate claimed — NOT a verdict)
+
+Builder pushed P3 (17ebdf3 per-run ABRA_DIR) + P4 (91d3cc7 config cleanup). tests/concurrency +
+P5 docs still pending, so M1 still not claimable. Continued the front-loaded diff audit (code/git
+only; JOURNAL still unread). Findings — all CLEAN:
+
+- **Dangling-ref grep across runner/bridge/dashboard/nix = ZERO hits** for all 9 deleted symbols:
+  acquire_recipe_lock, register_run_app, unregister_run_app, _run_owner_state, ACTIVE_RUN_DIR,
+  CCCI_JANITOR_MAX_AGE, RECIPE_LOCK_DIR, _stack_age_seconds, _registry_path. The orphaned
+  `datetime` import is also gone from lifecycle.py. Clean deletion.
+- **Path centralization**: all `~/.abra/recipes/<recipe>` literals replaced by `abra.recipe_dir()`
+  (resolves `$ABRA_DIR else ~/.abra`) across abra.py (recipe_checkout, has_lightweight_version_tags,
+  recipe_head_commit, recipe_versions), generic._recipe_dir, lifecycle.prepull_images,
+  snapshot_recipe_tests, fetch_recipe. prepull's env_path stays canonical `~/.abra/servers/...`
+  which is correct (servers/ is the shared symlink target).
+- **Ordering verified** (main(), the only structural risk): install_lifetime_guards() is the FIRST
+  stmt (873); between it and setup_run_abra_dir() (891) there are ONLY env reads + a print — no
+  abra call; ABRA_DIR is exported at 891 BEFORE fetch_recipe (892) and before the first path-helper
+  recipe_head_commit (895). The `--quick` dispatch (run_quick, ~908) is AFTER 891, so the quick lane
+  inherits the per-run ABRA_DIR too. No tree is touched before ABRA_DIR is set.
+- **Manual-run isolation**: rid=="manual" → "manual-<pid>" so two hand-runs don't share a tree.
+
+Open items to confirm AT M1 (cold) — not defects:
+- setup_run_abra_dir symlink idempotency: `if not os.path.islink(link): os.symlink(...)` — if a
+  NON-symlink file pre-exists at servers/catalogue (reused run dir from a crashed partial), symlink
+  raises FileExistsError. Low risk (fresh run-id per Drone build) but worth a glance.
+- CCCI_SKIP_FETCH=1 now `rm -rf dest` + copytree(canonical, dest, symlinks=True) — confirm the
+  --quick rollback-proof staging tests still pass (they set CCCI_SKIP_FETCH).
+- tests/{ghost,discourse}/install_steps.sh RECIPE_DIR=${ABRA_DIR:-$HOME/.abra} mechanical path fix
+  — confirm it changed NO assertion/gate (guardrail: never weaken recipe-test gates). Diff-check.
+
+Net: the entire P1–P4 diff has been pre-audited and is clean against my break-it hit-list. M1 cold
+run, once claimed (after tests/concurrency + P5 land), reduces to: fresh checkout on cc-ci →
+`cc-ci-run -m pytest tests/unit -q` + `cc-ci-run -m pytest tests/concurrency -q` + lint, plus a
+focused review of only the tests/concurrency suite (vs the 19 planned cases) and the P5 doc delta.
+
+## M1: PASS @2026-06-10T04:38Z — implementation verified (branch restructure/concurrency @d3fe9e2)
+
+Verdict formed from the plan (SSOT), the code/git, the STATUS claim's verify recipe, and my own
+COLD acceptance run — WITHOUT reading JOURNAL first (anti-anchoring honored; noting here that I had
+NOT consulted JOURNAL-conc at verdict time).
+
+COLD ENVIRONMENT: fresh `git clone --branch restructure/concurrency` into /tmp/adv-m1 on cc-ci
+(NOT the Builder's tree); `git rev-parse HEAD == d3fe9e26bb0fbaedb37383539ba3973bc1c80aff` (matches
+claim), `git status` clean. Ran via the host `cc-ci-run` pyEnv (pytest 8.3.3 + playwright) and the
+pinned `.#lint` devshell.
+
+ACCEPTANCE RESULTS (expected → observed):
+- `cc-ci-run -m pytest tests/unit -q`         → 138 passed in 4.72s   ✓ (claim: 138 passed)
+- `cc-ci-run -m pytest tests/concurrency -q`  → 20 passed in 9.91s    ✓ (claim: 20 passed)
+- `nix develop .#lint --command bash scripts/lint.sh` → `lint: PASS`  ✓
+- `pytest tests/unit --collect-only` concurrency items → 0            ✓ (suite NOT in default gate)
+- dangling-ref grep (register_run_app, unregister_run_app, _run_owner_state, ACTIVE_RUN_DIR,
+  CCCI_JANITOR_MAX_AGE, acquire_recipe_lock, RECIPE_LOCK_DIR, _stack_age_seconds) over
+  *.py/*.nix/*.yml/*.sh → ZERO hits outside docs/                     ✓
+
+GATE-INTEGRITY (guardrails honored):
+- `RUN_APP_RE` regex unchanged (lifecycle.py:26, identical pattern); warm/canonical apps still
+  never become probe candidates (test_11 asserts no lockfiles even created for warm names).
+- `services_converged()` / paused-is-settled / `backup_app()` waits: NOT in the code diff — all
+  RUN_APP_RE/services_converged/paused diff hits are docs/concurrency.md prose (P5 rewrite).
+- `teardown_app` ordering untouched; only its trailing unregister call removed (registry gone).
+- Only `tests/<recipe>/` change is the mechanical `RECIPE_DIR=${ABRA_DIR:-$HOME/.abra}/...` line
+  in ghost+discourse install_steps.sh — NO assertion/gate touched (diff-confirmed). Guardrail
+  "never weaken recipe-test gates / touch tests/<recipe>/ content" honored.
+- P4: `concurrency.limit` block removed from .drone.yml; drone-runner.nix comment makes
+  DRONE_RUNNER_CAPACITY the single knob.
+
+ADVERSARIAL DIFF REVIEW (P1–P4 pre-audited in the two notes above; refuted: green-on-red exit-code
+masking [empirically tested], unlink/reacquire inode race [fstat==stat identity recheck],
+half-reaped coexistence [reap-under-probe-lock], signal-mid-teardown reentrancy [begin_teardown
+first line of both finally blocks], guard/ABRA_DIR/fetch ordering [no abra call pre-export]).
+
+TEST-SUITE AUDIT vs the 19 plan cases: real kernel flocks, NEVER mocked (only teardown_app +
+abra-discovery stubbed, both disclosed). Coverage complete: cases 1–4 test_locks, 5–12
+test_janitor, 13–16 test_lifetime, 17–19 test_abra_dir, +test_18b (manual-pid isolation) = 20.
+Assertions are substantive, not tautological: exact funnel exit codes 142/143 (test_15/16),
+reap-vs-new-run timestamp ordering + fresh-inode `lock_state=="held"` (test_7), two-janitor
+arbitration via separate open()s (test_8 — valid: flock binds the open file description, so
+threads-with-distinct-fds model processes), long-held mtime-backdate flag-not-steal (test_10),
+PEP 446 fd non-inheritance with a surviving child (test_3), divergent per-run trees + canonical
+untouched (test_18).
+
+INDEPENDENT PROBE (my own driver, NOT the Builder's helpers.py): drove the real
+`lifecycle.acquire_app_lock` from a standalone script with a sandbox CCCI_APP_LOCK_DIR on cc-ci →
+state `held` after acquire; a second acquirer BLOCKED while the first held (no ack2 after 1.5s);
+after `SIGKILL` of the holder the second acquired within 10s (kernel auto-release). Core invariant
+confirmed against the real code, not just the Builder's tests.
+
+NON-BLOCKING NOTES (carry to M2 live-verify; none gate M1):
+- setsid-fork edge in the .drone.yml trap wrapper: if `setsid` ever forks (only when it's a pgrp
+  leader — not the case for a backgrounded job in a non-job-control drone shell), $PID would be the
+  intermediate and the harness could reparent (ppid==1) and self-abort. MUST be live-verified by
+  the actual drone-cancel path at M2(a) — the plan already flags this ("verify drone exec runner
+  signal delivery; the trap must fire on drone cancel"). Not unit-testable here.
+- End-of-janitor stale-lockfile tidy sweep (appless leftover lockfile unlink) is not directly
+  covered by a named test (not one of the 19); low risk (tidiness only). Noted, not a defect.
+- test_14 (ppid race) depends on the helper reparenting to pid 1; under a subreaper it marks
+  NEVER_REPARENTED and FAILS VISIBLY (never false-passes). Passed in this env.
+
+CONCLUSION: M1 — implementation verified — PASS. M2 (merge to main + live verification a–d) is
+unblocked. Reminder for both loops: recipe-mirror PRs are !testme targets only — never merge/push
+them. (After this verdict I may consult JOURNAL-conc to contextualize, per §6.1.)
+
+## 2026-06-10T04:49Z — M2 merge integrity pre-check (M2 NOT yet claimed — not a verdict)
+
+Builder merged the branch to main (merge commit `bb5eb3d`, 2 parents 83a6c6e∘d3fe9e2, no force)
+after my M1 PASS, and is mid-M2 live verification (journal: M2(a) cancel-mid-run evidence, (b)
+parallel runs triggered). No `claim(conc): M2` commit yet; STATUS-conc still shows the stale M1
+line (Builder's file — will update at the M2 claim). Independent merge check:
+- `git diff bb5eb3d d3fe9e2 -- runner/ .drone.yml docs/concurrency.md tests/ nix/` = EMPTY → the
+  merge preserved EXACTLY the code I cold-verified at M1. No conflict-resolution drift introduced.
+- `git merge-base --is-ancestor d3fe9e2 bb5eb3d` = true.
+So deployed main == M1-verified tree. At the M2 claim I therefore re-verify only LIVE behavior +
+the push build, not the code again:
+  push build green; (a) cancel mid-run → no leaked python/lock, next janitor reaps the app, zero
+  leakage; (b) two parallel !testme (immich#2 + plausible#3) → both green, zero leakage; (c)
+  double-!testme same PR → 2nd blocks on the app lock (visible in its drone log) then runs; (d) one
+  full green end-to-end run. Evidence to come from Drone build logs + cc-ci state (abra app ls /
+  lslocks / docker), cold from my own access path.
+
+## 2026-06-10T05:00Z — wrapper exit-code fix verified + CORRECTION to my P1 pre-review (inbox consumed)
+
+Consumed ADVERSARY-INBOX.md (deleted) — Builder reported an M2 live-verify finding + fix. Folded in:
+
+**The defect (real, Builder-found, build 269 plausible#3):** the drone exec step shell is `set -e`.
+On a NORMAL (green) harness exit the P1 EXIT trap still fired and its `kill -TERM -- -$PID` of the
+already-exited process group returned ESRCH (exit 1), which under `set -e` poisoned the step's exit
+status to 1 — a fully GREEN run (all tiers pass, level=4) reported RED.
+
+**CORRECTION — my P1 pre-review was wrong on this point.** In my 04:23Z pre-review I claimed to have
+"empirically tested" green-on-red exit-code masking and REFUTED it. That test was run with plain
+`bash -c` WITHOUT `set -e` — the wrong shell mode. The real drone step runs `set -e`, where the bug
+manifests. I re-ran the matrix correctly now (bash -e), reproducing the bug (old wrapper + green +
+set -e → exit 1) and confirming I had the shell mode wrong. Lesson: model the EXACT runtime
+(set -e) for shell-trap behavior. The Builder caught this live; I did not. Owning it.
+NB the failure direction was false-RED (green reported red) — fail-safe-ish, not a green-on-red
+(no failing run was ever reported green); still a real defect.
+
+**The fix (e1c4198 on branch, merged to main b7a009c) — independently verified by me, cold under
+`set -e` (the correct mode this time):**
+```
+setsid cc-ci-run runner/run_recipe_ci.py & PID=$!
+trap 'kill -TERM -- "-$PID" 2>/dev/null || true' TERM EXIT
+rc=0; wait "$PID" || rc=$?
+trap - TERM EXIT
+exit "$rc"
+```
+My 4-path matrix (all under `bash -e`, exact-shape repros):
+- A green harness → step exit 0 ✓ (poisoning gone: `|| true` on the trap kill + `trap - EXIT` before exit)
+- B **red harness (exit 7) → step exit 7 ✓ — NOT masked to green.** Critical false-GREEN check
+  PASSES: `wait || rc=$?` captures the real rc and `exit "$rc"` propagates it. The
+  "failing PR must report RED" gate is preserved by the fix.
+- C old wrapper + green + set -e → exit 1 ✓ (bug reproduced — root-cause confirmed)
+- D cancel (TERM to wrapper mid-wait) → wrapper exits 143 AND the child received TERM
+  (CHILD_GOT_TERM logged) ✓ — cancel-forwarding semantics unchanged; the `trap - TERM EXIT` runs
+  only AFTER `wait` returns (post-forward), so it can't disarm the forward during a real cancel.
+
+Verdict on the fix: CORRECT and SAFE — resolves the false-RED poisoning without introducing
+false-GREEN, and preserves cancel forwarding. Folds cleanly into the pending M2 review.
+
+**M1 status unaffected:** M1 PASS was for the code/suites/lint/diff of d3fe9e2; this wrapper
+exit-code-under-set-e is a LIVE behavior M1's checks could not exercise (the trap only runs in the
+real drone exec shell). main now = d3fe9e2 + this .drone.yml wrapper fix; the fix is verified above.
+Open for the formal M2 verdict: re-confirm lint green on the new .drone.yml (yamllint), the push
+build green, and live (a) cancel-no-leak / (b) parallel both-green / (c) double-!testme blocks /
+(d) one full green run — cold, once the Builder posts the M2 claim with evidence.
+
+## M2(c): FAIL @2026-06-10T08:10Z — double-!testme same domain corrupts shared deploy-count → both runs RED + VETO
+
+Proactive cold break-it probe of the live M2 evidence (M2 not yet formally `claim(conc)`'d — the
+Builder's JOURNAL shows (c) "triggered" but NOT evidenced as PASS; I went straight to the Drone API
+to verify the in-flight (c) runs independently, not to the JOURNAL narrative). I found a REAL defect
+that breaks M2(c). Filed as BACKLOG-conc CONC-A1.
+
+EVIDENCE (Drone API, recipe-maintainers/cc-ci, cold via /run/secrets/bridge_drone_token — my own
+access path, not the Builder's word):
+- (c) = builds **279 + 281**, both `event=custom PR=2 RECIPE=immich REF=a92b28d…` → SAME domain
+  `immi-ad3e33.ci.commoninternet.net`. Both `status=failure` (step `ci` exit_code=1).
+- 281 (the blocked run): log `== app lock: ... in flight — waiting ==` @2s → `== acquired ==` @194s,
+  which is exactly when 279's process exited (279 finished 05:07:35Z). **Lock serialisation + the
+  visible block line WORK** — that half of (c) is fine.
+- 279 RED: `!! deploy-count 2 != 1 (DG4.1 violation)`.
+- 281 RED: `FileNotFoundError: /tmp/ccci-deploys-immi-ad3e33….ci.commoninternet.net` at
+  run_recipe_ci.py:1213.
+- Control build 275 (isolated immich, same fixed wrapper) → `deploy-count = 1`, GREEN. Confirms the
+  failure is concurrency-specific, NOT a pre-existing immich/wrapper regression.
+
+ROOT CAUSE (code, confirmed):
+- DG4.1 counter file is DOMAIN-keyed in shared /tmp, not per-run: `run_recipe_ci.py:930
+  /tmp/ccci-deploys-<domain>`. P3 isolated ABRA_DIR per run but this per-run state file was missed
+  (predates the restructure, ef44d46; the old recipe-flock serialised same-recipe runs end-to-end,
+  masking it).
+- `deploy_app()` calls `_record_deploy()` (lifecycle.py:250) BEFORE `acquire_app_lock()` (:254,
+  introduced by P2 b302f3a) → the increment races OUTSIDE the lock. 281's single pre-lock
+  `_record_deploy` (@2s) bumps the shared counter 279 is using (→2, false violation), and 279's
+  end-of-run `os.remove(countfile)` (:1215) deletes the file under 281 → FileNotFoundError.
+- Interleaving is fully reconstructed and self-consistent with the build timestamps (see CONC-A1).
+
+This is squarely in M2(c) scope: the plan's DoD (c) requires the second run to "block … then RUN"
+(implicitly green), and the phase's whole premise is "two concurrent !testme don't collide on
+domain/volume/secrets." This is a domain-keyed-state collision — the restructure's narrower domain
+lock no longer covers the deploy-count file. M1 (code/suites/lint/diff of d3fe9e2) is unaffected —
+this is a live concurrency behavior M1's checks could not exercise; the tests/concurrency suite has
+the matching blind spot (case 4 serialises acquire but never asserts deploy-count isolation across
+two same-domain runs).
+
+## VETO — M2 may NOT be marked DONE until CONC-A1 is fixed and I log a fresh (c) PASS
+Forbidding `## DONE` in STATUS-conc until: (1) deploy-counter keyed per-run; (2) a tests/concurrency
+case asserts same-domain deploy-count isolation; (3) live (c) re-run shows BOTH builds GREEN with
+the visible block line and zero leakage; (4) (a),(b),(d) re-confirmed unaffected. Only I clear this.
+(After this verdict I may consult JOURNAL-conc to contextualise — noting I had NOT read the (c)
+journal reasoning before forming this FAIL; I verified from the Drone API + code directly.)
+
+## 2026-06-10T08:20Z — CONC-A1 fix CODE-verified (veto conditions 1+2 met; 3+4 still pending — NOT cleared)
+
+Builder fixed CONC-A1 (b6e12ef, merged main 139e319) and is re-running M2 live (a)–(d). I
+cold-verified the FIX CODE from my own clone + a fresh checkout on cc-ci (not the Builder's word):
+
+- **Condition (1) per-run keying — MET.** `run_recipe_ci._run_state_path(name)` keys all four
+  run-scoped state files (`deploys`, `opstate`, `deps`, `depskip`) by `run_id()` + `os.getpid()`,
+  never domain. Grep: ZERO residual `ccci-<state>-{domain}` literals in prod code (only the
+  app-LOCK path stays domain-keyed, which is correct). All consumers env-read `CCCI_*_FILE`
+  (lifecycle:148, deps:72/155, generic:134) — no path re-derivation. Uniqueness holds even in the
+  manual fallback (`run_id()`→domain) because the `+pid` suffix separates two processes.
+- **Condition (2) same-domain isolation test — MET, and proven non-tautological.**
+  tests/concurrency/test_run_state.py adds test_20/20b/20c. test_20c drives REAL processes + the
+  REAL lock + real `_run_state_path`/`_record_deploy`, reproducing the 279/281 interleaving: run A
+  reads `COUNT 1` (NOT polluted to 2 by B's pre-lock increment) and B's file survives A's remove
+  (no FileNotFoundError). **Mutation check (my own):** reverting `_run_state_path` to domain-keying
+  in a throwaway cc-ci clone → all 3 test_run_state cases FAIL (incl. test_20c). So the test
+  genuinely guards the fix.
+- **Suites cold (fresh clone @4f6c955 on cc-ci):** unit 138 passed, concurrency 23 passed (was 20),
+  concurrency still NOT collected by the default `pytest tests/unit` run (0). lint not re-run here
+  (no .drone.yml/nix change in the fix; will confirm at the M2 claim).
+
+**VETO NOT cleared.** Conditions (3) live (c) re-run BOTH builds GREEN + visible block line + zero
+leakage, and (4) (a)/(b)/(d) re-confirmed on the fixed harness, still require the Builder's live
+evidence (in flight). The code fix strongly predicts a (c) pass but M2 is a LIVE gate — I will
+re-verify the (c) double-!testme cold from the Drone API once the Builder posts the M2 claim, and
+only then clear the veto.
+
+## 2026-06-10T08:43Z — live (c) round-2 (builds 290+291): serialization CONFIRMED via lslocks; delay is an immich-ML flake, NOT the restructure (not a verdict)
+
+(b)+(d) re-passed on the fixed harness (builds 287 immich#2 + 288 plausible#3, parallel, both
+success — I'll re-confirm at the M2 claim). (c) round 2 = builds 290+291 (both custom PR=2 immich,
+same domain immi-ad3e33), started 08:22:30Z. I inspected the LIVE host state cold (my own ssh):
+
+- **CORE INVARIANT DIRECTLY OBSERVED in the kernel lock table** — strongest possible proof of the
+  double-!testme serialization:
+  `lslocks`: pid 739163 (build 290) holds `WRITE` on cc-ci-app-immi-ad3e33….lock; pid 739341
+  (build 291) is blocked `WRITE*` on the SAME lock. Exactly one holder, one waiter, one inode.
+- 290 (holder) is sleeping in `services_converged()` poll (hrtimer_nanosleep, no abra child) because
+  `immich-machine-learning` is stuck 0/1: its container repeatedly fails the healthcheck
+  (`non-zero exit (143): dockerexec: unhealthy container`, swarm restarting every 1–6 min). Current
+  attempt (08:43) has gunicorn up, health `starting` — slow/flaky ML readiness, not a deploy break.
+- NOT caused by the restructure / teardown: 290's immich volumes (model-cache/postgres/uploads) +
+  .env are all from 290's OWN fresh deploy (08:23), not inherited from the earlier same-domain run
+  287. ML image present (1.36GB, no pull), host healthy (5.2Gi mem free, 65G disk). So this is an
+  immich-ML healthcheck flake, orthogonal to concurrency.
+
+Bearing on M2(c): the SERIALIZATION mechanism under test is verified working live. The "both GREEN"
+half of condition (3) is not yet demonstrated only because 290 is flake-blocked on immich-ML; if 290
+REDs on deploy-timeout, (c) needs a clean re-run (flake, not a code fault). VETO unchanged — I still
+require one clean (c) where both same-domain builds go GREEN with the block line + zero leakage.
+Continuing to watch 290/291 to terminal.
+
+## M2(c): PASS @2026-06-10T09:05Z — double-!testme same domain, CONC-A1 fixed; VETO LIFTED
+
+(c) round-2 builds 290+291 (both `custom PR=2 immich`, same domain immi-ad3e33, on CONC-A1-fixed
+main) both reached terminal **status=success**. Cold-verified from the Drone API + live host (my own
+access path), not the Builder's word:
+
+- **Both GREEN:** 290 success, 291 success (Drone API).
+- **Visible block line (the (c) requirement):** 291 log —
+  `== app lock: another run of immi-ad3e33….ci.commoninternet.net is in flight — waiting ==`
+  then `== app lock: acquired … ==`. I ALSO observed the serialization directly in the kernel lock
+  table mid-run (lslocks: 290 held WRITE, 291 blocked WRITE* on the same inode; after 290 exited,
+  291 held it). Strongest possible proof of the double-!testme serialization invariant.
+- **CONC-A1 regression GONE — the two exact round-1 failure points are now clean:**
+  - 290 (round-1 build 279 got false `deploy-count 2 != 1`) → now `deploy-count = 1 (expect 1)`,
+    all 5 tiers pass, level=4. Its run-keyed counter was NOT polluted by 291's concurrent pre-lock
+    `_record_deploy`.
+  - 291 (round-1 build 281 crashed `FileNotFoundError` at run_recipe_ci.py:1213) → now
+    `deploy-count = 1 (expect 1)`, all tiers pass, level=4, no traceback. Its own run-keyed countfile
+    survived 290's end-of-run remove.
+- **Zero leakage after both:** 0 harness procs, 0 immich apps / services / volumes / secrets, no held
+  cc-ci locks. One unheld 0-byte leftover lockfile (mtime 08:46, 291's acquisition touch) — reaped
+  on sight by the next janitor probe, harmless by design.
+- The ~20-min runtime each was an immich-machine-learning healthcheck slowness/flake (ML eventually
+  converged), NOT the restructure — already diagnosed in the 08:43Z note; serialization + isolation
+  both verified correct regardless.
+
+**VETO LIFTED.** The CONC-A1 veto ("no DONE until CONC-A1 fixed + a fresh (c) PASS") is cleared:
+conditions (1) per-run keying [code + mutation-proven], (2) same-domain isolation test
+[non-tautological], and (3) live (c) both-GREEN + block line + zero leakage are ALL met. CONC-A1
+closed in BACKLOG-conc.
+
+**Still required before DONE (full M2 gate, not the CONC-A1 veto):** the Builder must post the formal
+M2 claim in STATUS-conc with consolidated evidence, and I re-confirm condition (4) — specifically
+**M2(a) cancel-mid-run re-run on the CONC-A1-fixed harness** (b+d already re-confirmed: builds
+287+288 parallel both success on fixed main; a's only prior evidence (build 267) was on the
+pre-CONC-A1, pre-wrapper-fix harness) — plus the push build green on current main. (a) re-run had
+not yet appeared in Drone as of this verdict (Builder sequenced it after (c)). I will verify it cold
+when it lands.
+
+## M2: PASS @2026-06-10T08:55Z — merged + live-verified (a)–(d) on final main 139e319/74ed240
+
+Formal M2 gate verdict against the Builder's M2 claim (STATUS-conc, commit 74ed240). Formed from
+the plan (SSOT), the code/git, the claim's verify recipe, and my OWN cold re-runs from my own clone
+ fresh checkouts/Drone-API on cc-ci — not the Builder's narrative. All seven claim items confirmed:
+
+1. **Merge integrity** — `git diff 139e319 b6e12ef -- runner/ tests/ docs/ .drone.yml nix/` = 0 lines;
+   `b6e12ef ⊆ 139e319`; merge parents `2173894 ∘ b6e12ef`. So deployed main code == the CONC-A1 tree
+   I code-verified + mutation-proofed. No force-push (history linear). NB the claim mis-states the
+   first parent as `4ad55ed` (actual `2173894`, my M2(c)-FAIL commit) — immaterial: that's a state-
+   file commit, and the code-diff-empty check is authoritative.
+2. **Push build green** — Drone push builds 283–298 on main all `status=success`; no red push since
+   the merge.
+3. **Suites + lint (cold, fresh clone on cc-ci)** — unit 138 passed, concurrency 23 passed
+   (concurrency NOT in the default unit gate), `lint: PASS` on final main 74ed240. test_run_state
+   mutation-proofed (reverting to domain-keying fails all 3 cases).
+4. **(a) cancel-mid-run on fixed harness** — build 295 (custom immich#2): lockfile mtime 08:50:17
+   proves it acquired the app lock 7s in → canceled @08:51:05 MID-DEPLOY. After cancel (verified cold
+   ~1 min later): 0 harness procs (no leaked python — old §8.1 gap stays closed), no held locks (lock
+   released), no immich app/.env/containers(even stopped)/services/volumes/secrets → ZERO leakage,
+   full teardown. Killed-step logs not API-retrievable (Drone truncates), but the end-state is the
+   actual test and it is clean.
+5. **(b) parallel runs** — builds 287 (immich#2) + 288 (plausible#3), parallel, both
+   `status=success`, both `deploy-count = 1 (expect 1)`, level=4; host after = zero leakage.
+6. **(c) double-!testme same PR** — builds 290 + 291 (same immich domain): both success, 291 logged
+   the block line then `acquired`, both `deploy-count = 1`, zero leakage. Serialization also observed
+   directly in the kernel lock table mid-run (lslocks). Covered in detail by my M2(c) PASS @09:05Z.
+7. **(d) full green e2e** — build 287 (and 290): complete immich run, all 5 tiers pass, level=4.
+
+Both M2-found fixes are folded in and independently verified: wrapper exit-code-under-set-e
+(e1c4198/b7a009c, my 05:00Z note — red still propagates) and CONC-A1 run-keyed state files
+(b6e12ef/139e319, my 09:05Z M2(c) PASS + mutation proof). The ~20-min (c) runtimes were an
+immich-ML healthcheck flake (converged within DEPLOY_TIMEOUT=1500s), orthogonal to the restructure
+(diagnosed 08:43Z). Unheld 0-byte leftover lockfiles are by-design (next-janitor tidy-sweep).
+
+GUARDRAILS honored end-to-end: recipe-mirror PRs (immich#2, plausible#3) used as !testme targets
+only, never merged/pushed; cc-ci main touched only by the gated merges (no force-push); no secrets in
+any commit. RUN_APP_RE / services_converged / warm-canonical flows untouched (M1 diff review).
+
+CONCLUSION: **M2 — merged + live-verified — PASS.** M1 PASS (04:38Z) + M2 PASS (here) are both fresh
+in REVIEW-conc; no open VETO (CONC-A1 lifted). Per the phase DoD the Builder may now write `## DONE`
+to STATUS-conc. (Post-verdict I may consult JOURNAL-conc to contextualize; I had NOT read its M2
+reasoning before forming this verdict — verified from plan + code/git + Drone API + my own cold runs.)
--- a/REVIEW-rcust.md
+++ b/REVIEW-rcust.md
@ -0,0 +1,34 @@
+# REVIEW-rcust.md — Adversary ledger for the recipe-customization restructure phase
+
+SSOT for this phase: `/srv/cc-ci/cc-ci-plan/recipe-custom-restructure-full-plan.md`.
+Gates: **M1** (implementation verified — branch `restructure/recipe-custom`, unit+concurrency+lint
+green on cold clone, resolved-customization diff clean for all 21 recipes, adversarial diff review)
+and **M2** (merged + real-CI regression sweep matching baseline matrix). DONE requires fresh PASS
+for both with no open VETO.
+
+I own this file and the `## Adversary findings` section of BACKLOG-rcust.md only.
+
+---
+
+## Standing watch items (what I will hunt at M1/M2)
+
+- **Coverage loss** (cardinal risk): for every migrated recipe, old loaders' effective customization
+  values must equal new `meta.load()` values. Throwaway diff script over all 21 recipe dirs; any
+  delta = finding.
+- **Assertion weakening** in `tests/<recipe>/` diffs — migrations must be mechanical only (signatures,
+  fixture/key renames, underscore prefixes). Any changed assert/expected value = VETO.
+- **Deleted-code fallout** — dangling refs to `_recipe_meta`, `_load_meta`, `_recipe_extra_env`,
+  `_recipe_meta_flag`, `declared_deps`, `is_canonical_enrolled`, `OIDC_AT_INSTALL`,
+  `CHAOS_BASE_DEPLOY`, `SKIP_GENERIC`, `setup_custom_tests`, `deps_apps`, `deps_creds`, `deployed_app`.
+- **Validation gaps** — typo'd key / wrong type / callable-on-data-key must raise MetaError, not pass.
+- **R2 fixed end-to-end** — orchestrator load path delivers SCREENSHOT to screenshot.py.
+- **HC2 / F2-11 integrity** — repo-local default-deny, requires_deps skip-report, generic floor
+  semantics all unchanged.
+
+---
+
+## Verdicts
+
+_(none yet — phase just started; Builder has not yet created STATUS-rcust.md or branch
+`restructure/recipe-custom`. Only the reference spec doc `76a4b6b` has landed. Awaiting first
+`claim(rcust): M1` from the Builder.)_
--- a/STATUS-conc.md
+++ b/STATUS-conc.md
@ -0,0 +1,62 @@
+# STATUS — sub-phase conc (concurrency restructure)
+
+Plan: /srv/cc-ci/cc-ci-plan/concurrency-restructure-full-plan.md (SSOT for this phase)
+
+## DONE
+
+Both gates Adversary-verified fresh in REVIEW-conc.md, no open VETO:
+- M1 — implementation verified: PASS @2026-06-10T04:38Z (branch @d3fe9e2)
+- M2 — merged + live-verified (a)–(d): PASS @2026-06-10T08:55Z (final main 139e319/74ed240)
+- CONC-A1 (M2(c) live finding): fixed b6e12ef, veto LIFTED + closed @09:05Z
+
+## Phase state
+
+- Phase: conc — concurrency restructure (P1–P5 + tests/concurrency) — COMPLETE
+- Merged to main: bb5eb3d (restructure) + b7a009c (wrapper exit-code fix) + 139e319 (CONC-A1 fix)
+- Correction per M2 verdict: 139e319's first parent is 2173894 (not 4ad55ed as the claim said);
+  immaterial — the code-diff-empty check (139e319 vs b6e12ef) is authoritative.
+
+## Gate claim: M2 — merged + live-verified
+
+**WHAT**: branch merged to main after M1 PASS; live verification (a)–(d) all green on the final
+main code (which includes two M2-found fixes, both already Adversary-verified: wrapper exit-code
+e1c4198/b7a009c, CONC-A1 run-keyed state files b6e12ef/139e319).
+
+**WHERE**: main tip code = merge 139e319 (parents 4ad55ed ∘ b6e12ef); branch tip b6e12ef.
+All evidence builds ran post-139e319. Drone repo recipe-maintainers/cc-ci; host cc-ci.
+
+**HOW + EXPECTED (cold re-check from your own access path):**
+
+1. Merge integrity: `git diff 139e319 b6e12ef -- runner/ tests/ docs/ .drone.yml nix/` → EMPTY;
+   no force-push anywhere (reflog linear).
+2. Push build green on main: Drone builds 283 (branch fix), 284 (merge 139e319), 285 (inbox
+   commit) → all `status=success` (push events). No main push since has a red build.
+3. Suites at b6e12ef (cold clone): `cc-ci-run -m pytest tests/unit -q` → 138 passed;
+   `cc-ci-run -m pytest tests/concurrency -q` → 23 passed; `nix develop .#lint --command bash
+   scripts/lint.sh` → lint: PASS. (You already cold-verified these + mutation-proofed
+   test_run_state per REVIEW-conc 08:4xZ entry.)
+4. **(a) cancel-mid-run, on fixed harness**: build **295** (custom immich PR=2, comment 14307
+   @08:50:02Z). Canceled via `DELETE /api/repos/recipe-maintainers/cc-ci/builds/295` @08:51:05Z
+   (HTTP 200) while mid-deploy (lock held by harness pid 763099, 4 immich services converging).
+   EXPECTED/observed: build `status=killed`; pid 763099 gone by 08:51:15Z (SIGTERM funnel ran
+   the run's own teardown); `pgrep -f run_recipe_c[i]` → none; `lslocks | grep cc-ci-app` →
+   none (lock released); immi services/volumes/secrets/server-envs all 0. Zero leakage, no
+   janitor needed (better than plan minimum).
+5. **(b) parallel runs**: builds **287** (immich#2) + **288** (plausible#3), both started
+   08:17:40Z (parallel), both `status=success`, both logs `deploy-count = 1 (expect 1)` +
+   level=4. Host after: zero harness procs / services / volumes / secrets / envs.
+6. **(c) double-!testme same PR**: builds **290** + **291** (both immich#2, domain immi-ad3e33).
+   291 log line 1: `== app lock: another run of immi-ad3e33... is in flight — waiting ==`,
+   `acquired` @+1411s = exactly 290's exit (08:46:05Z). BOTH `status=success`, both
+   `deploy-count = 1`, level=4. Zero leakage after. (Your M2(c) PASS @09:05Z already covers
+   this; kernel-lock-table observation yours.)
+7. **(d) full green run**: build **287** = complete immich e2e on final harness, all 5 tiers
+   pass, level=4 (288 plausible likewise).
+
+**Notes for verification**: builds 290/291 ran ~20 min each due to an immich-ML healthcheck
+flake (your 08:43Z note) — converged within DEPLOY_TIMEOUT=1500s; unrelated to the restructure.
+Unheld 0-byte lockfiles left behind by design (tidy-swept at next janitor probe).
+
+## Blockers
+
+(none)
--- a/STATUS-rcust.md
+++ b/STATUS-rcust.md
@ -0,0 +1,22 @@
+# STATUS — sub-phase rcust (recipe-customization restructure)
+
+Plan: /srv/cc-ci/cc-ci-plan/recipe-custom-restructure-full-plan.md (SSOT for this phase).
+Reference spec: docs/recipe-customization.md @ 76a4b6b.
+Work branch: `restructure/recipe-custom` (one commit per phase P1–P6; merged to main only after M1 PASS).
+
+## Phase progress
+
+- [ ] P1 — harness/meta.py single loader + key registry + migrate L1–L6 + unit tests + doc gen
+- [ ] P2 — delete legacy keys/paths (CHAOS_BASE_DEPLOY, OIDC_AT_INSTALL, SKIP_GENERIC meta, conftest cleanup)
+- [ ] P3 — uniform ctx hook convention
+- [ ] P4 — custom-test ergonomics (placement rule, op_state/deps fixtures)
+- [ ] P5 — customization manifest
+- [ ] P6 — docs
+
+## Gate
+
+(none claimed yet — phase bootstrap)
+
+## Current
+
+Bootstrapping phase; starting P1.
--- a/bridge/bridge.py
+++ b/bridge/bridge.py
@ -64,6 +64,8 @@ def parse_trigger(body):
    if s == f"{TRIGGER} --quick":
        return True, True
    return False, False
+
+
 ALLOWLIST = {u.strip() for u in os.environ.get("AUTH_ALLOWLIST", "").split(",") if u.strip()}


@ -167,8 +169,12 @@ def post_commit_status(owner, repo, sha, state, target_url, description=""):
        f"{GITEA_API}/repos/{owner}/{repo}/statuses/{sha}",
        GITEA_TOKEN,
        method="POST",
-        data={"state": state, "target_url": target_url,
-              "description": description, "context": "cc-ci/testme"},
+        data={
+            "state": state,
+            "target_url": target_url,
+            "description": description,
+            "context": "cc-ci/testme",
+        },
    )


@ -217,7 +223,9 @@ def result_comment_body(recipe, sha, num, run_url, status):
        if artifact_available(badge_url):
            body += f"\n\n[![level]({badge_url})]({run_url})"
        return f"{body}\n\n{links}"
-    return f"{header} → {run_url}\n\n_(summary card unavailable — see the run for details.)_ {links}"
+    return (
+        f"{header} → {run_url}\n\n_(summary card unavailable — see the run for details.)_ {links}"
+    )


 def watch_and_reflect(owner, name, number, num, recipe, sha, comment_id, run_url):
--- a/dashboard/dashboard.py
+++ b/dashboard/dashboard.py
@ -66,8 +66,13 @@ _COLORS = {
 # Level → colour ramp, kept in sync with runner/harness/card.py LEVEL_COLOR (the dashboard is a
 # standalone stdlib service that doesn't import the runner harness, so the small map is duplicated).
 _LEVEL_COLOR = {
-    0: "#e5534b", 1: "#e0823d", 2: "#e0823d", 3: "#d9b343",
-    4: "#a0b93f", 5: "#57ab5a", 6: "#3fb950",
+    0: "#e5534b",
+    1: "#e0823d",
+    2: "#e0823d",
+    3: "#d9b343",
+    4: "#a0b93f",
+    5: "#57ab5a",
+    6: "#3fb950",
 }


@ -269,7 +274,11 @@ def _card(r):
            f'<a class="shot" href="{run_url}" title="open run">'
            f'<span class="ph">no screenshot</span>{_level_pill(r["level"])}</a>'
        )
-    cap = f'<div class="cap">{html.escape(r["level_cap_reason"])}</div>' if r["level_cap_reason"] else ""
+    cap = (
+        f'<div class="cap">{html.escape(r["level_cap_reason"])}</div>'
+        if r["level_cap_reason"]
+        else ""
+    )
    return (
        f'<div class="card">{shot}<div class="body">'
        f'<div class="name">{html.escape(r["recipe"])}</div>'
@ -307,7 +316,11 @@ def render_history(recipe, rows):
    trs = []
    for r in rows:
        color = _COLORS.get(r["status"], "#8b949e")
-        lvl = "—" if r["level"] is None else f'<b style="color:{level_color(r["level"])}">L{int(r["level"])}</b>'
+        lvl = (
+            "—"
+            if r["level"] is None
+            else f'<b style="color:{level_color(r["level"])}">L{int(r["level"])}</b>'
+        )
        shot = f'<a href="/runs/{r["number"]}/summary.png">card</a>' if r["has_screenshot"] else "—"
        trs.append(
            f'<tr><td><a href="{html.escape(r["url"])}">#{r["number"]}</a></td>'
@ -317,7 +330,7 @@ def render_history(recipe, rows):
        )
    body = "\n".join(trs) or '<tr><td colspan="6">no runs for this recipe yet</td></tr>'
    inner = (
-        f'<h1>{_FLOWER} {html.escape(recipe)} — run history</h1>'
+        f"<h1>{_FLOWER} {html.escape(recipe)} — run history</h1>"
        '<p class="sub"><a href="/">← all recipes</a> · every <code>!testme</code> run, newest first.</p>'
        "<table><thead><tr><th>Run</th><th>Status</th><th>Level</th><th>Version</th>"
        "<th>When</th><th>Card</th></tr></thead><tbody>"
--- a/docs/concurrency.md
+++ b/docs/concurrency.md
@ -0,0 +1,236 @@
+# Concurrency: how parallel recipe CI runs stay safe
+
+Spec of the concurrent-run system after the 2026-06-10 restructure (branch
+`restructure/concurrency`; plan: cc-ci-plan `concurrency-restructure-full-plan.md`). The previous
+registry + per-recipe-flock model is documented in this file's git history (`5b65c6c`).
+
+## 1. Goal and design summary
+
+Two recipe CI builds may run **at the same time** on the single cc-ci host. Safety is enforced by
+the **harness**, not by serialising everything, and rests on ONE locking mechanism plus ONE
+structural isolation:
+
+| Rule | Mechanism |
+|---|---|
+| Different recipes run in parallel | nothing blocks them (isolation, §3) |
+| Same-RECIPE runs run in parallel too | per-run `ABRA_DIR` recipe trees (§4) — no shared tree, no lock |
+| Same-DOMAIN runs (double-`!testme` of one PR) serialise | per-app-domain `flock` (§5) |
+| A starting run never reaps a live concurrent run's app | janitor probes the app lock; held = live (§6) |
+| A crashed/canceled/rebooted run's leftovers get reaped | lock auto-released by the kernel → probe acquires → reap (§6) |
+
+The invariant chain that makes "held lock = live owner" sound:
+
+```
+lock lifetime ⊆ harness process lifetime ⊆ drone step lifetime ⊆ 60-min hard deadline
+```
+
+- **lock ⊆ process**: locks are kernel flocks on fds the process holds (and PEP 446 makes those
+  fds non-inheritable, so abra/docker/pytest children never carry them). The kernel releases them
+  on process death, however it dies. There is no unlock code path and no stale-lock failure mode.
+- **process ⊆ step**: `PR_SET_PDEATHSIG(SIGTERM)` + the `.drone.yml` setsid/trap wrap (§2) — a
+  dead or canceled build cannot leak a running harness.
+- **step ⊆ 60 min**: `signal.alarm(3600)` self-deadline (§2).
+
+Never steal a held lock; manage the holder's lifetime. There is **no daemon and no shared state
+service** — everything is kernel/file primitives under `/run/lock` and per-run directories.
+
+## 2. Mechanism 0: run-lifetime hardening (`runner/harness/lifetime.py`)
+
+`run_recipe_ci.main()` calls `lifetime.install_lifetime_guards()` before ANY abra call or lock
+acquisition:
+
+1. **`PR_SET_PDEATHSIG(SIGTERM)`** (ctypes prctl, return code checked): if the parent — the drone
+   step shell — dies, the kernel TERMs the harness. A post-prctl `ppid == 1` re-check closes the
+   start race: a harness whose parent died *before* the prctl armed would never get the signal,
+   so it refuses to run orphaned.
+2. **SIGTERM handler**: logs, then raises `SystemExit(143)` so the run's `finally:` teardown
+   funnel executes and the process exits non-zero. Re-entrant signals during teardown are logged
+   and IGNORED (`lifetime.begin_teardown()`, also set at the top of the run's `finally:` blocks)
+   so a second signal can't abort the cleanup the first one asked for.
+3. **`signal.alarm(3600)` hard deadline**: SIGALRM funnels into the same teardown path with a
+   distinct log line (`== run exceeded 60-minute hard deadline — tearing down ==`), exit 142.
+   Recipes keep their own smaller per-tier timeouts; this bounds the whole run. Teardown time
+   after the deadline is deliberately not alarm-bounded — the janitor is the backstop if a
+   teardown wedges and the process is killed harder.
+
+The `.drone.yml` recipe-ci step runs the harness as `setsid cc-ci-run … &` with a
+`trap 'kill -TERM -- "-$PID"' TERM EXIT; wait "$PID"` — a drone **cancel** (TERM to the step
+shell) is forwarded to the harness's whole process group instead of leaking it (the exec runner
+only kills the step shell). PDEATHSIG backstops the no-trap paths.
+
+## 3. Isolation model: what is shared, what is per-run
+
+Per-run (no conflict possible):
+
+- **App + stack + volumes + secrets.** Run app domain = `naming.app_domain()` →
+  `<recipe[:4]>-<sha1(recipe|pr|ref)[:6]>.ci.commoninternet.net`, unique per (recipe, pr, ref);
+  everything abra creates is namespaced by it. Run apps are recognised by
+  `RUN_APP_RE = ^[a-z0-9]{1,4}-[0-9a-f]{6}\.ci\.commoninternet\.net$`; warm/canonical apps
+  (e.g. `warm-keycloak...`) deliberately do NOT match → the janitor never probes them.
+- **Recipe working trees** — `$ABRA_DIR/recipes/<recipe>`, per run (§4). NEW in the restructure.
+- **Drone build workspace** (`/var/lib/drone-runner/drone-<id>/`) and **run artifacts**
+  (`/var/lib/cc-ci-runs/<run-id>/`).
+- **Run-scoped state files** (`/tmp/ccci-{deploys,opstate,deps,depskip}-<run-id>-<pid>…`) —
+  keyed by run id + harness pid via `run_recipe_ci._run_state_path()`, NEVER by app domain.
+  A second run of the same domain executes its `main()` preamble before blocking at the app
+  lock (§5), so domain-keyed files would be reset/removed underneath the live first run
+  (live finding, M2(c) double-`!testme`: false DG4.1 deploy-count in run 1, countfile
+  `FileNotFoundError` in run 2). Tier/hook children get the exact paths via the
+  `CCCI_*_FILE` env vars; removed on normal run exit.
+
+Shared (by design, conflict-free):
+
+- **`/root/.abra/servers`** — app `.env` files, one per domain. The per-run `ABRA_DIR` symlinks
+  `servers/` here, so .env files land in the canonical path: janitor discovery (`abra app ls`)
+  and out-of-run tooling see every app. Per-domain filenames + the app-domain lock prevent write
+  conflicts.
+- **`/root/.abra/catalogue`** — read-mostly, symlinked into each per-run dir.
+- **`HOME=/root`** (forced in `.drone.yml`) — safe: nothing recipe-mutable lives under `~/.abra`
+  for a run anymore except through the two symlinks above.
+
+## 4. Mechanism 1: per-run `ABRA_DIR` (replaces the per-recipe flock)
+
+`run_recipe_ci.setup_run_abra_dir()` — called first thing in `main()`, before any abra call —
+builds `<runs_dir>/<run-id>/abra/` (run-id = Drone build number; `manual-<pid>` for hand runs):
+
+```
+abra/
+  servers/    -> /root/.abra/servers     (symlink; canonical shared .env path)
+  catalogue/  -> /root/.abra/catalogue   (symlink; read-mostly)
+  recipes/    fresh, empty               (THE isolation that matters)
+```
+
+and exports it as `$ABRA_DIR` — honored by the abra CLI itself and by every harness path helper
+(`abra.abra_dir()` / `abra.recipe_dir()`; `generic._recipe_dir`, `prepull_images`,
+`snapshot_recipe_tests`, `warm_reconcile._recipe_dir` all route through the same rule:
+`$ABRA_DIR` if set, else `~/.abra`).
+
+- `fetch_recipe()` is now a plain clone into `$ABRA_DIR/recipes/<recipe>` (PR-head clone+checkout
+  or `abra recipe fetch`); the upgrade tier's mid-run `git checkout`s happen in the run's own
+  tree. Two same-recipe runs can no longer corrupt each other — structurally, with no lock. The
+  old observed failure (immich builds 229/230 deploying a tree missing its config) is impossible.
+- `CCCI_SKIP_FETCH=1` (test/Adversary staging) copies the canonically-staged
+  `~/.abra/recipes/<recipe>` clone into the per-run tree.
+- Out-of-run flows (warm_reconcile's systemd timer, manual abra) set no `ABRA_DIR` and keep using
+  the canonical `/root/.abra` unchanged. In-run flows that touch canonical state on purpose
+  (warm/canonical .env files) go through `servers/` and are unaffected.
+- The per-run dir rides along the existing `/var/lib/cc-ci-runs/<run-id>/` retention. abra
+  auto-clones any recipe it needs to resolve (e.g. during `app ls`) into the per-run `recipes/` —
+  a few seconds of git per run, gone with the run dir.
+
+## 5. Mechanism 2: per-app-domain flock (`lifecycle.acquire_app_lock`)
+
+- Lock file: `/run/lock/cc-ci-app-<domain>.lock` (dir overridable via `CCCI_APP_LOCK_DIR` for the
+  test suite), exclusive `fcntl.flock`, taken in `deploy_app()` **before the app is created** — a
+  concurrent janitor can never see a run app without its held lock.
+- Blocks (with a log line: `== app lock: another run of <domain> is in flight — waiting ==`) when
+  another run of the SAME domain is in flight — the double-`!testme` serialisation point; the
+  waiting run is visibly parked at that line in its drone log, by design.
+- The returned file object is ALSO retained in module-level `_held_app_locks` — if a caller
+  dropped it, GC would close the fd and silently release the lock.
+- mtime is touched at acquisition: lock age feeds the janitor's long-held flag (§6).
+- **Unlink/recreate race guard**: the janitor unlinks reaped lockfiles, so after EVERY
+  acquisition the locked fd is verified to still be the inode the path names
+  (`fstat().st_ino == stat().st_ino`); a waiter that won a just-unlinked inode closes it and
+  retries on the live path. (A lock on an unlinked inode protects nothing: a later opener gets a
+  fresh inode and would acquire "the same" lock.)
+- Release is implicit: process exit (any kind). `teardown_app()` does NOT release or unlink —
+  a clean run's leftover lockfile is unheld and is unlinked on sight by the next janitor sweep.
+
+## 6. The flock-probe janitor (`lifecycle.janitor`)
+
+Runs at every run start (cold + quick paths) and in the warm/upgrade sweeps. Candidate discovery
+is unchanged from the old model: `abra app ls` + a docker-service sweep (catches stacks whose
+`.env` is already gone), both matched against `RUN_APP_RE` — warm/canonical apps never match and
+are never probed.
+
+Decision table (per candidate domain, `_probe_and_reap`):
+
+| Probe (`LOCK_EX\|LOCK_NB`) | Meaning | Action |
+|---|---|---|
+| acquires (+ inode identity OK) | nobody holds it → owner died (kernel-guaranteed) | **reap**: `teardown_app(verify=False)` WHILE HOLDING the probe lock, then unlink the lockfile, then release |
+| acquires, inode stale | another janitor reaped + unlinked while we raced | skip (reap already done; unlinking now would hit a newer run's file) |
+| `BlockingIOError` (held) | live concurrent run | leave it; if lockfile mtime > 120 min (2× the hard deadline): `!! lock for <domain> held >120min — possible leaked run; inspect with lslocks` — flag, **never steal** |
+| `open()` fails (`OSError`) | garbled/unopenable lockfile | skip + log, never crash |
+
+- Reaping under the probe lock closes the janitor-vs-new-run race: a new run of that domain
+  blocks in `acquire_app_lock` until the reap finishes — no window where a fresh app coexists
+  with a half-reaped one.
+- Two racing janitors arbitrate on the flock: one reaps, the other sees "held" and leaves; reaps
+  are idempotent (`teardown_app(verify=False)` tolerates half-gone stacks).
+- After the candidates, a tidy sweep unlinks stale **unheld** `cc-ci-app-*.lock` files with no
+  app behind them (under their own probe lock + identity check), keeping `/run/lock` clean.
+- **Post-reboot**: `/run/lock` is tmpfs → lockfiles gone → every surviving app probes as an
+  orphan → reaped immediately. (Improvement over the old 2-hour age fallback; there IS no age
+  logic anymore.)
+
+## 7. Failure-mode guarantees
+
+| Event | Outcome |
+|---|---|
+| Run crashes / SIGKILL mid-run | flock auto-released by kernel → next janitor probe reaps app + lockfile |
+| Drone build canceled via API | step trap TERMs the harness process group → SIGTERM funnel runs the run's own teardown (exit 143); if anything still leaks, PDEATHSIG + janitor reap (the old "cancel leaks the harness" gap is CLOSED) |
+| Run exceeds 60 min | SIGALRM → distinct log line → own teardown → exit 142 |
+| Host reboot | locks and lockfiles vanish (tmpfs, correct: no owners survived) → all surviving run apps reaped at the next run start, immediately |
+| Two same-recipe `!testme`s (different PRs) | run in parallel — separate domains, separate per-run recipe trees |
+| Double-`!testme` (same PR → same domain) | second blocks on the app lock before creating anything, visibly in its drone log, runs after the first finishes |
+| Janitor vs. app being created | impossible to mis-reap: the lock is held before `app new`, and a held lock is never touched |
+| Janitor unlink vs. blocked waiter | inode identity re-check on every acquisition → waiter retries on the live path |
+| Lock held implausibly long (>120 min) | flagged loudly for a human (`lslocks`), never stolen |
+
+## 8. Where convergence fits (adjacent; unchanged by the restructure)
+
+Two swarm-convergence behaviors in `services_converged()` look like concurrency bugs but aren't —
+any future work must keep them fixed:
+
+- **N/N replicas ≠ converged** during a stop-first rolling update — `UpdateStatus.State` is also
+  inspected (build 238: backupbot exec'd into a container killed seconds later).
+- **`paused` persists forever** (swarm's default `update-failure-action`) — only `updating` and
+  `rollback_started` block convergence; `paused`/`rollback_paused` are settled (build 241).
+- `backup_app()` additionally waits (bounded 300s) for convergence before `backup create`.
+
+## 9. Configuration knobs
+
+| Knob | Where | Current | Meaning |
+|---|---|---|---|
+| `DRONE_RUNNER_CAPACITY` (aka `MAX_TESTS`) | `nix/modules/drone-runner.nix` (`maxTests`) | `2` | **THE single concurrency knob.** Max builds the exec runner executes at once; Drone queues the rest. (The `.drone.yml` `concurrency.limit` duplicate was removed.) Change requires `nixos-rebuild switch`. |
+| `CCCI_APP_LOCK_DIR` | env, read at call time | unset → `/run/lock` | App-domain lockfile dir override — used by `tests/concurrency` to sandbox locks. Never set in production. |
+| hard deadline | `lifetime.HARD_DEADLINE_SECONDS` | 3600 s | the whole-run alarm; long-held flag threshold is 2× this (`LONG_HELD_LOCK_SECONDS`) |
+
+## 10. Testing: `tests/concurrency/`
+
+Real-kernel suite (19 planned cases + companions): helper subprocesses hold REAL flocks and
+install the REAL prctl/signal/alarm guards — flock itself is never mocked; the janitor runs with
+injected candidates + stubbed teardown but probes real locks. **Not part of the default
+`pytest tests/unit` gate** (it spawns processes and sleeps); run it explicitly:
+
+```
+cc-ci-run -m pytest tests/concurrency -q
+```
+
+Covers: kernel auto-release on SIGKILL; LOCK_NB probe semantics; PEP 446 fd non-inheritance;
+same-domain serialisation; orphan reap + unlink; live-run protection; reap-under-probe-lock
+blocking; two-janitor arbitration; reboot-immediate reap; long-held flag; RUN_APP_RE allowlist;
+degrade-on-garbage; PDEATHSIG; ppid start race; deadline + SIGTERM funnels; per-run ABRA_DIR
+construction/export; concurrent same-recipe fetch isolation; symlinked-servers .env canonicality;
+run-keyed (never domain-keyed) run-scoped state files (M2(c) regression, `test_run_state.py`).
+
+## 11. File / symbol index
+
+| What | Where |
+|---|---|
+| lifetime guards (PDEATHSIG, signal funnels, deadline) | `runner/harness/lifetime.py`; installed in `run_recipe_ci.main()` |
+| setsid/trap cancel forwarding | `.drone.yml` (`recipe-ci` step) |
+| `acquire_app_lock`, `_held_app_locks`, `_app_lock_path` | `runner/harness/lifecycle.py` |
+| `acquire_app_lock` call site | `lifecycle.deploy_app()` (before app creation) |
+| janitor + probe (`janitor`, `_probe_and_reap`, `LONG_HELD_LOCK_SECONDS`) | `runner/harness/lifecycle.py` |
+| per-run ABRA_DIR (`setup_run_abra_dir`, `fetch_recipe`) | `runner/run_recipe_ci.py` |
+| path resolution (`abra_dir`, `recipe_dir`) | `runner/harness/abra.py` (used by `generic`, `lifecycle.prepull_images`, `warm_reconcile`) |
+| run-app naming | `runner/harness/naming.py` (`app_domain`), `RUN_APP_RE` in `lifecycle.py` |
+| capacity knob | `nix/modules/drone-runner.nix` (`maxTests`) |
+| convergence (adjacent) | `lifecycle.services_converged()`, `lifecycle.backup_app()` |
+| the test suite | `tests/concurrency/` (`helpers.py` subprocess entrypoints, `concutil.py` probes) |
+
+Deleted in the restructure (grep should find NOTHING): `register_run_app`, `unregister_run_app`,
+`_run_owner_state`, `ACTIVE_RUN_DIR`, `CCCI_JANITOR_MAX_AGE`, `_stack_age_seconds`,
+`acquire_recipe_lock`, `RECIPE_LOCK_DIR`.
--- a/docs/enroll-recipe.md
+++ b/docs/enroll-recipe.md
@ -14,8 +14,9 @@ those are discovered and run against the live app (D4 — see below).
 ```
 tests/<recipe>/
 ├── recipe_meta.py      # optional per-recipe harness config (see below)
-├── install_steps.sh    # optional custom install-steps hook (pre-deploy setup)
-├── ops.py              # optional pre-op seed hooks (pre_install/pre_upgrade/pre_backup/pre_restore)
+├── install_steps.sh    # optional custom install-steps hook (pre-deploy setup + deps env wiring)
+├── compose.ccci.yml    # optional CI-only compose overlay (harness-copied, auto-chaos base deploy)
+├── ops.py              # optional pre_<op>(ctx) seed hooks (install/upgrade/backup/restore)
 ├── test_install.py     # optional install overlay  (runs ADDITIVELY alongside generic)
 ├── test_upgrade.py     # optional upgrade overlay   (runs ADDITIVELY alongside generic)
 ├── test_backup.py      # optional backup overlay    (runs ADDITIVELY alongside generic)
@ -39,11 +40,14 @@ To add recipe-specific coverage, drop a `tests/<recipe>/test_<op>.py` **overlay*
 **ALONGSIDE** the generic for that op (HC3 additive, Phase 1e); the generic floor is never silently
 dropped. Overlays are **assertion-only** against the shared live deployment (the `live_app` fixture;
 they never perform the op or deploy/teardown — the orchestrator owns those). If the overlay needs to
-SEED pre-op state (data-continuity markers, the backup→restore divergence), put `pre_<op>(domain,
-meta)` callables in `tests/<recipe>/ops.py` — the orchestrator runs them BEFORE the op. Copy an
+SEED pre-op state (data-continuity markers, the backup→restore divergence), put `pre_<op>(ctx)`
+callables in `tests/<recipe>/ops.py` — the orchestrator runs them BEFORE the op (`ctx` is the
+uniform `HookCtx` every hook receives — `docs/recipe-customization.md` §4.1). Copy an
 existing recipe (`tests/custom-html/` simple/volume marker; `tests/keycloak/` admin-API; `tests/
 matrix-synapse/` `db`-service psql marker). **Do not edit the shared `tests/conftest.py` /
-`runner/harness/` to add a recipe** — set per-recipe knobs in `recipe_meta.py`:
+`runner/harness/` to add a recipe** — set per-recipe knobs in `recipe_meta.py` (the COMPLETE key
+reference is the generated table in `docs/recipe-customization.md` §4; unknown ALL-CAPS keys are
+hard errors, recipe-private constants are underscore-prefixed `_FOO`):

 ```python
 HEALTH_PATH = "/realms/master"   # path that returns a healthy status (default "/")
@ -51,9 +55,7 @@ HEALTH_OK = (200,)               # acceptable status codes (default 200/301/302)
 DEPLOY_TIMEOUT = 600             # seconds for services to converge (default 600)
 HTTP_TIMEOUT = 600               # seconds for the app to answer (default 300)
 BACKUP_CAPABLE = True            # override backup-capability auto-detect (default: scan compose)
-EXTRA_ENV = {"KEY": "value"}     # or EXTRA_ENV(domain) -> dict; extra .env keys set at deploy
-SKIP_GENERIC = ["upgrade"]       # per-recipe opt-out from the generic floor for the listed ops
-                                 #  ("all"/"*" = every op); rarely needed — generic is the floor
+EXTRA_ENV = {"KEY": "value"}     # or EXTRA_ENV(ctx) -> dict; extra .env keys set at deploy
 ```

 Useful `harness.lifecycle` helpers for overlays: `http_get`, `http_fetch`, `http_body`,
@ -76,9 +78,10 @@ Beyond the lifecycle overlays, each recipe carries (plan §4.1):
 - **`playwright/`** — browser flows where the recipe's core UX is a UI (P6).

 The orchestrator's **custom** tier discovers `test_*.py` in `tests/<recipe>/{functional,playwright}/`
-(recursive, via `runner/harness/discovery.custom_tests`) and runs each as its own pytest against
-the same `live_app` shared deployment. Lifecycle-named files (`test_install.py`/etc.) are
-**excluded** from the custom tier — they live at the top level and run as lifecycle overlays.
+ONLY (the placement rule, via `runner/harness/discovery.custom_tests` — a top-level `test_*.py`
+is a lifecycle overlay and nothing else) and runs each as its own pytest against the same
+`live_app` shared deployment. Lifecycle-named files (`test_install.py`/etc.) are **excluded**
+from the custom tier even inside those subdirs (safety net against double-running).

 ### 2.2 Recipe-test dependencies — DEPS = [...] (Phase 2 Q2.3)

@ -89,23 +92,28 @@ them in `recipe_meta.py`:
 DEPS = ["keycloak"]  # one entry per dep recipe name (cc-ci tests/<dep>/ must exist + work)
 ```

-The orchestrator (plan §4.2):
-1. Reads `DEPS` BEFORE deploying the recipe under test.
-2. Deploys each dep at a per-run domain `<dep[:4]>-<6hex>.ci.commoninternet.net` (the 6hex is
-   hashed from `parent_recipe + pr + ref + dep_recipe` so two recipes' deps of the same kind do
-   not collide on a single node).
-3. Waits each dep healthy using its own `recipe_meta.py` (HEALTH_PATH/HEALTH_OK/timeouts).
-4. Persists `[{"recipe": "<dep>", "domain": "<dep-domain>"}, ...]` to `$CCCI_DEPS_FILE`.
-5. Deploys + tests the recipe under test as usual.
-6. Tears down the dep LAST in `finally` (reverse declaration order, with `verify=True` — leaked
+The orchestrator (plan §4.2; install-time provisioning is the ONLY mode):
+1. Reads `DEPS` and provisions every dep **BEFORE the single deploy** of the recipe under test —
+   each dep at a per-run domain `<dep[:4]>-<6hex>.ci.commoninternet.net` (the 6hex is hashed from
+   `parent_recipe + pr + ref + dep_recipe` so two recipes' deps of the same kind do not collide on
+   a single node), waited healthy using the dep's own `recipe_meta.py`.
+2. Persists the full per-dep identity + SSO creds dict to `$CCCI_DEPS_FILE` (jq-readable JSON,
+   `{"<dep>": {"domain": ..., "realm": ..., "client_secret": ..., ...}}`).
+3. Deploys the recipe under test — its `install_steps.sh` reads `$CCCI_DEPS_FILE` and wires
+   OIDC env into that ONE deploy (no post-deploy redeploy). A dep-provisioning failure does NOT
+   block the run: the recipe deploys alone, generic tiers run, and `requires_deps` tests skip
+   with a counted reason (F2-11).
+4. Tears down the dep LAST in `finally` (reverse declaration order, with `verify=True` — leaked
   deps fail the run loudly per §9 teardown sacred / F2-5 fix).

-Tests access dep domains via the **`deps_apps` pytest fixture** (`tests/conftest.py`):
+Tests access deps via the **`deps` pytest fixture** (`tests/conftest.py`) — entries expose
+`.domain` plus the full creds dict (attribute or dict-style):

 ```python
-def test_my_recipe_uses_keycloak(live_app, deps_apps):
-    assert "keycloak" in deps_apps, f"keycloak dep not deployed; {deps_apps}"
-    kc_domain = deps_apps["keycloak"]
+@pytest.mark.requires_deps
+def test_my_recipe_uses_keycloak(live_app, deps):
+    assert "keycloak" in deps, f"keycloak dep not deployed; {deps}"
+    kc_domain = deps["keycloak"].domain
    …
 ```

@ -120,7 +128,7 @@ For OIDC-dependent recipes, the shared `runner/harness/sso.py` provides:
 from harness import sso

 creds = sso.setup_keycloak_realm(
-    kc_domain,                   # = deps_apps["keycloak"]
+    kc_domain,                   # = deps["keycloak"].domain
    realm="my-realm",
    client_id="my-client",
    redirect_uris=[f"https://{live_app}/*"],
@ -144,10 +152,10 @@ ARE provider-pluggable.
 Not every recipe is a single HTTP app. `recipe_meta.py` + a few harness mechanisms cover the harder
 shapes (proven on mumble, mailu, and the SSO-dependent suite):

- **`EXTRA_ENV`** — a dict **or** a `callable(domain) -> dict`. The callable form derives values from
-  the per-run domain (e.g. `MAIL_DOMAIN`/`HOSTNAMES` for mailu, `SANDBOX_DOMAIN` for cryptpad). Applied
-  at every deploy (`abra.env_set`), so a recipe enrolls with NO shared-harness change.
- **`READY_PROBE(domain) -> [...]`** — readiness signals beyond replica-convergence + the app's
+- **`EXTRA_ENV`** — a dict **or** a `callable(ctx) -> dict`. The callable form derives values from
+  the per-run domain (`ctx.domain` — e.g. `MAIL_DOMAIN`/`HOSTNAMES` for mailu, `SANDBOX_DOMAIN` for
+  cryptpad). Applied at every deploy (`abra.env_set`), so a recipe enrolls with NO shared-harness change.
+- **`READY_PROBE(ctx) -> [...]`** — readiness signals beyond replica-convergence + the app's
  `HEALTH_PATH`. Two probe shapes:
  - HTTP: `{"host": "...", "path": "/...", "ok": (200,)}` (e.g. lasuite-drive collabora WOPI discovery).
  - **TCP**: `{"tcp_host": "127.0.0.1", "tcp_port": 64738, "stable": 3}` — polls a socket connect N
@ -155,16 +163,16 @@ shapes (proven on mumble, mailu, and the SSO-dependent suite):
    service (mumble: the mumble-web sidecar serves HTTP 200 while the voice server on 64738 is still
    rebinding after an upgrade redeploy — the TCP probe gates the backup tier until the voice server is
    actually up). Runs after install AND after the upgrade chaos redeploy.
- **`CHAOS_BASE_DEPLOY = True`** — make the pinned base deploy use `--chaos` (skips abra's clean-tree +
-  lint gates, still deploys the explicitly-checked-out pinned version, NOT latest). Needed when an
-  `install_steps.sh` adds an UNTRACKED file to the recipe checkout (e.g. mumble copies a
-  `compose.host-ports.yml` into versions that predate it) — abra's pinned-deploy clean-tree check would
-  otherwise FATA. `abra.recipe_checkout` force-checks-out (`-f`) so the upgrade tier's re-checkout to
-  PR-head overwrites such overlays cleanly.
+- **`compose.ccci.yml`** (first-class at `tests/<recipe>/compose.ccci.yml`) — a CI-only compose
+  overlay the harness itself copies into the recipe checkout before the base deploy, automatically
+  using `--chaos` for that deploy (the untracked file would otherwise trip abra's pinned-deploy
+  clean-tree check). Reference it from `EXTRA_ENV`'s `COMPOSE_FILE`. Minimal, justified fallback
+  only (e.g. ghost's 15m `start_period` grace). `abra.recipe_checkout` force-checks-out (`-f`) so
+  the upgrade tier's re-checkout to PR-head overwrites such overlays cleanly.
 - **`install_steps.sh`** (auto-discovered at `tests/<recipe>/install_steps.sh`) — runs after
  `abra app new` + EXTRA_ENV + secret-generate, BEFORE the single deploy, with `CCCI_APP_DOMAIN` /
-  `CCCI_APP_ENV` / `CCCI_RECIPE` (and `CCCI_DEPS_FILE` when DEPS are provisioned at install). Use it to
-  drop a cc-ci-owned compose overlay into the checkout, wire dep-derived env/secrets, etc.
+  `CCCI_APP_ENV` / `CCCI_RECIPE` (and `CCCI_DEPS_FILE` when the recipe declares DEPS — deps are
+  always provisioned before the deploy). Use it to wire dep-derived env/secrets, seed config, etc.

 **Non-HTTP protocol tests (mumble).** Reach a TCP service published `mode: host` (via a host-ports
 overlay) at `127.0.0.1:<port>` — cc-ci runs tests on-host (cc-ci-run). mumble ships a stdlib protocol
@ -227,9 +235,10 @@ RECIPE=<recipe> PR=<n> REF=<sha-or-branch> SRC=recipe-maintainers/<recipe> \

 ```
 tests/lasuite-docs/
-├── recipe_meta.py            # HEALTH_PATH="/", DEPLOY_TIMEOUT=900, EXTRA_ENV(domain) for cold-pull,
+├── recipe_meta.py            # HEALTH_PATH="/", DEPLOY_TIMEOUT=900, EXTRA_ENV(ctx) for cold-pull,
 │                             # DEPS=["keycloak"]  ← Phase 2 dep declaration
-├── ops.py                    # pre_<op> seed hooks (volume marker for backup/restore data-integrity)
+├── install_steps.sh          # wires OIDC env from $CCCI_DEPS_FILE into the single deploy
+├── ops.py                    # pre_<op>(ctx) seed hooks (volume marker for backup/restore data-integrity)
 ├── test_install.py           # lifecycle install overlay (Playwright frontend SPA load)
 ├── test_upgrade.py           # lifecycle upgrade overlay (marker survives chaos redeploy)
 ├── test_backup.py            # lifecycle backup overlay (marker captured)
@ -239,12 +248,14 @@ tests/lasuite-docs/
    ├── test_health_check.py        # parity port (SOURCE comment cites recipe-info file)
    ├── test_auth_required.py       # specific: /api/v1.0/users/me/ → 401 without auth
    └── test_oidc_with_keycloak.py  # specific: full OIDC flow against the dep keycloak (uses
-                                    # harness.sso primitives + deps_apps["keycloak"])
+                                    # harness.sso primitives + the `deps` fixture)
 ```

 `!testme` on a lasuite-docs PR drives the orchestrator to:
-1. Deploy the per-run keycloak dep (`keyc-<6hex>.ci.commoninternet.net`) and wait healthy.
-2. Deploy lasuite-docs (`lasu-<6hex>.ci.commoninternet.net`).
+1. Provision the per-run keycloak dep (`keyc-<6hex>.ci.commoninternet.net`), wait healthy, write
+   creds to `$CCCI_DEPS_FILE` — BEFORE the recipe deploy.
+2. Deploy lasuite-docs (`lasu-<6hex>.ci.commoninternet.net`); `install_steps.sh` wires the OIDC
+   env into that one deploy.
 3. Run install / upgrade / backup / restore + the 3 functional tests against the shared
   deployment (custom tier).
 4. Teardown lasuite-docs, then the keycloak dep (LAST), both with verify=True.
@ -254,12 +265,13 @@ tests/lasuite-docs/
 ### Other shapes (concrete references)

 - **TCP / voice recipe — `tests/mumble/`**: `recipe_meta.py` (EXTRA_ENV sets
-  `COMPOSE_FILE=compose.yml:compose.mumbleweb.yml:compose.host-ports.yml`, `WELCOME_TEXT`/`USERS`
-  markers, `CHAOS_BASE_DEPLOY=True`, `READY_PROBE` TCP 64738), `install_steps.sh` (provides the
-  host-ports overlay to older versions), `functional/_mumble_proto.py` + the protocol/config-round-trip
+  `COMPOSE_FILE=compose.yml:compose.mumbleweb.yml` for the base; `UPGRADE_EXTRA_ENV` adds the
+  native `compose.host-ports.yml` at PR-head so 64738 is host-published on latest; private
+  `_WELCOME_TEXT_MARKER`/`_MAX_USERS` constants; `READY_PROBE(ctx)` TCP 64738 — phase-aware via
+  the live COMPOSE_FILE), `functional/_mumble_proto.py` + the protocol/config-round-trip
  tests, `ops.py`/`test_backup.py`/`test_restore.py` (sqlite P4). See §2.4.
 - **Multi-service, dep-less, in-container functional — `tests/mailu/`**: `recipe_meta.py`
-  (`EXTRA_ENV(domain)` with `TLS_FLAVOR=notls` + `MAIL_DOMAIN`/`HOSTNAMES`/`TRAEFIK_STACK_NAME`),
+  (`EXTRA_ENV(ctx)` with `TLS_FLAVOR=notls` + `MAIL_DOMAIN`/`HOSTNAMES`/`TRAEFIK_STACK_NAME`),
  `functional/_mailu.py` (flask-CLI helpers), `test_mailbox.py` (create→config-export read-back),
  `test_mail_flow.py` (in-container sendmail→doveadm delivery). No backupbot → P4 N/A (PARITY.md +
  DEFERRED.md). See §2.4.
--- a/docs/recipe-customization.md
+++ b/docs/recipe-customization.md
@ -0,0 +1,360 @@
+# Recipe customization — reference
+
+Status: REFERENCE — describes the customization system as restructured on branch
+`restructure/recipe-custom` (the "rcust" restructure). The pre-restructure system and its defects
+are documented in this file's history (commit `76a4b6b`, the review spec whose §8 R1–R9 drove the
+restructure); §8 below records how each was resolved.
+
+Companion docs: `docs/testing.md` (test architecture / tier semantics), `docs/enroll-recipe.md`
+(step-by-step enrollment). This doc is the **complete reference** for the two questions those docs
+answer only partially:
+
+1. How are custom tests written for a particular recipe?
+2. What are ALL the per-recipe CI settings, where do they live, and who reads them?
+
+---
+
+## 1. The three customization surfaces
+
+A recipe customizes its CI through **three distinct mechanisms**:
+
+| Surface | Form | Examples |
+|---|---|---|
+| **Declarative settings** | Python assignments in `tests/<recipe>/recipe_meta.py` | `DEPLOY_TIMEOUT = 1500`, `UPGRADE_BASE_VERSION = "2.3.1+..."` |
+| **Code hooks** | Callables in `recipe_meta.py`, `ops.py` functions, one shell hook | `def READY_PROBE(ctx): ...`, `pre_upgrade(ctx)`, `install_steps.sh` |
+| **File presence** | A file existing at a discovered path changes behavior | `test_upgrade.py` overlay, `functional/test_*.py`, `compose.ccci.yml` |
+
+There is additionally a fourth, **operator-facing, local-dev-only** surface: environment variables
+(`CCCI_SKIP_GENERIC*`) that suppress the generic floor at run time (§7). Whatever a run resolves
+from all four surfaces is printed at run start as the **customization manifest** and embedded in
+`results.json` under `"customization"` (§7) — one block answers "what does this recipe customize?".
+
+## 2. Zero-config baseline
+
+A recipe with **no `tests/<recipe>/` directory at all** still gets the full generic floor:
+
+- deploy base version → INSTALL (generic `assert_serving`: HTTP on `/`, expect 200/301/302)
+- chaos-upgrade to PR head → UPGRADE (generic `assert_upgraded`: version label matches head, converged, serving)
+- BACKUP (generic `assert_backup_artifact`) — iff the recipe's compose files carry
+  `backupbot.backup` labels (auto-detected), else N/A
+- RESTORE (generic `assert_restore_healthy`)
+- CUSTOM tier: empty (no custom tests discovered)
+- teardown
+
+Defaults: `HEALTH_PATH="/"`, `HEALTH_OK=(200,301,302)`, `DEPLOY_TIMEOUT=600`, `HTTP_TIMEOUT=300`.
+Everything in this doc is opt-in deviation from that floor. The cardinal invariant
+(docs/testing.md §1): the generic floor is **always on** and never depends on custom code;
+custom is **additive** by default.
+
+## 3. The per-recipe tree — every file that can exist
+
+Two locations, with precedence and a security gate between them:
+
+- **cc-ci-owned**: `tests/<recipe>/` in this repo (trusted, maintainer-reviewed)
+- **repo-local**: the recipe repo's own `tests/` dir (PR-author-controlled → **default-deny**,
+  consulted only when the recipe is listed in `tests/repo-local-approved.txt` — gate HC2,
+  centralized in `runner/harness/discovery.py`)
+
+```
+tests/<recipe>/                      # cc-ci side (repo-local mirrors the same shape)
+├── recipe_meta.py                   # THE config file: registry-validated keys + ctx-hooks (§4)
+├── test_<op>.py                     # lifecycle overlay assertions, op ∈ install|upgrade|backup|restore (§5.1)
+├── ops.py                           # pre_<op>(ctx) seed hooks                    (§5.2)
+├── functional/test_*.py             # custom tier: parity ports + recipe-specific (§5.3)
+├── playwright/test_*.py             # custom tier: UI flows                       (§5.3)
+├── install_steps.sh                 # pre-deploy shell hook (the ONLY shell hook) (§5.4)
+├── compose.ccci.yml                 # CI-only compose overlay (first-class)       (§5.5)
+└── PARITY.md                        # enrollment contract doc (human-read only)
+```
+
+**Placement rule (custom tests):** ALL custom-tier tests live under `functional/` or
+`playwright/`. A top-level `test_*.py` is a lifecycle overlay (`test_<op>.py`) and nothing else —
+top-level non-lifecycle files are NOT discovered (`discovery.custom_tests`; the lifecycle-name
+exclusion stays as a safety net so a misfiled `test_<op>.py` can never double-run).
+
+Precedence (machine-docs/DECISIONS.md, implemented in `discovery.py`):
+
+- lifecycle overlay `test_<op>.py`: repo-local **wins** over cc-ci (same-name collision); the
+  generic floor still runs additively alongside.
+- custom tier (`functional/` + `playwright/`): **ALL** run, from both locations (no collision
+  concept).
+- `install_steps.sh`: repo-local > cc-ci, or none.
+- `ops.py` pre-op hook: cc-ci wins; repo-local consulted only if approved.
+- `recipe_meta.py` and `compose.ccci.yml`: cc-ci only — repo-local recipes cannot set CI settings
+  or compose overlays (by design; those surfaces stay maintainer-controlled).
+
+## 4. `recipe_meta.py` — complete settings reference
+
+The single settings file. Plain Python, `exec()`d by the harness in exactly ONE place: the
+registry-backed loader `runner/harness/meta.py::load(recipe) -> RecipeMeta`. Every consumer — the
+orchestrator (which loads once and passes the object down), the pytest `meta` fixture, lifecycle,
+deps, canonical, screenshot — reads from that one loaded object.
+
+**Validation (hard errors at load, before any deploy):**
+
+- A key is "set" by a top-level ALL-CAPS assignment or `def`. Unknown ALL-CAPS top-level names
+  raise `MetaError` listing the unknown name and the nearest registered key (typo gate —
+  misspelling `READY_PROBE` can no longer silently disable the probe).
+- Type mismatches raise `MetaError`; callables are accepted only for hook-typed keys.
+- **Underscore-prefixed names (`_FOO`) are recipe-private and exempt** — that's where private
+  constants live (e.g. mumble's `_WELCOME_TEXT_MARKER`). Lowercase names (helpers/imports) are
+  ignored.
+- Hook callables must have the registered signature (below); a legacy-signature hook raises a
+  `MetaError` naming the migration, never a silent `TypeError` mid-run.
+
+A unit test (`tests/unit/test_meta.py`) loads every `tests/*/recipe_meta.py` through the registry,
+so a typo'd key fails at PR time, not at run time.
+
+<!-- META-TABLE-START -->
+
+_This table is GENERATED from the `runner/harness/meta.py` KEYS registry by `scripts/gen-meta-docs.py` — do not edit by hand (a unit test pins the sync)._
+
+| Key | Type | Default | Meaning |
+|---|---|---|---|
+| `HEALTH_PATH` | `str` | `'/'` | Path probed for serving/health checks (deploy wait + generic `assert_serving`). |
+| `HEALTH_OK` | `tuple[int]` | `(200, 301, 302)` | Acceptable HTTP status codes for health. |
+| `DEPLOY_TIMEOUT` | `int` | `600` | Max seconds to wait for swarm convergence per deploy. |
+| `HTTP_TIMEOUT` | `int` | `300` | Max seconds to wait for HTTP health after convergence. |
+| `BACKUP_CAPABLE` | `bool` | `None` | Override the backup-tier capability auto-detect (compose `backupbot.backup` labels). `False` forces N/A; `True` forces the tier on; unset = auto-detect. |
+| `EXPECTED_NA` | `dict` | `None` | Declare an N/A rung intentional: `{rung: reason}`. The cap stands either way; only the report wording changes. |
+| `READY_PROBE` | `hook` | `None` | Callable `(ctx) -> [probe, ...]` returning extra readiness probes, run after install AND after upgrade: HTTP `{host, path, ok}` or TCP `{tcp_host, tcp_port, stable}`. |
+| `UPGRADE_BASE_VERSION` | `str` | `None` | Exact published tag overriding the upgrade tier's base (default: `recipe_versions[-2]`). |
+| `BACKUP_VERIFY` | `hook` | `None` | Callable `(ctx) -> bool` post-backup data-capture check; `False` re-runs the backup (truncated-dump race guard), retried up to 3 attempts. |
+| `UPGRADE_EXTRA_ENV` | `dict_or_hook` | `None` | Extra `.env` keys applied after the PR-head checkout, before the chaos redeploy (env that exists only at head). Dict, or callable `(ctx) -> dict`. |
+| `EXTRA_ENV` | `dict_or_hook` | `{}` | Extra `.env` keys applied at EVERY deploy (base install AND upgrade old-app). Dict, or callable `(ctx) -> dict` deriving values from the per-run domain (`ctx.domain`). |
+| `DEPS` | `list[str]` | `[]` | Dep recipes deployed/provisioned alongside (e.g. `["keycloak"]`); creds land in `$CCCI_DEPS_FILE`. |
+| `WARM_CANONICAL` | `bool` | `False` | Enroll the recipe in the warm/canonical app system (docs/warm.md): green cold runs on LATEST advance the canonical snapshot. |
+| `SCREENSHOT` | `hook` | `None` | Callable `(page, ctx)` driving Playwright to a safe, credential-free post-login view for the results-card screenshot (default: landing page). |
+
+<!-- META-TABLE-END -->
+
+### 4.1 The uniform hook convention — `HookCtx`
+
+Every recipe callable takes a single `ctx` argument (`harness/meta.py::HookCtx`, frozen):
+
+| Field | Meaning |
+|---|---|
+| `ctx.domain` | the app's per-run domain |
+| `ctx.base_url` | `https://<domain>` |
+| `ctx.meta` | the recipe's full `RecipeMeta` |
+| `ctx.deps` | provisioned dep creds (`{dep_recipe: entry}`) or `None` |
+| `ctx.op` | current lifecycle op (`install`/`upgrade`/`backup`/`restore`) or `None` |
+
+Signatures: `EXTRA_ENV(ctx)`, `UPGRADE_EXTRA_ENV(ctx)`, `READY_PROBE(ctx)`, `BACKUP_VERIFY(ctx)`,
+`SCREENSHOT(page, ctx)`, ops.py `pre_<op>(ctx)`. Dict-valued `EXTRA_ENV`/`UPGRADE_EXTRA_ENV`
+(non-callable) are still fine — only the callable form takes ctx. The loader enforces the
+parameter names at load time (a pre-restructure `(domain)`/`(domain, meta)` hook gets a pointed
+`MetaError`, not a mid-run crash).
+
+Worked hook examples: cryptpad (`EXTRA_ENV(ctx)` derives `SANDBOX_DOMAIN` from `ctx.domain`),
+mumble (`READY_PROBE(ctx)` TCP voice-port probe, `UPGRADE_EXTRA_ENV(ctx)` adds a head-only compose
+overlay), ghost/discourse (`BACKUP_VERIFY(ctx)` dump-capture check).
+
+## 5. Writing custom tests & hooks
+
+### 5.1 Lifecycle overlay assertions — `test_<op>.py`
+
+One pytest file per lifecycle op (`install` / `upgrade` / `backup` / `restore`). The
+**orchestrator performs the op exactly once**; the overlay only *asserts* on the resulting state
+(HC3 op/assertion split — overlays never deploy, never restore, never mutate). The generic floor
+test runs additively against the same state.
+
+Conventions (see `tests/immich/test_backup.py` etc.):
+- use the `live_app` fixture (asserts `CCCI_APP_DOMAIN` is set, yields the domain)
+- use the `meta` fixture — the recipe's FULL validated `RecipeMeta` (attribute access)
+- use the `op_state` fixture for op context (versions, `snapshot_id`, artifact paths — the
+  orchestrator's run-scoped op record; skips with a clear reason outside an orchestrator run)
+- execute in-container checks via `harness.lifecycle.exec_in_app(domain, service, cmd)`
+
+### 5.2 Pre-op seed hooks — `ops.py`
+
+`def pre_<op>(ctx)` callables, imported and called by the orchestrator **before** performing the
+op. This is where data gets seeded so the post-op overlay can assert on it:
+
+```python
+# tests/immich/ops.py (pattern)
+def pre_upgrade(ctx):  _psql(ctx.domain, "INSERT ... 'upgrade-survives'")
+def pre_backup(ctx):   _psql(ctx.domain, "INSERT ... 'original'")
+def pre_restore(ctx):  _psql(ctx.domain, "DROP TABLE ci_marker")  # damage, restore must undo
+```
+
+Seed → op → assert is the whole pattern: `pre_backup` writes a marker, the orchestrator backs up,
+`pre_restore` destroys it, the orchestrator restores, `test_restore.py` asserts the marker is back.
+
+### 5.3 Custom tier — `functional/` and `playwright/` ONLY
+
+All custom-tier tests live under `tests/<recipe>/functional/` or `tests/<recipe>/playwright/`
+(discovery: `discovery.custom_tests`; the placement rule, §3). Run in the CUSTOM tier, after
+restore, against the post-upgrade (PR-head) app. ALL discovered files run — cc-ci's and (if
+HC2-approved) repo-local's, additively.
+
+Enrollment contract (`docs/enroll-recipe.md`): ≥2 NEW functional tests beyond ports of existing
+upstream checks; ported tests carry `SOURCE:` comments. Playwright tests get the shared
+browser/harness helpers (`harness.browser`); SSO recipes get `harness.sso`
+(`setup_keycloak_realm` — idempotent, `oidc_password_grant` — provider-pluggable). The documented
+import toolbox for custom tests is `from harness import lifecycle, sso, browser`.
+
+Tests needing deps use the `deps` fixture (entries expose `.domain` plus the full creds dict) and
+carry `@pytest.mark.requires_deps` — when dep provisioning failed they skip with reason
+`deps-not-ready` and the skip count is reported and FAILS a declared-deps run (F2-11; a green exit
+must not mask an unrun SSO test). Fixtures replace direct `os.environ` reads — after the
+restructure no recipe test parses env by hand.
+
+### 5.4 Pre-deploy shell hook — `install_steps.sh`
+
+The ONLY shell hook. Runs after `abra app new` + `EXTRA_ENV` application + secret generation,
+**before** the single base deploy. For setup that must precede the first deploy: writing extra
+config files into the recipe checkout, editing `.env` beyond simple key=val, and — for recipes
+with `DEPS` — wiring dep-derived OIDC env into the deploy (deps are always provisioned BEFORE the
+deploy; install-time wiring is the only mode, so there is exactly one deploy and no post-deploy
+redeploy hook).
+
+Env contract: `CCCI_APP_DOMAIN`, `CCCI_RECIPE`, `CCCI_APP_ENV` (path to the app's `.env`), and —
+when `DEPS` is declared — `CCCI_DEPS_FILE` (jq-readable JSON of dep creds/URLs; see
+lasuite-drive/-meet/-docs for the pattern). Must locate the recipe checkout ABRA_DIR-aware:
+`RECIPE_DIR="${ABRA_DIR:-${HOME}/.abra}/recipes/${CCCI_RECIPE}"` (per-run `ABRA_DIR` since the
+concurrency restructure — a hardcoded `~/.abra` writes to the wrong tree).
+
+Graceful-generic rule: a recipe needing a hook but not shipping one simply fails the generic
+install — a correct reported outcome, not a harness error.
+
+### 5.5 CI-only compose overlay — `compose.ccci.yml`
+
+**First-class:** if `tests/<recipe>/compose.ccci.yml` exists, the harness itself copies it into
+the recipe checkout (ABRA_DIR-aware) before the base deploy and automatically uses `--chaos` for
+that deploy (the untracked file would otherwise trip abra's clean-tree gate). No
+`install_steps.sh` copy boilerplate, no flag to remember (the old `CHAOS_BASE_DEPLOY` ⇄ overlay
+coupling is gone). The overlay is cc-ci-owned only.
+
+Policy unchanged: overlays are a minimal, justified fallback (ghost's is a 15m `start_period`
+grace — a literal, because abra validates `start_period` before env substitution). Reference the
+overlay from `EXTRA_ENV`'s `COMPOSE_FILE` as usual. Users: ghost, discourse.
+
+### 5.6 Environment & fixture contract (what custom code can read)
+
+Pytest fixtures (`tests/conftest.py` — the single fixture file):
+
+| Fixture | Yields |
+|---|---|
+| `recipe` | the recipe name (`$RECIPE`) |
+| `meta` | the FULL validated `RecipeMeta` (single loader) |
+| `live_app` | the shared deployment's domain (asserts it exists) |
+| `op_state` | the orchestrator's op-context dict (skips cleanly outside a run) |
+| `deps` | `{dep_recipe: entry}` — entries expose `.domain` + full SSO creds |
+
+Environment (hooks/shell, and approved repo-local code):
+
+| Var | Set for | Meaning |
+|---|---|---|
+| `CCCI_APP_DOMAIN` | all tests + hooks | the app's per-run domain |
+| `CCCI_BASE_URL` | approved repo-local code | `https://<domain>` |
+| `CCCI_RECIPE`, `CCCI_APP_ENV` | `install_steps.sh` | recipe name, app `.env` path |
+| `CCCI_OP_STATE_FILE` | overlay tests (via `op_state`) | JSON op context (versions, artifacts) |
+| `CCCI_DEPS_FILE` | `install_steps.sh` + harness | JSON dep creds dict |
+| `CCCI_DEPS_READY` / `CCCI_DEPS_NOT_READY_REASON` | custom tier (via `requires_deps`) | gate SSO tests, skip-with-reason |
+
+## 6. Run-model context (what the settings plug into)
+
+One deploy chain per run (full detail: `docs/testing.md` §2):
+
+```
+[DEPS? provision deps FIRST → $CCCI_DEPS_FILE]
+deploy BASE (UPGRADE_BASE_VERSION or recipe_versions[-2]; EXTRA_ENV; install_steps.sh;
+             compose.ccci.yml auto-copied + auto-chaos)
+  → INSTALL tier   (READY_PROBE; generic + overlay asserts)
+  → pre_upgrade(ctx) → chaos-deploy PR HEAD (UPGRADE_EXTRA_ENV)
+  → UPGRADE tier   (READY_PROBE; version-label == head_ref)
+  → pre_backup(ctx) → backup       (BACKUP_CAPABLE; BACKUP_VERIFY)
+  → BACKUP tier
+  → pre_restore(ctx) → restore
+  → RESTORE tier
+  → CUSTOM tier    (functional/ + playwright/; deps via the `deps` fixture)
+  → SCREENSHOT (best-effort, never affects the verdict)
+  → teardown (deps LAST)
+```
+
+Deploy-count guard (DG4.1): exactly `1 + len(DEPS)` deploys per run (chaos redeploys don't
+count); the per-run counter file is keyed by run since the concurrency restructure.
+
+## 7. Local iteration, the manifest, and the dev-only escape hatch
+
+```
+RECIPE=<recipe> PR=<n> REF=<sha> SRC=recipe-maintainers/<recipe> \
+  STAGES=install,upgrade,backup,restore,custom \
+  cc-ci-run runner/run_recipe_ci.py
+```
+
+(`docs/enroll-recipe.md` §5 for the full loop, including dep teardown caveats.)
+
+**Customization manifest.** Every run prints, right after meta load + discovery, one block:
+
+```
+===== customization manifest: <recipe> =====
+meta (non-default): DEPLOY_TIMEOUT=1500 DEPS=['keycloak'] EXTRA_ENV='<hook>'
+hooks: ops.py[pre_backup,pre_upgrade](cc-ci) install_steps.sh(cc-ci) compose.ccci.yml(cc-ci)
+overlays: test_backup.py(cc-ci) test_restore.py(repo-local)
+custom tests: functional/=5 playwright/=2 (cc-ci)
+env overrides: (none)
+```
+
+The same dict is embedded in `results.json` under `"customization"`. It is pure presentation —
+built from the SAME discovery/meta calls the run uses (so it cannot disagree with what executes,
+and it honors the HC2 gate) — and never influences a verdict.
+
+**Dev-only generic skip.** `CCCI_SKIP_GENERIC=1` (all ops) / `CCCI_SKIP_GENERIC_<OP>=1` (one op)
+suppress the generic floor — a LOCAL-DEV-ONLY escape hatch for iterating on one tier. There is no
+declarative equivalent (the old `SKIP_GENERIC` meta key is deleted). If the env form is active in
+a CI (drone) run, the run prints a loud `!!` warning and the manifest records it.
+
+## 8. Restructure outcomes (the review spec's R1–R9)
+
+How each defect identified in the review spec (commit `76a4b6b` §8) was resolved:
+
+- **R1 — six divergent meta loaders → RESOLVED.** One registry-backed loader
+  (`harness/meta.py::load`), the only `exec()` of `recipe_meta.py`. The orchestrator loads once
+  and passes the `RecipeMeta` down; conftest/lifecycle/deps/canonical all read the one object.
+- **R2 — dead `SCREENSHOT` knob → RESOLVED (kept + fixed).** The registry replaced the allowlist
+  that orphaned it; the orchestrator path now delivers the hook to `screenshot.py`
+  (proven end-to-end by `tests/unit/test_screenshot.py::test_screenshot_reachable_through_real_load_path`).
+- **R3 — 4-key pytest `meta` fixture → RESOLVED.** The fixture returns the full validated
+  `RecipeMeta`.
+- **R4 — three config languages → MITIGATED by the manifest** (§7): the surfaces stay (they serve
+  different actors), but every run resolves them into one visible block + results key.
+- **R5 — reference-doc drift → RESOLVED.** §4's key table is generated from the registry
+  (`scripts/gen-meta-docs.py`); a unit test fails CI on drift; `testing.md`/`enroll-recipe.md`
+  point here instead of keeping partial lists.
+- **R6 — silent typos → RESOLVED.** Unknown ALL-CAPS keys and type mismatches are hard
+  `MetaError`s; private constants are underscore-prefixed (exempt).
+- **R7 — `compose.ccci.yml` ⇄ `CHAOS_BASE_DEPLOY` coupling → RESOLVED.** The overlay is
+  first-class: harness-copied, auto-chaos. The flag is deleted.
+- **R8 — zero-user `SKIP_GENERIC` meta key → RESOLVED (deleted).** Env form remains, documented
+  dev-only, loudly flagged in CI runs (§7).
+- **R9 — `recipe_meta.py` is code, not config → REJECTED by decision.** No data/hooks file split:
+  registry validation gets the value (typed, validated keys) at lower cost; one file per recipe
+  remains the single config place. The expressiveness need is real (cryptpad derives env from the
+  per-run domain).
+
+Also settled in the restructure: install-time deps provisioning is the ONLY mode (the legacy
+post-deploy `setup_custom_tests.sh` machinery and its extra redeploy are deleted); the custom-test
+placement rule (§3); the uniform ctx hook convention (§4.1); the consolidated fixture surface
+(§5.6 — `deps` replaces `deps_apps`+`deps_creds`; dead `deployed`/`deployed_app`/`app_domain`
+fixtures deleted).
+
+## 9. File / symbol index
+
+| Concern | Where |
+|---|---|
+| THE meta loader + key registry + `HookCtx` + `MetaError` | `runner/harness/meta.py` (`load`, `KEYS`, `check_hook_signature`) |
+| Generated key table | `scripts/gen-meta-docs.py` → §4 above (sync pinned by `tests/unit/test_meta.py`) |
+| Customization manifest | `runner/harness/manifest.py` (`build`, `render`), printed by `runner/run_recipe_ci.py` |
+| Overlay/custom/hook discovery + HC2 gate + placement rule | `runner/harness/discovery.py` |
+| HC2 allowlist | `tests/repo-local-approved.txt` |
+| Generic assertions + `BACKUP_CAPABLE` detect | `runner/harness/generic.py` |
+| `compose.ccci.yml` auto-copy + auto-chaos | `runner/harness/lifecycle.py` (`provide_ccci_overlay`, `deploy_app`) |
+| `READY_PROBE` consumption | `runner/harness/lifecycle.py` (`wait_ready_probes`) |
+| `EXPECTED_NA` reporting | `runner/harness/results.py` |
+| `SCREENSHOT` consumer | `runner/harness/screenshot.py` |
+| Fixtures (`recipe`/`meta`/`live_app`/`op_state`/`deps`) + F2-11 skip-report | `tests/conftest.py` |
+| Skip-generic env logic (dev-only) | `runner/run_recipe_ci.py` (`_skip_generic`) |
+| Unit tests pinning all of the above | `tests/unit/test_meta.py`, `test_manifest.py`, `test_discovery*.py` |
+| Worked examples | `tests/ghost/` (overlay+compose.ccci.yml), `tests/mumble/` (TCP probe, UPGRADE_EXTRA_ENV, private `_` constants), `tests/lasuite-drive/` (DEPS + install-time OIDC wiring), `tests/immich/` (ops.py seed pattern) |
--- a/docs/testing.md
+++ b/docs/testing.md
@ -16,12 +16,13 @@ year from now, this is the one rule that should still hold.
  ship as the floor for every recipe. No SSO provider, no external deps, no per-recipe state
  scaffolding — just "does this recipe deploy and lifecycle work?"
 - **Generic must not depend on custom.** A custom test or a custom-tests setup (e.g. SSO/OIDC dep
-  provisioning) **can never be a precondition for the generic tier to pass.** Concretely: the
-  orchestrator runs all generic tiers (install → upgrade → backup → restore) against the recipe
-  **alone, with no deps deployed**, then runs the `setup_custom_tests` step (deps + post-deps
-  wiring) only after — and a failure there is **isolated** to the custom tier (tests tagged
-  `@pytest.mark.requires_deps` skip with reason `"deps-not-ready"`; generic tier reports
-  normally). See `cc-ci-plan/plan-sso-dep-testing.md` for the SSO-dep specifics.
+  provisioning) **can never be a precondition for the generic tier to pass.** Concretely: deps are
+  provisioned BEFORE the single deploy (so `install_steps.sh` can wire OIDC env into that one
+  deploy), but a dep-provisioning failure is **isolated** to the custom tier — the recipe still
+  deploys alone, every generic tier (install → upgrade → backup → restore) runs normally, and
+  tests tagged `@pytest.mark.requires_deps` skip with reason `"deps-not-ready"` (a counted,
+  reported skip — F2-11). A deps failure can never fail or block a generic tier. See
+  `cc-ci-plan/plan-sso-dep-testing.md` for the SSO-dep specifics.
 - **Custom tests are the thoroughness layer — and they cost more to maintain.** They're more
  thorough (authenticated APIs, multi-app flows, version-specific browser selectors, helper
  scripts, state-management) and *therefore* take more maintenance: an SSO provider's admin API
@ -113,9 +114,11 @@ repo-local  <recipe-repo>/tests/test_<op>.py     (upstream-authoritative; gated
 Only ONE overlay source wins for a given op (repo-local > cc-ci); the generic floor runs **in
 addition** unless explicitly opted out.

-**Custom (non-lifecycle) `test_*.py`** — any other `test_*.py` (e.g. `test_sso.py`) is **opt-in and
-additive**: it has no generic equivalent and runs only when present, discovered from both locations
-(repo-local gated by the HC2 allowlist).
+**Custom (non-lifecycle) tests** — e.g. `functional/test_sso.py` — are **opt-in and additive**:
+they have no generic equivalent and run only when present, discovered from both locations
+(repo-local gated by the HC2 allowlist). Placement rule: custom tests live ONLY under
+`functional/` or `playwright/`; a top-level `test_*.py` is a lifecycle overlay and nothing else
+(top-level non-lifecycle files are not discovered).

 ### Pre-op seed hooks (per-recipe `ops.py`)

@ -127,35 +130,38 @@ etc.). Since the orchestrator owns the op, overlays place their seed in an optio
 # tests/<recipe>/ops.py
 from harness import lifecycle

-def pre_upgrade(domain, meta):
+def pre_upgrade(ctx):
    # seed a marker before the harness performs the upgrade
-    lifecycle.exec_in_app(domain, ["sh", "-c", "echo upgrade-survives > /path/marker"])
+    lifecycle.exec_in_app(ctx.domain, ["sh", "-c", "echo upgrade-survives > /path/marker"])

-def pre_backup(domain, meta):
+def pre_backup(ctx):
    # establish a known "original" state before the backup op captures it
-    lifecycle.exec_in_app(domain, ["sh", "-c", "echo original > /path/marker"])
+    lifecycle.exec_in_app(ctx.domain, ["sh", "-c", "echo original > /path/marker"])

-def pre_restore(domain, meta):
+def pre_restore(ctx):
    # diverge from the backed-up state so a successful restore is observable
-    lifecycle.exec_in_app(domain, ["sh", "-c", "echo mutated > /path/marker"])
+    lifecycle.exec_in_app(ctx.domain, ["sh", "-c", "echo mutated > /path/marker"])
 ```

 The orchestrator imports `ops.py` in-process (with the recipe dir on `sys.path`, so it can import
-sibling helpers like `kc_admin.py`) and calls `pre_<op>(domain, meta)` immediately before performing
-the op. Then `test_<op>.py` asserts the post-op state. See `tests/custom-html/` (volume marker),
+sibling helpers like `kc_admin.py`) and calls `pre_<op>(ctx)` immediately before performing the
+op — `ctx` is the uniform `HookCtx` every recipe hook receives (`.domain`, `.base_url`, `.meta`,
+`.deps`, `.op` — `docs/recipe-customization.md` §4.1). Then `test_<op>.py` asserts the post-op
+state. See `tests/custom-html/` (volume marker),
 `tests/keycloak/` (admin-API/realm), `tests/matrix-synapse/`, `tests/lasuite-docs/` (psql in the `db`
 service) for worked examples.

-### Opting out of the generic floor
+### Opting out of the generic floor (LOCAL-DEV-ONLY)

-The generic runs additively by default. To skip it (e.g. when an overlay's recipe-specific check
-fully replaces the generic's mechanism check) set, in increasing specificity:
+The generic runs additively by default and there is **no declarative opt-out** — no recipe can
+ship without the floor. For local iteration only (e.g. re-running one tier while developing an
+overlay), two env escape hatches exist:

 - **env `CCCI_SKIP_GENERIC=1`** — skip generic for ALL ops (run-wide).
 - **env `CCCI_SKIP_GENERIC_<OP>=1`** — e.g. `CCCI_SKIP_GENERIC_UPGRADE=1` — skip generic for that one op.
- **declarative in `recipe_meta.py`** — `SKIP_GENERIC = ["upgrade"]` (per-op) or `SKIP_GENERIC = ["all"]`.

-Opting out is per-recipe and visible in git — not a hidden global. Truthy = `1`/`true`/`yes`/`on`.
+Truthy = `1`/`true`/`yes`/`on`. If either is active in a CI (drone) run, the run prints a loud
+`!!` warning and the customization manifest records it (`docs/recipe-customization.md` §7).

 ## Repo-local trust gate (HC2) — default-deny

@ -215,12 +221,14 @@ installs and stays 1.
   `tests/custom-html/test_upgrade.py`). Assert the POST-op state — reading app state through
   `lifecycle.exec_in_app` (volume/DB) for data checks, not HTTP. Generic + your overlay both run.
 3. If the overlay needs to seed PRE-op state (data-continuity markers, the backup→restore
-   divergence), drop `tests/<recipe>/ops.py` with `pre_upgrade/pre_backup/pre_restore(domain, meta)`.
+   divergence), drop `tests/<recipe>/ops.py` with `pre_upgrade/pre_backup/pre_restore(ctx)`.
 4. If the recipe needs install-time setup, add `tests/<recipe>/install_steps.sh`.
-5. Set per-recipe knobs (health path, timeouts, opt-out) in `recipe_meta.py`.
+5. Set per-recipe knobs (health path, timeouts) in `recipe_meta.py`.
 6. **Never weaken or skip an assertion to make a run pass** — a red tier is information.

-Per-recipe config (`tests/<recipe>/recipe_meta.py`, all optional):
+Per-recipe config (`tests/<recipe>/recipe_meta.py`, all optional — the COMPLETE key reference is
+the generated table in `docs/recipe-customization.md` §4; unknown keys are hard errors, private
+constants are underscore-prefixed):

 ```python
 HEALTH_PATH = "/realms/master"   # path that returns a healthy status (default "/")
@ -228,8 +236,7 @@ HEALTH_OK = (200,)               # acceptable status codes (default 200/301/302)
 DEPLOY_TIMEOUT = 600             # seconds for services to converge (default 600)
 HTTP_TIMEOUT = 600               # seconds for the app to answer (default 300)
 BACKUP_CAPABLE = True            # override backup-capability auto-detection (default: scan compose)
-EXTRA_ENV = {"KEY": "value"}     # or EXTRA_ENV(domain) -> dict; extra .env keys set at deploy
-SKIP_GENERIC = ["upgrade"]       # per-recipe declarative opt-out from generic ops ("all" = every op)
+EXTRA_ENV = {"KEY": "value"}     # or EXTRA_ENV(ctx) -> dict; extra .env keys set at deploy
 ```

 The harness self-tests for discovery / precedence / the HC2 allowlist live in `tests/unit/` (run:
--- a/flake.nix
+++ b/flake.nix
@ -31,34 +31,36 @@
      ];
    in
    {
-      # Canonical live host target: the Hetzner cc-ci server.
-      # Use `.#cc-ci` for the current production host.
-      nixosConfigurations.cc-ci = nixpkgs.lib.nixosSystem {
-        inherit system;
-        modules = [
-          sops-nix.nixosModules.sops
-          ./nix/hosts/cc-ci-hetzner/configuration.nix
-        ];
-      };
+      nixosConfigurations = {
+        # Canonical live host target: the Hetzner cc-ci server.
+        # Use `.#cc-ci` for the current production host.
+        cc-ci = nixpkgs.lib.nixosSystem {
+          inherit system;
+          modules = [
+            sops-nix.nixosModules.sops
+            ./nix/hosts/cc-ci-hetzner/configuration.nix
+          ];
+        };

-      # Legacy Incus VM host definition retained only for historical comparison and fallback.
-      # Do NOT use this target on the live Hetzner server.
-      nixosConfigurations.cc-ci-incus = nixpkgs.lib.nixosSystem {
-        inherit system;
-        modules = [
-          sops-nix.nixosModules.sops
-          ./nix/hosts/cc-ci/configuration.nix
-        ];
-      };
+        # Legacy Incus VM host definition retained only for historical comparison and fallback.
+        # Do NOT use this target on the live Hetzner server.
+        cc-ci-incus = nixpkgs.lib.nixosSystem {
+          inherit system;
+          modules = [
+            sops-nix.nixosModules.sops
+            ./nix/hosts/cc-ci/configuration.nix
+          ];
+        };

-      # Explicit alias for the live Hetzner host. Kept alongside `cc-ci` so the intended host target
-      # remains obvious in recovery/migration workflows.
-      nixosConfigurations.cc-ci-hetzner = nixpkgs.lib.nixosSystem {
-        inherit system;
-        modules = [
-          sops-nix.nixosModules.sops
-          ./nix/hosts/cc-ci-hetzner/configuration.nix
-        ];
+        # Explicit alias for the live Hetzner host. Kept alongside `cc-ci` so the intended host
+        # target remains obvious in recovery/migration workflows.
+        cc-ci-hetzner = nixpkgs.lib.nixosSystem {
+          inherit system;
+          modules = [
+            sops-nix.nixosModules.sops
+            ./nix/hosts/cc-ci-hetzner/configuration.nix
+          ];
+        };
      };

      devShells.${system} = {
--- a/machine-docs/DECISIONS.md
+++ b/machine-docs/DECISIONS.md
@ -1283,3 +1283,15 @@ the commit), which is the correct SCM integration.
  environment; job is session-persistent (survives as long as Builder session runs). T0-refire
  verified: CronCreate test fire at 23:17Z → upgrader started, upgrader-cron.log created, status
  RUNNING. (2026-06-01)
+
+## conc P3 (2026-06-10, Builder): install_steps.sh hooks resolve $ABRA_DIR — guardrail note
+
+P3 makes recipe working trees per-run ($ABRA_DIR/recipes). tests/{ghost,discourse}/install_steps.sh
+hard-coded `${HOME}/.abra/recipes/...` to copy their compose.ccci.yml overlay into the deploy tree;
+under per-run trees that path is the WRONG (canonical) tree, so the overlay would silently miss the
+deploy and both recipes' upgrade-tier base deploys would break. Fixed with ONE mechanical line per
+hook: `RECIPE_DIR="${ABRA_DIR:-${HOME}/.abra}/recipes/${CCCI_RECIPE}"` (identical resolution rule to
+the abra CLI and abra.recipe_dir()). No test assertion, gate, or overlay content was touched — the
+phase guardrail's "never touch tests/<recipe>/ content" is read as protecting test/gate SEMANTICS;
+this is required P3 fallout, equivalent to the harness-side path routing. Flagged here for the
+Adversary's gate-integrity review.
--- a/nix/hosts/cc-ci-hetzner/configuration.nix
+++ b/nix/hosts/cc-ci-hetzner/configuration.nix
@ -7,7 +7,7 @@
 #   git clone --recursive https://git.autonomic.zone/recipe-maintainers/cc-ci.git /etc/cc-ci
 #   install -m600 <age-private-key> /var/lib/sops-nix/key.txt
 #   nixos-rebuild switch --flake /etc/cc-ci#cc-ci-hetzner
-{ pkgs, lib, ... }:
+{ pkgs, ... }:
 {
  imports = [
    ./hardware.nix
--- a/nix/hosts/cc-ci-hetzner/hardware.nix
+++ b/nix/hosts/cc-ci-hetzner/hardware.nix
@ -11,13 +11,17 @@
 {
  imports = [ (modulesPath + "/profiles/qemu-guest.nix") ];

-  boot.loader = {
-    efi.efiSysMountPoint = "/boot/efi";
-    grub = {
-      efiSupport = true;
-      efiInstallAsRemovable = true;
-      device = "nodev";
+  boot = {
+    loader = {
+      efi.efiSysMountPoint = "/boot/efi";
+      grub = {
+        efiSupport = true;
+        efiInstallAsRemovable = true;
+        device = "nodev";
+      };
    };
+    initrd.availableKernelModules = [ "ata_piix" "uhci_hcd" "xen_blkfront" "vmw_pvscsi" ];
+    initrd.kernelModules = [ "nvme" ];
  };

  fileSystems."/boot/efi" = {
@ -25,9 +29,6 @@
    fsType = "vfat";
  };

-  boot.initrd.availableKernelModules = [ "ata_piix" "uhci_hcd" "xen_blkfront" "vmw_pvscsi" ];
-  boot.initrd.kernelModules = [ "nvme" ];
-
  fileSystems."/" = {
    device = "/dev/sda1";
    fsType = "ext4";
--- a/nix/modules/drone-runner.nix
+++ b/nix/modules/drone-runner.nix
@ -8,14 +8,19 @@
 { pkgs, config, lib, ... }:
 let
  # MAX_TESTS (plan §4.2/§4.3 resource safety): max CI builds the exec runner runs at once. Drone
-  # queues the rest in its native pending-build queue (no custom queue). THE concurrency cap that
-  # bounds how many test apps can be live at once — kept LOW (1) on this single 28GiB node since
-  # recipes are heavy (immich/matrix large volumes). With capacity=1 there is never a concurrent
-  # in-flight run, so the run-start janitor can safely reap *any* orphan (a SIGKILL'd build runs no
-  # teardown) and the "at most MAX_TESTS apps live" bound holds exactly. Raise to 2 only if the node
-  # is shown to handle two light recipes at once (then the janitor MUST stay age-based to avoid
-  # reaping a concurrent run — see DECISIONS.md "Resource safety").
-  maxTests = "1";
+  # queues the rest in its native pending-build queue (no custom queue). THE SINGLE concurrency
+  # knob — nothing else caps recipe-ci parallelism (the .drone.yml concurrency.limit was removed:
+  # one knob, one place). Bounds how many test apps can be live at once.
+  #
+  # Raised to 2 (operator request 2026-06-09) so two recipes can be tested in parallel (e.g. immich
+  # and plausible under active development at once). Verified safe on the current node (Hetzner cpx22,
+  # ~7.6 GiB / 4 vCPU — NOTE: smaller than the original 28 GiB this was written for): a full immich CI
+  # stack measured ~1 GiB (server+ML+pg+redis) with multiple GiB free, so two concurrent recipes fit.
+  # Concurrent-run safety is the harness's job at ANY capacity (docs/concurrency.md): per-run
+  # ABRA_DIR recipe trees, per-app-domain flocks, and a flock-probe janitor that reaps a crashed
+  # build's orphan immediately (held lock = live run, never touched). Revert to "1" if OOM /
+  # disk-I/O contention is observed under load.
+  maxTests = "2";
 in
 {
  # Drone ships under the Polyform Small Business license (nixpkgs marks it unfree);
--- a/nix/modules/nightly-sweep.nix
+++ b/nix/modules/nightly-sweep.nix
@ -29,7 +29,7 @@ in
    serviceConfig = {
      Type = "oneshot";
      # A full sweep across several recipes (each a cold deploy/test/teardown) is long; bound it.
-      TimeoutStartSec = "21600";  # 6h ceiling
+      TimeoutStartSec = "21600"; # 6h ceiling
      ExecStart = "${sweep}/bin/cc-ci-nightly-sweep";
    };
  };
@ -39,7 +39,7 @@ in
    wantedBy = [ "timers.target" ];
    timerConfig = {
      OnCalendar = "*-*-* 03:00:00";
-      Persistent = true;   # catch up a missed nightly after downtime
+      Persistent = true; # catch up a missed nightly after downtime
      RandomizedDelaySec = "600";
    };
  };
--- a/nix/modules/reports.nix
+++ b/nix/modules/reports.nix
@ -3,10 +3,49 @@
 # no secrets — just static files behind traefik + the wildcard TLS (same pattern as dashboard.nix,
 # but a plain nginx:alpine since there's nothing to render server-side). Content is updated by writing
 # files into /var/lib/cc-ci-reports; nginx serves them live (no redeploy needed).
+#
+# It ALSO serves a same-origin realtime PR-status proxy at /pr/<recipe>/<n>: the report's STATUS
+# column fetches it client-side to show each PR's live state (open vs. ✓). Same-origin means no
+# dependency on the Gitea CORS allow-list; the recipe mirrors are public so no token is needed. The
+# proxy is pinned to recipe-maintainers + a safe recipe-name charset and is read-only (GET/HEAD).
 { pkgs, ... }:
 let
  reportsDir = "/var/lib/cc-ci-reports";

+  # Custom nginx server: static report files + the /pr/<recipe>/<n> → Gitea-API proxy. Replaces the
+  # stock /etc/nginx/conf.d/default.conf (which the image's nginx.conf includes inside http{}).
+  nginxConf = pkgs.writeText "cc-ci-reports-default.conf" ''
+    server {
+        listen 80;
+        server_name _;
+        root /usr/share/nginx/html;
+        index index.html;
+
+        # Realtime PR-status proxy for the Recipe Report STATUS column.
+        # GET /pr/<recipe>/<n> -> the PUBLIC Gitea PR JSON ({state, merged, ...}). Same-origin from
+        # the browser's view, so no CORS dependency; unauthenticated, since the recipe mirrors are
+        # public. The repo owner is hard-pinned to recipe-maintainers and the recipe name to a
+        # slashless charset, so the proxied path can only ever address recipe-maintainers/<name>/pulls
+        # (it cannot be coerced to another org or path). Only safe read methods are allowed.
+        location ~ ^/pr/([a-z0-9._-]+)/([0-9]+)$ {
+            limit_except GET HEAD { deny all; }
+            resolver 127.0.0.11 ipv6=off valid=30s;   # docker embedded DNS (forwards external names)
+            proxy_ssl_server_name on;
+            proxy_set_header Host git.autonomic.zone;
+            proxy_set_header Accept "application/json";
+            proxy_pass https://git.autonomic.zone/api/v1/repos/recipe-maintainers/$1/pulls/$2;
+            proxy_intercept_errors off;
+            proxy_connect_timeout 5s;
+            proxy_read_timeout 10s;
+            add_header Cache-Control "no-store" always;  # always fetch live state, never cache in the browser
+        }
+
+        location / {
+            try_files $uri $uri/ =404;
+        }
+    }
+  '';
+
  stack = pkgs.writeText "cc-ci-reports-stack.yml" ''
    version: "3.8"
    services:
@ -17,6 +56,10 @@ let
            source: ${reportsDir}
            target: /usr/share/nginx/html
            read_only: true
+          - type: bind
+            source: ${nginxConf}
+            target: /etc/nginx/conf.d/default.conf
+            read_only: true
        networks:
          - proxy
        deploy:
--- a/runner/harness/abra.py
+++ b/runner/harness/abra.py
@ -10,6 +10,7 @@ Bakes in the known abra gotchas (re-verify per installed abra version, currently
 from __future__ import annotations

 import json
+import os
 import subprocess

 ABRA = "abra"
@ -19,6 +20,20 @@ class AbraError(RuntimeError):
    pass


+def abra_dir() -> str:
+    """abra's state dir, resolved the same way the abra CLI resolves it: $ABRA_DIR if set, else
+    ~/.abra. Inside a CI run, run_recipe_ci exports a PER-RUN $ABRA_DIR (fresh recipes/, shared
+    servers/+catalogue/ symlinks) before any abra call, so every helper here and every abra
+    subprocess agree on the same tree; outside a run (warm_reconcile's systemd timer, manual use)
+    both fall back to the canonical /root/.abra."""
+    return os.environ.get("ABRA_DIR") or os.path.expanduser("~/.abra")
+
+
+def recipe_dir(recipe: str) -> str:
+    """The current ABRA_DIR's working tree for a recipe (per-run inside a CI run)."""
+    return os.path.join(abra_dir(), "recipes", recipe)
+
+
 def _run_pty(
    args: list[str], timeout: int = 900, check: bool = True
 ) -> subprocess.CompletedProcess:
@ -77,9 +92,7 @@ def recipe_checkout(recipe: str, version: str) -> None:
    a chaos (`-C`) deploy ignores ENV VERSION and uses the current checkout — together that silently
    deployed LATEST for a 'previous-version' base, making the upgrade a no-op (Adversary F1d-2). With
    this checkout + a non-chaos deploy, a pinned deploy genuinely deploys that version."""
-    import os
-
-    path = os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    path = recipe_dir(recipe)
    # -f (force): the version-pinning checkout must yield the EXACT ref tree. Without it, a cc-ci
    # install_steps-provided overlay (e.g. discourse's compose.ccci.yml, copied into the pinned base)
    # is an UNTRACKED file that collides with the same path TRACKED in a later ref, and
@ -100,9 +113,7 @@ def has_lightweight_version_tags(recipe: str) -> bool:
    'reference not found'.) The caller (deploy_app) uses this to fall back to a chaos base deploy
    (which skips lint and deploys the explicitly-checked-out pinned version — see lifecycle.deploy_app).
    Read-only: just `git tag` + `cat-file -t`; no fetch/mutation, so it can't trigger abra's revert."""
-    import os
-
-    path = os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    path = recipe_dir(recipe)
    tags = subprocess.run(
        ["git", "-C", path, "tag", "-l"], capture_output=True, text=True
    ).stdout.split()
@ -168,7 +179,9 @@ def secret_generate(domain: str, timeout: int = 300) -> None:
    )


-def deploy(domain: str, chaos: bool = True, timeout: int = 900, no_converge_checks: bool = False) -> None:
+def deploy(
+    domain: str, chaos: bool = True, timeout: int = 900, no_converge_checks: bool = False
+) -> None:
    args = ["app", "deploy", domain, "-o", "-n"]
    if chaos:
        args.append("-C")
@ -203,7 +216,10 @@ def backup_create(domain: str, timeout: int = 900) -> str:
    # remote and fails "authentication required: Unauthorized". Returns the captured output, whose
    # restic JSON summary line carries the produced "snapshot_id" (the backup artifact, DG3) — note
    # `abra app backup snapshots` needs a TTY and is awkward to script, so we read the create output.
-    out = _run_pty(["app", "backup", "create", domain, "-n", "-C", "-o"], timeout=timeout).stdout or ""
+    out = (
+        _run_pty(["app", "backup", "create", domain, "-n", "-C", "-o"], timeout=timeout).stdout
+        or ""
+    )
    # Echo the backup output (incl. backupbot's pre-hook run / any "Failed to run command" or
    # "Container ... not running" ERROR) into the run log. Backup is otherwise opaque: a pre-hook that
    # fails to register/run leaves the DB dump out of the snapshot, surfacing only as a downstream
@ -226,9 +242,7 @@ def recipe_head_commit(recipe: str) -> str | None:
    """The current HEAD commit of the recipe checkout — captured right after fetch (the PR head, or
    the catalogue current) so the upgrade tier can re-checkout it for the chaos redeploy after the
    prev-tag base deploy reset the working tree (HC1)."""
-    import os
-
-    path = os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    path = recipe_dir(recipe)
    proc = subprocess.run(["git", "-C", path, "rev-parse", "HEAD"], capture_output=True, text=True)
    out = proc.stdout.strip()
    return out or None
@ -236,10 +250,7 @@ def recipe_head_commit(recipe: str) -> str | None:

 def recipe_versions(recipe: str) -> list[str]:
    """Published versions of a recipe, oldest→newest (from the recipe git tags)."""
-    import os
-    import subprocess
-
-    path = os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    path = recipe_dir(recipe)
    proc = subprocess.run(
        ["git", "-C", path, "tag", "--sort=creatordate"], capture_output=True, text=True
    )
--- a/runner/harness/browser.py
+++ b/runner/harness/browser.py
@ -13,8 +13,15 @@ from __future__ import annotations
 import time


-def goto_with_retry(page, url, *, deadline_seconds: int = 120, accept_statuses=(200, 304),
-                    goto_timeout_ms: int = 30_000, wait_until: str = "domcontentloaded"):
+def goto_with_retry(
+    page,
+    url,
+    *,
+    deadline_seconds: int = 120,
+    accept_statuses=(200, 304),
+    goto_timeout_ms: int = 30_000,
+    wait_until: str = "domcontentloaded",
+):
    """Poll `page.goto(url)` until status is in `accept_statuses` OR the deadline expires.

    Returns the final Playwright response. Raises AssertionError if the deadline expires without
--- a/runner/harness/canonical.py
+++ b/runner/harness/canonical.py
@ -30,17 +30,13 @@ import subprocess
 import time

 from . import abra, warm, warmsnap
+from . import meta as meta_mod


 def is_enrolled(recipe: str) -> bool:
-    """True if `tests/<recipe>/recipe_meta.py` sets `WARM_CANONICAL = True`. Missing meta → False."""
-    path = os.path.join(os.path.dirname(__file__), "..", "..", "tests", recipe, "recipe_meta.py")
-    if not os.path.exists(path):
-        return False
-    ns: dict = {}
-    with open(path) as fh:
-        exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
-    return bool(ns.get("WARM_CANONICAL"))
+    """True if `tests/<recipe>/recipe_meta.py` sets `WARM_CANONICAL = True`. Missing meta → False.
+    Reads through the single meta loader (rcust P1 — no per-module exec)."""
+    return bool(meta_mod.load(recipe).WARM_CANONICAL)


 def canonical_domain(recipe: str) -> str:
@ -51,11 +47,13 @@ def canonical_domain(recipe: str) -> str:
 def enrolled_recipes() -> list[str]:
    """All recipes enrolled as data-warm canonicals (recipe_meta.WARM_CANONICAL=True), sorted. Used
    by the WC6 nightly sweep to know which canonicals to refresh via a green cold run on latest."""
-    tests_dir = os.path.join(os.path.dirname(__file__), "..", "..", "tests")
+    tests_dir = meta_mod.TESTS_DIR
    out = []
    try:
        for name in sorted(os.listdir(tests_dir)):
-            if os.path.isfile(os.path.join(tests_dir, name, "recipe_meta.py")) and is_enrolled(name):
+            if os.path.isfile(os.path.join(tests_dir, name, "recipe_meta.py")) and is_enrolled(
+                name
+            ):
                out.append(name)
    except OSError:
        pass
@ -122,11 +120,15 @@ def deploy_canonical(recipe: str, timeout: int = 900) -> None:
    abra.recipe_checkout(recipe, version)
    r = subprocess.run(
        ["abra", "app", "deploy", domain, version, "-o", "-n", "-f"],
-        capture_output=True, text=True, timeout=timeout,
+        capture_output=True,
+        text=True,
+        timeout=timeout,
    )
    if r.returncode != 0:
-        raise RuntimeError(f"deploy canonical {domain} {version} failed: "
-                           f"{(r.stderr + ' ' + r.stdout).strip()[:300]}")
+        raise RuntimeError(
+            f"deploy canonical {domain} {version} failed: "
+            f"{(r.stderr + ' ' + r.stdout).strip()[:300]}"
+        )
    _set_status(recipe, "warm")


--- a/runner/harness/card.py
+++ b/runner/harness/card.py
@ -148,7 +148,9 @@ RUNG_LABEL = {
    "backup_restore": "backup/restore",
    "functional": "functional",
 }
-SKIP_GREEN = "#57ab5a"  # muted green — an intentional skip reads like a pass (but labelled, never inflating)
+SKIP_GREEN = (
+    "#57ab5a"  # muted green — an intentional skip reads like a pass (but labelled, never inflating)
+)


 def _skip_rows(skips: dict) -> str:
@ -159,14 +161,16 @@ def _skip_rows(skips: dict) -> str:
    for rung, reason in (skips.get("intentional") or {}).items():
        rows.append(
            f'<tr class="stage"><td colspan="2"><span class="mark" style="color:{SKIP_GREEN}">⊘</span>'
-            f'<b>{html.escape(RUNG_LABEL.get(rung, rung))}</b></td>'
+            f"<b>{html.escape(RUNG_LABEL.get(rung, rung))}</b></td>"
            f'<td class="st" style="color:{SKIP_GREEN}">intentional skip</td></tr>'
        )
-        rows.append(f'<tr class="skipreason"><td></td><td colspan="2">{html.escape(reason)}</td></tr>')
+        rows.append(
+            f'<tr class="skipreason"><td></td><td colspan="2">{html.escape(reason)}</td></tr>'
+        )
    for rung in skips.get("unintentional") or []:
        rows.append(
            f'<tr class="stage"><td colspan="2"><span class="mark" style="color:{GAP_COLOR}">⊘</span>'
-            f'<b>{html.escape(RUNG_LABEL.get(rung, rung))}</b></td>'
+            f"<b>{html.escape(RUNG_LABEL.get(rung, rung))}</b></td>"
            f'<td class="st" style="color:{GAP_COLOR}">unintentional skip</td></tr>'
        )
        rows.append(
--- a/runner/harness/deps.py
+++ b/runner/harness/deps.py
@ -20,7 +20,7 @@ Per Phase-2 DECISIONS:
 Run state:
 - `$CCCI_DEPS_FILE` — JSON file written by the orchestrator after each dep deploys; each entry is
  `{"recipe": "<dep-recipe>", "domain": "<dep-domain>", "version": null}`. Tests access via the
-  `deps_apps` pytest fixture defined in `tests/conftest.py`.
+  `deps` pytest fixture defined in `tests/conftest.py`.
 """

 from __future__ import annotations
@ -28,24 +28,10 @@ from __future__ import annotations
 import contextlib
 import json
 import os
-from typing import Iterable
+from collections.abc import Iterable

 from . import lifecycle, naming
-
-
-def declared_deps(recipe: str) -> list[str]:
-    """Read `DEPS` from `tests/<recipe>/recipe_meta.py` — a list of recipe names this recipe needs
-    deployed alongside it. Returns [] if none."""
-    path = os.path.join(
-        os.path.dirname(__file__), "..", "..", "tests", recipe, "recipe_meta.py"
-    )
-    if not os.path.exists(path):
-        return []
-    ns: dict = {}
-    with open(path) as fh:
-        exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
-    deps = ns.get("DEPS") or []
-    return [str(d) for d in deps if d]
+from . import meta as meta_mod


 def dep_domain(parent_recipe: str, pr: str, ref: str | None, dep_recipe: str) -> str:
@ -64,11 +50,11 @@ def write_run_state(deps_state) -> None:
    """Write the deps state file ($CCCI_DEPS_FILE). Two shapes supported (canonical=keyed dict):

    1. **Legacy list-of-entries:** `[{"recipe": "<dep>", "domain": "<d>"}, ...]` (Q2.3 original).
-       Still accepted by `load_run_state` for backwards compat — `deps_apps` fixture flattens.
+       Still accepted by `load_run_state` for backwards compat — the `deps` fixture flattens.
    2. **NEW per-spec dict (operator-2026-05-28 SSO-dep plan §3.2):**
       `{"<dep_recipe>": {"recipe": "<dep>", "domain": "<d>", "realm": "...",
       "client_id": "...", "client_secret": "...", "admin_user": "...", "admin_password": "..."}}`.
-       The `setup_custom_tests.sh` per-recipe hook reads this via `jq` to wire OIDC env.
+       The per-recipe `install_steps.sh` hook reads this via `jq` to wire OIDC env.

    No-op if `$CCCI_DEPS_FILE` isn't set."""
    path = os.environ.get("CCCI_DEPS_FILE")
@ -83,11 +69,12 @@ def deploy_deps(
    pr: str,
    ref: str | None,
    deps: Iterable[str],
-    meta_for: dict[str, dict] | None = None,
+    meta_for: dict | None = None,
 ) -> list[dict]:
    """Deploy each declared dep, sequentially, at its per-run domain. Returns the list of state
-    dicts (one per dep). `meta_for` maps dep_recipe -> meta (HEALTH_PATH/HEALTH_OK/timeouts) so the
-    readiness wait uses per-dep config; missing dep meta falls back to (/, 200/301/302, 600s)."""
+    dicts (one per dep). `meta_for` maps dep_recipe -> RecipeMeta (HEALTH_PATH/HEALTH_OK/timeouts)
+    so the readiness wait uses per-dep config; a missing dep meta is loaded via meta.load()
+    (defaults: /, 200/301/302, 600s)."""
    meta_for = meta_for or {}
    state: list[dict] = []
    for dep in deps:
@ -96,20 +83,21 @@ def deploy_deps(
        # NB: each dep_app gets a fresh deploy_count entry only on `_record_deploy` which fires
        # inside `lifecycle.deploy_app`. For Phase 2 the deploy-count guard (DG4.1) counts the
        # parent + its deps as distinct install events — by design, since each is a separate app.
-        dm = meta_for.get(dep, {})
+        dm = meta_for.get(dep) or meta_mod.load(dep)
        lifecycle.deploy_app(
            dep,
            domain,
            secrets=True,
-            deploy_timeout=int(dm.get("DEPLOY_TIMEOUT", 900)),
+            deploy_timeout=int(dm.DEPLOY_TIMEOUT),
+            meta=dm,
        )
        try:
            lifecycle.wait_healthy(
                domain,
-                ok_codes=tuple(dm.get("HEALTH_OK", (200, 301, 302))),
-                path=dm.get("HEALTH_PATH", "/"),
-                deploy_timeout=int(dm.get("DEPLOY_TIMEOUT", 600)),
-                http_timeout=int(dm.get("HTTP_TIMEOUT", 600)),
+                ok_codes=tuple(dm.HEALTH_OK),
+                path=dm.HEALTH_PATH,
+                deploy_timeout=int(dm.DEPLOY_TIMEOUT),
+                http_timeout=int(dm.HTTP_TIMEOUT),
            )
        except Exception:
            # If a dep fails to converge, abort the whole resolve — let the caller teardown
@ -165,7 +153,7 @@ def load_run_state():


 def deps_as_dict(state) -> dict[str, dict]:
-    """Coerce either shape (legacy list or new dict) into a recipe→entry dict for the deps_apps
+    """Coerce either shape (legacy list or new dict) into a recipe→entry dict for the `deps`
    fixture + dependent-tests consumption."""
    if isinstance(state, dict):
        return state
--- a/runner/harness/discovery.py
+++ b/runner/harness/discovery.py
@ -11,7 +11,8 @@ hook; the orchestrator decides additive-vs-skip. Sources, in precedence order
      > cc-ci       tests/<recipe>/test_<op>.py
    (the generic tests/_generic/test_<op>.py is the always-present floor, run separately by default)

-    custom (non-lifecycle) test_*.py — ALL run, additively, from BOTH locations (opt-in).
+    custom test_*.py (functional/ + playwright/ ONLY, rcust P4 placement rule) — ALL run,
+        additively, from BOTH locations (opt-in).

    install-steps hook — install_steps.sh: repo-local > cc-ci, or none.

@ -100,29 +101,22 @@ def resolve_op(recipe: str, op: str, repo_local_dir: str | None) -> tuple[str, s


 def custom_tests(recipe: str, repo_local_dir: str | None) -> list[tuple[str, str]]:
-    """All non-lifecycle test_*.py from cc-ci's tests/<recipe>/ and (if approved) the recipe's
-    repo-local tests/. Discovered locations (Phase 2 §4.1):
-      - the top-level dir   tests/<recipe>/test_*.py  (legacy + cross-cutting)
-      - functional/         tests/<recipe>/functional/test_*.py  (parity ports + recipe-specific)
-      - playwright/         tests/<recipe>/playwright/test_*.py  (UI flows P6)
-    Files named `test_<op>.py` (lifecycle ops) are excluded from this list — the orchestrator runs
-    those in their lifecycle tier, not the custom one. Repo-local is consulted only for
-    allowlist-approved recipes (HC2)."""
+    """All custom-tier test_*.py from cc-ci's tests/<recipe>/ and (if approved) the recipe's
+    repo-local tests/. PLACEMENT RULE (rcust P4): custom tests live ONLY under
+      - functional/   tests/<recipe>/functional/test_*.py  (parity ports + recipe-specific)
+      - playwright/   tests/<recipe>/playwright/test_*.py  (UI flows)
+    A top-level test_*.py is a LIFECYCLE OVERLAY (test_<op>.py) and nothing else — top-level
+    non-lifecycle files are NOT discovered (zero users at the time of the change; the lifecycle-
+    name exclusion below stays as a safety net so a misfiled test_<op>.py can never double-run).
+    Repo-local is consulted only for allowlist-approved recipes (HC2)."""
    lifecycle_names = {f"test_{op}.py" for op in LIFECYCLE_OPS}
    subdirs = ("functional", "playwright")
    found: list[tuple[str, str]] = []
    for source, d in (("cc-ci", cc_ci_dir(recipe)), ("repo-local", _gated(recipe, repo_local_dir))):
        if not d or not os.path.isdir(d):
            continue
-        # top-level (legacy / cross-cutting tests not under functional/playwright)
-        for p in sorted(glob.glob(os.path.join(d, "test_*.py"))):
-            if os.path.basename(p) not in lifecycle_names:
-                found.append((source, p))
-        # functional/ and playwright/ subdirs (Phase 2 §4.1)
        for sub in subdirs:
            for p in sorted(glob.glob(os.path.join(d, sub, "test_*.py"))):
-                # Phase-2 layout: lifecycle ops never live under functional/playwright, but be
-                # explicit so a misfiled file doesn't silently get double-run.
                if os.path.basename(p) not in lifecycle_names:
                    found.append((source, p))
    return found
@ -144,7 +138,7 @@ def install_steps(recipe: str, repo_local_dir: str | None) -> tuple[str, str] |

 def pre_op_hook(recipe: str, op: str, repo_local_dir: str | None) -> tuple[str, str] | None:
    """The pre-op seed hook for `op`: the path to a recipe `ops.py` module that defines a
-    `pre_<op>(domain, meta)` callable, or None. cc-ci's tests/<recipe>/ops.py wins; the repo-local
+    `pre_<op>(ctx)` callable, or None. cc-ci's tests/<recipe>/ops.py wins; the repo-local
    ops.py is consulted only for allowlist-approved recipes (HC2). The orchestrator imports the
    module and calls pre_<op> BEFORE performing the op (HC3 op/assertion split — overlays seed
    pre-op state here, then assert post-op in test_<op>.py)."""
--- a/runner/harness/generic.py
+++ b/runner/harness/generic.py
@ -19,22 +19,24 @@ import ssl
 import time

 from . import abra, lifecycle
+from . import meta as meta_mod

 # A recipe is backup-capable iff a compose file carries a truthy backupbot.backup label.
 _BACKUPBOT_RE = re.compile(r"backupbot\.backup\b[^\n]*\btrue\b", re.IGNORECASE)


 def _recipe_dir(recipe: str) -> str:
-    return os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    return abra.recipe_dir(recipe)  # the per-run tree inside a CI run ($ABRA_DIR)


-def backup_capable(recipe: str, meta: dict | None = None) -> bool:
+def backup_capable(recipe: str, meta=None) -> bool:
    """Whether the harness should run the backup/restore tiers (else they are a clean N/A skip, DG3).

-    `recipe_meta.BACKUP_CAPABLE` (bool) overrides; otherwise auto-detect by scanning the recipe's
-    compose*.yml for a truthy `backupbot.backup` label (the Co-op Cloud backup convention)."""
-    if meta and "BACKUP_CAPABLE" in meta:
-        return bool(meta["BACKUP_CAPABLE"])
+    `recipe_meta.BACKUP_CAPABLE` (bool) overrides when explicitly set (RecipeMeta default is None =
+    unset); otherwise auto-detect by scanning the recipe's compose*.yml for a truthy
+    `backupbot.backup` label (the Co-op Cloud backup convention)."""
+    if meta is not None and meta.BACKUP_CAPABLE is not None:
+        return bool(meta.BACKUP_CAPABLE)
    for path in glob.glob(os.path.join(_recipe_dir(recipe), "compose*.yml")):
        try:
            with open(path) as fh:
@ -75,7 +77,7 @@ def served_cert(domain: str, port: int = 443) -> tuple[bool, str]:
    return (True, f"CN={cn} SAN={sans}")


-def assert_serving(domain: str, meta: dict) -> None:
+def assert_serving(domain: str, meta) -> None:
    """The single generic "is the app really serving?" assertion (DG1).

    The app-vs-Traefik-fallback proof is steps 1+2 (both load-bearing, verified by the Adversary):
@ -90,14 +92,14 @@ def assert_serving(domain: str, meta: dict) -> None:

    Steps 1–2 are BOUNDED POLLS (no bare sleep), so a state-mutating op (upgrade/restore) that leaves
    the app briefly reconverging settles, while a persistent failure still fails within the timeout."""
-    deadline = time.time() + meta["DEPLOY_TIMEOUT"]
+    deadline = time.time() + meta.DEPLOY_TIMEOUT
    while time.time() < deadline and not lifecycle.services_converged(domain):
        time.sleep(5)
    assert lifecycle.services_converged(domain), f"{domain}: services did not converge"

-    path = meta["HEALTH_PATH"]
-    ok = tuple(meta["HEALTH_OK"])
-    deadline = time.time() + meta["HTTP_TIMEOUT"]
+    path = meta.HEALTH_PATH
+    ok = tuple(meta.HEALTH_OK)
+    deadline = time.time() + meta.HTTP_TIMEOUT
    served = False
    status, body = 0, ""
    while time.time() < deadline:
@ -141,7 +143,7 @@ def op_state() -> dict:
        return {}


-def assert_upgraded(domain: str, meta: dict) -> None:
+def assert_upgraded(domain: str, meta) -> None:
    """Generic UPGRADE assertion (post-op): the orchestrator already performed the upgrade once via
    `abra app deploy --chaos` of the PR-head checkout. Assert it reconverged + still serves AND that
    the deployment is genuinely the PR-head code under test (HC1) — non-vacuously (guarding F1d-2).
@ -212,7 +214,7 @@ def assert_backup_artifact(domain: str) -> str:
    return snap_id


-def assert_restore_healthy(domain: str, meta: dict) -> None:
+def assert_restore_healthy(domain: str, meta) -> None:
    """Generic RESTORE assertion (post-op): the orchestrator already restored. Assert the app is
    healthy + serving again (assert_serving polls, so the post-restore reconverge settles)."""
    assert_serving(domain, meta)
@ -222,7 +224,11 @@ def assert_restore_healthy(domain: str, meta: dict) -> None:


 def perform_upgrade(
-    domain: str, recipe: str, head_ref: str | None, deploy_timeout: int = 900, meta: dict | None = None
+    domain: str,
+    recipe: str,
+    head_ref: str | None,
+    deploy_timeout: int = 900,
+    meta=None,
 ) -> dict[str, str | None]:
    """Perform the UPGRADE op once, in place, to the PR-HEAD code under test (HC1): re-checkout the
    PR head (the prev-tag base deploy reset the recipe working tree), then `abra app deploy --chaos`
@ -240,7 +246,8 @@ def perform_upgrade(
    STRICTER convergence+health wait here: services N/N (wait_healthy) + app HEALTH_PATH healthy +
    any recipe READY_PROBE (collabora WOPI discovery 200). This bounds readiness by OUR generous
    deadline, not abra's impatient one — and is stronger evidence than abra's monitor."""
-    meta = meta or {}
+    if meta is None:
+        meta = meta_mod.load(recipe)
    before = lifecycle.deployed_identity(domain)
    if head_ref:
        lifecycle.recipe_checkout_ref(recipe, head_ref)
@ -249,9 +256,7 @@ def perform_upgrade(
    # (target) version, so the base deploys minimally WITHOUT it and the upgrade adds it to COMPOSE_FILE
    # here, after the PR-head checkout (which ships the overlay) and before the chaos redeploy that
    # picks up the new .env. Dict or callable(domain)->dict. No-op for recipes without it.
-    upgrade_env = meta.get("UPGRADE_EXTRA_ENV") or {}
-    if callable(upgrade_env):
-        upgrade_env = upgrade_env(domain) or {}
+    upgrade_env = meta_mod.upgrade_extra_env(meta, meta_mod.hook_ctx(domain, meta, op="upgrade"))
    for k, v in upgrade_env.items():
        print(f"  upgrade-env: {k}={v}", flush=True)
        abra.env_set(domain, k, v)
@ -262,12 +267,12 @@ def perform_upgrade(
    # Own the convergence verification (abra's monitor was skipped via -c).
    lifecycle.wait_healthy(
        domain,
-        ok_codes=tuple(meta.get("HEALTH_OK", (200, 301, 302))),
-        path=meta.get("HEALTH_PATH", "/"),
-        deploy_timeout=int(meta.get("DEPLOY_TIMEOUT", deploy_timeout)),
-        http_timeout=int(meta.get("HTTP_TIMEOUT", 300)),
+        ok_codes=tuple(meta.HEALTH_OK),
+        path=meta.HEALTH_PATH,
+        deploy_timeout=int(meta.DEPLOY_TIMEOUT),
+        http_timeout=int(meta.HTTP_TIMEOUT),
    )
-    lifecycle.wait_ready_probes(meta, domain, timeout=int(meta.get("DEPLOY_TIMEOUT", deploy_timeout)))
+    lifecycle.wait_ready_probes(meta, domain, timeout=int(meta.DEPLOY_TIMEOUT), op="upgrade")
    after = lifecycle.deployed_identity(domain)
    # Evidence (HC1): the chaos-version label = the deployed recipe commit; it should match the
    # PR-head we checked out — proving the upgrade deployed the code under test, not a published tag.
--- a/runner/harness/http.py
+++ b/runner/harness/http.py
@ -73,7 +73,7 @@ def http_post(
    `data` is JSON-encoded if content_type='application/json',
    form-encoded if 'application/x-www-form-urlencoded' (the OIDC token endpoint form),
    or sent raw bytes if data is already bytes."""
-    if isinstance(data, (bytes, bytearray)):
+    if isinstance(data, bytes | bytearray):
        body: bytes | None = bytes(data)
    elif content_type == "application/json" and data is not None:
        body = json.dumps(data).encode()
@ -107,7 +107,7 @@ def http_request(
 ) -> tuple[int, object | None]:
    """Arbitrary-method HTTP (PUT/DELETE/PATCH) for parity tests that mutate. Same shape as
    http_post (returns (status, json_or_None))."""
-    if isinstance(data, (bytes, bytearray)):
+    if isinstance(data, bytes | bytearray):
        body: bytes | None = bytes(data)
    elif content_type == "application/json" and data is not None:
        body = json.dumps(data).encode()
@ -142,7 +142,7 @@ def post_with_headers(
    """Like http_post but ALSO returns the response headers as a dict — for APIs that hand back an
    auth token in a response header rather than the body (e.g. mattermost login → `Token` header).
    Returns (status, parsed_json_or_None, response_headers). status=0 + {} on transport failure."""
-    if isinstance(data, (bytes, bytearray)):
+    if isinstance(data, bytes | bytearray):
        body: bytes | None = bytes(data)
    elif content_type == "application/json" and data is not None:
        body = json.dumps(data).encode()
@ -252,13 +252,16 @@ def retry_http_post(
 ) -> tuple[int, object | None]:
    """POST with retry until expect_fn(status, json) is truthy. Defaults to any 2xx."""
    if expect_fn is None:
+
        def expect_fn(s, _j):  # noqa: ARG001
            return 200 <= s < 300

    result: list[tuple[int, object | None]] = [(0, None)]

    def _check():
-        s, j = http_post(url, data=data, headers=headers, content_type=content_type, timeout=timeout)
+        s, j = http_post(
+            url, data=data, headers=headers, content_type=content_type, timeout=timeout
+        )
        result[0] = (s, j)
        return expect_fn(s, j)

--- a/runner/harness/lifecycle.py
+++ b/runner/harness/lifecycle.py
@ -7,17 +7,20 @@ next run. Callers wrap deploy()/teardown() in try/finally (or a pytest finalizer
 from __future__ import annotations

 import contextlib
-import datetime
+import fcntl
+import glob
 import json
 import os
 import re
+import shutil
 import socket
 import ssl
 import subprocess
 import time
 import urllib.request

-from . import abra
+from . import abra, lifetime
+from . import meta as meta_mod

 GATEWAY_IP = "143.244.213.108"  # *.ci.commoninternet.net -> gateway (TLS passthrough to cc-ci)
 # A run app domain is "<recipe[:4]>-<6hex>.ci.commoninternet.net" (see DECISIONS.md). Used by the
@ -29,6 +32,68 @@ class TeardownError(RuntimeError):
    pass


+# --- Concurrent-run safety (capacity=2) -------------------------------------------------------
+# ONE mechanism, process-lifetime-scoped so SIGKILL can't leak a stale claim: every run holds an
+# exclusive kernel flock on its app DOMAIN (/run/lock/cc-ci-app-<domain>.lock) for the whole run.
+# A held lock implies a live owner — the kernel releases a flock when the holding process dies,
+# however it dies. The janitor probes the lock (LOCK_NB) to tell a live concurrent run (held →
+# leave it) from a crashed run's orphan (acquirable → reap it); it never inspects pids and never
+# steals a held lock. Recipe-tree corruption between same-recipe runs is gone structurally (each
+# run deploys from its own per-run ABRA_DIR — there is no shared recipe tree and no recipe lock),
+# and same-domain runs (double-!testme of one PR) serialise on this app lock.
+# See docs/concurrency.md.
+
+# Acquired app-lock file objects are retained here for the REMAINING PROCESS LIFETIME: if the
+# caller drops the returned file object, GC would close the fd and silently release the lock —
+# this list is the lock's owner of record. Never cleared; release is process exit.
+_held_app_locks: list = []
+
+
+def _app_lock_dir() -> str:
+    """The app-domain lockfile dir. /run/lock (tmpfs: a reboot clears locks AND lockfiles, so
+    post-reboot apps probe as orphans and are reaped immediately). Env-overridable so the
+    tests/concurrency suite (and its helper subprocesses) can use a sandbox dir."""
+    return os.environ.get("CCCI_APP_LOCK_DIR", "/run/lock")
+
+
+def _app_lock_path(domain: str) -> str:
+    return os.path.join(_app_lock_dir(), f"cc-ci-app-{domain}.lock")
+
+
+def acquire_app_lock(domain: str):
+    """Take the per-app-domain exclusive lock; blocks (with a log line) if another run of the
+    same domain is in flight (double-!testme serialisation). Returns the open lock file, which is
+    ALSO retained in _held_app_locks so the flock lives exactly as long as the process.
+
+    Unlink/recreate race guard: the janitor unlinks a reaped orphan's lockfile while holding its
+    flock, so a waiter blocked on the OLD inode can win a lock no later opener can observe (a new
+    open() at the path creates a FRESH inode). After every acquisition, verify the locked fd is
+    still the file at the path (st_ino match); if not, drop it and retry on the live path."""
+    path = _app_lock_path(domain)
+    waited = False
+    while True:
+        # PEP 446: the fd is non-inheritable, so subprocess children never carry the lock.
+        f = open(path, "a")  # noqa: SIM115 — deliberately held for the rest of the process
+        try:
+            fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)
+        except BlockingIOError:
+            if not waited:
+                print(f"== app lock: another run of {domain} is in flight — waiting ==", flush=True)
+                waited = True
+            fcntl.flock(f, fcntl.LOCK_EX)
+        try:
+            if os.fstat(f.fileno()).st_ino == os.stat(path).st_ino:
+                break  # we hold the lock on the inode the path names — done
+        except FileNotFoundError:
+            pass
+        f.close()  # locked a stale (unlinked) inode — retry on the live path
+    os.utime(f.fileno())  # mtime = acquisition time = lock age (janitor's long-held flag)
+    _held_app_locks.append(f)
+    if waited:
+        print(f"== app lock: acquired {path} ==", flush=True)
+    return f
+
+
 def _docker_names(kind: str, stack: str) -> list[str]:
    """docker <kind> ls names filtered to a stack (kind: service|volume|secret)."""
    proc = subprocess.run(
@ -48,62 +113,6 @@ def _residual(domain: str) -> dict:
    }


-def _stack_age_seconds(stack: str) -> float | None:
-    """Age of the stack's oldest service, or None if not present."""
-    svcs = _docker_names("service", stack)
-    if not svcs:
-        return None
-    oldest = None
-    for s in svcs:
-        p = subprocess.run(
-            ["docker", "service", "inspect", s, "--format", "{{.CreatedAt}}"],
-            capture_output=True,
-            text=True,
-        )
-        ts = p.stdout.strip()
-        try:
-            # docker emits e.g. 2026-05-27 00:12:33.123 +0000 UTC -> take the leading 19 chars
-            dt = datetime.datetime.strptime(ts[:19], "%Y-%m-%d %H:%M:%S").replace(
-                tzinfo=datetime.UTC
-            )
-        except ValueError:
-            continue
-        age = (datetime.datetime.now(datetime.UTC) - dt).total_seconds()
-        oldest = age if oldest is None else max(oldest, age)
-    return oldest
-
-
-def _recipe_extra_env(recipe: str, domain: str) -> dict[str, str]:
-    """Per-recipe extra .env keys, applied at every deploy (install + upgrade's old_app) so a recipe
-    with multi-domain / config needs is enrolled with NO shared-harness change (D5/M6.5). A recipe
-    declares `EXTRA_ENV` in tests/<recipe>/recipe_meta.py as either a dict or a callable
-    `EXTRA_ENV(domain) -> dict` (callable form lets it derive values from the per-run domain, e.g.
-    cryptpad's SANDBOX_DOMAIN). Returns {} if none."""
-    path = os.path.join(os.path.dirname(__file__), "..", "..", "tests", recipe, "recipe_meta.py")
-    if not os.path.exists(path):
-        return {}
-    ns: dict = {}
-    with open(path) as fh:
-        exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
-    ee = ns.get("EXTRA_ENV")
-    if callable(ee):
-        ee = ee(domain)
-    return {str(k): str(v) for k, v in (ee or {}).items()}
-
-
-def _recipe_meta_flag(recipe: str, key: str) -> bool:
-    """Read a boolean flag from tests/<recipe>/recipe_meta.py (e.g. CHAOS_BASE_DEPLOY). Returns
-    False if the recipe ships no meta or the flag is absent/falsey. Trusted in-repo exec, same as
-    _recipe_extra_env."""
-    path = os.path.join(os.path.dirname(__file__), "..", "..", "tests", recipe, "recipe_meta.py")
-    if not os.path.exists(path):
-        return False
-    ns: dict = {}
-    with open(path) as fh:
-        exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
-    return bool(ns.get(key))
-
-
 def _record_deploy() -> None:
    """Increment the per-run deploy counter (DG4.1: one deploy per run). No-op unless the
    orchestrator set CCCI_DEPLOY_COUNT_FILE — so it never affects standalone/manual use."""
@ -117,6 +126,34 @@ def _record_deploy() -> None:
        f.write(str(n + 1))


+def ccci_overlay_path(recipe: str) -> str:
+    """The cc-ci-owned compose overlay for a recipe (rcust P2a: first-class, auto-discovered)."""
+    return os.path.join(meta_mod.TESTS_DIR, recipe, "compose.ccci.yml")
+
+
+def has_ccci_overlay(recipe: str) -> bool:
+    return os.path.isfile(ccci_overlay_path(recipe))
+
+
+def provide_ccci_overlay(recipe: str) -> None:
+    """Copy tests/<recipe>/compose.ccci.yml into THIS run's recipe checkout (ABRA_DIR-aware), so
+    the recipe's COMPOSE_FILE reference resolves (rcust P2a — the harness owns the copy; recipes
+    no longer ship install_steps.sh boilerplate for it). No-op for recipes without an overlay."""
+    src = ccci_overlay_path(recipe)
+    if not os.path.isfile(src):
+        return
+    dest_dir = abra.recipe_dir(recipe)
+    if not os.path.isdir(dest_dir):
+        print(f"  ccci-overlay: recipe dir {dest_dir} missing — cannot provide overlay", flush=True)
+        raise RuntimeError(f"recipe checkout missing for {recipe}: {dest_dir}")
+    shutil.copy(src, os.path.join(dest_dir, "compose.ccci.yml"))
+    print(
+        f"  ccci-overlay: provided compose.ccci.yml to the {recipe} checkout "
+        "(first-class overlay; base deploy auto-chaos)",
+        flush=True,
+    )
+
+
 def _run_install_steps(hook: tuple[str, str], recipe: str, domain: str) -> None:
    """Run a recipe's custom install-steps hook (install_steps.sh) during the install tier — after
    `abra app new` + env defaults + secret generate, before deploy (Phase 1d DG5). The hook gets the
@ -149,9 +186,9 @@ def prepull_images(recipe: str, domain: str) -> None:
    app-INIT time (slow-init apps like collabora/immich still need their recipe healthcheck/READY_PROBE).
    Best-effort on resolution failure (skip + let the deploy pull as usual); HARD-fails on a real
    pull error (don't mask it)."""
-    import os
-
-    recipe_dir = os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    recipe_dir = abra.recipe_dir(recipe)  # per-run tree inside a CI run
+    # The app .env lives in the CANONICAL servers path (the per-run ABRA_DIR's servers/ is a
+    # symlink to it, so abra and this path agree on the same file).
    env_path = os.path.expanduser(f"~/.abra/servers/default/{domain}.env")
    if not os.path.isdir(recipe_dir) or not os.path.isfile(env_path):
        print(f"  prepull: recipe dir or .env missing for {recipe} — skipping", flush=True)
@ -161,7 +198,8 @@ def prepull_images(recipe: str, domain: str) -> None:
    # --env-file supplies $VERSION-style interpolation so pinned tags resolve correctly.
    cf = subprocess.run(
        ["bash", "-c", f'set -a; . "{env_path}"; printf "%s" "${{COMPOSE_FILE:-compose.yml}}"'],
-        capture_output=True, text=True,
+        capture_output=True,
+        text=True,
    ).stdout.strip()
    files = [f for f in cf.split(":") if f] or ["compose.yml"]
    args = ["docker", "compose", "--env-file", env_path]
@ -199,16 +237,28 @@ def deploy_app(
    secrets: bool = True,
    install_steps_hook: tuple[str, str] | None = None,
    deploy_timeout: int = 900,
+    meta=None,
 ) -> None:
    """Create + configure + deploy an app. Forces LETS_ENCRYPT_ENV='' so traefik serves the
    wildcard cert via the file provider and NEVER attempts ACME (adversary finding A1). Applies any
-    per-recipe EXTRA_ENV (recipe_meta.py) and the custom install-steps hook (Phase 1d) before deploy.
+    per-recipe EXTRA_ENV (recipe_meta.py), the custom install-steps hook (Phase 1d), and the
+    first-class `tests/<recipe>/compose.ccci.yml` overlay (rcust P2a) before deploy.
+
+    `meta` is the recipe's loaded RecipeMeta (EXTRA_ENV); the orchestrator loads once and passes
+    it down. Callers without one in hand (fixtures, warm reconcile) may omit it — it is then
+    loaded here via the single meta.load() path.

    `deploy_timeout` is the subprocess timeout for `abra app deploy`. Caller (orchestrator) passes
    `recipe_meta.DEPLOY_TIMEOUT` so heavy recipes (ghost, matrix-synapse, lasuite-meet) can extend
    past the 900s default. abra's INTERNAL TIMEOUT (recipe's TIMEOUT env, default 300s) is set via
    EXTRA_ENV; this is the Python subprocess wrapper's timeout so abra doesn't get SIGKILLed mid-deploy."""
+    if meta is None:
+        meta = meta_mod.load(recipe)
    _record_deploy()
+    # Lock BEFORE the app exists: a concurrent run's janitor must never see this app without a
+    # held app lock (it would probe it as an orphan and reap an in-flight deploy). Also the
+    # double-!testme serialisation point: a second run of the same domain blocks here.
+    acquire_app_lock(domain)
    abra.app_config_remove(domain)  # clear any stale .env from a prior crashed run
    abra.app_new(recipe, domain, version=version, secrets=secrets)
    # A pinned version must actually deploy that version: check the recipe out to the tag so the
@ -231,16 +281,18 @@ def deploy_app(
                flush=True,
            )
            chaos = True
-        # A recipe may force a chaos base deploy via recipe_meta CHAOS_BASE_DEPLOY=True when an
-        # install_steps hook adds an untracked compose overlay to the recipe checkout (e.g. discourse's
-        # compose.ccci.yml, provided by install_steps for the pinned base). The untracked file makes
-        # abra's pinned-deploy clean-tree check FATA ('has locally unstaged changes'); chaos skips lint +
-        # the clean-tree gate and deploys the EXPLICITLY-checked-out pinned version (we already ran
-        # recipe_checkout(version) above) — NOT latest. Same mechanism as the lightweight-tag branch.
-        elif _recipe_meta_flag(recipe, "CHAOS_BASE_DEPLOY"):
+        # A first-class cc-ci compose overlay (tests/<recipe>/compose.ccci.yml, copied into the
+        # checkout below — rcust P2a) is an UNTRACKED file in the recipe checkout, which makes
+        # abra's pinned-deploy clean-tree check FATA ('has locally unstaged changes'). Auto-chaos:
+        # chaos skips lint + the clean-tree gate and deploys the EXPLICITLY-checked-out pinned
+        # version (we already ran recipe_checkout(version) above) — NOT latest. Same mechanism as
+        # the lightweight-tag branch. (Replaces the deleted CHAOS_BASE_DEPLOY meta flag — the
+        # overlay's presence IS the signal, killing the R7 implicit coupling.)
+        elif has_ccci_overlay(recipe):
            print(
-                f"  deploy_app({recipe}@{version}): CHAOS_BASE_DEPLOY set → chaos base deploy of the "
-                "checked-out pinned version (skips clean-tree/lint; deploys version, not LATEST)",
+                f"  deploy_app({recipe}@{version}): compose.ccci.yml overlay present → chaos base "
+                "deploy of the checked-out pinned version (skips clean-tree/lint; deploys version, "
+                "not LATEST)",
                flush=True,
            )
            chaos = True
@ -250,12 +302,18 @@ def deploy_app(
    # it ourselves is recipe-agnostic and canonical (the run domain IS the app's domain).
    abra.env_set(domain, "DOMAIN", domain)
    abra.env_set(domain, "LETS_ENCRYPT_ENV", "")
-    for k, v in _recipe_extra_env(recipe, domain).items():
+    for k, v in meta_mod.extra_env(meta, meta_mod.hook_ctx(domain, meta)).items():
        abra.env_set(domain, k, v)
    if secrets:
        abra.secret_generate(domain)
    if install_steps_hook:
        _run_install_steps(install_steps_hook, recipe, domain)
+    # First-class cc-ci compose overlay (rcust P2a): if the recipe ships
+    # tests/<recipe>/compose.ccci.yml, copy it into THIS run's recipe checkout (ABRA_DIR-aware)
+    # so the COMPOSE_FILE reference in the recipe's EXTRA_ENV resolves. Untracked, so it persists
+    # across the later PR-head checkout (idempotent when the head ships the same fix). Replaces
+    # the per-recipe install_steps.sh copy boilerplate + CHAOS_BASE_DEPLOY flag (auto-chaos above).
+    provide_ccci_overlay(recipe)
    # HQ1: warm the local image store before the (real, unchanged) abra deploy.
    prepull_images(recipe, domain)
    abra.deploy(domain, chaos=chaos, timeout=deploy_timeout)
@ -268,18 +326,22 @@ def _stack_name(domain: str) -> str:


 def services_converged(domain: str) -> bool:
-    """True when every service in the stack reports replicas N/N (N>0)."""
+    """True when every service in the stack reports replicas N/N (N>0) AND no service is
+    mid-rolling-update (swarm UpdateStatus settled)."""
    stack = _stack_name(domain)
    proc = subprocess.run(
-        ["docker", "stack", "services", stack, "--format", "{{.Replicas}}"],
+        ["docker", "stack", "services", stack, "--format", "{{.Name}} {{.Replicas}}"],
        capture_output=True,
        text=True,
    )
    rows = [r for r in proc.stdout.split("\n") if r.strip()]
    if not rows:
        return False
+    names = []
    for r in rows:
-        cur, _, want = r.partition("/")
+        name, _, replicas = r.partition(" ")
+        names.append(name)
+        cur, _, want = replicas.partition("/")
        # A service at its DESIRED replica count is converged — including a `replicas: 0`
        # on-demand one-shot (e.g. lasuite-drive's `minio-createbuckets`, which is scaled up
        # manually only when buckets need (re)creating), which reports "0/0". The earlier
@ -288,6 +350,34 @@ def services_converged(domain: str) -> bool:
        # still spinning up shows e.g. "0/1" (cur != want) and is correctly not-yet-converged.
        if not want or cur != want:
            return False
+    # N/N alone is NOT convergence during a stop-first rolling update: a chaos redeploy that changes
+    # a non-app service image (e.g. immich's db pin) registers the update immediately, but swarm may
+    # not have cycled that service's task yet — the OLD task still shows 1/1, then dies seconds later
+    # (immich CI 238: backupbot exec'd the db pre-hook into the just-killed container → 409). Require
+    # every service's UpdateStatus to be settled too, so the wait spans the whole rolling update.
+    proc = subprocess.run(
+        [
+            "docker",
+            "service",
+            "inspect",
+            *names,
+            "--format",
+            "{{if .UpdateStatus}}{{.UpdateStatus.State}}{{end}}",
+        ],
+        capture_output=True,
+        text=True,
+    )
+    if proc.returncode != 0:
+        return False  # a service vanished mid-check — not settled
+    for state in proc.stdout.split("\n"):
+        # Only ACTIVE states block convergence. 'paused'/'rollback_paused' are terminal-without-
+        # intervention: swarm's default update-failure-action pauses the update on one task flicker
+        # and the flag then persists FOREVER (immich CI 241: app service 'paused' from a restart
+        # during restore, service back at 1/1 and healthy — the wait hung to its deadline). With
+        # N/N already required above, a paused update is settled for our purposes; the HTTP-health
+        # and tier assertions still gate whether the app actually works.
+        if state.strip() in ("updating", "rollback_started"):
+            return False
    return True


@ -415,7 +505,9 @@ def recipe_checkout_ref(recipe: str, ref: str) -> None:
    abra.recipe_checkout(recipe, ref)


-def chaos_redeploy(domain: str, deploy_timeout: int = 900, no_converge_checks: bool = False) -> None:
+def chaos_redeploy(
+    domain: str, deploy_timeout: int = 900, no_converge_checks: bool = False
+) -> None:
    """In-place `abra app deploy --chaos`: redeploy the running app at the CURRENT recipe checkout
    (HC1: the PR-head code under test). This is the upgrade op, not a fresh install — it does NOT go
    through deploy_app, so the deploy-count guard (DG4.1) is not incremented.
@ -433,7 +525,7 @@ def chaos_redeploy(domain: str, deploy_timeout: int = 900, no_converge_checks: b
    abra.deploy(domain, chaos=True, timeout=deploy_timeout, no_converge_checks=no_converge_checks)


-def wait_ready_probes(meta: dict, domain: str, timeout: int = 600) -> None:
+def wait_ready_probes(meta, domain: str, timeout: int = 600, op: str | None = None) -> None:
    """Poll a recipe's optional READY_PROBE endpoints until each returns an accepted status, or raise.

    A recipe_meta may define `READY_PROBE(domain) -> [{"host":..., "path":..., "ok":(200,)}, ...]`
@ -450,10 +542,10 @@ def wait_ready_probes(meta: dict, domain: str, timeout: int = 600) -> None:
    must be released by the old task + rebound by the new) the voice server can be down while
    HTTP-200 still passes — and backup-bot then execs into a not-running app container (409). Requiring
    the voice port to be stably listening before proceeding closes that window."""
-    probe_fn = meta.get("READY_PROBE")
+    probe_fn = meta.READY_PROBE
    if not callable(probe_fn):
        return
-    probes = probe_fn(domain) or []
+    probes = probe_fn(meta_mod.hook_ctx(domain, meta, op=op)) or []
    for probe in probes:
        if "tcp_port" in probe:
            host = probe.get("tcp_host", "127.0.0.1")
@ -498,6 +590,16 @@ def wait_ready_probes(meta: dict, domain: str, timeout: int = 600) -> None:

 def backup_app(domain: str) -> str:
    """Create a backup; return the abra/restic output (carries the produced snapshot_id)."""
+    # Never back up a stack that is still converging/rolling-updating: backupbot resolves each
+    # service's hook container ONCE up front, so a task that cycles between that lookup and the
+    # pre-hook exec crashes the whole backup with a 409 (immich CI 238). Bounded wait — on timeout
+    # we still attempt the backup and let the tier's assertion deliver the verdict.
+    deadline = time.time() + 300
+    while time.time() < deadline and not services_converged(domain):
+        print(
+            f"  backup: {domain} stack not settled yet — waiting before backup create", flush=True
+        )
+        time.sleep(5)
    return abra.backup_create(domain)


@ -603,17 +705,84 @@ def teardown_app(domain: str, verify: bool = True) -> None:
        residual = _residual(domain)
        if any(residual.values()):
            raise TeardownError(f"teardown left residual for {domain}: {residual}")
+    # No unregistration step: the app lock releases implicitly at process exit. The clean run's
+    # leftover lockfile (unheld) is unlinked on sight by the next janitor's stale-lockfile sweep.


-def janitor(max_age_seconds: int | None = None) -> None:
-    """Reap orphaned run apps from crashed/rebooted runs. Matches the real naming scheme and only
-    reaps apps older than max_age_seconds (so concurrent in-flight runs are never killed). Reaps via
-    docker primitives so it works even when the .env is gone (A2/A3). Default 2h, env-overridable
-    via CCCI_JANITOR_MAX_AGE (e.g. 0 to reap all matching orphans immediately)."""
-    import os
+# A lock held longer than 2x the 60-min hard deadline can only be a leaked run (the deadline
+# bounds every healthy run). Flag it for a human — NEVER steal a held lock.
+LONG_HELD_LOCK_SECONDS = 2 * lifetime.HARD_DEADLINE_SECONDS

-    if max_age_seconds is None:
-        max_age_seconds = int(os.environ.get("CCCI_JANITOR_MAX_AGE", "7200"))
+
+def _probe_and_reap(domain: str) -> None:
+    """Probe one run app's lock; reap iff nobody holds it (kernel-guaranteed orphan).
+
+    Reaping happens WHILE HOLDING the probe lock, closing the janitor-vs-new-run race: a new run
+    of the same domain blocks in acquire_app_lock until the reap finishes, so a fresh app never
+    coexists with a half-reaped one. The lockfile is unlinked before release (still holding the
+    lock); a waiter that blocked on the unlinked inode re-checks identity and retries. Two racing
+    janitors arbitrate on the same flock: one reaps, the other sees 'held' and leaves —
+    teardown_app(verify=False) is idempotent either way."""
+    path = _app_lock_path(domain)
+    try:
+        # PEP 446: non-inheritable fd, same as acquire_app_lock.
+        f = open(path, "a")  # noqa: SIM115 — closed in the finally below, lock released with it
+    except OSError as e:
+        print(f"!! janitor: cannot open lockfile {path} ({e}) — skipping {domain}", flush=True)
+        return
+    try:
+        try:
+            fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)
+        except BlockingIOError:
+            # Held -> live run. Never steal; flag if it has been held implausibly long.
+            try:
+                held_for = time.time() - os.stat(path).st_mtime
+            except OSError:
+                held_for = 0
+            if held_for > LONG_HELD_LOCK_SECONDS:
+                print(
+                    f"!! lock for {domain} held >{LONG_HELD_LOCK_SECONDS // 60}min — possible "
+                    "leaked run; inspect with lslocks",
+                    flush=True,
+                )
+            else:
+                print(
+                    f"  janitor: {domain} lock held — live concurrent run, leaving it", flush=True
+                )
+            return
+        # Acquired — but only the inode the PATH names counts (another janitor may have reaped
+        # and unlinked this inode while we raced; a lock on an unlinked inode protects nothing
+        # and unlinking the path now would delete a NEWER run's lockfile).
+        try:
+            if os.fstat(f.fileno()).st_ino != os.stat(path).st_ino:
+                return
+        except FileNotFoundError:
+            return
+        # Orphan: no live owner (the kernel released the lock when the owner died). Reap while
+        # holding the probe lock, then unlink the lockfile before releasing.
+        print(f"  janitor: {domain} lock acquirable — orphan, reaping", flush=True)
+        with contextlib.suppress(Exception):
+            teardown_app(domain, verify=False)
+        with contextlib.suppress(OSError):
+            os.unlink(path)
+    finally:
+        f.close()
+
+
+def janitor() -> None:
+    """Reap orphaned run apps from crashed/rebooted runs; the kernel flock is the only liveness
+    oracle. For every candidate run app, probe its app-domain lock (LOCK_NB):
+
+      acquirable -> nobody holds it -> orphan -> reap under the probe lock + unlink lockfile
+      held       -> live concurrent run -> leave it (warn if held >2x the hard deadline)
+
+    Candidate discovery is unchanged: `abra app ls` + a docker-service sweep (catches stacks
+    whose .env is already gone), both matched against RUN_APP_RE — warm/canonical apps never
+    match and are never probed. Post-reboot, /run/lock (tmpfs) is empty, so every surviving app
+    probes as an orphan and is reaped immediately (no age threshold). Stale lockfiles with no
+    app behind them are unlinked on sight. Degrades safely: an unreadable lockfile/dir is
+    skipped with a log line, never a crash. Reaps via docker primitives so it works even when
+    the .env is gone (A2/A3)."""
    seen = set()
    for app in abra.app_ls():
        name = app.get("appName") or app.get("domain") or ""
@ -627,9 +796,22 @@ def janitor(max_age_seconds: int | None = None) -> None:
            seen.add(f"{m.group(1)}.ci.commoninternet.net")

    for name in seen:
-        stack = _stack_name(name)
-        age = _stack_age_seconds(stack)
-        if age is not None and age < max_age_seconds:
-            continue  # likely a concurrent in-flight run; leave it
-        with contextlib.suppress(Exception):
-            teardown_app(name, verify=False)
+        _probe_and_reap(name)
+
+    # Tidy /run/lock: a clean run's leftover lockfile is unheld and appless — unlink it (under
+    # its own probe lock, with the same identity check as above).
+    with contextlib.suppress(OSError):
+        for path in glob.glob(os.path.join(_app_lock_dir(), "cc-ci-app-*.lock")):
+            domain = os.path.basename(path)[len("cc-ci-app-") : -len(".lock")]
+            if domain in seen:
+                continue  # handled (or deliberately left) above
+            with contextlib.suppress(OSError):
+                f = open(path, "a")  # noqa: SIM115 — closed below, lock released with it
+                try:
+                    fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)
+                    if os.fstat(f.fileno()).st_ino == os.stat(path).st_ino:
+                        os.unlink(path)
+                except (BlockingIOError, FileNotFoundError):
+                    pass  # held (live run pre-deploy) or already gone — leave it
+                finally:
+                    f.close()
--- a/runner/harness/lifetime.py
+++ b/runner/harness/lifetime.py
@ -0,0 +1,95 @@
+"""Run-lifetime hardening (concurrency restructure P1).
+
+The concurrency model's invariant chain is:
+
+    lock lifetime ⊆ harness process lifetime ⊆ drone step lifetime ⊆ 60-min hard deadline
+
+Locks are kernel flocks released on process exit, so the only thing that needs managing is the
+PROCESS lifetime. Three guards, installed at run startup (before any abra call) by
+`install_lifetime_guards()`:
+
+  1. `PR_SET_PDEATHSIG(SIGTERM)`: if the parent (the drone step shell) dies — cancel, runner
+     crash, host shutdown of the step — the kernel delivers SIGTERM to the harness, so a dead
+     build can never leak a running harness that holds locks. Paired with a ppid==1 re-check
+     AFTER the prctl: a parent that died BEFORE the prctl took effect would never trigger the
+     death signal, so a harness that finds itself already reparented refuses to run.
+  2. SIGTERM handler: raise SystemExit so the run's `finally:` teardown funnel executes and the
+     process exits non-zero. Re-entrant deliveries during teardown are logged and IGNORED so a
+     second signal can't abort the cleanup the first one asked for (`begin_teardown()` guards
+     this; the run's own `finally:` blocks also call it so a signal landing mid-normal-teardown
+     can't abort that either).
+  3. `signal.alarm(3600)`: self-imposed hard deadline. SIGALRM funnels into the same teardown
+     path with a distinct log line. Teardown time after the deadline is not alarm-bounded —
+     interrupting a teardown buys nothing; the janitor (flock probe) is the backstop if a
+     teardown wedges and the process is killed harder.
+"""
+
+from __future__ import annotations
+
+import ctypes
+import os
+import signal
+import sys
+
+HARD_DEADLINE_SECONDS = 60 * 60
+
+_PR_SET_PDEATHSIG = 1  # linux/prctl.h
+
+_state = {"tearing_down": False}
+
+
+def begin_teardown() -> None:
+    """Mark the teardown funnel as running. From here on SIGTERM/SIGALRM must NOT raise — it
+    would abort the very cleanup it asks for — so the handlers log and return instead. Called by
+    the handlers themselves before raising, and at the top of the run's `finally:` blocks."""
+    _state["tearing_down"] = True
+
+
+def _funnel_handler(log_line: str, exit_code: int):
+    """A signal handler that routes into the teardown funnel exactly once: log, then raise
+    SystemExit (propagates through the run's try/finally → teardown executes → non-zero exit).
+    While teardown is already running, further signals are logged and swallowed."""
+
+    def handler(signum: int, frame) -> None:  # noqa: ARG001
+        print(log_line, flush=True)
+        if _state["tearing_down"]:
+            print(
+                f"== signal {signum} during teardown — ignored (teardown continues, "
+                "exit stays non-zero) ==",
+                flush=True,
+            )
+            return
+        begin_teardown()
+        raise SystemExit(exit_code)
+
+    return handler
+
+
+def install_lifetime_guards(deadline_seconds: int = HARD_DEADLINE_SECONDS) -> None:
+    """Install all three lifetime guards (see module docstring). Must run at harness startup,
+    before any abra call and before any lock is taken."""
+    libc = ctypes.CDLL("libc.so.6", use_errno=True)
+    if libc.prctl(_PR_SET_PDEATHSIG, signal.SIGTERM, 0, 0, 0) != 0:
+        err = ctypes.get_errno()
+        raise OSError(err, f"prctl(PR_SET_PDEATHSIG, SIGTERM) failed: {os.strerror(err)}")
+    # The prctl is armed now — but only fires for a parent death AFTER this point. If the parent
+    # already died, we are reparented (ppid 1) and would never get the signal: refuse to run, an
+    # orphaned harness would hold locks/apps with nothing managing its lifetime.
+    if os.getppid() == 1:
+        sys.exit("parent died before prctl(PR_SET_PDEATHSIG) — refusing to run orphaned")
+    signal.signal(
+        signal.SIGTERM,
+        _funnel_handler(
+            "== SIGTERM received (drone cancel / parent death) — tearing down ==",
+            128 + signal.SIGTERM,
+        ),
+    )
+    minutes = deadline_seconds // 60
+    signal.signal(
+        signal.SIGALRM,
+        _funnel_handler(
+            f"== run exceeded {minutes}-minute hard deadline — tearing down ==",
+            128 + signal.SIGALRM,
+        ),
+    )
+    signal.alarm(deadline_seconds)
--- a/runner/harness/manifest.py
+++ b/runner/harness/manifest.py
@ -0,0 +1,153 @@
+"""Customization manifest (rcust P5; spec §8 R4 mitigation).
+
+One block at run start answering "what does this recipe customize?" across ALL the surfaces
+(recipe_meta keys, hook files, file-presence, run-time env overrides) — printed to the run log and
+embedded verbatim in results.json under "customization". PURE PRESENTATION: building or printing
+the manifest must never influence any verdict (R7-class invariant).
+"""
+
+from __future__ import annotations
+
+import os
+import re
+
+from . import discovery, lifecycle
+from . import meta as meta_mod
+
+_PRE_OP_RE = re.compile(r"^def (pre_[a-z]+)\(", re.MULTILINE)
+
+# Meta values are repo-public by construction (recipe_meta.py is committed; real secrets are
+# class-B generated, never meta), but the manifest lands on the dashboard — mask values whose
+# key NAME is secret-shaped so a field literally called SECRET_KEY_BASE never shows a value
+# (defense in depth + keeps dashboard secret-scans quiet). `KEY` matches only as a word segment
+# (API_KEY yes, KEYCLOAK_URL no).
+_SENSITIVE_NAME_RE = re.compile(r"SECRET|PASSWORD|TOKEN|CREDENTIAL|(^|_)KEY(_|$)", re.IGNORECASE)
+
+
+def _jsonable(v, name=""):
+    """Manifest values must be JSON-serializable + deterministic: hooks render as '<hook>',
+    tuples become lists, secret-named entries (by key name, incl. nested dict keys) as
+    '<redacted>'."""
+    if callable(v):
+        return "<hook>"
+    if name and _SENSITIVE_NAME_RE.search(name):
+        return "<redacted>"
+    if isinstance(v, tuple):
+        return list(v)
+    if isinstance(v, dict):
+        return {k: _jsonable(x, name=str(k)) for k, x in v.items()}
+    return v
+
+
+def _pre_ops(path: str) -> list[str]:
+    """The pre_<op> hook names an ops.py defines (cheap source scan, same approach as
+    discovery._module_defines — no import)."""
+    try:
+        with open(path) as fh:
+            return sorted(set(_PRE_OP_RE.findall(fh.read())))
+    except OSError:
+        return []
+
+
+def _custom_counts(recipe: str, repo_local: str | None) -> dict[str, dict[str, int]]:
+    out: dict[str, dict[str, int]] = {}
+    for source, path in discovery.custom_tests(recipe, repo_local):
+        sub = os.path.basename(os.path.dirname(path))  # functional | playwright
+        out.setdefault(source, {}).setdefault(sub, 0)
+        out[source][sub] += 1
+    return out
+
+
+def build(recipe: str, meta, repo_local: str | None) -> dict:
+    """Collect the run's resolved customization into one deterministic, JSON-serializable dict.
+
+    Keys: meta_non_default (explicitly-customized recipe_meta keys), hooks (ops.py pre-ops +
+    install_steps.sh + compose.ccci.yml with their source), overlays (lifecycle overlay files by
+    op + source), custom_tests (counts per source/subdir), env_overrides (active
+    CCCI_SKIP_GENERIC* — the dev-only escape hatch, flagged when riding a CI run)."""
+    hooks: dict = {}
+    pre_ops: dict[str, list[str]] = {}
+    for source, d in (
+        ("cc-ci", discovery.cc_ci_dir(recipe)),
+        ("repo-local", discovery._gated(recipe, repo_local)),  # noqa: SLF001 — same HC2 gate
+    ):
+        if not d:
+            continue
+        p = os.path.join(d, "ops.py")
+        if os.path.isfile(p):
+            ops = _pre_ops(p)
+            if ops:
+                pre_ops[source] = ops
+    if pre_ops:
+        hooks["ops.py"] = pre_ops
+    ist = discovery.install_steps(recipe, repo_local)
+    if ist:
+        hooks["install_steps.sh"] = ist[0]
+    if lifecycle.has_ccci_overlay(recipe):
+        hooks["compose.ccci.yml"] = "cc-ci"
+
+    overlays = {}
+    for op in discovery.LIFECYCLE_OPS:
+        ov = discovery.resolve_overlay_op(recipe, op, repo_local)
+        if ov:
+            overlays[op] = ov[0]
+
+    env_overrides = sorted(
+        k
+        for k in os.environ
+        if k.startswith("CCCI_SKIP_GENERIC")
+        and str(os.environ.get(k) or "").strip().lower() in ("1", "true", "yes", "on")
+    )
+
+    return {
+        "meta_non_default": {
+            k: _jsonable(v, name=k) for k, v in sorted(meta_mod.non_default(meta).items())
+        },
+        "hooks": hooks,
+        "overlays": overlays,
+        "custom_tests": _custom_counts(recipe, repo_local),
+        "env_overrides": env_overrides,
+    }
+
+
+def render(recipe: str, manifest: dict) -> str:
+    """The human block printed at run start (same content as the results.json key)."""
+    lines = [f"===== customization manifest: {recipe} ====="]
+    nd = manifest["meta_non_default"]
+    lines.append(
+        "meta (non-default): "
+        + (" ".join(f"{k}={v!r}" for k, v in nd.items()) if nd else "(none — zero-config floor)")
+    )
+    hk = manifest["hooks"]
+    parts = []
+    for source, ops in hk.get("ops.py", {}).items():
+        parts.append(f"ops.py[{','.join(ops)}]({source})")
+    if "install_steps.sh" in hk:
+        parts.append(f"install_steps.sh({hk['install_steps.sh']})")
+    if "compose.ccci.yml" in hk:
+        parts.append(f"compose.ccci.yml({hk['compose.ccci.yml']})")
+    lines.append("hooks: " + (" ".join(parts) if parts else "(none)"))
+    ov = manifest["overlays"]
+    lines.append(
+        "overlays: "
+        + (" ".join(f"test_{op}.py({src})" for op, src in ov.items()) if ov else "(none)")
+    )
+    ct = manifest["custom_tests"]
+    lines.append(
+        "custom tests: "
+        + (
+            " ".join(
+                " ".join(f"{sub}/={n}" for sub, n in sorted(counts.items())) + f" ({source})"
+                for source, counts in sorted(ct.items())
+            )
+            if ct
+            else "(none)"
+        )
+    )
+    eo = manifest["env_overrides"]
+    if eo:
+        suffix = "   !! dev-only override active in CI" if os.environ.get("DRONE") else ""
+        lines.append("env overrides: " + " ".join(f"{k}=1" for k in eo) + suffix)
+    else:
+        lines.append("env overrides: (none)")
+    return "\n".join(lines)
--- a/runner/harness/meta.py
+++ b/runner/harness/meta.py
@ -0,0 +1,320 @@
+"""Single recipe-meta loader + declarative key registry (recipe-custom restructure P1; spec
+docs/recipe-customization.md §8 R1).
+
+THE one place `tests/<recipe>/recipe_meta.py` is `exec()`d. Every consumer (orchestrator, pytest
+`meta` fixture, deploy env shaping, deps, warm-canonical enrollment, screenshot) reads the ONE
+loaded `RecipeMeta` object instead of re-exec'ing the file and cherry-picking keys — that drift
+(six divergent loaders, spec §4 L1–L6) is what made `SCREENSHOT` an unreachable knob (R2) and let
+key typos silently disable coverage (R6).
+
+Validation (locked decision, recipe-custom-restructure-full-plan.md):
+- unknown ALL-CAPS top-level name → MetaError (hard error, fails fast at load; the all-recipes
+  unit test catches it at PR time). Underscore-prefixed names (`_FOO`) are recipe-private and
+  exempt; lowercase names (helper functions/imports) are ignored.
+- type mismatch → MetaError. Callables are accepted ONLY for hook-typed keys.
+
+The KEYS registry is the single source of truth for the key set: it drives validation, the
+RecipeMeta dataclass fields, and the generated reference table in docs/recipe-customization.md §4
+(scripts/gen-meta-docs.py; a unit test asserts the committed table matches).
+"""
+
+from __future__ import annotations
+
+import copy
+import dataclasses
+import difflib
+import inspect
+import json
+import os
+from collections.abc import Callable
+
+ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
+TESTS_DIR = os.path.join(ROOT, "tests")
+
+
+class MetaError(Exception):
+    """A recipe_meta.py failed registry validation (unknown key / type mismatch / callable on a
+    data key). Hard error by design: a typo'd key must fail the run at load, not silently reduce
+    coverage (spec §8 R6 — the worst failure mode for a CI harness)."""
+
+
+@dataclasses.dataclass(frozen=True)
+class Key:
+    """One registered recipe_meta key: name, type tag, default, one-line doc (rendered into the
+    generated reference table), optional extra validator, and a deprecation marker (deprecated
+    keys still load+validate but are scheduled for deletion)."""
+
+    name: str
+    type: str  # "int"|"str"|"tuple[int]"|"bool"|"dict_or_hook"|"hook"|"list[str]"|"dict"
+    default: object
+    doc: str
+    validate: Callable[[object], None] | None = None
+    deprecated: bool = False
+    # Expected positional-parameter names for a callable value (rcust P3 uniform ctx convention).
+    # Enforced at load so a legacy-signature hook (e.g. `def READY_PROBE(domain)`) fails with a
+    # CLEAR MetaError naming the migration — never a silent TypeError mid-run.
+    hook_params: tuple[str, ...] | None = None
+
+
+KEYS: tuple[Key, ...] = (
+    Key(
+        "HEALTH_PATH",
+        "str",
+        "/",
+        "Path probed for serving/health checks (deploy wait + generic `assert_serving`).",
+    ),
+    Key("HEALTH_OK", "tuple[int]", (200, 301, 302), "Acceptable HTTP status codes for health."),
+    Key("DEPLOY_TIMEOUT", "int", 600, "Max seconds to wait for swarm convergence per deploy."),
+    Key("HTTP_TIMEOUT", "int", 300, "Max seconds to wait for HTTP health after convergence."),
+    Key(
+        "BACKUP_CAPABLE",
+        "bool",
+        None,
+        "Override the backup-tier capability auto-detect (compose `backupbot.backup` labels). `False` forces N/A; `True` forces the tier on; unset = auto-detect.",
+    ),
+    Key(
+        "EXPECTED_NA",
+        "dict",
+        None,
+        "Declare an N/A rung intentional: `{rung: reason}`. The cap stands either way; only the report wording changes.",
+    ),
+    Key(
+        "READY_PROBE",
+        "hook",
+        None,
+        "Callable `(ctx) -> [probe, ...]` returning extra readiness probes, run after install AND after upgrade: HTTP `{host, path, ok}` or TCP `{tcp_host, tcp_port, stable}`.",
+        hook_params=("ctx",),
+    ),
+    Key(
+        "UPGRADE_BASE_VERSION",
+        "str",
+        None,
+        "Exact published tag overriding the upgrade tier's base (default: `recipe_versions[-2]`).",
+    ),
+    Key(
+        "BACKUP_VERIFY",
+        "hook",
+        None,
+        "Callable `(ctx) -> bool` post-backup data-capture check; `False` re-runs the backup (truncated-dump race guard), retried up to 3 attempts.",
+        hook_params=("ctx",),
+    ),
+    Key(
+        "UPGRADE_EXTRA_ENV",
+        "dict_or_hook",
+        None,
+        "Extra `.env` keys applied after the PR-head checkout, before the chaos redeploy (env that exists only at head). Dict, or callable `(ctx) -> dict`.",
+        hook_params=("ctx",),
+    ),
+    Key(
+        "EXTRA_ENV",
+        "dict_or_hook",
+        {},
+        "Extra `.env` keys applied at EVERY deploy (base install AND upgrade old-app). Dict, or callable `(ctx) -> dict` deriving values from the per-run domain (`ctx.domain`).",
+        hook_params=("ctx",),
+    ),
+    Key(
+        "DEPS",
+        "list[str]",
+        [],
+        'Dep recipes deployed/provisioned alongside (e.g. `["keycloak"]`); creds land in `$CCCI_DEPS_FILE`.',
+    ),
+    Key(
+        "WARM_CANONICAL",
+        "bool",
+        False,
+        "Enroll the recipe in the warm/canonical app system (docs/warm.md): green cold runs on LATEST advance the canonical snapshot.",
+    ),
+    Key(
+        "SCREENSHOT",
+        "hook",
+        None,
+        "Callable `(page, ctx)` driving Playwright to a safe, credential-free post-login view for the results-card screenshot (default: landing page).",
+        hook_params=("page", "ctx"),
+    ),
+    # (CHAOS_BASE_DEPLOY, OIDC_AT_INSTALL and SKIP_GENERIC were deleted in restructure P2:
+    # compose.ccci.yml is first-class + auto-chaos; install-time deps wiring is the only mode;
+    # the generic floor is suppressible only via the dev-only CCCI_SKIP_GENERIC* env form.)
+)
+
+_REGISTRY: dict[str, Key] = {k.name: k for k in KEYS}
+
+# The one validated, attribute-access view of a recipe's customization. Generated from KEYS so the
+# field set can never drift from the registry (frozen: consumers share one immutable object).
+RecipeMeta = dataclasses.make_dataclass(
+    "RecipeMeta",
+    [(k.name, object, dataclasses.field(default=None)) for k in KEYS],
+    frozen=True,
+)
+RecipeMeta.__doc__ = (
+    "Validated per-recipe customization (one field per registered key; attribute access). "
+    "Built ONLY by meta.load()."
+)
+
+
+def meta_path(recipe: str, tests_dir: str | None = None) -> str:
+    """Canonical path of a recipe's meta file (pure)."""
+    return os.path.join(tests_dir or TESTS_DIR, recipe, "recipe_meta.py")
+
+
+def check_hook_signature(fn, expected: tuple[str, ...], where: str) -> None:
+    """Enforce the uniform ctx hook convention (rcust P3): a hook callable's positional parameters
+    must be exactly `expected` (e.g. ("ctx",) or ("page", "ctx")). A legacy-signature hook (the
+    pre-restructure `(domain)` / `(domain, meta)` / `(page, domain, meta)` forms) raises a CLEAR
+    MetaError naming the migration — never a silent TypeError mid-run."""
+    try:
+        params = [
+            p.name
+            for p in inspect.signature(fn).parameters.values()
+            if p.kind in (p.POSITIONAL_ONLY, p.POSITIONAL_OR_KEYWORD)
+        ]
+    except (TypeError, ValueError):  # builtins/odd callables — let the call site surface it
+        return
+    if tuple(params) != expected:
+        raise MetaError(
+            f"{where}: hook signature is ({', '.join(params)}) — the recipe-customization "
+            f"restructure (P3) changed ALL recipe hook signatures to ({', '.join(expected)}); "
+            f"read fields off the HookCtx (ctx.domain, ctx.base_url, ctx.meta, ctx.deps, ctx.op). "
+            f"See docs/recipe-customization.md §5."
+        )
+
+
+def _coerce(key: Key, value: object, path: str) -> object:
+    """Validate `value` against `key`'s declared type; normalize containers (tuple[int]/list[str]).
+    Raises MetaError on mismatch — including a callable supplied for a data-typed key."""
+    t = key.type
+    if callable(value) and t not in ("hook", "dict_or_hook"):
+        raise MetaError(
+            f"{path}: {key.name} is a data key (type {t}) — callables are accepted only for "
+            f"hook-typed keys"
+        )
+    if t == "int":
+        if isinstance(value, int) and not isinstance(value, bool):
+            return value
+    elif t == "str":
+        if isinstance(value, str):
+            return value
+    elif t == "bool":
+        if isinstance(value, bool):
+            return value
+    elif t == "tuple[int]":
+        if isinstance(value, tuple | list) and all(
+            isinstance(x, int) and not isinstance(x, bool) for x in value
+        ):
+            return tuple(value)
+    elif t == "list[str]":
+        if isinstance(value, tuple | list) and all(isinstance(x, str) for x in value):
+            return list(value)
+    elif t == "dict":
+        if isinstance(value, dict):
+            return value
+    elif (
+        t == "hook"
+        and callable(value)
+        or t == "dict_or_hook"
+        and (isinstance(value, dict) or callable(value))
+    ):
+        return value
+    raise MetaError(f"{path}: {key.name} must be {t}, got {type(value).__name__} ({value!r})")
+
+
+def load(recipe: str, tests_dir: str | None = None):
+    """Load + validate a recipe's customization -> RecipeMeta. THE only exec() of recipe_meta.py.
+
+    Missing file -> all registry defaults (the zero-config baseline, spec §2). Unknown
+    non-underscore ALL-CAPS top-level name or type mismatch -> MetaError (hard error).
+    `tests_dir` overrides the recipe-meta root (unit tests / fixtures)."""
+    path = meta_path(recipe, tests_dir)
+    values = {k.name: copy.copy(k.default) for k in KEYS}
+    if os.path.exists(path):
+        ns: dict = {}
+        with open(path) as fh:
+            exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
+        for name in sorted(ns):
+            if name.startswith("_") or not name.isupper():
+                continue  # _FOO = recipe-private (exempt); lowercase = helpers/imports (ignored)
+            key = _REGISTRY.get(name)
+            if key is None:
+                near = difflib.get_close_matches(name, _REGISTRY, n=1)
+                hint = f" — did you mean {near[0]!r}?" if near else ""
+                raise MetaError(
+                    f"{path}: unknown recipe_meta key {name!r}{hint}. Registered keys: "
+                    f"{', '.join(sorted(_REGISTRY))}. Recipe-private constants must be "
+                    f"underscore-prefixed (e.g. _{name})."
+                )
+            values[name] = _coerce(key, ns[name], path)
+            if key.hook_params and callable(values[name]):
+                check_hook_signature(values[name], key.hook_params, f"{path}: {name}")
+            if key.validate:
+                key.validate(values[name])
+    return RecipeMeta(**values)
+
+
+def as_dict(meta) -> dict:
+    """RecipeMeta -> {key: value} (every registered key, defaults included)."""
+    return dataclasses.asdict(meta)
+
+
+def non_default(meta) -> dict:
+    """The keys a recipe explicitly customized: {key: value} where value differs from the registry
+    default. Hooks compare by identity-vs-None (a set hook is always non-default). Feeds the run's
+    customization manifest (P5)."""
+    out = {}
+    for k in KEYS:
+        v = getattr(meta, k.name)
+        if v != k.default:
+            out[k.name] = v
+    return out
+
+
+@dataclasses.dataclass(frozen=True)
+class HookCtx:
+    """The single argument every recipe hook receives (rcust P3 uniform ctx convention):
+    `EXTRA_ENV(ctx)`, `UPGRADE_EXTRA_ENV(ctx)`, `READY_PROBE(ctx)`, `BACKUP_VERIFY(ctx)`,
+    `SCREENSHOT(page, ctx)`, ops.py `pre_<op>(ctx)`."""
+
+    domain: str  # the app's per-run domain
+    base_url: str  # https://<domain>
+    meta: object  # the recipe's full RecipeMeta
+    deps: dict | None  # provisioned dep creds ({dep_recipe: entry}) or None if absent/empty
+    op: str | None  # current lifecycle op (install|upgrade|backup|restore) or None
+
+
+def _run_deps() -> dict | None:
+    """The current run's provisioned dep creds from $CCCI_DEPS_FILE (either shape), or None.
+    Read directly (not via harness.deps) to keep meta.py import-cycle-free."""
+    path = os.environ.get("CCCI_DEPS_FILE")
+    if not path or not os.path.exists(path):
+        return None
+    try:
+        with open(path) as f:
+            data = json.load(f)
+    except (OSError, ValueError):
+        return None
+    if isinstance(data, dict):
+        return data or None
+    if isinstance(data, list):
+        out = {e["recipe"]: e for e in data if isinstance(e, dict) and e.get("recipe")}
+        return out or None
+    return None
+
+
+def hook_ctx(domain: str, meta, *, op: str | None = None) -> HookCtx:
+    """Build the HookCtx for a hook call site. Dep creds are picked up from the run's
+    $CCCI_DEPS_FILE when present (None otherwise)."""
+    return HookCtx(domain=domain, base_url=f"https://{domain}", meta=meta, deps=_run_deps(), op=op)
+
+
+def _env_map(value, ctx: HookCtx) -> dict[str, str]:
+    if callable(value):
+        value = value(ctx)
+    return {str(k): str(v) for k, v in (value or {}).items()}
+
+
+def extra_env(meta, ctx: HookCtx) -> dict[str, str]:
+    """Resolve EXTRA_ENV (dict or callable(ctx)->dict) to the concrete per-run env map."""
+    return _env_map(meta.EXTRA_ENV, ctx)
+
+
+def upgrade_extra_env(meta, ctx: HookCtx) -> dict[str, str]:
+    """Resolve UPGRADE_EXTRA_ENV (dict or callable(ctx)->dict) to the concrete env map."""
+    return _env_map(meta.UPGRADE_EXTRA_ENV, ctx)
--- a/runner/harness/results.py
+++ b/runner/harness/results.py
@ -203,6 +203,7 @@ def build_results(
    screenshot: str | None = None,
    summary_card: str | None = None,
    expected_na: dict | None = None,
+    customization: dict | None = None,
 ) -> dict:
    """Assemble the full results.json dict (no I/O). `finished_ts` is passed in (the orchestrator
    stamps it) so this stays pure and deterministic for unit tests. `expected_na` is the recipe's
@ -236,6 +237,9 @@ def build_results(
        },
        "screenshot": screenshot,
        "summary_card": summary_card,
+        # rcust P5: the run's resolved customization manifest (pure presentation — consumers must
+        # never derive a verdict from it).
+        "customization": customization,
    }


--- a/runner/harness/screenshot.py
+++ b/runner/harness/screenshot.py
@ -8,7 +8,7 @@ Secret-safety (R7, the cardinal screenshot guardrail): the screenshot step must
 that displays generated credentials (an install wizard showing the initial admin password, a secrets
 page, etc.). The DEFAULT capture is the app's **landing page** (a login form shows fields, not the
 password) — safe for every recipe. A recipe that needs a post-login view opts in via a recipe-meta
-`SCREENSHOT` hook: a callable `screenshot(page, domain, meta) -> None` that drives Playwright to a
+`SCREENSHOT` hook: a callable `SCREENSHOT(page, ctx) -> None` that drives Playwright to a
 safe, credential-free view and is responsible for not landing on a secrets page. The harness never
 auto-fills a wizard.

@ -21,6 +21,7 @@ from __future__ import annotations
 import os

 from . import browser as harness_browser
+from . import meta as meta_mod

 # Default viewport for the captured screenshot — a desktop-ish frame that crops well into the card.
 VIEWPORT = {"width": 1280, "height": 800}
@ -33,12 +34,19 @@ def screenshot_path(run_artifact_dir: str) -> str:
    return os.path.join(run_artifact_dir, "screenshot.png")


-def _load_screenshot_hook(recipe_meta: dict | None):
+def _load_screenshot_hook(recipe_meta):
    """Return the recipe's optional SCREENSHOT hook (a callable) if it declared one, else None.
-    The hook drives Playwright to a safe post-login view; default is the landing page."""
-    if not recipe_meta:
+    The hook drives Playwright to a safe post-login view; default is the landing page.
+
+    `recipe_meta` is the loaded RecipeMeta (rcust P1 — the single loader actually delivers
+    SCREENSHOT now; under the old L1 allowlist the key never arrived, spec §8 R2). A plain dict
+    is still accepted for direct/manual callers."""
+    if recipe_meta is None:
        return None
-    hook = recipe_meta.get("SCREENSHOT")
+    if isinstance(recipe_meta, dict):
+        hook = recipe_meta.get("SCREENSHOT")
+    else:
+        hook = getattr(recipe_meta, "SCREENSHOT", None)
    return hook if callable(hook) else None


@ -67,8 +75,9 @@ def capture(domain: str, out_path: str, *, recipe_meta: dict | None = None) -> s
                if hook is not None:
                    # Recipe-specific safe view (post-login etc.). The hook owns navigation +
                    # the no-secret-page guarantee; it should call page.screenshot itself, but if
-                    # it doesn't, we still snap the resulting page below.
-                    hook(page, domain, recipe_meta)
+                    # it doesn't, we still snap the resulting page below. SCREENSHOT(page, ctx) —
+                    # the uniform ctx convention (rcust P3).
+                    hook(page, meta_mod.hook_ctx(domain, recipe_meta))
                    if not os.path.exists(out_path):
                        page.screenshot(path=out_path, full_page=False)
                else:
--- a/runner/harness/warmsnap.py
+++ b/runner/harness/warmsnap.py
@ -113,7 +113,9 @@ def _assert_undeployed(domain: str) -> None:
        )


-def snapshot(recipe: str, domain: str, commit: str | None = None, version: str | None = None) -> dict:
+def snapshot(
+    recipe: str, domain: str, commit: str | None = None, version: str | None = None
+) -> dict:
    """Take a last-known-good snapshot of every data volume of <domain>'s stack. The app MUST be
    undeployed. Atomically replaces the prior last-good. Returns the written meta dict."""
    _assert_undeployed(domain)
@ -169,7 +171,9 @@ def restore(recipe: str, domain: str) -> dict:
    for vol in meta.get("volumes", []):
        tar_path = os.path.join(volumes_dir(recipe), f"{vol}.tar")
        if vol not in current:
-            raise SnapshotError(f"snapshot volume {vol} absent from current stack {sorted(current)}")
+            raise SnapshotError(
+                f"snapshot volume {vol} absent from current stack {sorted(current)}"
+            )
        mp = _volume_mountpoint(vol)
        # Clear the volume contents (incl. dotfiles) without removing the mountpoint itself.
        r = _run(["sh", "-c", f'rm -rf -- "{mp}"/* "{mp}"/.[!.]* "{mp}"/..?* 2>/dev/null; true'])
--- a/runner/nightly_sweep.py
+++ b/runner/nightly_sweep.py
@ -60,14 +60,17 @@ def sweep() -> int:
    for r in recipes:
        print(f"\n===== nightly: full-cold {r} (latest) =====", flush=True)
        env = dict(os.environ, RECIPE=r)
-        env.pop("REF", None)      # latest, not a PR head
+        env.pop("REF", None)  # latest, not a PR head
        env.pop("CCCI_QUICK", None)
        env.pop("MODE", None)
        rc = subprocess.run(
            [sys.executable, os.path.join(_here(), "run_recipe_ci.py")], env=env
        ).returncode
        results[r] = rc
-        print(f"nightly: {r} rc={rc} ({'green→canonical refreshed' if rc == 0 else 'red'})", flush=True)
+        print(
+            f"nightly: {r} rc={rc} ({'green→canonical refreshed' if rc == 0 else 'red'})",
+            flush=True,
+        )
    # WC8 disk hygiene: drop warm data for de-enrolled canonicals; log the disk budget.
    pruned = canonical.prune_stale()
    if pruned:
--- a/runner/run_recipe_ci.py
+++ b/runner/run_recipe_ci.py
@ -44,24 +44,39 @@ sys.path.insert(0, os.path.join(ROOT, "runner"))
 from harness import (  # noqa: E402
    abra,
    canonical,
-    card as card_mod,
-    deps as deps_mod,
    discovery,
    generic,
    lifecycle,
+    lifetime,
    naming,
-    results as results_mod,
-    screenshot as screenshot_mod,
    warm,
    warmsnap,
 )
+from harness import (  # noqa: E402
+    card as card_mod,
+)
+from harness import (  # noqa: E402
+    deps as deps_mod,
+)
+from harness import (  # noqa: E402
+    manifest as manifest_mod,
+)
+from harness import (  # noqa: E402
+    meta as meta_mod,
+)
+from harness import (  # noqa: E402
+    results as results_mod,
+)
+from harness import (  # noqa: E402
+    screenshot as screenshot_mod,
+)

 ALL_STAGES = ("install", "upgrade", "backup", "restore", "custom")


 def sso_dep_unverified(declared, deps_ready: bool, requires_deps_skipped: int) -> bool:
    """F2-11 gate predicate (pure, unit-tested). True when a recipe declares DEPS but its
-    setup_custom_tests failed (deps not ready) AND that caused ≥1 `requires_deps` (SSO/OIDC) test
+    dep provisioning failed (deps not ready) AND that caused ≥1 `requires_deps` (SSO/OIDC) test
    to SKIP. In that case the recipe's characteristic SSO claim was NOT verified, so the run must
    NOT report GREEN — even though a skip-only pytest file exits 0 and leaves every tier 'pass'.
    Generic-tier failure-isolation is preserved (those results stand); only the green SIGNAL is
@ -129,18 +144,73 @@ def _gitea_token() -> str | None:
    return tok or None


+def _run_state_path(name: str) -> str:
+    """Run-scoped state file in the tempdir, keyed by run id + harness pid — NEVER by app domain.
+    A second run of the SAME domain overlaps this process (its main() preamble executes before it
+    blocks at the app lock inside deploy_app), so domain-keyed files get reset/removed under the
+    live run: M2(c) double-!testme produced a false DG4.1 deploy-count=2 in run 1 and a countfile
+    FileNotFoundError crash in run 2. Children never re-derive these paths — they receive them
+    via the CCCI_*_FILE env vars, so the key only has to be unique per harness process."""
+    rid = results_mod.run_id()
+    return os.path.join(tempfile.gettempdir(), f"ccci-{name}-{rid}-{os.getpid()}")
+
+
+def setup_run_abra_dir() -> str:
+    """P3: build + export this run's PER-RUN ABRA_DIR — structural isolation of recipe trees.
+
+    `<runs_dir>/<run-id>/abra/` with:
+      servers/   -> symlink to the canonical ~/.abra/servers. App .env files land in the shared
+                    canonical path, so janitor discovery (`abra app ls`) and env-based teardown
+                    work unchanged from any process; per-domain filenames + the app-domain lock
+                    prevent write conflicts.
+      catalogue/ -> symlink to the canonical ~/.abra/catalogue (read-mostly).
+      recipes/   fresh + empty — THE isolation that matters: each run clones and git-checkouts
+                 its own recipe trees, so concurrent runs (same recipe included) can never
+                 corrupt each other's deploy tree. Replaces the per-recipe flock.
+    Exported as $ABRA_DIR — honored by the abra CLI and by every harness path helper
+    (abra.abra_dir()) — BEFORE any abra call. Rides along the existing run-dir retention."""
+    canonical = os.path.expanduser("~/.abra")
+    rid = results_mod.run_id()
+    if rid == "manual":
+        rid = f"manual-{os.getpid()}"  # two concurrent hand-runs must not share a tree
+    run_abra_dir = os.path.join(results_mod.runs_dir(), rid, "abra")
+    os.makedirs(os.path.join(run_abra_dir, "recipes"), exist_ok=True)
+    for shared in ("servers", "catalogue"):
+        link = os.path.join(run_abra_dir, shared)
+        if not os.path.islink(link):
+            os.symlink(os.path.join(canonical, shared), link)
+    os.environ["ABRA_DIR"] = run_abra_dir
+    print(
+        f"== per-run ABRA_DIR: {run_abra_dir} (servers/catalogue -> canonical; fresh recipes/) ==",
+        flush=True,
+    )
+    return run_abra_dir
+
+
 def fetch_recipe(recipe: str, ref: str | None, src: str | None) -> None:
-    """Make the recipe available at the code under test. If SRC+REF point at the mirror PR,
+    """Make the recipe available at the code under test in THIS RUN's recipe tree
+    ($ABRA_DIR/recipes/<recipe>): a plain clone — no locking needed, no rm-rf of any shared
+    state (the rm below only clears this run's own leftovers, e.g. a janitor-triggered
+    `abra app ls` auto-clone or a Drone build-number reuse). If SRC+REF point at the mirror PR,
    clone it at that ref; otherwise fetch the catalogue copy. Private mirror repos need the bot
    token — passed via a per-command http.extraHeader (not persisted in .git/config, not printed)."""
-    recipes_dir = os.path.expanduser("~/.abra/recipes")
-    os.makedirs(recipes_dir, exist_ok=True)
-    dest = os.path.join(recipes_dir, recipe)
-    # CCCI_SKIP_FETCH=1: use the local recipe clone as-is (lets a test/Adversary stage a fake/broken
-    # ref — e.g. a simulated broken PR head for the --quick rollback proof — without it being clobbered
-    # by a re-fetch). Never set in production CI.
+    dest = abra.recipe_dir(recipe)
+    os.makedirs(os.path.dirname(dest), exist_ok=True)
+    # CCCI_SKIP_FETCH=1: use the locally STAGED recipe clone as-is (lets a test/Adversary stage a
+    # fake/broken ref — e.g. a simulated broken PR head for the --quick rollback proof — without it
+    # being clobbered by a re-fetch). Staging happens in the canonical ~/.abra/recipes/<recipe>;
+    # copy it into the per-run tree so the rest of the run reads the staged state. Never set in
+    # production CI.
    if os.environ.get("CCCI_SKIP_FETCH") == "1":
-        print(f"[fetch] CCCI_SKIP_FETCH=1 — using local {recipe} recipe clone as-is", flush=True)
+        canonical = os.path.expanduser(f"~/.abra/recipes/{recipe}")
+        subprocess.run(["rm", "-rf", dest], check=False)
+        if os.path.isdir(canonical):
+            shutil.copytree(canonical, dest, symlinks=True)
+        print(
+            f"[fetch] CCCI_SKIP_FETCH=1 — using staged {recipe} clone as-is "
+            f"(copied {canonical} -> per-run tree)",
+            flush=True,
+        )
        return
    if src and ref:
        url = f"https://git.autonomic.zone/{src}.git"
@ -169,7 +239,7 @@ def fetch_recipe(recipe: str, ref: str | None, src: str | None) -> None:
 def snapshot_recipe_tests(recipe: str) -> str | None:
    """Copy the recipe-shipped tests/ to a stable temp dir, immune to abra re-checking-out the
    recipe to a version tag during the run. Returns the snapshot path, or None if no tests/."""
-    src = os.path.expanduser(f"~/.abra/recipes/{recipe}/tests")
+    src = os.path.join(abra.recipe_dir(recipe), "tests")
    if not os.path.isdir(src):
        return None
    has_overlay = glob.glob(os.path.join(src, "test_*.py")) or os.path.isfile(
@ -183,52 +253,29 @@ def snapshot_recipe_tests(recipe: str) -> str | None:
    return dst


-def _load_meta(recipe: str) -> dict:
-    """Mirror tests/conftest._recipe_meta so the orchestrator's deploy/wait uses the same per-recipe
-    config the tiers see (timeouts, health path/codes)."""
-    meta = {
-        "HEALTH_PATH": "/",
-        "HEALTH_OK": (200, 301, 302),
-        "DEPLOY_TIMEOUT": 600,
-        "HTTP_TIMEOUT": 300,
-    }
-    path = os.path.join(ROOT, "tests", recipe, "recipe_meta.py")
-    if os.path.exists(path):
-        ns: dict = {}
-        with open(path) as fh:
-            exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
-        for k in list(meta) + [
-            "BACKUP_CAPABLE",
-            "SKIP_GENERIC",
-            "EXPECTED_NA",
-            "OIDC_AT_INSTALL",
-            "READY_PROBE",
-            "UPGRADE_BASE_VERSION",
-            "BACKUP_VERIFY",
-            "UPGRADE_EXTRA_ENV",
-        ]:
-            if k in ns:
-                meta[k] = ns[k]
-    return meta
-
-
 def _tier_env(domain: str) -> dict:
    return dict(os.environ, CCCI_APP_DOMAIN=domain, CCCI_BASE_URL=f"https://{domain}")


-def _skip_generic(op: str, meta: dict) -> bool:
+def skip_generic_env_overrides() -> list[str]:
+    """Active CCCI_SKIP_GENERIC* env overrides (rcust P2c: the meta key is deleted; the env form
+    is a documented LOCAL-DEV-ONLY escape hatch). Surfaced loudly when set in a CI (drone) run —
+    it reduces generic-floor coverage and must never silently ride a CI verdict."""
+    return sorted(
+        k for k in os.environ if k.startswith("CCCI_SKIP_GENERIC") and _truthy(os.environ.get(k))
+    )
+
+
+def _skip_generic(op: str) -> bool:
    """Whether the generic assertion for `op` is opted out (Phase 1e HC3). Default: run (additive).
-    Opt-out, any of: env CCCI_SKIP_GENERIC (all ops), env CCCI_SKIP_GENERIC_<OP>, or the recipe's
-    declarative recipe_meta.SKIP_GENERIC list (op name, or "all"/"*")."""
+    Opt-out via env only (dev-only escape hatch, P2c): CCCI_SKIP_GENERIC (all ops) or
+    CCCI_SKIP_GENERIC_<OP>. The recipe_meta SKIP_GENERIC key is deleted (zero users)."""
    if _truthy(os.environ.get("CCCI_SKIP_GENERIC")):
        return True
-    if _truthy(os.environ.get(f"CCCI_SKIP_GENERIC_{op.upper()}")):
-        return True
-    sg = [str(s).lower() for s in (meta.get("SKIP_GENERIC") or [])]
-    return "all" in sg or "*" in sg or op in sg
+    return _truthy(os.environ.get(f"CCCI_SKIP_GENERIC_{op.upper()}"))


-def _run_pre_hook(recipe: str, op: str, repo_local: str | None, domain: str, meta: dict) -> None:
+def _run_pre_hook(recipe: str, op: str, repo_local: str | None, domain: str, meta) -> None:
    """Run the optional pre-op seed hook (recipe ops.py `pre_<op>`) BEFORE the harness performs the
    op (HC3 op/assertion split): overlays seed data-continuity markers / the backup→restore mutation
    here, then assert post-op in test_<op>.py. cc-ci's ops.py is trusted; a repo-local ops.py is
@ -245,7 +292,11 @@ def _run_pre_hook(recipe: str, op: str, repo_local: str | None, domain: str, met
        mod = importlib.util.module_from_spec(spec)
        spec.loader.exec_module(mod)
        print(f"  pre-op seed ({source}): {os.path.relpath(path, ROOT)}::pre_{op}", flush=True)
-        getattr(mod, f"pre_{op}")(domain, meta)
+        fn = getattr(mod, f"pre_{op}")
+        # Uniform ctx convention (rcust P3): pre_<op>(ctx). A legacy (domain, meta) hook fails
+        # HERE with a clear migration message, not a TypeError mid-call.
+        meta_mod.check_hook_signature(fn, ("ctx",), f"{os.path.relpath(path, ROOT)}::pre_{op}")
+        fn(meta_mod.hook_ctx(domain, meta, op=op))
    finally:
        if d in sys.path:
            sys.path.remove(d)
@ -258,7 +309,7 @@ def _perform_op(
    head_ref: str | None,
    op_state: dict,
    deploy_timeout: int = 900,
-    meta: dict | None = None,
+    meta=None,
 ) -> None:
    """Perform the single mutating op ONCE (the harness owns the op, HC3). install has no op. Records
    what the assertions need (pre-upgrade identity, backup snapshot_id) into op_state. None of these
@ -281,9 +332,10 @@ def _perform_op(
        # verify fails we re-run the WHOLE backup (fresh restic snapshot) with a re-stabilised DB, up to
        # 3 attempts. Recipes without BACKUP_VERIFY are unaffected (single backup, as before).
        snap = generic.perform_backup(domain)
-        verify = meta.get("BACKUP_VERIFY") if meta else None
+        verify = meta.BACKUP_VERIFY if meta else None
+        verify_ctx = meta_mod.hook_ctx(domain, meta, op="backup") if meta else None
        attempt = 1
-        while callable(verify) and not verify(domain) and attempt < 3:
+        while callable(verify) and not verify(verify_ctx) and attempt < 3:
            attempt += 1
            print(
                f"  backup-verify FAILED (attempt {attempt - 1}/3) — backup did not capture the "
@ -291,7 +343,7 @@ def _perform_op(
                flush=True,
            )
            snap = generic.perform_backup(domain)
-        if callable(verify) and not verify(domain):
+        if callable(verify) and not verify(verify_ctx):
            print(
                f"  !! backup-verify still FAILED after {attempt} attempts — backup is incomplete",
                flush=True,
@ -307,7 +359,7 @@ def run_lifecycle_tier(
    op: str,
    repo_local: str | None,
    domain: str,
-    meta: dict,
+    meta,
    head_ref: str | None,
    op_state: dict,
    records: list[dict] | None = None,
@ -322,7 +374,7 @@ def run_lifecycle_tier(
    a {tier,source,file,rc,junit} record appended, so the run can assemble per-stage/per-test
    results.json + the level afterwards. Purely additive — does not change the verdict."""
    overlay = discovery.resolve_overlay_op(recipe, op, repo_local)
-    skip_gen = _skip_generic(op, meta)
+    skip_gen = _skip_generic(op)
    files: list[tuple[str, str]] = []
    if not skip_gen:
        files.append(discovery.generic_op(op))
@ -347,7 +399,7 @@ def run_lifecycle_tier(
            recipe,
            head_ref,
            op_state,
-            deploy_timeout=int(meta.get("DEPLOY_TIMEOUT", 900)),
+            deploy_timeout=int(meta.DEPLOY_TIMEOUT),
            meta=meta,
        )
        with open(os.environ["CCCI_OP_STATE_FILE"], "w") as f:
@ -385,7 +437,7 @@ def run_lifecycle_tier(
 def _enrich_deps_with_sso(parent_recipe: str, parent_domain: str, deps_list) -> dict[str, dict]:
    """For each dep, set up a fresh realm/client + test user via the harness's provider-specific
    setup function, then return a recipe→entry dict carrying domain + admin + realm/client/user
-    info — the shape the `setup_custom_tests.sh` hook (and dependent tests) read.
+    info — the shape the `install_steps.sh` hook (and dependent tests) read.

    Provider routing: today only `keycloak` is supported. authentik will need a parallel
    `setup_authentik_realm` when an authentik-dep recipe enrolls (DEFERRED.md #9).
@ -399,7 +451,7 @@ def _enrich_deps_with_sso(parent_recipe: str, parent_domain: str, deps_list) ->
        if not dep_recipe or not dep_domain:
            continue
        if dep_recipe != "keycloak":
-            # Provider not yet supported — record bare entry; setup_custom_tests.sh / tests will
+            # Provider not yet supported — record bare entry; install_steps.sh / tests will
            # raise if they need realm/client info they don't see.
            out[dep_recipe] = entry
            continue
@ -443,12 +495,10 @@ def _provision_deps(

    Splits deps into live-warm (shared provider at a stable domain + a per-run realm) vs cold
    (co-deployed per run), provisions each dep's SSO realm/client/user, and persists the enriched
-    dict the `setup_custom_tests.sh`/`install_steps.sh` hooks + dependent tests read. Raises on any
-    failure (the caller marks deps-not-ready). Used by BOTH wiring paths:
-    - post-deploy (legacy): provision AFTER generic tiers, then `setup_custom_tests.sh` does an
-      in-place OIDC redeploy.
-    - install-time (`OIDC_AT_INSTALL`, Q3.2a): provision BEFORE the single deploy so the
-      install-tier `install_steps.sh` hook wires OIDC env into that one deploy — no reconverge.
+    dict the `install_steps.sh` hooks + dependent tests read. Raises on any failure (the caller
+    marks deps-not-ready). Install-time wiring is the ONLY mode (rcust P2b): provision BEFORE the
+    single deploy so the install-tier `install_steps.sh` hook wires OIDC env into that one deploy —
+    no reconverge, no post-deploy `setup_custom_tests.sh` machinery.
    """
    warm_deps, cold_deps = [], []
    for d in declared:
@ -459,7 +509,7 @@ def _provision_deps(
            if wd:
                print(f"  dep: {d} warm provider {wd} not up — cold fallback", flush=True)
            cold_deps.append(d)
-    dep_metas = {d: _load_meta(d) for d in cold_deps}
+    dep_metas = {d: meta_mod.load(d) for d in cold_deps}
    deps_list = (
        deps_mod.deploy_deps(recipe, os.environ.get("PR", "0"), ref, cold_deps, meta_for=dep_metas)
        if cold_deps
@ -477,32 +527,6 @@ def _provision_deps(
    return deps_state


-def _run_setup_custom_tests_hook(recipe: str, domain: str, deps_file: str) -> None:
-    """Run `tests/<recipe>/setup_custom_tests.sh` if present (operator-2026-05-28 SSO-dep plan
-    §3.2). The hook reads `$CCCI_DEPS_FILE`, sets OIDC env via `abra app config set` + secret
-    insert, and triggers an in-place `abra app deploy --force --chaos`. Failure here propagates
-    to mark deps-not-ready (caught in main())."""
-    path = os.path.join(ROOT, "tests", recipe, "setup_custom_tests.sh")
-    if not os.path.isfile(path):
-        # No hook = recipe doesn't need post-deps wiring; deps are deployed + creds available
-        # via deps_apps fixture as-is.
-        print(
-            f"  setup_custom_tests: no hook at {os.path.relpath(path, ROOT)} (deps creds ready in $CCCI_DEPS_FILE)",
-            flush=True,
-        )
-        return
-    print(f"  setup_custom_tests hook: {os.path.relpath(path, ROOT)}", flush=True)
-    rc = subprocess.run(
-        ["bash", path],
-        check=False,
-        env=dict(os.environ, CCCI_APP_DOMAIN=domain, CCCI_RECIPE=recipe, CCCI_DEPS_FILE=deps_file),
-    )
-    if rc.returncode != 0:
-        raise RuntimeError(
-            f"setup_custom_tests.sh exited {rc.returncode} (deps env not wired into parent)"
-        )
-
-
 def run_custom(
    recipe: str,
    repo_local: str | None,
@ -545,7 +569,7 @@ def _wait_undeployed(domain: str, timeout: int = 120) -> None:


 def run_quick(
-    recipe: str, ref: str | None, head_ref: str | None, repo_local: str | None, meta: dict
+    recipe: str, ref: str | None, head_ref: str | None, repo_local: str | None, meta
 ) -> int:
    """WC4 `--quick` opt-in fast lane (plan §2). Reattach the data-warm canonical (known-good volume)
    → upgrade IN PLACE to the PR head (chaos) → assert generic UPGRADE (reconverge+moved+serving) +
@ -566,22 +590,22 @@ def run_quick(
        flush=True,
    )

-    statefile = os.path.join(tempfile.gettempdir(), f"ccci-opstate-{domain}.json")
+    statefile = _run_state_path("opstate") + ".json"
    with open(statefile, "w") as f:
        json.dump({}, f)
    os.environ["CCCI_OP_STATE_FILE"] = statefile
-    depsfile = os.path.join(tempfile.gettempdir(), f"ccci-deps-{domain}.json")
+    depsfile = _run_state_path("deps") + ".json"
    with open(depsfile, "w") as f:
        json.dump({}, f)
    os.environ["CCCI_DEPS_FILE"] = depsfile
-    skipfile = os.path.join(tempfile.gettempdir(), f"ccci-depskip-{domain}.txt")
+    skipfile = _run_state_path("depskip") + ".txt"
    with contextlib.suppress(OSError):
        os.remove(skipfile)
    os.environ["CCCI_DEPS_SKIP_REPORT"] = skipfile

    op_state: dict = {}
    results: dict[str, str] = {}
-    declared = deps_mod.declared_deps(recipe)
+    declared = list(meta.DEPS)
    deps_state: dict = {}
    deps_ready = True
    deps_not_ready_reason = ""
@ -593,28 +617,32 @@ def run_quick(
    try:
        # 1) reattach the canonical (warm boot at the known-good version + retained volume)
        try:
-            canonical.deploy_canonical(recipe, timeout=int(meta.get("DEPLOY_TIMEOUT", 900)))
+            canonical.deploy_canonical(recipe, timeout=int(meta.DEPLOY_TIMEOUT))
            lifecycle.wait_healthy(
                domain,
-                ok_codes=tuple(meta["HEALTH_OK"]),
-                path=meta["HEALTH_PATH"],
-                deploy_timeout=meta["DEPLOY_TIMEOUT"],
-                http_timeout=meta["HTTP_TIMEOUT"],
+                ok_codes=tuple(meta.HEALTH_OK),
+                path=meta.HEALTH_PATH,
+                deploy_timeout=meta.DEPLOY_TIMEOUT,
+                http_timeout=meta.HTTP_TIMEOUT,
            )
            warm_ok = True
        except Exception as e:  # noqa: BLE001
            print(f"!! canonical reattach/readiness failed: {_scrub(str(e))}", flush=True)

        if warm_ok:
-            # 2) deps (warm keycloak + per-run realm) — mirrors main()'s warm/cold split
+            # 2) deps (warm keycloak + per-run realm) — mirrors main()'s warm/cold split. NB
+            # (rcust P2b): deps are provisioned (realm/creds in $CCCI_DEPS_FILE) but quick mode
+            # cannot do install-time OIDC env wiring — the canonical app pre-exists its per-run
+            # realm. No quick-enrolled recipe declares DEPS today; if one ever does, its
+            # requires_deps tests will exercise creds-only flows or skip (F2-11 keeps the signal).
            if declared:
-                print(f"\n===== setup_custom_tests (quick): deps {declared} =====", flush=True)
+                print(f"\n===== deps (quick): {declared} =====", flush=True)
                try:
                    warm_deps, cold_deps = [], []
                    for d in declared:
                        wd = warm.warm_domain(d)
                        (warm_deps if (wd and warm.is_warm_up(d, wd)) else cold_deps).append(d)
-                    dep_metas = {d: _load_meta(d) for d in cold_deps}
+                    dep_metas = {d: meta_mod.load(d) for d in cold_deps}
                    deps_list = (
                        deps_mod.deploy_deps(
                            recipe, os.environ.get("PR", "0"), ref, cold_deps, meta_for=dep_metas
@ -629,12 +657,11 @@ def run_quick(
                        print(f"  dep: using live-warm {d} @ {wd} (per-run realm)", flush=True)
                    deps_state = _enrich_deps_with_sso(recipe, domain, deps_list)
                    deps_mod.write_run_state(deps_state)
-                    _run_setup_custom_tests_hook(recipe, domain, depsfile)
                except Exception as e:  # noqa: BLE001
                    deps_ready = False
                    deps_not_ready_reason = _scrub(str(e))[:300]
                    print(
-                        f"!! setup_custom_tests failed (deps-not-ready): {deps_not_ready_reason}",
+                        f"!! dep provisioning failed (deps-not-ready): {deps_not_ready_reason}",
                        flush=True,
                    )

@ -650,6 +677,8 @@ def run_quick(
            results["upgrade"] = "fail"
            results["custom"] = "skip"
    finally:
+        # Teardown funnel running: further SIGTERM/SIGALRM are logged + ignored (lifetime.py).
+        lifetime.begin_teardown()
        # F2-11 skip count (read before deciding pass/fail)
        requires_deps_skipped = 0
        try:
@ -747,7 +776,7 @@ def run_quick(
        overall = 1
    if sso_unverified:
        print(
-            f"!! DEPS={declared} but setup_custom_tests failed and {requires_deps_skipped} "
+            f"!! DEPS={declared} but dep provisioning failed and {requires_deps_skipped} "
            "requires_deps SKIPPED — SSO NOT verified (F2-11)",
            file=sys.stderr,
        )
@ -782,7 +811,7 @@ def promote_canonical(recipe: str, head_ref: str | None) -> None:
    if not latest:
        print(f"WC5 promote: no version tags for {recipe} — skip", flush=True)
        return
-    meta = _load_meta(recipe)
+    meta = meta_mod.load(recipe)
    # The cold run's deploy-count was already asserted + the countfile removed; don't perturb it.
    os.environ.pop("CCCI_DEPLOY_COUNT_FILE", None)
    print(
@ -794,14 +823,15 @@ def promote_canonical(recipe: str, head_ref: str | None) -> None:
        domain,
        version=latest,
        secrets=True,
-        deploy_timeout=int(meta.get("DEPLOY_TIMEOUT", 900)),
+        deploy_timeout=int(meta.DEPLOY_TIMEOUT),
+        meta=meta,
    )
    lifecycle.wait_healthy(
        domain,
-        ok_codes=tuple(meta["HEALTH_OK"]),
-        path=meta["HEALTH_PATH"],
-        deploy_timeout=meta["DEPLOY_TIMEOUT"],
-        http_timeout=meta["HTTP_TIMEOUT"],
+        ok_codes=tuple(meta.HEALTH_OK),
+        path=meta.HEALTH_PATH,
+        deploy_timeout=meta.DEPLOY_TIMEOUT,
+        http_timeout=meta.HTTP_TIMEOUT,
    )
    abra.undeploy(domain)
    _wait_undeployed(domain)
@ -813,6 +843,9 @@ def promote_canonical(recipe: str, head_ref: str | None) -> None:


 def main() -> int:
+    # P1 lock-lifetime hardening: PDEATHSIG + SIGTERM/SIGALRM teardown funnel + 60-min hard
+    # deadline, armed before ANY abra call or lock acquisition (see harness/lifetime.py).
+    lifetime.install_lifetime_guards()
    recipe = os.environ.get("RECIPE")
    if not recipe:
        print("RECIPE env is required", file=sys.stderr)
@ -827,13 +860,34 @@ def main() -> int:
    print(
        f"== cc-ci run: recipe={recipe} ref={ref} pr={os.environ.get('PR', '0')} stages={sorted(stages)}"
    )
+    # P2c: the CCCI_SKIP_GENERIC* env escape hatch is LOCAL-DEV-ONLY. If it rides a CI (drone)
+    # run, shout — generic-floor coverage is reduced and the verdict must not look routine.
+    for ov in skip_generic_env_overrides():
+        if os.environ.get("DRONE"):
+            print(
+                f"!! {ov}=1 — dev-only generic-floor override ACTIVE IN A CI RUN; generic "
+                "assertions are suppressed for the affected op(s). This must never gate a merge.",
+                flush=True,
+            )
+        else:
+            print(f"== {ov}=1 (dev-only generic-floor override active)", flush=True)
+    # Concurrent-run safety is structural: this run's recipe trees live in its own ABRA_DIR
+    # (exported here, before ANY abra call), so no recipe-tree lock exists; same-DOMAIN runs
+    # serialise on the app-domain flock taken in deploy_app (see docs/concurrency.md).
+    setup_run_abra_dir()
    fetch_recipe(recipe, ref, src)
    # The PR-head commit the upgrade tier re-checks out for the chaos redeploy to the code under test
    # (HC1). Prefer the explicit PR head sha ($REF) — robust + exact; fall back to the recipe checkout
    # HEAD (the catalogue current) for a non-PR `!testme`. Captured before any version-tag checkout.
    head_ref = ref or lifecycle.recipe_head_commit(recipe)
    repo_local = snapshot_recipe_tests(recipe)
-    meta = _load_meta(recipe)
+    meta = meta_mod.load(recipe)
+
+    # Customization manifest (rcust P5, R4): ONE block answering "what does this recipe
+    # customize?" across all surfaces — printed here and embedded verbatim in results.json under
+    # "customization". Pure presentation; never influences a verdict.
+    customization = manifest_mod.build(recipe, meta, repo_local)
+    print("\n" + manifest_mod.render(recipe, customization) + "\n", flush=True)

    # WC4/WC7: opt-in `--quick` fast lane. Requires an existing data-warm canonical; if none, fall
    # back cleanly to the full COLD run below so the PR is still tested (DECISIONS Phase-2w).
@ -856,16 +910,14 @@ def main() -> int:
    # override must be an exact published version tag (deployed as a pinned base). (Adversary §7.1.)
    want_upgrade = "upgrade" in stages
    prev = (
-        (meta.get("UPGRADE_BASE_VERSION") or lifecycle.previous_version(recipe))
-        if want_upgrade
-        else None
+        (meta.UPGRADE_BASE_VERSION or lifecycle.previous_version(recipe)) if want_upgrade else None
    )
    base = prev or target
    backup_cap = generic.backup_capable(recipe, meta)
    hook = discovery.install_steps(recipe, repo_local)

    # Deploy-count guard (DG4.1): exactly one deploy_app() per run.
-    countfile = os.path.join(tempfile.gettempdir(), f"ccci-deploys-{domain}")
+    countfile = _run_state_path("deploys")
    with open(countfile, "w") as f:
        f.write("0")
    os.environ["CCCI_DEPLOY_COUNT_FILE"] = countfile
@ -881,35 +933,27 @@ def main() -> int:

    # Run-scoped op state (HC3): the orchestrator records op results (pre-upgrade identity, backup
    # snapshot_id) here for the assertion tiers (generic + overlay) to read via generic.op_state().
-    statefile = os.path.join(tempfile.gettempdir(), f"ccci-opstate-{domain}.json")
+    statefile = _run_state_path("opstate") + ".json"
    with open(statefile, "w") as f:
        json.dump({}, f)
    os.environ["CCCI_OP_STATE_FILE"] = statefile
    op_state: dict = {}

-    # Run-scoped dep state (Phase 2 Q2.3, refined per operator-2026-05-28 SSO-dep plan §1):
-    # deps now deploy AFTER generic tiers (between RESTORE and CUSTOM) so a failed dep deploy
-    # cannot break the generic-tier signal. The `setup_custom_tests` step deploys each dep + runs
-    # `tests/<recipe>/setup_custom_tests.sh` to wire OIDC env via in-place redeploy.
+    # Run-scoped dep state (Phase 2 Q2.3; install-time-only since rcust P2b): deps are provisioned
+    # BEFORE the single deploy so install_steps.sh wires OIDC env into that one deploy.
    # `$CCCI_DEPS_FILE` is written with the full creds dict the hook script needs (jq-readable).
-    depsfile = os.path.join(tempfile.gettempdir(), f"ccci-deps-{domain}.json")
+    depsfile = _run_state_path("deps") + ".json"
    with open(depsfile, "w") as f:
        json.dump({}, f)
    os.environ["CCCI_DEPS_FILE"] = depsfile
    # F2-11: conftest appends the count of requires_deps tests it skips (deps-not-ready) here.
-    skipfile = os.path.join(tempfile.gettempdir(), f"ccci-depskip-{domain}.txt")
+    skipfile = _run_state_path("depskip") + ".txt"
    with contextlib.suppress(OSError):
        os.remove(skipfile)
    os.environ["CCCI_DEPS_SKIP_REPORT"] = skipfile
-    declared = deps_mod.declared_deps(recipe)
-    # Q3.2a: a recipe that tolerates OIDC env at first boot AND whose deps are live-warm wires OIDC
-    # at INSTALL time (provision the realm BEFORE the single deploy; install_steps.sh writes the env
-    # into it) instead of the post-deploy in-place `--chaos` redeploy — which is flaky on the heavy
-    # 12-service lasuite-drive stack (collabora WOPI race; see JOURNAL Step 0). Opt-in per recipe.
-    oidc_at_install = bool(meta.get("OIDC_AT_INSTALL")) and bool(declared)
+    declared = list(meta.DEPS)
    if declared:
-        when = "BEFORE deploy (install-time OIDC)" if oidc_at_install else "AFTER generic tiers"
-        print(f"\n===== DEPS declared (provision {when}): {declared} =====", flush=True)
+        print(f"\n===== DEPS declared (provision BEFORE deploy): {declared} =====", flush=True)
    deps_state: dict[str, dict] = {}  # new shape: recipe→entry dict (sso-dep plan §1)
    deps_ready = True
    deps_not_ready_reason: str = ""
@ -923,7 +967,7 @@ def main() -> int:
        # install_steps.sh can read $CCCI_DEPS_FILE and wire the OIDC env into that one deploy. On
        # failure we mark deps-not-ready but STILL deploy the recipe alone (install_steps.sh no-ops
        # on an empty deps file) so the generic tiers run; the OIDC custom test then skips → F2-11. ----
-        if oidc_at_install:
+        if declared:
            print(
                f"\n===== install-time OIDC: provisioning deps {declared} BEFORE deploy =====",
                flush=True,
@ -950,18 +994,21 @@ def main() -> int:
                version=base,
                secrets=True,
                install_steps_hook=hook,
-                deploy_timeout=int(meta.get("DEPLOY_TIMEOUT", 900)),
+                deploy_timeout=int(meta.DEPLOY_TIMEOUT),
+                meta=meta,
            )
            lifecycle.wait_healthy(
                domain,
-                ok_codes=tuple(meta["HEALTH_OK"]),
-                path=meta["HEALTH_PATH"],
-                deploy_timeout=meta["DEPLOY_TIMEOUT"],
-                http_timeout=meta["HTTP_TIMEOUT"],
+                ok_codes=tuple(meta.HEALTH_OK),
+                path=meta.HEALTH_PATH,
+                deploy_timeout=meta.DEPLOY_TIMEOUT,
+                http_timeout=meta.HTTP_TIMEOUT,
            )
            # Recipe READY_PROBE (e.g. lasuite-drive collabora WOPI discovery) — readiness beyond
            # replica convergence + app HEALTH_PATH; no-op for recipes without one.
-            lifecycle.wait_ready_probes(meta, domain, timeout=int(meta.get("DEPLOY_TIMEOUT", 900)))
+            lifecycle.wait_ready_probes(
+                meta, domain, timeout=int(meta.DEPLOY_TIMEOUT), op="install"
+            )
            deploy_ok = True
        except Exception as e:  # noqa: BLE001 — a failed deploy is a reported INSTALL failure
            print(f"!! deploy/readiness failed: {e}", flush=True)
@ -1058,41 +1105,11 @@ def main() -> int:
                    if backup_cap
                    else "skip"
                )
-            # ---- setup_custom_tests step (NEW, operator-2026-05-28 SSO-dep plan §3.2) ----
-            # Deploy each declared dep + wire OIDC env into the parent app via the per-recipe
-            # setup_custom_tests.sh hook + in-place redeploy. Failure here marks deps-not-ready
-            # but does NOT abort the run — @pytest.mark.requires_deps tests skip with reason;
-            # non-deps custom tests still run normally.
-            if declared and not oidc_at_install:
-                # LEGACY post-deploy path: provision deps AFTER generic tiers, then wire OIDC env
-                # into the parent via the setup_custom_tests.sh hook + an in-place `--chaos` redeploy.
-                print("\n===== setup_custom_tests: deps + OIDC wiring =====", flush=True)
-                try:
-                    deps_state = _provision_deps(recipe, domain, ref, declared)
-                    # Run the per-recipe post-deps hook (jq-driven OIDC wiring + in-place redeploy)
-                    _run_setup_custom_tests_hook(recipe, domain, depsfile)
-                except Exception as e:  # noqa: BLE001 — setup failure is ISOLATED to dep-marked tests
-                    deps_ready = False
-                    deps_not_ready_reason = _scrub(str(e))[:300]
-                    print(
-                        f"!! setup_custom_tests failed (deps-not-ready): {deps_not_ready_reason}",
-                        flush=True,
-                    )
-            elif declared and oidc_at_install and deps_ready:
-                # INSTALL-TIME path (Q3.2a): deps were provisioned BEFORE the single deploy and the
-                # install-tier install_steps.sh hook already wired OIDC env into that one deploy —
-                # so NO re-provision, NO reconverge here. Run only the post-deploy setup hook
-                # (e.g. lasuite-drive's minio-createbuckets one-shot), which needs the live stack.
-                print("\n===== post-deploy setup (OIDC already wired at install) =====", flush=True)
-                try:
-                    _run_setup_custom_tests_hook(recipe, domain, depsfile)
-                except Exception as e:  # noqa: BLE001 — isolated to dep-marked / state-dependent tests
-                    deps_ready = False
-                    deps_not_ready_reason = _scrub(str(e))[:300]
-                    print(
-                        f"!! post-deploy setup failed: {deps_not_ready_reason}",
-                        flush=True,
-                    )
+            # (rcust P2b: install-time deps wiring is the ONLY mode — deps were provisioned BEFORE
+            # the single deploy and install_steps.sh wired the OIDC env into it. The legacy
+            # post-deploy provisioning + setup_custom_tests.sh redeploy machinery is deleted; a
+            # recipe's post-deploy seeding belongs in ops.py pre_install, e.g. lasuite-drive's
+            # MinIO bucket one-shot.)

            # ---- CUSTOM tier ----
            if "custom" in stages:
@ -1109,6 +1126,9 @@ def main() -> int:
                if op in stages:
                    results[op] = "skip"
    finally:
+        # From here the teardown funnel runs: a SIGTERM/SIGALRM landing now is logged + ignored
+        # (lifetime.py) so a second signal can't abort the cleanup the first one asked for.
+        lifetime.begin_teardown()
        # Teardown the recipe under test FIRST, then deps in reverse declaration order.
        # Parent verify=False (Phase 1d): keep as-is so a parent residual doesn't mask a tier
        # failure. Dep teardown uses verify=True via teardown_deps (F2-5 fix); failures are
@ -1164,8 +1184,7 @@ def main() -> int:

    # ---- per-op summary (DG6 feed) ----
    # SSO-dep plan §1: DG4.1 generalised — one `abra app new` per app in the run (recipe + each
-    # COLD dep). In-place reconfigure-and-redeploy (the setup_custom_tests step's
-    # `abra app deploy --force --chaos`) is NOT a fresh `app_new` and does NOT increment the count.
+    # COLD dep). Chaos redeploys are NOT a fresh `app_new` and do NOT increment the count.
    # WC1: a live-warm dep (keycloak) is NOT deployed by the run — it only gets a per-run realm — so
    # warm deps contribute 0. So expected = 1 + (number of COLD deps that actually got deployed).
    _dep_entries = deps_state.values() if isinstance(deps_state, dict) else (deps_state or [])
@ -1206,12 +1225,12 @@ def main() -> int:
        overall = 1
    if any(v == "fail" for v in results.values()):
        overall = 1
-    # F2-11: a deps-declaring recipe whose setup_custom_tests failed has NOT verified its SSO/OIDC
+    # F2-11: a deps-declaring recipe whose dep provisioning failed has NOT verified its SSO/OIDC
    # claim — its requires_deps tests SKIPPED (a skip-only file exits 0, so without this the run
    # would report GREEN). Fail the run for that recipe; generic-tier results above are untouched.
    if sso_dep_unverified(declared, deps_ready, requires_deps_skipped):
        print(
-            f"!! recipe declares DEPS={declared} but setup_custom_tests failed and "
+            f"!! recipe declares DEPS={declared} but dep provisioning failed and "
            f"{requires_deps_skipped} requires_deps (SSO) test(s) were SKIPPED — SSO claim NOT "
            f"verified; failing run (F2-11). deps-not-ready: {deps_not_ready_reason}",
            file=sys.stderr,
@ -1238,7 +1257,8 @@ def main() -> int:
            no_secret_leak=True,  # narrowed below by an actual scan of the serialised artifact
            screenshot=screenshot_rel,  # Phase 3 U1 (R4): relative PNG name iff capture succeeded
            finished_ts=time.time(),
-            expected_na=meta.get("EXPECTED_NA"),  # declared intentional-skip map (recipe_meta)
+            expected_na=meta.EXPECTED_NA,  # declared intentional-skip map (recipe_meta)
+            customization=customization,  # rcust P5: the run-start manifest, verbatim
        )
        # Real (if narrow) leak check: no known infra-secret value may appear in the artifact (R7).
        blob = json.dumps(data)
@ -1285,8 +1305,10 @@ def main() -> int:
            capped = data.get("level_cap_rung")
            sk = data.get("skips", {})
            cap_skip = (
-                "intentional" if capped in (sk.get("intentional") or {})
-                else "unintentional" if capped in (sk.get("unintentional") or [])
+                "intentional"
+                if capped in (sk.get("intentional") or {})
+                else "unintentional"
+                if capped in (sk.get("unintentional") or [])
                else ""
            )
            with open(os.path.join(run_artifact_dir, "badge.svg"), "w", encoding="utf-8") as f:
--- a/runner/warm_reconcile.py
+++ b/runner/warm_reconcile.py
@ -43,11 +43,16 @@ def _traefik_setup(recipe: str, domain: str, version: str) -> None:
    ssl_cert/ssl_key swarm secrets; NO ACME). Uses the proven abra.env_set (newline-safe, unlike the
    bash set_env that bit keycloak)."""
    cert_dir = "/var/lib/ci-certs/live"
-    if not (os.path.isfile(f"{cert_dir}/fullchain.pem") and os.path.isfile(f"{cert_dir}/privkey.pem")):
+    if not (
+        os.path.isfile(f"{cert_dir}/fullchain.pem") and os.path.isfile(f"{cert_dir}/privkey.pem")
+    ):
        raise RuntimeError(f"FATAL: wildcard cert missing at {cert_dir} (sops decrypt broken?)")
    if not os.path.isfile(env_file(domain)):
-        _run(["abra", "app", "new", recipe, "-s", "default", "-D", domain, version, "-o", "-n"],
-             timeout=120, check=True)
+        _run(
+            ["abra", "app", "new", recipe, "-s", "default", "-D", domain, version, "-o", "-n"],
+            timeout=120,
+            check=True,
+        )
    abra.env_set(domain, "DOMAIN", domain)
    abra.env_set(domain, "LETS_ENCRYPT_ENV", "")
    abra.env_set(domain, "WILDCARDS_ENABLED", "1")
@ -61,11 +66,39 @@ def _traefik_setup(recipe: str, domain: str, version: str) -> None:
        return any(s.endswith(f"_{name}_v1") for s in have)

    if not _has("ssl_cert"):
-        _run(["abra", "app", "secret", "insert", domain, "ssl_cert", "v1",
-              f"{cert_dir}/fullchain.pem", "-f", "-n"], timeout=120, check=True)
+        _run(
+            [
+                "abra",
+                "app",
+                "secret",
+                "insert",
+                domain,
+                "ssl_cert",
+                "v1",
+                f"{cert_dir}/fullchain.pem",
+                "-f",
+                "-n",
+            ],
+            timeout=120,
+            check=True,
+        )
    if not _has("ssl_key"):
-        _run(["abra", "app", "secret", "insert", domain, "ssl_key", "v1",
-              f"{cert_dir}/privkey.pem", "-f", "-n"], timeout=120, check=True)
+        _run(
+            [
+                "abra",
+                "app",
+                "secret",
+                "insert",
+                domain,
+                "ssl_key",
+                "v1",
+                f"{cert_dir}/privkey.pem",
+                "-f",
+                "-n",
+            ],
+            timeout=120,
+            check=True,
+        )


 SPECS: dict[str, dict] = {
@ -166,7 +199,13 @@ def _run(cmd, timeout=120, check=False):


 def _recipe_dir(recipe: str) -> str:
-    return os.path.expanduser(f"~/.abra/recipes/{recipe}")
+    # Resolve like the abra CLI does: $ABRA_DIR (the per-run tree when imported by a CI run,
+    # e.g. promote_canonical) else the canonical ~/.abra (this module's own systemd-timer runs,
+    # which set no ABRA_DIR). Keeps fetch_recipe (an `abra` subprocess) and the git readers
+    # below pointed at the SAME tree in both contexts.
+    return os.path.join(
+        os.environ.get("ABRA_DIR") or os.path.expanduser("~/.abra"), "recipes", recipe
+    )


 def recipe_tags(recipe: str) -> list[str]:
@ -218,8 +257,17 @@ def health_code(spec: dict) -> int:
    domain = spec.get("health_domain", spec["domain"])
    r = _run(
        [
-            "curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "--max-time", "10",
-            "--resolve", f"{domain}:443:127.0.0.1", f"https://{domain}{spec['health_path']}",
+            "curl",
+            "-sk",
+            "-o",
+            "/dev/null",
+            "-w",
+            "%{http_code}",
+            "--max-time",
+            "10",
+            "--resolve",
+            f"{domain}:443:127.0.0.1",
+            f"https://{domain}{spec['health_path']}",
        ],
        timeout=20,
    )
@ -230,7 +278,6 @@ def health_code(spec: dict) -> int:


 def wait_healthy(spec: dict, timeout: int | None = None) -> bool:
-    domain = spec["domain"]
    deadline = time.time() + (timeout or spec["health_timeout"])
    while time.time() < deadline:
        if health_code(spec) in tuple(spec["health_ok"]):
@ -325,15 +372,18 @@ def ensure_server() -> None:

 def ensure_app_config(recipe: str, domain: str, version: str) -> None:
    if not os.path.isfile(env_file(domain)):
-        _run(["abra", "app", "new", recipe, "-s", "default", "-D", domain, version, "-o", "-n"],
-             timeout=120, check=True)
+        _run(
+            ["abra", "app", "new", recipe, "-s", "default", "-D", domain, version, "-o", "-n"],
+            timeout=120,
+            check=True,
+        )
    abra.env_set(domain, "DOMAIN", domain)
    abra.env_set(domain, "LETS_ENCRYPT_ENV", "")


 def ensure_secrets(domain: str) -> None:
    stack = lifecycle._stack_name(domain)  # noqa: SLF001
-    have = {n for n in lifecycle._docker_names("secret", stack)}  # noqa: SLF001
+    have = set(lifecycle._docker_names("secret", stack))  # noqa: SLF001
    if not any(n.endswith("_admin_password_v1") for n in have):
        abra.secret_generate(domain)

@ -393,8 +443,9 @@ def reconcile(app: str) -> str:
        write_alert(app, "held-major", current=current, latest=latest, release_notes=notes[:4000])
        return f"held-major:{current}->{latest}"
    if notes_flag_manual_migration(notes):
-        write_alert(app, "held-manual-migration", current=current, latest=latest,
-                    release_notes=notes[:4000])
+        write_alert(
+            app, "held-manual-migration", current=current, latest=latest, release_notes=notes[:4000]
+        )
        return f"held-manual-migration:{current}->{latest}"

    # WC1.1 health-gated upgrade with rollback.
@ -428,8 +479,14 @@ def reconcile(app: str) -> str:
        warmsnap.restore(recipe, domain)
    deploy_version(recipe, domain, last_good, dt)
    recovered = wait_healthy(spec)
-    write_alert(app, "rollback", last_good=last_good, attempted=latest, recovered=recovered,
-                release_notes=notes[:2000])
+    write_alert(
+        app,
+        "rollback",
+        last_good=last_good,
+        attempted=latest,
+        recovered=recovered,
+        release_notes=notes[:2000],
+    )
    if not recovered:
        raise RuntimeError(f"{app} rollback to {last_good} did not become healthy")
    return f"rolled-back:{latest}->{last_good}"
--- a/scripts/gen-meta-docs.py
+++ b/scripts/gen-meta-docs.py
@ -0,0 +1,71 @@
+#!/usr/bin/env python3
+"""Render the harness.meta KEYS registry to the markdown key-reference table in
+docs/recipe-customization.md §4 (rcust P1.5; kills the R5 doc-drift class).
+
+Usage:
+  python3 scripts/gen-meta-docs.py            # rewrite the table in-place between the markers
+  python3 scripts/gen-meta-docs.py --print    # print the rendered table to stdout (used by the
+                                              # doc-sync unit test, tests/unit/test_meta.py)
+
+The table lives between `<!-- META-TABLE-START -->` / `<!-- META-TABLE-END -->` markers; a unit
+test asserts the committed table equals this rendering, so editing it by hand fails CI.
+"""
+
+from __future__ import annotations
+
+import os
+import sys
+
+ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
+sys.path.insert(0, os.path.join(ROOT, "runner"))
+from harness.meta import KEYS  # noqa: E402
+
+DOC = os.path.join(ROOT, "docs", "recipe-customization.md")
+START = "<!-- META-TABLE-START -->"
+END = "<!-- META-TABLE-END -->"
+
+
+def _default_repr(v) -> str:
+    if v is None:
+        return "`None`"
+    return f"`{v!r}`"
+
+
+def render() -> str:
+    lines = [
+        START,
+        "",
+        "_This table is GENERATED from the `runner/harness/meta.py` KEYS registry by"
+        " `scripts/gen-meta-docs.py` — do not edit by hand (a unit test pins the sync)._",
+        "",
+        "| Key | Type | Default | Meaning |",
+        "|---|---|---|---|",
+    ]
+    for k in KEYS:
+        doc = k.doc.replace("|", "\\|")
+        name = f"`{k.name}`" + (" **(deprecated)**" if k.deprecated else "")
+        lines.append(f"| {name} | `{k.type}` | {_default_repr(k.default)} | {doc} |")
+    lines += ["", END]
+    return "\n".join(lines)
+
+
+def main() -> int:
+    table = render()
+    if "--print" in sys.argv:
+        print(table)
+        return 0
+    with open(DOC) as f:
+        text = f.read()
+    if START not in text or END not in text:
+        print(f"{DOC}: missing {START}/{END} markers", file=sys.stderr)
+        return 1
+    head, _, rest = text.partition(START)
+    _, _, tail = rest.partition(END)
+    with open(DOC, "w") as f:
+        f.write(head + table + tail)
+    print(f"{DOC}: key table rewritten from the registry ({len(KEYS)} keys)")
+    return 0
+
+
+if __name__ == "__main__":
+    raise SystemExit(main())
--- a/tests/bluesky-pds/_p4.py
+++ b/tests/bluesky-pds/_p4.py
@ -15,7 +15,8 @@ import shlex
 import sys

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
-from harness import http as harness_http, lifecycle  # noqa: E402
+from harness import http as harness_http  # noqa: E402
+from harness import lifecycle

 PDS_HOST_LOCAL = "http://localhost:3000"
 _PW = "ccci-P4-marker-pw-2026"
--- a/tests/bluesky-pds/functional/test_account_and_post.py
+++ b/tests/bluesky-pds/functional/test_account_and_post.py
@ -27,6 +27,7 @@ CRUD). A wedged PDS subsystem fails AT its layer.

 from __future__ import annotations

+import contextlib
 import os
 import re
 import secrets
@ -35,7 +36,8 @@ import sys
 import uuid

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "..", "runner"))
-from harness import http as harness_http, lifecycle  # noqa: E402
+from harness import http as harness_http  # noqa: E402
+from harness import lifecycle

 PDS_HOST_LOCAL = "http://localhost:3000"

@ -58,14 +60,18 @@ def _goat_admin(domain: str, args: str) -> str:
    return _in_container(domain, cmd)


-def _xrpc_post(domain: str, nsid: str, data: dict, token: str | None = None) -> tuple[int, dict | None]:
+def _xrpc_post(
+    domain: str, nsid: str, data: dict, token: str | None = None
+) -> tuple[int, dict | None]:
    headers = {}
    if token:
        headers["Authorization"] = f"Bearer {token}"
    return harness_http.http_post(f"https://{domain}/xrpc/{nsid}", data=data, headers=headers)


-def _xrpc_get(domain: str, nsid: str, query: str, token: str | None = None) -> tuple[int, dict | None]:
+def _xrpc_get(
+    domain: str, nsid: str, query: str, token: str | None = None
+) -> tuple[int, dict | None]:
    headers = {}
    if token:
        headers["Authorization"] = f"Bearer {token}"
@ -82,9 +88,9 @@ def test_account_lifecycle_and_post_roundtrip(live_app):

    # Step 1: PDS describe via goat — recipe self-identifies as did:web:<domain>
    out = _in_container(domain, f"goat pds describe {PDS_HOST_LOCAL} 2>&1")
-    assert f"did:web:{domain}" in out, (
-        f"goat pds describe did not contain expected DID 'did:web:{domain}'. Output:\n{out[:500]!r}"
-    )
+    assert (
+        f"did:web:{domain}" in out
+    ), f"goat pds describe did not contain expected DID 'did:web:{domain}'. Output:\n{out[:500]!r}"

    # Step 2: Create account (UUID-suffixed handle = no run-to-run collision)
    out = _goat_admin(
@ -127,9 +133,9 @@ def test_account_lifecycle_and_post_roundtrip(live_app):
        assert s == 200, f"createRecord HTTP {s}: {body!r}"
        record_uri = (body or {}).get("uri", "")
        # URI format: at://<did>/app.bsky.feed.post/<rkey>
-        assert record_uri.startswith(f"at://{new_did}/app.bsky.feed.post/"), (
-            f"unexpected record uri: {record_uri!r}"
-        )
+        assert record_uri.startswith(
+            f"at://{new_did}/app.bsky.feed.post/"
+        ), f"unexpected record uri: {record_uri!r}"
        rkey = record_uri.rsplit("/", 1)[-1]
        assert rkey, f"no rkey in uri: {record_uri!r}"

@ -142,15 +148,13 @@ def test_account_lifecycle_and_post_roundtrip(live_app):
        )
        assert s == 200, f"getRecord HTTP {s}: {body!r}"
        record_value = (body or {}).get("value", {})
-        assert record_value.get("text") == marker, (
-            f"post text did not round-trip: created={marker!r}, fetched={record_value.get('text')!r}"
-        )
+        assert (
+            record_value.get("text") == marker
+        ), f"post text did not round-trip: created={marker!r}, fetched={record_value.get('text')!r}"
        assert record_value.get("$type") == "app.bsky.feed.post"
    finally:
        # Step 6: Best-effort cleanup. (The per-run domain teardown will discard the volume
        # too, but we exercise the delete-account path because it's part of §4.3.)
        if cleanup_did:
-            try:
+            with contextlib.suppress(Exception):
                _goat_admin(domain, f"account delete {cleanup_did}")
-            except Exception:  # noqa: BLE001
-                pass
--- a/tests/bluesky-pds/functional/test_describe_server.py
+++ b/tests/bluesky-pds/functional/test_describe_server.py
@ -26,6 +26,6 @@ def test_describe_server_returns_atproto_envelope(live_app):
    # At least one of these atproto-spec fields must be present
    expected_any = ("availableUserDomains", "inviteCodeRequired", "links", "did")
    present = [k for k in expected_any if k in body]
-    assert present, (
-        f"describe-server missing all of {expected_any}; got keys: {sorted(body.keys())[:20]}"
-    )
+    assert (
+        present
+    ), f"describe-server missing all of {expected_any}; got keys: {sorted(body.keys())[:20]}"
--- a/tests/bluesky-pds/functional/test_health_check.py
+++ b/tests/bluesky-pds/functional/test_health_check.py
@ -17,6 +17,6 @@ def test_pds_health_returns_version(live_app):
    url = f"https://{live_app}/xrpc/_health"
    status, body = harness_http.retry_http_get(url, expect_status=200, max_wait=60, interval=3)
    assert status == 200, f"GET {url} HTTP {status} (expected 200)"
-    assert isinstance(body, dict) and isinstance(body.get("version"), str) and body["version"], (
-        f"GET {url} response is not the expected health envelope: {body!r}"
-    )
+    assert (
+        isinstance(body, dict) and isinstance(body.get("version"), str) and body["version"]
+    ), f"GET {url} response is not the expected health envelope: {body!r}"
--- a/tests/bluesky-pds/functional/test_session_auth.py
+++ b/tests/bluesky-pds/functional/test_session_auth.py
@ -30,6 +30,6 @@ def test_get_session_requires_auth(live_app):
        f"body: {body!r}"
    )
    # The XRPC error envelope is JSON with an `error` field per the atproto spec.
-    assert isinstance(body, dict) and body.get("error"), (
-        f"expected XRPC JSON error envelope; got: {body!r}"
-    )
+    assert isinstance(body, dict) and body.get(
+        "error"
+    ), f"expected XRPC JSON error envelope; got: {body!r}"
--- a/tests/bluesky-pds/install_steps.sh
+++ b/tests/bluesky-pds/install_steps.sh
@ -22,12 +22,12 @@ echo "  bluesky-pds install_steps: generating secp256k1 PLC rotation key..."
 # same shape the PDS expects (32-byte hex). Equivalent for atproto PDS bootstrap.
 KEY_HEX=$(cc-ci-run -c 'import secrets; print(secrets.token_bytes(32).hex())')
 if [ -z "${KEY_HEX}" ] || [ "${#KEY_HEX}" != "64" ]; then
-    echo "  install_steps: failed to generate PLC rotation key (KEY_HEX length=${#KEY_HEX})" >&2
-    exit 1
+  echo "  install_steps: failed to generate PLC rotation key (KEY_HEX length=${#KEY_HEX})" >&2
+  exit 1
 fi

 # Insert via abra under TTY-wrap (`abra app secret insert` requires a TTY on this version).
 # We DON'T log the key value — abra also doesn't print it.
 script -qec "abra app secret insert ${CCCI_APP_DOMAIN} pds_plc_rotation_key v1 ${KEY_HEX} --no-input" /dev/null \
-    >/dev/null 2>&1
+  >/dev/null 2>&1
 echo "  bluesky-pds install_steps: PLC rotation key inserted (v1)."
--- a/tests/bluesky-pds/ops.py
+++ b/tests/bluesky-pds/ops.py
@ -9,14 +9,14 @@ sys.path.insert(0, os.path.dirname(__file__))
 import _p4  # noqa: E402


-def pre_upgrade(domain, meta):
-    _p4.create_account(domain)
+def pre_upgrade(ctx):
+    _p4.create_account(ctx.domain)


-def pre_backup(domain, meta):
-    _p4.create_account(domain)
+def pre_backup(ctx):
+    _p4.create_account(ctx.domain)


-def pre_restore(domain, meta):
-    _p4.delete_account(domain)
-    assert not _p4.account_exists(domain), "marker account delete did not take (pre_restore)"
+def pre_restore(ctx):
+    _p4.delete_account(ctx.domain)
+    assert not _p4.account_exists(ctx.domain), "marker account delete did not take (pre_restore)"
--- a/tests/bluesky-pds/test_restore.py
+++ b/tests/bluesky-pds/test_restore.py
@ -11,6 +11,6 @@ import _p4  # noqa: E402


 def test_restore_returns_state(live_app):
-    assert _p4.account_exists(live_app), (
-        "restore did not bring back the seeded marker account (PDS data did not survive restore)"
-    )
+    assert _p4.account_exists(
+        live_app
+    ), "restore did not bring back the seeded marker account (PDS data did not survive restore)"
--- a/tests/concurrency/concutil.py
+++ b/tests/concurrency/concutil.py
@ -0,0 +1,108 @@
+"""Shared utilities for the real-kernel concurrency suite (imported by the test modules; the
+fixtures in conftest.py wrap these). No flock mocking anywhere — probes use real LOCK_NB."""
+
+from __future__ import annotations
+
+import contextlib
+import fcntl
+import os
+import signal
+import subprocess
+import sys
+import time
+
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
+from harness import lifecycle  # noqa: E402
+
+HELPERS = os.path.join(os.path.dirname(__file__), "helpers.py")
+DOMAIN = "test-abc123.ci.commoninternet.net"  # matches RUN_APP_RE
+
+
+class HelperPool:
+    """Spawns helpers.py subprocesses and GUARANTEES their cleanup (incl. recorded grandchild
+    pids from `hold-with-child`/`wrapper` markers) — no leaked children in the test VM."""
+
+    def __init__(self, out_dir: str):
+        self.out_dir = out_dir
+        self.procs: list[subprocess.Popen] = []
+        self.extra_pids: list[int] = []
+        self._n = 0
+
+    def spawn(self, *args: str, env_extra: dict | None = None) -> tuple[subprocess.Popen, str]:
+        """Start `helpers.py <args...>`; returns (proc, marker_file)."""
+        self._n += 1
+        out = os.path.join(self.out_dir, f"helper-{self._n}.out")
+        env = dict(os.environ, CCCI_HELPER_OUT=out, **(env_extra or {}))
+        p = subprocess.Popen(  # noqa: S603
+            [sys.executable, HELPERS, *args],
+            env=env,
+            stdout=subprocess.DEVNULL,
+            stderr=subprocess.STDOUT,
+        )
+        self.procs.append(p)
+        return p, out
+
+    def track_pid(self, pid: int) -> None:
+        self.extra_pids.append(pid)
+
+    def cleanup(self) -> None:
+        for p in self.procs:
+            if p.poll() is None:
+                p.kill()
+            with contextlib.suppress(subprocess.TimeoutExpired):
+                p.wait(timeout=10)
+        for pid in self.extra_pids:
+            with contextlib.suppress(OSError):
+                os.kill(pid, signal.SIGKILL)
+
+
+def wait_marker(out: str, token: str, timeout: float = 15.0) -> str | None:
+    """Poll a helper's marker file for a line containing `token`; returns the line or None."""
+    deadline = time.time() + timeout
+    while time.time() < deadline:
+        try:
+            with open(out) as f:
+                for line in f:
+                    if token in line:
+                        return line.strip()
+        except OSError:
+            pass
+        time.sleep(0.1)
+    return None
+
+
+def lock_state(domain: str) -> str:
+    """'held' | 'free' | 'absent' for the domain's lockfile, probed with a REAL LOCK_NB."""
+    path = lifecycle._app_lock_path(domain)  # noqa: SLF001
+    if not os.path.exists(path):
+        return "absent"
+    with open(path, "a") as f:
+        try:
+            fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)
+            return "free"
+        except BlockingIOError:
+            return "held"
+
+
+def wait_lock_state(domain: str, want: str, timeout: float = 10.0) -> str:
+    """Poll until lock_state(domain) == want (kernel release on process death is fast, but give
+    the scheduler room). Returns the final observed state."""
+    deadline = time.time() + timeout
+    state = lock_state(domain)
+    while state != want and time.time() < deadline:
+        time.sleep(0.1)
+        state = lock_state(domain)
+    return state
+
+
+def pid_alive(pid: int) -> bool:
+    return os.path.exists(f"/proc/{pid}")
+
+
+def wait_pid_gone(pid: int, timeout: float = 15.0) -> bool:
+    deadline = time.time() + timeout
+    while time.time() < deadline:
+        if not pid_alive(pid):
+            return True
+        time.sleep(0.1)
+    return False
--- a/tests/concurrency/conftest.py
+++ b/tests/concurrency/conftest.py
@ -0,0 +1,34 @@
+"""Fixtures for the real-kernel concurrency suite (concurrency-restructure plan, 19 cases).
+
+NOT part of the default `pytest tests/unit` gate — run explicitly with `pytest tests/concurrency
+-q` (docs/concurrency.md). Locks live in a per-test tmp dir (CCCI_APP_LOCK_DIR); helper
+subprocesses hold REAL flocks / install the REAL prctl+signal guards and are always reaped in
+fixture finalizers (no leaked children in the test VM).
+"""
+
+from __future__ import annotations
+
+import os
+import sys
+
+import pytest
+
+sys.path.insert(0, os.path.dirname(__file__))
+from concutil import HelperPool  # noqa: E402
+
+
+@pytest.fixture
+def lock_dir(tmp_path, monkeypatch):
+    """Sandbox lock dir, exported so BOTH this process's lifecycle calls and helper subprocesses
+    (which inherit os.environ) resolve their lockfiles here — never /run/lock."""
+    d = tmp_path / "locks"
+    d.mkdir()
+    monkeypatch.setenv("CCCI_APP_LOCK_DIR", str(d))
+    return str(d)
+
+
+@pytest.fixture
+def pool(tmp_path):
+    hp = HelperPool(str(tmp_path))
+    yield hp
+    hp.cleanup()
--- a/tests/concurrency/helpers.py
+++ b/tests/concurrency/helpers.py
@ -0,0 +1,149 @@
+#!/usr/bin/env python3
+"""Subprocess helpers for tests/concurrency — REAL kernel locks and the REAL lifetime guards in
+separate processes (flock/prctl are never mocked; tests assert on actual kernel behavior).
+
+Invoked as:  python3 helpers.py <command> <args...>
+
+Env contract (set by the spawning test):
+  CCCI_APP_LOCK_DIR  sandbox lock dir (never /run/lock in tests)
+  CCCI_HELPER_OUT    marker file this helper APPENDS progress lines to (ACQUIRED/READY/...)
+
+Commands:
+  hold <domain>                 acquire the app lock, mark `ACQUIRED <ts>`, sleep forever
+  hold-with-child <domain>      acquire the lock, spawn a plain sleeping subprocess child, mark
+                                `ACQUIRED <ts>` + `CHILD <pid>` (PEP 446: the child must NOT
+                                inherit the lock fd), sleep forever
+  guarded <domain> <deadline>   install the REAL lifetime guards (alarm=<deadline>s), acquire the
+                                lock, mark `READY`; when the teardown funnel runs (`finally:`),
+                                mark `TEARDOWN` before exiting
+  wrapper <domain>              spawn `guarded <domain> 3600` as MY child, mark `WRAPPED <pid>`,
+                                sleep — the test kills me to prove PDEATHSIG TERMs the child
+  orphan-probe                  wait (bounded) until reparented (ppid==1), then install the
+                                guards; mark `REFUSED` if they exit (expected) or `GUARDS_OK`
+  fetch-checkout <recipe> <ref> run run_recipe_ci.fetch_recipe (the test sets CCCI_SKIP_FETCH=1
+                                + a per-"run" ABRA_DIR), git-checkout <ref>, mark
+                                `RESULT <head> <data.txt content>`
+"""
+
+from __future__ import annotations
+
+import os
+import subprocess
+import sys
+import time
+
+sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "..", "runner"))
+from harness import abra, lifecycle, lifetime  # noqa: E402
+
+OUT = os.environ.get("CCCI_HELPER_OUT")
+
+
+def mark(line: str) -> None:
+    if OUT:
+        with open(OUT, "a") as f:
+            f.write(line + "\n")
+            f.flush()
+    print(line, flush=True)
+
+
+def cmd_hold(domain: str) -> None:
+    lifecycle.acquire_app_lock(domain)
+    mark(f"ACQUIRED {time.time()}")
+    time.sleep(3600)
+
+
+def cmd_hold_with_child(domain: str) -> None:
+    lifecycle.acquire_app_lock(domain)
+    child = subprocess.Popen([sys.executable, "-c", "import time; time.sleep(3600)"])
+    mark(f"ACQUIRED {time.time()}")
+    mark(f"CHILD {child.pid}")
+    time.sleep(3600)
+
+
+def cmd_guarded(domain: str, deadline: str) -> None:
+    lifetime.install_lifetime_guards(deadline_seconds=int(deadline))
+    lifecycle.acquire_app_lock(domain)
+    mark("READY")
+    try:
+        time.sleep(3600)
+    finally:
+        mark("TEARDOWN")
+
+
+def cmd_wrapper(domain: str) -> None:
+    p = subprocess.Popen(  # noqa: S603
+        [sys.executable, os.path.abspath(__file__), "guarded", domain, "3600"],
+        env=os.environ.copy(),
+    )
+    mark(f"WRAPPED {p.pid}")
+    time.sleep(3600)
+
+
+def cmd_orphan_probe() -> None:
+    # Our spawner exits immediately after fork; wait (bounded) until we are reparented so the
+    # prctl is installed with the parent ALREADY dead — the exact race the ppid check closes.
+    for _ in range(200):
+        if os.getppid() == 1:
+            break
+        time.sleep(0.05)
+    else:
+        mark("NEVER_REPARENTED")  # e.g. a subreaper environment — test will fail visibly
+        return
+    try:
+        lifetime.install_lifetime_guards()
+    except SystemExit:
+        mark("REFUSED")
+        raise
+    mark("GUARDS_OK")
+
+
+def cmd_fetch_checkout(recipe: str, ref: str) -> None:
+    import run_recipe_ci
+
+    run_recipe_ci.fetch_recipe(recipe, None, None)
+    abra.recipe_checkout(recipe, ref)
+    head = abra.recipe_head_commit(recipe)
+    with open(os.path.join(abra.recipe_dir(recipe), "data.txt")) as f:
+        content = f.read().strip()
+    mark(f"RESULT {head} {content}")
+
+
+def cmd_deploy_count_run(domain: str, gate: str) -> None:
+    """Mirror the REAL run flow for the DG4.1 counter (CONC-A1 regression): countfile init
+    (main() preamble) → _record_deploy (deploy_app fires it BEFORE the app lock) → acquire
+    the app lock → wait for `gate` (file path; '' = no wait) → read + remove own countfile.
+    Two of these on the SAME domain must each see COUNT 1 and never lose their file."""
+    import run_recipe_ci
+
+    countfile = run_recipe_ci._run_state_path("deploys")
+    with open(countfile, "w") as f:
+        f.write("0")
+    os.environ["CCCI_DEPLOY_COUNT_FILE"] = countfile
+    lifecycle._record_deploy()  # pre-lock, exactly like lifecycle.deploy_app()
+    mark("PRELOCK")
+    lifecycle.acquire_app_lock(domain)
+    mark("ACQUIRED")
+    if gate:
+        deadline = time.time() + 15
+        while not os.path.exists(gate) and time.time() < deadline:
+            time.sleep(0.05)
+    try:
+        with open(countfile) as f:
+            n = int(f.read().strip() or "0")
+        os.remove(countfile)
+        mark(f"COUNT {n}")
+    except FileNotFoundError:
+        mark("COUNT_FILE_MISSING")
+
+
+if __name__ == "__main__":
+    cmd, *args = sys.argv[1:]
+    {
+        "hold": cmd_hold,
+        "hold-with-child": cmd_hold_with_child,
+        "guarded": cmd_guarded,
+        "wrapper": cmd_wrapper,
+        "orphan-probe": cmd_orphan_probe,
+        "fetch-checkout": cmd_fetch_checkout,
+        "deploy-count-run": cmd_deploy_count_run,
+    }[cmd](*args)
--- a/tests/concurrency/test_abra_dir.py
+++ b/tests/concurrency/test_abra_dir.py
@ -0,0 +1,175 @@
+"""Per-run ABRA_DIR isolation (concurrency-restructure plan, cases 17-19). Real directories,
+real symlinks, real git — abra itself is replaced by a recording stub where a CLI call is
+involved (case 17), because these cases test OUR dir/env plumbing, not abra."""
+
+from __future__ import annotations
+
+import os
+import stat
+import subprocess
+import sys
+
+sys.path.insert(0, os.path.dirname(__file__))
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
+import run_recipe_ci  # noqa: E402
+from concutil import wait_marker  # noqa: E402
+from harness import abra  # noqa: E402
+
+RECIPE = "fakerecipe"
+
+
+def _git(cwd, *args):
+    subprocess.run(
+        ["git", "-c", "user.email=t@t", "-c", "user.name=t", *args],
+        cwd=cwd,
+        check=True,
+        capture_output=True,
+    )
+
+
+def _make_fake_home(tmp_path):
+    """A fake $HOME with a canonical ~/.abra: servers/default + catalogue dirs, and a recipe git
+    repo with two tags whose data.txt differs (v1 -> 'one', v2 -> 'two', HEAD at v2)."""
+    home = tmp_path / "home"
+    (home / ".abra" / "servers" / "default").mkdir(parents=True)
+    (home / ".abra" / "catalogue").mkdir(parents=True)
+    repo = home / ".abra" / "recipes" / RECIPE
+    repo.mkdir(parents=True)
+    _git(repo, "init", "-q")
+    (repo / "data.txt").write_text("one\n")
+    _git(repo, "add", "data.txt")
+    _git(repo, "commit", "-qm", "v1")
+    _git(repo, "tag", "v1")
+    (repo / "data.txt").write_text("two\n")
+    _git(repo, "add", "data.txt")
+    _git(repo, "commit", "-qm", "v2")
+    _git(repo, "tag", "v2")
+    return home
+
+
+def test_17_per_run_dir_built_and_exported_before_abra(tmp_path, monkeypatch):
+    """Case 17: setup_run_abra_dir builds the per-run dir correctly (servers/catalogue symlinks
+    resolve to the canonical tree, recipes/ empty + writable) and $ABRA_DIR is exported before
+    the first abra call — proven by a stub `abra` on PATH that records the env it saw."""
+    home = _make_fake_home(tmp_path)
+    monkeypatch.setenv("HOME", str(home))
+    monkeypatch.setenv("CCCI_RUNS_DIR", str(tmp_path / "runs"))
+    monkeypatch.setenv("DRONE_BUILD_NUMBER", "777")
+    monkeypatch.setenv("ABRA_DIR", "sentinel-to-be-overwritten")  # so monkeypatch restores it
+
+    d = run_recipe_ci.setup_run_abra_dir()
+    assert d == str(tmp_path / "runs" / "777" / "abra")
+    assert os.environ["ABRA_DIR"] == d
+    assert os.readlink(os.path.join(d, "servers")) == str(home / ".abra" / "servers")
+    assert os.readlink(os.path.join(d, "catalogue")) == str(home / ".abra" / "catalogue")
+    # symlinks RESOLVE (targets exist) and recipes/ is empty + writable
+    assert os.path.isdir(os.path.join(d, "servers", "default"))
+    assert os.path.isdir(os.path.join(d, "catalogue"))
+    assert os.listdir(os.path.join(d, "recipes")) == []
+    probe = os.path.join(d, "recipes", ".write-probe")
+    open(probe, "w").close()
+    os.remove(probe)
+    # idempotent re-entry (Drone build-number retry): must not raise on existing symlinks
+    assert run_recipe_ci.setup_run_abra_dir() == d
+
+    # stub abra records $ABRA_DIR at call time; fetch_recipe's catalogue branch invokes it
+    stub_dir = tmp_path / "bin"
+    stub_dir.mkdir()
+    log = tmp_path / "abra-env.log"
+    stub = stub_dir / "abra"
+    stub.write_text(f'#!/bin/sh\necho "$ABRA_DIR" >> {log}\nexit 0\n')
+    stub.chmod(stub.stat().st_mode | stat.S_IEXEC)
+    monkeypatch.setenv("PATH", f"{stub_dir}{os.pathsep}{os.environ['PATH']}")
+    monkeypatch.delenv("CCCI_SKIP_FETCH", raising=False)
+    run_recipe_ci.fetch_recipe(RECIPE, None, None)
+    assert log.read_text().strip() == d, "abra was called without the per-run ABRA_DIR exported"
+
+
+def test_18_concurrent_same_recipe_fetch_no_cross_talk(tmp_path, monkeypatch, pool):
+    """Case 18: two CONCURRENT fetch+checkout flows of the SAME recipe into different ABRA_DIRs
+    produce two correct, divergent trees (v1 vs v2) — the old shared-tree corruption scenario,
+    now structurally safe with no lock. The canonical staged clone is untouched."""
+    home = _make_fake_home(tmp_path)
+    canonical_repo = home / ".abra" / "recipes" / RECIPE
+    head_before = subprocess.run(
+        ["git", "-C", canonical_repo, "rev-parse", "HEAD"], capture_output=True, text=True
+    ).stdout.strip()
+
+    runs = {}
+    for name, ref in (("runA", "v1"), ("runB", "v2")):
+        abra_dir = tmp_path / name / "abra"
+        abra_dir.mkdir(parents=True)
+        _, out = pool.spawn(
+            "fetch-checkout",
+            RECIPE,
+            ref,
+            env_extra={
+                "HOME": str(home),
+                "ABRA_DIR": str(abra_dir),
+                "CCCI_SKIP_FETCH": "1",
+            },
+        )
+        runs[name] = (out, ref, abra_dir)
+
+    expect = {"v1": "one", "v2": "two"}
+    for name, (out, ref, abra_dir) in runs.items():
+        line = wait_marker(out, "RESULT", timeout=30)
+        assert line, f"{name} never produced a RESULT"
+        _, head, content = line.split()
+        assert content == expect[ref], f"{name}@{ref}: tree content {content!r}"
+        tree = abra_dir / "recipes" / RECIPE
+        assert (tree / "data.txt").read_text().strip() == expect[ref]
+        assert (
+            head
+            == subprocess.run(
+                ["git", "-C", tree, "rev-parse", "HEAD"], capture_output=True, text=True
+            ).stdout.strip()
+        )
+
+    # the two trees genuinely diverge AND the canonical staged clone is untouched
+    a = (runs["runA"][2] / "recipes" / RECIPE / "data.txt").read_text()
+    b = (runs["runB"][2] / "recipes" / RECIPE / "data.txt").read_text()
+    assert a != b
+    head_after = subprocess.run(
+        ["git", "-C", canonical_repo, "rev-parse", "HEAD"], capture_output=True, text=True
+    ).stdout.strip()
+    assert head_after == head_before, "canonical clone must not be touched by per-run fetches"
+
+
+def test_19_env_written_through_servers_symlink_lands_canonical(tmp_path, monkeypatch):
+    """Case 19: an app .env written through the per-run servers/ symlink (what abra does under
+    $ABRA_DIR) lands in the CANONICAL shared path — so janitor discovery and every
+    expanduser('~/.abra/servers/...') reader keep working unchanged."""
+    home = _make_fake_home(tmp_path)
+    monkeypatch.setenv("HOME", str(home))
+    monkeypatch.setenv("CCCI_RUNS_DIR", str(tmp_path / "runs"))
+    monkeypatch.setenv("DRONE_BUILD_NUMBER", "778")
+    monkeypatch.setenv("ABRA_DIR", "sentinel-to-be-overwritten")
+    d = run_recipe_ci.setup_run_abra_dir()
+
+    domain = "test-abc123.ci.commoninternet.net"
+    via_symlink = os.path.join(d, "servers", "default", f"{domain}.env")
+    with open(via_symlink, "w") as f:
+        f.write("TYPE=fakerecipe:1.0.0\nDOMAIN=placeholder\n")
+
+    canonical = home / ".abra" / "servers" / "default" / f"{domain}.env"
+    assert canonical.is_file(), ".env written via the symlink must land in the canonical path"
+    # the canonical-path readers/writers (abra.env_get/env_set use ~/.abra) see the same file
+    assert abra.env_get(domain, "TYPE") == "fakerecipe:1.0.0"
+    abra.env_set(domain, "DOMAIN", domain)
+    with open(via_symlink) as f:
+        assert f"DOMAIN={domain}" in f.read()
+
+
+def test_18b_run_id_manual_fallback_is_per_process(tmp_path, monkeypatch):
+    """Companion to case 18: two concurrent MANUAL runs (no DRONE_BUILD_NUMBER) must not share an
+    abra dir either — the manual fallback is pid-suffixed."""
+    home = _make_fake_home(tmp_path)
+    monkeypatch.setenv("HOME", str(home))
+    monkeypatch.setenv("CCCI_RUNS_DIR", str(tmp_path / "runs"))
+    monkeypatch.delenv("DRONE_BUILD_NUMBER", raising=False)
+    monkeypatch.delenv("CCCI_APP_DOMAIN", raising=False)
+    monkeypatch.delenv("CCCI_RUN_ID", raising=False)
+    monkeypatch.setenv("ABRA_DIR", "sentinel-to-be-overwritten")
+    d = run_recipe_ci.setup_run_abra_dir()
+    assert f"manual-{os.getpid()}" in d
--- a/tests/concurrency/test_janitor.py
+++ b/tests/concurrency/test_janitor.py
@ -0,0 +1,189 @@
+"""Janitor / flock-probe semantics (concurrency-restructure plan, cases 5-12).
+
+The janitor runs IN-PROCESS with its discovery monkeypatched (candidates injected via a stubbed
+abra.app_ls + empty docker sweep) and teardown_app stubbed to record calls — but the LOCKS are
+real kernel flocks, held by real helper subprocesses where a live owner is needed."""
+
+from __future__ import annotations
+
+import os
+import sys
+import threading
+import time
+
+sys.path.insert(0, os.path.dirname(__file__))
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
+from concutil import DOMAIN, lock_state, wait_marker  # noqa: E402
+from harness import lifecycle  # noqa: E402
+
+
+def _inject_candidates(monkeypatch, domains):
+    """Point janitor discovery at exactly `domains`: abra lists them, docker sweep is empty.
+    teardown_app is stubbed to a recorder; returns the calls list."""
+    calls = []
+    monkeypatch.setattr(lifecycle.abra, "app_ls", lambda: [{"appName": d} for d in domains])
+    monkeypatch.setattr(lifecycle, "_docker_names", lambda kind, stack: [])
+    monkeypatch.setattr(lifecycle, "teardown_app", lambda d, verify=True: calls.append(d))
+    return calls
+
+
+def test_5_orphan_reaped_lockfile_unlinked(lock_dir, pool, monkeypatch):
+    """Case 5: an orphan (lockfile exists, no holder — its run was SIGKILL'd) is reaped exactly
+    once and its lockfile unlinked."""
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    p.kill()
+    p.wait(timeout=10)
+    calls = _inject_candidates(monkeypatch, [DOMAIN])
+    lifecycle.janitor()
+    assert calls == [DOMAIN], f"teardown calls: {calls} (expected exactly one)"
+    assert lock_state(DOMAIN) == "absent", "reaped orphan's lockfile must be unlinked"
+
+
+def test_6_live_run_never_reaped(lock_dir, pool, monkeypatch, capsys):
+    """Case 6: a held lock (live helper) is never reaped and is logged as live."""
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    calls = _inject_candidates(monkeypatch, [DOMAIN])
+    lifecycle.janitor()
+    assert calls == []
+    assert "live concurrent run" in capsys.readouterr().out
+    assert lock_state(DOMAIN) == "held"
+
+
+def test_7_new_run_blocks_until_reap_finishes(lock_dir, pool, monkeypatch):
+    """Case 7: the janitor reaps WHILE HOLDING the probe lock, so a new run of the same domain
+    blocks in acquire_app_lock until the reap completes — no window where a fresh app coexists
+    with a half-reaped one."""
+    # Make an orphan.
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    p.kill()
+    p.wait(timeout=10)
+
+    state = {"teardown_end": None, "acquirer_out": None}
+
+    def slow_teardown(domain, verify=True):
+        # While the janitor holds the probe lock mid-reap, a new run starts acquiring.
+        _, aout = pool.spawn("hold", DOMAIN)
+        state["acquirer_out"] = aout
+        time.sleep(2.0)
+        state["teardown_end"] = time.time()
+
+    monkeypatch.setattr(lifecycle.abra, "app_ls", lambda: [{"appName": DOMAIN}])
+    monkeypatch.setattr(lifecycle, "_docker_names", lambda kind, stack: [])
+    monkeypatch.setattr(lifecycle, "teardown_app", slow_teardown)
+    lifecycle.janitor()
+
+    line = wait_marker(state["acquirer_out"], "ACQUIRED", timeout=15)
+    assert line, "new run never acquired after the reap"
+    acquired_ts = float(line.split()[1])
+    assert (
+        acquired_ts >= state["teardown_end"]
+    ), f"new run acquired at {acquired_ts} BEFORE the reap finished at {state['teardown_end']}"
+    # The new run must hold a lock the next probe can SEE (fresh inode at the path).
+    assert lock_state(DOMAIN) == "held"
+
+
+def test_8_two_janitors_exactly_one_reaps(lock_dir, pool, monkeypatch):
+    """Case 8: two concurrent janitors arbitrate on the probe flock — exactly one reaps (the
+    other sees 'held' and leaves). Teardown is slowed so the runs genuinely overlap."""
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    p.kill()
+    p.wait(timeout=10)
+
+    calls = []
+    calls_lock = threading.Lock()
+
+    def slow_teardown(domain, verify=True):
+        with calls_lock:
+            calls.append(domain)
+        time.sleep(2.0)
+
+    monkeypatch.setattr(lifecycle.abra, "app_ls", lambda: [{"appName": DOMAIN}])
+    monkeypatch.setattr(lifecycle, "_docker_names", lambda kind, stack: [])
+    monkeypatch.setattr(lifecycle, "teardown_app", slow_teardown)
+
+    barrier = threading.Barrier(2)
+
+    def run_janitor():
+        barrier.wait()
+        lifecycle.janitor()
+
+    t1, t2 = threading.Thread(target=run_janitor), threading.Thread(target=run_janitor)
+    t1.start(), t2.start()
+    t1.join(timeout=30), t2.join(timeout=30)
+    assert calls == [DOMAIN], f"expected exactly one reap, got {calls}"
+    assert lock_state(DOMAIN) == "absent"
+
+
+def test_9_reboot_lockfile_absent_reaped_immediately(lock_dir, monkeypatch):
+    """Case 9: post-reboot simulation — the app exists but its lockfile is gone (/run/lock is
+    tmpfs). The probe trivially acquires -> immediate reap, NO age threshold (improvement over
+    the old 2h fallback)."""
+    assert lock_state(DOMAIN) == "absent"
+    calls = _inject_candidates(monkeypatch, [DOMAIN])
+    t0 = time.time()
+    lifecycle.janitor()
+    assert calls == [DOMAIN]
+    assert time.time() - t0 < 5, "reap must be immediate (no age wait)"
+
+
+def test_10_long_held_lock_flagged_never_stolen(lock_dir, pool, monkeypatch, capsys):
+    """Case 10: a lock held with mtime older than 120min is flagged as a possible leaked run —
+    and NOT reaped (never steal a held lock)."""
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    path = lifecycle._app_lock_path(DOMAIN)  # noqa: SLF001
+    backdate = time.time() - (130 * 60)
+    os.utime(path, (backdate, backdate))
+    calls = _inject_candidates(monkeypatch, [DOMAIN])
+    lifecycle.janitor()
+    assert calls == []
+    out_text = capsys.readouterr().out
+    assert "possible leaked run" in out_text and "lslocks" in out_text
+    assert lock_state(DOMAIN) == "held"
+
+
+def test_11_warm_canonical_names_never_probed(lock_dir, monkeypatch):
+    """Case 11: RUN_APP_RE allowlist — warm/canonical-shaped names never become candidates, so
+    they are never probed (no lockfile is even created for them) and never reaped."""
+    warmish = [
+        "warm-keycloak.ci.commoninternet.net",
+        "keycloak.ci.commoninternet.net",
+        "warm-hedgedoc.ci.commoninternet.net",
+        "drone.ci.commoninternet.net",
+    ]
+    calls = []
+    monkeypatch.setattr(lifecycle.abra, "app_ls", lambda: [{"appName": d} for d in warmish])
+    monkeypatch.setattr(
+        lifecycle,
+        "_docker_names",
+        lambda kind, stack: ["warm-keycloak_ci_commoninternet_net_app"]
+        if kind == "service"
+        else [],
+    )
+    monkeypatch.setattr(lifecycle, "teardown_app", lambda d, verify=True: calls.append(d))
+    lifecycle.janitor()
+    assert calls == []
+    lockdir = os.environ["CCCI_APP_LOCK_DIR"]
+    assert [
+        f for f in os.listdir(lockdir) if f.startswith("cc-ci-app-")
+    ] == [], "janitor must not create lockfiles for non-run-app names"
+
+
+def test_12_degrades_safely_on_bad_lockfile_and_missing_dir(lock_dir, monkeypatch, capsys):
+    """Case 12: a garbled/unopenable lockfile (here: a DIRECTORY at the lockfile path) is skipped
+    with a log line; a missing lock dir doesn't crash the janitor either. Never a crash."""
+    path = lifecycle._app_lock_path(DOMAIN)  # noqa: SLF001
+    os.makedirs(path)  # open(path, "a") -> IsADirectoryError (an OSError)
+    calls = _inject_candidates(monkeypatch, [DOMAIN])
+    lifecycle.janitor()  # must not raise
+    assert calls == []
+    assert "skipping" in capsys.readouterr().out
+
+    os.rmdir(path)
+    monkeypatch.setenv("CCCI_APP_LOCK_DIR", os.path.join(os.environ["CCCI_APP_LOCK_DIR"], "gone"))
+    lifecycle.janitor()  # missing dir: probe open fails -> skip; tidy glob -> empty. No crash.
+    assert calls == []
--- a/tests/concurrency/test_lifetime.py
+++ b/tests/concurrency/test_lifetime.py
@ -0,0 +1,82 @@
+"""Lifetime hardening (concurrency-restructure plan, cases 13-16): the REAL prctl/signal/alarm
+guards installed by helper subprocesses; tests assert teardown ran, exit was non-zero, and the
+lock was released."""
+
+from __future__ import annotations
+
+import os
+import signal
+import sys
+
+sys.path.insert(0, os.path.dirname(__file__))
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
+from concutil import (  # noqa: E402
+    DOMAIN,
+    wait_lock_state,
+    wait_marker,
+    wait_pid_gone,
+)
+
+
+def test_13_pdeathsig_parent_kill_terms_harness(lock_dir, pool):
+    """Case 13: wrapper-parent spawns a guarded harness-child; the parent is SIGKILL'd (the
+    harness gets no courtesy signal) -> the kernel's PDEATHSIG TERMs the child, its teardown
+    funnel runs, it exits, and the lock is released."""
+    p, out = pool.spawn("wrapper", DOMAIN)
+    line = wait_marker(out, "WRAPPED")
+    assert line, "wrapper never spawned its child"
+    child_pid = int(line.split()[1])
+    pool.track_pid(child_pid)
+    assert wait_marker(out, "READY"), "guarded child never got ready"
+
+    p.kill()  # parent dies WITHOUT signalling the child — only PDEATHSIG can save us
+    p.wait(timeout=10)
+    assert wait_pid_gone(child_pid), "guarded child must exit on parent death (PDEATHSIG)"
+    assert wait_marker(out, "TEARDOWN", timeout=5), "teardown funnel did not run"
+    assert wait_lock_state(DOMAIN, "free") == "free"
+
+
+def test_14_already_orphaned_helper_refuses_to_run(lock_dir, pool):
+    """Case 14 (ppid race): a helper whose parent died BEFORE the prctl was armed (it starts
+    already reparented to pid 1) must refuse to run — PDEATHSIG would never fire for it."""
+    # Spawn an intermediate parent that forks orphan-probe and exits immediately.
+    import subprocess
+
+    out = os.path.join(pool.out_dir, "orphan.out")
+    intermediate = (
+        "import subprocess, sys, os; "
+        "subprocess.Popen([sys.executable, os.environ['CCCI_HELPERS'], 'orphan-probe']); "
+    )
+    env = dict(
+        os.environ,
+        CCCI_HELPER_OUT=out,
+        CCCI_HELPERS=os.path.join(os.path.dirname(__file__), "helpers.py"),
+    )
+    subprocess.run([sys.executable, "-c", intermediate], env=env, timeout=15, check=True)
+    line = wait_marker(out, "REFUSED", timeout=20)
+    assert line, "orphaned helper did not refuse to run (or never reparented to pid 1)"
+
+
+def test_15_deadline_alarm_fires_teardown_and_releases(lock_dir, pool):
+    """Case 15: the self-deadline (alarm). A guarded helper with a 2s deadline tears down via
+    the funnel (finally: ran), exits NON-zero, and its lock is released."""
+    p, out = pool.spawn("guarded", DOMAIN, "2")
+    assert wait_marker(out, "READY")
+    rc = p.wait(timeout=20)
+    assert rc != 0, f"deadline exit must be non-zero (got {rc})"
+    assert rc == 128 + signal.SIGALRM, f"expected 142 (128+SIGALRM), got {rc}"
+    assert wait_marker(out, "TEARDOWN", timeout=5), "teardown funnel did not run on deadline"
+    assert wait_lock_state(DOMAIN, "free") == "free"
+
+
+def test_16_sigterm_runs_teardown_funnel_and_releases(lock_dir, pool):
+    """Case 16: SIGTERM (drone cancel path) -> the finally: teardown funnel runs, exit is
+    non-zero, lock released."""
+    p, out = pool.spawn("guarded", DOMAIN, "3600")
+    assert wait_marker(out, "READY")
+    p.send_signal(signal.SIGTERM)
+    rc = p.wait(timeout=20)
+    assert rc != 0, f"SIGTERM exit must be non-zero (got {rc})"
+    assert rc == 128 + signal.SIGTERM, f"expected 143 (128+SIGTERM), got {rc}"
+    assert wait_marker(out, "TEARDOWN", timeout=5), "teardown funnel did not run on SIGTERM"
+    assert wait_lock_state(DOMAIN, "free") == "free"
--- a/tests/concurrency/test_locks.py
+++ b/tests/concurrency/test_locks.py
@ -0,0 +1,85 @@
+"""Lock fundamentals (concurrency-restructure plan, cases 1-4). Real kernel flocks held by real
+subprocesses — nothing mocked."""
+
+from __future__ import annotations
+
+import fcntl
+import os
+import sys
+import time
+
+sys.path.insert(0, os.path.dirname(__file__))
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
+from concutil import (  # noqa: E402
+    DOMAIN,
+    lock_state,
+    wait_lock_state,
+    wait_marker,
+)
+from harness import lifecycle  # noqa: E402
+
+
+def test_1_sigkill_releases_lock(lock_dir, pool):
+    """Case 1: acquire -> holder SIGKILL'd -> lock immediately acquirable (kernel auto-release).
+    The exact property the old pidfile registry approximated with /proc checks."""
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED"), "holder never acquired"
+    assert lock_state(DOMAIN) == "held"
+    p.kill()
+    p.wait(timeout=10)
+    assert wait_lock_state(DOMAIN, "free") == "free"
+
+
+def test_2_nb_probe_held_vs_unheld(lock_dir, pool):
+    """Case 2: LOCK_NB probe raises BlockingIOError against a held lock; succeeds when unheld."""
+    p, out = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    path = lifecycle._app_lock_path(DOMAIN)  # noqa: SLF001
+    with open(path, "a") as f:
+        try:
+            fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)
+            raise AssertionError("LOCK_NB succeeded against a held lock")
+        except BlockingIOError:
+            pass
+    p.kill()
+    p.wait(timeout=10)
+    assert wait_lock_state(DOMAIN, "free") == "free"
+    with open(path, "a") as f:
+        fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB)  # must not raise now
+
+
+def test_3_lock_fd_not_inherited_by_children(lock_dir, pool):
+    """Case 3 (PEP 446): the holder spawns a subprocess child, the holder dies, the child lives —
+    and the lock is STILL released (the child never inherited the lock fd). This is what makes
+    'held lock == live HARNESS owner' sound even though runs spawn abra/docker/pytest children."""
+    p, out = pool.spawn("hold-with-child", DOMAIN)
+    assert wait_marker(out, "ACQUIRED")
+    child_line = wait_marker(out, "CHILD")
+    assert child_line, "holder never reported its child pid"
+    child_pid = int(child_line.split()[1])
+    pool.track_pid(child_pid)
+    p.kill()
+    p.wait(timeout=10)
+    assert os.path.exists(f"/proc/{child_pid}"), "child should outlive the holder"
+    assert (
+        wait_lock_state(DOMAIN, "free") == "free"
+    ), "lock must release on holder death even with a live child (PEP 446 non-inheritable fd)"
+
+
+def test_4_second_acquire_blocks_until_first_exits(lock_dir, pool):
+    """Case 4: a second same-domain acquire blocks until the first holder exits — the
+    double-!testme serialisation property."""
+    p1, out1 = pool.spawn("hold", DOMAIN)
+    assert wait_marker(out1, "ACQUIRED")
+    p2, out2 = pool.spawn("hold", DOMAIN)
+    # p2 must NOT acquire while p1 holds.
+    time.sleep(1.5)
+    assert wait_marker(out2, "ACQUIRED", timeout=0.1) is None, "second acquire did not block"
+    t_kill = time.time()
+    p1.kill()
+    p1.wait(timeout=10)
+    line = wait_marker(out2, "ACQUIRED", timeout=15)
+    assert line, "second acquire never completed after first holder exited"
+    acquired_ts = float(line.split()[1])
+    assert acquired_ts >= t_kill - 0.05, "second holder acquired before the first exited"
+    assert lock_state(DOMAIN) == "held"
--- a/tests/concurrency/test_run_state.py
+++ b/tests/concurrency/test_run_state.py
@ -0,0 +1,79 @@
+"""Run-scoped state files — M2(c) live-verify regression (not one of the 19 plan cases).
+
+The four CCCI state files (deploys countfile, opstate, deps, depskip) must be keyed by
+run id + harness pid, NEVER by app domain: a second run of the SAME domain executes its
+main() preamble (state-file init, deploy_app's _record_deploy) BEFORE it blocks at the
+app lock, so domain-keyed files in the shared tempdir get reset/removed underneath the
+live first run. Observed live (builds 279/281): false DG4.1 deploy-count=2 in run 1,
+countfile FileNotFoundError crash in run 2. Children never re-derive these paths — they
+receive them via the CCCI_*_FILE env vars, so per-process uniqueness is sufficient.
+"""
+
+from __future__ import annotations
+
+import os
+import sys
+import tempfile
+
+sys.path.insert(0, os.path.dirname(__file__))
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
+import run_recipe_ci  # noqa: E402
+from concutil import wait_marker  # noqa: E402
+
+DOMAIN = "fake-abc123.ci.commoninternet.net"
+
+
+def test_20_state_paths_keyed_by_run_and_pid_never_by_domain(monkeypatch):
+    domain = "immi-ad3e33.ci.commoninternet.net"
+    monkeypatch.setenv("CCCI_APP_DOMAIN", domain)
+
+    monkeypatch.setenv("DRONE_BUILD_NUMBER", "279")
+    p279 = run_recipe_ci._run_state_path("deploys")
+    monkeypatch.setenv("DRONE_BUILD_NUMBER", "281")
+    p281 = run_recipe_ci._run_state_path("deploys")
+
+    # the double-!testme invariant: two runs (same domain) share NO state file
+    assert p279 != p281
+    # keyed by run id + pid, under the tempdir
+    base = os.path.basename(p279)
+    assert base == f"ccci-deploys-279-{os.getpid()}"
+    assert os.path.dirname(p279) == tempfile.gettempdir()
+    # the app domain must not appear in the path at all
+    assert domain not in p279 and domain not in p281
+
+
+def test_20c_same_domain_runs_each_keep_their_own_count(tmp_path, lock_dir, pool):
+    """The live CONC-A1 interleaving, with REAL processes + the REAL lock and counter code:
+    run A holds the app lock; run B (same domain) fires its pre-lock _record_deploy and
+    blocks; A then reads its counter — must still be 1 (not polluted by B) — and removes
+    its own file; B acquires and must find ITS file intact (no FileNotFoundError)."""
+    gate = tmp_path / "gate"
+    env_a = {"TMPDIR": str(tmp_path), "DRONE_BUILD_NUMBER": "9001"}
+    env_b = {"TMPDIR": str(tmp_path), "DRONE_BUILD_NUMBER": "9002"}
+
+    pa, out_a = pool.spawn("deploy-count-run", DOMAIN, str(gate), env_extra=env_a)
+    assert wait_marker(out_a, "ACQUIRED")
+    pb, out_b = pool.spawn("deploy-count-run", DOMAIN, "", env_extra=env_b)
+    # B's main()-preamble + pre-lock increment have fired; B is now blocked on the app lock
+    assert wait_marker(out_b, "PRELOCK")
+    assert wait_marker(out_b, "ACQUIRED", timeout=1.0) is None  # still serialised behind A
+
+    gate.touch()  # let A read its counter only AFTER B's pre-lock work landed
+    line_a = wait_marker(out_a, "COUNT")
+    assert line_a is not None and line_a.strip() == "COUNT 1", line_a  # not 2: B didn't pollute A
+    pa.wait(timeout=15)
+
+    line_b = wait_marker(out_b, "COUNT")
+    assert (
+        line_b is not None and line_b.strip() == "COUNT 1"
+    ), line_b  # B's file survived A's remove
+    pb.wait(timeout=15)
+
+
+def test_20b_manual_runs_distinct_via_pid(monkeypatch):
+    # no DRONE_BUILD_NUMBER and no domain/run-id env → run_id() falls back to "manual";
+    # the pid suffix still separates two concurrent hand-runs of the same domain.
+    for var in ("DRONE_BUILD_NUMBER", "CCCI_APP_DOMAIN", "CCCI_RUN_ID"):
+        monkeypatch.delenv(var, raising=False)
+    p = run_recipe_ci._run_state_path("opstate")
+    assert os.path.basename(p) == f"ccci-opstate-manual-{os.getpid()}"
--- a/tests/conftest.py
+++ b/tests/conftest.py
@ -13,32 +13,8 @@ import sys
 import pytest

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "runner"))
-from harness import deps as deps_mod, lifecycle, naming  # noqa: E402
-
-
-def _short(s: str, n: int = 8) -> str:
-    return "".join(c for c in s if c.isalnum())[:n] or "local"
-
-
-def _recipe_meta(recipe: str) -> dict:
-    """Optional per-recipe config so enrolling a recipe needs NO shared-harness change (D5).
-    A recipe may ship tests/<recipe>/recipe_meta.py with any of: HEALTH_PATH (str),
-    HEALTH_OK (tuple of status codes), DEPLOY_TIMEOUT (int), HTTP_TIMEOUT (int)."""
-    path = os.path.join(os.path.dirname(__file__), recipe, "recipe_meta.py")
-    meta = {
-        "HEALTH_PATH": "/",
-        "HEALTH_OK": (200, 301, 302),
-        "DEPLOY_TIMEOUT": 600,
-        "HTTP_TIMEOUT": 300,
-    }
-    if os.path.exists(path):
-        ns: dict = {}
-        with open(path) as fh:
-            exec(compile(fh.read(), path, "exec"), ns)  # noqa: S102 (trusted, in-repo)
-        for k in meta:
-            if k in ns:
-                meta[k] = ns[k]
-    return meta
+from harness import deps as deps_mod  # noqa: E402
+from harness import meta as meta_mod  # noqa: E402


@pytest.fixture(scope="session")
@ -47,18 +23,10 @@ def recipe() -> str:


@pytest.fixture(scope="session")
-def app_domain(recipe) -> str:
-    # Docker swarm config/secret names = <stackname>_<res>_<ver> must be <= 64 chars, and
-    # stackname is the sanitized domain. ".ci.commoninternet.net" alone is 22 chars, so the
-    # subdomain label must stay short. Use <recipe[:4]>-<6hex(recipe|pr|ref)> — unique per run,
-    # collision-safe across recipes (full recipe in the hash), readable context lives in the
-    # Drone build params + PR comment. (Deviation from plan §4.0 long name; see DECISIONS.md.)
-    return naming.app_domain(recipe, os.environ.get("PR", "0"), os.environ.get("REF"))
-
-
-@pytest.fixture(scope="session")
-def meta(recipe) -> dict:
-    return _recipe_meta(recipe)
+def meta(recipe):
+    """The recipe's FULL validated customization (RecipeMeta, attribute access) via the single
+    loader (rcust P1 — previously this fixture saw only the 4 base keys, spec §8 R3)."""
+    return meta_mod.load(recipe)


@pytest.fixture(scope="session")
@ -72,32 +40,55 @@ def live_app() -> str:
    return domain


-@pytest.fixture(scope="session")
-def deps_apps() -> dict[str, str]:
-    """Phase 2 Q2.3 dependency-resolver contract (refined operator-2026-05-28 SSO-dep plan §1):
-    when a recipe declares `DEPS = [...]` in its `recipe_meta.py`, the orchestrator deploys each
-    dep AFTER the generic tiers (between RESTORE and CUSTOM) and persists their per-run identity
-    + SSO creds to `$CCCI_DEPS_FILE`. Tests access the dep's per-run domain via this fixture.
-    For full SSO creds (realm/client/secret/admin) use the `deps_creds` fixture instead.
+@pytest.fixture
+def op_state() -> dict:
+    """The orchestrator's run-scoped op context (rcust P4): versions, artifact paths — written to
+    `$CCCI_OP_STATE_FILE` after each lifecycle op (e.g. `{"upgrade": {"before": {...},
+    "head_ref": ...}, "backup": {"snapshot_id": ...}}`). Overlay tests read op facts from here
+    instead of hand-parsing env/JSON. Skips with a clear reason outside an orchestrator run."""
+    import json

-    Returns `{dep_recipe: domain}` (str→str). Empty when no deps declared OR deps-not-ready."""
+    path = os.environ.get("CCCI_OP_STATE_FILE")
+    if not path:
+        pytest.skip(
+            "CCCI_OP_STATE_FILE not set — op_state is only available under the orchestrator"
+        )
+    if not os.path.exists(path):
+        pytest.skip(f"op-state file missing ({path}) — orchestrator has not performed an op yet")
+    try:
+        with open(path) as f:
+            return json.load(f)
+    except ValueError:
+        pytest.skip(f"op-state file unreadable/not JSON ({path})")
+
+
+class _DepEntry(dict):
+    """One provisioned dep (full creds dict) with attribute sugar: `entry.domain`, `entry.realm`,
+    `entry.client_secret`, ... — dict-style access works too (rcust P2d)."""
+
+    def __getattr__(self, name):
+        try:
+            return self[name]
+        except KeyError as e:
+            raise AttributeError(name) from e
+
+
+@pytest.fixture(scope="session")
+def deps() -> dict[str, _DepEntry]:
+    """The recipe's provisioned deps (rcust P2d — consolidates the old `deps_apps`+`deps_creds`
+    pair). When a recipe declares `DEPS = [...]` in its `recipe_meta.py`, the orchestrator
+    provisions each dep BEFORE the single deploy and persists per-run identity + SSO creds to
+    `$CCCI_DEPS_FILE`. `deps["keycloak"]` carries domain/realm/client_id/client_secret/user/
+    password/email/admin_user/admin_password/discovery_url/token_url/... (`.domain` etc. work as
+    attributes). Empty when no deps declared OR deps-not-ready — pair with
+    `@pytest.mark.requires_deps` so the F2-11 skip-report keeps the green signal honest."""
    state = deps_mod.deps_as_dict(deps_mod.load_run_state())
-    return {r: e["domain"] for r, e in state.items() if e.get("domain")}
-
-
-@pytest.fixture(scope="session")
-def deps_creds() -> dict[str, dict]:
-    """Full SSO-creds dict for each declared dep (operator-2026-05-28 SSO-dep plan §1).
-    `deps_creds["keycloak"]` returns the entry written by setup_custom_tests with keys
-    domain/realm/client_id/client_secret/user/password/email/admin_user/admin_password/
-    discovery_url/token_url/.... Use this in `@pytest.mark.requires_deps` tests that need to
-    authenticate via OIDC."""
-    return deps_mod.deps_as_dict(deps_mod.load_run_state())
+    return {r: _DepEntry(e) for r, e in state.items()}


 def pytest_collection_modifyitems(config, items):
    """SSO-dep plan §4: tests marked `@pytest.mark.requires_deps` are skipped with reason
-    `deps-not-ready: <captured-err>` when the orchestrator's setup_custom_tests step failed
+    `deps-not-ready: <captured-err>` when the orchestrator's dep provisioning failed
    (orchestrator sets CCCI_DEPS_READY=0 in env). Non-deps custom tests are unaffected.

    This is failure-isolation per plan §1 — generic tiers cannot break the SSO-marked tests'
@ -130,40 +121,5 @@ def pytest_configure(config):
    """Register the `requires_deps` marker so pytest doesn't warn about it."""
    config.addinivalue_line(
        "markers",
-        "requires_deps: test requires DEPS-declared services + setup_custom_tests success.",
+        "requires_deps: test requires DEPS-declared services + dep provisioning success.",
    )
-
-
-def _wait_healthy(domain, meta):
-    lifecycle.wait_healthy(
-        domain,
-        ok_codes=tuple(meta["HEALTH_OK"]),
-        path=meta["HEALTH_PATH"],
-        deploy_timeout=meta["DEPLOY_TIMEOUT"],
-        http_timeout=meta["HTTP_TIMEOUT"],
-    )
-
-
-@pytest.fixture
-def deployed(recipe, app_domain, meta, request):
-    """Function-scoped: deploy the current/$REF version healthy, guaranteed teardown after.
-    Used by stages that start from current (install/backup)."""
-    version = os.environ.get("VERSION") or None
-    lifecycle.janitor()
-    request.addfinalizer(lambda: lifecycle.teardown_app(app_domain))
-    lifecycle.deploy_app(recipe, app_domain, version=version)
-    _wait_healthy(app_domain, meta)
-    return app_domain
-
-
-@pytest.fixture(scope="session")
-def deployed_app(recipe, app_domain, meta):
-    """Install stage: deploy the recipe and wait until healthy; tear down at session end."""
-    version = os.environ.get("VERSION") or None
-    lifecycle.janitor()  # sweep orphans from crashed runs first
-    try:
-        lifecycle.deploy_app(recipe, app_domain, version=version, secrets=True)
-        _wait_healthy(app_domain, meta)
-        yield app_domain
-    finally:
-        lifecycle.teardown_app(app_domain)
--- a/tests/cryptpad/ops.py
+++ b/tests/cryptpad/ops.py
@ -15,13 +15,13 @@ def _write(domain, val):
    lifecycle.exec_in_app(domain, ["sh", "-c", f"echo {val} > {MARKER}"])


-def pre_upgrade(domain, meta):
-    _write(domain, "upgrade-survives")
+def pre_upgrade(ctx):
+    _write(ctx.domain, "upgrade-survives")


-def pre_backup(domain, meta):
-    _write(domain, "original")
+def pre_backup(ctx):
+    _write(ctx.domain, "original")


-def pre_restore(domain, meta):
-    _write(domain, "mutated")  # diverge so a successful restore is observable
+def pre_restore(ctx):
+    _write(ctx.domain, "mutated")  # diverge so a successful restore is observable
--- a/tests/cryptpad/playwright/test_pad_content_roundtrip.py
+++ b/tests/cryptpad/playwright/test_pad_content_roundtrip.py
@ -26,6 +26,7 @@ Transient `net::ERR_NETWORK_CHANGED` is handled by the shared `goto_with_retry`

 from __future__ import annotations

+import contextlib
 import os
 import sys
 import uuid
@ -39,7 +40,11 @@ def _open_pad(ctx, url):
    bar once CryptPad has created/loaded the fragment-keyed pad (`#/2/pad/edit/<key>/`)."""
    page = ctx.new_page()
    harness_browser.goto_with_retry(
-        page, url, accept_statuses=(200,), goto_timeout_ms=60_000, wait_until="load",
+        page,
+        url,
+        accept_statuses=(200,),
+        goto_timeout_ms=60_000,
+        wait_until="load",
        deadline_seconds=150,
    )
    pad_url = url
@ -53,13 +58,15 @@ def _open_pad(ctx, url):
            pad_url = page.url
            break
        if i == 40:
-            try:
+            with contextlib.suppress(Exception):  # best-effort unstick
                harness_browser.goto_with_retry(
-                    page, url, accept_statuses=(200,), goto_timeout_ms=60_000,
-                    wait_until="load", deadline_seconds=120,
+                    page,
+                    url,
+                    accept_statuses=(200,),
+                    goto_timeout_ms=60_000,
+                    wait_until="load",
+                    deadline_seconds=120,
                )
-            except Exception:  # noqa: BLE001 — best-effort unstick
-                pass
    return page, pad_url


@ -74,18 +81,22 @@ def _ckeditor_frame(page, deadline_polls=90, reload_at=22, reload_url=None):
            if "ckeditor-inner" in f.url:
                return f
        if i == reload_at and reload_url is not None:
-            try:
+            with contextlib.suppress(Exception):  # reload is a best-effort unstick
                harness_browser.goto_with_retry(
-                    page, reload_url, accept_statuses=(200,), goto_timeout_ms=60_000,
-                    wait_until="load", deadline_seconds=120,
+                    page,
+                    reload_url,
+                    accept_statuses=(200,),
+                    goto_timeout_ms=60_000,
+                    wait_until="load",
+                    deadline_seconds=120,
                )
-            except Exception:  # noqa: BLE001 — reload is a best-effort unstick
-                pass
        page.wait_for_timeout(2000)
    return None


-def _poll_any_frame_for_text(page, needle, deadline_polls=120, reload_at=(20, 45, 75, 100), reload_url=None):
+def _poll_any_frame_for_text(
+    page, needle, deadline_polls=120, reload_at=(20, 45, 75, 100), reload_url=None
+):
    """Robust read-back (F2-13): poll EVERY frame's body text for `needle`, returning True as soon as
    it appears. The fresh cold-cache read-back context's deeply-nested CKEditor frame is slow/flaky to
    *attach* by URL (the prior `_ckeditor_frame` wait timed out on the Adversary's cold run), but the
@ -101,13 +112,15 @@ def _poll_any_frame_for_text(page, needle, deadline_polls=120, reload_at=(20, 45
            except Exception:  # noqa: BLE001 — frame not ready / detached; keep polling
                pass
        if reload_url and i in reload_at:
-            try:
+            with contextlib.suppress(Exception):  # best-effort unstick
                harness_browser.goto_with_retry(
-                    page, reload_url, accept_statuses=(200,), goto_timeout_ms=60_000,
-                    wait_until="load", deadline_seconds=120,
+                    page,
+                    reload_url,
+                    accept_statuses=(200,),
+                    goto_timeout_ms=60_000,
+                    wait_until="load",
+                    deadline_seconds=120,
                )
-            except Exception:  # noqa: BLE001 — best-effort unstick
-                pass
        page.wait_for_timeout(2000)
    return False

@ -137,9 +150,9 @@ def test_cryptpad_pad_content_survives_fresh_session(live_app):
            # --- session 1: create the pad + write the marker ---
            ctx1 = browser.new_context(ignore_https_errors=True)
            page, pad_url = _open_pad(ctx1, f"https://{live_app}/pad/")
-            assert "#/2/pad/edit/" in pad_url, (
-                f"CryptPad did not create a fragment-keyed pad URL; got {pad_url!r}"
-            )
+            assert (
+                "#/2/pad/edit/" in pad_url
+            ), f"CryptPad did not create a fragment-keyed pad URL; got {pad_url!r}"
            ck = _ckeditor_frame(page, reload_url=pad_url)
            assert ck is not None, "CKEditor content frame never attached (pad editor not ready)"
            _dismiss_store_modal(page)
@ -148,9 +161,9 @@ def test_cryptpad_pad_content_survives_fresh_session(live_app):
            page.wait_for_timeout(1000)
            body.type(marker, delay=40)
            page.wait_for_timeout(12000)  # let CryptPad encrypt + sync the update to the server
-            assert marker in ck.locator("body").inner_text(), (
-                "marker not present in the editor after typing — type did not land"
-            )
+            assert (
+                marker in ck.locator("body").inner_text()
+            ), "marker not present in the editor after typing — type did not land"
            ctx1.close()

            # --- session 2: FRESH context (no shared storage/localStorage) reads the pad back by URL.
--- a/tests/cryptpad/playwright/test_pad_create.py
+++ b/tests/cryptpad/playwright/test_pad_create.py
@ -51,9 +51,9 @@ def test_cryptpad_spa_renders_with_no_console_errors(live_app):
            title = (page.title() or "").lower()
            body = page.content()
            blower = body.lower()
-            assert "cryptpad" in title or "cryptpad" in blower, (
-                f"CryptPad SPA does not carry brand. title={title!r}, body excerpt: {body[:200]!r}"
-            )
+            assert (
+                "cryptpad" in title or "cryptpad" in blower
+            ), f"CryptPad SPA does not carry brand. title={title!r}, body excerpt: {body[:200]!r}"

            # Canonical CryptPad asset references in the rendered DOM
            canonical = ("/customize/", "/components/", "main.js", "/api/broadcast")
--- a/tests/cryptpad/recipe_meta.py
+++ b/tests/cryptpad/recipe_meta.py
@ -7,9 +7,9 @@ DEPLOY_TIMEOUT = 600
 HTTP_TIMEOUT = 600


-def EXTRA_ENV(domain):
+def EXTRA_ENV(ctx):
    """cryptpad needs a SANDBOX_DOMAIN distinct from the main DOMAIN (it serves user content from a
    separate origin; the web router routes both). Derive a sibling subdomain under the same wildcard
    (covered by the wildcard cert, so no cert work)."""
-    label, _, rest = domain.partition(".")
+    label, _, rest = ctx.domain.partition(".")
    return {"SANDBOX_DOMAIN": f"{label}-sb.{rest}"}
--- a/tests/cryptpad/test_install.py
+++ b/tests/cryptpad/test_install.py
@ -8,7 +8,8 @@ import os
 import sys

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
-from harness import browser as harness_browser, generic, lifecycle  # noqa: E402
+from harness import browser as harness_browser  # noqa: E402
+from harness import generic, lifecycle


 def test_serving_and_content(live_app, meta):
--- a/tests/custom-html-bkp-bad/ops.py
+++ b/tests/custom-html-bkp-bad/ops.py
@ -12,8 +12,8 @@ from harness import lifecycle
 MARKER_PATH = "/usr/share/nginx/html/ci-marker.txt"


-def pre_restore(domain: str, meta: dict) -> None:
+def pre_restore(ctx) -> None:
    """Write 'mutated' to the marker before restore runs. If restore brings back the
    snapshot (which has no marker — never seeded by pre_backup), the marker ends up
    MISSING or 'mutated' after restore → test_restore_returns_state FAILS → restore=RED."""
-    lifecycle.exec_in_app(domain, ["sh", "-c", f"echo mutated > {MARKER_PATH}"])
+    lifecycle.exec_in_app(ctx.domain, ["sh", "-c", f"echo mutated > {MARKER_PATH}"])
--- a/tests/custom-html-bkp-bad/test_backup.py
+++ b/tests/custom-html-bkp-bad/test_backup.py
@ -20,7 +20,9 @@ def test_backup_captures_state(live_app):
    Since custom-html-bkp-bad has no ops.py::pre_backup to seed the marker, this file does NOT
    exist at backup time — exec_in_app returns empty or raises → assertion fails → backup tier RED.
    This models a recipe that declares backup capability but omits the data-seeding hook."""
-    result = lifecycle.exec_in_app(live_app, ["sh", "-c", f"cat {MARKER_PATH} 2>/dev/null || echo MISSING"]).strip()
+    result = lifecycle.exec_in_app(
+        live_app, ["sh", "-c", f"cat {MARKER_PATH} 2>/dev/null || echo MISSING"]
+    ).strip()
    assert result == "original", (
        f"backup did not capture the expected marker at {MARKER_PATH}: got {result!r}. "
        "Expected 'original' (seeded by pre_backup). If the marker is 'MISSING', the pre_backup "
--- a/tests/custom-html-rst-bad/ops.py
+++ b/tests/custom-html-rst-bad/ops.py
@ -11,5 +11,5 @@ from harness import lifecycle
 MARKER_PATH = "/usr/share/nginx/html/ci-marker.txt"


-def pre_restore(domain: str, meta: dict) -> None:
-    lifecycle.exec_in_app(domain, ["sh", "-c", f"echo mutated > {MARKER_PATH}"])
+def pre_restore(ctx) -> None:
+    lifecycle.exec_in_app(ctx.domain, ["sh", "-c", f"echo mutated > {MARKER_PATH}"])
--- a/tests/custom-html-tiny/functional/test_serves_content.py
+++ b/tests/custom-html-tiny/functional/test_serves_content.py
@ -79,9 +79,9 @@ def test_static_file_roundtrip_and_404(live_app):
        # A random non-existent path must 404 — proves real static-file semantics, distinguishing a
        # working server from a 200-everything stub or a mis-routed Traefik fallback.
        miss_status, _ = _get(f"https://{live_app}/ccci-missing-{uuid.uuid4().hex}.txt")
-        assert miss_status == 404, (
-            f"missing path returned {miss_status} (expected 404 — generic 200-returner / mis-route?)"
-        )
+        assert (
+            miss_status == 404
+        ), f"missing path returned {miss_status} (expected 404 — generic 200-returner / mis-route?)"
    finally:
        with contextlib.suppress(OSError):
            os.remove(path)
--- a/tests/custom-html/functional/test_content_roundtrip.py
+++ b/tests/custom-html/functional/test_content_roundtrip.py
@ -15,7 +15,8 @@ import sys
 import uuid

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "..", "runner"))
-from harness import http as harness_http, lifecycle  # noqa: E402
+from harness import http as harness_http  # noqa: E402
+from harness import lifecycle


 def test_content_roundtrip(live_app):
--- a/tests/custom-html/functional/test_content_type_header.py
+++ b/tests/custom-html/functional/test_content_type_header.py
@ -53,9 +53,9 @@ def test_content_type_html_and_txt(live_app):
    ct_txt = h_txt.get("content-type", "")

    # nginx default: "text/html" for .html and "text/plain" for .txt (may include "; charset=utf-8")
-    assert ct_html.startswith("text/html"), (
-        f"{html_name} Content-Type={ct_html!r}, expected text/html (nginx MIME config broken?)"
-    )
-    assert ct_txt.startswith("text/plain"), (
-        f"{txt_name} Content-Type={ct_txt!r}, expected text/plain (nginx MIME config broken?)"
-    )
+    assert ct_html.startswith(
+        "text/html"
+    ), f"{html_name} Content-Type={ct_html!r}, expected text/html (nginx MIME config broken?)"
+    assert ct_txt.startswith(
+        "text/plain"
+    ), f"{txt_name} Content-Type={ct_txt!r}, expected text/plain (nginx MIME config broken?)"
--- a/tests/custom-html/ops.py
+++ b/tests/custom-html/ops.py
@ -1,4 +1,4 @@
-"""custom-html — pre-op seed hooks (Phase 1e HC3). The orchestrator runs `pre_<op>(domain, meta)`
+"""custom-html — pre-op seed hooks (Phase 1e HC3). The orchestrator runs `pre_<op>(ctx)`
 BEFORE it performs the op; the matching test_<op>.py asserts the post-op state (assertion-only).

 nginx serves the volume at /usr/share/nginx/html, so the marker file survives an upgrade / a
@ -17,16 +17,16 @@ def _write(domain: str, val: str) -> None:
    lifecycle.exec_in_app(domain, ["sh", "-c", f"echo {val} > {MARKER_PATH}"])


-def pre_upgrade(domain, meta):
+def pre_upgrade(ctx):
    # seed a marker before the upgrade so the overlay can prove the data survives it
-    _write(domain, "upgrade-survives")
+    _write(ctx.domain, "upgrade-survives")


-def pre_backup(domain, meta):
+def pre_backup(ctx):
    # establish a known original state before the backup op captures it
-    _write(domain, "original")
+    _write(ctx.domain, "original")


-def pre_restore(domain, meta):
+def pre_restore(ctx):
    # diverge from the backed-up state so a successful restore (back to "original") is observable
-    _write(domain, "mutated")
+    _write(ctx.domain, "mutated")
--- a/tests/custom-html/test_install.py
+++ b/tests/custom-html/test_install.py
@ -9,7 +9,8 @@ import os
 import sys

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
-from harness import browser as harness_browser, generic  # noqa: E402
+from harness import browser as harness_browser  # noqa: E402
+from harness import generic


 def test_serving_and_content(live_app, meta):
--- a/tests/discourse/functional/_discourse.py
+++ b/tests/discourse/functional/_discourse.py
@ -53,7 +53,7 @@ def mint_admin(domain: str) -> tuple[str, str]:
    cmd = (
        "cd /opt/bitnami/discourse && "
        "RUBY=$(command -v ruby || echo /opt/bitnami/ruby/bin/ruby) && "
-        f"RAILS_ENV=production \"$RUBY\" bin/rails runner \"{_BOOTSTRAP_RB}\""
+        f'RAILS_ENV=production "$RUBY" bin/rails runner "{_BOOTSTRAP_RB}"'
    )
    out = lifecycle.exec_in_app(domain, ["bash", "-c", cmd], service="app", timeout=240)
    key = user = None
@ -63,9 +63,9 @@ def mint_admin(domain: str) -> tuple[str, str]:
            key = line.split("=", 1)[1].strip()
        elif line.startswith("CCCI_API_USER="):
            user = line.split("=", 1)[1].strip()
-    assert key and user, (
-        f"could not bootstrap discourse admin/API key; rails output tail:\n{out[-1000:]}"
-    )
+    assert (
+        key and user
+    ), f"could not bootstrap discourse admin/API key; rails output tail:\n{out[-1000:]}"
    return key, user


--- a/tests/discourse/functional/test_create_topic.py
+++ b/tests/discourse/functional/test_create_topic.py
@ -48,21 +48,23 @@ def test_create_topic_roundtrip(live_app):
        headers=hdrs,
        timeout=60,
    )
-    assert status in (200, 201) and isinstance(body, dict), (
-        f"create topic failed: HTTP {status}, body={body!r}"
-    )
+    assert status in (200, 201) and isinstance(
+        body, dict
+    ), f"create topic failed: HTTP {status}, body={body!r}"
    topic_id = body.get("topic_id")
    assert topic_id, f"create topic returned no topic_id: {body!r}"

    # 4) Read the topic back and assert title + first-post body round-trip.
    status, got = harness_http.http_get(f"{base}/t/{topic_id}.json", headers=hdrs, timeout=30)
-    assert status == 200 and isinstance(got, dict), f"read topic failed: HTTP {status}, body={got!r}"
-    assert got.get("title") == title, (
-        f"topic title did not round-trip: sent {title!r}, got {got.get('title')!r}"
-    )
+    assert status == 200 and isinstance(
+        got, dict
+    ), f"read topic failed: HTTP {status}, body={got!r}"
+    assert (
+        got.get("title") == title
+    ), f"topic title did not round-trip: sent {title!r}, got {got.get('title')!r}"
    posts = (got.get("post_stream") or {}).get("posts") or []
    assert posts, f"topic has no posts on read-back: {got!r}"
    first_cooked = posts[0].get("cooked", "")
-    assert marker in first_cooked, (
-        f"topic body did not round-trip: marker {marker!r} not in first post {first_cooked!r}"
-    )
+    assert (
+        marker in first_cooked
+    ), f"topic body did not round-trip: marker {marker!r} not in first post {first_cooked!r}"
--- a/tests/discourse/functional/test_site_basic.py
+++ b/tests/discourse/functional/test_site_basic.py
@ -20,12 +20,12 @@ def test_site_json_has_discourse_config(live_app):
    status, body = harness_http.retry_http_get(
        f"https://{live_app}/site.json", expect_status=200, max_wait=120, interval=5
    )
-    assert status == 200 and isinstance(body, dict), (
-        f"GET /site.json failed: HTTP {status}, body type={type(body).__name__}"
-    )
+    assert status == 200 and isinstance(
+        body, dict
+    ), f"GET /site.json failed: HTTP {status}, body type={type(body).__name__}"
    # /site.json carries Discourse-specific structure — `categories` (a list) and `groups` are always
    # present in a booted Discourse. A non-Discourse 200 (placeholder page) would not parse to this.
    assert "categories" in body, f"/site.json missing 'categories' key: keys={list(body)[:20]}"
-    assert isinstance(body["categories"], list), (
-        f"/site.json 'categories' not a list: {type(body['categories']).__name__}"
-    )
+    assert isinstance(
+        body["categories"], list
+    ), f"/site.json 'categories' not a list: {type(body['categories']).__name__}"
--- a/tests/discourse/install_steps.sh
+++ b/tests/discourse/install_steps.sh
@ -1,26 +0,0 @@
-#!/usr/bin/env bash
-# discourse — INSTALL-TIME hook (Phase 2 Q4.6). Runs during the install tier AFTER `abra app new` +
-# EXTRA_ENV + `abra app secret generate` and BEFORE the single `abra app deploy`
-# (lifecycle.py::_run_install_steps), with CCCI_RECIPE / CCCI_APP_DOMAIN in env.
-#
-# Purpose: provide the cc-ci re-pin+grace overlay (compose.ccci.yml) to the recipe checkout so the
-# UPGRADE-tier BASE deploy (published 0.7.0+3.3.1, whose compose pins the Docker-Hub-removed
-# `bitnami/discourse:3.3.1` and ships a too-tight 5m start_period) is deployable and can survive the
-# 15-25min Rails cold boot — so upgrade-to-latest can run. See compose.ccci.yml's header for the full
-# rationale. The overlay is referenced by recipe_meta COMPOSE_FILE; it is a cc-ci file (not part of the
-# recipe), so copying it here makes it resolvable. It persists across the later `git checkout <head>`
-# (untracked) so the head deploy also merges it (idempotent — the PR head already re-pins + ships 20m).
-# CHAOS_BASE_DEPLOY=True is set so abra's pinned-deploy clean-tree check doesn't FATA on the overlay.
-set -euo pipefail
-
-: "${CCCI_RECIPE:?missing CCCI_RECIPE}"
-SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
-RECIPE_DIR="${HOME}/.abra/recipes/${CCCI_RECIPE}"
-
-if [ ! -d "$RECIPE_DIR" ]; then
-  echo "  discourse install_steps: recipe dir $RECIPE_DIR missing — cannot provide compose.ccci.yml" >&2
-  exit 1
-fi
-
-cp "$SCRIPT_DIR/compose.ccci.yml" "$RECIPE_DIR/compose.ccci.yml"
-echo "  discourse install_steps: provided compose.ccci.yml (bitnamilegacy re-pin + 20m start_period grace) to recipe checkout (${CCCI_RECIPE})"
--- a/tests/discourse/ops.py
+++ b/tests/discourse/ops.py
@ -15,8 +15,7 @@ from harness import lifecycle  # noqa: E402

 def _psql(domain, sql):
    cmd = (
-        'PGPASSWORD=$(cat /run/secrets/db_password) '
-        f'psql -U discourse -d discourse -tAc "{sql}"'
+        "PGPASSWORD=$(cat /run/secrets/db_password) " f'psql -U discourse -d discourse -tAc "{sql}"'
    )
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()

@ -31,17 +30,18 @@ def _seed(domain, value):
    assert got == value, f"seed did not commit (read back {got!r}, expected {value!r})"


-def pre_upgrade(domain, meta):
-    _seed(domain, "upgrade-survives")
+def pre_upgrade(ctx):
+    _seed(ctx.domain, "upgrade-survives")


-def pre_backup(domain, meta):
-    _seed(domain, "original")
+def pre_backup(ctx):
+    _seed(ctx.domain, "original")


-def pre_restore(domain, meta):
+def pre_restore(ctx):
    # diverge from the backup so a successful restore is observable
-    _psql(domain, "DROP TABLE IF EXISTS ci_marker;")
-    assert _psql(domain, "SELECT to_regclass('public.ci_marker');") in ("", "NULL"), (
-        "drop did not take"
-    )
+    _psql(ctx.domain, "DROP TABLE IF EXISTS ci_marker;")
+    assert _psql(ctx.domain, "SELECT to_regclass('public.ci_marker');") in (
+        "",
+        "NULL",
+    ), "drop did not take"
--- a/tests/discourse/recipe_meta.py
+++ b/tests/discourse/recipe_meta.py
@ -6,7 +6,9 @@
 # app is actually serving (the canonical "is discourse up" signal — NOT "/", which may redirect to setup).
 HEALTH_PATH = "/srv/status"
 HEALTH_OK = (200,)
-DEPLOY_TIMEOUT = 3600  # slow Rails cold boot (15-25min) on the 7-GiB single node; bumped 2400→3600 for
+DEPLOY_TIMEOUT = (
+    3600  # slow Rails cold boot (15-25min) on the 7-GiB single node; bumped 2400→3600 for
+)
 # headroom after full4's base deploy timed out at 2400s (RAM/CPU-constrained boot + image re-pull).
 HTTP_TIMEOUT = 1200

@ -27,11 +29,11 @@ HTTP_TIMEOUT = 1200
 #   (1) it pins the Docker-Hub-removed `bitnami/discourse:3.3.1` (404) → overlay re-pins app+sidekiq to
 #       `bitnamilegacy/discourse:3.3.1` (namespace-only, identical image), the same re-pin the PR makes;
 #   (2) its 5m start_period is too tight for the 15-25min Rails boot → overlay widens it to 20m (grace).
-# install_steps.sh provides the overlay; CHAOS_BASE_DEPLOY skips the clean-tree gate on the untracked
-# overlay; it persists across the head checkout (idempotent — the PR head already re-pins + ships 20m).
+# The harness auto-provides the overlay to the checkout and auto-chaoses the base deploy
+# (first-class compose.ccci.yml, rcust P2a); it persists across the head checkout (idempotent — the
+# PR head already re-pins + ships 20m).
 # Upgrade crossover: 0.7.0 (re-pinned base) → PR head; full assertions run on the HEAD. The 0.7.0
 # *custom* tests are not separately run (custom tier runs once, on the head — policy §1 allows skip+record).
-CHAOS_BASE_DEPLOY = True
 UPGRADE_BASE_VERSION = "0.7.0+3.3.1"
 EXTRA_ENV = {
    "TIMEOUT": "3600",  # abra's internal convergence wait; matches DEPLOY_TIMEOUT (slow Rails boot headroom)
@ -39,7 +41,7 @@ EXTRA_ENV = {
 }


-def BACKUP_VERIFY(domain):
+def BACKUP_VERIFY(ctx):
    """Post-backup integrity check (Q4.6, same race ghost F2-14b hit). The recipe's backupbot db
    pre-hook (`/pg_backup.sh backup`) dumps the discourse postgres DB to `/var/lib/postgresql/data/
    backup.sql` (gzip), then restic captures that path. On the loaded single CI node the db container
@ -58,8 +60,12 @@ def BACKUP_VERIFY(domain):

    try:
        out = lifecycle.exec_in_app(
-            domain,
-            ["sh", "-c", "gzip -t /var/lib/postgresql/data/backup.sql && wc -c < /var/lib/postgresql/data/backup.sql"],
+            ctx.domain,
+            [
+                "sh",
+                "-c",
+                "gzip -t /var/lib/postgresql/data/backup.sql && wc -c < /var/lib/postgresql/data/backup.sql",
+            ],
            service="db",
            timeout=60,
        ).strip()
--- a/tests/discourse/test_backup.py
+++ b/tests/discourse/test_backup.py
@ -14,13 +14,12 @@ from harness import lifecycle  # noqa: E402

 def _psql(domain, sql):
    cmd = (
-        'PGPASSWORD=$(cat /run/secrets/db_password) '
-        f'psql -U discourse -d discourse -tAc "{sql}"'
+        "PGPASSWORD=$(cat /run/secrets/db_password) " f'psql -U discourse -d discourse -tAc "{sql}"'
    )
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()


 def test_backup_captures_state(live_app):
-    assert _psql(live_app, "SELECT v FROM ci_marker;") == "original", (
-        "the seeded discourse postgres state was not present at backup time"
-    )
+    assert (
+        _psql(live_app, "SELECT v FROM ci_marker;") == "original"
+    ), "the seeded discourse postgres state was not present at backup time"
--- a/tests/discourse/test_restore.py
+++ b/tests/discourse/test_restore.py
@ -14,13 +14,12 @@ from harness import lifecycle  # noqa: E402

 def _psql(domain, sql):
    cmd = (
-        'PGPASSWORD=$(cat /run/secrets/db_password) '
-        f'psql -U discourse -d discourse -tAc "{sql}"'
+        "PGPASSWORD=$(cat /run/secrets/db_password) " f'psql -U discourse -d discourse -tAc "{sql}"'
    )
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()


 def test_restore_returns_state(live_app):
-    assert _psql(live_app, "SELECT v FROM ci_marker;") == "original", (
-        "restore did not return the pre-mutation discourse postgres state (data-integrity failure)"
-    )
+    assert (
+        _psql(live_app, "SELECT v FROM ci_marker;") == "original"
+    ), "restore did not return the pre-mutation discourse postgres state (data-integrity failure)"
--- a/tests/ghost/functional/_ghost.py
+++ b/tests/ghost/functional/_ghost.py
@ -93,9 +93,10 @@ class GhostAdmin:
        status, body = self.req(
            "POST", "/session/", {"username": ADMIN_EMAIL, "password": ADMIN_PW}
        )
-        assert status in (200, 201), (
-            f"ghost admin session login failed: HTTP {status}, body={body!r}"
-        )
+        assert status in (
+            200,
+            201,
+        ), f"ghost admin session login failed: HTTP {status}, body={body!r}"

    def create_post(self, title: str, html: str) -> dict:
        status, body = self.req(
--- a/tests/ghost/functional/test_admin_redirect.py
+++ b/tests/ghost/functional/test_admin_redirect.py
@ -53,13 +53,15 @@ def test_ghost_admin_route_is_wired(live_app):
        return None

    status_body = harness_http.assert_converges(
-        _ready, f"GET {url} returns Ghost admin (200) or setup redirect (302)",
-        max_wait=60, interval=3,
+        _ready,
+        f"GET {url} returns Ghost admin (200) or setup redirect (302)",
+        max_wait=60,
+        interval=3,
    )
    status, body = status_body
    assert status in (200, 302), f"unexpected status: {status}"
    if status == 200:
        # The admin SPA references /ghost-assets/ or contains "ghost" in title/body
-        assert "ghost" in body.lower(), (
-            f"GET {url} 200 but body has no Ghost markers: {body[:200]!r}"
-        )
+        assert (
+            "ghost" in body.lower()
+        ), f"GET {url} 200 but body has no Ghost markers: {body[:200]!r}"
--- a/tests/ghost/functional/test_content_api.py
+++ b/tests/ghost/functional/test_content_api.py
@ -35,10 +35,10 @@ def test_content_api_settings_endpoint(live_app):
    assert body is not None, f"GET {url} returned non-JSON body"
    # On success: {"settings": {...}}. On error: {"errors": [...]}. Either shape is valid.
    if status == 200:
-        assert isinstance(body, dict) and "settings" in body, (
-            f"200 response missing 'settings' envelope: {body!r}"
-        )
+        assert (
+            isinstance(body, dict) and "settings" in body
+        ), f"200 response missing 'settings' envelope: {body!r}"
    else:
-        assert isinstance(body, dict) and ("errors" in body or "message" in body or body), (
-            f"error response not a proper Ghost error envelope: {body!r}"
-        )
+        assert isinstance(body, dict) and (
+            "errors" in body or "message" in body or body
+        ), f"error response not a proper Ghost error envelope: {body!r}"
--- a/tests/ghost/functional/test_post_roundtrip.py
+++ b/tests/ghost/functional/test_post_roundtrip.py
@ -43,17 +43,17 @@ def test_create_post_roundtrip(live_app):
    title = f"ccci-marker-{uniq}"
    marker = f"ccci-body-marker-{uniq}-roundtrip"
    created = admin.create_post(title, f"<p>{marker}</p>")
-    assert created.get("title") == title, (
-        f"created post title mismatch: sent {title!r}, got {created.get('title')!r}"
-    )
+    assert (
+        created.get("title") == title
+    ), f"created post title mismatch: sent {title!r}, got {created.get('title')!r}"

    # 4) Read it back by id and assert the post survived the round-trip (title always returned;
    #    html returned because we requested ?formats=html).
    got = admin.get_post(created["id"])
-    assert got.get("title") == title, (
-        f"post title did not round-trip: sent {title!r}, got {got.get('title')!r}"
-    )
+    assert (
+        got.get("title") == title
+    ), f"post title did not round-trip: sent {title!r}, got {got.get('title')!r}"
    html = got.get("html") or ""
-    assert marker in html, (
-        f"post body did not round-trip: marker {marker!r} not in read-back html {html!r}"
-    )
+    assert (
+        marker in html
+    ), f"post body did not round-trip: marker {marker!r} not in read-back html {html!r}"
--- a/tests/ghost/install_steps.sh
+++ b/tests/ghost/install_steps.sh
@ -1,26 +0,0 @@
-#!/usr/bin/env bash
-# ghost — INSTALL-TIME hook (Phase 2 F2-14b). Runs during the install tier AFTER `abra app new` +
-# EXTRA_ENV + `abra app secret generate` and BEFORE the single `abra app deploy`
-# (lifecycle.py::_run_install_steps), with CCCI_RECIPE / CCCI_APP_DOMAIN in env.
-#
-# Purpose: provide the cc-ci start_period-grace overlay (compose.ccci.yml) to the recipe checkout so
-# the UPGRADE-tier BASE deploy (a previous published version whose app healthcheck still ships the
-# too-tight 1m start_period) can survive ghost's ~6-9min fresh-DB migration and converge. See
-# compose.ccci.yml's header for the full rationale. The overlay is referenced by recipe_meta
-# COMPOSE_FILE; copying it here (it is a cc-ci file, not part of the recipe) makes it resolvable.
-# It persists across the later `git checkout <head>` (untracked) so the head deploy also merges it
-# (idempotent — the PR head already ships 15m). CHAOS_BASE_DEPLOY=True is set so abra's pinned-deploy
-# clean-tree check doesn't FATA on the untracked overlay.
-set -euo pipefail
-
-: "${CCCI_RECIPE:?missing CCCI_RECIPE}"
-SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
-RECIPE_DIR="${HOME}/.abra/recipes/${CCCI_RECIPE}"
-
-if [ ! -d "$RECIPE_DIR" ]; then
-  echo "  ghost install_steps: recipe dir $RECIPE_DIR missing — cannot provide compose.ccci.yml" >&2
-  exit 1
-fi
-
-cp "$SCRIPT_DIR/compose.ccci.yml" "$RECIPE_DIR/compose.ccci.yml"
-echo "  ghost install_steps: provided compose.ccci.yml (app start_period grace) to recipe checkout (${CCCI_RECIPE})"
--- a/tests/ghost/ops.py
+++ b/tests/ghost/ops.py
@ -22,10 +22,7 @@ from harness import lifecycle  # noqa: E402


 def _mysql(domain, sql):
-    cmd = (
-        'MYSQL_PWD="$(cat /run/secrets/db_password)" '
-        f'mysql -u root -N -s ghost -e "{sql}"'
-    )
+    cmd = 'MYSQL_PWD="$(cat /run/secrets/db_password)" ' f'mysql -u root -N -s ghost -e "{sql}"'
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()


@ -39,19 +36,19 @@ def _seed(domain, value):
    assert got == value, f"seed did not commit (read back {got!r}, expected {value!r})"


-def pre_upgrade(domain, meta):
-    _seed(domain, "upgrade-survives")
+def pre_upgrade(ctx):
+    _seed(ctx.domain, "upgrade-survives")


-def pre_backup(domain, meta):
-    _seed(domain, "original")
+def pre_backup(ctx):
+    _seed(ctx.domain, "original")


-def pre_restore(domain, meta):
+def pre_restore(ctx):
    # diverge from the backup so a successful restore is observable: drop the marker table.
-    _mysql(domain, "DROP TABLE IF EXISTS ci_marker;")
+    _mysql(ctx.domain, "DROP TABLE IF EXISTS ci_marker;")
    got = _mysql(
-        domain,
+        ctx.domain,
        "SELECT COUNT(*) FROM information_schema.tables "
        "WHERE table_schema='ghost' AND table_name='ci_marker';",
    )
--- a/tests/ghost/recipe_meta.py
+++ b/tests/ghost/recipe_meta.py
@ -31,23 +31,22 @@ HTTP_TIMEOUT = 900
 # (plan-ccci-compose-overlay-policy.md §1), so the harness base-deploys the previous PUBLISHED version
 # (1.1.1+6-alpine) — which predates the PR and still ships the too-tight 1m start_period → it would
 # deadlock on the same migration kill. compose.ccci.yml re-applies the 15m grace to the BASE so the
-# from-version is deployable; install_steps.sh provides it to the checkout; CHAOS_BASE_DEPLOY skips the
-# clean-tree gate on that untracked overlay. It persists across the head checkout (idempotent — the PR
-# head already ships 15m). This is the policy-blessed "minimal overlay on the from-version so
+# from-version is deployable; the harness auto-provides it to the checkout and auto-chaoses the base
+# deploy (first-class compose.ccci.yml, rcust P2a). It persists across the head checkout (idempotent —
+# the PR head already ships 15m). This is the policy-blessed "minimal overlay on the from-version so
 # upgrade-to-latest can run" — grace-only, masks no defect, weakens no test.
 # TIMEOUT/DEPLOY_TIMEOUT 2400s: the BASE cold boot's wall-time is mysql fresh-dir init (~6min, during
 # which the app crash-loops harmlessly on `ECONNREFUSED 3306` until mysql accepts connections — no
 # migration progress lost, it hasn't started) PLUS the ~9-15min schema migration (round-trip-bound,
 # slower under host load). 1200s was too tight (full4 killed at the near-final `email_recipients`
 # tables while still 0/1); 2400s gives headroom while still bounding a genuine hang (matches discourse).
-CHAOS_BASE_DEPLOY = True
 EXTRA_ENV = {
    "TIMEOUT": "2400",
    "COMPOSE_FILE": "compose.yml:compose.ccci.yml",
 }


-def BACKUP_VERIFY(domain):
+def BACKUP_VERIFY(ctx):
    """Post-backup integrity check (F2-14b). The recipe's backupbot db pre-hook dumps the ghost MySQL
    DB to `/var/lib/mysql/backup.sql.gz` (then restic captures that path). On the loaded single CI node
    the db container intermittently CYCLES mid-dump (observed: full5/6/7 RED, full8 green — pure race;
@ -62,8 +61,12 @@ def BACKUP_VERIFY(domain):

    try:
        out = lifecycle.exec_in_app(
-            domain,
-            ["sh", "-c", "gzip -t /var/lib/mysql/backup.sql.gz && wc -c < /var/lib/mysql/backup.sql.gz"],
+            ctx.domain,
+            [
+                "sh",
+                "-c",
+                "gzip -t /var/lib/mysql/backup.sql.gz && wc -c < /var/lib/mysql/backup.sql.gz",
+            ],
            service="db",
            timeout=60,
        ).strip()
--- a/tests/ghost/test_backup.py
+++ b/tests/ghost/test_backup.py
@ -15,14 +15,11 @@ from harness import lifecycle  # noqa: E402


 def _mysql(domain, sql):
-    cmd = (
-        'MYSQL_PWD="$(cat /run/secrets/db_password)" '
-        f'mysql -u root -N -s ghost -e "{sql}"'
-    )
+    cmd = 'MYSQL_PWD="$(cat /run/secrets/db_password)" ' f'mysql -u root -N -s ghost -e "{sql}"'
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()


 def test_backup_captures_state(live_app):
-    assert _mysql(live_app, "SELECT v FROM ci_marker;") == "original", (
-        "the seeded ghost MySQL marker was not present at backup time"
-    )
+    assert (
+        _mysql(live_app, "SELECT v FROM ci_marker;") == "original"
+    ), "the seeded ghost MySQL marker was not present at backup time"
--- a/tests/ghost/test_restore.py
+++ b/tests/ghost/test_restore.py
@ -22,10 +22,7 @@ from harness import lifecycle  # noqa: E402


 def _mysql(domain, sql):
-    cmd = (
-        'MYSQL_PWD="$(cat /run/secrets/db_password)" '
-        f'mysql -u root -N -s ghost -e "{sql}"'
-    )
+    cmd = 'MYSQL_PWD="$(cat /run/secrets/db_password)" ' f'mysql -u root -N -s ghost -e "{sql}"'
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()


--- a/tests/ghost/test_upgrade.py
+++ b/tests/ghost/test_upgrade.py
@ -14,14 +14,11 @@ from harness import lifecycle  # noqa: E402


 def _mysql(domain, sql):
-    cmd = (
-        'MYSQL_PWD="$(cat /run/secrets/db_password)" '
-        f'mysql -u root -N -s ghost -e "{sql}"'
-    )
+    cmd = 'MYSQL_PWD="$(cat /run/secrets/db_password)" ' f'mysql -u root -N -s ghost -e "{sql}"'
    return lifecycle.exec_in_app(domain, ["sh", "-c", cmd], service="db").strip()


 def test_upgrade_preserves_state(live_app):
-    assert _mysql(live_app, "SELECT v FROM ci_marker;") == "upgrade-survives", (
-        "the seeded ghost MySQL marker did not survive the upgrade redeploy (data loss on upgrade)"
-    )
+    assert (
+        _mysql(live_app, "SELECT v FROM ci_marker;") == "upgrade-survives"
+    ), "the seeded ghost MySQL marker did not survive the upgrade redeploy (data loss on upgrade)"
--- a/tests/hedgedoc/functional/test_branding.py
+++ b/tests/hedgedoc/functional/test_branding.py
@ -14,7 +14,6 @@ import urllib.request
 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "..", "runner"))
 from harness import http as harness_http  # noqa: E402

-
 _CTX = ssl.create_default_context()
 _CTX.check_hostname = False
 _CTX.verify_mode = ssl.CERT_NONE
--- a/tests/hedgedoc/functional/test_health_check.py
+++ b/tests/hedgedoc/functional/test_health_check.py
@ -15,7 +15,5 @@ from harness import http as harness_http  # noqa: E402
 def test_hedgedoc_root_serves(live_app):
    """GET / → 200 or 302 (login/new redirect)."""
    url = f"https://{live_app}/"
-    status, _ = harness_http.retry_http_get(
-        url, expect_status=(200, 302), max_wait=90, interval=5
-    )
+    status, _ = harness_http.retry_http_get(url, expect_status=(200, 302), max_wait=90, interval=5)
    assert status in (200, 302), f"GET {url} HTTP {status} (expected 200 or 302)"
--- a/tests/immich/functional/test_asset_processing.py
+++ b/tests/immich/functional/test_asset_processing.py
@ -111,13 +111,13 @@ def test_immich_processes_uploaded_asset_metadata_and_statistics(live_app):
        if exif and exif.get("exifImageWidth"):
            break
        time.sleep(5)
-    assert exif and exif.get("exifImageWidth") == 1 and exif.get("exifImageHeight") == 1, (
-        f"immich metadata-extraction did not populate the 1x1 PNG dimensions in exifInfo: {exif!r}"
-    )
+    assert (
+        exif and exif.get("exifImageWidth") == 1 and exif.get("exifImageHeight") == 1
+    ), f"immich metadata-extraction did not populate the 1x1 PNG dimensions in exifInfo: {exif!r}"

    # the asset is catalogued into the owner's library statistics (list-back in aggregate)
    sst, stats = harness_http.http_request("GET", f"{base}/api/assets/statistics", headers=auth)
    assert sst == 200 and isinstance(stats, dict), f"statistics HTTP {sst}: {stats!r}"
-    assert stats.get("images", 0) >= 1 and stats.get("total", 0) >= 1, (
-        f"uploaded asset not reflected in library statistics: {stats!r}"
-    )
+    assert (
+        stats.get("images", 0) >= 1 and stats.get("total", 0) >= 1
+    ), f"uploaded asset not reflected in library statistics: {stats!r}"
--- a/tests/immich/functional/test_asset_upload.py
+++ b/tests/immich/functional/test_asset_upload.py
@ -121,6 +121,6 @@ def test_immich_upload_asset_readback_and_thumbnail(live_app):
        if thumb == 200:
            break
        time.sleep(5)
-    assert thumb == 200, (
-        f"immich did not generate a thumbnail/derivative for the uploaded asset (last HTTP {thumb})"
-    )
+    assert (
+        thumb == 200
+    ), f"immich did not generate a thumbnail/derivative for the uploaded asset (last HTTP {thumb})"
--- a/tests/immich/functional/test_health_check.py
+++ b/tests/immich/functional/test_health_check.py
@ -16,5 +16,11 @@ from harness import http as harness_http  # noqa: E402

 def test_immich_returns_200(live_app):
    url = f"https://{live_app}/"
-    status, _ = harness_http.retry_http_get(url, expect_status=(200, 301, 302), max_wait=60, interval=3)
-    assert status in (200, 301, 302), f"immich at {url} returned HTTP {status} (expected 200/301/302)"
+    status, _ = harness_http.retry_http_get(
+        url, expect_status=(200, 301, 302), max_wait=60, interval=3
+    )
+    assert status in (
+        200,
+        301,
+        302,
+    ), f"immich at {url} returned HTTP {status} (expected 200/301/302)"
--- a/tests/immich/ops.py
+++ b/tests/immich/ops.py
@ -25,14 +25,17 @@ def _seed(domain, value):
    assert _psql(domain, "SELECT v FROM ci_marker;") == value


-def pre_upgrade(domain, meta):
-    _seed(domain, "upgrade-survives")
+def pre_upgrade(ctx):
+    _seed(ctx.domain, "upgrade-survives")


-def pre_backup(domain, meta):
-    _seed(domain, "original")
+def pre_backup(ctx):
+    _seed(ctx.domain, "original")


-def pre_restore(domain, meta):
-    _psql(domain, "DROP TABLE ci_marker;")
-    assert _psql(domain, "SELECT to_regclass('public.ci_marker');") in ("", "NULL"), "drop did not take"
+def pre_restore(ctx):
+    _psql(ctx.domain, "DROP TABLE ci_marker;")
+    assert _psql(ctx.domain, "SELECT to_regclass('public.ci_marker');") in (
+        "",
+        "NULL",
+    ), "drop did not take"
--- a/tests/immich/test_backup.py
+++ b/tests/immich/test_backup.py
@ -14,4 +14,6 @@ def _psql(domain, sql):


 def test_backup_captures_state(live_app):
-    assert _psql(live_app, "SELECT v FROM ci_marker;") == "original", "seeded postgres state not present at backup time"
+    assert (
+        _psql(live_app, "SELECT v FROM ci_marker;") == "original"
+    ), "seeded postgres state not present at backup time"
--- a/tests/immich/test_install.py
+++ b/tests/immich/test_install.py
@ -7,7 +7,8 @@ import os
 import sys

 sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "runner"))
-from harness import browser as harness_browser, generic, lifecycle  # noqa: E402
+from harness import browser as harness_browser  # noqa: E402
+from harness import generic, lifecycle


 def test_serving_and_frontend(live_app, meta):
@ -25,7 +26,11 @@ def test_serving_and_frontend(live_app, meta):
            resp = harness_browser.goto_with_retry(
                page, url, accept_statuses=(200, 301, 302), goto_timeout_ms=60_000
            )
-            assert resp is not None and resp.status in (200, 301, 302), f"page status {resp and resp.status}"
+            assert resp is not None and resp.status in (
+                200,
+                301,
+                302,
+            ), f"page status {resp and resp.status}"
            assert "<html" in page.content().lower(), "no HTML served by the immich frontend"
        finally:
            browser.close()
--- a/tests/immich/test_restore.py
+++ b/tests/immich/test_restore.py
@ -14,4 +14,6 @@ def _psql(domain, sql):


 def test_restore_returns_state(live_app):
-    assert _psql(live_app, "SELECT v FROM ci_marker;") == "original", "restore did not return the pre-mutation postgres state"
+    assert (
+        _psql(live_app, "SELECT v FROM ci_marker;") == "original"
+    ), "restore did not return the pre-mutation postgres state"
--- a/tests/immich/test_upgrade.py
+++ b/tests/immich/test_upgrade.py
@ -14,4 +14,6 @@ def _psql(domain, sql):


 def test_upgrade_preserves_data(live_app):
-    assert _psql(live_app, "SELECT v FROM ci_marker;") == "upgrade-survives", "postgres data did not survive the upgrade"
+    assert (
+        _psql(live_app, "SELECT v FROM ci_marker;") == "upgrade-survives"
+    ), "postgres data did not survive the upgrade"
--- a/tests/keycloak/functional/test_create_client_and_use.py
+++ b/tests/keycloak/functional/test_create_client_and_use.py
@ -120,9 +120,9 @@ def test_create_confidential_client_and_obtain_token(live_app):
        "clientId": client_id,
        "enabled": True,
        "secret": client_secret,
-        "publicClient": False,            # confidential client
-        "serviceAccountsEnabled": True,    # required for client_credentials grant
-        "standardFlowEnabled": False,      # not needed for service-account-only client
+        "publicClient": False,  # confidential client
+        "serviceAccountsEnabled": True,  # required for client_credentials grant
+        "standardFlowEnabled": False,  # not needed for service-account-only client
        "directAccessGrantsEnabled": False,
        "protocol": "openid-connect",
    }
@ -144,25 +144,25 @@ def test_create_confidential_client_and_obtain_token(live_app):

        # Use the client to obtain its own token (client_credentials grant)
        tok_status, tok_resp = _client_credentials_token(live_app, client_id, client_secret)
-        assert tok_status == 200, (
-            f"client_credentials token returned HTTP {tok_status}: {tok_resp!r}"
-        )
+        assert (
+            tok_status == 200
+        ), f"client_credentials token returned HTTP {tok_status}: {tok_resp!r}"
        access_token = tok_resp.get("access_token") if isinstance(tok_resp, dict) else None
-        assert isinstance(access_token, str) and access_token.count(".") == 2, (
-            f"client_credentials access_token not a JWT: {access_token!r}"
-        )
+        assert (
+            isinstance(access_token, str) and access_token.count(".") == 2
+        ), f"client_credentials access_token not a JWT: {access_token!r}"

        # Decode the JWT payload; assert azp matches the new client
        payload = json.loads(_b64url_decode(access_token.split(".")[1]))
-        assert payload.get("azp") == client_id, (
-            f"client_credentials JWT azp={payload.get('azp')!r} != client_id={client_id!r}"
-        )
+        assert (
+            payload.get("azp") == client_id
+        ), f"client_credentials JWT azp={payload.get('azp')!r} != client_id={client_id!r}"
        # Service-account token does NOT carry a session-scoped user (azp + clientId differ from
        # admin-cli token). The presence of azp + iss == per-run-domain proves the issuance flow.
        expected_iss = f"https://{live_app}/realms/master"
-        assert payload.get("iss") == expected_iss, (
-            f"JWT iss={payload.get('iss')!r} != {expected_iss!r}"
-        )
+        assert (
+            payload.get("iss") == expected_iss
+        ), f"JWT iss={payload.get('iss')!r} != {expected_iss!r}"
    finally:
        # Idempotent cleanup
        if cleanup_id:
--- a/Show More
+++ b/Show More