orchestrator: reboot-resilience + session auto-resume + full session plan/tooling

Reboot survival for the Pi orchestrator host:
- systemd unit cc-ci-plan/systemd/cc-ci-loops.service (installed + enabled): on boot
  records the reboot, starts loops+watchdog (RESUME_PHASE=1), and resumes the
  orchestrator session.
- reboot-log.sh: boot_id-gated reboot record -> REBOOTS.md (manual restarts don't count).
- launch-orchestrator.sh: injects an AGENTS.md startup nudge so an auto-resumed
  orchestrator announces itself (PushNotification) + reports reboots.
- AGENTS.md: on-startup notify routine documented.

Plans/tooling accumulated this session:
- plan-phase1d (generic suite), 1e (harness corrections), phase4 (final review),
  sso-dep-testing, orchestrator-migration (parked), test-e2e-testme-acceptance.
- launch.sh: 1d/1e/2/2b/3/4 phase sequence, machine-docs-aware state resolution,
  limit-stall re-nudge, INBOX side-channel detection.
- plan.md §6.1/§7: artifact-layer isolation, INBOX, 5-min long-run polling, DEFERRED.
- prompts: isolation discipline + INBOX + pacing.
- .gitignore: harden (.sops/, cc-ci-secrets/, .claude/, *.tmp.*).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-05-28 20:28:10 +01:00
co-authored by Claude Opus 4.8
parent 5681438b0f
commit 36a6c9872a
20 changed files with 1395 additions and 19 deletions
+58 -3
View File
@@ -126,7 +126,13 @@ without the auth key.
- **Wildcard TLS cert — PROVIDED, not a token.** The operator has pre-issued the wildcard SAN cert
(`*.ci.commoninternet.net` + `ci.commoninternet.net`) and placed it on cc-ci at
`/var/lib/ci-certs/live/{fullchain.pem,privkey.pem}` (§4.0). The agent feeds these into the
`/var/lib/ci-certs/live/{fullchain.pem,privkey.pem}` (§4.0).
> **Phase-1c update (supersedes the cert references in §1.5/§4.0/§4.4 below):** the cert is no longer
> an out-of-band operator file-drop — it is now **sops-encrypted in the private `cc-ci-secrets` repo**
> (a git submodule) and **decrypted at activation to that same path** by sops-nix. Issuance stays
> operator-only (LE/Gandi, no token on the box); to rotate, the operator re-issues then re-encrypts
> the cert into `cc-ci-secrets` and rebuilds. The ONE out-of-band secret is now the bootstrap age key
> at `/var/lib/sops-nix/key.txt`. Authoritative model: `cc-ci/docs/secrets.md` + `docs/install.md`. The agent feeds these into the
`coop-cloud/traefik` recipe as its `ssl_cert`/`ssl_key` swarm secrets (wildcard/file-provider
mode) and runs **no ACME** for this domain. **Do not request or expect a `commoninternet.net` DNS
token** — issuance/renewal is handled out-of-band by the operator (LE 90-day cert; next renewal
@@ -597,8 +603,37 @@ its own pacing. To make concurrent writes conflict-free:
merges the two cleanly. Closing an item = checking the box *in your own section*; the Builder
fixes an `[adversary]` finding and notes the fix in JOURNAL, but only the Adversary ticks it
closed after re-test.
- `DEFERRED.md` (in `machine-docs/`) is the **single canonical registry for things the loops
have deliberately decided not to do autonomously and that need operator input to move on.**
Append-only; either agent may file. Each entry should clearly say *what's needed from the
operator* to lift the deferral (an opt-in flag, a resource decision, an architectural call,
plain "go ahead"). The list is **open-ended** — items can sit indefinitely, **no obligation
to close every item**, closure is operator-driven. A re-entry trigger / IDEA cross-link is
**optional** (include when there's a natural mechanism, e.g. an opt-in flag in
`cc-ci-plan/IDEAS.md`). Don't park deferrals as a vague "Q4 follow-up" / buried JOURNAL note
— file them here so the operator can review the whole list. The Phase-4 cleanup pass should
**surface** DEFERRED.md to the operator at least once but does **not** force closure.
Future-aspirational ideas (out of current scope) still go to `cc-ci-plan/IDEAS.md`; DEFERRED
is for considered-and-parked work the loops won't tackle without operator input.
- **Append-only where possible.** `JOURNAL.md` and `REVIEW.md` are append-only logs → they never
conflict. Prefer appending over rewriting.
- **Artifact-layer isolation — facts in STATUS, reasoning in JOURNAL (anti-anchoring).** Rigorous
adversarial verification requires the Adversary NOT to consume the Builder's rationalisations
before forming its verdict. The split:
- `STATUS.md` MUST carry **everything the Adversary needs to verify the claim** — withholding
verification context defeats the verification: **WHAT** is claimed (gate id, DoD items), **HOW**
to verify (the exact command/check the Adversary can re-run from its own clone), the
**EXPECTED** outcome (build hashes, file contents, leaf fingerprints, status codes), and
**WHERE** the inputs live (commit shas, paths). If it's essential for the Adversary to verify,
it goes in STATUS.
- `STATUS.md` MUST NOT carry rationalisations / "why I think this passes" / design narrative /
dead-ends explored. Those go in `JOURNAL.md` (Builder-private to write).
- The Adversary reads STATUS for the claim + verification info, the plan as SSOT, and the code /
git history; it forms its verdict from those + its own **cold** acceptance run, and does **not**
read `JOURNAL.md` before the verdict. After an independent verdict, consulting JOURNAL is fine
(e.g. to contextualise a finding) — note in REVIEW that you did.
In short: **WHAT + HOW + EXPECTED + WHERE = STATUS; WHY = JOURNAL.**
- **Git discipline (both loops, every write):** `git pull --rebase` before editing, make the
smallest change, commit, `git push`. On a rebase conflict, it will be inside the *other* agent's
file/section only if a rule was broken — re-pull and keep to your own files. Never `--force`.
@@ -613,6 +648,22 @@ its own pacing. To make concurrent writes conflict-free:
- **Liveness.** If the Adversary sees a gate `CLAIMED` for too long with no Builder progress, or
the Builder sees no Adversary verdict on a standing claim, note it in your own ledger and keep
doing independent work — neither loop blocks idle waiting on the other beyond its gate.
- **INBOX — explicit cross-loop messaging beyond CLAIMS.** Sometimes you have something to say to
the other loop that isn't a gate claim or a REVIEW verdict (a heads-up, a request for
early-look, a "I refactored X, please re-verify Y", an observation outside the normal flow). For
those, use the inbox files in `machine-docs/`:
- **Builder → Adversary:** the Builder writes/appends `machine-docs/ADVERSARY-INBOX.md` in its
own clone, commits, pushes.
- **Adversary → Builder:** the Adversary writes/appends `machine-docs/BUILDER-INBOX.md` in its
own clone, commits, pushes.
- The watchdog edge-triggers on **newly-present** inbox files in the relevant clone and pings
the receiver. The receiver, on receipt, reads + processes the message, then **deletes the
inbox file** (commits + pushes) — deletion is the "message consumed" signal. Single-writer
discipline: only the sender writes their counterpart's inbox; only the receiver deletes it.
- **Use for:** non-gate signals — "heads-up I'm about to refactor X," "please cold-verify this
while I keep going," "I observed Y outside our normal flow," "I'm taking a long e2e now."
**Do NOT use for:** formal gate claims (`STATUS.md` still owns those) or verdicts (`REVIEW.md`
still owns those). The inbox is a side-channel, not a replacement.
(If you are ever forced to run with a single process, the degraded fallback is to alternate
roles per iteration and keep `JOURNAL.md` and `REVIEW.md` strictly separate — but two loops is
@@ -649,8 +700,12 @@ every wake, `git pull --rebase` first, then:
**Pacing.** Use `/loop` (self-paced) or `ScheduleWakeup`. Most waits here are for things the
harness can't notify you about — a Drone build, a `nixos-rebuild`, a deploy converging — so poll
the *specific* thing. Three cases:
1. **Something in flight** (build/deploy/`nixos-rebuild`) → re-check on a short cadence (≈4 min) to
stay cache-warm; keep polling *it*, don't treat it as idle, and don't spin on a minutes-long build.
1. **Something in flight** (build/deploy/`nixos-rebuild`/e2e/heavy test) → **poll every ~5 min** to
stay cache-warm and to **see failures as they happen**, not at the end of a 25-minute sleep. Do
**NOT** `ScheduleWakeup` for the expected total runtime of the task in a single big sleep — a 25
min e2e gets 5 short cache-warm polls, not one 25-min cache-cold blackout. The wakeup that wakes
you mid-task is *cheap* (one cache hit, one quick status check); the value of catching a deploy
that died at minute 4 of a 25-min budget is large. Keep polling *it*, don't treat it as idle.
2. **Blocked on the *other* loop** — Builder parked at a `CLAIMED` gate awaiting the Adversary, or
Adversary waiting for the Builder to fix an `[adversary]` finding. **You don't need to busy-poll
here: the watchdog signals across the handoff.** The moment the Builder writes a `CLAIMED` gate,