orchestrator: reboot-resilience + session auto-resume + full session plan/tooling
Reboot survival for the Pi orchestrator host: - systemd unit cc-ci-plan/systemd/cc-ci-loops.service (installed + enabled): on boot records the reboot, starts loops+watchdog (RESUME_PHASE=1), and resumes the orchestrator session. - reboot-log.sh: boot_id-gated reboot record -> REBOOTS.md (manual restarts don't count). - launch-orchestrator.sh: injects an AGENTS.md startup nudge so an auto-resumed orchestrator announces itself (PushNotification) + reports reboots. - AGENTS.md: on-startup notify routine documented. Plans/tooling accumulated this session: - plan-phase1d (generic suite), 1e (harness corrections), phase4 (final review), sso-dep-testing, orchestrator-migration (parked), test-e2e-testme-acceptance. - launch.sh: 1d/1e/2/2b/3/4 phase sequence, machine-docs-aware state resolution, limit-stall re-nudge, INBOX side-channel detection. - plan.md §6.1/§7: artifact-layer isolation, INBOX, 5-min long-run polling, DEFERRED. - prompts: isolation discipline + INBOX + pacing. - .gitignore: harden (.sops/, cc-ci-secrets/, .claude/, *.tmp.*). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
+58
-3
@@ -126,7 +126,13 @@ without the auth key.
|
||||
|
||||
- **Wildcard TLS cert — PROVIDED, not a token.** The operator has pre-issued the wildcard SAN cert
|
||||
(`*.ci.commoninternet.net` + `ci.commoninternet.net`) and placed it on cc-ci at
|
||||
`/var/lib/ci-certs/live/{fullchain.pem,privkey.pem}` (§4.0). The agent feeds these into the
|
||||
`/var/lib/ci-certs/live/{fullchain.pem,privkey.pem}` (§4.0).
|
||||
> **Phase-1c update (supersedes the cert references in §1.5/§4.0/§4.4 below):** the cert is no longer
|
||||
> an out-of-band operator file-drop — it is now **sops-encrypted in the private `cc-ci-secrets` repo**
|
||||
> (a git submodule) and **decrypted at activation to that same path** by sops-nix. Issuance stays
|
||||
> operator-only (LE/Gandi, no token on the box); to rotate, the operator re-issues then re-encrypts
|
||||
> the cert into `cc-ci-secrets` and rebuilds. The ONE out-of-band secret is now the bootstrap age key
|
||||
> at `/var/lib/sops-nix/key.txt`. Authoritative model: `cc-ci/docs/secrets.md` + `docs/install.md`. The agent feeds these into the
|
||||
`coop-cloud/traefik` recipe as its `ssl_cert`/`ssl_key` swarm secrets (wildcard/file-provider
|
||||
mode) and runs **no ACME** for this domain. **Do not request or expect a `commoninternet.net` DNS
|
||||
token** — issuance/renewal is handled out-of-band by the operator (LE 90-day cert; next renewal
|
||||
@@ -597,8 +603,37 @@ its own pacing. To make concurrent writes conflict-free:
|
||||
merges the two cleanly. Closing an item = checking the box *in your own section*; the Builder
|
||||
fixes an `[adversary]` finding and notes the fix in JOURNAL, but only the Adversary ticks it
|
||||
closed after re-test.
|
||||
- `DEFERRED.md` (in `machine-docs/`) is the **single canonical registry for things the loops
|
||||
have deliberately decided not to do autonomously and that need operator input to move on.**
|
||||
Append-only; either agent may file. Each entry should clearly say *what's needed from the
|
||||
operator* to lift the deferral (an opt-in flag, a resource decision, an architectural call,
|
||||
plain "go ahead"). The list is **open-ended** — items can sit indefinitely, **no obligation
|
||||
to close every item**, closure is operator-driven. A re-entry trigger / IDEA cross-link is
|
||||
**optional** (include when there's a natural mechanism, e.g. an opt-in flag in
|
||||
`cc-ci-plan/IDEAS.md`). Don't park deferrals as a vague "Q4 follow-up" / buried JOURNAL note
|
||||
— file them here so the operator can review the whole list. The Phase-4 cleanup pass should
|
||||
**surface** DEFERRED.md to the operator at least once but does **not** force closure.
|
||||
Future-aspirational ideas (out of current scope) still go to `cc-ci-plan/IDEAS.md`; DEFERRED
|
||||
is for considered-and-parked work the loops won't tackle without operator input.
|
||||
- **Append-only where possible.** `JOURNAL.md` and `REVIEW.md` are append-only logs → they never
|
||||
conflict. Prefer appending over rewriting.
|
||||
- **Artifact-layer isolation — facts in STATUS, reasoning in JOURNAL (anti-anchoring).** Rigorous
|
||||
adversarial verification requires the Adversary NOT to consume the Builder's rationalisations
|
||||
before forming its verdict. The split:
|
||||
- `STATUS.md` MUST carry **everything the Adversary needs to verify the claim** — withholding
|
||||
verification context defeats the verification: **WHAT** is claimed (gate id, DoD items), **HOW**
|
||||
to verify (the exact command/check the Adversary can re-run from its own clone), the
|
||||
**EXPECTED** outcome (build hashes, file contents, leaf fingerprints, status codes), and
|
||||
**WHERE** the inputs live (commit shas, paths). If it's essential for the Adversary to verify,
|
||||
it goes in STATUS.
|
||||
- `STATUS.md` MUST NOT carry rationalisations / "why I think this passes" / design narrative /
|
||||
dead-ends explored. Those go in `JOURNAL.md` (Builder-private to write).
|
||||
- The Adversary reads STATUS for the claim + verification info, the plan as SSOT, and the code /
|
||||
git history; it forms its verdict from those + its own **cold** acceptance run, and does **not**
|
||||
read `JOURNAL.md` before the verdict. After an independent verdict, consulting JOURNAL is fine
|
||||
(e.g. to contextualise a finding) — note in REVIEW that you did.
|
||||
|
||||
In short: **WHAT + HOW + EXPECTED + WHERE = STATUS; WHY = JOURNAL.**
|
||||
- **Git discipline (both loops, every write):** `git pull --rebase` before editing, make the
|
||||
smallest change, commit, `git push`. On a rebase conflict, it will be inside the *other* agent's
|
||||
file/section only if a rule was broken — re-pull and keep to your own files. Never `--force`.
|
||||
@@ -613,6 +648,22 @@ its own pacing. To make concurrent writes conflict-free:
|
||||
- **Liveness.** If the Adversary sees a gate `CLAIMED` for too long with no Builder progress, or
|
||||
the Builder sees no Adversary verdict on a standing claim, note it in your own ledger and keep
|
||||
doing independent work — neither loop blocks idle waiting on the other beyond its gate.
|
||||
- **INBOX — explicit cross-loop messaging beyond CLAIMS.** Sometimes you have something to say to
|
||||
the other loop that isn't a gate claim or a REVIEW verdict (a heads-up, a request for
|
||||
early-look, a "I refactored X, please re-verify Y", an observation outside the normal flow). For
|
||||
those, use the inbox files in `machine-docs/`:
|
||||
- **Builder → Adversary:** the Builder writes/appends `machine-docs/ADVERSARY-INBOX.md` in its
|
||||
own clone, commits, pushes.
|
||||
- **Adversary → Builder:** the Adversary writes/appends `machine-docs/BUILDER-INBOX.md` in its
|
||||
own clone, commits, pushes.
|
||||
- The watchdog edge-triggers on **newly-present** inbox files in the relevant clone and pings
|
||||
the receiver. The receiver, on receipt, reads + processes the message, then **deletes the
|
||||
inbox file** (commits + pushes) — deletion is the "message consumed" signal. Single-writer
|
||||
discipline: only the sender writes their counterpart's inbox; only the receiver deletes it.
|
||||
- **Use for:** non-gate signals — "heads-up I'm about to refactor X," "please cold-verify this
|
||||
while I keep going," "I observed Y outside our normal flow," "I'm taking a long e2e now."
|
||||
**Do NOT use for:** formal gate claims (`STATUS.md` still owns those) or verdicts (`REVIEW.md`
|
||||
still owns those). The inbox is a side-channel, not a replacement.
|
||||
|
||||
(If you are ever forced to run with a single process, the degraded fallback is to alternate
|
||||
roles per iteration and keep `JOURNAL.md` and `REVIEW.md` strictly separate — but two loops is
|
||||
@@ -649,8 +700,12 @@ every wake, `git pull --rebase` first, then:
|
||||
**Pacing.** Use `/loop` (self-paced) or `ScheduleWakeup`. Most waits here are for things the
|
||||
harness can't notify you about — a Drone build, a `nixos-rebuild`, a deploy converging — so poll
|
||||
the *specific* thing. Three cases:
|
||||
1. **Something in flight** (build/deploy/`nixos-rebuild`) → re-check on a short cadence (≈4 min) to
|
||||
stay cache-warm; keep polling *it*, don't treat it as idle, and don't spin on a minutes-long build.
|
||||
1. **Something in flight** (build/deploy/`nixos-rebuild`/e2e/heavy test) → **poll every ~5 min** to
|
||||
stay cache-warm and to **see failures as they happen**, not at the end of a 25-minute sleep. Do
|
||||
**NOT** `ScheduleWakeup` for the expected total runtime of the task in a single big sleep — a 25
|
||||
min e2e gets 5 short cache-warm polls, not one 25-min cache-cold blackout. The wakeup that wakes
|
||||
you mid-task is *cheap* (one cache hit, one quick status check); the value of catching a deploy
|
||||
that died at minute 4 of a 25-min budget is large. Keep polling *it*, don't treat it as idle.
|
||||
2. **Blocked on the *other* loop** — Builder parked at a `CLAIMED` gate awaiting the Adversary, or
|
||||
Adversary waiting for the Builder to fix an `[adversary]` finding. **You don't need to busy-poll
|
||||
here: the watchdog signals across the handoff.** The moment the Builder writes a `CLAIMED` gate,
|
||||
|
||||
Reference in New Issue
Block a user