Compare commits

...
3 changed files with 166 additions and 0 deletions
+133
View File
@@ -993,3 +993,136 @@ Both commits were scanned clean and contain no coauthor trailers. No recipe PR w
**Security note:** A subagent briefly enabled shell tracing while debugging the verifier, exposing **Security note:** A subagent briefly enabled shell tracing while debugging the verifier, exposing
runtime credentials in its private agent trace. No values were committed or put in this journal, runtime credentials in its private agent trace. No values were committed or put in this journal,
but rotate the affected `/srv/cc-ci/.testenv` credentials as a precaution. but rotate the affected `/srv/cc-ci/.testenv` credentials as a precaution.
## Session 2026-09-07 19:30 UTC — Claude Fable 5.1 orchestrator (re)launch, startup check
**What happened:** Orchestrator relaunched on the `claude` backend (`agents.toml` now says
`backend = "claude"`, `model = "claude-fable-5-1"`, operator change today, uncommitted). Ran the
AGENTS.md on-startup routine. NOT a reboot: host uptime 15 days, REBOOTS.md still shows 5 reboots
(last 2026-08-23 03:11 UTC). `cc-ci-loops.service` was restarted at 14:50 and 15:14 UTC today by a
`nixos-rebuild test --flake /srv/notplants-nix#notplants-orchestrator`, which re-ran `launch.sh start`;
the phase sequence immediately re-concluded (all 15 phases DONE, "entire build finished"), so
builder/adversary/watchdog being stopped is the expected terminal state. Did NOT relaunch the loops.
**Current state:**
- Weekly `/upgrade-all` 2026-09-04 completed: 8 upgrade PRs extended (custom-html, ghost,
lasuite-docs/drive/meet, matrix-synapse, mattermost-lts, n8n), 0 failed, nothing merged. Report
`week-2026-09-04.html` returns 200. Next timer run Fri 2026-09-11 02:00 UTC.
- Hourly supervisor (XX:07) fires and stands down in ~1s — nothing to drive.
- Open operator items from the 09-04 run: review/merge the 8 PRs; `warm-gitea` canonical
crash-looping on read-only `/etc/gitea` (pre-existing); deployed `/root/cc-ci/tests` on the CI host
lags server-repo `main` (missing `tests/wordpress`).
- Uncommitted in this checkout (left alone, operator WIP): `agents.toml` backend switch,
auto-appended 2026-08-23 line in `REBOOTS.md`, and the untracked `plan-agent-orchestrator.md` /
`plan-phase-ao*.md` / `cc-ci-conc/` set.
## Session 2026-09-07 20:00 UTC — start of the cc-ci + orchestrator consolidation onto one Hetzner host
**Operator request:** move the cc-ci CI server AND the orchestrator to a new Hetzner box
(`195.201.88.249`, 8 GB), leave everything notplants-side on this host, keep cc-ci's nix in the
cc-ci repo and the orchestrator's in cc-ci-orchestrator with the latter including the former,
add `archive/` + a from-scratch deploy README, and (last) move to `autonomic.zone` subdomains.
Plan + live log: `cc-ci-plan/plan-cc-ci-combined-host.md` (on the branch; copy here).
**Done this session:**
- New ssh key `notplants-orchestrator` (`/secrets/files/notplants-orchestrator-ed25519`), on the new box.
- nixos-infect on the new box (Debian 13 → NixOS 26.05). Gotcha: `/tmp` is tmpfs on that image,
nixos-infect's temp swapfile fails → `NO_SWAP=true`. It built and rebooted ~19:50 UTC and had
NOT come back by 20:00 (no ping) — operator to check the Hetzner console / give an API token.
- cc-ci branch `feat/nixos-module-export` (9b99f81, pushed): `nixosModules.cc-ci-server`
(`nix/modules/default.nix`), options `cc-ci.publicIPv4` + `cc-ci.sopsFile`; standalone `#cc-ci`
drv byte-identical before/after.
- cc-ci-orchestrator branch `feat/combined-cc-ci-host` (31af820, pushed): flake input `cc-ci`
(follows), `nixosConfigurations.cc-ci`, `nix/modules/orchestrator-host.nix`, `nix/hosts/cc-ci/`
(hardware/networking PROVISIONAL until the infect output is captured), README deploy guide,
`archive/` (old host configs, terraform, migration plans), AGENTS.md + update-skill refs.
`#cc-ci` evaluates. Work is in git worktrees under the session scratchpad, not in this checkout.
**Next:** box reachable → capture hardware/networking → stage secrets → `nixos-rebuild test`
→ data copy → DNS cutover → move the orchestrator → notplants-nix PR dropping cc-ci → autonomic.zone.
## Session 2026-09-07 20:30 UTC — new combined host is UP, pre-cutover
- nixos-infect trouble root-caused from Hetzner rescue mode (operator gave an API token, stored
at `/srv/cc-ci/.hcloud-token`, server id 165014541, cpx32 nbg1): (1) `NO_SWAP=true` for tmpfs
/tmp; (2) 26.05's systemd initrd did NOT lustrate — Debian's units shadowed NixOS's, every
service failed; fixed by moving the old root to `/old-root` by hand; (3) bare-string
`defaultGateway` → no default route; fixed + chroot `nixos-rebuild boot --option sandbox false`.
All documented in the new README §2a.
- cc-ci PR #32 merged (module export). cc-ci-orchestrator PR #19 merged (combined host). Both
branches scanned clean by the commit hook.
- New box: `nixos-rebuild test` → verified → `switch`; reboot test OK. Data restored: acme (+
acme-dns account), acme-dns, ci-certs, reports, runs, ci-warm, /root/.abra, Drone volume (with
drone scaled to 0 during the copy). Dashboard/reports/drone answer on the new IP with the valid
LE cert; acme-dns answers on public 53.
- Pre-cutover quarantine on the new box: `ccci-bridge_app` scaled to 0, both cc-ci timers
`mask --runtime`, cc-ci-orchestrator/loops units stopped (these do NOT survive a reboot — redo).
- Staged for loops: ~/.claude, opencode config+state, ssh keys, .testenv, upgrader.env,
.sops/master-age.txt, .cc-ci-logs; nginx oc-* files (root:nginx 0640).
- Open: tailscale auth key revoked (`invalid key: API key does not exist`) → operator issues a
new one. DNS cutover at Gandi (ci, *.ci, ns-acme → 195.201.88.249) → operator.
## 2026-09-07 22:15 UTC — first weekly upgrade run on the new host: GREEN, report published
Started by hand 21:23 UTC (`systemctl start cc-ci-upgrade-all` on 195.201.88.249, opencode /
deepseek-v4-flash); `UPGRADE RUN COMPLETE` 22:02 (39 min). Everything ran on the new host — old
server's Drone/bridge at 0/0, no new run dirs or report there. 20 recipes surveyed, 2 upgrade PRs
extended and `!testme` GREEN on the new Drone (lasuite-docs #8 → v5.6.1, build 1338; n8n #7 →
2.38.4, build 1339), 1 PR closed as merged upstream (custom-html #7), 18 skipped as up-to-date or
covered. Summary: `.cc-ci-logs/upgrades/upgrade-all-2026-09-07.md`. Report agent published
https://report.ci.commoninternet.net/week-2026-09-07.html (200, 42 KB, indexed) at 22:11.
One side effect: the run's orphan sweep removed the `opencode-ui` swarm stack (traefik route to
the opencode web UI) — redeployed, renamed `ccci-opencode-ui`, added to the sweep keep-list.
## Session 2026-09-28 20:00 UTC — operator-broken cc-ci recovered by plain hard reset
- Operator reported ci.autonomic.zone down after their own change, supplied a Hetzner API token
in chat (token is now in the transcript — SHOULD BE ROTATED). Staged at /tmp/opencode/hcloud-token
(0600) instead of echoing it.
- Triage: SSH (port 22) timed out, ICMP 100% loss, tailscale 100.95.31.88 no reply — yet Hetzner
reported "running". Old recovery note's server id 134485294 is GONE; current cc-ci is id
165014541, public 195.201.88.249 (token project also holds 114514766 autonomic-cc-testing).
Last Hetzner action was 2026-09-07 (rescue cycles during the rebuild), so the outage was
OS-internal, not API-driven.
- Fix: single hard reset via `POST /servers/165014541/actions/reset`. ICMP after ~60s, SSH after
~90s. Box booted the default profile nixos-system-cc-ci-26.05.20260906.c257840 — no rescue/
GRUB generation-picking needed this time.
- Post-checks: nginx + gitea active, drone-runner-exec active (NOT drone-runner-docker — wrong
guess), disk 41%, https://ci.autonomic.zone → 200. One failed unit:
acme-order-renew-ci.autonomic.zone.service — renewal itself fine (cert valid to 2026-12-20),
it died on `chmod: out/acme-dns-accounts.json: Operation not permitted` because the file was
root:root (touched today 19:54, likely by whatever the operator did) while the unit runs as
acme. chown acme:acme (matching the healthy ci.commoninternet.net dir) + restart → unit green,
zero failed units.
- NOTE: no tailscale on this host (`tailscale: command not found`) — the AGENTS.md "ssh cc-ci"
alias + 100.90.116.4 peer notes are stale post-rebuild; public-IP SSH is the access path.
Recovery scripts in scripts/recovery/ still reference old server id 134485294 — worth updating.
## Session 2026-09-28 21:00 UTC — cc-ci ssh keys: claude keys out, notplants + sandbox keys in, rebuilt + verified
- Operator asked: authorized_keys must include notplants.pub (both their notplants identities —
mfowler.email@protonmail.com CONFIRMED by operator as "the other notplants.pub", plus
notplants-orchestrator) and the sandbox's cc-ci-root-ed25519, with every key mentioning claude
removed, then rebuild + verify access.
- Authoritative source = THIS repo's `nix/hosts/cc-ci/ssh-keys` (feeds root AND loops
authorizedKeys; the deployed gen's /etc/ssh/authorized_keys.d/root matched it exactly — the
/etc/cc-ci clone (cc-ci repo) is NOT the deploy source for keys). Change (PR-able branch
fix/root-authorized-keys-notplants, fast-forwarded to main as 154b8ce):
removed `claude@claude-vm` (Ok8NaeBd, foreign); relabelled the Csp key (was
`claude-cc-ci-sandbox@20260526` — the key MATERIAL is the sandbox's cc-ci-root-ed25519 and
stays, comment now `cc-ci-root-ed25519@cc-ci-orchestrator-sandbox`); added MEPO
`notplants-orchestrator` (the /root/.ssh/notplants-orchestrator.pub sandbox identity).
Zero claude mentions remain. trav@/aadil@/unnamed keys untouched.
- Deployed on the cc-ci server: /srv/cc-ci-orch ff to 154b8ce → `nixos-rebuild test` → verified
(authorized_keys.d root==loops, 0 claude, 0 failed units, fresh ssh OK with BOTH held keys:
cc-ci-root-ed25519 AND notplants-orchestrator) → `nixos-rebuild switch` (boot profile =
2cra4nk3aa7…; running gen 2cra4nk, booted nn1vwiv until next reboot) → ci.autonomic.zone +
report.ci.commoninternet.net both 200.
- NOTE: this session box (notplants-orchestrator, 168.119.126.100) is a DIFFERENT host from the
cc-ci server (195.201.88.249) — its live /etc/ssh/authorized_keys.d/root still holds the old
3-key list (claude@claude-vm + mfowler + claude-labelled sandbox key); its config is no longer
in this flake (stale artifact only on old branches). Left untouched per operator clarification;
fix manually if that box matters going forward.
- SECURITY: the server-side /srv/cc-ci-orch remote embeds autonomic-bot credentials in the URL
(visible in git remote -v) — consider switching it to the ssh remote. Hetzner token from this
morning's incident STILL needs rotation.
+12
View File
@@ -103,6 +103,18 @@
2026-08-15 (upstream main still pins 10.11.22 = EXPIRED ESR → the 10→11 ESR move PR #2 carries 2026-08-15 (upstream main still pins 10.11.22 = EXPIRED ESR → the 10→11 ESR move PR #2 carries
remains required; Mattermost docs: ESR→ESR is "fully supported and tested"). postgres 15-alpine remains required; Mattermost docs: ESR→ESR is "fully supported and tested"). postgres 15-alpine
still HELD (DB-major out of scope, operator dump/pg_upgrade). still HELD (DB-major out of scope, operator dump/pg_upgrade).
- **2026-09-04 re-check** (endoflife.date/api/mattermost.json 2026-09-04; Mattermost release-policy
docs `https://docs.mattermost.com/product-overview/release-policy.html`; `mattermost-server-releases.html`;
GitHub releases `v11.7.10`): **11.7 ESR is STILL the current supported ESR/LTS line** — "v11.7 &
Desktop App v6.2 Extended Support: 2026-05-15 → 2027-05-15" (the chart on the release-policy page;
ESR cadence = every 9 months, supported 12 months). Latest 11.7.x patch **11.7.10** (2026-08-26,
"Mattermost Platform Extended Support Release 11.7.10 contains various bug fixes") — NOT a
prerelease; target confirmed. 11.8/11.9/11.10 remain Feature/innovation releases (EOL 2026-09-15 /
10-15 / 11-15, `lts:false`), NOT ESR — do NOT target; wait for the NEXT official ESR (expected
~Feb 2027 on the 9-month cadence). No newer 11.7.x ESR patch exists as of this week, so PR #2's
head (`59e8c2c`, app image `11.7.10`) is still the correct target → this run RE-VERIFIES PR #2
(no new app bump). 11.11.0-rc1/rc2 seen on GitHub but innovation + pre-release — not a target.
postgres 15-alpine still HELD (DB-major out of scope, operator dump/pg_upgrade).
## NVD CPE fallback ## NVD CPE fallback
This project publishes nothing machine-readable we can reach — no GitHub advisory feed, This project publishes nothing machine-readable we can reach — no GitHub advisory feed,
+21
View File
@@ -137,3 +137,24 @@
2.37.5 withdrawn). 2.36.9 holds the Stable/Latest badge; 2.37.x remains Pre-release on GitHub 2.37.5 withdrawn). 2.36.9 holds the Stable/Latest badge; 2.37.x remains Pre-release on GitHub
(consistent precedent). Re-verified 2.37.3→2.37.6 (pure core bugfixes), no breaking changes beyond (consistent precedent). Re-verified 2.37.3→2.37.6 (pure core bugfixes), no breaking changes beyond
the already-flagged 2.37.0 API behavior pair. Rolling upgrade safe. Recommended release: `-y`. the already-flagged 2.37.0 API behavior pair. Rolling upgrade safe. Recommended release: `-y`.
- 2.37.7 (2026-09-02, **stable**): empty changelog (auto-generated release, no bugfixes listed).
- **2.37.8**: plain Docker Hub tag exists (2026-09-03) but **no GitHub release page** (like 2.36.1/
2.37.2/2.37.5 for the release notes; the tag itself is real). Treat as no separate changelog.
- 2.37.9 (2026-09-03, **now the stable/Latest badge** — `stable` tag points here): 1 core bugfix
(restore mutating array methods on $json data in expressions).
- **2.38.0**: plain Docker Hub tag exists (2026-09-01) but **no GitHub release page**; the 2.38.1
release body compares `2.37.0...2.38.1` (it absorbs the 2.38.0 changes).
- 2.38.1 (2026-09-01, Pre-release): the large 2.38.x line changelog (compare basis 2.37.0). Mostly
bugfixes + features: expression-engine fixes (copy-on-write writes on VM lazy proxies, one shared
time budget across nested expressions, validate engine timeout/memory settings), distroless runners
image fixes (glibc/libatomic — only relevant if using n8n's community/distroless runner image, not
the recipe), model-provider additions (Moonshot/MiniMax/Qwen Cloud), editor improvements. **No
breaking compose/config/migration changes, no `N8N_*` env renames.**
- 2.38.2 (2026-09-02, Pre-release): empty changelog (auto-generated, no bugfixes listed).
- 2.38.3 (2026-09-03, Pre-release; **newest 2.x tag abra lists**): 1 core bugfix (ensure running job
cleanup when a workflow run rejects).
- 2026-09-04 run: PR #7 extended **2.34.4 → 2.38.3** (newest tag abra lists = 2.38.3/2.38.2/2.38.1/
2.38.0/2.37.9/2.37.8/2.37.7/…). 2.37.9 holds the Stable/Latest badge; 2.38.x remains Pre-release on
GitHub (consistent newest-tag precedent). 2.37.7/2.37.9 patched the stable line; 2.38.x carries the
expression-engine / agent-runtime fixes. No breaking changes beyond the already-flagged 2.37.0 API
behavior pair. Rolling upgrade safe. Recommended release: `-y`.