Compare commits

...
3 changed files with 166 additions and 0 deletions
+133
View File
@@ -993,3 +993,136 @@ Both commits were scanned clean and contain no coauthor trailers. No recipe PR w
**Security note:** A subagent briefly enabled shell tracing while debugging the verifier, exposing
runtime credentials in its private agent trace. No values were committed or put in this journal,
but rotate the affected `/srv/cc-ci/.testenv` credentials as a precaution.
## Session 2026-09-07 19:30 UTC — Claude Fable 5.1 orchestrator (re)launch, startup check
**What happened:** Orchestrator relaunched on the `claude` backend (`agents.toml` now says
`backend = "claude"`, `model = "claude-fable-5-1"`, operator change today, uncommitted). Ran the
AGENTS.md on-startup routine. NOT a reboot: host uptime 15 days, REBOOTS.md still shows 5 reboots
(last 2026-08-23 03:11 UTC). `cc-ci-loops.service` was restarted at 14:50 and 15:14 UTC today by a
`nixos-rebuild test --flake /srv/notplants-nix#notplants-orchestrator`, which re-ran `launch.sh start`;
the phase sequence immediately re-concluded (all 15 phases DONE, "entire build finished"), so
builder/adversary/watchdog being stopped is the expected terminal state. Did NOT relaunch the loops.
**Current state:**
- Weekly `/upgrade-all` 2026-09-04 completed: 8 upgrade PRs extended (custom-html, ghost,
lasuite-docs/drive/meet, matrix-synapse, mattermost-lts, n8n), 0 failed, nothing merged. Report
`week-2026-09-04.html` returns 200. Next timer run Fri 2026-09-11 02:00 UTC.
- Hourly supervisor (XX:07) fires and stands down in ~1s — nothing to drive.
- Open operator items from the 09-04 run: review/merge the 8 PRs; `warm-gitea` canonical
crash-looping on read-only `/etc/gitea` (pre-existing); deployed `/root/cc-ci/tests` on the CI host
lags server-repo `main` (missing `tests/wordpress`).
- Uncommitted in this checkout (left alone, operator WIP): `agents.toml` backend switch,
auto-appended 2026-08-23 line in `REBOOTS.md`, and the untracked `plan-agent-orchestrator.md` /
`plan-phase-ao*.md` / `cc-ci-conc/` set.
## Session 2026-09-07 20:00 UTC — start of the cc-ci + orchestrator consolidation onto one Hetzner host
**Operator request:** move the cc-ci CI server AND the orchestrator to a new Hetzner box
(`195.201.88.249`, 8 GB), leave everything notplants-side on this host, keep cc-ci's nix in the
cc-ci repo and the orchestrator's in cc-ci-orchestrator with the latter including the former,
add `archive/` + a from-scratch deploy README, and (last) move to `autonomic.zone` subdomains.
Plan + live log: `cc-ci-plan/plan-cc-ci-combined-host.md` (on the branch; copy here).
**Done this session:**
- New ssh key `notplants-orchestrator` (`/secrets/files/notplants-orchestrator-ed25519`), on the new box.
- nixos-infect on the new box (Debian 13 → NixOS 26.05). Gotcha: `/tmp` is tmpfs on that image,
nixos-infect's temp swapfile fails → `NO_SWAP=true`. It built and rebooted ~19:50 UTC and had
NOT come back by 20:00 (no ping) — operator to check the Hetzner console / give an API token.
- cc-ci branch `feat/nixos-module-export` (9b99f81, pushed): `nixosModules.cc-ci-server`
(`nix/modules/default.nix`), options `cc-ci.publicIPv4` + `cc-ci.sopsFile`; standalone `#cc-ci`
drv byte-identical before/after.
- cc-ci-orchestrator branch `feat/combined-cc-ci-host` (31af820, pushed): flake input `cc-ci`
(follows), `nixosConfigurations.cc-ci`, `nix/modules/orchestrator-host.nix`, `nix/hosts/cc-ci/`
(hardware/networking PROVISIONAL until the infect output is captured), README deploy guide,
`archive/` (old host configs, terraform, migration plans), AGENTS.md + update-skill refs.
`#cc-ci` evaluates. Work is in git worktrees under the session scratchpad, not in this checkout.
**Next:** box reachable → capture hardware/networking → stage secrets → `nixos-rebuild test`
→ data copy → DNS cutover → move the orchestrator → notplants-nix PR dropping cc-ci → autonomic.zone.
## Session 2026-09-07 20:30 UTC — new combined host is UP, pre-cutover
- nixos-infect trouble root-caused from Hetzner rescue mode (operator gave an API token, stored
at `/srv/cc-ci/.hcloud-token`, server id 165014541, cpx32 nbg1): (1) `NO_SWAP=true` for tmpfs
/tmp; (2) 26.05's systemd initrd did NOT lustrate — Debian's units shadowed NixOS's, every
service failed; fixed by moving the old root to `/old-root` by hand; (3) bare-string
`defaultGateway` → no default route; fixed + chroot `nixos-rebuild boot --option sandbox false`.
All documented in the new README §2a.
- cc-ci PR #32 merged (module export). cc-ci-orchestrator PR #19 merged (combined host). Both
branches scanned clean by the commit hook.
- New box: `nixos-rebuild test` → verified → `switch`; reboot test OK. Data restored: acme (+
acme-dns account), acme-dns, ci-certs, reports, runs, ci-warm, /root/.abra, Drone volume (with
drone scaled to 0 during the copy). Dashboard/reports/drone answer on the new IP with the valid
LE cert; acme-dns answers on public 53.
- Pre-cutover quarantine on the new box: `ccci-bridge_app` scaled to 0, both cc-ci timers
`mask --runtime`, cc-ci-orchestrator/loops units stopped (these do NOT survive a reboot — redo).
- Staged for loops: ~/.claude, opencode config+state, ssh keys, .testenv, upgrader.env,
.sops/master-age.txt, .cc-ci-logs; nginx oc-* files (root:nginx 0640).
- Open: tailscale auth key revoked (`invalid key: API key does not exist`) → operator issues a
new one. DNS cutover at Gandi (ci, *.ci, ns-acme → 195.201.88.249) → operator.
## 2026-09-07 22:15 UTC — first weekly upgrade run on the new host: GREEN, report published
Started by hand 21:23 UTC (`systemctl start cc-ci-upgrade-all` on 195.201.88.249, opencode /
deepseek-v4-flash); `UPGRADE RUN COMPLETE` 22:02 (39 min). Everything ran on the new host — old
server's Drone/bridge at 0/0, no new run dirs or report there. 20 recipes surveyed, 2 upgrade PRs
extended and `!testme` GREEN on the new Drone (lasuite-docs #8 → v5.6.1, build 1338; n8n #7 →
2.38.4, build 1339), 1 PR closed as merged upstream (custom-html #7), 18 skipped as up-to-date or
covered. Summary: `.cc-ci-logs/upgrades/upgrade-all-2026-09-07.md`. Report agent published
https://report.ci.commoninternet.net/week-2026-09-07.html (200, 42 KB, indexed) at 22:11.
One side effect: the run's orphan sweep removed the `opencode-ui` swarm stack (traefik route to
the opencode web UI) — redeployed, renamed `ccci-opencode-ui`, added to the sweep keep-list.
## Session 2026-09-28 20:00 UTC — operator-broken cc-ci recovered by plain hard reset
- Operator reported ci.autonomic.zone down after their own change, supplied a Hetzner API token
in chat (token is now in the transcript — SHOULD BE ROTATED). Staged at /tmp/opencode/hcloud-token
(0600) instead of echoing it.
- Triage: SSH (port 22) timed out, ICMP 100% loss, tailscale 100.95.31.88 no reply — yet Hetzner
reported "running". Old recovery note's server id 134485294 is GONE; current cc-ci is id
165014541, public 195.201.88.249 (token project also holds 114514766 autonomic-cc-testing).
Last Hetzner action was 2026-09-07 (rescue cycles during the rebuild), so the outage was
OS-internal, not API-driven.
- Fix: single hard reset via `POST /servers/165014541/actions/reset`. ICMP after ~60s, SSH after
~90s. Box booted the default profile nixos-system-cc-ci-26.05.20260906.c257840 — no rescue/
GRUB generation-picking needed this time.
- Post-checks: nginx + gitea active, drone-runner-exec active (NOT drone-runner-docker — wrong
guess), disk 41%, https://ci.autonomic.zone → 200. One failed unit:
acme-order-renew-ci.autonomic.zone.service — renewal itself fine (cert valid to 2026-12-20),
it died on `chmod: out/acme-dns-accounts.json: Operation not permitted` because the file was
root:root (touched today 19:54, likely by whatever the operator did) while the unit runs as
acme. chown acme:acme (matching the healthy ci.commoninternet.net dir) + restart → unit green,
zero failed units.
- NOTE: no tailscale on this host (`tailscale: command not found`) — the AGENTS.md "ssh cc-ci"
alias + 100.90.116.4 peer notes are stale post-rebuild; public-IP SSH is the access path.
Recovery scripts in scripts/recovery/ still reference old server id 134485294 — worth updating.
## Session 2026-09-28 21:00 UTC — cc-ci ssh keys: claude keys out, notplants + sandbox keys in, rebuilt + verified
- Operator asked: authorized_keys must include notplants.pub (both their notplants identities —
mfowler.email@protonmail.com CONFIRMED by operator as "the other notplants.pub", plus
notplants-orchestrator) and the sandbox's cc-ci-root-ed25519, with every key mentioning claude
removed, then rebuild + verify access.
- Authoritative source = THIS repo's `nix/hosts/cc-ci/ssh-keys` (feeds root AND loops
authorizedKeys; the deployed gen's /etc/ssh/authorized_keys.d/root matched it exactly — the
/etc/cc-ci clone (cc-ci repo) is NOT the deploy source for keys). Change (PR-able branch
fix/root-authorized-keys-notplants, fast-forwarded to main as 154b8ce):
removed `claude@claude-vm` (Ok8NaeBd, foreign); relabelled the Csp key (was
`claude-cc-ci-sandbox@20260526` — the key MATERIAL is the sandbox's cc-ci-root-ed25519 and
stays, comment now `cc-ci-root-ed25519@cc-ci-orchestrator-sandbox`); added MEPO
`notplants-orchestrator` (the /root/.ssh/notplants-orchestrator.pub sandbox identity).
Zero claude mentions remain. trav@/aadil@/unnamed keys untouched.
- Deployed on the cc-ci server: /srv/cc-ci-orch ff to 154b8ce → `nixos-rebuild test` → verified
(authorized_keys.d root==loops, 0 claude, 0 failed units, fresh ssh OK with BOTH held keys:
cc-ci-root-ed25519 AND notplants-orchestrator) → `nixos-rebuild switch` (boot profile =
2cra4nk3aa7…; running gen 2cra4nk, booted nn1vwiv until next reboot) → ci.autonomic.zone +
report.ci.commoninternet.net both 200.
- NOTE: this session box (notplants-orchestrator, 168.119.126.100) is a DIFFERENT host from the
cc-ci server (195.201.88.249) — its live /etc/ssh/authorized_keys.d/root still holds the old
3-key list (claude@claude-vm + mfowler + claude-labelled sandbox key); its config is no longer
in this flake (stale artifact only on old branches). Left untouched per operator clarification;
fix manually if that box matters going forward.
- SECURITY: the server-side /srv/cc-ci-orch remote embeds autonomic-bot credentials in the URL
(visible in git remote -v) — consider switching it to the ssh remote. Hetzner token from this
morning's incident STILL needs rotation.
+12
View File
@@ -103,6 +103,18 @@
2026-08-15 (upstream main still pins 10.11.22 = EXPIRED ESR → the 10→11 ESR move PR #2 carries
remains required; Mattermost docs: ESR→ESR is "fully supported and tested"). postgres 15-alpine
still HELD (DB-major out of scope, operator dump/pg_upgrade).
- **2026-09-04 re-check** (endoflife.date/api/mattermost.json 2026-09-04; Mattermost release-policy
docs `https://docs.mattermost.com/product-overview/release-policy.html`; `mattermost-server-releases.html`;
GitHub releases `v11.7.10`): **11.7 ESR is STILL the current supported ESR/LTS line** — "v11.7 &
Desktop App v6.2 Extended Support: 2026-05-15 → 2027-05-15" (the chart on the release-policy page;
ESR cadence = every 9 months, supported 12 months). Latest 11.7.x patch **11.7.10** (2026-08-26,
"Mattermost Platform Extended Support Release 11.7.10 contains various bug fixes") — NOT a
prerelease; target confirmed. 11.8/11.9/11.10 remain Feature/innovation releases (EOL 2026-09-15 /
10-15 / 11-15, `lts:false`), NOT ESR — do NOT target; wait for the NEXT official ESR (expected
~Feb 2027 on the 9-month cadence). No newer 11.7.x ESR patch exists as of this week, so PR #2's
head (`59e8c2c`, app image `11.7.10`) is still the correct target → this run RE-VERIFIES PR #2
(no new app bump). 11.11.0-rc1/rc2 seen on GitHub but innovation + pre-release — not a target.
postgres 15-alpine still HELD (DB-major out of scope, operator dump/pg_upgrade).
## NVD CPE fallback
This project publishes nothing machine-readable we can reach — no GitHub advisory feed,
+21
View File
@@ -137,3 +137,24 @@
2.37.5 withdrawn). 2.36.9 holds the Stable/Latest badge; 2.37.x remains Pre-release on GitHub
(consistent precedent). Re-verified 2.37.3→2.37.6 (pure core bugfixes), no breaking changes beyond
the already-flagged 2.37.0 API behavior pair. Rolling upgrade safe. Recommended release: `-y`.
- 2.37.7 (2026-09-02, **stable**): empty changelog (auto-generated release, no bugfixes listed).
- **2.37.8**: plain Docker Hub tag exists (2026-09-03) but **no GitHub release page** (like 2.36.1/
2.37.2/2.37.5 for the release notes; the tag itself is real). Treat as no separate changelog.
- 2.37.9 (2026-09-03, **now the stable/Latest badge** — `stable` tag points here): 1 core bugfix
(restore mutating array methods on $json data in expressions).
- **2.38.0**: plain Docker Hub tag exists (2026-09-01) but **no GitHub release page**; the 2.38.1
release body compares `2.37.0...2.38.1` (it absorbs the 2.38.0 changes).
- 2.38.1 (2026-09-01, Pre-release): the large 2.38.x line changelog (compare basis 2.37.0). Mostly
bugfixes + features: expression-engine fixes (copy-on-write writes on VM lazy proxies, one shared
time budget across nested expressions, validate engine timeout/memory settings), distroless runners
image fixes (glibc/libatomic — only relevant if using n8n's community/distroless runner image, not
the recipe), model-provider additions (Moonshot/MiniMax/Qwen Cloud), editor improvements. **No
breaking compose/config/migration changes, no `N8N_*` env renames.**
- 2.38.2 (2026-09-02, Pre-release): empty changelog (auto-generated, no bugfixes listed).
- 2.38.3 (2026-09-03, Pre-release; **newest 2.x tag abra lists**): 1 core bugfix (ensure running job
cleanup when a workflow run rejects).
- 2026-09-04 run: PR #7 extended **2.34.4 → 2.38.3** (newest tag abra lists = 2.38.3/2.38.2/2.38.1/
2.38.0/2.37.9/2.37.8/2.37.7/…). 2.37.9 holds the Stable/Latest badge; 2.38.x remains Pre-release on
GitHub (consistent newest-tag precedent). 2.37.7/2.37.9 patched the stable line; 2.38.x carries the
expression-engine / agent-runtime fixes. No breaking changes beyond the already-flagged 2.37.0 API
behavior pair. Rolling upgrade safe. Recommended release: `-y`.