JOURNAL: 2026-09-28 cc-ci outage recovered via Hetzner hard reset (server 165014541)

This commit is contained in:
autonomic-bot
2026-09-28 19:57:50 +00:00
parent aa2c6dbf98
commit bafa9be01c
+104
View File
@@ -993,3 +993,107 @@ Both commits were scanned clean and contain no coauthor trailers. No recipe PR w
**Security note:** A subagent briefly enabled shell tracing while debugging the verifier, exposing **Security note:** A subagent briefly enabled shell tracing while debugging the verifier, exposing
runtime credentials in its private agent trace. No values were committed or put in this journal, runtime credentials in its private agent trace. No values were committed or put in this journal,
but rotate the affected `/srv/cc-ci/.testenv` credentials as a precaution. but rotate the affected `/srv/cc-ci/.testenv` credentials as a precaution.
## Session 2026-09-07 19:30 UTC — Claude Fable 5.1 orchestrator (re)launch, startup check
**What happened:** Orchestrator relaunched on the `claude` backend (`agents.toml` now says
`backend = "claude"`, `model = "claude-fable-5-1"`, operator change today, uncommitted). Ran the
AGENTS.md on-startup routine. NOT a reboot: host uptime 15 days, REBOOTS.md still shows 5 reboots
(last 2026-08-23 03:11 UTC). `cc-ci-loops.service` was restarted at 14:50 and 15:14 UTC today by a
`nixos-rebuild test --flake /srv/notplants-nix#notplants-orchestrator`, which re-ran `launch.sh start`;
the phase sequence immediately re-concluded (all 15 phases DONE, "entire build finished"), so
builder/adversary/watchdog being stopped is the expected terminal state. Did NOT relaunch the loops.
**Current state:**
- Weekly `/upgrade-all` 2026-09-04 completed: 8 upgrade PRs extended (custom-html, ghost,
lasuite-docs/drive/meet, matrix-synapse, mattermost-lts, n8n), 0 failed, nothing merged. Report
`week-2026-09-04.html` returns 200. Next timer run Fri 2026-09-11 02:00 UTC.
- Hourly supervisor (XX:07) fires and stands down in ~1s — nothing to drive.
- Open operator items from the 09-04 run: review/merge the 8 PRs; `warm-gitea` canonical
crash-looping on read-only `/etc/gitea` (pre-existing); deployed `/root/cc-ci/tests` on the CI host
lags server-repo `main` (missing `tests/wordpress`).
- Uncommitted in this checkout (left alone, operator WIP): `agents.toml` backend switch,
auto-appended 2026-08-23 line in `REBOOTS.md`, and the untracked `plan-agent-orchestrator.md` /
`plan-phase-ao*.md` / `cc-ci-conc/` set.
## Session 2026-09-07 20:00 UTC — start of the cc-ci + orchestrator consolidation onto one Hetzner host
**Operator request:** move the cc-ci CI server AND the orchestrator to a new Hetzner box
(`195.201.88.249`, 8 GB), leave everything notplants-side on this host, keep cc-ci's nix in the
cc-ci repo and the orchestrator's in cc-ci-orchestrator with the latter including the former,
add `archive/` + a from-scratch deploy README, and (last) move to `autonomic.zone` subdomains.
Plan + live log: `cc-ci-plan/plan-cc-ci-combined-host.md` (on the branch; copy here).
**Done this session:**
- New ssh key `notplants-orchestrator` (`/secrets/files/notplants-orchestrator-ed25519`), on the new box.
- nixos-infect on the new box (Debian 13 → NixOS 26.05). Gotcha: `/tmp` is tmpfs on that image,
nixos-infect's temp swapfile fails → `NO_SWAP=true`. It built and rebooted ~19:50 UTC and had
NOT come back by 20:00 (no ping) — operator to check the Hetzner console / give an API token.
- cc-ci branch `feat/nixos-module-export` (9b99f81, pushed): `nixosModules.cc-ci-server`
(`nix/modules/default.nix`), options `cc-ci.publicIPv4` + `cc-ci.sopsFile`; standalone `#cc-ci`
drv byte-identical before/after.
- cc-ci-orchestrator branch `feat/combined-cc-ci-host` (31af820, pushed): flake input `cc-ci`
(follows), `nixosConfigurations.cc-ci`, `nix/modules/orchestrator-host.nix`, `nix/hosts/cc-ci/`
(hardware/networking PROVISIONAL until the infect output is captured), README deploy guide,
`archive/` (old host configs, terraform, migration plans), AGENTS.md + update-skill refs.
`#cc-ci` evaluates. Work is in git worktrees under the session scratchpad, not in this checkout.
**Next:** box reachable → capture hardware/networking → stage secrets → `nixos-rebuild test`
→ data copy → DNS cutover → move the orchestrator → notplants-nix PR dropping cc-ci → autonomic.zone.
## Session 2026-09-07 20:30 UTC — new combined host is UP, pre-cutover
- nixos-infect trouble root-caused from Hetzner rescue mode (operator gave an API token, stored
at `/srv/cc-ci/.hcloud-token`, server id 165014541, cpx32 nbg1): (1) `NO_SWAP=true` for tmpfs
/tmp; (2) 26.05's systemd initrd did NOT lustrate — Debian's units shadowed NixOS's, every
service failed; fixed by moving the old root to `/old-root` by hand; (3) bare-string
`defaultGateway` → no default route; fixed + chroot `nixos-rebuild boot --option sandbox false`.
All documented in the new README §2a.
- cc-ci PR #32 merged (module export). cc-ci-orchestrator PR #19 merged (combined host). Both
branches scanned clean by the commit hook.
- New box: `nixos-rebuild test` → verified → `switch`; reboot test OK. Data restored: acme (+
acme-dns account), acme-dns, ci-certs, reports, runs, ci-warm, /root/.abra, Drone volume (with
drone scaled to 0 during the copy). Dashboard/reports/drone answer on the new IP with the valid
LE cert; acme-dns answers on public 53.
- Pre-cutover quarantine on the new box: `ccci-bridge_app` scaled to 0, both cc-ci timers
`mask --runtime`, cc-ci-orchestrator/loops units stopped (these do NOT survive a reboot — redo).
- Staged for loops: ~/.claude, opencode config+state, ssh keys, .testenv, upgrader.env,
.sops/master-age.txt, .cc-ci-logs; nginx oc-* files (root:nginx 0640).
- Open: tailscale auth key revoked (`invalid key: API key does not exist`) → operator issues a
new one. DNS cutover at Gandi (ci, *.ci, ns-acme → 195.201.88.249) → operator.
## 2026-09-07 22:15 UTC — first weekly upgrade run on the new host: GREEN, report published
Started by hand 21:23 UTC (`systemctl start cc-ci-upgrade-all` on 195.201.88.249, opencode /
deepseek-v4-flash); `UPGRADE RUN COMPLETE` 22:02 (39 min). Everything ran on the new host — old
server's Drone/bridge at 0/0, no new run dirs or report there. 20 recipes surveyed, 2 upgrade PRs
extended and `!testme` GREEN on the new Drone (lasuite-docs #8 → v5.6.1, build 1338; n8n #7 →
2.38.4, build 1339), 1 PR closed as merged upstream (custom-html #7), 18 skipped as up-to-date or
covered. Summary: `.cc-ci-logs/upgrades/upgrade-all-2026-09-07.md`. Report agent published
https://report.ci.commoninternet.net/week-2026-09-07.html (200, 42 KB, indexed) at 22:11.
One side effect: the run's orphan sweep removed the `opencode-ui` swarm stack (traefik route to
the opencode web UI) — redeployed, renamed `ccci-opencode-ui`, added to the sweep keep-list.
## Session 2026-09-28 20:00 UTC — operator-broken cc-ci recovered by plain hard reset
- Operator reported ci.autonomic.zone down after their own change, supplied a Hetzner API token
in chat (token is now in the transcript — SHOULD BE ROTATED). Staged at /tmp/opencode/hcloud-token
(0600) instead of echoing it.
- Triage: SSH (port 22) timed out, ICMP 100% loss, tailscale 100.95.31.88 no reply — yet Hetzner
reported "running". Old recovery note's server id 134485294 is GONE; current cc-ci is id
165014541, public 195.201.88.249 (token project also holds 114514766 autonomic-cc-testing).
Last Hetzner action was 2026-09-07 (rescue cycles during the rebuild), so the outage was
OS-internal, not API-driven.
- Fix: single hard reset via `POST /servers/165014541/actions/reset`. ICMP after ~60s, SSH after
~90s. Box booted the default profile nixos-system-cc-ci-26.05.20260906.c257840 — no rescue/
GRUB generation-picking needed this time.
- Post-checks: nginx + gitea active, drone-runner-exec active (NOT drone-runner-docker — wrong
guess), disk 41%, https://ci.autonomic.zone → 200. One failed unit:
acme-order-renew-ci.autonomic.zone.service — renewal itself fine (cert valid to 2026-12-20),
it died on `chmod: out/acme-dns-accounts.json: Operation not permitted` because the file was
root:root (touched today 19:54, likely by whatever the operator did) while the unit runs as
acme. chown acme:acme (matching the healthy ci.commoninternet.net dir) + restart → unit green,
zero failed units.
- NOTE: no tailscale on this host (`tailscale: command not found`) — the AGENTS.md "ssh cc-ci"
alias + 100.90.116.4 peer notes are stale post-rebuild; public-IP SSH is the access path.
Recovery scripts in scripts/recovery/ still reference old server id 134485294 — worth updating.