Merge journal/weekly-upgrade-report: 2026-09-28 outage recovery + Sep upstream re-check notes

# Conflicts:
#	cc-ci-plan/JOURNAL.md
#	cc-ci-plan/upstream/n8n.md
This commit is contained in:
autonomic-bot
2026-09-28 19:58:53 +00:00
2 changed files with 35 additions and 0 deletions
+23
View File
@@ -1202,3 +1202,26 @@ the host: `opencode-go/deepseek-v4-flash` and `opencode-go/glm-5.3-flash` answer
`upgrader.env` (`LOOP_TIER=go` maps to the `opencode-go` auth entry; `LOOP_MODEL` overrides the
tier default). Next fire Fri 2026-09-11 02:00 UTC.
- The steering orchestrator agent stays on `opencode-go/glm-5.2` (not asked to change).
## Session 2026-09-28 20:00 UTC — operator-broken cc-ci recovered by plain hard reset
- Operator reported ci.autonomic.zone down after their own change, supplied a Hetzner API token
in chat (token is now in the transcript — SHOULD BE ROTATED). Staged at /tmp/opencode/hcloud-token
(0600) instead of echoing it.
- Triage: SSH (port 22) timed out, ICMP 100% loss, tailscale 100.95.31.88 no reply — yet Hetzner
reported "running". Old recovery note's server id 134485294 is GONE; current cc-ci is id
165014541, public 195.201.88.249 (token project also holds 114514766 autonomic-cc-testing).
Last Hetzner action was 2026-09-07 (rescue cycles during the rebuild), so the outage was
OS-internal, not API-driven.
- Fix: single hard reset via `POST /servers/165014541/actions/reset`. ICMP after ~60s, SSH after
~90s. Box booted the default profile nixos-system-cc-ci-26.05.20260906.c257840 — no rescue/
GRUB generation-picking needed this time.
- Post-checks: nginx + gitea active, drone-runner-exec active (NOT drone-runner-docker — wrong
guess), disk 41%, https://ci.autonomic.zone → 200. One failed unit:
acme-order-renew-ci.autonomic.zone.service — renewal itself fine (cert valid to 2026-12-20),
it died on `chmod: out/acme-dns-accounts.json: Operation not permitted` because the file was
root:root (touched today 19:54, likely by whatever the operator did) while the unit runs as
acme. chown acme:acme (matching the healthy ci.commoninternet.net dir) + restart → unit green,
zero failed units.
- NOTE: no tailscale on this host (`tailscale: command not found`) — the AGENTS.md "ssh cc-ci"
alias + 100.90.116.4 peer notes are stale post-rebuild; public-IP SSH is the access path.
Recovery scripts in scripts/recovery/ still reference old server id 134485294 — worth updating.