The 2026-08-03 cc-ci outage was recovered with ad-hoc tooling living in /tmp (leftover from a PREVIOUS incident, half-evaporated). Promoted to the repo: - scripts/recovery/hetzner.py — Hetzner API helper (status/actions/reboot/reset/ power/rescue-on|off/console), knows cc-ci=134485294 + orchestrator=134487234 by name; token from HCLOUD_TOKEN or /srv/cc-ci/.hcloud-token (0600, never in git). - scripts/recovery/hetzner-console.sh — shell-only VGA console: fresh console session -> websocat bridge -> vncdotool (venv auto-bootstrapped in ~/.cache). screenshot / key / type subcommands; encodes the reset-invalidates-session and single-connection-bridge gotchas. - scripts/recovery/README.md — the condensed 10-minute unreachable-server drill, incl. the GRUB submenu 1>N ids + clear-grubenv-after-switch rule. - hetzner-server-recovery skill: console/API sections now point at the repo tools instead of describing /tmp rebuilds. Smoke-tested: hetzner.py cc-ci status OK.
2.5 KiB
2.5 KiB
Recovery tooling — permanent home
Promoted from /tmp ad-hocery after the 2026-08-03 cc-ci 26.05 outage (see
.cc-ci-logs/server-update-2026-08-03.md and the hetzner-server-recovery skill, which is
the procedure; these are the tools it uses).
| Tool | What it does |
|---|---|
hetzner.py <server> <cmd> |
Hetzner Cloud API: status, actions, reboot/reset/poweroff/poweron, rescue-on/off, console credentials. Knows cc-ci (134485294) and orchestrator (134487234) by name. |
hetzner-console.sh <server> screenshot|key|type |
Shell-only access to the VGA console: fresh console session → websocat bridge → vncdotool (venv auto-bootstrapped at ~/.cache/hetzner-console-venv). |
Token: HCLOUD_TOKEN env, or /srv/cc-ci/.hcloud-token (chmod 600). Not in git, not in
.testenv. Prefer per-incident tokens and revoke after — and never paste tokens into chat
transcripts (the 2026-08-03 incident token was pasted mid-incident and had to be flagged for
rotation).
The 10-minute unreachable-server drill (condensed from 2026-08-03)
hetzner.py cc-ci status— "running" + no SSH/ping means booted-but-broken, not crashed.hetzner-console.sh cc-ci screenshot /tmp/console.png— look at the actual screen: which generation booted? login prompt or emergency shell?- Plain
rebootfirst. If the default boot is the broken generation, DON'T fight GRUB timing over VNC — go to rescue:rescue-on→poweroff→ waitstatus=off→poweron→ssh root@<public-ip>(key 113082420 =~/.ssh/cc-ci-root-ed25519; freshUserKnownHostsFile). - In rescue:
mount /dev/sda1 /mntand fix the boot default:- Generations live in a GRUB submenu: entry ids are
1>N(top level: 0 = default entry, 1 = the submenu). A bare index like2silently falls back to the broken default. - Persistent:
grub-editenv /mnt/boot/grub/grubenv set 'default=1>N'(survives reboots). - Clear it after the next successful
nixos-rebuild switch— the regenerated grub.cfg shifts indices and a stale override points at the wrong generation. - Journal of the failed boot:
journalctl -D /mnt/var/log/journal --list-boots/-b <id>.
- Generations live in a GRUB submenu: entry ids are
rescue-off→poweroff→poweron→ verify → write the incident up incc-ci-plan/JOURNAL.mdand (if server).cc-ci-logs/.
Prevention: nixos-rebuild test before switch, always (see AGENTS.md / the update
skills) — test leaves the bootloader alone, so a power-cycle recovers by itself.