Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
aa2c6dbf98 | ||
|
|
bc26c64065 |
@@ -13,8 +13,8 @@ a health gate, not as silent drift.
|
||||
|
||||
> **Two hosts, two flakes — don't confuse them.** This skill updates the **orchestrator** host:
|
||||
> the machine this session runs on (`cc-ci-orchestrator-1`, Hetzner cpx22 **server 134487234**,
|
||||
> tailnet `cc-ci`, public `195.201.88.249` — the SAME host as the cc-ci CI server since 2026-09), flake checkout **`/srv/cc-ci-orch`** (repo
|
||||
> `recipe-maintainers/cc-ci-orchestrator`), target **`.#cc-ci`** (which now also rebuilds the CI server half, from the cc-ci repo flake input). The **cc-ci
|
||||
> tailnet `100.84.190.30`, public `168.119.126.100`), flake checkout **`/srv/cc-ci-orch`** (repo
|
||||
> `recipe-maintainers/cc-ci-orchestrator`), target **`.#cc-ci-orchestrator-hetzner`**. The **cc-ci
|
||||
> CI server** (`ssh cc-ci`, repo `recipe-maintainers/cc-ci`, target `.#cc-ci`) is a different
|
||||
> machine — that's `/cc-ci-server-update`, NOT this skill.
|
||||
|
||||
@@ -72,7 +72,7 @@ deploy).
|
||||
### 3. Build (catch errors before any activation)
|
||||
|
||||
```
|
||||
cd /srv/cc-ci-orch && nixos-rebuild build --flake .#cc-ci 2>&1 | tail -15
|
||||
cd /srv/cc-ci-orch && nixos-rebuild build --flake .#cc-ci-orchestrator-hetzner 2>&1 | tail -15
|
||||
readlink -f result
|
||||
```
|
||||
Build failure → fix on the branch (option renames etc.) before going further. Never activate a
|
||||
@@ -85,7 +85,7 @@ activation breaks the host (cf. the cc-ci server's 2026-08-03 no-default-route o
|
||||
reboot — Hetzner API power-cycle on server **134487234** if SSH is gone (see
|
||||
`hetzner-server-recovery`) — lands back on the last-known-good generation.
|
||||
```
|
||||
cd /srv/cc-ci-orch && setsid nohup nixos-rebuild test --flake .#cc-ci \
|
||||
cd /srv/cc-ci-orch && setsid nohup nixos-rebuild test --flake .#cc-ci-orchestrator-hetzner \
|
||||
> /tmp/orchestrator-test-switch.log 2>&1 < /dev/null & echo launched
|
||||
# after it settles (poll; tailscaled/sshd may blip):
|
||||
readlink /run/current-system # should be the new store path
|
||||
@@ -100,7 +100,7 @@ switch.
|
||||
### 5. Switch (make permanent — only after 4 is healthy)
|
||||
|
||||
```
|
||||
cd /srv/cc-ci-orch && nixos-rebuild switch --flake .#cc-ci 2>&1 | tail -10
|
||||
cd /srv/cc-ci-orch && nixos-rebuild switch --flake .#cc-ci-orchestrator-hetzner 2>&1 | tail -10
|
||||
```
|
||||
(If it fails with "Unit nixos-rebuild-switch-to-configuration.service was already loaded", the
|
||||
detached test's transient unit is still running — wait or `systemctl stop` it, then retry.)
|
||||
@@ -127,7 +127,7 @@ git add flake.lock # flake.nix too if the channel
|
||||
git commit -m "flake: bump nixpkgs (nixos-26.05, $(date -u +%Y-%m-%d))
|
||||
|
||||
nixpkgs: <old-rev[:8]> -> <new-rev[:8]> (nixos-26.05 tip)
|
||||
Deployed to the cc-ci host (.#cc-ci): build + test + switch + health gate green."
|
||||
Deployed to cc-ci-orchestrator-hetzner: build + test + switch + health gate green."
|
||||
git push -u origin HEAD
|
||||
```
|
||||
Open the PR on `recipe-maintainers/cc-ci-orchestrator` (Gitea API with the `GITEA_*` creds from
|
||||
|
||||
@@ -30,18 +30,15 @@ the orchestrator watches from outside.
|
||||
|
||||
Reboot resilience is handled by **`cc-ci-loops.service`** (system unit): on boot it logs the reboot
|
||||
to `REBOOTS.md` (boot_id-gated) and runs `launch.sh start` with `RESUME_PHASE=1`, so the loops +
|
||||
watchdog auto-resume the saved phase. The orchestrator session itself is relaunched by
|
||||
`cc-ci-orchestrator.service` (`agents.py up orchestrator`) — the operator reconnects to it (that's
|
||||
why the startup notification matters). Since 2026-09 the orchestrator runs on the **same Hetzner
|
||||
host as the cc-ci CI server** (`cc-ci`, public `195.201.88.249`, tailnet `cc-ci`), declared by
|
||||
`nixosConfigurations.cc-ci` in this repo's `flake.nix`, which imports the CI server from the cc-ci
|
||||
repo's `nixosModules.cc-ci-server`. `ssh cc-ci` from the loops user therefore goes to loopback.
|
||||
The full provisioning + deploy guide is `README.md`; the move is recorded in
|
||||
`cc-ci-plan/plan-cc-ci-combined-host.md`; the previous hosts (Pi → Incus VM → Hetzner `cpx22`
|
||||
shared with notplants) are in `archive/`. Rebuild this host with
|
||||
`nixos-rebuild switch --flake .#cc-ci` from `/srv/cc-ci-orch` — but **always
|
||||
watchdog auto-resume the saved phase. The orchestrator session itself is NOT auto-started — the
|
||||
operator reconnects to it (that's why the startup notification matters). The orchestrator now runs on
|
||||
a **Hetzner `cpx22`** cloud server (`cc-ci-orchestrator-1`, tailnet `100.84.190.30`, public
|
||||
`168.119.126.100`, flake host `cc-ci-orchestrator-hetzner`) — see
|
||||
`cc-ci-plan/plan-orchestrator-hetzner-migration.md`. The earlier Pi→Incus-VM move is the historical
|
||||
`cc-ci-plan/plan-orchestrator-migration.md`. Rebuild this host with
|
||||
`nixos-rebuild switch --flake .#cc-ci-orchestrator-hetzner` from `/srv/cc-ci-orch` — but **always
|
||||
`nixos-rebuild test` the same flake target first and verify the host is still healthy/reachable
|
||||
before the `switch`** (general policy for nix deploys to this host: `test`
|
||||
before the `switch`** (general policy for nix deploys to this host and the cc-ci server: `test`
|
||||
leaves the bootloader and system profile untouched, so a reboot always recovers to the
|
||||
last-known-good generation; the 2026-08-03 cc-ci 26.05 bump outage is the cautionary tale, see
|
||||
`.cc-ci-logs/server-update-2026-08-03.md`).
|
||||
|
||||
@@ -1,344 +1,47 @@
|
||||
# cc-ci-orchestrator
|
||||
|
||||
The **cc-ci orchestrator**: the agent loops that built the cc-ci Co-op Cloud recipe CI server,
|
||||
the operator's steering session, and the weekly autonomous recipe-upgrade run — plus the NixOS
|
||||
host they run on. Since 2026-09 that host is **the same Hetzner server as the CI server itself**:
|
||||
one `nixos-rebuild` from this repo builds both, because this flake imports the CI server as a
|
||||
module from the [cc-ci](https://git.autonomic.zone/recipe-maintainers/cc-ci) repo.
|
||||
Orchestrator workspace for building the **cc-ci** Co-op Cloud recipe CI server. The plan, launch
|
||||
tooling, and loop prompts live in [`cc-ci-plan/`](cc-ci-plan/); see [`AGENTS.md`](AGENTS.md) for the
|
||||
roles and operating model. Secrets (`.testenv`) are gitignored — never commit them.
|
||||
|
||||
| | where |
|
||||
|---|---|
|
||||
| Orchestrator loops, timers (weekly upgrader, hourly supervisor) | `nix/modules/cc-ci.nix` → `nixosModules.cc-ci-orchestrator` |
|
||||
| The host contract those need (loops user, claude/opencode CLIs, opencode web UI) | `nix/modules/orchestrator-host.nix` → `nixosModules.orchestrator-host` |
|
||||
| The CI server (swarm, traefik, drone, runner, `!testme` bridge, dashboard, reports, acme-dns) | cc-ci repo `nix/modules/` → `nixosModules.cc-ci-server` (flake input `cc-ci`) |
|
||||
| The machine: hardware, networking, tailscale, root keys | `nix/hosts/cc-ci/` → `nixosConfigurations.cc-ci` |
|
||||
| Plans, launch tooling, loop prompts, journal | `cc-ci-plan/` (see `AGENTS.md` for roles) |
|
||||
| Skills the orchestrator runs (`/upgrade-all`, `/recipe-upgrade`, `/cc-ci-status`, …) | `.claude/skills/`, `.opencode/skills/` |
|
||||
| How it used to be built (Pi → Incus VM → shared Hetzner box) | `archive/` |
|
||||
## Run the orchestrator in tmux (survives disconnects + closing your laptop)
|
||||
|
||||
Secrets (`.testenv`, `upgrader.env`, `.sops/`, everything under `/secrets`) are gitignored — never
|
||||
commit them.
|
||||
|
||||
---
|
||||
|
||||
# Deploying a cc-ci host from scratch
|
||||
|
||||
This is the whole path from "nothing" to a working CI server + orchestrator on one Hetzner
|
||||
server. It was last done on 2026-09-07 for `195.201.88.249` and is written so a person or an LLM
|
||||
can repeat it. Read it once before starting; the order matters.
|
||||
|
||||
## 0. What you need in hand
|
||||
|
||||
- A **Hetzner Cloud** project you can create servers in (console login or an API token).
|
||||
- **SSH keys**: yours, and the orchestrator's own key so the automation can reach the box. The
|
||||
public keys that get root are tracked in `nix/hosts/cc-ci/ssh-keys` (one per line).
|
||||
- Read access to `recipe-maintainers/cc-ci`, `recipe-maintainers/cc-ci-orchestrator` (both public
|
||||
read) and the **private** `recipe-maintainers/cc-ci-secrets` (the `autonomic-bot` deploy key,
|
||||
`autonomic-bot-gitea-ed25519`, has it).
|
||||
- The out-of-band secrets listed in §4. If you are migrating, they come from the old host; if
|
||||
you are starting fresh you create them (each row says how).
|
||||
- Control of the DNS zone (Gandi for `commoninternet.net`) for the cutover in §7.
|
||||
|
||||
## 1. Provision the server on Hetzner (Debian image)
|
||||
|
||||
In the Hetzner Cloud console (or with `hcloud server create`):
|
||||
|
||||
| setting | value | why |
|
||||
|---|---|---|
|
||||
| Image | **Debian 13** (any recent Debian/Ubuntu works with nixos-infect) | it is replaced by NixOS in §2 |
|
||||
| Type | **x86**, **8 GB RAM**, 4 vCPU — e.g. `cpx32` (dedicated AMD) or `cx33`. **Never `cax*`** (ARM): the flakes are `x86_64-linux`. | swarm + recipe deploys + 3–6 agent sessions; 4 GB is too small |
|
||||
| Disk | the type's default 150+ GB NVMe | docker layers alone are ~60 GB after a few weeks |
|
||||
| Network | public **IPv4** required; IPv6 optional (leave enabled or not, NixOS config ignores it) | cc-ci serves 80/443 and DNS on 53 publicly |
|
||||
| SSH keys | add every key from `nix/hosts/cc-ci/ssh-keys` you want to log in with, at least the orchestrator's | nixos-infect carries `/root/.ssh/authorized_keys` over |
|
||||
| Name | `cc-ci` | becomes the hostname |
|
||||
| Firewall | if a Hetzner Cloud Firewall is attached it must allow **22/tcp, 80/tcp, 443/tcp, 53/tcp, 53/udp** in, and ICMP | the NixOS firewall is separate and is configured by the flake |
|
||||
|
||||
Check you can log in: `ssh root@<ip> hostname`.
|
||||
|
||||
## 2. Convert Debian → NixOS with nixos-infect
|
||||
|
||||
[nixos-infect](https://github.com/elitak/nixos-infect) installs NixOS over the running Debian and
|
||||
reboots. Run it detached so the SSH session dropping does not kill it:
|
||||
Keep this supervising session alive on the host with tmux, and use `--remote-control` so you can
|
||||
watch/steer it from **claude.ai/code** (or the mobile app).
|
||||
|
||||
```bash
|
||||
ssh root@<ip> 'cat > /root/infect.sh <<"EOF"
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
# Pinned nixos-infect revision (same one that built the previous cc-ci hosts).
|
||||
INFECT_SHA="40f62a680bb0e8f2f607d79abfaaecd99d59401c"
|
||||
export NIX_CHANNEL="nixos-26.05" # must match the nixpkgs channel in flake.nix
|
||||
export PROVIDER="hetznercloud" # GRUB + Hetzner networking
|
||||
export NIXOS_IMPORT="" # the real config comes from the flake in §5
|
||||
# The Debian 13 cloud image mounts /tmp as tmpfs; nixos-infect makes a temporary swapfile
|
||||
# there and swapon fails with "Invalid argument". 8 GB RAM needs no extra swap: skip it.
|
||||
export NO_SWAP=true
|
||||
curl -fsSL "https://raw.githubusercontent.com/elitak/nixos-infect/${INFECT_SHA}/nixos-infect" | bash -x
|
||||
EOF
|
||||
chmod +x /root/infect.sh
|
||||
nohup /root/infect.sh > /var/log/nixos-infect.log 2>&1 &'
|
||||
# 0. Exit any running orchestrator session first — a conversation can't be resumed while it's live:
|
||||
# /exit (inside Claude) or Ctrl-D
|
||||
|
||||
# 1. Start a detachable tmux session on this host
|
||||
tmux new -s orchestrator
|
||||
|
||||
# 2. Inside tmux, resume the orchestrator conversation WITH remote control:
|
||||
claude --resume autonomous-orchestrator \
|
||||
--remote-control "autonomous-orchestrator" \
|
||||
--dangerously-skip-permissions
|
||||
# - If name-resume opens a picker instead of resuming directly, choose "autonomous-orchestrator".
|
||||
# - Or resume by the stable session id (more deterministic in a fresh pane):
|
||||
# claude --resume 34a80a99-b37e-4809-b8da-ccc9fafe785e \
|
||||
# --remote-control "autonomous-orchestrator" --dangerously-skip-permissions
|
||||
|
||||
# 3. Detach — the process keeps running: press Ctrl-b, then d
|
||||
```
|
||||
|
||||
It downloads Nix, builds a NixOS system (5–10 min; follow with
|
||||
`ssh root@<ip> tail -f /var/log/nixos-infect.log`), then reboots. The SSH host key changes:
|
||||
`ssh-keygen -R <ip>` and confirm `ssh root@<ip> nixos-version` prints a 26.05 version.
|
||||
**Reconnect later**
|
||||
- On this host: `tmux attach -t orchestrator`
|
||||
- From anywhere: **claude.ai/code** → the `autonomous-orchestrator` session
|
||||
|
||||
### 2a. What went wrong on 2026-09-07, and the fixes (Debian 13 image, NixOS 26.05)
|
||||
**Why it survives:** tmux keeps the `claude` process alive across SSH disconnects and your laptop
|
||||
closing; remote-control runs *outbound* from this host to Anthropic, so it stays connected
|
||||
regardless of the viewer. After a host reboot, re-run steps 1–2.
|
||||
|
||||
All three bit on the first attempt; the script above and §3 already include the fixes, this is
|
||||
so you recognise them if they come back in another form.
|
||||
|
||||
1. **`swapon: /tmp/nixos-infect.XXXX.swp: Invalid argument`** right at the start, script exits.
|
||||
The Debian 13 cloud image mounts `/tmp` as tmpfs and a swapfile cannot live there.
|
||||
Fix: `NO_SWAP=true` (in the script above). An 8 GB box does not need the temporary swap.
|
||||
2. **The box never comes back after the reboot: it boots NixOS, but nearly every unit fails**
|
||||
(`dbus`, `systemd-logind`, `sshd`, networking …) with
|
||||
`Could not start dynamically linked executable: /usr/bin/dbus-daemon` in the journal.
|
||||
nixos-infect leaves the old Debian root in place and relies on NixOS's first boot to move it
|
||||
to `/old-root` (`/etc/NIXOS_LUSTRATE`). With NixOS 26.05's systemd-based initrd that
|
||||
lustration did not happen, so Debian's `/etc/systemd/system/*.service` files shadowed the
|
||||
NixOS units and started Debian binaries. Fix, from Hetzner **rescue mode**
|
||||
(`enable_rescue` + `reset` in the API/console, ssh in, `mount /dev/sda1 /mnt/root`):
|
||||
move everything except `nix`, `boot`, `swapfile`, `lost+found`, `var/log`, `var/empty`,
|
||||
`etc/nixos`, `etc/resolv.conf`, `etc/NIXOS`, `etc/machine-id`, `etc/ssh/ssh_host_*`,
|
||||
`root/.nix-*`, `root/.ssh` into `/mnt/root/old-root`, delete `etc/NIXOS_LUSTRATE`, unmount,
|
||||
`disable_rescue`, `reset`. (`/old-root`, ~1 GB, can be deleted once the host is in service.)
|
||||
3. **Boots, units fine, but no network.** The generated `networking.nix` has
|
||||
`defaultGateway = "172.31.1.1";` — a bare string. Since NixOS 25.05 that yields no default
|
||||
route. Fix: `defaultGateway = { address = "172.31.1.1"; interface = "eth0"; };` (this is what
|
||||
`nix/hosts/cc-ci/networking.nix` carries). To apply it from rescue mode, chroot into the
|
||||
mounted root and rebuild the boot entry — the nix sandbox cannot `pivot_root` inside a chroot,
|
||||
so turn it off for that one build:
|
||||
```bash
|
||||
for d in proc sys dev dev/pts; do mount --bind /$d /mnt/root/$d; done
|
||||
mount -t tmpfs tmpfs /mnt/root/run; cp -L /etc/resolv.conf /mnt/root/etc/resolv.conf
|
||||
chroot /mnt/root /nix/var/nix/profiles/system/sw/bin/bash -c '
|
||||
export PATH=/nix/var/nix/profiles/system/sw/bin NIX_REMOTE= HOME=/root
|
||||
export NIX_PATH=nixos-config=/etc/nixos/configuration.nix:nixpkgs=/root/.nix-defexpr/channels/nixos
|
||||
ln -sfn /nix/var/nix/profiles/system /run/current-system
|
||||
nixos-rebuild boot --option sandbox false'
|
||||
```
|
||||
The `journalctl -D /mnt/root/var/log/journal -b 0` trick (reading the dead system's journal
|
||||
from rescue mode) is what told these apart.
|
||||
|
||||
> Rescue mode without a console: `POST /servers/<id>/actions/enable_rescue` with your ssh key
|
||||
> id, then `…/actions/reset`; afterwards `disable_rescue` **and check `rescue_enabled` is false
|
||||
> before** the next `reset`, or it boots the rescue image again. `scripts/recovery/hetzner.py`
|
||||
> wraps these (token in `/srv/cc-ci/.hcloud-token`).
|
||||
|
||||
## 3. Capture the machine-specific config into this repo
|
||||
|
||||
nixos-infect wrote `/etc/nixos/{hardware-configuration,networking,configuration}.nix`. Only the
|
||||
first two matter; the flake replaces `configuration.nix`.
|
||||
|
||||
```bash
|
||||
scp root@<ip>:/etc/nixos/hardware-configuration.nix nix/hosts/cc-ci/hardware.nix
|
||||
scp root@<ip>:/etc/nixos/networking.nix nix/hosts/cc-ci/networking.nix
|
||||
```
|
||||
|
||||
Then in `nix/hosts/cc-ci/`:
|
||||
|
||||
- `hardware.nix`: keep as generated (GRUB EFI with `efiInstallAsRemovable`, `/boot/efi` by UUID,
|
||||
`/dev/sda1` root). Do not copy another host's file — the UUIDs are per machine.
|
||||
- `networking.nix`: keep the static IPv4 + Hetzner gateway `172.31.1.1`. Make sure
|
||||
`networking.defaultGateway` has **both** `address` and `interface = "eth0"` (§2a item 3). If
|
||||
the generated IPv6 block has an empty address, delete the IPv6 parts; a real global address
|
||||
(as on the 2026-09 box) can stay.
|
||||
- `configuration.nix`: set `cc-ci.publicIPv4` to the server's IPv4 and check `system.stateVersion`
|
||||
is the release you installed (never change it later).
|
||||
- `ssh-keys`: the root keys.
|
||||
|
||||
Commit on a branch; the rebuild in §5 can use the local checkout before the PR merges.
|
||||
|
||||
## 4. Stage the workspace and secrets on the new host
|
||||
|
||||
Everything in this section is **outside git**. Do it as root over SSH, in this order.
|
||||
|
||||
### 4a. Tailscale
|
||||
|
||||
```bash
|
||||
# a reusable (or fresh) tailnet auth key from the tailscale admin console
|
||||
install -m600 /dev/stdin /etc/ts-auth-key <<<'tskey-auth-…'
|
||||
```
|
||||
|
||||
### 4b. The CI server's checkout and its one out-of-band secret
|
||||
|
||||
```bash
|
||||
# root's deploy key for the private cc-ci-secrets submodule
|
||||
install -d -m700 /root/.ssh
|
||||
install -m600 <autonomic-bot-gitea-ed25519> /root/.ssh/autonomic-bot-gitea-ed25519
|
||||
cat > /root/.ssh/config <<'EOF'
|
||||
Host git.autonomic.zone
|
||||
Port 2222
|
||||
User git
|
||||
IdentityFile /root/.ssh/autonomic-bot-gitea-ed25519
|
||||
IdentitiesOnly yes
|
||||
EOF
|
||||
# the deployed checkout: nightly-sweep runs from it, sops reads secrets/secrets.yaml from it
|
||||
git clone --recursive https://git.autonomic.zone/recipe-maintainers/cc-ci.git /etc/cc-ci
|
||||
# the master (recovery) age key — the only sops recipient a fresh host can be
|
||||
install -d -m700 /var/lib/sops-nix
|
||||
install -m600 <master-age.txt> /var/lib/sops-nix/key.txt
|
||||
```
|
||||
|
||||
`/etc/cc-ci/secrets/secrets.yaml` is encrypted to the master key and the *old* host's SSH host
|
||||
key. That is enough to deploy. Afterwards (optional, tidier) add the new host as a recipient:
|
||||
`ssh-to-age < /etc/ssh/ssh_host_ed25519_key.pub`, add it to `secrets/.sops.yaml` in cc-ci-secrets,
|
||||
`sops updatekeys secrets.yaml`, push, `git -C /etc/cc-ci submodule update --remote`.
|
||||
|
||||
### 4c. The orchestrator's workspace (as the `loops` user — it exists after the first rebuild, so
|
||||
run §5 once first if this is a fresh host, then come back)
|
||||
|
||||
```bash
|
||||
sudo -iu loops
|
||||
git clone --recursive https://git.autonomic.zone/recipe-maintainers/cc-ci-orchestrator.git /srv/cc-ci-orch
|
||||
sudo ln -sfn /srv/cc-ci-orch /srv/cc-ci # every script and unit says /srv/cc-ci
|
||||
cd /srv/cc-ci-orch
|
||||
git clone https://git.autonomic.zone/recipe-maintainers/cc-ci.git cc-ci # Builder clone
|
||||
git clone https://git.autonomic.zone/recipe-maintainers/cc-ci.git cc-ci-adv # Adversary clone
|
||||
mkdir -p .cc-ci-logs .sops
|
||||
```
|
||||
|
||||
Then the files below (`install -m600 -o loops -g users`):
|
||||
|
||||
| file | what | source |
|
||||
|---|---|---|
|
||||
| `/srv/cc-ci/.testenv` | `TS_AUTH_KEY`, `GITEA_PASSWORD` (autonomic-bot), `DOCKERHUB_USERNAME/TOKEN`, model API keys | old host `/secrets/files/cc-ci.testenv`; fresh: create each credential |
|
||||
| `/srv/cc-ci/upgrader.env` | `LOOP_TIER`, `LOOP_MODEL`, `REPORT_MODEL` for the weekly run (no secrets) | old host, or copy the example in `AGENTS.md` |
|
||||
| `/srv/cc-ci/.sops/master-age.txt` | the same master age key as 4b (skills that re-key secrets use it) | old host |
|
||||
| `~loops/.ssh/cc-ci-root-ed25519` (+`.pub`) | `ssh cc-ci` as root — to loopback on this host | old host; fresh: `ssh-keygen -t ed25519` and add the pub to `nix/hosts/cc-ci/ssh-keys` |
|
||||
| `~loops/.ssh/autonomic-bot-gitea-ed25519` (+`.pub`) | pushes recipe branches / PRs as `autonomic-bot` | old host; fresh: new key added to the bot's Gitea account |
|
||||
| `~loops/.ssh/tangled-ed25519` | optional, tangled.org mirrors | old host |
|
||||
| `~loops/.claude/` | Claude Code auth + settings + the orchestrator session history | old host (`rsync -a`); fresh: `claude auth login` as loops (device code, interactive) |
|
||||
| `~loops/.local/share/opencode/auth.json`, `~loops/.config/opencode/` | opencode provider auth (the weekly upgrader runs on opencode) | old host; fresh: `opencode auth login` |
|
||||
| `/etc/nginx/oc-selfsigned.{crt,key}`, `/etc/nginx/oc-htpasswd` | the tailnet-only opencode UI; **nginx refuses to start without them**, and its config check runs as the `nginx` user, so: `root:nginx`, crt `0644`, key + htpasswd `0640` (the `nginx` group exists after the first rebuild — fix ownership then and `systemctl restart nginx`) | old host, or generate (commands in `nix/modules/orchestrator-host.nix`) |
|
||||
|
||||
`~loops/.ssh/config` is written by the activation script on first rebuild (`Host cc-ci` →
|
||||
`127.0.0.1`, `git.autonomic.zone`, `tangled.org`); it is not overwritten if present.
|
||||
|
||||
## 5. Build and activate
|
||||
|
||||
From the checkout with the §3 commit (root can build from the loops-owned checkout via sudo):
|
||||
|
||||
```bash
|
||||
# as root, detached (the activation restarts sshd/tailscale; a dropped session must not kill it).
|
||||
# Three things the FIRST rebuild on a bare infect system needs, none of which the converged
|
||||
# host needs afterwards: `git` on PATH (nix's flake fetcher shells out to it and the infect
|
||||
# system has none — hence nix-shell), HOME=/root (so root's `git config --global
|
||||
# safe.directory '*'` applies to the loops-owned checkout), and a login shell (`bash -l`, for
|
||||
# NIX_SSL_CERT_FILE and friends from /etc/set-environment).
|
||||
git config --global --add safe.directory '*'
|
||||
systemd-run --unit=ccci-rebuild --collect -E HOME=/root -p WorkingDirectory=/srv/cc-ci-orch \
|
||||
bash -lc 'nix-shell -p git --run "nixos-rebuild test --flake /srv/cc-ci-orch#cc-ci"'
|
||||
journalctl -fu ccci-rebuild # ~10 min the first time (image pulls + two OCI image builds)
|
||||
```
|
||||
|
||||
`test` first, always: it activates WITHOUT touching the bootloader, so if the activation breaks
|
||||
networking or sshd a reboot from the Hetzner console lands on the last known-good generation.
|
||||
Later rebuilds are simply `sudo nixos-rebuild test|switch --flake .#cc-ci` from the checkout.
|
||||
|
||||
The first activation takes a while: it pulls the traefik/drone/keycloak images, builds the bridge
|
||||
and dashboard OCI images with Nix, initialises the swarm and runs the serialized reconcile
|
||||
oneshots (`swarm-init → deploy-proxy → deploy-drone → deploy-bridge → deploy-dashboard →
|
||||
deploy-reports`, `deploy-backupbot`, `warm-keycloak`). Verify:
|
||||
|
||||
```bash
|
||||
systemctl is-system-running # running — or list-units --failed and read journalctl -u <unit>
|
||||
tailscale status | head -3
|
||||
docker service ls # traefik app+socket-proxy, drone, bridge, dashboard, reports, backups: 1/1
|
||||
systemctl status cc-ci-loops cc-ci-orchestrator opencode-web nginx acme-dns
|
||||
systemctl list-timers 'cc-ci-*' nightly-sweep
|
||||
sudo -iu loops tmux ls # cc-ci-orchestrator (+ loops sessions if a phase is active)
|
||||
# the CI front doors, before DNS points here (expect 200 / 200 / 303 and ssl_verify=0 once
|
||||
# /var/lib/acme is restored or a cert has been issued):
|
||||
curl -s --resolve ci.commoninternet.net:443:127.0.0.1 -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' https://ci.commoninternet.net/
|
||||
curl -s --resolve report.ci.commoninternet.net:443:127.0.0.1 -o /dev/null -w '%{http_code}\n' https://report.ci.commoninternet.net/
|
||||
curl -s --resolve drone.ci.commoninternet.net:443:127.0.0.1 -o /dev/null -w '%{http_code}\n' https://drone.ci.commoninternet.net/
|
||||
dig +short @<ip> ns-acme.commoninternet.net # acme-dns answering on the public 53
|
||||
```
|
||||
|
||||
Seen on 2026-09-07: `tailscaled-autoconnect` failed with `invalid key: API key does not exist` —
|
||||
the reusable auth key had been revoked. Generate a fresh one in the tailscale admin console, put
|
||||
it in `/etc/ts-auth-key`, `systemctl restart tailscaled-autoconnect`. Nothing else depends on it
|
||||
during the install; the box is reachable on its public IP throughout.
|
||||
|
||||
When it is healthy: `sudo nixos-rebuild switch --flake .#cc-ci` (same config, now also the boot
|
||||
default). **If you are migrating from another host, do §6 before letting it serve anything**: right
|
||||
after the first activation scale the `!testme` bridge to 0 and mask the two orchestrator timers so
|
||||
the new box does not process PR comments or start a second weekly run while the old host is live:
|
||||
|
||||
```bash
|
||||
docker service scale ccci-bridge_app=0
|
||||
systemctl mask --now cc-ci-upgrade-all.timer cc-ci-upgrade-supervisor.timer
|
||||
```
|
||||
|
||||
## 6. Migrating: restore state from the previous host
|
||||
|
||||
Over tailscale (`rsync -aHAX --numeric-ids root@<old>:<path> <path>`), with the matching service
|
||||
stopped on the new host while its directory is copied:
|
||||
|
||||
| path | holds | notes |
|
||||
|---|---|---|
|
||||
| `/var/lib/cc-ci-reports` | the published weekly report pages (`report.ci…`) | |
|
||||
| `/var/lib/cc-ci-runs` | per-run artifacts the dashboard shows | |
|
||||
| `/var/lib/ci-warm` | warm-canonical state + alerts | recipe warm *volumes* are caches: not copied, rebuilt by the Sunday sweep / first use |
|
||||
| `/var/lib/acme` | the Let's Encrypt cert + account **and `acme-dns-accounts.json`** — the account the permanent `_acme-challenge` CNAME points at | without it a fresh registration + a new CNAME at Gandi is needed (registration is disabled in `acme-dns.nix`) |
|
||||
| `/var/lib/acme-dns` | the acme-dns zone DB | |
|
||||
| `/var/lib/ci-certs` | the copy traefik is handed | then `systemctl restart cc-ci-acme-traefik-handoff` |
|
||||
| `/root/.abra` | abra's per-app env files for the deployed stacks | |
|
||||
| Drone data volume `/var/lib/docker/volumes/drone_ci_commoninternet_net_data` | Drone's DB: the Gitea OAuth grant, repo activation, build history | `docker service scale drone_ci_commoninternet_net_app=0` on the new host, copy, scale back to 1. Otherwise run `scripts/bootstrap-drone-oauth.sh` (cc-ci repo) with the bot password and re-activate repos |
|
||||
| `/srv/cc-ci-orch/.cc-ci-logs`, `/srv/cc-ci-orch/cc-ci-plan/upstream/`, `REBOOTS.md`, `JOURNAL.md` | orchestrator history, the upgrader's per-recipe release-note registry | as loops; do the final sync after stopping the orchestrator on the old host |
|
||||
|
||||
## 7. Cutover and verification
|
||||
|
||||
1. **DNS** (operator, Gandi zone `commoninternet.net`): A records `ci`, `*.ci` and `ns-acme` →
|
||||
the new IPv4. `acme NS ns-acme` and `_acme-challenge.ci CNAME <account>.acme…` stay as they
|
||||
are. Wait for propagation (`dig +short ci.commoninternet.net`).
|
||||
2. Check the new host answers on the new IP before DNS moves: `dig @<new-ip> ns-acme.commoninternet.net`
|
||||
(acme-dns), `curl --resolve ci.commoninternet.net:443:<new-ip> https://ci.commoninternet.net/`
|
||||
(dashboard, valid cert), same for `report.ci` and `drone.ci`.
|
||||
3. Old host: `docker service scale ccci-bridge_app=0 drone_ci_commoninternet_net_app=0`;
|
||||
`systemctl disable --now cc-ci-upgrade-all.timer cc-ci-upgrade-supervisor.timer` on the old
|
||||
orchestrator. New host: `docker service scale ccci-bridge_app=1`;
|
||||
`systemctl unmask cc-ci-upgrade-all.timer cc-ci-upgrade-supervisor.timer && systemctl start` both.
|
||||
4. End to end: post `!testme` on an open recipe PR and watch it turn green on the new Drone;
|
||||
open `https://ci.commoninternet.net` and `https://report.ci.commoninternet.net`.
|
||||
5. The orchestrator: as loops on the new host `cd /srv/cc-ci-orch && python3 cc-ci-plan/agents.py up orchestrator`
|
||||
(or just `systemctl restart cc-ci-orchestrator`), attach with `claude --resume` or from
|
||||
claude.ai/code. Its startup routine (AGENTS.md) reports phase + reboot count.
|
||||
6. Keep the old host as a cold standby for a week, then delete it and its tailnet node.
|
||||
|
||||
## 8. Day 2
|
||||
|
||||
- **Update the host** (nixpkgs bump for both halves): `/cc-ci-orchestrator-update`, which is
|
||||
`nix flake update` → `nixos-rebuild test` → verify → `switch` → PR. The `cc-ci` input follows
|
||||
this flake's nixpkgs, so the CI server is rebuilt on the same nixpkgs.
|
||||
- **Update only cc-ci's code** (harness/tests/modules): merge in the cc-ci repo, then
|
||||
`nix flake update cc-ci` here and rebuild; also `git -C /etc/cc-ci pull --recurse-submodules`
|
||||
so the deployed checkout the sweep runs from matches.
|
||||
- **Something is down**: `systemctl --failed`, `journalctl -u deploy-<x>`, `docker service ps <svc>`;
|
||||
the cc-ci repo's `docs/runbook.md`. Host unreachable: Hetzner console → reboot lands on the last
|
||||
`switch`ed generation; rescue mode + `nixos-enter` for anything worse (skill
|
||||
`/hetzner-server-recovery`).
|
||||
|
||||
---
|
||||
|
||||
# Operating the orchestrator session
|
||||
|
||||
The steering session is a long-lived interactive Claude Code session under tmux with
|
||||
`--remote-control`, so it can be watched and steered from **claude.ai/code** (or the mobile app).
|
||||
`cc-ci-orchestrator.service` relaunches it on boot via `cc-ci-plan/agents.py up orchestrator`
|
||||
(backend + model in `cc-ci-plan/agents.toml`).
|
||||
|
||||
```bash
|
||||
# attach on the host
|
||||
sudo -iu loops tmux attach -t cc-ci-orchestrator
|
||||
# or resume the conversation by hand in a fresh tmux pane
|
||||
claude --resume autonomous-orchestrator --remote-control "autonomous-orchestrator" --dangerously-skip-permissions
|
||||
# already inside a live session and just want the web surface? /remote-control
|
||||
```
|
||||
|
||||
`--resume <name|id>` selects the *conversation* to restore; the `--remote-control "<name>"` value is
|
||||
only the web display label. Don't pass `--fork-session` unless you mean to branch.
|
||||
> Two different "names": `--resume <name|id>` selects the *conversation* to restore (shown in the
|
||||
> `/resume` picker); the `--remote-control "<name>"` value is only the web display label and resumes
|
||||
> nothing. Resuming reuses the same session id each time (stays `34a8…`) — don't pass
|
||||
> `--fork-session` unless you intend to branch a new conversation.
|
||||
>
|
||||
> Already inside a live session and just want the web surface? Run `/remote-control` — no exit/resume.
|
||||
|
||||
## Kick off / supervise the loops
|
||||
|
||||
@@ -350,5 +53,5 @@ cd /srv/cc-ci/cc-ci-plan
|
||||
./launch.sh stop
|
||||
```
|
||||
|
||||
Full supervision guide, credential map and history are in `cc-ci-plan/kickoff.md`,
|
||||
`cc-ci-plan/plan.md` §1.5 and `cc-ci-plan/JOURNAL.md`.
|
||||
Full supervision guide, credential map, and the Incus VM fallback are in
|
||||
[`cc-ci-plan/kickoff.md`](cc-ci-plan/kickoff.md) and [`cc-ci-plan/plan.md`](cc-ci-plan/plan.md) §1.5.
|
||||
|
||||
@@ -1,22 +0,0 @@
|
||||
# archive/ — how cc-ci and its orchestrator were built and moved, before the combined host
|
||||
|
||||
Historical record only. Nothing in here is deployed or evaluated. It was moved out of the live
|
||||
tree on 2026-09-07 when the CI server and the orchestrator were consolidated onto one Hetzner
|
||||
host (`nixosConfigurations.cc-ci` in `../flake.nix`; deploy guide in `../README.md`; the plan
|
||||
that did it is `../cc-ci-plan/plan-cc-ci-combined-host.md`).
|
||||
|
||||
| path | what it was |
|
||||
|---|---|
|
||||
| `nix/configuration-incus-vm.nix` | Channel-based NixOS config of the first orchestrator VM on b1 (Incus, 2 GB). Ran the loops as root; hard-coded the dead Incus cc-ci IP. Replaced by the Hetzner host 2026-05-31. |
|
||||
| `nix/README.md` | The README for that Incus VM config. |
|
||||
| `nix/cc-ci-orchestrator-hetzner/` | The orchestrator's own Hetzner `cpx22` host (`168.119.126.100`, tailnet `cc-ci-orchestrator-1`), 2026-05-31 → 2026-09. From 2026-08-20 the live copy of this config was `notplants-nix`'s `notplants-orchestrator` host (the box became a shared agent host for several projects); this one had drifted and still carried lichen/project-orchestrator units. Superseded by `../nix/hosts/cc-ci` + `../nix/modules/orchestrator-host.nix`. |
|
||||
| `nix/atproto-likes.nix` | A notplants (not cc-ci) service that lived on the shared box; kept by notplants-nix. |
|
||||
| `terraform/` | OpenTofu for the `cpx22` orchestrator server (Debian 12 → nixos-infect at `nixos-24.11`). The combined host was provisioned by hand instead; the README documents that path. Note its `user-data.sh` would fail on the Debian 13 image (nixos-infect's temp swapfile on a tmpfs `/tmp`) — see the README's `NO_SWAP=true` note. |
|
||||
| `plans/plan-orchestrator-migration.md` | Pi → Incus VM move of the orchestrator (2026-05). |
|
||||
| `plans/plan-orchestrator-hetzner-migration.md` | Incus VM → Hetzner `cpx22` move of the orchestrator (2026-05-31). Has the reboot-resilience design (`cc-ci-loops.service`). |
|
||||
| `plans/plan-migrate-cc-ci-to-hetzner.md`, `plans/plan-cc-ci-hetzner-migration.md`, `plans/plan-cc-ci-hetzner-terraform.md` | The CI server's own move from the `cc-nix-test` Incus VM to Hetzner `cpx32` (`91.98.47.73`, 2026-05-31), and the terraform that provisioned it (lives in the cc-ci repo). |
|
||||
| `plans/plan-repo-consolidation.md` | The earlier repo layout consolidation. |
|
||||
|
||||
The cc-ci server's own history (machine-docs, decisions, the clean-room rebuild that proved
|
||||
"two repos + one age key + one `nixos-rebuild switch`") is in the cc-ci repo under
|
||||
`machine-docs/` and `docs/`.
|
||||
@@ -1,102 +0,0 @@
|
||||
# Plan — one Hetzner host for cc-ci (CI server) + cc-ci-orchestrator
|
||||
|
||||
**Status:** IN PROGRESS (started 2026-09-07). Operator request: move the cc-ci CI server AND the
|
||||
cc-ci orchestrator onto one new Hetzner server (`195.201.88.249`, 8 GB, 150 GB, Debian 13 image),
|
||||
cleanly split off from the shared `notplants-orchestrator` box, which keeps everything else
|
||||
(lichen, project-orchestrator, notplants agents). Nix config ownership: cc-ci's config in
|
||||
`recipe-maintainers/cc-ci`, the orchestrator's in `recipe-maintainers/cc-ci-orchestrator`, and the
|
||||
orchestrator flake **includes** cc-ci's module so one `nixos-rebuild` produces the combined host.
|
||||
Last step (separate, after everything works on the current names): move both to `autonomic.zone`
|
||||
subdomains.
|
||||
|
||||
## Facts (2026-09-07)
|
||||
|
||||
| | old cc-ci server | old orchestrator host (stays, becomes notplants-only) | **new combined host** |
|
||||
|---|---|---|---|
|
||||
| public IP | 91.98.47.73 (fsn1, Hetzner 134485294) | 168.119.126.100 (nbg1, Hetzner 134487234) | **195.201.88.249** |
|
||||
| tailnet | `cc-ci` 100.95.31.88 | `cc-ci-orchestrator-1` 100.84.190.30 | `cc-ci` (new node) |
|
||||
| RAM / disk | 8 GB / 150 GB (83 GB used, 59 GB docker) | 4 GB + 4 GB swap / 75 GB + 250 GB `/mnt/data` | 8 GB / 150 GB, one disk |
|
||||
| built by | `cc-ci` flake `#cc-ci` (nixpkgs 26.05 rev 531670d) | `notplants-nix` flake `#notplants-orchestrator` (26.05 channel), importing `cc-ci-orchestrator`'s `nixosModules.cc-ci` | `cc-ci-orchestrator` flake `#cc-ci` importing `cc-ci`'s `nixosModules.cc-ci-server` |
|
||||
| DNS | `ci.`, `*.ci.`, `ns-acme.commoninternet.net` → 91.98.47.73 (Gandi, direct, no gateway) | `oc.commoninternet.net` → 100.84.190.30 | operator repoints at cutover |
|
||||
|
||||
Data on the old cc-ci server that must move: `/var/lib/cc-ci-reports` (published reports),
|
||||
`/var/lib/cc-ci-runs` (dashboard artifacts, 1.7 G), `/var/lib/ci-warm` (1.4 G), `/var/lib/acme`
|
||||
(LE cert valid to 2026-11-29 + **acme-dns account json** that the `_acme-challenge` CNAME points at),
|
||||
`/var/lib/acme-dns` (the authoritative zone DB), `/var/lib/ci-certs`, `/root/.abra` (app env files),
|
||||
`/etc/cc-ci` (deployed checkout the Sunday sweep runs from), Drone's `drone_ci_commoninternet_net_data`
|
||||
volume (Gitea OAuth grant + repo activation + build history). Warm recipe volumes are caches and get
|
||||
rebuilt on first use / the Sunday sweep. Docker swarm secrets/configs cannot be copied; the reconcile
|
||||
oneshots recreate them from sops.
|
||||
|
||||
Out-of-band secrets the new host needs (never in git): `/var/lib/sops-nix/key.txt` (= the master age
|
||||
key, `/srv/cc-ci/.sops/master-age.txt` here — the new host's SSH host key is not a sops recipient),
|
||||
`/etc/ts-auth-key`, `/srv/cc-ci/.testenv`, `/srv/cc-ci/upgrader.env`, `/srv/cc-ci/.sops/master-age.txt`,
|
||||
`~loops/.ssh/{cc-ci-root,autonomic-bot-gitea,tangled}-ed25519`, `/etc/nginx/oc-*` (self-signed cert +
|
||||
htpasswd for the opencode UI), claude/opencode/codex auth under `~loops`.
|
||||
|
||||
## Design
|
||||
|
||||
**cc-ci repo** (`feat/nixos-module-export`):
|
||||
- `nixosModules.cc-ci-server` = `nix/modules/default.nix`: imports all service modules + the
|
||||
host-generic cc-ci settings that used to sit in the host file (UTC, docker/swarm firewall 80/443,
|
||||
`environment.systemPackages = ccciRuntimeTools`, allowUnfree). No hardware, no networking, no
|
||||
tailscale, no root keys, no stateVersion — the host supplies those.
|
||||
- New options under `cc-ci.*`: `publicIPv4` (acme-dns listen + the `ns-acme` A record),
|
||||
`sopsFile` (absolute path to the decrypted-at-activation `secrets.yaml`, default the submodule
|
||||
path so `#cc-ci` keeps working), `repoPath` (`/etc/cc-ci`, used by nightly-sweep).
|
||||
- `nixosConfigurations.cc-ci` (old host) keeps building unchanged via the same module.
|
||||
|
||||
**cc-ci-orchestrator repo** (`feat/combined-cc-ci-host`):
|
||||
- flake input `cc-ci` (https, public) with `nixpkgs`/`sops-nix` `follows` so one nixpkgs + one sops-nix.
|
||||
- `nixosModules.cc-ci-orchestrator` (the existing `nix/modules/cc-ci.nix`, kept exported as
|
||||
`nixosModules.cc-ci` too so notplants-nix keeps evaluating until it drops the input) — the loops,
|
||||
orchestrator session and the weekly/hourly timers.
|
||||
- `nix/modules/orchestrator-host.nix`: the host contract the module assumes — `loops` user + sudo,
|
||||
nix-ld, claude/opencode/codex installers, `opencode-web`, the tailnet-only nginx `oc.` vhost
|
||||
(on the tailscale IP, port **8443**, because traefik owns 80/443), tool packages, PATH.
|
||||
- `nixosConfigurations.cc-ci` = `nix/hosts/cc-ci/{configuration,hardware,networking}.nix` importing
|
||||
both modules. `/srv` is a plain directory (no `/mnt/data`), 8 GB swapfile, root keys, tailscale
|
||||
`--hostname=cc-ci`, firewall 22 (+ what cc-ci-server opens: 80, 443, 53).
|
||||
- `loops`' ssh config `Host cc-ci` → `127.0.0.1` so every `ssh cc-ci …` in skills/scripts keeps working.
|
||||
- `archive/`: the retired Incus/Hetzner-orchestrator host config, old terraform, historical plans.
|
||||
- `README.md`: provisioning (Hetzner Debian → nixos-infect → NixOS), secrets staging, the one
|
||||
`nixos-rebuild`, data restore, cutover, verification — written so a person or an LLM can redo it.
|
||||
|
||||
**notplants-nix** (`chore/drop-cc-ci`, after cutover): remove the `cc-ci` input, module import, the
|
||||
four cc-ci units' mount gating, `loopsSshConfig`, `opencode-web` + the `oc.` vhost (unless something
|
||||
notplants-side uses it), tailscale hostname → `notplants-orchestrator`.
|
||||
|
||||
## Steps
|
||||
|
||||
1. [ ] nixos-infect the new box (`NIX_CHANNEL=nixos-26.05 PROVIDER=hetzner`); capture
|
||||
`hardware-configuration.nix` + `networking.nix`.
|
||||
2. [ ] cc-ci: module export + options; verify `#cc-ci` still evaluates; PR.
|
||||
3. [ ] cc-ci-orchestrator: input + host + modules + archive/ + README + terraform refresh; verify
|
||||
`#cc-ci` evaluates; PR.
|
||||
4. [ ] Stage secrets + clones on the new host; `nixos-rebuild test` → verify → `switch`.
|
||||
Immediately after: scale the new `ccci-bridge_app` to 0 and mask the two cc-ci timers so the
|
||||
new host does not double-process `!testme` or run a second weekly upgrade before cutover.
|
||||
5. [ ] Copy data (rsync over tailscale): reports, runs, ci-warm, acme, acme-dns, ci-certs,
|
||||
/root/.abra, /etc/cc-ci; Drone volume with Drone scaled to 0 during the copy.
|
||||
6. [ ] Pre-cutover verification on the new IP (`curl --resolve`, port 53, dashboard, reports,
|
||||
drone, one direct `cc-ci-run` on custom-html-tiny).
|
||||
7. [ ] Operator: Gandi A records `ci`, `*.ci`, `ns-acme` → 195.201.88.249. Then: old bridge +
|
||||
drone + timers off, new bridge up, one real `!testme` end-to-end, a `!testme`-driven report page.
|
||||
8. [ ] Move the orchestrator: stop cc-ci units here, final rsync of `/srv/cc-ci-orch` + agent
|
||||
state, enable on the new host, operator reconnects there; notplants-nix PR removing cc-ci.
|
||||
9. [ ] Old cc-ci server: cold standby ~1 week, then operator deletes it and the stale tailnet node.
|
||||
10. [ ] Domain move to `autonomic.zone` — separate plan, after 1–9 are proven.
|
||||
|
||||
## Log
|
||||
|
||||
- 2026-09-07 19:40 UTC — recon done, plan written, ssh to the new box verified as root with
|
||||
`notplants-orchestrator-ed25519`.
|
||||
- 2026-09-07 20:05 UTC — nixos-infect started on 195.201.88.249 (rev 40f62a6, nixos-26.05,
|
||||
PROVIDER=hetznercloud). Two false starts: the Debian 13 image has /tmp on tmpfs, so
|
||||
nixos-infect's temp swapfile fails `swapon: Invalid argument`; fixed with `NO_SWAP=true`.
|
||||
Build ran, box rebooted ~20:11 UTC and has not answered ping/ssh since (>25 min) — needs the
|
||||
Hetzner console (no API token for that project on this host).
|
||||
- 2026-09-07 20:40 UTC — cc-ci branch `feat/nixos-module-export` (9b99f81) pushed: the standalone
|
||||
`#cc-ci` drv is byte-identical before/after. Orchestrator branch `feat/combined-cc-ci-host`:
|
||||
`#cc-ci` evaluates (gcnwq4fy…-nixos-system-cc-ci-26.05.20260803.531670d.drv) with PROVISIONAL
|
||||
hardware/networking copied from the old CI server — to be replaced by the infect output.
|
||||
@@ -103,6 +103,18 @@
|
||||
2026-08-15 (upstream main still pins 10.11.22 = EXPIRED ESR → the 10→11 ESR move PR #2 carries
|
||||
remains required; Mattermost docs: ESR→ESR is "fully supported and tested"). postgres 15-alpine
|
||||
still HELD (DB-major out of scope, operator dump/pg_upgrade).
|
||||
- **2026-09-04 re-check** (endoflife.date/api/mattermost.json 2026-09-04; Mattermost release-policy
|
||||
docs `https://docs.mattermost.com/product-overview/release-policy.html`; `mattermost-server-releases.html`;
|
||||
GitHub releases `v11.7.10`): **11.7 ESR is STILL the current supported ESR/LTS line** — "v11.7 &
|
||||
Desktop App v6.2 Extended Support: 2026-05-15 → 2027-05-15" (the chart on the release-policy page;
|
||||
ESR cadence = every 9 months, supported 12 months). Latest 11.7.x patch **11.7.10** (2026-08-26,
|
||||
"Mattermost Platform Extended Support Release 11.7.10 contains various bug fixes") — NOT a
|
||||
prerelease; target confirmed. 11.8/11.9/11.10 remain Feature/innovation releases (EOL 2026-09-15 /
|
||||
10-15 / 11-15, `lts:false`), NOT ESR — do NOT target; wait for the NEXT official ESR (expected
|
||||
~Feb 2027 on the 9-month cadence). No newer 11.7.x ESR patch exists as of this week, so PR #2's
|
||||
head (`59e8c2c`, app image `11.7.10`) is still the correct target → this run RE-VERIFIES PR #2
|
||||
(no new app bump). 11.11.0-rc1/rc2 seen on GitHub but innovation + pre-release — not a target.
|
||||
postgres 15-alpine still HELD (DB-major out of scope, operator dump/pg_upgrade).
|
||||
|
||||
## NVD CPE fallback
|
||||
This project publishes nothing machine-readable we can reach — no GitHub advisory feed,
|
||||
|
||||
@@ -137,3 +137,24 @@
|
||||
2.37.5 withdrawn). 2.36.9 holds the Stable/Latest badge; 2.37.x remains Pre-release on GitHub
|
||||
(consistent precedent). Re-verified 2.37.3→2.37.6 (pure core bugfixes), no breaking changes beyond
|
||||
the already-flagged 2.37.0 API behavior pair. Rolling upgrade safe. Recommended release: `-y`.
|
||||
- 2.37.7 (2026-09-02, **stable**): empty changelog (auto-generated release, no bugfixes listed).
|
||||
- **2.37.8**: plain Docker Hub tag exists (2026-09-03) but **no GitHub release page** (like 2.36.1/
|
||||
2.37.2/2.37.5 for the release notes; the tag itself is real). Treat as no separate changelog.
|
||||
- 2.37.9 (2026-09-03, **now the stable/Latest badge** — `stable` tag points here): 1 core bugfix
|
||||
(restore mutating array methods on $json data in expressions).
|
||||
- **2.38.0**: plain Docker Hub tag exists (2026-09-01) but **no GitHub release page**; the 2.38.1
|
||||
release body compares `2.37.0...2.38.1` (it absorbs the 2.38.0 changes).
|
||||
- 2.38.1 (2026-09-01, Pre-release): the large 2.38.x line changelog (compare basis 2.37.0). Mostly
|
||||
bugfixes + features: expression-engine fixes (copy-on-write writes on VM lazy proxies, one shared
|
||||
time budget across nested expressions, validate engine timeout/memory settings), distroless runners
|
||||
image fixes (glibc/libatomic — only relevant if using n8n's community/distroless runner image, not
|
||||
the recipe), model-provider additions (Moonshot/MiniMax/Qwen Cloud), editor improvements. **No
|
||||
breaking compose/config/migration changes, no `N8N_*` env renames.**
|
||||
- 2.38.2 (2026-09-02, Pre-release): empty changelog (auto-generated, no bugfixes listed).
|
||||
- 2.38.3 (2026-09-03, Pre-release; **newest 2.x tag abra lists**): 1 core bugfix (ensure running job
|
||||
cleanup when a workflow run rejects).
|
||||
- 2026-09-04 run: PR #7 extended **2.34.4 → 2.38.3** (newest tag abra lists = 2.38.3/2.38.2/2.38.1/
|
||||
2.38.0/2.37.9/2.37.8/2.37.7/…). 2.37.9 holds the Stable/Latest badge; 2.38.x remains Pre-release on
|
||||
GitHub (consistent newest-tag precedent). 2.37.7/2.37.9 patched the stable line; 2.38.x carries the
|
||||
expression-engine / agent-runtime fixes. No breaking changes beyond the already-flagged 2.37.0 API
|
||||
behavior pair. Rolling upgrade safe. Recommended release: `-y`.
|
||||
|
||||
Generated
-24
@@ -1,28 +1,5 @@
|
||||
{
|
||||
"nodes": {
|
||||
"cc-ci": {
|
||||
"inputs": {
|
||||
"nixpkgs": [
|
||||
"nixpkgs"
|
||||
],
|
||||
"sops-nix": [
|
||||
"sops-nix"
|
||||
]
|
||||
},
|
||||
"locked": {
|
||||
"lastModified": 1788812004,
|
||||
"narHash": "sha256-Vc7RSeqFHCwlVRhnEEjavuJoyIID+RQdJSygHNA8s8Y=",
|
||||
"ref": "refs/heads/main",
|
||||
"rev": "f6dbfa368995f4d45de09f4052631fd433c87d5b",
|
||||
"revCount": 1533,
|
||||
"type": "git",
|
||||
"url": "https://git.autonomic.zone/recipe-maintainers/cc-ci.git"
|
||||
},
|
||||
"original": {
|
||||
"type": "git",
|
||||
"url": "https://git.autonomic.zone/recipe-maintainers/cc-ci.git"
|
||||
}
|
||||
},
|
||||
"nixpkgs": {
|
||||
"locked": {
|
||||
"lastModified": 1785734586,
|
||||
@@ -41,7 +18,6 @@
|
||||
},
|
||||
"root": {
|
||||
"inputs": {
|
||||
"cc-ci": "cc-ci",
|
||||
"nixpkgs": "nixpkgs",
|
||||
"sops-nix": "sops-nix"
|
||||
}
|
||||
|
||||
@@ -1,52 +1,37 @@
|
||||
{
|
||||
description = "cc-ci-orchestrator — the cc-ci orchestrator (loops, steering session, weekly upgrader) and the NixOS host it shares with the cc-ci CI server";
|
||||
description = "cc-ci-orchestrator — NixOS host for the cc-ci loops runtime (Builder/Adversary/Watchdog)";
|
||||
|
||||
inputs = {
|
||||
# Stable release channel (operator 2026-08-01). `nix flake update` moves it; the cc-ci input
|
||||
# below FOLLOWS it, so one nixpkgs builds the whole combined host and CVEs get patched once.
|
||||
# Follow the current stable release channel (operator 2026-08-01), was a hard rev pin at
|
||||
# nixpkgs 24.11 (50ab7937, 2025-06-30) kept "the same as the cc-ci server". This host runs
|
||||
# agents/tmux/nginx/docker, not recipe CI, so it does not need to match that server — and a
|
||||
# frozen rev only accrues unpatched CVEs. `nix flake update` now actually moves.
|
||||
nixpkgs.url = "github:NixOS/nixpkgs/nixos-26.05";
|
||||
|
||||
# sops-nix follows nixpkgs below, so it no longer needs its own matching pin.
|
||||
sops-nix.url = "github:Mic92/sops-nix";
|
||||
sops-nix.inputs.nixpkgs.follows = "nixpkgs";
|
||||
|
||||
# The cc-ci CI server, as a NixOS module (`nixosModules.cc-ci-server`). HTTPS, anonymous read:
|
||||
# nix evaluates every input for every output, so the input must be fetchable without
|
||||
# credentials. The private secrets submodule is deliberately NOT fetched through this input —
|
||||
# the host reads the deployed --recursive checkout's secrets.yaml at activation instead
|
||||
# (`cc-ci.sopsFile`). Both `follows` are REQUIRED: without them cc-ci's own nixpkgs/sops-nix
|
||||
# pins would produce a second sops-nix module tree and a second nixpkgs in one system.
|
||||
cc-ci.url = "git+https://git.autonomic.zone/recipe-maintainers/cc-ci.git";
|
||||
cc-ci.inputs.nixpkgs.follows = "nixpkgs";
|
||||
cc-ci.inputs.sops-nix.follows = "sops-nix";
|
||||
};
|
||||
|
||||
outputs = { self, nixpkgs, sops-nix, cc-ci, ... }:
|
||||
outputs = { nixpkgs, sops-nix, ... }:
|
||||
let
|
||||
system = "x86_64-linux";
|
||||
in
|
||||
{
|
||||
nixosModules = {
|
||||
# The orchestrator itself: loops supervisor, steering session, weekly/hourly timers.
|
||||
cc-ci-orchestrator = ./nix/modules/cc-ci.nix;
|
||||
# The host contract those units assume: loops user, claude/opencode CLIs, opencode web
|
||||
# server + tailnet UI, nix-ld, tool set, `ssh cc-ci` config.
|
||||
orchestrator-host = ./nix/modules/orchestrator-host.nix;
|
||||
# Old name of cc-ci-orchestrator, kept while notplants-nix still imports it (2026-09).
|
||||
cc-ci = ./nix/modules/cc-ci.nix;
|
||||
};
|
||||
# The cc-ci part of a host, on its own, so a host that runs cc-ci can import just this and
|
||||
# keep its own (unrelated) configuration separate. Split out 2026-08-20; consumed by
|
||||
# notplants-nix's `notplants-orchestrator` host.
|
||||
nixosModules.cc-ci = ./nix/modules/cc-ci.nix;
|
||||
|
||||
nixosConfigurations = {
|
||||
# THE live host: cc-ci CI server + cc-ci orchestrator on one Hetzner cpx32-class box
|
||||
# (195.201.88.249, since 2026-09). README.md is the deploy guide.
|
||||
cc-ci = nixpkgs.lib.nixosSystem {
|
||||
inherit system;
|
||||
modules = [
|
||||
cc-ci.nixosModules.cc-ci-server
|
||||
self.nixosModules.cc-ci-orchestrator
|
||||
self.nixosModules.orchestrator-host
|
||||
./nix/hosts/cc-ci/configuration.nix
|
||||
];
|
||||
};
|
||||
# Hetzner cpx11 host (nixos-infect generated hardware.nix + orchestrator config).
|
||||
# Provision with terraform/ then run Stage 2 per terraform/README.md.
|
||||
nixosConfigurations.cc-ci-orchestrator-hetzner = nixpkgs.lib.nixosSystem {
|
||||
inherit system;
|
||||
modules = [
|
||||
sops-nix.nixosModules.sops
|
||||
./nix/hosts/cc-ci-orchestrator-hetzner/hardware.nix
|
||||
./nix/hosts/cc-ci-orchestrator-hetzner/configuration.nix
|
||||
];
|
||||
};
|
||||
};
|
||||
}
|
||||
|
||||
@@ -10,10 +10,9 @@ metadata:
|
||||
The cc-ci orchestrator (loops + watchdog + this session) runs on a **Hetzner cpx22** as of
|
||||
2026-05-31, replacing the Incus VM (100.116.55.106).
|
||||
|
||||
- Since 2026-09-07: ONE Hetzner host for CI server + orchestrator, public **195.201.88.249**,
|
||||
tailnet **cc-ci**, flake host **`.#cc-ci`** (this repo). Before: orchestrator on Hetzner
|
||||
134487234 (168.119.126.100 / 100.84.190.30, `cc-ci-orchestrator-hetzner`), shared with notplants.
|
||||
- Rebuild: `sudo nixos-rebuild switch --flake .#cc-ci` from `/srv/cc-ci-orch`
|
||||
- Hetzner server **134487234**, public **168.119.126.100**, tailnet **cc-ci-orchestrator-1** @
|
||||
**100.84.190.30**. Flake host **cc-ci-orchestrator-hetzner**.
|
||||
- Rebuild: `sudo nixos-rebuild switch --flake .#cc-ci-orchestrator-hetzner` from `/srv/cc-ci-orch`
|
||||
(`/srv/cc-ci` is a symlink to it). The Bash tool runs as user **loops** (uid 1000, passwordless
|
||||
sudo) — plain `nixos-rebuild switch` fails on the profile symlink; use `sudo`.
|
||||
- Reboot-resilience: `cc-ci-loops.service` is **enabled** (wantedBy multi-user.target); ExecStartPre
|
||||
@@ -24,4 +23,4 @@ The cc-ci orchestrator (loops + watchdog + this session) runs on a **Hetzner cpx
|
||||
identity unknown". Set per-repo to match prior commits: `autonomic-bot
|
||||
<autonomic-bot@git.autonomic.zone>`.
|
||||
|
||||
Full record: `archive/plans/plan-orchestrator-hetzner-migration.md`.
|
||||
Full record: `cc-ci-plan/plan-orchestrator-hetzner-migration.md`.
|
||||
|
||||
@@ -1,68 +0,0 @@
|
||||
# cc-ci — ONE Hetzner Cloud host running both the cc-ci CI server and the cc-ci orchestrator.
|
||||
#
|
||||
# This file is only what is physical or identity about the machine: hardware, networking, the
|
||||
# tailscale node, root SSH keys, swap, stateVersion. Everything functional comes from modules:
|
||||
# cc-ci.nixosModules.cc-ci-server recipe-maintainers/cc-ci — swarm, traefik, drone,
|
||||
# runner, bridge, dashboard, reports, acme-dns, harness
|
||||
# self.nixosModules.cc-ci-orchestrator nix/modules/cc-ci.nix — loops, orchestrator, timers
|
||||
# self.nixosModules.orchestrator-host nix/modules/orchestrator-host.nix — loops user, CLIs
|
||||
# See README.md for provisioning (Hetzner Debian → nixos-infect → this flake) and staging.
|
||||
{ lib, pkgs, ... }:
|
||||
{
|
||||
imports = [
|
||||
./hardware.nix
|
||||
./networking.nix
|
||||
];
|
||||
|
||||
networking.hostName = "cc-ci";
|
||||
|
||||
# ---- cc-ci server identity --------------------------------------------------------------
|
||||
# Public address: acme-dns binds to it and publishes it as the `ns-acme` glue record; the
|
||||
# Gandi A records for ci / *.ci / ns-acme .commoninternet.net point here.
|
||||
cc-ci.publicIPv4 = "195.201.88.249";
|
||||
# cc-ci is a plain flake input here (no private submodule), so the sops file is the one in
|
||||
# the deployed --recursive checkout the weekly sweep runs from (README "Stage the workspace").
|
||||
cc-ci.sopsFile = "/etc/cc-ci/secrets/secrets.yaml";
|
||||
|
||||
# ---- orchestrator identity --------------------------------------------------------------
|
||||
# The CI server is this very host, so `ssh cc-ci` goes to loopback (the module default).
|
||||
cc-ci-orchestrator.ciSshHost = "127.0.0.1";
|
||||
|
||||
# ---- tailscale — auth key staged out of band at /etc/ts-auth-key -----------------------
|
||||
services.tailscale = {
|
||||
enable = true;
|
||||
authKeyFile = "/etc/ts-auth-key";
|
||||
extraUpFlags = [ "--hostname=cc-ci" ];
|
||||
};
|
||||
|
||||
# ---- ssh ----------------------------------------------------------------------------------
|
||||
services.openssh = {
|
||||
enable = true;
|
||||
settings.PermitRootLogin = "yes";
|
||||
};
|
||||
# Root keys: PUBLIC keys, tracked deliberately in ./ssh-keys (one per line, blank lines ok).
|
||||
users.users.root.openssh.authorizedKeys.keys =
|
||||
builtins.filter (s: s != "") (lib.splitString "\n" (builtins.readFile ./ssh-keys));
|
||||
# The loops user can also be reached directly (same keys) — handy for rsync of its workspace.
|
||||
users.users.loops.openssh.authorizedKeys.keys =
|
||||
builtins.filter (s: s != "") (lib.splitString "\n" (builtins.readFile ./ssh-keys));
|
||||
|
||||
# ---- firewall -------------------------------------------------------------------------------
|
||||
# 80/443 (traefik) and 53 (acme-dns) are opened by the cc-ci-server module. The tailscale
|
||||
# interface is trusted, which is what makes the opencode UI on 8443 tailnet-only.
|
||||
networking.firewall = {
|
||||
enable = true;
|
||||
trustedInterfaces = [ "tailscale0" ];
|
||||
allowedTCPPorts = [ 22 ];
|
||||
};
|
||||
networking.nameservers = [ "1.1.1.1" "8.8.8.8" ];
|
||||
|
||||
# ---- memory: 8 GB RAM shared by the swarm (recipe deploys) and 3–6 agent sessions ---------
|
||||
swapDevices = [ { device = "/swapfile"; size = 8192; } ];
|
||||
|
||||
# ssh client for root (the orchestrator's `ssh cc-ci` goes through the loops user's own config).
|
||||
environment.systemPackages = [ pkgs.openssh ];
|
||||
|
||||
# Fresh NixOS 26.05 install (nixos-infect, 2026-09-07). Never change this on an existing host.
|
||||
system.stateVersion = "26.05";
|
||||
}
|
||||
@@ -1,19 +0,0 @@
|
||||
# Generated by nixos-infect on this machine (2026-09-07), captured verbatim per README §3.
|
||||
# The ESP UUID is specific to THIS server; a new server gets a new file.
|
||||
{ modulesPath, ... }:
|
||||
{
|
||||
imports = [ (modulesPath + "/profiles/qemu-guest.nix") ];
|
||||
boot.loader = {
|
||||
efi.efiSysMountPoint = "/boot/efi";
|
||||
grub = {
|
||||
efiSupport = true;
|
||||
efiInstallAsRemovable = true;
|
||||
device = "nodev";
|
||||
};
|
||||
};
|
||||
fileSystems."/boot/efi" = { device = "/dev/disk/by-uuid/E079-7D41"; fsType = "vfat"; };
|
||||
boot.initrd.availableKernelModules = [ "ata_piix" "uhci_hcd" "xen_blkfront" "vmw_pvscsi" ];
|
||||
boot.initrd.kernelModules = [ "nvme" ];
|
||||
fileSystems."/" = { device = "/dev/sda1"; fsType = "ext4"; };
|
||||
|
||||
}
|
||||
@@ -1,38 +0,0 @@
|
||||
# Generated by nixos-infect on this machine (2026-09-07), captured per README §3, with ONE edit:
|
||||
# `defaultGateway` as an attrset WITH `interface = "eth0"`. The generated bare-string form leaves
|
||||
# NixOS ≥25.05 without a default route (the host boots and is unreachable) — see README §2.
|
||||
{ lib, ... }: {
|
||||
# This file was populated at runtime with the networking
|
||||
# details gathered from the active system.
|
||||
networking = {
|
||||
nameservers = [ "2a01:4ff:ff00::add:2"
|
||||
"2a01:4ff:ff00::add:1"
|
||||
"185.12.64.2"
|
||||
];
|
||||
defaultGateway = { address = "172.31.1.1"; interface = "eth0"; };
|
||||
defaultGateway6 = {
|
||||
address = "fe80::1";
|
||||
interface = "eth0";
|
||||
};
|
||||
dhcpcd.enable = false;
|
||||
usePredictableInterfaceNames = lib.mkForce false;
|
||||
interfaces = {
|
||||
eth0 = {
|
||||
ipv4.addresses = [
|
||||
{ address="195.201.88.249"; prefixLength=32; }
|
||||
];
|
||||
ipv6.addresses = [
|
||||
{ address="2a01:4f8:1c1c:a9b::1"; prefixLength=64; }
|
||||
{ address="fe80::2ff8:e3ea:bbb8:aa39"; prefixLength=64; }
|
||||
];
|
||||
ipv4.routes = [ { address = "172.31.1.1"; prefixLength = 32; } ];
|
||||
ipv6.routes = [ { address = "fe80::1"; prefixLength = 128; } ];
|
||||
};
|
||||
|
||||
};
|
||||
};
|
||||
services.udev.extraRules = ''
|
||||
ATTR{address}=="92:00:09:d5:ec:0d", NAME="eth0"
|
||||
|
||||
'';
|
||||
}
|
||||
@@ -1,10 +0,0 @@
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGZGp/DQTFuD1GvsyTzCVBUTmoWqcb5T+Z7zZo5nYLXO
|
||||
ssh-rsa AAAAB3NzaC1yc2EAAAADAQABAAABgQDhgo41nt8/L+Cr0PKd8jQK45mw/A+h041j6LQ8JWZisEVaQOzr6s9rxPL8VT5ML4P3/4bMblzdDiXWlJxymcb+yk5S5TnVrMavzHEDhWHwEvTRMe6xNTmsU6cmmhRw7PJqqQ+0GTlQalu3I4jkC0kTF7kuPwduUOgUuSpJqxvDTwYiXoyVnOQHAIygh+BmQvYUz0PBfQgIhgcbYmGZ++T0DnMzdGFzW2UB/iy5mymnpmbaZCgLy0w8AoDE+0YLtUc4gwTXc183nvqO1i7LQr+3jBYkv5ZthCCc52vXFHDSw9xZ5ohsOrBvoi5foRbqinmU5/t0aTK7SSrat7xXm/odIOyS+S7PJyeEcsXN6d5zdxbabAy5vLfodEaKGZd4rqQeDCxOTPAS/BlrBV/EV714n4E+fSOAllAuMBO4IibJM/gLJrh2Dql3co50QW9HEDeSC7iqp2lxRBDxvUs3rIEzy7o4HSN8chqBUK1bbBY6B17fuNHIpBAw4akRVVvPnVM= trav@trav480sweet
|
||||
ssh-rsa AAAAB3NzaC1yc2EAAAADAQABAAABgQC6jrKj7iZUNRLBTZG0vZM1D/BXtARhhB4+GrvpyuqmPb9iw2ifT9YqRUwgyGrOW9U6nIAR9yFnfp9+FkyhEKWByqEBbe/zYKlGLRGjfsIdDdW29QQ3hvmqNyboCkXLxZGat93poYhnoomqicmGD/xST4s0OUhcK9E494lUmenlD9dcMZW1aKpJ+9O4Dq6A7nk2z1e4KFcZdrZDI2Hgg+gfEdsKZQqd/R3Mls/eVKpzhfv3Y8BiNoHssUChVf8IGESqTOBOR7Dk7FsU5Z2ZcnQ1coxY7VlBn4fPjTWmz/Ac0jLqgcpCLpNyQzFPDVMYZKYrPVoqBeKVhN5YnfwR5OVP8YsakT/obLwC43sx/esXfjhVGcsRoGpiLOfazzNw/eC8s6FlS8cesOubEM37a7F25z4UEG3d487oM7EjQ39gBCCj/KRgUimCKMWsm6yIas4OSctBWEAo/NhZp0gwulSRxleW6eJCNNwzOmWjdzYIVWoVP0EIeM95Tq8PVUN7gpc= aadil@t480
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIMyHSi12R0+HCVBz7+d9fyOBnoJi8Nsj5D7vQ9UQO8a5
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIJVlfoLBPseQ9fA9534KmRg2KWcksKZGzAJIpHJ2JpsI mfowler.email@protonmail.com
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIAQFuqUB2qNZSDNjDsjjhVA/WnnQNVAMmsUscW6OgMDN
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIHOcLo0YBa0UYi7i/l8K/Y/7cF2OclmDqSTlAsHM0dOS notplants-orchestrator
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIMniNzAzuI527bfk/EipqFILFayUCwYXDoZ3R7+QgYq6
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIOk8NaeBdPbS2gfUvbny8h0AkZlVjGYHzx4QPXSJ38gd claude@claude-vm
|
||||
ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIAcyTGb/wVgdhg5oBCZZvBaR1RuUQRY/3WHnOQpNDCsp claude-cc-ci-sandbox@20260526
|
||||
+16
-16
@@ -1,16 +1,15 @@
|
||||
# cc-ci.nix — the cc-ci ORCHESTRATOR: the Builder/Adversary loops supervisor, the operator's
|
||||
# steering session, and the weekly-upgrade + hourly-supervisor timers. Nothing else.
|
||||
# cc-ci.nix — everything on this host that exists FOR cc-ci, and nothing else.
|
||||
#
|
||||
# Exported from this repo's flake as `nixosModules.cc-ci-orchestrator` (and, for the host that
|
||||
# used to import it under the old name, `nixosModules.cc-ci`). Split out of the shared agent
|
||||
# host config on 2026-08-20; since 2026-09 it runs on the same Hetzner host as the CI server
|
||||
# itself (`#cc-ci` in flake.nix), next to recipe-maintainers/cc-ci's `nixosModules.cc-ci-server`.
|
||||
# Split out of the orchestrator host config on 2026-08-20. The host it runs on is a general
|
||||
# agent/orchestration box that also serves several unrelated projects; this module is the cc-ci
|
||||
# part of it, so that the two can evolve (and be reviewed) independently. It is exported from this
|
||||
# repo's flake as `nixosModules.cc-ci` and imported by whichever host runs cc-ci.
|
||||
#
|
||||
# All of it assumes the cc-ci workspaces exist on the host:
|
||||
# /srv/cc-ci the loops workspace (+ .cc-ci-logs, upgrader.env) — a symlink to
|
||||
# /srv/cc-ci-orch this repo (the orchestrator's own working dir), with cc-ci/ checked out
|
||||
# /srv/cc-ci the loops workspace (+ .cc-ci-logs, upgrader.env)
|
||||
# /srv/cc-ci-orch this repo (the orchestrator's own working dir)
|
||||
# and that a `loops` user, tmux, python3 and the standalone claude/opencode CLIs are present —
|
||||
# those are host concerns, provided by nix/modules/orchestrator-host.nix, not by this module.
|
||||
# those are host concerns, provided by the host config, not by this module.
|
||||
{ config, pkgs, lib, ... }:
|
||||
{
|
||||
# cc-ci-loops supervisor — workspace staged 2026-05-31, so ENABLED for reboot-resilience.
|
||||
@@ -49,14 +48,15 @@
|
||||
# cc-ci-orchestrator supervisor — the operator's steering session. Same shape as
|
||||
# lichen-orchestrator / project-orchestrator above: this unit only LAUNCHES the orchestrator's
|
||||
# tmux session via the agent-orchestrator harness (cc-ci-plan/agents.py); it does not own the
|
||||
# session or the tmux server. The orchestrator agent is declared in cc-ci-plan/agents.toml
|
||||
# (backend/model chosen there — Claude Code under Remote Control since 2026-09-07; before that
|
||||
# opencode/glm-5.2 attached to the shared opencode web server, opencode-web.service in
|
||||
# orchestrator-host.nix, which the upgrader still uses). The harness watchdog (started by
|
||||
# `agents.py up`) keeps it alive: heal-only (no stall reboots — a persistent supervisor must not
|
||||
# be killed just for idling). Added 2026-08-03 for reboot-resilience.
|
||||
# session or the tmux server. The orchestrator agent is declared in cc-ci-plan/agents.toml on
|
||||
# the OPencode backend (backend = "opencode", model = "opencode/glm-5.2"), so on boot it
|
||||
# attaches to the shared opencode web server (opencode-web.service below) and is reachable for
|
||||
# Remote Control at https://oc.commoninternet.net under the /srv/cc-ci-orch project. The harness
|
||||
# watchdog (started by `agents.py up`) keeps it alive: heal-only (no stall reboots — a persistent
|
||||
# supervisor must not be killed just for idling). Added 2026-08-03 to give the cc-ci orchestrator
|
||||
# the same reboot-resilience the other two orchestrators already have.
|
||||
systemd.services.cc-ci-orchestrator = {
|
||||
description = "cc-ci orchestrator (operator steering session) — agents.py up orchestrator";
|
||||
description = "cc-ci orchestrator (operator steering session) — agents.py up orchestrator, opencode backend";
|
||||
wantedBy = [ "multi-user.target" ];
|
||||
after = [ "network-online.target" "tailscaled.service" "opencode-web.service" ];
|
||||
wants = [ "network-online.target" ];
|
||||
|
||||
@@ -1,203 +0,0 @@
|
||||
# orchestrator-host.nix — the host contract that nix/modules/cc-ci.nix (the orchestrator's
|
||||
# loops/timers) silently assumes, made explicit and reusable: the `loops` user the agents run as,
|
||||
# the standalone claude/opencode CLIs, the shared opencode web server and its tailnet-only UI,
|
||||
# nix-ld so foreign binaries run on NixOS, and the tool set agents reach for.
|
||||
#
|
||||
# Exported from flake.nix as `nixosModules.orchestrator-host`. A host imports this together with
|
||||
# `nixosModules.cc-ci-orchestrator`; the combined CI-server + orchestrator host (`#cc-ci`) also
|
||||
# imports recipe-maintainers/cc-ci's `nixosModules.cc-ci-server`.
|
||||
#
|
||||
# History: until 2026-09 this lived (twice, drifting) in nix/hosts/cc-ci-orchestrator-hetzner/
|
||||
# configuration.nix here and in notplants-nix's hosts/notplants-orchestrator/configuration.nix,
|
||||
# the shared agent box that also ran lichen + project-orchestrator. The cc-ci half moved to its
|
||||
# own host; this file is that half.
|
||||
{ config, lib, pkgs, ... }:
|
||||
let
|
||||
cfg = config.cc-ci-orchestrator;
|
||||
in
|
||||
{
|
||||
options.cc-ci-orchestrator = {
|
||||
ciSshHost = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "127.0.0.1";
|
||||
example = "100.95.31.88";
|
||||
description = ''
|
||||
Where `ssh cc-ci` (used by every skill and script that drives the CI server) connects to,
|
||||
as root with ~loops/.ssh/cc-ci-root-ed25519. On the combined host the CI server IS this
|
||||
machine, so the default is loopback; a standalone orchestrator points it at the CI
|
||||
server's tailnet address.
|
||||
'';
|
||||
};
|
||||
|
||||
opencodeUiPort = lib.mkOption {
|
||||
type = lib.types.port;
|
||||
default = 8443;
|
||||
description = ''
|
||||
TLS port of the nginx front door for the opencode web UI. Not 443: on the combined host
|
||||
Traefik (docker swarm) owns 80/443. The port is not opened in the firewall, so it is
|
||||
reachable only over the trusted tailscale interface.
|
||||
'';
|
||||
};
|
||||
|
||||
opencodeUiHost = lib.mkOption {
|
||||
type = lib.types.str;
|
||||
default = "oc.commoninternet.net";
|
||||
description = "nginx server_name for the opencode web UI (self-signed, basic auth).";
|
||||
};
|
||||
};
|
||||
|
||||
config = {
|
||||
# ---- the loops user -------------------------------------------------------------------
|
||||
# claude sessions run as non-root (--dangerously-skip-permissions is refused for root).
|
||||
users.users.loops = {
|
||||
isNormalUser = true;
|
||||
uid = 1000; # fixed: workspace files are rsynced between hosts by uid
|
||||
home = "/home/loops";
|
||||
shell = pkgs.bash;
|
||||
extraGroups = [ "wheel" "docker" ];
|
||||
};
|
||||
security.sudo.wheelNeedsPassword = false;
|
||||
security.sudo.extraRules = [{
|
||||
users = [ "loops" ];
|
||||
commands = [{ command = "ALL"; options = [ "NOPASSWD" ]; }];
|
||||
}];
|
||||
|
||||
# /home/loops/.local/bin holds the standalone claude + opencode binaries; it must be first on
|
||||
# every PATH (interactive shells, tmux, the systemd units in cc-ci.nix prepend it too).
|
||||
environment.variables.PATH = lib.mkForce
|
||||
"/home/loops/.local/bin:/run/current-system/sw/bin:/run/wrappers/bin:/usr/bin:/bin";
|
||||
|
||||
# ---- nix-ld: the standalone Claude Code / opencode CLIs are foreign dynamic ELF binaries ---
|
||||
programs.nix-ld.enable = true;
|
||||
programs.nix-ld.libraries = with pkgs; [ stdenv.cc.cc.lib zlib openssl curl glibc ];
|
||||
|
||||
# ---- the toolbox every agent on this box gets ----------------------------------------
|
||||
# Bar for adding something: an agent doing ordinary work would otherwise waste a turn
|
||||
# discovering it is absent.
|
||||
environment.systemPackages = with pkgs; [
|
||||
git tmux python3 jq curl cacert
|
||||
gnused gawk coreutils gnugrep findutils util-linux nettools openssh
|
||||
age sops ssh-to-age
|
||||
wget gnutar gzip unzip zip xz
|
||||
ripgrep fd tree file less which
|
||||
procps psmisc htop lsof strace ncdu
|
||||
dnsutils socat netcat-gnu iproute2 iputils
|
||||
openssl gnumake gcc pkg-config
|
||||
yq-go diffutils patch rsync bubblewrap
|
||||
];
|
||||
|
||||
# ---- ssh config for the loops user: `ssh cc-ci` = the CI server (root) -----------------
|
||||
# Written only if absent so a manual customisation survives rebuilds.
|
||||
system.activationScripts.loopsSshConfig = ''
|
||||
mkdir -p /home/loops/.ssh && chown loops:users /home/loops/.ssh && chmod 700 /home/loops/.ssh
|
||||
if [ ! -f /home/loops/.ssh/config ]; then
|
||||
cat > /home/loops/.ssh/config <<'SSHCFG'
|
||||
Host cc-ci
|
||||
HostName ${cfg.ciSshHost}
|
||||
User root
|
||||
IdentityFile /home/loops/.ssh/cc-ci-root-ed25519
|
||||
IdentitiesOnly yes
|
||||
StrictHostKeyChecking accept-new
|
||||
ServerAliveInterval 30
|
||||
|
||||
Host git.autonomic.zone
|
||||
HostName git.autonomic.zone
|
||||
Port 2222
|
||||
User git
|
||||
IdentityFile /home/loops/.ssh/autonomic-bot-gitea-ed25519
|
||||
IdentitiesOnly yes
|
||||
|
||||
Host tangled.org
|
||||
IdentityFile /home/loops/.ssh/tangled-ed25519
|
||||
IdentitiesOnly yes
|
||||
SSHCFG
|
||||
chmod 600 /home/loops/.ssh/config
|
||||
chown loops:users /home/loops/.ssh/config
|
||||
fi
|
||||
'';
|
||||
|
||||
# ---- standalone CLIs (idempotent installers; re-run on every activation, no-op if present) --
|
||||
systemd.services.claude-install = {
|
||||
description = "Install Claude Code CLI for loops user (idempotent)";
|
||||
wantedBy = [ "multi-user.target" ];
|
||||
after = [ "network-online.target" ];
|
||||
wants = [ "network-online.target" ];
|
||||
serviceConfig = { Type = "oneshot"; RemainAfterExit = true; User = "loops"; Group = "users"; };
|
||||
environment = { HOME = "/home/loops"; };
|
||||
path = [ pkgs.curl pkgs.bash pkgs.coreutils pkgs.gnutar pkgs.gzip ];
|
||||
script = ''
|
||||
if [ ! -x "$HOME/.local/bin/claude" ]; then
|
||||
echo "installing Claude Code CLI for loops user..."
|
||||
curl -fsSL https://claude.ai/install.sh | bash || echo "install failed — retry on next activation"
|
||||
fi
|
||||
'';
|
||||
};
|
||||
|
||||
systemd.services.opencode-install = {
|
||||
description = "Install opencode CLI for loops user (idempotent)";
|
||||
wantedBy = [ "multi-user.target" ];
|
||||
after = [ "network-online.target" ];
|
||||
wants = [ "network-online.target" ];
|
||||
serviceConfig = { Type = "oneshot"; RemainAfterExit = true; User = "loops"; Group = "users"; };
|
||||
environment = { HOME = "/home/loops"; };
|
||||
path = [ pkgs.curl pkgs.bash pkgs.coreutils pkgs.gnutar pkgs.gzip pkgs.unzip ];
|
||||
script = ''
|
||||
if [ ! -x "$HOME/.local/bin/opencode" ]; then
|
||||
echo "installing opencode CLI for loops user..."
|
||||
curl -fsSL https://opencode.ai/install | bash || echo "install failed — retry on next activation"
|
||||
# The installer puts the binary in ~/.opencode/bin; every unit here expects ~/.local/bin.
|
||||
if [ -x "$HOME/.opencode/bin/opencode" ]; then
|
||||
mkdir -p "$HOME/.local/bin" && ln -sfn "$HOME/.opencode/bin/opencode" "$HOME/.local/bin/opencode"
|
||||
fi
|
||||
fi
|
||||
'';
|
||||
};
|
||||
|
||||
# ---- opencode web server: one shared instance the opencode-backed agents attach to -------
|
||||
# Provider creds come from /srv/cc-ci/.testenv (out of band, see README).
|
||||
systemd.services.opencode-web = {
|
||||
description = "opencode web server for cc-ci agents";
|
||||
wantedBy = [ "multi-user.target" ];
|
||||
after = [ "network-online.target" "tailscaled.service" "opencode-install.service" ];
|
||||
wants = [ "network-online.target" ];
|
||||
serviceConfig = {
|
||||
Type = "simple";
|
||||
User = "loops"; Group = "users";
|
||||
WorkingDirectory = "/srv/cc-ci-orch/cc-ci";
|
||||
EnvironmentFile = [ "-/srv/cc-ci/cc-ci/.env.public" "/srv/cc-ci/.testenv" ];
|
||||
ExecStartPre = "${pkgs.coreutils}/bin/rm -rf /tmp/opencode";
|
||||
ExecStart = "/home/loops/.local/bin/opencode serve --hostname 127.0.0.1 --port 4096";
|
||||
Restart = "on-failure";
|
||||
RestartSec = "5s";
|
||||
};
|
||||
environment = {
|
||||
HOME = "/home/loops";
|
||||
PATH = lib.mkForce "/run/wrappers/bin:/home/loops/.local/bin:/run/current-system/sw/bin:/usr/bin:/bin:/etc/profiles/per-user/loops/bin:/nix/var/nix/profiles/default/bin";
|
||||
};
|
||||
path = [ pkgs.bash pkgs.coreutils pkgs.git pkgs.python3 pkgs.openssh pkgs.tmux pkgs.nettools ];
|
||||
};
|
||||
|
||||
# ---- tailnet-only nginx front door for the opencode UI -------------------------------
|
||||
# Self-signed cert + basic auth, both created out of band (a store path would be world
|
||||
# readable) — see README "Secrets to stage". nginx FAILS TO START if they are missing.
|
||||
# /etc/nginx/oc-selfsigned.crt root:nginx 0644
|
||||
# /etc/nginx/oc-selfsigned.key root:nginx 0640
|
||||
# /etc/nginx/oc-htpasswd root:nginx 0640 (`oc:<bcrypt>`; plaintext in /secrets)
|
||||
services.nginx = {
|
||||
enable = true;
|
||||
recommendedProxySettings = true;
|
||||
virtualHosts.${cfg.opencodeUiHost} = {
|
||||
listen = [ { addr = "0.0.0.0"; port = cfg.opencodeUiPort; ssl = true; } ];
|
||||
# onlySSL flags the vhost as SSL so the module renders ssl_certificate for the listener.
|
||||
onlySSL = true;
|
||||
sslCertificate = "/etc/nginx/oc-selfsigned.crt";
|
||||
sslCertificateKey = "/etc/nginx/oc-selfsigned.key";
|
||||
basicAuthFile = "/etc/nginx/oc-htpasswd";
|
||||
locations."/" = {
|
||||
proxyPass = "http://127.0.0.1:4096";
|
||||
proxyWebsockets = true;
|
||||
};
|
||||
};
|
||||
};
|
||||
};
|
||||
}
|
||||
Reference in New Issue
Block a user