terraform+nix: Hetzner orchestrator server (cpx11, nixos-infect, cc-ci-orchestrator-hetzner flake host)

Adds terraform/ to provision a Hetzner cpx11 (2 vCPU / 2 GB dedicated AMD / 40 GB NVMe)
for the loops runtime, and a flake + NixOS host config to converge it — replacing the slow
b1 Incus VM. Mirrors the cc-ci server terraform (same nixos-infect pin, same pattern).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
autonomic-bot
2026-05-31 02:11:30 +00:00
co-authored by Claude Sonnet 4.6
parent 4c418765c8
commit 0103f369ad
9 changed files with 422 additions and 0 deletions
+123
View File
@@ -0,0 +1,123 @@
# terraform — Hetzner cc-ci-orchestrator server
Provisions a Hetzner **cpx11** (2 vCPU / 2 GB dedicated AMD / 40 GB NVMe) for the cc-ci loops
runtime (Builder + Adversary + Watchdog + Orchestrator sessions), replacing the slow b1 Incus VM.
Uses nixos-infect to convert Debian → NixOS, then converges via the cc-ci-orchestrator flake.
---
## Stage 1 — provision the server
```bash
# from /srv/cc-ci/terraform/
source /srv/cc-ci/.testenv # loads HCLOUD_TOKEN
export TF_VAR_ssh_public_key="$(cat /home/loops/.ssh/cc-ci-root-ed25519.pub)"
tofu init
tofu plan
tofu apply
```
Note the `server_ipv4` output. nixos-infect runs on first boot — wait ~5 min, then:
```bash
# confirm NixOS is up (may need to retry while infect reboots)
ssh root@<server_ipv4> 'nixos-version'
```
---
## Stage 2 — converge to cc-ci-orchestrator-hetzner
### 2a. Capture hardware config
```bash
ssh root@<server_ipv4> 'cat /etc/nixos/hardware-configuration.nix'
```
Copy the output to `nix/hosts/cc-ci-orchestrator-hetzner/hardware.nix` in this repo, commit, push.
### 2b. Stage workspace on the new server
```bash
ssh root@<server_ipv4>
# Install Tailscale auth key (from .testenv TS_AUTH_KEY)
echo "<TS_AUTH_KEY>" > /etc/ts-auth-key && chmod 600 /etc/ts-auth-key
# Clone this repo as the loops user workspace
git clone --recursive \
https://autonomic-bot:<token>@git.autonomic.zone/recipe-maintainers/cc-ci-orchestrator.git \
/srv/cc-ci-orch
ln -sfn /srv/cc-ci-orch /srv/cc-ci # loops expect /srv/cc-ci
# Place master age key (copied from current VM .sops/master-age.txt)
mkdir -p /srv/cc-ci/.sops
scp loops@<old-vm>:/srv/cc-ci/.sops/master-age.txt /srv/cc-ci/.sops/master-age.txt
chmod 600 /srv/cc-ci/.sops/master-age.txt
```
### 2c. Run nixos-rebuild
```bash
# on the new server
cd /srv/cc-ci
nixos-rebuild switch --flake .#cc-ci-orchestrator-hetzner
```
### 2d. Stage credentials (not in git — placed once)
```bash
# SSH key for reaching cc-ci
mkdir -p /home/loops/.ssh && chmod 700 /home/loops/.ssh
# scp cc-ci-root-ed25519 from current VM or copy content
chmod 600 /home/loops/.ssh/cc-ci-root-ed25519
# .testenv (GITEA creds, etc.)
cp /path/to/.testenv /srv/cc-ci/.testenv && chmod 600 /srv/cc-ci/.testenv
```
### 2e. Auth claude and start loops
```bash
# as loops user on new server
sudo -u loops /home/loops/.local/bin/claude auth login # device code — operator step
# start the loops
cd /srv/cc-ci && sudo -u loops ./cc-ci-plan/launch.sh start
```
### 2f. Verify
```bash
tmux ls # should show cc-ci-builder, cc-ci-adv, cc-ci-watchdog
```
---
## Cutover
Once the new server is running and the loops are verified:
1. Update the `Host cc-ci` entry in the current VM's `/home/loops/.ssh/config` if needed
2. Stop the old Incus VM (or just leave it idle — it costs nothing in disk)
---
## Variables
| Variable | Default | Notes |
|---|---|---|
| `location` | `nbg1` | Nuremberg |
| `server_type` | `cpx11` | 2 vCPU / 2 GB dedicated AMD. Upgrade to `cpx21` (4 GB) if OOM. |
| `image` | `debian-12` | nixos-infect base |
| `server_name` | `cc-ci-orchestrator` | |
| `ssh_public_key` | required | Pass via `TF_VAR_ssh_public_key` |
---
## State
`terraform.tfstate` and `terraform.tfstate.backup` are gitignored. Keep the state file locally or
in a remote backend — losing it means `tofu destroy` can't find the server (use `tofu import` to
recover, or delete directly via the Hetzner console).
+32
View File
@@ -0,0 +1,32 @@
resource "hcloud_ssh_key" "cc_ci_orch" {
name = "cc-ci-orchestrator-deploy"
public_key = var.ssh_public_key
labels = {
project = "cc-ci-orchestrator"
managed = "terraform"
}
}
resource "hcloud_server" "cc_ci_orch" {
name = var.server_name
server_type = var.server_type
image = var.image
location = var.location
ssh_keys = [hcloud_ssh_key.cc_ci_orch.id]
# Stage 1: cloud-init runs nixos-infect on first boot, converting Debian to NixOS, then reboots.
# Wait ~5 min after apply, then SSH in and run Stage 2 per README.md.
user_data = file("${path.module}/user-data.sh")
public_net {
ipv4_enabled = true
ipv6_enabled = false
}
labels = {
project = "cc-ci-orchestrator"
managed = "terraform"
stage = "infect"
}
}
+19
View File
@@ -0,0 +1,19 @@
output "server_ipv4" {
description = "Public IPv4 address of the cc-ci-orchestrator Hetzner server"
value = hcloud_server.cc_ci_orch.ipv4_address
}
output "server_id" {
description = "Hetzner internal server ID"
value = hcloud_server.cc_ci_orch.id
}
output "ssh_connect" {
description = "SSH command to connect as root (after nixos-infect)"
value = "ssh root@${hcloud_server.cc_ci_orch.ipv4_address}"
}
output "nixos_infect_log" {
description = "Check infect progress"
value = "ssh root@${hcloud_server.cc_ci_orch.ipv4_address} 'cat /var/log/nixos-infect.log'"
}
+20
View File
@@ -0,0 +1,20 @@
#!/usr/bin/env bash
# Stage 1 — convert Debian 12 → NixOS via nixos-infect (pinned revision).
#
# nixos-infect generates /etc/nixos/{configuration.nix,hardware-configuration.nix,networking.nix}
# with Hetzner-correct bootloader (GRUB) and networking, then reboots into NixOS.
#
# After the reboot SSH as root is available. Run Stage 2 per terraform/README.md.
# Logs: /var/log/nixos-infect.log
set -euo pipefail
# Same pinned revision as the cc-ci server terraform (2026-03-22).
INFECT_SHA="40f62a680bb0e8f2f607d79abfaaecd99d59401c"
export NIX_CHANNEL="nixos-24.11"
export PROVIDER="hetzner"
export NIXOS_IMPORT=""
curl -fsSL "https://raw.githubusercontent.com/elitak/nixos-infect/${INFECT_SHA}/nixos-infect" \
| bash -x 2>&1 | tee /var/log/nixos-infect.log
+38
View File
@@ -0,0 +1,38 @@
variable "location" {
description = "Hetzner datacenter (nbg1=Nuremberg, fsn1=Falkenstein, hel1=Helsinki)"
type = string
default = "nbg1"
}
variable "server_type" {
description = <<-EOT
Hetzner server type. Must be x86 — the flake is x86_64-linux; NEVER use cax* (ARM).
cpx11 = AMD 2 vCPU / 2 GB (default; dedicated vCPU, NVMe — the orchestrator loops runtime).
cpx21 = AMD 3 vCPU / 4 GB (upgrade if claude sessions OOM under cpx11).
cx22 = AMD 2 vCPU / 4 GB (shared vCPU, cheaper alternative with more RAM).
EOT
type = string
default = "cpx11"
validation {
condition = !startswith(var.server_type, "cax")
error_message = "ARM server types (cax*) are not supported — the flake is x86_64-linux only."
}
}
variable "image" {
description = "Base OS image. nixos-infect supports debian-12 and ubuntu-24.04. debian-12 preferred."
type = string
default = "debian-12"
}
variable "ssh_public_key" {
description = "SSH public key content (the full line). Registered with Hetzner for root access post-infect. Pass via TF_VAR_ssh_public_key."
type = string
}
variable "server_name" {
description = "Hetzner server name and initial NixOS hostname"
type = string
default = "cc-ci-orchestrator"
}
+14
View File
@@ -0,0 +1,14 @@
terraform {
required_version = ">= 1.0"
required_providers {
hcloud = {
source = "hetznercloud/hcloud"
version = "1.64.0"
}
}
}
# The hcloud provider reads HCLOUD_TOKEN from the environment automatically.
# Never put the token value in any .tf file or .tfvars — keep it in the shell
# environment (export HCLOUD_TOKEN=...) or pass via TF_VAR_hcloud_token.
provider "hcloud" {}