nix/modules/auto-update.nix (own module, no notplants-nix dependency): Tuesday 03:00 UTC, busy-gated (CI run, weekly upgrader, report, sweep, running Drone builds → skip), `nix flake update` → build → switch-to-configuration test → cc-ci health checks (sshd, 0 failed units, core units, every swarm service at replica count, sops decrypted, dashboard/reports/drone 200, opencode UI 401) → profile + bootloader → flake.lock committed and pushed to main → /etc/cc-ci fast-forwarded; revert + lock restore on failure; one-line state file for /cc-ci-status. Skills (.opencode canonical, .claude pointers' descriptions synced): - cc-ci-orchestrator-update: THE host update — drives the auto-update unit by hand; --cc-ci-only for a cc-ci-main-only move; failure playbook. - cc-ci-server-update: delegates to it and explains why the old procedure (rebuilding the cc-ci repo's standalone #cc-ci) must not be run on this host. - cc-ci-update: chains orchestrator-update then tests-update. - cc-ci-status: §5/§6 for one host — auto-update state, generation vs boot generation, front doors (oc.ci = 401), sops, fail2ban, timers, orchestrator agent session, secrets inventory; verdict updated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FqkQq3CDmFWcQ7u1LzoyRz
187 lines
13 KiB
Markdown
187 lines
13 KiB
Markdown
---
|
||
name: cc-ci-status
|
||
description: Comprehensive read-only status check of the whole cc-ci system - how the recent weekly upgrade runs went and whether their reports published, which recipes/tests are stale, how long since the server + orchestrator host flake updates, open recipe PRs (flagging CVE-carrying PRs that have been open too long), host health (failed units, disk, timers, bridge/!testme path), ending with a verdict (ALL HEALTHY or a findings list) and recommended next steps mapped to the skills to invoke. Never changes anything - it only reads and reports. Invoke as /cc-ci-status.
|
||
---
|
||
|
||
# cc-ci-status
|
||
|
||
One comprehensive, **read-only** status pass over the cc-ci system. Output ends with either
|
||
**`ALL HEALTHY`** or a prioritized findings list, each finding paired with the skill to invoke.
|
||
Nothing here mutates state — no restarts, no deploys, no merges.
|
||
|
||
## Checks (run all; collect findings, don't stop at the first)
|
||
|
||
### 1. Weekly upgrade runs — recency + outcome
|
||
|
||
```
|
||
ls -t /srv/cc-ci/.cc-ci-logs/upgrades/upgrade-all-*.md | head -3
|
||
head -20 <latest> # the Summary block: Considered/green/stale/Failed/Skipped
|
||
systemctl list-timers cc-ci-upgrade-all.timer --no-pager | head -3
|
||
```
|
||
- **Overdue** if the newest report is >8 days old, or the timer is inactive/missing → recommend
|
||
`systemctl start cc-ci-upgrade-all.service` (or investigate the timer) / `/upgrade-all`.
|
||
- **Failed entries** in the latest report → recommend `/recipe-upgrade <recipe>` per entry (or
|
||
`/ci-test-review` if the failure is harness-side).
|
||
- A run currently in flight (tmux `cc-ci-upgrader` session live) is NOT a finding — report it as
|
||
in-progress and skip staleness checks that depend on its output.
|
||
|
||
### 2. Report publishing — report.ci.commoninternet.net
|
||
|
||
```
|
||
curl -s -o /dev/null -w '%{http_code}' https://report.ci.commoninternet.net/
|
||
ls -t /var/lib/cc-ci-reports/week-*.html | head -2 # via ssh cc-ci
|
||
```
|
||
- Index must be 200 and there must be a `week-*.html` at least as new as the last **completed**
|
||
upgrade run (a completed run without a matching page = report generation broke → recommend
|
||
`/recipe-report` / inspect `launch-report.py`).
|
||
|
||
### 3. Stale recipes / stale tests
|
||
|
||
- Latest report's "PRs where a test looks stale" section + carry-over notes.
|
||
- Any entries → recommend `/cc-ci-tests-update` (fleet) or `/recipe-upgrade <recipe>
|
||
--with-tests` (single).
|
||
- Also check `/srv/cc-ci/.cc-ci-logs/tests-update-*.md` recency — if stale tests were reported
|
||
weeks ago and no tests-update run since, say so.
|
||
|
||
### 4. Open recipe PRs — especially CVE-carrying ones that linger
|
||
|
||
Enumerate open PRs across `recipe-maintainers/*` (Gitea API, creds in `/srv/cc-ci/.testenv`):
|
||
```
|
||
GET /repos/recipe-maintainers/<repo>/pulls?state=open # repos = the used-recipes.md inventory + cc-ci
|
||
```
|
||
For each open PR: age (now − created_at), and whether the PR title/body/report row mentions
|
||
**CVE** / security patch.
|
||
- **CVE-carrying PR open >14 days** → HIGH-priority finding: name the PR, the CVE context, and
|
||
the blocker (commonly a stale test — check the report row) → recommend the unblocking skill
|
||
(`/cc-ci-tests-update`) plus "operator: review + merge <PR>".
|
||
- Non-CVE PRs open >30 days → low-priority note ("operator review backlog: N PRs").
|
||
- Verified-green PRs awaiting operator merge are normal — list them as "ready to merge", not
|
||
as failures.
|
||
|
||
### 5. The cc-ci host: update recency + the weekly auto-update
|
||
|
||
Since 2026-09-07 the CI server and the orchestrator are ONE Hetzner host (`cc-ci`,
|
||
195.201.88.249, no tailnet), built from `recipe-maintainers/cc-ci-orchestrator`'s
|
||
`nixosConfigurations.cc-ci` (which imports cc-ci's `nixosModules.cc-ci-server`). From the
|
||
notplants-orchestrator box `ssh cc-ci` reaches it as root; on the host itself `ssh cc-ci` is loopback.
|
||
|
||
```
|
||
ssh cc-ci 'nixos-version; readlink /run/current-system; readlink /nix/var/nix/profiles/system'
|
||
ssh cc-ci 'cd /srv/cc-ci-orch && nix flake metadata --json | jq -r ".locks.nodes | to_entries[] | \"\(.key) \(.value.locked.rev // \"-\") \(.value.locked.lastModified // \"-\")\""'
|
||
git ls-remote https://github.com/NixOS/nixpkgs nixos-26.05 | cut -c1-12 # channel tip
|
||
# the weekly self-update (nix/modules/auto-update.nix, Tue 03:00 UTC, busy-gated):
|
||
ssh cc-ci 'cat /srv/cc-ci/.cc-ci-logs/auto-update-state; systemctl list-timers --no-pager cc-ci-auto-update.timer | sed -n 2p; journalctl -u cc-ci-auto-update --no-pager -n 3'
|
||
```
|
||
- The state file is one line: `<UTC time> result=ok|unchanged|skipped|failed generation=<store path> note=<why>`.
|
||
- `ok` / `unchanged` within the last 8 days → healthy.
|
||
- `skipped` (busy) two weeks in a row, `failed`, or no state line / timer inactive for >8 days →
|
||
finding: *"host auto-update <state> since <date>: <note>"* → recommend
|
||
`/cc-ci-orchestrator-update` (it runs the same unit by hand and reads its journal).
|
||
- `failed` also means the host is running the previous generation and `flake.lock` was restored —
|
||
say so; the journal (`journalctl -u cc-ci-auto-update`) names the health check that failed.
|
||
- Running `/run/current-system` must equal the profile's boot generation (`system` link) — a
|
||
mismatch means a `test`-only activation was left behind (finding: switch or reboot decision).
|
||
- nixpkgs rev lagging the `nixos-26.05` channel tip by >30 days → finding → `/cc-ci-orchestrator-update`.
|
||
- `cc-ci` (the CI server input) lagging cc-ci `main` → note; the auto-update moves it weekly, or
|
||
`/cc-ci-orchestrator-update --cc-ci-only`.
|
||
|
||
### 6. Host + service health (one machine)
|
||
|
||
```
|
||
ssh cc-ci 'systemctl --failed --no-legend; df -h / | tail -1; free -m | sed -n 2p'
|
||
ssh cc-ci 'docker service ls --format "{{.Name}} {{.Replicas}}"; docker stack ls --format "{{.Name}}"'
|
||
ssh cc-ci 'for h in ci report.ci drone.ci oc.ci; do printf "%s " $h; curl -s -o /dev/null -w "%{http_code} ssl=%{ssl_verify_result}\n" https://$h.commoninternet.net/; done'
|
||
ssh cc-ci 'ls -s /run/secrets/test_secret; systemctl is-active acme-dns fail2ban nginx opencode-web drone-runner-exec; fail2ban-client status sshd | grep -E "Currently banned|Total banned"; fail2ban-client status nginx-http-auth | grep -E "Currently banned|Total banned"'
|
||
ssh cc-ci 'systemctl list-timers --no-pager cc-ci-upgrade-all.timer cc-ci-upgrade-supervisor.timer cc-ci-auto-update.timer nightly-sweep.timer | tail -n +2 | head -4'
|
||
ssh cc-ci 'sudo -iu loops tmux ls' # cc-ci-orchestrator must exist (the opencode orchestrator agent); cc-ci-upgrader only during a run
|
||
```
|
||
- Failed units, swarm services not at their replica count (`warm-*` spares flapping is benign —
|
||
note, don't page), a front door not 200 (`oc.ci` must be **401**: basic auth in front of the
|
||
opencode UI; 404 there = the `ccci-opencode-ui` stack is gone, `systemctl restart opencode-ui-route`
|
||
and check the weekly sweep's keep-list), `ssl≠0` (cert), `/run/secrets` empty (sops), acme-dns /
|
||
fail2ban / nginx / opencode-web / drone-runner-exec inactive, any timer missing, the orchestrator
|
||
tmux session missing (`systemctl restart cc-ci-orchestrator`) → findings. Disk **>65%** → finding.
|
||
Host unreachable → HIGH: `hetzner-server-recovery` (token in `/srv/cc-ci/.hcloud-token` on the
|
||
notplants box while the operator keeps it active).
|
||
> **65%, not 80% — it is not a steady-state measure.** Two concurrent recipe runs pull images and
|
||
> write volumes worth tens of GB, so a host sitting at 73% still hits 100% mid-run (2026-08-11:
|
||
> 63GB of unused images, `OSError: [Errno 28] No space left on device`). Remedy: `docker image
|
||
> prune -af` — it spares anything a container references. Do NOT `docker volume prune`: warm-*
|
||
> canonical volumes are data-warm and legitimately dangling.
|
||
- fail2ban: a non-zero "Currently banned" is information, not a finding (that is the jail
|
||
working); the notplants box's own address showing up banned IS a finding (it means a skill
|
||
misbehaved) — `fail2ban-client set <jail> unbanip <ip>`.
|
||
- Secrets hygiene: `ssh cc-ci 'ls /secrets/files'` must list ONLY cc-ci.testenv, cc-ci-local-ed25519(.pub),
|
||
autonomic-bot-cc-ci-ed25519(.pub), opencode-auth.json. Anything else there (a tailscale key, a
|
||
master age key, another host's ssh key…) → finding: *"secret sprawl on the cc-ci host"*.
|
||
- **!testme actually produces results** (the check that would have caught the above days earlier):
|
||
the newest few `/var/lib/cc-ci-runs/<build>/` dirs must each contain `results.json`. A build that
|
||
dies before the harness writes one leaves an EMPTY dir — and the PR badge still says "failure", so
|
||
it reads as a recipe regression rather than a sick host. Builds 1236–1242 all failed that way.
|
||
Finding: *"N recent builds produced no results.json — the harness is dying at startup, check disk
|
||
and the drone step log"*. The step log lives in drone's sqlite
|
||
(`/var/lib/docker/volumes/drone_ci_commoninternet_net_data/_data/database.sqlite`) — copy it and
|
||
read `logs.log_data` for the failing `steps.step_id`; the bridge's drone token is not extractable
|
||
(distroless container, swarm secret).
|
||
> **If the error is ENOSPC but the disk is fine**, it is not disk. Seen 2026-08-11: builds 1244-1249
|
||
> died on `mkdir /var/lib/cc-ci-runs/<build>` with **110GB free and 16% inodes**, while the identical
|
||
> mkdir succeeded as root over ssh, inside the runner's own mount namespace, and 61/61 times in a
|
||
> stress loop — and the same harness run by hand with a numeric run id worked fine. Restarting
|
||
> `drone-runner-exec` did NOT help, and neither did recreating the runs directory with a fresh
|
||
> inode (it recurred afterwards — that apparent fix was coincidence).
|
||
>
|
||
> **It is INTERMITTENT and tracks concurrent activity**, which is the useful signal: every failure
|
||
> landed while a second run or a manual deploy was in flight (1252 was triggered while 1251 was
|
||
> still finishing), and every build on a quiet host succeeded (1243, 1250, 1251, 1253). Free space
|
||
> never moved during a failing build. So on ENOSPC-with-free-disk: **wait for the host to go quiet
|
||
> and re-trigger** before treating it as a recipe failure. Root cause is still NOT established;
|
||
> `DRONE_RUNNER_CAPACITY=2` allows the overlap, so lowering it to 1 is the obvious next experiment
|
||
> if it becomes disruptive.
|
||
- **Bridge / !testme path**: `docker service ls` shows `ccci-bridge_app 1/1` AND the bridge log
|
||
has no auth errors (`docker service logs --since 24h ccci-bridge_app 2>&1 | grep -ci "401\|user does not exist"` == 0).
|
||
A silently-401ing bridge drops `!testme` (seen 2026-08-03, stale rotated Gitea secret) →
|
||
recommend refreshing the bridge secret + redeploy (test-before-switch).
|
||
|
||
### 7. Maintained-set consistency (quick)
|
||
|
||
- Every `weekly` row in `used-recipes.md` has `tests/<recipe>/` on cc-ci AND a
|
||
`recipe-maintainers/<recipe>` mirror AND is in bridge `POLL_REPOS`; mismatches → recommend
|
||
finishing enrollment (`/recipe-enroll <recipe>` covers all touchpoints).
|
||
|
||
## Output format
|
||
|
||
```
|
||
# cc-ci status — <date>
|
||
## Verdict: ALL HEALTHY | N findings (M high-priority)
|
||
## Weekly upgrades: <last run date + one-line outcome; next timer firing>
|
||
## Report site: <ok/broken + latest page>
|
||
## Stale tests: <none | list>
|
||
## Open PRs: <count; CVE-carrying + age flagged FIRST; ready-to-merge list>
|
||
## Host updates: server <rev, N days behind tip> · orchestrator <rev, N days>
|
||
## Health: server <failed/disk/services> · orchestrator <failed/disk/sessions> · bridge <ok/401s>
|
||
## Recommended next steps
|
||
1. <finding> → /<skill> (or operator action)
|
||
…
|
||
```
|
||
When a finding is that the fleet's **security exposure is unknown** — the last weekly run failed or
|
||
is stale, so nobody has scanned for CVEs recently — the recommended step is **`/cve-check`** (read-only,
|
||
minutes, no PRs). If it is instead that a known CVE is sitting unpatched, recommend
|
||
**`/cve-check-and-upgrade`** (add `--min-severity high` when only the urgent ones matter). Prefer
|
||
`/cve-check` over waiting for the next weekly run whenever the question is "are we exposed?".
|
||
|
||
`ALL HEALTHY` requires: recent successful weekly run + published report, no stale tests, no
|
||
CVE PR open >14 days, the host <30 days behind its channel with the weekly auto-update `ok`/`unchanged`
|
||
within 8 days and running generation == boot generation, zero failed units, every front door
|
||
answering (oc.ci with 401), sops decrypted, fail2ban up, the orchestrator agent session present,
|
||
recent builds all producing results.json, disk under threshold, bridge clean, maintained-set
|
||
consistent, secrets inventory exact. Anything else is a finding — even minor
|
||
ones get a recommended next step. Order findings by priority (CVE/unreachable-host first).
|
||
|
||
## Guardrails
|
||
|
||
- **Read-only.** This skill diagnoses and recommends; it invokes nothing and changes nothing.
|
||
- Use the pseudo-TTY wrap for any abra call; plain ssh for everything else. `ssh cc-ci` from the
|
||
notplants box is root on 195.201.88.249 (loops ssh config); there is no tailnet.
|
||
- Don't double-count: a finding that explains another (bridge 401 → !testme "failures") gets
|
||
reported once, at the root cause.
|