CI server health: image pruning, disk thresholds, reconcile-upstream #5

Merged
autonomic-bot merged 4 commits from review/4-ci-health into review/3-upstream-resolution 2026-08-11 19:03:55 +00:00
2 changed files with 42 additions and 2 deletions
Showing only changes of commit ab88e59c21 - Show all commits
@@ -73,6 +73,29 @@ done
# 5) Stray exited containers (debug one-shots) — best-effort prune.
docker container prune -f >/dev/null 2>&1 || true
# 6) Unused IMAGES — the one that actually took CI down. Every run pulls each recipe's images and
# nothing ever removed the old ones: on 2026-08-11 they had grown to 72GB (63GB of it unused),
# the root filesystem hit 100% under two concurrent runs, and the harness died at startup with
# `OSError: [Errno 28] No space left on device: '/var/lib/cc-ci-runs/<build>'`. Every !testme
# from build 1236 to 1242 failed that way — with no results.json, so the PR badges just read
# "failure" and looked like recipe regressions.
#
# Only prune above a threshold, so a healthy host keeps its layer cache and runs stay fast.
# `image prune -a` removes only images no container references, so anything deployed (infra +
# warm-* canonicals) is untouched; anything else is re-pulled on demand.
#
# Volumes are deliberately NOT pruned here — see the KEEP_RE guard in (3): warm-* canonicals are
# data-warm and their volumes are legitimately dangling between runs.
DISK_PRUNE_PCT="${DISK_PRUNE_PCT:-60}"
used_pct="$(df --output=pcent / 2>/dev/null | tail -1 | tr -dc '0-9')"
if [ -n "$used_pct" ] && [ "$used_pct" -ge "$DISK_PRUNE_PCT" ]; then
echo " disk ${used_pct}% >= ${DISK_PRUNE_PCT}% -> pruning unused images"
freed="$(docker image prune -af 2>/dev/null | awk '/Total reclaimed space/ {print $4, $5}')"
echo " reclaimed: ${freed:-0B}; disk now $(df -h / | tail -1 | awk '{print $5" used, "$4" free"}')"
else
echo " disk ${used_pct:-?}% < ${DISK_PRUNE_PCT}% -> keeping image cache"
fi
if [ "$removed" -eq 0 ]; then
echo "== orphan sweep: clean (nothing to remove) =="
else
+19 -2
View File
@@ -78,8 +78,24 @@ ssh cc-ci 'systemctl --failed --no-legend; df -h / | tail -1; docker service ls
systemctl --failed --no-legend; df -h / | tail -1; tmux ls
```
- Failed units, core swarm services not 1/1 (warm-* spares flapping is a known benign pattern —
note, don't page), disk >80% (server) / >85% (orchestrator) → findings. Server unreachable →
note, don't page), disk **>65% (server)** / >85% (orchestrator) → findings. Server unreachable →
HIGH: recommend `hetzner-server-recovery`.
> **65%, not 80%, on the server — it is not a steady-state measure.** Two concurrent recipe runs
> pull images and write volumes worth tens of GB, so a host sitting at 73% still hits 100% mid-run.
> That is exactly what happened on 2026-08-11: 63GB of unused images had accumulated (nothing ever
> pruned them), the filesystem filled during a run, and the harness died at startup with
> `OSError: [Errno 28] No space left on device`. Remedy: `docker image prune -af` on cc-ci — it
> spares anything a container references, so infra and warm-* canonicals are untouched. Do NOT
> `docker volume prune`: warm-* canonical volumes are data-warm and legitimately dangling.
- **!testme actually produces results** (the check that would have caught the above days earlier):
the newest few `/var/lib/cc-ci-runs/<build>/` dirs must each contain `results.json`. A build that
dies before the harness writes one leaves an EMPTY dir — and the PR badge still says "failure", so
it reads as a recipe regression rather than a sick host. Builds 12361242 all failed that way.
Finding: *"N recent builds produced no results.json — the harness is dying at startup, check disk
and the drone step log"*. The step log lives in drone's sqlite
(`/var/lib/docker/volumes/drone_ci_commoninternet_net_data/_data/database.sqlite`) — copy it and
read `logs.log_data` for the failing `steps.step_id`; the bridge's drone token is not extractable
(distroless container, swarm secret).
- **Bridge / !testme path**: `docker service ls` shows `ccci-bridge_app 1/1` AND the bridge log
has no auth errors (`docker service logs --since 24h ccci-bridge_app 2>&1 | grep -ci "401\|user does not exist"` == 0).
A silently-401ing bridge drops `!testme` (seen 2026-08-03, stale rotated Gitea secret) →
@@ -113,7 +129,8 @@ minutes, no PRs). If it is instead that a known CVE is sitting unpatched, recomm
`/cve-check` over waiting for the next weekly run whenever the question is "are we exposed?".
`ALL HEALTHY` requires: recent successful weekly run + published report, no stale tests, no
CVE PR open >14 days, both hosts <30 days behind their channel, zero failed units, disk under
CVE PR open >14 days, both hosts <30 days behind their channel, zero failed units, recent builds all
producing results.json, disk under
thresholds, bridge clean, maintained-set consistent. Anything else is a finding — even minor
ones get a recommended next step. Order findings by priority (CVE/unreachable-host first).