The CI server filled up and every !testme from build 1236 to 1242 died at harness startup with ENOSPC on /var/lib/cc-ci-runs/<build>. Because the harness never got far enough to write results.json, the PR badges just said 'failure' — so it read as recipe regressions, and plausible's genuinely-fixed suite looked still-broken. Cause: every run pulls each recipe's images and nothing ever removed the old ones. 72GB of images, 63GB of it unused. Reclaimed 69.8GB; the host went 73% -> 22%. Two changes so it does not recur: - sweep-orphans.sh (runs at the start AND end of every /upgrade-all) now prunes unused images when the disk is >=60% (DISK_PRUNE_PCT). Below that it keeps the layer cache so runs stay fast. 'docker image prune -a' spares anything a container references, so infra and warm-* canonicals are safe. Volumes are still NOT pruned — warm-* canonical volumes are data-warm and legitimately dangling. - /cc-ci-status flags server disk at >65% rather than >80%, because this is not a steady-state measure: the host was at 73% when runs started failing. It also now checks that recent builds actually produced results.json — an empty run dir is the fingerprint of a host problem masquerading as a recipe failure — and records how to read a drone step log out of its sqlite when the API token is unreachable.
106 lines
5.8 KiB
Bash
Executable File
106 lines
5.8 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Sweep orphaned test deployments + debug debris left by PREVIOUS cc-ci runs.
|
|
#
|
|
# Runs ON the cc-ci host (root) — pipe it in: ssh cc-ci 'bash -s' < sweep-orphans.sh
|
|
# Invoked at the START of /upgrade-all so a leaked stack/container/volume/process from a prior run
|
|
# (a teardown that crashed, a manual debug probe, a killed agent) does not contend for the shared
|
|
# Swarm or skew the survey. Idempotent + safe to run anytime: a no-op when the host is already clean.
|
|
#
|
|
# SAFE BY ALLOWLIST. It removes ONLY things NOT on the keep-list, so it can never take down infra or
|
|
# the warm canonicals. The keep-list (leading name prefix) is:
|
|
# - traefik, drone, backups : Swarm + CI infra
|
|
# - ccci-bridge / -dashboard / -reports: the cc-ci control plane
|
|
# - warm-* : warm canonicals (idle persistent deps reused across runs;
|
|
# their retained volumes are spared too)
|
|
# Everything else deployed on the Swarm is a per-run test stack and is fair game.
|
|
set -uo pipefail
|
|
export PATH=/run/current-system/sw/bin:$PATH
|
|
|
|
KEEP_RE='^(traefik|drone|backups|ccci-(bridge|dashboard|reports)|warm-)'
|
|
removed=0
|
|
|
|
echo "== orphan sweep: scanning (keep-list: infra + warm-* canonicals) =="
|
|
|
|
# 1) Orphan Swarm stacks — any deployed stack not on the keep-list is a leftover per-run test deploy
|
|
# (the harness deploys each recipe under its own per-run stack; a clean run tears it down).
|
|
mapfile -t STACKS < <(docker stack ls --format '{{.Name}}' 2>/dev/null)
|
|
for s in "${STACKS[@]}"; do
|
|
[ -z "$s" ] && continue
|
|
if printf '%s' "$s" | grep -Eq "$KEEP_RE"; then continue; fi
|
|
echo " orphan stack -> removing: $s"
|
|
docker stack rm "$s" >/dev/null 2>&1 || true
|
|
removed=$((removed + 1))
|
|
done
|
|
# wait (bounded) for removed stacks' services to drain so their volumes free up for step 3
|
|
for _ in $(seq 1 30); do
|
|
[ -z "$(docker service ls --format '{{.Name}}' 2>/dev/null | grep -Ev "$KEEP_RE")" ] && break
|
|
sleep 2
|
|
done
|
|
|
|
# 2) Orphan standalone containers — running/exited containers NOT managed by Swarm (no
|
|
# com.docker.swarm.service.id label) are debug `docker run` leftovers (e.g. the plausible
|
|
# clickhouse entrypoint probes). Remove only ones started > 30 min ago, so an in-flight manual
|
|
# probe started right before the run is spared.
|
|
now=$(date +%s)
|
|
for c in $(docker ps -aq 2>/dev/null); do
|
|
sid=$(docker inspect -f '{{ index .Config.Labels "com.docker.swarm.service.id" }}' "$c" 2>/dev/null)
|
|
[ -n "$sid" ] && continue # Swarm-managed → keep
|
|
started=$(docker inspect -f '{{.State.StartedAt}}' "$c" 2>/dev/null)
|
|
st=$(date -d "$started" +%s 2>/dev/null || echo "$now")
|
|
[ $((now - st)) -lt 1800 ] && continue # younger than 30 min → spare a fresh manual probe
|
|
echo " orphan container -> removing: $(docker inspect -f '{{.Name}} ({{.Config.Image}})' "$c" 2>/dev/null)"
|
|
docker rm -f "$c" >/dev/null 2>&1 || true
|
|
removed=$((removed + 1))
|
|
done
|
|
|
|
# 3) Leaked volumes — dangling (referenced by no container) volumes left by removed test stacks.
|
|
# Spare warm-* volumes: a warm canonical idles UNDEPLOYED with its data volume retained, so its
|
|
# volume is legitimately dangling and must NOT be pruned.
|
|
for v in $(docker volume ls -qf dangling=true 2>/dev/null); do
|
|
printf '%s' "$v" | grep -Eq "$KEEP_RE" && continue
|
|
docker volume rm "$v" >/dev/null 2>&1 && { echo " leaked volume -> removed: $v"; removed=$((removed + 1)); }
|
|
done
|
|
|
|
# 4) Orphan debug host-processes — reparented (ppid==1) `timeout … docker run …` wrappers left by
|
|
# manual recipe probes; they outlive their container and never self-reap. Kill by explicit PID
|
|
# (NEVER pkill -f, which would self-match this script's own command line).
|
|
for p in $(ps -eo pid=,ppid=,cmd= 2>/dev/null | awk '$2==1 && /timeout/ && /docker[[:space:]]+run/ {print $1}'); do
|
|
echo " orphan debug wrapper -> killing pid $p"
|
|
kill "$p" 2>/dev/null || true
|
|
removed=$((removed + 1))
|
|
done
|
|
|
|
# 5) Stray exited containers (debug one-shots) — best-effort prune.
|
|
docker container prune -f >/dev/null 2>&1 || true
|
|
|
|
# 6) Unused IMAGES — the one that actually took CI down. Every run pulls each recipe's images and
|
|
# nothing ever removed the old ones: on 2026-08-11 they had grown to 72GB (63GB of it unused),
|
|
# the root filesystem hit 100% under two concurrent runs, and the harness died at startup with
|
|
# `OSError: [Errno 28] No space left on device: '/var/lib/cc-ci-runs/<build>'`. Every !testme
|
|
# from build 1236 to 1242 failed that way — with no results.json, so the PR badges just read
|
|
# "failure" and looked like recipe regressions.
|
|
#
|
|
# Only prune above a threshold, so a healthy host keeps its layer cache and runs stay fast.
|
|
# `image prune -a` removes only images no container references, so anything deployed (infra +
|
|
# warm-* canonicals) is untouched; anything else is re-pulled on demand.
|
|
#
|
|
# Volumes are deliberately NOT pruned here — see the KEEP_RE guard in (3): warm-* canonicals are
|
|
# data-warm and their volumes are legitimately dangling between runs.
|
|
DISK_PRUNE_PCT="${DISK_PRUNE_PCT:-60}"
|
|
used_pct="$(df --output=pcent / 2>/dev/null | tail -1 | tr -dc '0-9')"
|
|
if [ -n "$used_pct" ] && [ "$used_pct" -ge "$DISK_PRUNE_PCT" ]; then
|
|
echo " disk ${used_pct}% >= ${DISK_PRUNE_PCT}% -> pruning unused images"
|
|
freed="$(docker image prune -af 2>/dev/null | awk '/Total reclaimed space/ {print $4, $5}')"
|
|
echo " reclaimed: ${freed:-0B}; disk now $(df -h / | tail -1 | awk '{print $5" used, "$4" free"}')"
|
|
else
|
|
echo " disk ${used_pct:-?}% < ${DISK_PRUNE_PCT}% -> keeping image cache"
|
|
fi
|
|
|
|
if [ "$removed" -eq 0 ]; then
|
|
echo "== orphan sweep: clean (nothing to remove) =="
|
|
else
|
|
echo "== orphan sweep: removed/killed $removed orphan(s) =="
|
|
fi
|
|
echo "-- surviving Swarm services (should be infra + warm-* only) --"
|
|
docker service ls --format '{{.Name}}' 2>/dev/null | sort
|