Files
cc-ci-orchestrator/.claude/skills/upgrade-all/SKILL.md
T
autonomic-botandClaude Opus 4.8 1bd156e7e6 weekly-run: pre-reclaim stale cc-ci images + hourly glm-5.2 supervisor
Root-cause fix for the 2026-07-03 run stalling: the cc-ci host disk filled to
100% (ENOSPC) mid-run (Wave 6, lasuite-drive), the agent stopped to reclaim
space, and nothing resumed it — the log-idle/429 watchdog only covers opencode-go
usage-limit stalls, not an environmental wedge.

- launch-upgrader.py: step-0 prereclaim_cc_ci() prunes STALE cc-ci docker images
  (unused AND older than a week, so this week's likely-reused images stay) before
  each weekly run. Best-effort; env-tunable (UPGRADER_PRERECLAIM*).
- launch-supervisor.py (new): hourly glm-5.2 orchestrator wake-up. Cheap
  deterministic gate — no-ops (zero tokens) when the run is complete or
  progressing; only when a run stalled/died before completing does it launch a
  short-lived glm-5.2 agent to diagnose + drive it to a clean DONE. Progress is
  judged by live run-proc + log mtime (session_busy() is claude-tuned and misreads
  a headless opencode run as idle).
- configuration.nix: cc-ci-upgrade-supervisor service + hourly timer (:07).
- upgrade-all SKILL §0: note the stale-image reclaim for manual runs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WxbpH3DquKzoSTSwGvGuET
2026-07-04 04:33:05 +00:00

20 KiB
Raw Blame History

name, description
name description
upgrade-all Weekly autonomous upgrade run for the cc-ci CI server. Surveys every enrolled recipe (except those tagged `external` in cc-ci-plan/used-recipes.md — used/tested but maintained elsewhere, e.g. uptime-kuma) for available upstream upgrades, then runs /recipe-upgrade on each upgradeable one via a subagent — plan, implement, verify green on cc-ci, open a recipe PR (and, only if a cc-ci test went stale, a verified cc-ci test PR). Collects results into one summary listing every PR to review. Rolling pool by default — works through recipes ALPHABETICALLY keeping DRONE_RUNNER_CAPACITY (the drone runner's slots, currently 2) subagents running at once, starting the next as each finishes; --sequential for one-at-a-time, --capacity N to override the pool size, --parallel to start all at once, --dry-run to preview. NEVER merges. Built to run once weekly on a cron. Invoke as /upgrade-all.

upgrade-all

The cc-ci analogue of recipe-maintainer's /recipe-upgrade-cron-all: an unattended weekly pass that keeps every enrolled recipe current, with cc-ci as the test gate and a human in the loop only where it matters — PR review. It surveys upgrades, runs /recipe-upgrade <recipe> per upgradeable recipe, and writes one summary of every PR to review. It never pushes upstream and never merges.

Drives cc-ci over ssh cc-ci. Logs/summary go to /srv/cc-ci/.cc-ci-logs/upgrades/.

Runs as the cc-ci-upgrader agent. This skill is normally executed by a dedicated, observable one-shot job agent — cc-ci-upgrader — spun up under remote-control (viewable/steerable at claude.ai/code, like the Builder) by cc-ci-plan/launch-upgrader.sh. That agent runs this skill to completion, then stops and stays idle so the run + summary remain reviewable in the web UI (it does NOT self-terminate). The weekly cron just invokes launch-upgrader.sh start; the next week's run clears the idle session and starts fresh. You can also run /upgrade-all inline in any /srv/cc-ci session, but the agent is the intended path so the weekly run isn't buried in headless output.

Arguments (optional $ARGUMENTS)

  • A space-separated list of recipe names → only those (else all enrolled recipes).
  • --dry-run → survey + print what WOULD upgrade; spawn nothing.
  • Default = rolling pool, alphabetical: work through the recipes in alphabetical order keeping DRONE_RUNNER_CAPACITY subagents running at once (the drone runner's slots — currently 2), starting the next recipe the moment one finishes. The 2026-06-10 concurrency restructure (docs/concurrency.md) makes concurrent recipe runs SAFE (per-run recipe trees + app-domain locks + isolation), and the capacity knob is the operator's resource-tuned ceiling — so matching the subagent pool to it uses all available concurrency without oversubscribing.
  • --capacity N → override the pool size (else the live DRONE_RUNNER_CAPACITY).
  • --sequential → force one subagent at a time (the old cautious default; use if the host is also busy with the build loops).
  • --parallel → unbounded: fan out ALL subagents at once, ignoring capacity (heaviest host load; drone still queues !testme beyond capacity, but every subagent's step-2b chaos deploy runs concurrently — only when the box is otherwise idle).

0. Sweep orphans from previous runs (FIRST, before anything else)

A prior run's teardown can crash, an agent can be killed mid-deploy, or a manual debug probe can be left running — leaving an orphan test stack, a standalone debug container, leaked volumes, or a stuck docker run wrapper that contends for the shared Swarm and skews the survey. Clear them before surveying. The sweep is safe by allowlist — it removes only what is NOT infra (traefik/drone/ backups/the ccci-* control plane) and NOT a warm-* canonical (whose retained volumes are spared), so it can never take down a live service:

ssh cc-ci 'bash -s' < /srv/cc-ci/.claude/skills/upgrade-all/sweep-orphans.sh

It is idempotent (a no-op when the host is already clean) and prints what it removed plus the surviving Swarm services (which should be infra + warm-* only — eyeball that before continuing). If anything legitimate looks at risk, stop and investigate rather than proceeding.

Then reap any leftover step-2b dev deploys from a prior run (the dev-* stacks /recipe-upgrade step 2b creates to debug an upgrade with live logs). The full sweep above already removes them, but run the dedicated reaper too so the start/end cleanup is explicit and symmetric (THRESHOLD=0 = clear ALL dev-*, since the run is quiescent now):

ssh cc-ci 'THRESHOLD=0 bash -s' < /srv/cc-ci/.claude/skills/upgrade-all/reap-dev-deploys.sh

Then reclaim leaked overlay IPs and guard against proxy VIP exhaustion. The shared proxy overlay (a /24 = 254 VIPs that EVERY recipe deploy joins) leaks endpoints under concurrent stack rm (a Swarm endpoint-GC race); over many days of churn the pool exhausts and new test deploys hang in Swarm New state with could not find an available IP while allocating VIP — which looks exactly like a recipe failure but is infra (root-caused 2026-06-12; see cc-ci-plan plan-proxy-vip-exhaustion-fix.md). Reclaim before the run:

# 1. reclaim leaked per-stack overlay networks (cheap, always safe)
ssh cc-ci 'docker network prune -f'
# 2. proxy VIP-exhaustion guard: if the allocator recently failed to assign a VIP, the leak has hit
#    the ceiling — rebuild the allocator with a docker restart (the box is QUIESCENT at run start, so
#    this is a ~30s infra blip that auto-recovers). Only fires when actually needed.
VIPFAIL=$(ssh cc-ci 'journalctl -u docker --since "26 hours ago" --no-pager 2>/dev/null | grep -c "available IP while allocating VIP"')
if [ "${VIPFAIL:-0}" -gt 0 ]; then
  echo "!! proxy VIP exhaustion detected ($VIPFAIL recent failures) — restarting docker to reclaim leaked endpoints"
  ssh cc-ci 'sudo systemctl restart docker'; sleep 25
  ssh cc-ci 'docker node ls && docker service ls --format "{{.Replicas}}" | grep -c "/"'   # sanity: node Ready, infra back
fi

(The durable fix — enlarging the proxy subnet to a /16 — landed 2026-06-13 (phase pvfix). This guard remains as belt-and-suspenders even after the /16 fix: it fires on the exact error signature and restarts docker to reclaim leaked endpoints if VIP exhaustion ever recurs despite the larger subnet.)

Then reclaim STALE docker images so the run can't fill the disk mid-flight. A full run deploys ~16 recipes; their images accumulate week over week and can run the cc-ci root FS to 100% (ENOSPC), which killed the 2026-07-03 run mid-way (lasuite-drive, Wave 6). Clear only stale images — unused by any container AND older than a week — so this week's likely-reused images are kept:

ssh cc-ci 'docker image prune -af --filter until=168h 2>&1 | tail -1; df -h / | tail -1'

(When the run is launched via launch-upgrader.py this is done automatically as step 0 — the prereclaim_cc_ci() pre-step — so you only run it by hand for a manual /upgrade-all.)

1. Build the candidate list

Enrolled recipes = the cc-ci tests/<recipe>/ dirs (same set ci-test-review sweeps), MINUS any recipe tagged external in cc-ci-plan/used-recipes.md — recipes cc-ci uses/tests but does NOT maintain (someone else upgrades them, e.g. uptime-kuma). used-recipes.md is the canonical inventory of every recipe we use; only its weekly rows get an upgrade survey + PR here.

EXTERNAL=$(awk '!/^[[:space:]]*#/ && $2=="external"{print $1}' /srv/cc-ci/cc-ci-plan/used-recipes.md)
ssh cc-ci 'cd /root/cc-ci/tests && ls -d */' | sed 's#/##' \
  | grep -vE '^(_generic|unit|__pycache__)$' \
  | grep -vxF -f <(printf '%s\n' "$EXTERNAL")    # drop externally-maintained recipes

(or the names passed in $ARGUMENTS — an explicit recipe name overrides the external skip, so you can still upgrade one on request.) For each candidate, on cc-ci, check availability — skip dirty/up-to-date. (If /root/cc-ci isn't present on the host, stage it first — see cc-ci-plan/plan-proxy-vip-exhaustion-fix.md / the host-rebuild memory for the staging step.)

⚠️ Four things that silently skip recipes — handle ALL FOUR per recipe before the version check:

  1. pseudo-TTY: abra FATAs inappropriate ioctl for device under plain ssh — wrap every abra call in script: ssh cc-ci 'script -qec "abra <args> -n" /dev/null' (git/other commands need none).
  2. go-git auth to git.autonomic.zone: recipes on the private mirror FATA unable to fetch tags … authentication required: Unauthorized — abra/go-git reads remote.origin.url literally and ignores git url.insteadOf. Bake creds into origin first (idempotent, only if origin is git.autonomic.zone): git -C ~/.abra/recipes/<r> remote set-url origin "https://$GITEA_USERNAME:$GITEA_PASSWORD@git.autonomic.zone/recipe-maintainers/<r>.git".
  3. dirty worktree is usually just the untracked cc-ci overlay (compose.ccci.yml) — git stash -u before / stash pop after; only skip dirty-worktree if TRACKED changes remain.
  4. tag+digest image pins abra can't parse — DO NOT SKIP, check upstream directly. abra FATAs failed to parse image <ref>, saw: Docker references with both a tag and digest are currently not supported and aborts the WHOLE recipe — even images it already parsed (e.g. immich: immich-server parses fine, then it dies on postgres:14-vectorchord…@sha256:…). This recipe is NOT not-fetchable. Catch the failure and check the upstream registry yourself:
    • Use abra's partial output for the images it DID parse.
    • Enumerate every image: ref in the compose; for each image abra could not parse (the FATA names it; plus any it never reached), strip the @sha256:… mentally and look up the live upstream tags directly — Docker Hub https://hub.docker.com/v2/repositories/<repo>/tags?page_size=100, ghcr https://ghcr.io/v2/<repo>/tags/list (anon bearer from https://ghcr.io/token?scope=repository:<repo>:pull&service=ghcr.io), or script -qec "abra ... " /dev/null-free docker buildx imagetools inspect <image>:<tag>.
    • Use judgment for non-semver tags: match the variant the app version supports, do NOT blindly take the highest (immich pins <pgmajor>-vectorchord<x>-pgvectors<y> to a combo immich-server supports — pick the newest tag on the SAME scheme/compatibility, not pg-18 just because it exists).
    • A recipe is upgradeable if abra OR this direct check finds a newer tag. The actual bump + digest re-pin happens in /recipe-upgrade (see its Implement step).
ssh cc-ci 'export PATH=/run/current-system/sw/bin:$PATH; R=<r>; \
  url=$(git -C ~/.abra/recipes/$R config --get remote.origin.url); \
  case "$url" in https://git.autonomic.zone/*) \
    git -C ~/.abra/recipes/$R remote set-url origin "https://'"$GITEA_USERNAME"':'"$GITEA_PASSWORD"'@${url#https://}";; esac; \
  git -C ~/.abra/recipes/$R stash -u >/dev/null 2>&1 || true; \
  script -qec "abra recipe fetch $R --force -n" /dev/null; \
  script -qec "abra recipe upgrade $R -m -n" /dev/null; \
  git -C ~/.abra/recipes/$R stash pop >/dev/null 2>&1 || true'

Build RECIPES_TO_UPGRADE = recipes with ≥1 available upgrade (after the stash, so the untracked overlay doesn't count as dirty). Others go to SKIPPED_UPFRONT with a reason (dirty-worktree only for real TRACKED edits, up-to-date, not-fetchable). A tag and digest … not supported parse FATA is NOT not-fetchable — per box item 4, catch it and resolve availability via the direct upstream-tag check; only not-fetchable when the mirror genuinely can't be fetched (network/auth), never for a digest-pinned image.

Reconcile every candidate's mirror during the survey (even the up-to-date ones), so merged-upstream PRs are closed and every mirror main tracks true upstream — fleet-wide, not just where there's a new upgrade:

set -a; . /srv/cc-ci/.testenv; set +a
ssh cc-ci "GITEA_USERNAME='$GITEA_USERNAME' GITEA_PASSWORD='$GITEA_PASSWORD' GITEA_URL='$GITEA_URL' bash -s <recipe> --reconcile-only" \
   < /srv/cc-ci/.claude/skills/recipe-upgrade/open-recipe-pr.sh

(The per-recipe /recipe-upgrade also reconciles, so this is belt-and-suspenders for skipped recipes — count closed-merged PRs in the summary.)

2. Determine concurrency, then print the plan

Read the live capacity (unless --capacity N / --sequential / --parallel was passed):

CAP=$(ssh cc-ci 'systemctl show drone-runner-exec -p Environment 2>/dev/null' | grep -oE 'DRONE_RUNNER_CAPACITY=[0-9]+' | cut -d= -f2)
CAP=${CAP:-2}   # fallback to the documented default if the query fails

--sequential → CAP=1; --capacity N → CAP=N; --parallel → CAP=∞ (all at once). Then print a table — Recipe | Status (will upgrade / skipped:reason) | Available upgrade(s) — plus the mode (Rolling pool (N=<CAP>, alphabetical) default / Sequential / Parallel). If --dry-run, stop here.

3. Upgrade each recipe via a subagent

For each recipe in RECIPES_TO_UPGRADE, spawn an Agent (subagent_type: "general-purpose", description "Upgrade <recipe> on cc-ci") with a prompt like:

Run the /recipe-upgrade <recipe> skill end-to-end — DEFAULT mode, do NOT pass --with-tests (/srv/cc-ci/.claude/skills/recipe-upgrade/SKILL.md): plan, implement, directly deploy + inspect the WIP on cc-ci for live feedback (step 2b, --chaos, tear it down after), open a recipe PR, and verify it by posting !testme on the PR (results visible in the PR; iterate ≤3×). Use the existing tests — if a test fails because it is genuinely stale, leave an explanatory comment on the PR for the operator and do NOT modify any test. Drive cc-ci over ssh cc-ci. Do NOT prompt. Do NOT push upstream. Do NOT merge. Print exactly one RESULT: line as your final line (SUCCESS / SUCCESS-PENDING-TESTS / FAILED / SKIPPED — SUCCESS+TESTPR will not occur in default mode).

Why default (no --with-tests): the weekly cron must never auto-edit cc-ci tests unattended — a test change deserves a human decision. So the cron opens recipe PRs only; where a recipe's existing test looks stale, the operator sees the explanation in the PR comment and can re-run that one recipe with /recipe-upgrade <recipe> --with-tests to also get a verified test-update PR.

Rolling pool (default), CAP running at once, recipes in ALPHABETICAL order. Keep exactly CAP recipe subagents in flight at all times — as soon as one finishes, immediately start the next alphabetical recipe. No waves, no heavy/light classification — just two (=CAP) always running. How:

  1. Sort RECIPES_TO_UPGRADE alphabetically.
  2. Start the first CAP as background Agents (run_in_background: true) — you are notified when each completes and can read its final output for the RESULT: line.
  3. On each completion: read that agent's RESULT: and record it; if recipes remain, immediately start the next alphabetical one as a background Agent (keeping CAP in flight).
  4. Repeat until every recipe has been started AND every subagent has completed, then go to §4. CAP=1 → strictly one-at-a-time (sequential). An Agent that dies/errors → record FAILED — agent tool error and still start its replacement; one failure never aborts the run.
  • Parallel (--parallel): no pool cap — start ALL recipes as background Agents at once. Heaviest load (every step-2b chaos deploy concurrent); only when the box is otherwise idle.

4. Collect results

Parse each final RESULT: line into SUCCESS / SUCCESS-PENDING-TESTS / FAILED / SKIPPED (default mode won't emit SUCCESS+TESTPR). A subagent that emitted no RESULT: line → FAILED — no result emitted.

4b. Reap dev deploys (END of run)

Every /recipe-upgrade subagent is required to tear down its own step-2b dev-<recipe> deploy, but once ALL recipes are done, reap any that leaked (a crashed/killed subagent) — the symmetric end-of-run cleanup to Step 0. The run is quiescent here, so clear ALL dev-* unconditionally (THRESHOLD=0):

ssh cc-ci 'THRESHOLD=0 bash -s' < /srv/cc-ci/.claude/skills/upgrade-all/reap-dev-deploys.sh

(Scoped to dev-* only — never touches the recipe PRs, CI per-run stacks, warm-*, or infra.)

5. Write + print the summary

Write /srv/cc-ci/.cc-ci-logs/upgrades/upgrade-all-<YYYY-MM-DD>.md and print it, leading with the PR list (the actionable output):

# cc-ci Weekly Upgrade Run — <YYYY-MM-DD>
## Summary
- Considered: N · PR green (!testme): N · PR open but tests look stale (commented): N · Failed: N · Skipped: N
## PRs to review (NOT merged)
- <recipe> <old> → <new> — recipe PR: <url> — !testme GREEN
## PRs where a test looks stale (operator: re-run `--with-tests` to update tests)
- <recipe> <old> → <new> — recipe PR: <url> — !testme RED on <test>; see PR comment
## Failed (investigate)
- <recipe> — at <step>: <reason>   (log: .cc-ci-logs/upgrades/<recipe>-upgrade-<date>.md)
## Skipped
| recipe | reason |

End with the report path and a reminder that nothing was merged.

6. Launch the public Recipe Report

Once the summary is written, kick off the weekly public report (its own agent; model overridable via REPORT_BACKEND/REPORT_MODEL, defaults to opencode-go/glm-5.2): python3 /srv/cc-ci/cc-ci-plan/launch-report.py fresh (use fresh, not start — the report is a one-shot and must always run a NEW session for THIS week, even if a previous report session is still around). It runs /recipe-report, reviews this run + the live recipe/PR state, and publishes to https://report.ci.commoninternet.net. Fire-and-forget — it runs independently; you can then go idle.

Safety / coordination (this matters — shared host with the build loops)

  • Concurrency is bounded to the drone capacity, and that's deliberate. The 2026-06-10 concurrency restructure (docs/concurrency.md) made concurrent recipe runs correct — per-run recipe trees, app-domain locks, isolation — so two recipes no longer collide/corrupt. The remaining limit is host RESOURCES (memory), which is exactly what DRONE_RUNNER_CAPACITY (=2) is tuned to. Running CAP subagents matches the runner's slots; do NOT exceed it on the shared box (that's what --parallel is for, and only when the box is idle). Each recipe-upgrade still tears down its own step-2b deploy; the end-of-run reap (§4b) clears any that leaked.
  • If the host is ALSO running the build loops, prefer --sequential (CAP=1) — the loops and a capacity-2 upgrade together can oversubscribe the 7.6 GB box (a Swarm task can wedge in New state under load). The orchestrator should serialize: don't run a capacity>1 upgrade concurrently with active phase loops.
  • Single-writer: every PR (recipe or cc-ci test) is on a dedicated branch; never push main, never touch the build loops' /cc-ci /cc-ci-adv working clones or their in-flight state.
  • Contention with active loop development: while the loops are still building cc-ci, this run competes for the host. Prefer a quiescent window; if a recipe fails due to contention it's simply retried next week. (Once cc-ci is built and the loops are idle, this is the steady-state weekly job.)
  • Never merges; failures/ skips are surfaced and retried next week — safe to re-run anytime.

Cron

Designed for a weekly Claude Code scheduled task that runs cc-ci-plan/launch-upgrader.sh start — which spins up the cc-ci-upgrader remote-control agent to run this skill to completion (the agent then stays idle/viewable). The cron does NOT invoke /upgrade-all inline.

Schedule (operator, 2026-05-29): the weekly cron is installed as the closing action of the final build phase (Phase 5), and its first run is ~1 hour after the build completes, weekly from then on — i.e. anchored to actual completion, not a fixed clock slot. Compute T0 = completion + 1h and install a weekly job at T0's day-of-week + HH:MM (MM HH * * DOW) running launch-upgrader.sh start. Then verify the first kickoff actually launches the cc-ci-upgrader agent. Do NOT install it earlier — while the loops are still building cc-ci it would contend for the shared host; until then run it manually via launch-upgrader.sh start/fresh. Full procedure: Phase 5 plan §4 (cc-ci-plan/plan-phase5-verify-upgrade-flow.md).

Re-running is idempotent: already-current recipes report SKIPPED — up-to-date; recipes with an open PR for the same branch report the existing PR rather than duplicating it.