Make explicit that ALL formatting/HTML is owned by recipe-report.py render() and the model's only artifact is the spec JSON — never hand-write/edit HTML. Matters now that glm-5.2 drives the report. Also fix stale 'default opus' refs (report now defaults to opencode-go/glm-5.2, overridable via REPORT_BACKEND/REPORT_MODEL). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
20 KiB
name, description
| name | description |
|---|---|
| upgrade-all | Weekly autonomous upgrade run for the cc-ci CI server. Surveys every enrolled recipe (except those tagged `external` in cc-ci-plan/used-recipes.md — used/tested but maintained elsewhere, e.g. uptime-kuma) for available upstream upgrades, then runs /recipe-upgrade on each upgradeable one via a subagent — plan, implement, verify green on cc-ci, open a recipe PR (and, only if a cc-ci test went stale, a verified cc-ci test PR). Collects results into one summary listing every PR to review. Rolling pool by default — works through recipes ALPHABETICALLY keeping DRONE_RUNNER_CAPACITY (the drone runner's slots, currently 2) subagents running at once, starting the next as each finishes; --sequential for one-at-a-time, --capacity N to override the pool size, --parallel to start all at once, --dry-run to preview. NEVER merges. Built to run once weekly on a cron. Invoke as /upgrade-all. |
upgrade-all
The cc-ci analogue of recipe-maintainer's /recipe-upgrade-cron-all: an unattended weekly pass
that keeps every enrolled recipe current, with cc-ci as the test gate and a human in the loop only
where it matters — PR review. It surveys upgrades, runs /recipe-upgrade <recipe> per upgradeable
recipe, and writes one summary of every PR to review. It never pushes upstream and never merges.
Drives cc-ci over ssh cc-ci. Logs/summary go to /srv/cc-ci/.cc-ci-logs/upgrades/.
Runs as the cc-ci-upgrader agent. This skill is normally executed by a dedicated, observable
one-shot job agent — cc-ci-upgrader — spun up under remote-control (viewable/steerable at
claude.ai/code, like the Builder) by cc-ci-plan/launch-upgrader.sh. That agent runs this skill to
completion, then stops and stays idle so the run + summary remain reviewable in the web UI (it does
NOT self-terminate). The weekly cron just invokes launch-upgrader.sh start; the next week's run
clears the idle session and starts fresh. You can also run /upgrade-all inline in any /srv/cc-ci
session, but the agent is the intended path so the weekly run isn't buried in headless output.
Arguments (optional $ARGUMENTS)
- A space-separated list of recipe names → only those (else all enrolled recipes).
--dry-run→ survey + print what WOULD upgrade; spawn nothing.- Default = rolling pool, alphabetical: work through the recipes in alphabetical order keeping
DRONE_RUNNER_CAPACITYsubagents running at once (the drone runner's slots — currently2), starting the next recipe the moment one finishes. The 2026-06-10 concurrency restructure (docs/concurrency.md) makes concurrent recipe runs SAFE (per-run recipe trees + app-domain locks + isolation), and the capacity knob is the operator's resource-tuned ceiling — so matching the subagent pool to it uses all available concurrency without oversubscribing. --capacity N→ override the pool size (else the liveDRONE_RUNNER_CAPACITY).--sequential→ force one subagent at a time (the old cautious default; use if the host is also busy with the build loops).--parallel→ unbounded: fan out ALL subagents at once, ignoring capacity (heaviest host load; drone still queues!testmebeyond capacity, but every subagent's step-2b chaos deploy runs concurrently — only when the box is otherwise idle).
0. Sweep orphans from previous runs (FIRST, before anything else)
A prior run's teardown can crash, an agent can be killed mid-deploy, or a manual debug probe can be
left running — leaving an orphan test stack, a standalone debug container, leaked volumes, or a stuck
docker run wrapper that contends for the shared Swarm and skews the survey. Clear them before
surveying. The sweep is safe by allowlist — it removes only what is NOT infra (traefik/drone/
backups/the ccci-* control plane) and NOT a warm-* canonical (whose retained volumes are spared),
so it can never take down a live service:
ssh cc-ci 'bash -s' < /srv/cc-ci/.claude/skills/upgrade-all/sweep-orphans.sh
It is idempotent (a no-op when the host is already clean) and prints what it removed plus the
surviving Swarm services (which should be infra + warm-* only — eyeball that before continuing). If
anything legitimate looks at risk, stop and investigate rather than proceeding.
Then reap any leftover step-2b dev deploys from a prior run (the dev-* stacks /recipe-upgrade
step 2b creates to debug an upgrade with live logs). The full sweep above already removes them, but run
the dedicated reaper too so the start/end cleanup is explicit and symmetric (THRESHOLD=0 = clear ALL
dev-*, since the run is quiescent now):
ssh cc-ci 'THRESHOLD=0 bash -s' < /srv/cc-ci/.claude/skills/upgrade-all/reap-dev-deploys.sh
Then reclaim leaked overlay IPs and guard against proxy VIP exhaustion. The shared proxy
overlay (a /24 = 254 VIPs that EVERY recipe deploy joins) leaks endpoints under concurrent stack
rm (a Swarm endpoint-GC race); over many days of churn the pool exhausts and new test deploys hang
in Swarm New state with could not find an available IP while allocating VIP — which looks exactly
like a recipe failure but is infra (root-caused 2026-06-12; see cc-ci-plan plan-proxy-vip-exhaustion-fix.md).
Reclaim before the run:
# 1. reclaim leaked per-stack overlay networks (cheap, always safe)
ssh cc-ci 'docker network prune -f'
# 2. proxy VIP-exhaustion guard: if the allocator recently failed to assign a VIP, the leak has hit
# the ceiling — rebuild the allocator with a docker restart (the box is QUIESCENT at run start, so
# this is a ~30s infra blip that auto-recovers). Only fires when actually needed.
VIPFAIL=$(ssh cc-ci 'journalctl -u docker --since "26 hours ago" --no-pager 2>/dev/null | grep -c "available IP while allocating VIP"')
if [ "${VIPFAIL:-0}" -gt 0 ]; then
echo "!! proxy VIP exhaustion detected ($VIPFAIL recent failures) — restarting docker to reclaim leaked endpoints"
ssh cc-ci 'sudo systemctl restart docker'; sleep 25
ssh cc-ci 'docker node ls && docker service ls --format "{{.Replicas}}" | grep -c "/"' # sanity: node Ready, infra back
fi
(The durable fix — enlarging the proxy subnet to a /16 — landed 2026-06-13 (phase pvfix). This guard
remains as belt-and-suspenders even after the /16 fix: it fires on the exact error signature and restarts
docker to reclaim leaked endpoints if VIP exhaustion ever recurs despite the larger subnet.)
1. Build the candidate list
Enrolled recipes = the cc-ci tests/<recipe>/ dirs (same set ci-test-review sweeps), MINUS any
recipe tagged external in cc-ci-plan/used-recipes.md — recipes cc-ci uses/tests but does NOT
maintain (someone else upgrades them, e.g. uptime-kuma). used-recipes.md is the canonical
inventory of every recipe we use; only its weekly rows get an upgrade survey + PR here.
EXTERNAL=$(awk '!/^[[:space:]]*#/ && $2=="external"{print $1}' /srv/cc-ci/cc-ci-plan/used-recipes.md)
ssh cc-ci 'cd /root/cc-ci/tests && ls -d */' | sed 's#/##' \
| grep -vE '^(_generic|unit|__pycache__)$' \
| grep -vxF -f <(printf '%s\n' "$EXTERNAL") # drop externally-maintained recipes
(or the names passed in $ARGUMENTS — an explicit recipe name overrides the external skip, so you
can still upgrade one on request.) For each candidate, on cc-ci, check availability — skip
dirty/up-to-date. (If /root/cc-ci isn't present on the host, stage it first — see
cc-ci-plan/plan-proxy-vip-exhaustion-fix.md / the host-rebuild memory for the staging step.)
⚠️ Four things that silently skip recipes — handle ALL FOUR per recipe before the version check:
- pseudo-TTY: abra FATAs
inappropriate ioctl for deviceunder plain ssh — wrap every abra call inscript:ssh cc-ci 'script -qec "abra <args> -n" /dev/null'(git/other commands need none).- go-git auth to git.autonomic.zone: recipes on the private mirror FATA
unable to fetch tags … authentication required: Unauthorized— abra/go-git readsremote.origin.urlliterally and ignoresgit url.insteadOf. Bake creds into origin first (idempotent, only if origin is git.autonomic.zone):git -C ~/.abra/recipes/<r> remote set-url origin "https://$GITEA_USERNAME:$GITEA_PASSWORD@git.autonomic.zone/recipe-maintainers/<r>.git".- dirty worktree is usually just the untracked cc-ci overlay (
compose.ccci.yml) —git stash -ubefore /stash popafter; only skipdirty-worktreeif TRACKED changes remain.- tag+digest image pins abra can't parse — DO NOT SKIP, check upstream directly. abra FATAs
failed to parse image <ref>, saw: Docker references with both a tag and digest are currently not supportedand aborts the WHOLE recipe — even images it already parsed (e.g. immich:immich-serverparses fine, then it dies onpostgres:14-vectorchord…@sha256:…). This recipe is NOTnot-fetchable. Catch the failure and check the upstream registry yourself:
- Use abra's partial output for the images it DID parse.
- Enumerate every
image:ref in the compose; for each image abra could not parse (the FATA names it; plus any it never reached), strip the@sha256:…mentally and look up the live upstream tags directly — Docker Hubhttps://hub.docker.com/v2/repositories/<repo>/tags?page_size=100, ghcrhttps://ghcr.io/v2/<repo>/tags/list(anon bearer fromhttps://ghcr.io/token?scope=repository:<repo>:pull&service=ghcr.io), orscript -qec "abra ... " /dev/null-freedocker buildx imagetools inspect <image>:<tag>.- Use judgment for non-semver tags: match the variant the app version supports, do NOT blindly take the highest (immich pins
<pgmajor>-vectorchord<x>-pgvectors<y>to a combo immich-server supports — pick the newest tag on the SAME scheme/compatibility, not pg-18 just because it exists).- A recipe is upgradeable if abra OR this direct check finds a newer tag. The actual bump + digest re-pin happens in
/recipe-upgrade(see its Implement step).
ssh cc-ci 'export PATH=/run/current-system/sw/bin:$PATH; R=<r>; \
url=$(git -C ~/.abra/recipes/$R config --get remote.origin.url); \
case "$url" in https://git.autonomic.zone/*) \
git -C ~/.abra/recipes/$R remote set-url origin "https://'"$GITEA_USERNAME"':'"$GITEA_PASSWORD"'@${url#https://}";; esac; \
git -C ~/.abra/recipes/$R stash -u >/dev/null 2>&1 || true; \
script -qec "abra recipe fetch $R --force -n" /dev/null; \
script -qec "abra recipe upgrade $R -m -n" /dev/null; \
git -C ~/.abra/recipes/$R stash pop >/dev/null 2>&1 || true'
Build RECIPES_TO_UPGRADE = recipes with ≥1 available upgrade (after the stash, so the untracked
overlay doesn't count as dirty). Others go to SKIPPED_UPFRONT with a reason (dirty-worktree only for
real TRACKED edits, up-to-date, not-fetchable). A tag and digest … not supported parse FATA is
NOT not-fetchable — per box item 4, catch it and resolve availability via the direct upstream-tag
check; only not-fetchable when the mirror genuinely can't be fetched (network/auth), never for a
digest-pinned image.
Reconcile every candidate's mirror during the survey (even the up-to-date ones), so merged-upstream
PRs are closed and every mirror main tracks true upstream — fleet-wide, not just where there's a new
upgrade:
set -a; . /srv/cc-ci/.testenv; set +a
ssh cc-ci "GITEA_USERNAME='$GITEA_USERNAME' GITEA_PASSWORD='$GITEA_PASSWORD' GITEA_URL='$GITEA_URL' bash -s <recipe> --reconcile-only" \
< /srv/cc-ci/.claude/skills/recipe-upgrade/open-recipe-pr.sh
(The per-recipe /recipe-upgrade also reconciles, so this is belt-and-suspenders for skipped recipes —
count closed-merged PRs in the summary.)
2. Determine concurrency, then print the plan
Read the live capacity (unless --capacity N / --sequential / --parallel was passed):
CAP=$(ssh cc-ci 'systemctl show drone-runner-exec -p Environment 2>/dev/null' | grep -oE 'DRONE_RUNNER_CAPACITY=[0-9]+' | cut -d= -f2)
CAP=${CAP:-2} # fallback to the documented default if the query fails
--sequential → CAP=1; --capacity N → CAP=N; --parallel → CAP=∞ (all at once).
Then print a table — Recipe | Status (will upgrade / skipped:reason) | Available upgrade(s) — plus
the mode (Rolling pool (N=<CAP>, alphabetical) default / Sequential / Parallel). If --dry-run,
stop here.
3. Upgrade each recipe via a subagent
For each recipe in RECIPES_TO_UPGRADE, spawn an Agent (subagent_type: "general-purpose",
description "Upgrade <recipe> on cc-ci") with a prompt like:
Run the
/recipe-upgrade <recipe>skill end-to-end — DEFAULT mode, do NOT pass--with-tests(/srv/cc-ci/.claude/skills/recipe-upgrade/SKILL.md): plan, implement, directly deploy + inspect the WIP on cc-ci for live feedback (step 2b,--chaos, tear it down after), open a recipe PR, and verify it by posting!testmeon the PR (results visible in the PR; iterate ≤3×). Use the existing tests — if a test fails because it is genuinely stale, leave an explanatory comment on the PR for the operator and do NOT modify any test. Drive cc-ci overssh cc-ci. Do NOT prompt. Do NOT push upstream. Do NOT merge. Print exactly oneRESULT:line as your final line (SUCCESS / SUCCESS-PENDING-TESTS / FAILED / SKIPPED —SUCCESS+TESTPRwill not occur in default mode).
Why default (no --with-tests): the weekly cron must never auto-edit cc-ci tests unattended — a
test change deserves a human decision. So the cron opens recipe PRs only; where a recipe's existing
test looks stale, the operator sees the explanation in the PR comment and can re-run that one recipe
with /recipe-upgrade <recipe> --with-tests to also get a verified test-update PR.
Rolling pool (default), CAP running at once, recipes in ALPHABETICAL order. Keep exactly CAP
recipe subagents in flight at all times — as soon as one finishes, immediately start the next
alphabetical recipe. No waves, no heavy/light classification — just two (=CAP) always running.
How:
- Sort
RECIPES_TO_UPGRADEalphabetically. - Start the first
CAPas background Agents (run_in_background: true) — you are notified when each completes and can read its final output for theRESULT:line. - On each completion: read that agent's
RESULT:and record it; if recipes remain, immediately start the next alphabetical one as a background Agent (keepingCAPin flight). - Repeat until every recipe has been started AND every subagent has completed, then go to §4.
CAP=1→ strictly one-at-a-time (sequential). An Agent that dies/errors → recordFAILED — agent tool errorand still start its replacement; one failure never aborts the run.
- Parallel (
--parallel): no pool cap — start ALL recipes as background Agents at once. Heaviest load (every step-2b chaos deploy concurrent); only when the box is otherwise idle.
4. Collect results
Parse each final RESULT: line into SUCCESS / SUCCESS-PENDING-TESTS / FAILED / SKIPPED (default mode
won't emit SUCCESS+TESTPR). A subagent that emitted no RESULT: line → FAILED — no result emitted.
4b. Reap dev deploys (END of run)
Every /recipe-upgrade subagent is required to tear down its own step-2b dev-<recipe> deploy, but
once ALL recipes are done, reap any that leaked (a crashed/killed subagent) — the symmetric end-of-run
cleanup to Step 0. The run is quiescent here, so clear ALL dev-* unconditionally (THRESHOLD=0):
ssh cc-ci 'THRESHOLD=0 bash -s' < /srv/cc-ci/.claude/skills/upgrade-all/reap-dev-deploys.sh
(Scoped to dev-* only — never touches the recipe PRs, CI per-run stacks, warm-*, or infra.)
5. Write + print the summary
Write /srv/cc-ci/.cc-ci-logs/upgrades/upgrade-all-<YYYY-MM-DD>.md and print it, leading with the
PR list (the actionable output):
# cc-ci Weekly Upgrade Run — <YYYY-MM-DD>
## Summary
- Considered: N · PR green (!testme): N · PR open but tests look stale (commented): N · Failed: N · Skipped: N
## PRs to review (NOT merged)
- <recipe> <old> → <new> — recipe PR: <url> — !testme GREEN
## PRs where a test looks stale (operator: re-run `--with-tests` to update tests)
- <recipe> <old> → <new> — recipe PR: <url> — !testme RED on <test>; see PR comment
## Failed (investigate)
- <recipe> — at <step>: <reason> (log: .cc-ci-logs/upgrades/<recipe>-upgrade-<date>.md)
## Skipped
| recipe | reason |
End with the report path and a reminder that nothing was merged.
6. Launch the public Recipe Report
Once the summary is written, kick off the weekly public report (its own agent; model overridable via
REPORT_BACKEND/REPORT_MODEL, defaults to opencode-go/glm-5.2): python3 /srv/cc-ci/cc-ci-plan/launch-report.py fresh (use fresh, not start — the report
is a one-shot and must always run a NEW session for THIS week, even if a previous report session is
still around). It runs /recipe-report, reviews this run + the live recipe/PR state, and publishes to
https://report.ci.commoninternet.net. Fire-and-forget — it runs independently; you can then go idle.
Safety / coordination (this matters — shared host with the build loops)
- Concurrency is bounded to the drone capacity, and that's deliberate. The 2026-06-10
concurrency restructure (
docs/concurrency.md) made concurrent recipe runs correct — per-run recipe trees, app-domain locks, isolation — so two recipes no longer collide/corrupt. The remaining limit is host RESOURCES (memory), which is exactly whatDRONE_RUNNER_CAPACITY(=2) is tuned to. RunningCAPsubagents matches the runner's slots; do NOT exceed it on the shared box (that's what--parallelis for, and only when the box is idle). Eachrecipe-upgradestill tears down its own step-2b deploy; the end-of-run reap (§4b) clears any that leaked. - If the host is ALSO running the build loops, prefer
--sequential(CAP=1) — the loops and a capacity-2 upgrade together can oversubscribe the 7.6 GB box (a Swarm task can wedge inNewstate under load). The orchestrator should serialize: don't run a capacity>1 upgrade concurrently with active phase loops. - Single-writer: every PR (recipe or cc-ci test) is on a dedicated branch; never push
main, never touch the build loops'/cc-ci/cc-ci-advworking clones or their in-flight state. - Contention with active loop development: while the loops are still building cc-ci, this run competes for the host. Prefer a quiescent window; if a recipe fails due to contention it's simply retried next week. (Once cc-ci is built and the loops are idle, this is the steady-state weekly job.)
- Never merges; failures/ skips are surfaced and retried next week — safe to re-run anytime.
Cron
Designed for a weekly Claude Code scheduled task that runs cc-ci-plan/launch-upgrader.sh start —
which spins up the cc-ci-upgrader remote-control agent to run this skill to completion (the agent
then stays idle/viewable). The cron does NOT invoke /upgrade-all inline.
Schedule (operator, 2026-05-29): the weekly cron is installed as the closing action of the final
build phase (Phase 5), and its first run is ~1 hour after the build completes, weekly from then on
— i.e. anchored to actual completion, not a fixed clock slot. Compute T0 = completion + 1h and
install a weekly job at T0's day-of-week + HH:MM (MM HH * * DOW) running launch-upgrader.sh start. Then verify the first kickoff actually launches the cc-ci-upgrader agent. Do NOT install
it earlier — while the loops are still building cc-ci it would contend for the shared host; until then
run it manually via launch-upgrader.sh start/fresh. Full procedure: Phase 5 plan §4
(cc-ci-plan/plan-phase5-verify-upgrade-flow.md).
Re-running is idempotent: already-current recipes report SKIPPED — up-to-date; recipes with an open
PR for the same branch report the existing PR rather than duplicating it.