Files
cc-ci-orchestrator/.claude/skills/ci-test-review/SKILL.md
T
autonomic-bot bb7ebb4a27 reconcile-upstream.sh: one deterministic entry point, mandated before PR work
Working against a stale mirror has cost us three different ways:

- mailu #6 was linked as the fix for two internet-facing Roundcube CVEs while
  upstream had already merged AND released it (3.1.3+2024.06.57). The work was
  done; only our mirror was behind. Reconciling closed the PR automatically.
- a stale mirror makes a survey report 'no upgrades available', so the recipe
  silently drops out of the weekly run.
- reading the wrong branch: several coopcloud recipes keep a stale 'main' beside
  the real default 'master'. gitea's main is 1.24.2-rootless while master has
  1.27.1-rootless and the merged PRs, so reading main manufactures a false
  'three releases behind, missing two CVSS-9.8 RCEs' finding.

The reconcile logic already existed inside open-recipe-pr.sh --reconcile-only and
already resolves the default branch itself. What was missing was a single obvious
entry point and a rule saying to run it. reconcile-upstream.sh takes recipes or
--all, and is idempotent — recipe work lives in branches, never on mirror main, so
force-syncing main discards nothing.

/ci-test-review and /cc-ci-tests-update had NO reconcile step at all; both now
require it. /cve-check, /recipe-upgrade and /upgrade-all already reconciled and now
point at the shared script.
2026-08-11 18:38:09 +00:00

9.5 KiB

name, description
name description
ci-test-review On-demand AI review + auto-fix of the cc-ci server. Runs the deterministic recipe test suite across ALL enrolled recipes; for any failure it diagnoses the root cause, classifies it as a RECIPE bug or a CI-SERVER bug, AUTHORS the fix, OPENS a PR (recipe repo or cc-ci repo), and VERIFIES that PR green on the CI server (deploy + full suite against the PR head, dogfood) — but NEVER merges (the operator merges). If everything is green and current it reports "all passed, recipes + tests up to date". Use for a full health check that ends in verified, ready-to-merge fix PRs. Invoke as /ci-test-review.

ci-test-review

An on-demand, AI-driven review + auto-fix layer over the cc-ci recipe CI server. cc-ci's primary use is deterministic testing (normal !testme / nightly — no AI). This skill runs the full suite across all recipes, then for any failure: diagnoses the root cause, classifies it (recipe vs CI-server), authors the fix, opens a PR, and verifies that PR green on the CI server — the same deploy-and-test dogfood the build loops use. It leaves a verified, ready-to-merge PR; it does NOT merge. If everything's green and current it reports "all passed".

This skill drives its own CI server (cc-ci) the way the loops do: it can push branches, deploy recipes, and run the harness there to verify a fix end-to-end.

Hard boundaries:

  • The test execution and PR verification are deterministic (the cc-ci harness decides pass/fail). AI never decides pass/fail — it only diagnoses, authors the fix, and drives the flow.
  • Create + verify, never merge. The operator reviews the cc-ci-green PR and merges.
  • Never weaken a test to make a PR go green.

Access (this skill runs from the orchestration repo, drives cc-ci over the tailnet)

  • cc-ci is reachable via ssh cc-ci (root, through the userspace-tailscaled SOCKS proxy on 127.0.0.1:1055). If ssh cc-ci fails, restart the proxy (systemctl restart cc-ci-tailscaled) per the orchestrator's access notes, then retry.
  • The cc-ci repo + harness live at /root/cc-ci on the cc-ci VM; the runner is cc-ci-run runner/run_recipe_ci.py with RECIPE=<name> (+ optional STAGES=...).

Procedure

1. Run the deterministic sweep (no AI)

Run the helper, which enumerates enrolled recipes and runs the full suite per recipe on cc-ci, emitting one structured result line per recipe (per-tier pass/fail/skip + the tested vs latest published version):

bash "$(dirname "$0")/run-all-recipes.sh"            # or: .claude/skills/ci-test-review/run-all-recipes.sh

It writes a results table to stdout and a machine-readable summary to /srv/cc-ci/.cc-ci-logs/ci-test-review-<runid>.json. Do NOT hand-judge pass/fail — trust the harness's per-tier results.

2. If everything is green AND current → report all-clear

If every recipe's tiers are pass/skip(N/A) (no fail) and every recipe was tested at its latest published version with no stale test flagged, output:

ALL PASSED — recipes up to date, tests up to date. <N recipes, versions listed>

Stop here. Nothing to fix.

3. For each failure → diagnose root cause + CLASSIFY (AI)

For every fail (after confirming it's not a flake — see below), read the run log, the recipe's compose/code (on cc-ci ~/.abra/recipes/<recipe>), and the harness, then determine the root cause and classify it as exactly one of:

  • RECIPE bug → recipe-side fix. The recipe itself is broken/fragile (bad healthcheck, missing backup hook, startup race, etc.). The fix is a change to the recipe (e.g. "add a collabora WOPI healthcheck + start_period", "add a pg_dump backup hook").
  • CI-SERVER bug → cc-ci-side fix. The harness/test/infra is wrong (a test asserts the wrong thing, a readiness gate is off, an infra config bug). The fix is a change to the cc-ci repo.
  • TEST out-of-date → cc-ci-side fix. The recipe legitimately changed upstream and the cc-ci test/overlay needs updating to match.
  • FLAKY / infra-transient. Distinguish from a real failure: re-run the recipe once or twice; if it then passes, classify as flaky (note it; suggest a robustness fix if there's a pattern) — do NOT author a recipe/CI "fix" for a flake.

Freshness is part of this: compare each recipe's tested version to the latest published version (the sweep records both). A recipe tested at a stale pin, or a cc-ci test lagging an upstream recipe change, is itself a finding — "all passed" must mean "passes against current upstream".

4. Author the fix + OPEN a PR (AI authors; never merge)

For each real (non-flaky) finding, write the actual fix and open a PR. Never merge.

  • Recipe-side fix → recipe PR. Branch the recipe, apply the fix, and open the PR via the recipe-create-pr skill (/srv/recipe-maintainer/.opencode/skills/recipe-create-pr/SKILL.md) — it handles the mirror to git.autonomic.zone/recipe-maintainers/<recipe> (upstream git.coopcloud.tech). Keep the change bounded to the diagnosed root cause; don't rewrite the recipe.
  • Before editing any test, read tests/STYLE.md in the cc-ci repo. It encodes the rules a test change must satisfy — set state up through the app's interface rather than its database, gate on version instead of branching, correct the fixture/wait but NEVER the assertion, and diagnose from the app's own telemetry before concluding a test is stale.
  • CI-server-side fix → cc-ci PR. Branch the cc-ci product repo (recipe-maintainers/cc-ci), apply the fix, and open the PR via the Gitea API (use the GITEA_* creds from /srv/cc-ci/.testenv). Single-writer discipline: work on a dedicated branch in a SEPARATE clone — never push main, never touch the build loops' working clones (/cc-ci, /cc-ci-adv) or their in-flight state.

⚠️ RECONCILE FROM UPSTREAM FIRST — always, before any PR work or upgrade check

cc-ci-plan/reconcile-upstream.sh <recipe>...   # or --all

Deterministic, idempotent, and safe (recipe work lives in branches, never on mirror main). It force-syncs each mirror to coopcloud's default branch — resolved from the API, main OR master — and closes any mirror PR whose changes upstream already merged. Skipping it has cost us three distinct ways: mailu #6 was reported as the fix for two internet-facing CVEs while upstream had already merged AND released it; a stale mirror makes a survey report "no upgrades available" so the recipe drops out of the weekly run; and reading the wrong branch on a recipe with a stale main beside a live master (gitea) manufactures a false "three releases behind, missing two CVSS-9.8 RCEs" finding.

5. VERIFY each PR on the CI server (deterministic; still never merge)

A PR is only "working" once cc-ci verifies it green (operator rule) — dogfood the CI that found the bug. Verification is deterministic (the harness), not an AI judgement.

  • Recipe PR: run the full suite COLD against the PR head on cc-ci — the same path !testme uses — via the helper:
    RECIPE=<recipe> REF=<pr-head-branch-or-sha> .claude/skills/ci-test-review/verify-pr.sh
    
    Green ⇔ the harness exits 0. General bar = one cold green. Use REPEAT=3 (repeated-green) only for a recipe already known to be FLAKY (e.g. lasuite-drive) as a flakiness proof — this is NOT the general standard.
  • CI-server PR: check the branch out in a clone on cc-ci, rebuild if the change requires it (nixos-rebuild / flake), then re-run the previously-failing recipe(s) through the harness plus a small regression sample of unrelated recipes. Green ⇔ the fix passes AND nothing regressed.
  • If verification is RED: iterate the fix on the same branch (bounded — ≤3 attempts) and re-verify. If still red, leave the PR open and report it as "opened, NOT yet green — needs work" with the failing evidence. Never abandon silently; never weaken a test to force green.

6. Output a structured report

  • A per-recipe status table (recipe | tiers | version tested / latest | verdict).
  • For each finding: root cause → classification (recipe / CI-server) → PR link → verification verdict (GREEN — cold-verified on cc-ci, ready for operator merge / RED — opened, needs work).
  • Or the all-clear summary from step 2.
  • Always end with the explicit reminder that nothing was merged — the PRs await operator review.

Guardrails

  • Deterministic execution + verification stay AI-free — the harness decides pass/fail; AI only diagnoses, authors the fix, and drives the flow.
  • Create + verify, NEVER merge. Leave a cc-ci-green, ready-to-merge PR; the operator merges (recipe PRs especially are operator-merged only after cc-ci verifies them green).
  • Don't weaken any test to make a PR go green — the fix must make the recipe/CI genuinely correct.
  • Flake ≠ bug — re-run before authoring any fix.
  • Real abra path throughout (no docker-level bypass).
  • Bounded fixes — each PR targets one diagnosed root cause; not a rewrite.
  • Coordination / single-writer: while the build loops are actively developing cc-ci, the server and the recipe-maintainers/cc-ci repo are in use. Always use a dedicated branch + separate clone; never push main or disturb the loops' clones. Recipe deploys are stateful on the shared Swarm, so run verification when you have effectively-exclusive use of the host (or serialize), and tear down whatever you deploy.