Working against a stale mirror has cost us three different ways: - mailu #6 was linked as the fix for two internet-facing Roundcube CVEs while upstream had already merged AND released it (3.1.3+2024.06.57). The work was done; only our mirror was behind. Reconciling closed the PR automatically. - a stale mirror makes a survey report 'no upgrades available', so the recipe silently drops out of the weekly run. - reading the wrong branch: several coopcloud recipes keep a stale 'main' beside the real default 'master'. gitea's main is 1.24.2-rootless while master has 1.27.1-rootless and the merged PRs, so reading main manufactures a false 'three releases behind, missing two CVSS-9.8 RCEs' finding. The reconcile logic already existed inside open-recipe-pr.sh --reconcile-only and already resolves the default branch itself. What was missing was a single obvious entry point and a rule saying to run it. reconcile-upstream.sh takes recipes or --all, and is idempotent — recipe work lives in branches, never on mirror main, so force-syncing main discards nothing. /ci-test-review and /cc-ci-tests-update had NO reconcile step at all; both now require it. /cve-check, /recipe-upgrade and /upgrade-all already reconciled and now point at the shared script.
9.5 KiB
name, description
| name | description |
|---|---|
| ci-test-review | On-demand AI review + auto-fix of the cc-ci server. Runs the deterministic recipe test suite across ALL enrolled recipes; for any failure it diagnoses the root cause, classifies it as a RECIPE bug or a CI-SERVER bug, AUTHORS the fix, OPENS a PR (recipe repo or cc-ci repo), and VERIFIES that PR green on the CI server (deploy + full suite against the PR head, dogfood) — but NEVER merges (the operator merges). If everything is green and current it reports "all passed, recipes + tests up to date". Use for a full health check that ends in verified, ready-to-merge fix PRs. Invoke as /ci-test-review. |
ci-test-review
An on-demand, AI-driven review + auto-fix layer over the cc-ci recipe CI server. cc-ci's primary
use is deterministic testing (normal !testme / nightly — no AI). This skill runs the full suite
across all recipes, then for any failure: diagnoses the root cause, classifies it (recipe
vs CI-server), authors the fix, opens a PR, and verifies that PR green on the CI server — the same
deploy-and-test dogfood the build loops use. It leaves a verified, ready-to-merge PR; it does
NOT merge. If everything's green and current it reports "all passed".
This skill drives its own CI server (cc-ci) the way the loops do: it can push branches, deploy recipes, and run the harness there to verify a fix end-to-end.
Hard boundaries:
- The test execution and PR verification are deterministic (the cc-ci harness decides pass/fail). AI never decides pass/fail — it only diagnoses, authors the fix, and drives the flow.
- Create + verify, never merge. The operator reviews the cc-ci-green PR and merges.
- Never weaken a test to make a PR go green.
Access (this skill runs from the orchestration repo, drives cc-ci over the tailnet)
- cc-ci is reachable via
ssh cc-ci(root, through the userspace-tailscaled SOCKS proxy on 127.0.0.1:1055). Ifssh cc-cifails, restart the proxy (systemctl restart cc-ci-tailscaled) per the orchestrator's access notes, then retry. - The cc-ci repo + harness live at
/root/cc-cion the cc-ci VM; the runner iscc-ci-run runner/run_recipe_ci.pywithRECIPE=<name>(+ optionalSTAGES=...).
Procedure
1. Run the deterministic sweep (no AI)
Run the helper, which enumerates enrolled recipes and runs the full suite per recipe on cc-ci, emitting one structured result line per recipe (per-tier pass/fail/skip + the tested vs latest published version):
bash "$(dirname "$0")/run-all-recipes.sh" # or: .claude/skills/ci-test-review/run-all-recipes.sh
It writes a results table to stdout and a machine-readable summary to
/srv/cc-ci/.cc-ci-logs/ci-test-review-<runid>.json. Do NOT hand-judge pass/fail — trust the
harness's per-tier results.
2. If everything is green AND current → report all-clear
If every recipe's tiers are pass/skip(N/A) (no fail) and every recipe was tested at its
latest published version with no stale test flagged, output:
ALL PASSED — recipes up to date, tests up to date. <N recipes, versions listed>
Stop here. Nothing to fix.
3. For each failure → diagnose root cause + CLASSIFY (AI)
For every fail (after confirming it's not a flake — see below), read the run log, the recipe's
compose/code (on cc-ci ~/.abra/recipes/<recipe>), and the harness, then determine the root cause
and classify it as exactly one of:
- RECIPE bug → recipe-side fix. The recipe itself is broken/fragile (bad healthcheck, missing
backup hook, startup race, etc.). The fix is a change to the recipe (e.g. "add a collabora WOPI
healthcheck +
start_period", "add apg_dumpbackup hook"). - CI-SERVER bug → cc-ci-side fix. The harness/test/infra is wrong (a test asserts the wrong thing, a readiness gate is off, an infra config bug). The fix is a change to the cc-ci repo.
- TEST out-of-date → cc-ci-side fix. The recipe legitimately changed upstream and the cc-ci test/overlay needs updating to match.
- FLAKY / infra-transient. Distinguish from a real failure: re-run the recipe once or twice; if it then passes, classify as flaky (note it; suggest a robustness fix if there's a pattern) — do NOT author a recipe/CI "fix" for a flake.
Freshness is part of this: compare each recipe's tested version to the latest published version (the sweep records both). A recipe tested at a stale pin, or a cc-ci test lagging an upstream recipe change, is itself a finding — "all passed" must mean "passes against current upstream".
4. Author the fix + OPEN a PR (AI authors; never merge)
For each real (non-flaky) finding, write the actual fix and open a PR. Never merge.
- Recipe-side fix → recipe PR. Branch the recipe, apply the fix, and open the PR via the
recipe-create-prskill (/srv/recipe-maintainer/.opencode/skills/recipe-create-pr/SKILL.md) — it handles the mirror togit.autonomic.zone/recipe-maintainers/<recipe>(upstreamgit.coopcloud.tech). Keep the change bounded to the diagnosed root cause; don't rewrite the recipe. - Before editing any test, read
tests/STYLE.mdin the cc-ci repo. It encodes the rules a test change must satisfy — set state up through the app's interface rather than its database, gate on version instead of branching, correct the fixture/wait but NEVER the assertion, and diagnose from the app's own telemetry before concluding a test is stale. - CI-server-side fix → cc-ci PR. Branch the cc-ci product repo
(
recipe-maintainers/cc-ci), apply the fix, and open the PR via the Gitea API (use theGITEA_*creds from/srv/cc-ci/.testenv). Single-writer discipline: work on a dedicated branch in a SEPARATE clone — never pushmain, never touch the build loops' working clones (/cc-ci,/cc-ci-adv) or their in-flight state.
⚠️ RECONCILE FROM UPSTREAM FIRST — always, before any PR work or upgrade check
cc-ci-plan/reconcile-upstream.sh <recipe>... # or --allDeterministic, idempotent, and safe (recipe work lives in branches, never on mirror
main). It force-syncs each mirror to coopcloud's default branch — resolved from the API,mainORmaster— and closes any mirror PR whose changes upstream already merged. Skipping it has cost us three distinct ways: mailu #6 was reported as the fix for two internet-facing CVEs while upstream had already merged AND released it; a stale mirror makes a survey report "no upgrades available" so the recipe drops out of the weekly run; and reading the wrong branch on a recipe with a stalemainbeside a livemaster(gitea) manufactures a false "three releases behind, missing two CVSS-9.8 RCEs" finding.
5. VERIFY each PR on the CI server (deterministic; still never merge)
A PR is only "working" once cc-ci verifies it green (operator rule) — dogfood the CI that found the bug. Verification is deterministic (the harness), not an AI judgement.
- Recipe PR: run the full suite COLD against the PR head on cc-ci — the same path
!testmeuses — via the helper:Green ⇔ the harness exits 0. General bar = one cold green. UseRECIPE=<recipe> REF=<pr-head-branch-or-sha> .claude/skills/ci-test-review/verify-pr.shREPEAT=3(repeated-green) only for a recipe already known to be FLAKY (e.g. lasuite-drive) as a flakiness proof — this is NOT the general standard. - CI-server PR: check the branch out in a clone on cc-ci, rebuild if the change requires it
(
nixos-rebuild/ flake), then re-run the previously-failing recipe(s) through the harness plus a small regression sample of unrelated recipes. Green ⇔ the fix passes AND nothing regressed. - If verification is RED: iterate the fix on the same branch (bounded — ≤3 attempts) and re-verify. If still red, leave the PR open and report it as "opened, NOT yet green — needs work" with the failing evidence. Never abandon silently; never weaken a test to force green.
6. Output a structured report
- A per-recipe status table (recipe | tiers | version tested / latest | verdict).
- For each finding: root cause → classification (recipe / CI-server) → PR link → verification
verdict (
GREEN — cold-verified on cc-ci, ready for operator merge/RED — opened, needs work). - Or the all-clear summary from step 2.
- Always end with the explicit reminder that nothing was merged — the PRs await operator review.
Guardrails
- Deterministic execution + verification stay AI-free — the harness decides pass/fail; AI only diagnoses, authors the fix, and drives the flow.
- Create + verify, NEVER merge. Leave a cc-ci-green, ready-to-merge PR; the operator merges (recipe PRs especially are operator-merged only after cc-ci verifies them green).
- Don't weaken any test to make a PR go green — the fix must make the recipe/CI genuinely correct.
- Flake ≠ bug — re-run before authoring any fix.
- Real abra path throughout (no docker-level bypass).
- Bounded fixes — each PR targets one diagnosed root cause; not a rewrite.
- Coordination / single-writer: while the build loops are actively developing cc-ci, the server
and the
recipe-maintainers/cc-cirepo are in use. Always use a dedicated branch + separate clone; never pushmainor disturb the loops' clones. Recipe deploys are stateful on the shared Swarm, so run verification when you have effectively-exclusive use of the host (or serialize), and tear down whatever you deploy.