Compare commits

..
20 changed files with 48 additions and 2442 deletions
-97
View File
@@ -1,97 +0,0 @@
---
name: cc-ci-cleanup
description: Tidy the fleet's open recipe PRs. Reconciles every mirror from TRUE upstream first (which alone closes PRs upstream already merged), then surveys every open PR deterministically, CLOSES the ones that can no longer be merged or were never meant to be (CI sweep artifacts, obsolete bumps, superseded duplicates) with a reason, and reports prioritised action items for the ones that SHOULD merge — what specifically is blocking each. NEVER merges a recipe PR. Invoke as /cc-ci-cleanup [recipe ...] [--dry-run].
---
# cc-ci-cleanup
Open recipe PRs accumulate and rot. Some were never meant to merge (CI sweep artifacts), some were
overtaken (upstream merged the same change, or a newer PR supersedes them), and some genuinely should
land but are quietly blocked. Left alone the list becomes noise, and a real CVE fix hides in it.
This skill separates those three, acts on the first two, and hands you a short list for the third.
**Boundaries.** It **CLOSES** irrelevant PRs and **NEVER MERGES** any recipe PR — those change what
deploys on other people's infrastructure, so a human merges them (see AGENTS.md). Closing is the only
write it performs, always with a comment saying why.
## Arguments
- `<recipe> …` — limit to these recipes (else every recipe in `cc-ci-plan/used-recipes.md`).
- `--dry-run` — classify and report, close nothing.
## Procedure
### 1. Reconcile every mirror from TRUE upstream — MANDATORY, FIRST
```
cc-ci-plan/reconcile-upstream.sh --all # or: reconcile-upstream.sh <recipe>...
```
**Do not skip this and do not reorder it.** Every signal in step 2 is measured against the mirror's
`main`; against a stale mirror they are all wrong. This step also does a chunk of the cleanup by
itself — it closes any PR whose changes upstream has already merged.
> On the first real run (2026-08-11) this alone closed **three** PRs that looked pending and were
> already merged upstream: discourse #6 (carrying **140 CVEs**), keycloak #6 (**12 CVEs**), n8n #5.
> All three had been reported to the operator as outstanding work. mailu #6 went the same way earlier
> the same day. Reconciling is not hygiene, it is how you avoid recommending work that is already done.
### 2. Survey every open PR (deterministic — no judgement yet)
```
python3 cc-ci-plan/pr-survey.py [recipe ...] # add --json for the raw facts
```
Per PR it measures: `behind_main`, `ahead`, `mergeable`, `diff_files`, the images it **adds**, which
of those are **already in main**, `obsolete`, the newest `!testme` verdict + build, `branch_kind`,
and age/idle days. It decides nothing — that is this skill's job.
### 3. Classify
**CLOSE — cannot merge, or was never meant to.** Each needs a *positive* reason, not an absence:
| signal | why it is closeable |
|---|---|
| `branch_kind: ci-artifact` (`ci/*`) | regall/cfold sweeps and `!testme` probes — harness artifacts, never intended to merge |
| `obsolete: true` | every image it adds is **already pinned in main** — it has nothing left to contribute |
| superseded | a newer PR on the same recipe makes the same bump (name both numbers in the comment) |
| `diff_files: 0` | genuinely empty diff — nothing to merge |
**NEVER close on:**
- `DIFF-UNREADABLE` — the diff could not be fetched, which is NOT an empty diff. gitea #4 reads that
way (force-pushed branch) while being a verified, green, needed fix.
- any field that came back `null`/unknown.
- a PR that carries a **CVE fix** and is the only thing carrying it, even if it looks stale — report it
instead. Losing a security fix to tidiness is far worse than a long PR list.
- `--dry-run`.
**NEEDS WORK — should merge, something blocks it.** Give the *specific* next action:
| signal | action item |
|---|---|
| `mergeable: false` | conflicts — rebase the branch on `main` and re-run `!testme` |
| `behind_main > 0` | out of date — rebase, then re-verify (a green from before main moved proves nothing) |
| `ci: failed` | diagnose via `/ci-test-review`; classify recipe-bug vs stale test |
| `ci: never-run` | run `!testme` |
| blocked on the operator | say exactly what is needed (a secret, an upstream release, a decision) |
**READY — green, current, no conflicts.** Action item is simply: review and merge.
### 4. Close the CLOSE set (skip entirely under `--dry-run`)
Comment first, then close. The comment must say **which signal** made it closeable and **what to do
if that is wrong** ("reopen if …"), so a wrong call is cheap to undo. Never close silently.
### 5. Report
Order by what deserves attention, not by recipe name:
1. **CVE-carrying PRs that should merge** — most severe first, with the CVE ids.
2. Other **READY** PRs (green + current).
3. **NEEDS WORK**, each with its one specific action.
4. **CLOSED this run**, with the reason for each.
5. Anything **deliberately left alone** despite looking stale, and why.
End with a one-line summary: `N open → C closed, R ready to merge, W need work`.
## Guardrails
- **Never merge a recipe PR.** Create/verify/close only; the operator merges.
- **Reconcile first, always.** Judging a PR against a stale mirror is how you close good work or
recommend work that is already done.
- **Close only on a positive signal**, never on "looks old". Age alone is not a reason — several
60-day-old PRs here are green and mergeable.
- **Never close a lone CVE fix.** Report it, however stale.
- Every close gets a comment with its reason and a reopen hint.
-17
View File
@@ -79,29 +79,12 @@ For each real (non-flaky) finding, write the actual fix and open a PR. **Never m
it handles the mirror to `git.autonomic.zone/recipe-maintainers/<recipe>` (upstream
`git.coopcloud.tech`). Keep the change **bounded** to the diagnosed root cause; don't rewrite the
recipe.
- **Before editing any test, read `tests/STYLE.md` in the cc-ci repo.** It encodes the rules a test
change must satisfy — set state up through the app's interface rather than its database, gate on
version instead of branching, correct the fixture/wait but NEVER the assertion, and diagnose from
the app's own telemetry before concluding a test is stale.
- **CI-server-side fix → cc-ci PR.** Branch the cc-ci product repo
(`recipe-maintainers/cc-ci`), apply the fix, and open the PR via the Gitea API (use the
`GITEA_*` creds from `/srv/cc-ci/.testenv`). **Single-writer discipline:** work on a dedicated
branch in a SEPARATE clone — **never push `main`, never touch the build loops' working clones**
(`/cc-ci`, `/cc-ci-adv`) or their in-flight state.
> ### ⚠️ RECONCILE FROM UPSTREAM FIRST — always, before any PR work or upgrade check
> ```
> cc-ci-plan/reconcile-upstream.sh <recipe>... # or --all
> ```
> Deterministic, idempotent, and safe (recipe work lives in branches, never on mirror `main`). It
> force-syncs each mirror to coopcloud's **default branch — resolved from the API, `main` OR
> `master`** — and closes any mirror PR whose changes upstream already merged. Skipping it has cost
> us three distinct ways: mailu #6 was reported as the fix for two internet-facing CVEs while
> upstream had already merged AND released it; a stale mirror makes a survey report "no upgrades
> available" so the recipe drops out of the weekly run; and reading the wrong branch on a recipe with
> a stale `main` beside a live `master` (gitea) manufactures a false "three releases behind, missing
> two CVSS-9.8 RCEs" finding.
### 5. VERIFY each PR on the CI server (deterministic; still never merge)
A PR is only "working" once **cc-ci verifies it green** (operator rule) — dogfood the CI that found
the bug. Verification is deterministic (the harness), not an AI judgement.
@@ -1,91 +0,0 @@
---
name: cve-check-and-upgrade
description: Security-driven upgrade run. Does a full /cve-check sweep first (per-image advisory scan of every recipe's available upgrade, with adjudication), then runs /recipe-upgrade ONLY on the recipes whose upgrade fixes at least one CVE — worst severity first — opening a verified recipe PR for each, and finally publishes one report covering both the sweep and the PRs. Recipes with no CVEs are left alone; that is the point. NEVER merges. Invoke as /cve-check-and-upgrade [recipe ...] [--min-severity high] [--capacity N] [--dry-run].
---
# cve-check-and-upgrade
`/upgrade-all` upgrades everything that *has* an upgrade. **This upgrades what has a reason.** It runs
the `/cve-check` sweep, then spends CI time only on the recipes where an upgrade actually closes a
vulnerability, handling the worst first.
Use it when you want to act on security rather than churn the whole fleet: after a vendor announcement,
when CI capacity is short, or between weekly runs. When you only want to *know*, use `/cve-check`. When
you want everything current regardless of CVEs, use `/upgrade-all`.
**Creates PRs. Never merges.** Every PR is verified green on cc-ci and left for a human.
## Arguments
- `<recipe> …` — restrict the whole run to these recipes.
- `--min-severity critical|high|medium|low` — only upgrade recipes whose fixed CVEs reach this
severity. Default **`low`** (any CVE at all justifies the upgrade). `--min-severity high` is the
useful "just the urgent ones" setting.
- `--capacity N` — subagent pool size; defaults to the live `DRONE_RUNNER_CAPACITY` (the drone
runner's slots), matching `/upgrade-all`'s rolling-pool behaviour.
- `--dry-run` — do the whole sweep and print exactly which recipes *would* be upgraded and why, then
stop without spawning a single upgrade. **Publishes no report and opens no PR.**
## Procedure
### 1. Sweep — run `/cve-check` in full
Follow `.claude/skills/cve-check/SKILL.md` steps 15 exactly: candidate list, per-image upgrade windows,
the advisory scan per recipe, pass-2 adjudication of anything undecided, and the severity classification.
**Do not publish its report** — this run produces one combined report at the end instead.
Keep, per recipe: the windows scanned, the CVE count, the CVE ids with severities, and whether the count
is a floor (undetermined advisories remain) or unknown.
### 2. Decide what to upgrade
`RECIPES_TO_UPGRADE` = recipes where **the scan found ≥1 CVE** at or above `--min-severity`.
Deliberate exclusions, each recorded in the report with its reason:
- **0 CVEs** — an upgrade may exist, but nothing security-relevant. Left alone; that is the point of
this skill. `/upgrade-all` is what sweeps those up.
- **`external` tier** — swept for visibility, **never upgraded here**; someone else maintains it. Flag
it loudly in the report if it has a critical, since the action is to tell them, not to open a PR.
- **`UPTODATE`** — nothing available.
- **count `?` / UNKNOWN** — do **not** upgrade blind, and do **not** treat it as clean. Put it in the
Addendum as needing a look. An unknown is a gap in our knowledge, not evidence of safety.
Order the queue by **worst severity first** (critical → high → …), count breaking ties. If `--dry-run`,
print this queue with each recipe's CVE ids and severities, and STOP here.
### 3. Upgrade each one — via `/recipe-upgrade` subagents
Run `/recipe-upgrade <recipe>` per queued recipe as a **subagent**, in the queue order above, as a
**rolling pool** keeping `--capacity` (default `DRONE_RUNNER_CAPACITY`) running at once and starting the
next as each finishes — the same concurrency discipline as `/upgrade-all` §3, and safe for the same
reason (per-run recipe trees + app-domain locks).
Each subagent does the full job: plan, implement the bump, verify green on cc-ci with `!testme`, and
open a recipe PR. **Default mode — no `--with-tests`**: a genuinely stale test gets an explanatory PR
comment, not a test edit.
**Tell each subagent which CVEs justify its upgrade**, with ids and severities, so the PR description
says why it exists. That is most of this skill's value to a reviewer: a PR that names the CVSS-9.8 RCE
it closes gets merged today, an unexplained version bump waits a fortnight.
Collect per recipe: PR url + number, the `!testme` verdict and build number, and any failure.
### 4. Report — one page covering sweep AND PRs
Write `/tmp/cve-spec-<DATE>.json` per `/cve-check` step 6, with these differences:
- Rows for upgraded recipes carry the real `ci` (`build N ✓` / `RED N · <stage>`) + `ci_url`, and
`pr`/`pr_url`. `status` is the CI verdict (`GREEN`/`FAILED`/`STALE`); the live PR-status column
derives itself from `recipe` + `pr`.
- Rows for swept-but-not-upgraded recipes keep `PENDING`/`UPTODATE` with empty `ci`/`pr`, and a
`notes` reason (`0 CVEs — not upgraded`, `external — maintained elsewhere`, `below --min-severity`).
- Include `changes[]` — one entry per recipe that got a PR, describing what the upgrade changes **and
the CVEs it closes**.
- Keep `"kind": "cve"`: it titles the page "The Recipe Report — CVE check" and files it as
`cve-<DATE>.html`, alongside the weekly editions in the same archive index.
Then render + publish exactly as `/cve-check` step 7, and verify as its step 8. Print the report URL,
`N swept · M upgraded · K PRs green · J failed`, and `CVE CHECK AND UPGRADE COMPLETE`.
## Guardrails
- **NEVER merge.** Create and verify; a human merges. Never push to true upstream.
- **Never weaken a test** to make a PR green, and never edit a test without `--with-tests`.
- **Never upgrade a recipe whose CVE count is unknown** on the assumption it is fine — surface it.
- **Never upgrade an `external` recipe** here, even with a critical; report it instead.
- **Public-safe report only** — no secrets, tokens, internal hostnames, raw logs, or spend figures.
- If the sweep finds **nothing** at or above `--min-severity`, that is a good outcome: publish the
report saying so and open no PRs. Do not manufacture work.
-204
View File
@@ -1,204 +0,0 @@
---
name: cve-check
description: Fleet-wide CVE sweep WITHOUT upgrading anything. For every recipe cc-ci deploys, works out what upgrade is available (current pinned tag → newest supported tag, per image including sidecars), runs the deterministic advisory scan over that window, adjudicates whatever the scan could not decide, and publishes a CVE report to report.ci.commoninternet.net as cve-<DATE>.html. READ-ONLY — opens no PRs, edits no recipes, runs no CI, merges nothing. Answers "what are we exposed to that an upgrade would fix?" in minutes rather than the hours a full upgrade run takes. Invoke as /cve-check [recipe ...] [--weekly-only].
---
# cve-check
A **security sweep, not an upgrade run.** It answers one question for every recipe cc-ci deploys:
> If we upgraded this recipe today, how many CVEs would that fix, and how bad are they?
It is the cheap, safe half of `/upgrade-all`: the same version research and the same advisory scan,
with **no implementation, no CI, and no PRs**. Use it when you want the security picture now — after a
vendor announcement, before deciding what to prioritise, or between weekly runs. When you want the PRs
too, use **`/cve-check-and-upgrade`**.
**Read-only, absolutely.** Never edit a recipe, never open or comment on a PR, never merge, never
deploy. The only thing it writes is its own log and the published report page.
## Arguments
- `<recipe> …` — sweep only these recipes (else every recipe in `cc-ci-plan/used-recipes.md`).
- `--weekly-only` — skip rows tagged `external`. **Off by default on purpose**: an `external` recipe is
still deployed and still exposes us, so a security sweep that silently skipped it would misreport the
fleet's exposure. Externals are swept and clearly marked "maintained elsewhere" in the report.
## Procedure
> ### ⚠️ Run abra over a pseudo-TTY (or it FATAs `inappropriate ioctl for device`)
> `abra` needs a TTY. Wrap every abra call: `ssh cc-ci 'script -qec "abra <args> -n" /dev/null'`.
> (`git` and other commands do NOT need the wrapper.)
### 1. Build the candidate list
Read `cc-ci-plan/used-recipes.md` — the canonical inventory. Take every row (both tiers), recording the
tier per recipe; with `--weekly-only`, drop the `external` rows. An explicit recipe argument overrides
any skip.
### 2. Per recipe — establish the upgrade window WITHOUT upgrading
This is `/recipe-upgrade` step 1's research, stopping before it implements anything.
> ⚠️ **The same four things that silently skip recipes apply here — handle ALL FOUR:**
> 1. **pseudo-TTY** — per the box above.
> 2. **go-git auth to git.autonomic.zone** — recipes on the private mirror FATA
> `authentication required: Unauthorized`. Bake creds into origin first (idempotent, only when
> origin is on git.autonomic.zone):
> `git -C ~/.abra/recipes/<r> remote set-url origin "https://$GITEA_USERNAME:$GITEA_PASSWORD@git.autonomic.zone/recipe-maintainers/<r>.git"`
> 3. **dirty worktree** — usually just the untracked cc-ci overlay; `git stash -u` before, `stash pop`
> after. Only a genuinely dirty TRACKED tree is a skip.
> 4. **tag+digest pins abra cannot parse** — abra FATAs and aborts the WHOLE recipe (immich). Do not
> hand-check the registry; run the resolver, which is abra-independent and covers every image:
> ```
> python3 /srv/cc-ci/cc-ci-plan/resolve-images.py <recipe> --ssh cc-ci --table
> ```
> It reports, per image, `newest_within_major` (the compatibility-safe pick) and
> `newest_same_shape` (the newest of that tag's form). **Use `newest_within_major` unless you have
> checked the app supports the major jump** — immich's postgres tag encodes the pg major plus the
> vectorchord/pgvectors versions immich-server is built against, so taking the newest would break
> the deploy. `all_resolved: false` means an image could NOT be resolved — that is a `?`, never a 0.
**Reconcile the mirror from true upstream FIRST — ALWAYS, no exceptions** — one command,
`cc-ci-plan/reconcile-upstream.sh <recipe>... | --all`. This is the same reconcile
`/upgrade-all` does. Do not skip it in the name of keeping the sweep read-only: skipping it makes you
research a stale checkout, and on the first real run that produced **two recipes with no survey output
at all**, which is indistinguishable from "no upgrades" unless you check. It is safe — recipe work
lives in **branches**, never directly on `main`, so a force-sync of `main` to upstream discards
nothing; it also auto-closes mirror PRs whose changes upstream has already merged.
> ### ⚠️ The default branch may be `master`, not `main` — check, do not assume
> Several coopcloud recipes keep a **stale `main` alongside the real default `master`**. gitea is one:
> `main` sits at 1.24.2-rootless while `master` has 1.27.1-rootless plus the merged PRs and the 3.6.3
> release. Reading `main` there tells you the recipe is three releases behind and missing two CVSS-9.8
> RCE fixes — a false alarm that reads exactly like a real one. Resolve the default branch from the
> API (`/api/v1/repos/coop-cloud/<recipe>` → `default_branch`) before reading any file, and never
> `git reset --hard origin/main` on a checkout that tracks `master`.
**Cross-check abra with the resolver.** abra is the primary source, but it silently contributes
nothing for images it cannot parse, and it reported "no new versions" for images that did have them
(mumble v1.6.870-0 → -4). Run `resolve-images.py` for every recipe and take the UNION of the two: on
the first real sweep the resolver found upgrades abra missed entirely in five recipes, one of which
(plausible's clickhouse) carried four CVEs.
Then read versions:
```
set -a; . /srv/cc-ci/.testenv; set +a
ssh cc-ci "GITEA_USERNAME='$GITEA_USERNAME' GITEA_PASSWORD='$GITEA_PASSWORD' GITEA_URL='$GITEA_URL' bash -s <recipe> --reconcile-only" \
< /srv/cc-ci/.claude/skills/recipe-upgrade/open-recipe-pr.sh
ssh cc-ci 'export PATH=/run/current-system/sw/bin:$PATH; R=<recipe>; \
git -C ~/.abra/recipes/$R stash -u >/dev/null 2>&1 || true; \
script -qec "abra recipe fetch $R --force -n" /dev/null; \
script -qec "abra recipe upgrade $R -m -n" /dev/null; \
git -C ~/.abra/recipes/$R stash pop >/dev/null 2>&1 || true'
```
For each recipe produce **one window per image**: `current pinned tag → newest supported tag`. You need
the sidecars (redis, postgres, nginx …), not just the app — a sidecar bump is where discourse's only
CRITICAL came from, and an image with no window is not counted at all.
- **No upgrade available** → the recipe is `UPTODATE`; its CVE count is **`0`**, not `?`. There is
nothing an upgrade could fix. Record it and move on.
- **No output at all is NOT "no upgrade".** An abra call that times out, FATAs, or prints nothing
leaves the recipe **unverified** — treat it as a distinct outcome, never fold it into up-to-date.
Re-run it, and if it still yields nothing, resolve the versions by direct registry check (box item 4).
Only report `?` once BOTH the abra check and the direct check have failed. On the first real run this
distinction was the difference between two false zeros and the truth (both recipes turned out fine,
but nothing in the survey said so).
### 2c. Know which recipes CANNOT see CVEs at all
```
python3 cc-ci-plan/audit-sources.py --security-sources
```
A recipe whose sources yield **no CVE data at all** cannot produce a meaningful `0` — nothing was
measured, the same way a missing registry file cannot. Render those as **`?`**, not `0`.
**The fleet is currently at zero such recipes.** The last two — `mattermost-lts` (empty advisory
feed, client-side-rendered bulletins) and `mumble` (nothing published anywhere) — were fixed by
declaring an NVD CPE in their registry:
```
- nvd-cpe: mattermost-team-edition = cpe:2.3:a:mattermost:mattermost_server:*:*:*:*:*:*:*:*
```
**If this sweep ever reports a blind recipe again, that is the fix**: find the product's CPE at
nvd.nist.gov and add the line. Prefer a real advisory feed or an attributable changelog when one
exists — NVD lags the vendor — but a lagging source beats no source, and it turns a `?` into a
number.
An *unparseable page* is NOT the same thing: it is harmless when the same project also publishes an
advisory feed (redis, gitea, minio, clickhouse all do). Only "no usable source for this image" counts.
### 3. Run the advisory scan over that window
```
python3 /srv/cc-ci/cc-ci-plan/advisory-scan.py <recipe> --from <old-app> --to <new-app> \
[--image <name>=<old>:<new>]...
```
**One call per recipe with every image in it** — the count is a union across images, and the
UNKNOWN guarantee only holds when a single run sees them all. Paste the markdown block verbatim into
the per-recipe log at `/srv/cc-ci/.cc-ci-logs/cve-check/<DATE>/<recipe>.md`.
### 4. Adjudicate what the scan could not decide (pass 2)
If the block reports advisories it **could NOT judge**, or the count is **UNKNOWN**, re-run with
`--adjudicate` and decide each open case yourself:
```
python3 /srv/cc-ci/cc-ci-plan/advisory-scan.py <recipe> … --adjudicate
```
Answer **FIXED / NOT-FIXED / STILL-UNKNOWN** per case, each with a one-line reason **citing the
evidence shown** — never from memory of the project, which is the exact failure that let two CVSS-9.8
gitea RCEs be published as "none". Every FIXED is added to the count; pass 1's number is a floor. The
block also lists what pass 1 already decided — if a verdict looks wrong given its evidence, say so.
Record your verdicts in the per-recipe log so the number is auditable.
### 5. Classify severity and priority
For each recipe collect the CVE ids with **severities** (the scan gives them, with GHSA ids). Sort the
report rows by what an operator should deal with first:
1. recipes with a **critical**, then **high**, then anything else with CVEs (more CVEs higher within a band);
2. then `?` (a count that could not be established — investigate, do not ignore);
3. then recipes with an upgrade available but **0** CVEs;
4. then `UPTODATE`.
Severity outranks raw count: 2 CVSS-9.8 RCEs matter more than 120 medium plugin advisories.
### 6. Write the report spec
`/tmp/cve-spec-<DATE>.json`, same shape as `/recipe-report` (see `recipe-report.py`'s header), with:
- **`"kind": "cve"`** — titles the page "The Recipe Report — CVE check" and files it as
`cve-<DATE>.html`. It appears in the SAME archive index as the weekly editions, suffixed
"— CVE check" so the two are told apart at a glance. Without this field you would overwrite that
date's weekly edition.
- `date`, `subtitle` "CVE check <human date>",
- `lead`**one short paragraph**: fleet exposure in a sentence and what to do first.
- `table[]` — every recipe swept. `recipe`; `change` = the window you scanned, e.g.
`1.27.0 → 1.27.1 · redis 7.4 → 8.10`; `status` = `UPTODATE` when nothing is available, else
`PENDING` (an upgrade exists and is not yet taken); **`cve`** = the count (integer, `?` only per the
rules below); `notes` = severity mix, whether the number is a floor, and `maintained elsewhere` for
`external` rows. **Leave `ci`/`pr` empty — nothing was built and no PR exists.**
- `addendum[]` — real anomalies only: registry URLs that failed, recipes whose window could not be
established, a scan whose count is a floor with many undetermined advisories.
- `security[]` — one entry per **critical/high** finding: recipe · CVE id(s) + severity · what it fixes
· **which image** it is in. Name the image: `CVE-2025-49844` is a redis flaw, and an operator reading
"discourse" needs to know that.
- `changes[]`**omit** (nothing changed; there are no PRs).
**`?` must stay RARE.** Use it only when a scan ran and reported genuinely failed sources, or the count
came back UNKNOWN and adjudication could not settle it. Never `none` for an unknown — a blank reads as
clean. A recipe with no upgrade available is `0`, not `?`. Many `?` is a bug for the Addendum.
### 7. Render and publish — via the script only
```
python3 /srv/cc-ci/cc-ci-plan/recipe-report.py render /tmp/cve-spec-<DATE>.json /tmp/cve-<DATE>.html
python3 /srv/cc-ci/cc-ci-plan/recipe-report.py publish /tmp/cve-<DATE>.html <DATE> cve
```
All layout is owned by `recipe-report.py`. Never hand-write or post-process HTML; if `render` errors,
fix the spec JSON and re-render. **Public page — no secrets, tokens, internal hostnames, raw logs, or
any billing/spend figures.**
### 8. Verify and stop
`curl -fsS https://report.ci.commoninternet.net/cve-<DATE>.html` renders and the index lists it. Print
the URL, a one-line summary (`N recipes swept · M with CVEs · K critical`), and `CVE CHECK COMPLETE`,
then go idle. One-shot — do not loop, and do not start upgrading anything.
## Guardrails
- **Read-only.** No PRs, no edits, no merges, no deploys, no CI runs. If a recipe looks urgent, say so
in the report — do not act on it. `/cve-check-and-upgrade` is the skill that acts.
- **Never report `0` for something you could not scan.** `0` means checked-and-clean; unknown is `?`.
- A count with undetermined advisories is a **floor** — say so in the notes rather than rounding away.
- **Public-safe output only.**
-6
View File
@@ -312,12 +312,6 @@ test change, and a test change is **gated by `--with-tests`**:
Do **NOT** modify any test. Report `SUCCESS-PENDING-TESTS` (recipe PR open; `!testme` red on a
stale test; operator to decide).
- **`--with-tests` — open + verify a cc-ci test PR.** Make it the `ci-test-review` way:
0. **READ `tests/STYLE.md` in the cc-ci repo FIRST.** It is the rulebook for changing a test, and
it is written against the failures this pipeline has actually produced. The two that matter most
here: **set state up through the app's own interface, never its database** (a plausible fixture
that INSERTed rows passed on v2 and silently broke on v3, holding the recipe RED for six weeks),
and **gate on version rather than writing a fixture that supports both** — old-version tests can
simply be deleted, since the older version is only exercised through the upgrade tier.
1. Branch `recipe-maintainers/cc-ci` in a **separate clone** (single-writer: never push `main`,
never touch the build loops' `/cc-ci` `/cc-ci-adv` clones); update the test/overlay.
2. **Verify the recipe upgrade WITH the updated test applied.** `!testme` on the recipe PR uses the
@@ -73,29 +73,6 @@ done
# 5) Stray exited containers (debug one-shots) — best-effort prune.
docker container prune -f >/dev/null 2>&1 || true
# 6) Unused IMAGES — the one that actually took CI down. Every run pulls each recipe's images and
# nothing ever removed the old ones: on 2026-08-11 they had grown to 72GB (63GB of it unused),
# the root filesystem hit 100% under two concurrent runs, and the harness died at startup with
# `OSError: [Errno 28] No space left on device: '/var/lib/cc-ci-runs/<build>'`. Every !testme
# from build 1236 to 1242 failed that way — with no results.json, so the PR badges just read
# "failure" and looked like recipe regressions.
#
# Only prune above a threshold, so a healthy host keeps its layer cache and runs stay fast.
# `image prune -a` removes only images no container references, so anything deployed (infra +
# warm-* canonicals) is untouched; anything else is re-pulled on demand.
#
# Volumes are deliberately NOT pruned here — see the KEEP_RE guard in (3): warm-* canonicals are
# data-warm and their volumes are legitimately dangling between runs.
DISK_PRUNE_PCT="${DISK_PRUNE_PCT:-60}"
used_pct="$(df --output=pcent / 2>/dev/null | tail -1 | tr -dc '0-9')"
if [ -n "$used_pct" ] && [ "$used_pct" -ge "$DISK_PRUNE_PCT" ]; then
echo " disk ${used_pct}% >= ${DISK_PRUNE_PCT}% -> pruning unused images"
freed="$(docker image prune -af 2>/dev/null | awk '/Total reclaimed space/ {print $4, $5}')"
echo " reclaimed: ${freed:-0B}; disk now $(df -h / | tail -1 | awk '{print $5" used, "$4" free"}')"
else
echo " disk ${used_pct:-?}% < ${DISK_PRUNE_PCT}% -> keeping image cache"
fi
if [ "$removed" -eq 0 ]; then
echo "== orphan sweep: clean (nothing to remove) =="
else
+2 -39
View File
@@ -78,38 +78,8 @@ ssh cc-ci 'systemctl --failed --no-legend; df -h / | tail -1; docker service ls
systemctl --failed --no-legend; df -h / | tail -1; tmux ls
```
- Failed units, core swarm services not 1/1 (warm-* spares flapping is a known benign pattern —
note, don't page), disk **>65% (server)** / >85% (orchestrator) → findings. Server unreachable →
note, don't page), disk >80% (server) / >85% (orchestrator) → findings. Server unreachable →
HIGH: recommend `hetzner-server-recovery`.
> **65%, not 80%, on the server — it is not a steady-state measure.** Two concurrent recipe runs
> pull images and write volumes worth tens of GB, so a host sitting at 73% still hits 100% mid-run.
> That is exactly what happened on 2026-08-11: 63GB of unused images had accumulated (nothing ever
> pruned them), the filesystem filled during a run, and the harness died at startup with
> `OSError: [Errno 28] No space left on device`. Remedy: `docker image prune -af` on cc-ci — it
> spares anything a container references, so infra and warm-* canonicals are untouched. Do NOT
> `docker volume prune`: warm-* canonical volumes are data-warm and legitimately dangling.
- **!testme actually produces results** (the check that would have caught the above days earlier):
the newest few `/var/lib/cc-ci-runs/<build>/` dirs must each contain `results.json`. A build that
dies before the harness writes one leaves an EMPTY dir — and the PR badge still says "failure", so
it reads as a recipe regression rather than a sick host. Builds 12361242 all failed that way.
Finding: *"N recent builds produced no results.json — the harness is dying at startup, check disk
and the drone step log"*. The step log lives in drone's sqlite
(`/var/lib/docker/volumes/drone_ci_commoninternet_net_data/_data/database.sqlite`) — copy it and
read `logs.log_data` for the failing `steps.step_id`; the bridge's drone token is not extractable
(distroless container, swarm secret).
> **If the error is ENOSPC but the disk is fine**, it is not disk. Seen 2026-08-11: builds 1244-1249
> died on `mkdir /var/lib/cc-ci-runs/<build>` with **110GB free and 16% inodes**, while the identical
> mkdir succeeded as root over ssh, inside the runner's own mount namespace, and 61/61 times in a
> stress loop — and the same harness run by hand with a numeric run id worked fine. Restarting
> `drone-runner-exec` did NOT help, and neither did recreating the runs directory with a fresh
> inode (it recurred afterwards — that apparent fix was coincidence).
>
> **It is INTERMITTENT and tracks concurrent activity**, which is the useful signal: every failure
> landed while a second run or a manual deploy was in flight (1252 was triggered while 1251 was
> still finishing), and every build on a quiet host succeeded (1243, 1250, 1251, 1253). Free space
> never moved during a failing build. So on ENOSPC-with-free-disk: **wait for the host to go quiet
> and re-trigger** before treating it as a recipe failure. Root cause is still NOT established;
> `DRONE_RUNNER_CAPACITY=2` allows the overlap, so lowering it to 1 is the obvious next experiment
> if it becomes disruptive.
- **Bridge / !testme path**: `docker service ls` shows `ccci-bridge_app 1/1` AND the bridge log
has no auth errors (`docker service logs --since 24h ccci-bridge_app 2>&1 | grep -ci "401\|user does not exist"` == 0).
A silently-401ing bridge drops `!testme` (seen 2026-08-03, stale rotated Gitea secret) →
@@ -136,15 +106,8 @@ systemctl --failed --no-legend; df -h / | tail -1; tmux ls
1. <finding> → /<skill> (or operator action)
```
When a finding is that the fleet's **security exposure is unknown** — the last weekly run failed or
is stale, so nobody has scanned for CVEs recently — the recommended step is **`/cve-check`** (read-only,
minutes, no PRs). If it is instead that a known CVE is sitting unpatched, recommend
**`/cve-check-and-upgrade`** (add `--min-severity high` when only the urgent ones matter). Prefer
`/cve-check` over waiting for the next weekly run whenever the question is "are we exposed?".
`ALL HEALTHY` requires: recent successful weekly run + published report, no stale tests, no
CVE PR open >14 days, both hosts <30 days behind their channel, zero failed units, recent builds all
producing results.json, disk under
CVE PR open >14 days, both hosts <30 days behind their channel, zero failed units, disk under
thresholds, bridge clean, maintained-set consistent. Anything else is a finding — even minor
ones get a recommended next step. Order findings by priority (CVE/unreachable-host first).
@@ -83,29 +83,8 @@ failure (AI — this is the `ci-test-review` step-3 diagnosis):
changed upstream, what the test currently asserts.
- **FLAKY** → re-run once or twice; if it passes, drop it (not stale, just flaky).
> ### ⚠️ RECONCILE FROM UPSTREAM FIRST — always, before any PR work or upgrade check
> ```
> cc-ci-plan/reconcile-upstream.sh <recipe>... # or --all
> ```
> Deterministic, idempotent, and safe (recipe work lives in branches, never on mirror `main`). It
> force-syncs each mirror to coopcloud's **default branch — resolved from the API, `main` OR
> `master`** — and closes any mirror PR whose changes upstream already merged. Skipping it has cost
> us three distinct ways: mailu #6 was reported as the fix for two internet-facing CVEs while
> upstream had already merged AND released it; a stale mirror makes a survey report "no upgrades
> available" so the recipe drops out of the weekly run; and reading the wrong branch on a recipe with
> a stale `main` beside a live `master` (gitea) manufactures a false "three releases behind, missing
> two CVSS-9.8 RCEs" finding.
### 2. For each stale test — author the minimal test update (AI; never weaken)
> **Read `tests/STYLE.md` in the cc-ci repo before writing the update.** It is the rulebook for test
> changes, written from failures this pipeline actually produced. Most load-bearing: set state up
> through the app's **own interface, never its database** (a plausible fixture that INSERTed rows
> passed on v2 and silently broke on v3 — 202 acks, rows in postgres, nothing ingested — and held the
> recipe RED for six weeks), **gate on version rather than supporting both** (old-version tests can be
> deleted; the older version is only exercised via the upgrade tier), and correct the fixture or the
> wait but **never the assertion**.
Work on **one recipe at a time** (serialize — each verification deploys a recipe on the shared
Swarm). For each `STALE_TESTS` entry:
-18
View File
@@ -31,20 +31,6 @@ Then present the roster grouped as follows, and close with the situation guide.
PR). `--with-tests` also fixes that recipe's stale test.
- **/recipe-report** — (re)generate the weekly report page for report.ci.commoninternet.net.
**Keeping the PR list honest**
- **/cc-ci-cleanup** — reconciles every mirror from true upstream (which alone closes PRs upstream
already merged), then closes the open recipe PRs that can no longer merge or were never meant to
(CI sweep artifacts, obsolete bumps, superseded duplicates) and reports what is actually blocking
the ones that should land. Never merges.
**Security (CVEs)**
- **/cve-check** — fleet-wide CVE sweep with **no upgrading**: for every recipe, work out what
upgrade is available (per image, sidecars included), scan it for CVEs, and publish a CVE report.
Read-only and quick — the "what are we exposed to?" answer without an upgrade run.
- **/cve-check-and-upgrade** — the same sweep, then open verified PRs **only** for the recipes whose
upgrade actually fixes a CVE, worst severity first. `--min-severity high` for just the urgent ones.
Never merges.
**Tests**
- **/cc-ci-tests-update** — fleet-wide stale-test cleanup: find tests broken by legitimate
upstream changes, fix without weakening, verify, merge the test PRs.
@@ -87,10 +73,6 @@ ARM skills never touch cc-ci infra. After a submodule bump run `scripts/gen-ccte
| "Run the weekly upgrades now" | `/upgrade-all` (or `systemctl start cc-ci-upgrade-all.service`) |
| "Upgrade just <recipe>" | `/recipe-upgrade <recipe>` |
| "The report site is stale/missing a week" | `/recipe-report` |
| "The open PR list is a mess / what should I merge?" | `/cc-ci-cleanup` |
| "What CVEs are we exposed to right now?" | `/cve-check` (read-only, no PRs) |
| "A CVE just dropped — check and patch it" | `/cve-check-and-upgrade` (add `--min-severity high` to skip the noise) |
| "Is <recipe> vulnerable?" | `/cve-check <recipe>` |
| "Tests are red because upstream changed" | `/cc-ci-tests-update` (fleet) or `/recipe-upgrade <r> --with-tests` |
| "A CI run failed and I don't know why" | `/ci-test-review` |
| "Update the CI server OS/deps" | `/cc-ci-server-update` |
-26
View File
@@ -115,29 +115,3 @@ When the orchestrator, Builder, or assistant makes intentional repository change
promptly and push them to `git.autonomic.zone` in append-only fashion (never force-push). Match the
existing commit author and message style in this repo. Do not bundle unrelated worktree changes you
did not make; stage only the intended files.
## Ship as PRs, merge them yourself, operator reviews retrospectively
**This applies to the two INFRASTRUCTURE repos — `recipe-maintainers/cc-ci-orchestrator` (here) and
`recipe-maintainers/cc-ci` (the CI product).** For work in either:
1. Branch, don't commit straight to `main`.
2. Open a PR with a description written to be read **after** the fact: what changed, why, and what
evidence says it works (test output, a verified run, a before/after number). The PR *is* the
review artifact and the historical record.
3. **Merge it yourself once it is verified** — do not wait for review. The invocation is the
authorization; blocking on review would stall the pipeline these repos exist to run.
4. The operator reviews **retrospectively**, from the PR.
So the PR is not a gate — it is how the work stays legible. A PR that merely says "fix scanner" has
failed at its only job.
> ### This does NOT extend to RECIPE repos
> Recipe PRs — any `coop-cloud/<recipe>` or its `recipe-maintainers/<recipe>` mirror — are
> **created and verified but NEVER merged by an agent**. Those change what deploys on other people's
> infrastructure, so a human merges them. The split is deliberate: agents own the tooling, the
> operator owns the recipes.
If work has already landed on `main` without a PR, do not rewrite published history to fix it.
Create a branch pinned at the pre-work commit and open the PR against that, so the diff is still
reviewable and merging only advances the pointer (see PRs #2-#5, 2026-08-11).
+3 -76
View File
@@ -39,8 +39,6 @@ keeps landing in pass 2, the fix is a new deterministic method in pass 1. §4c i
```
advisory-scan.py <recipe> [--from <version>] [--to <version>]
[--image <name>=<from>:<to>]... [--adjudicate] [--json] [--registry DIR]
advisory-scan.py <recipe> --compose-to <URL> [--compose-from <URL>] # windows derived, not typed
```
| Input | Meaning |
@@ -48,8 +46,6 @@ advisory-scan.py <recipe> --compose-to <URL> [--compose-from <URL>] # windows
| `<recipe>` | Recipe name; selects `cc-ci-plan/upstream/<recipe>.md` (the per-recipe URL registry) |
| `--from` / `--to` | The **primary app image's** version window being upgraded across |
| `--image NAME=FROM:TO` | A **sidecar image and the versions it moved between** (repeatable, all in ONE call). `NAME` is substring-matched against source repo names. Malformed values warn on stderr and are skipped. Without it that image's advisories stay unclassified. |
| `--compose-to URL` | **Derive every window by diffing this compose against its baseline**, instead of typing `--from/--to/--image`. Point it at a PR's `compose.yml`. |
| `--compose-from URL` | Baseline for the above. Default: the same repo's **default branch, resolved from the API** — never assumed to be `main`. |
| `--adjudicate` | Run pass 2: append the evidence dossier for judgement |
| `--registry` | Registry dir; also `CCCI_UPSTREAM_REGISTRY` |
| `GITHUB_TOKEN` / `GITHUB_TOKEN_FILE` | Read-only token; **rate limit only** (60/hr anonymous → 5000/hr). Default file `/srv/cc-ci/.github-token`, mode 600. Public advisories need **no scopes**. |
@@ -107,43 +103,10 @@ URLs containing `<`, `>`, `{`, `}`, `VERSION`, or `vX.Y.Z` are **skipped as temp
human documentation (`…/changelog/v<VERSION>/`), not fetchable, and counting them as failures is wrong.
This is the source that would have caught gitea: the vendor blog names both CVEs, the GitHub release
page names neither.
page names neither. A CVE found **only** here carries no version data, so pass 1 cannot place it — it
goes to pass 2 (§6).
**When the page is a changelog organised by release, each CVE is attributed to the release heading it
appears under** (`Changes with nginx 1.31.3`, `## v1.31.3`, …) and that becomes its fixed-in version.
Without this, a project that publishes no advisory feed can never contribute a CVE:
> **nginx publishes NO GitHub security advisories.** Every nginx CVE we can see comes from
> `nginx.org/en/CHANGES`. Scraping ids out of it without attributing them to a release left them with
> no patched version, so they were never classifiable — and every nginx bump in the fleet reported
> **0** forever. nginx is a sidecar in most recipes. Measured: `1.31.1 → 1.31.3` fixes **six** CVEs
> (three in .2, three in .3); lasuite-docs#7 went 0 → 6 and lasuite-drive#6 went 0 → 3 on this alone.
A changelog CVE is tied to a window by the **image name appearing in the page URL** (window `nginx` ↔
`nginx.org/...`). A CVE found on a vendor page with no attributable release still has no version data,
so pass 1 cannot place it — it goes to pass 2 (§6).
### 2c. NVD by CPE — the fallback for projects that publish nothing
Declared per recipe in the registry as `nvd-cpe: <image-key> = <cpe:2.3:...>`.
> **Why it exists.** Two recipes could not see CVEs *at all*: `mattermost-lts` (empty GitHub advisory
> feed, security bulletins rendered client-side so a text sweep finds nothing) and `mumble` (nothing
> published anywhere the registry points). Their scans returned `?` — nothing measured. NVD is
> CPE-indexed and carries structured ranges, so it answers where the vendor does not: mattermost
> 10.5.0 → 10.12.4 now scores **165**, and mumble finds `CVE-2025-71264` (fixed 1.6.870).
Two range forms, both used:
| NVD field | meaning | how it is judged |
|---|---|---|
| `versionEndExcluding X` | fixed in X exactly | a normal patched version (§4a) |
| `versionEndIncluding X` | affected **up to and including** X; fix version unpublished | fixed when the upgrade crosses X, i.e. `from ≤ X < to` |
**NVD lags the vendor** — it had neither gitea CVSS-9.8 RCE at publication — so this is a fallback,
never a replacement for 2a/2b. Unauthenticated calls are rate-limited (~5/30s), hence the retry.
### 2d. OSV.dev — supplementary
### 2c. OSV.dev — supplementary
Only when the recipe has an entry in `OSV_PACKAGES` (ecosystem + package) and a version is given.
@@ -177,44 +140,8 @@ Two invariants govern this step, both learned from a wrong answer in production.
> `null` / `UNKNOWN`, never `0`. A `0` in a security column asserts safety. Equally, an advisory that
> cannot be judged is **indeterminate** (§4d) — never silently counted as "not fixed".
### 3b. Deriving the windows from a compose diff (`--compose-to`)
Typing `--from/--to/--image` by hand means someone has to remember that the recipe also bumped its
redis. That is how sidecar CVEs went uncounted for months. This mode reads the windows off the diff:
1. Fetch both compose files (baseline = the repo's **default branch from the API**, since several
recipes keep a stale `main` beside a live `master`).
2. Parse `{service: (image-repo, tag)}` — keyed by **service, not image repo**, because an upgrade
may change the repo itself (plausible moved `plausible/analytics` →
`ghcr.io/plausible/community-edition`; keyed by repo that reads as one image vanishing and an
unrelated one appearing, losing the app window entirely).
3. Every service whose tag or repo changed becomes a window. The `app` service drives `--from/--to`
(coop-cloud convention: it is the recipe's primary image); the rest become `--image` windows.
Unchanged images produce no window — inventing one would be a false count.
4. The derived windows are printed to stderr before the scan, so the inputs are auditable.
Image names are matched against advisory sources **both ways** — an image name is often longer than
its source repo (`clickhouse/clickhouse-server` vs `ClickHouse/ClickHouse`) and sometimes shorter
(`redis` vs `redis/redis`).
Verified on plausible PR #5: from the compose URL alone it derives `v2.0.0 → v3.2.1` plus
`clickhouse-server 23.4.2.11-alpine → 24.12-alpine`, and reports **6** — identical to the
hand-specified args.
`--from/--to/--image` remain available for finer-grained checks (scanning a window that is not a
literal compose diff, e.g. "what would the compatibility-safe target fix?").
### 4a. By patched version (preferred — exact)
**A fix on the line you are upgrading FROM was already yours.** Projects that maintain several lines
patch them all at once: mattermost fixed `CVE-2025-11794` in 10.11.4, 10.12.1 *and* 10.5.12. An
upgrade 10.11.22 → 10.12.4 crosses 10.12.1, so a naive window test counts it — but 10.11.22 is
already past 10.11.4, so the deployment had the fix before the upgrade. Counting it credits the
upgrade with work it did not do. This check is **skipped for placeholder versions** (`7.4.X` parses
to a bare `7.4`, which would read as "already fixed at 7.4" and silently drop a real fix — exactly
how redis `CVE-2024-46981` was lost when the rule was first added).
`patched_versions` is a **range expression** (`">= 2.18.1"`), possibly several joined by `;`. Extract
every version-looking token; the advisory is **fixed-by-this-upgrade** if **any** patched version `p`
satisfies `from < p <= to` — exclusive lower (a fix already in the version you were on is not this
+21 -445
View File
@@ -44,9 +44,7 @@ import json
import os
import re
import sys
import time
import urllib.error
import urllib.parse
import urllib.request
REGISTRY_DIR = os.environ.get("CCCI_UPSTREAM_REGISTRY", "/srv/cc-ci/cc-ci-plan/upstream")
@@ -166,42 +164,6 @@ def _vkey(v: str | None) -> tuple:
return tuple(out)
def _already_fixed_on_from_line(kf: tuple, cands: list[tuple]) -> bool:
"""Was it ALREADY fixed on the line we are upgrading FROM?
The mirror image of _superseded_on_target_line, and just as necessary. mattermost fixes each CVE
across several maintained lines at once — CVE-2025-11794 is patched in 10.11.4, 10.12.1 and
10.5.12. Upgrading 10.11.22 -> 10.12.4 crosses 10.12.1, so a naive window test counts it; but
10.11.22 is already past 10.11.4, so the deployment HAD the fix before the upgrade. Counting it
credits the upgrade with work it did not do."""
if len(kf) < 2:
return False
line = kf[:2]
for c in cands:
if len(c) < 2 or c[:2] != line:
continue
n = max(len(kf), len(c))
pad = lambda z: z + (0,) * (n - len(z))
if pad(c) <= pad(kf):
return True
return False
def _superseded_on_target_line(kt: tuple, cands: list[tuple]) -> bool:
"""Does a patched version on the TARGET's own release line sit ABOVE the target?
Projects maintain several branches at once and backport per branch, so "some patched version is
inside the numeric window" is not the same as "the version we land on has the fix". ClickHouse
fixed CVE-2023-48704 in 23.9.6.20 AND 23.10.5.20; an upgrade landing on 23.10.4.25 crosses the
23.9 fix numerically but is still BELOW its own line's fix, so it does NOT have it. When the
advisory names a fix on the target's own line and the target is older than it, that is proof of
absence and outranks any other candidate."""
if len(kt) < 2:
return False
line = kt[:2]
return any(c[:2] == line and c > kt for c in cands if len(c) >= 2)
def _within(kf: tuple, kt: tuple, c: tuple) -> bool:
"""Is patched-version `c` inside the window (kf, kt] — exclusive lower, inclusive upper?
@@ -295,51 +257,6 @@ def github_advisories(urls: list[str]) -> list[dict]:
return results
# Release headings in a vendor changelog. nginx's CHANGES uses "Changes with nginx 1.31.3", most
# markdown changelogs use "## 1.31.3" / "## v1.31.3".
_HEADING_RE = re.compile(
r"^\s*(?:#{1,4}\s*)?(?:Changes with\s+\S+\s+|Version\s+|Release\s+)?v?(\d+\.\d+(?:\.\d+)*)\s*$"
r"|^\s*Changes with\s+\S+\s+(\d+\.\d+(?:\.\d+)*)", re.I)
def _changelog_versions(text: str) -> dict:
"""{cve: version} for a changelog that is ORGANISED BY RELEASE.
Why this exists: nginx publishes NO GitHub security advisories. Every nginx CVE we can see comes
from nginx.org/en/CHANGES, and scraping ids out of it without attributing them to a release
leaves them with no patched version — so they can never be classified, and an nginx bump reports
0 CVEs forever. nginx 1.31.1 -> 1.31.3 in fact fixes SIX (three in .2, three in .3), and nginx is
a sidecar in most of the fleet, so that was a fleet-wide blind spot.
Attributes each CVE to the nearest PRECEDING release heading — the release that fixed it.
"""
plain = re.sub(r"<[^>]+>", " ", text)
out, cur = {}, None
for line in plain.splitlines():
m = _HEADING_RE.match(line)
if m:
cur = m.group(1) or m.group(2)
continue
if cur:
for cve in CVE_RE.findall(line):
out.setdefault(cve, cur)
return out
_BLOB_RE = re.compile(r"^https://github\.com/([^/]+)/([^/]+)/blob/(.+)$")
def _raw_if_blob(url: str) -> str:
"""A GitHub *blob* URL is an HTML viewer, not the file.
The registry pointed ONLYOFFICE's CHANGELOG.md at its blob page. Fetching that returns 636KB of
markup in which the release headings do not survive HTML-stripping, so 24 CVEs were visible and
NONE attributable to a release — the same shape of blind spot as nginx. The raw URL attributes
all 24. Normalising here fixes every registry entry at once, present and future."""
m = _BLOB_RE.match(url)
return f"https://raw.githubusercontent.com/{m.group(1)}/{m.group(2)}/{m.group(3)}" if m else url
def vendor_pages(urls: list[str]) -> list[dict]:
"""Fetch each registry URL and regex out CVE ids, with a little surrounding context."""
out = []
@@ -352,93 +269,20 @@ def vendor_pages(urls: list[str]) -> list[dict]:
# correct; counting them as failures would wrongly mark the recipe's count unreliable.
out.append({"source": u, "status": "skipped: template URL (not fetchable)", "cves": [], "context": {}})
continue
entry = {"source": u, "status": "ok", "cves": [], "context": {}, "fixed_in": {}}
entry = {"source": u, "status": "ok", "cves": [], "context": {}}
try:
text = _fetch(_raw_if_blob(u))
text = _fetch(u)
plain = re.sub(r"<[^>]+>", " ", text)
for cve in sorted(set(CVE_RE.findall(plain))):
entry["cves"].append(cve)
i = plain.find(cve)
entry["context"][cve] = re.sub(r"\s+", " ", plain[max(0, i - 160) : i + 200]).strip()
entry["fixed_in"] = _changelog_versions(text)
except Exception as e: # noqa: BLE001
entry["status"] = f"error: {type(e).__name__}: {e}"
out.append(entry)
return out
NVD_API = "https://services.nvd.nist.gov/rest/json/cves/2.0"
NVD_CPE_RE = re.compile(r"^\s*[-*]?\s*nvd-cpe:\s*(\S+)\s*=\s*(cpe:2\.3:[^\s`]+)", re.M | re.I)
def registry_cpes(recipe: str, registry_dir: str) -> list[tuple[str, str]]:
"""[(image-key, cpe)] declared in the recipe's registry as `nvd-cpe: <key> = <cpe>`."""
path = os.path.join(registry_dir, f"{recipe}.md")
try:
return [(m.group(1), m.group(2)) for m in NVD_CPE_RE.finditer(open(path).read())]
except OSError:
return []
def nvd_advisories(cpe: str, key: str) -> dict:
"""CVEs for a CPE from NVD, with the version data the classifier needs.
THE FALLBACK FOR PROJECTS THAT PUBLISH NOTHING MACHINE-READABLE. mattermost's GitHub advisory
feed is empty and its security bulletins are client-side rendered; mumble publishes neither. Both
scanned as `?` — nothing measured — until here. NVD is CPE-indexed and carries structured ranges:
versionEndExcluding X -> fixed in X exactly (a patched version)
versionEndIncluding X -> affected up to and INCLUDING X, fixed in some later release. The
exact fix version is unknown, but the upgrade fixes it whenever it
crosses X — recorded as `affected_max` and judged in the classifier.
NVD LAGS the vendor (it had neither gitea CVSS-9.8 RCE at publication), so this is a fallback,
never a replacement for 2a/2b. Unauthenticated calls are rate-limited to ~5/30s, hence the retry.
"""
entry = {"source": f"nvd:{key}", "status": "ok", "advisories": []}
url = f"{NVD_API}?resultsPerPage=2000&virtualMatchString={urllib.parse.quote(cpe)}"
data = None
for attempt in range(3):
try:
data = json.loads(_fetch(url))
break
except Exception as e: # noqa: BLE001
if attempt == 2:
entry["status"] = f"error: {type(e).__name__}"
return entry
time.sleep(8)
for v in (data or {}).get("vulnerabilities", []):
c = v.get("cve") or {}
cid = c.get("id")
if not cid:
continue
fixed, affected_max = set(), set()
for cfg in c.get("configurations", []):
for node in cfg.get("nodes", []):
for m in node.get("cpeMatch", []):
if m.get("versionEndExcluding"):
fixed.add(m["versionEndExcluding"])
elif m.get("versionEndIncluding"):
affected_max.add(m["versionEndIncluding"])
sev = None
for mk in ("cvssMetricV31", "cvssMetricV30", "cvssMetricV2"):
got = (c.get("metrics") or {}).get(mk) or []
if got:
sev = (got[0].get("cvssData") or {}).get("baseSeverity")
break
entry["advisories"].append({
"cve": cid, "ghsa": None, "severity": (sev or "").lower() or None,
"summary": next((d.get("value") for d in c.get("descriptions", [])
if d.get("lang") == "en"), "")[:200],
"vulnerable_range": None,
"patched": "; ".join(sorted(fixed)) or None,
"affected_max": "; ".join(sorted(affected_max)) or None,
"url": f"https://nvd.nist.gov/vuln/detail/{cid}",
"published_at": c.get("published"), "description": None, "cvss": None,
})
return entry
def osv(recipe: str, version: str | None) -> dict | None:
pkg = OSV_PACKAGES.get(recipe)
if not pkg or not version:
@@ -489,19 +333,6 @@ def _releases(owner: str, repo: str, max_pages: int = 4) -> list[tuple[str, str]
return out
def _source_repo(source: str) -> tuple[str, str] | None:
"""(owner, repo) for a source, whether it is an advisory feed or a vendor page on GitHub.
A CVE that appears ONLY on a vendor page still deserves the release-note method when that page
lives on GitHub — mailu announces its Roundcube CVEs on github.com/Mailu/Mailu/releases and
nowhere structured, so requiring an advisory feed sent a deterministic case to pass 2."""
if source.startswith("github-advisories:"):
owner, _, repo = source.split(":", 1)[1].partition("/")
return (owner, repo) if owner and repo else None
m = re.match(r"https?://github\.com/([^/]+)/([^/#?]+)", source)
return (m.group(1), m.group(2).removesuffix(".git")) if m else None
def release_fix_versions(source: str, cve: str) -> list[str]:
"""Release tags whose notes NAME this CVE — a deterministic fix version when the advisory has none.
@@ -511,10 +342,10 @@ def release_fix_versions(source: str, cve: str) -> list[str]:
(CVE-2025-32023 → 6.2.19, 7.2.10, 7.4.5, 8.0.3, 8.2.0). Ignoring that evidence undercounted
discourse by 12 CVEs, so this is checked BEFORE giving up on an advisory.
"""
ref = _source_repo(source)
if not ref:
if not source.startswith("github-advisories:"):
return []
return [tag for tag, body in _releases(*ref) if cve in body]
owner, _, repo = source.split(":", 1)[1].partition("/")
return [tag for tag, body in _releases(owner, repo) if cve in body]
def advisory_text(ghsa: str, source: str | None = None) -> dict:
@@ -736,8 +567,7 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
e = report["cves"].setdefault(cve, {"sources": [], "severity": None, "ghsa": None,
"vulnerable_range": None, "patched": None,
"context": None, "published_at": None,
"description": None, "url": None, "cvss": None,
"changelog_fixed_in": None, "affected_max": None})
"description": None, "url": None, "cvss": None})
if src not in e["sources"]:
e["sources"].append(src)
for k, v in extra.items():
@@ -754,22 +584,11 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
context=a.get("summary"), published_at=a.get("published_at"),
description=a.get("description"), url=a.get("url"), cvss=a.get("cvss"))
for key, cpe in registry_cpes(recipe, registry_dir):
entry = nvd_advisories(cpe, key)
report["sources"].append({"source": entry["source"], "status": entry["status"],
"found": len(entry.get("advisories") or [])})
for a in entry.get("advisories", []):
record(a["cve"], entry["source"], severity=a.get("severity"),
patched=a.get("patched"), affected_max=a.get("affected_max"),
context=a.get("summary"), published_at=a.get("published_at"),
url=a.get("url"))
for entry in vendor_pages(urls):
report["sources"].append({"source": entry["source"], "status": entry["status"],
"found": len(entry.get("cves", []))})
for cve in entry.get("cves", []):
record(cve, entry["source"], context=entry["context"].get(cve),
changelog_fixed_in=(entry.get("fixed_in") or {}).get(cve))
record(cve, entry["source"], context=entry["context"].get(cve))
for version in filter(None, (v_from, v_to)):
o = osv(recipe, version)
@@ -804,39 +623,17 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
#
# A source with no window is not classified: its advisories are listed as unclassified so they
# stay visible without inflating the count.
gh_sources = [x["source"] for x in report["sources"]
if x["source"].startswith(("github-advisories:", "nvd:"))]
gh_sources = [x["source"] for x in report["sources"] if x["source"].startswith("github-advisories:")]
primary = gh_sources[0] if (gh_sources and (v_from or v_to)) else None
report["primary_source"] = primary
windows = {} # source name -> (from, to)
window_key = {} # source name -> the image name it covers
if primary:
windows[primary] = (v_from, v_to)
window_key[primary] = primary.split("/")[-1]
# The app's window must also cover its NVD entry. NVD sources are keyed by IMAGE name
# (`mattermost-team-edition`) while the advisory feed is keyed by REPO (`mattermost/
# mattermost`), so without this the fallback source that exists precisely because the feed
# is empty would itself go unwindowed — and mumble/mattermost would still report nothing.
pname = primary.split("/")[-1].lower()
for src in gh_sources:
if src.startswith("nvd:") and src not in windows:
k = src.split(":", 1)[1].lower()
if pname in k or k in pname:
windows[src] = (v_from, v_to)
window_key[src] = k
for key, wf, wt in (images or []):
for src in gh_sources:
if src in windows:
continue
# Match BOTH ways: an image name is often longer than its source repo
# (`clickhouse/clickhouse-server` vs source `ClickHouse/ClickHouse`) and sometimes
# shorter (`redis` vs `redis/redis`). One-directional matching silently dropped the
# clickhouse window when the key was derived from a compose file.
k, name = key.lower(), src.split("/")[-1].lower()
if k in src.lower() or name in k:
if key.lower() in src.lower() and src not in windows:
windows[src] = (wf, wt)
window_key[src] = key
report["windows"] = {k: {"from": f, "to": t} for k, (f, t) in windows.items()}
def _classify_window(src, wf, wt):
@@ -855,31 +652,8 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
continue
patched = e.get("patched") or ""
cands = [_vkey(t) for t in re.findall(r"\d+(?:\.\d+)*", patched)]
# NEVER on a placeholder: "7.4.X" parses to the bare 7.4, which then reads as
# "already fixed at 7.4" and silently drops a real fix (redis CVE-2024-46981).
# A placeholder means the fix version is unknown — that is the indeterminate path.
if (kf and kt and not PLACEHOLDER_RE.search(patched)
and _already_fixed_on_from_line(kf, cands)):
# already had it before the upgrade
e.setdefault("classification", "outside-window")
continue
if kf and kt and _superseded_on_target_line(kt, cands):
# The target's own line got the fix LATER than the target: not fixed here.
e.setdefault("classification", "outside-window")
continue
if kf and kt and any(_within(kf, kt, c) for c in cands):
got.add(cve)
elif kf and kt and e.get("affected_max"):
# NVD's `versionEndIncluding X`: affected up to and INCLUDING X, fixed in some
# later release. The exact fix version is unpublished, but the upgrade delivers
# it whenever it crosses X — i.e. from <= X < to.
for t in re.findall(r"\d+(?:\.\d+)*", e["affected_max"]):
x = _vkey(t)
n = max(len(kf), len(kt), len(x))
pad = lambda z: z + (0,) * (n - len(z))
if x and pad(kf) <= pad(x) < pad(kt):
got.add(cve)
break
elif not patched or PLACEHOLDER_RE.search(patched):
# No fix version published ("TBD") or only a placeholder ("7.4.X" — which could
# be 7.4.1, inside the window). We cannot say either way, so say so.
@@ -913,32 +687,6 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
report["cves"][cve]["classification"] = f"fixed-by-this-upgrade ({method}) via {src}"
fixed_set.add(cve)
# A CVE seen only in a vendor CHANGELOG has no advisory feed behind it, but the changelog says
# which release fixed it (see _changelog_versions). Tie it to a window by the image name
# appearing in the page URL — nginx's window is `nginx`, and its changelog is nginx.org/... .
# Without this, projects that publish no GitHub advisories (nginx being the big one) can never
# contribute a CVE, and every nginx bump in the fleet silently reports 0.
from_changelog = {}
for cve, e in report["cves"].items():
if cve in fixed_set or not e.get("changelog_fixed_in"):
continue
for src, (wf, wt) in windows.items():
key = (window_key.get(src) or "").lower()
if not key:
continue
if not any(key in s_.lower() for s_ in e["sources"] if s_.startswith("http")):
continue
kf, kt = _vkey(wf), _vkey(wt)
cand = _vkey(e["changelog_fixed_in"])
if kf and kt and cand and _within(kf, kt, cand):
e["classification"] = (f"fixed-by-this-upgrade (named under {e['changelog_fixed_in']} "
f"in the vendor changelog) via {src}")
fixed_set.add(cve)
from_changelog[cve] = e["changelog_fixed_in"]
break
if from_changelog:
report["resolved_by_changelog"] = from_changelog
unknown = []
for cve, e in report["cves"].items():
if cve in fixed_set:
@@ -959,31 +707,17 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
# THIRD METHOD: before declaring an advisory undecidable, look for the CVE id in the project's
# own release notes. A tag that names it, inside the window, IS the fix version the advisory
# failed to publish. Deterministic and citable — not a judgement call.
# A window is keyed by advisory-feed source; map it to its repo so a vendor page on the SAME
# repo can be judged by the same window.
win_by_repo = {}
for wsrc, wv in windows.items():
ref = _source_repo(wsrc)
if ref:
win_by_repo[ref] = wv
resolved_by_release, already_fixed = {}, set()
candidates = set(indeterminate) | {
c for c in unknown
if not any(s in windows for s in report["cves"][c]["sources"])
and any(_source_repo(s) in win_by_repo for s in report["cves"][c]["sources"])
}
for cve in sorted(candidates - fixed_set):
resolved_by_release = {}
for cve in sorted(indeterminate - fixed_set):
e = report["cves"][cve]
for src in e["sources"]:
ref = _source_repo(src)
if src not in windows and ref not in win_by_repo:
if src not in windows:
continue
wf, wt = windows[src] if src in windows else win_by_repo[ref]
wf, wt = windows[src]
kf, kt = _vkey(wf), _vkey(wt)
if not (kf and kt):
continue
naming = release_fix_versions(src, cve)
hits = [t for t in naming if _within(kf, kt, _vkey(t))]
hits = [t for t in release_fix_versions(src, cve) if _within(kf, kt, _vkey(t))]
if hits:
fixed_set.add(cve)
resolved_by_release[cve] = sorted(hits)
@@ -991,23 +725,10 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
f"{', '.join(sorted(hits))}) via {src}")
e["fix_versions_from_release_notes"] = sorted(hits)
break
# Naming releases exist but ALL predate the version we were already on: the fix shipped
# before this upgrade, so the upgrade did not deliver it. That is a DECISION, not an
# unknown — mailu's redis 8.8.0 → 8.10.0 crosses 12 advisories all fixed by 8.6.3 or
# earlier, and reporting them as "could not judge" overstates the uncertainty.
if naming and all(_vkey(t) and not _within(kf, kt, _vkey(t)) for t in naming) \
and max(_vkey(t) for t in naming if _vkey(t)) <= kf:
e["classification"] = ("outside-window: fixed in "
f"{', '.join(sorted(naming))}, all at or before {wf}")
e["fix_versions_from_release_notes"] = sorted(naming)
already_fixed.add(cve)
break
if resolved_by_release:
report["resolved_by_release_notes"] = resolved_by_release
indeterminate -= fixed_set | already_fixed
if already_fixed:
report["already_fixed_before_upgrade"] = sorted(already_fixed)
indeterminate -= fixed_set
for cve in indeterminate:
report["cves"][cve]["classification"] = "indeterminate: no fix version published"
unknown = [c for c in unknown if c not in fixed_set]
@@ -1017,19 +738,8 @@ def scan(recipe: str, v_from: str | None, v_to: str | None, registry_dir: str,
report["unclassified"] = sorted(unknown)
# NEVER report 0 for something we could not determine — a 0 asserts safety. If ANY requested
# window could not be ordered at all, the total is UNKNOWN rather than a partial number.
# A scan with NO usable source has not measured anything, so it must not report a number —
# least of all 0, which asserts safety. mumble had no cc-ci-plan/upstream/mumble.md at all and
# still produced "0 identified", which was then published as a clean 0 in a CVE report.
usable_sources = [
s for s in report["sources"]
if s["status"] == "ok" or s["status"].startswith("no-advisories-published")
]
no_sources = not usable_sources
if no_sources:
report["no_usable_sources"] = True
report["count_known"] = not unresolved_any and not no_sources
report["cve_count_fixed"] = (len(fixed_set)
if (not unresolved_any and not no_sources) else None)
report["count_known"] = not unresolved_any
report["cve_count_fixed"] = len(fixed_set) if not unresolved_any else None
report["cve_count_total_seen"] = len(report["cves"])
# Only GENUINE failures make a count unreliable. "no-advisories-published" (404: the repo has
# no advisory feed) and "skipped: template URL" are benign and must not degrade the verdict.
@@ -1052,16 +762,10 @@ def markdown(rep: dict) -> str:
f"{rep.get('from') or '?'}{rep.get('to') or '?'}"]
if not rep.get("count_known", True):
L.append("\n**CVEs fixed by this upgrade: UNKNOWN — the scan could NOT determine a count.**")
if rep.get("no_usable_sources"):
L.append("\n⚠ This is NOT zero. **No usable source was checked at all** — the registry "
"file `cc-ci-plan/upstream/<recipe>.md` is missing or every source failed, so "
"nothing was measured. Render this recipe's cve cell as `?`, never `0`, and add "
"the registry file.")
else:
L.append("\n⚠ This is NOT zero. A version-scheme change (e.g. semver → calver) makes "
"numeric ordering meaningless across this jump, so no advisory could be "
"classified. Render this recipe's cve cell as `?`, never `0`. Read the vendor's "
"release notes for the jump and count by hand.")
L.append("\n⚠ This is NOT zero. A version-scheme change (e.g. semver → calver) makes numeric "
"ordering meaningless across this jump, so no advisory could be classified. Render "
"this recipe's cve cell as `?`, never `0`. Read the vendor's release notes for the "
"jump and count by hand.")
if rep["unclassified"]:
L.append(f"\nAdvisories seen but unclassifiable ({len(rep['unclassified'])}) — includes "
f"other images in this recipe: " + ", ".join(rep["unclassified"][:12]))
@@ -1113,113 +817,6 @@ def markdown(rep: dict) -> str:
return "\n".join(L)
def _gitea_auth(url: str) -> dict:
"""Basic auth for the private mirror, from /srv/cc-ci/.testenv.
Sent as a HEADER, never embedded in the URL: in-URL credentials leak into shell history, process
lists and error messages, and urllib mis-parses a password containing a colon."""
host = re.sub(r"^https?://", "", url).split("/")[0]
env = {}
try:
for ln in open(os.environ.get("CCCI_TESTENV", "/srv/cc-ci/.testenv")):
if "=" in ln and not ln.strip().startswith("#"):
k, v = ln.strip().split("=", 1)
env[k] = v.strip().strip("\"'")
except OSError:
return {}
if host != env.get("GITEA_URL", "git.autonomic.zone"):
return {}
u, pw = env.get("GITEA_USERNAME"), env.get("GITEA_PASSWORD")
if not (u and pw):
return {}
import base64 as _b64
return {"Authorization": "Basic " + _b64.b64encode(f"{u}:{pw}".encode()).decode()}
def _compose_images(url: str) -> dict[str, tuple[str, str]]:
"""{service: (image-repo, tag)} for a compose file.
Keyed by SERVICE, not by image repo, because an upgrade may change the repo itself: plausible
moved `plausible/analytics` -> `ghcr.io/plausible/community-edition`. Keyed by repo that reads
as one image vanishing and an unrelated one appearing, and the app's version window is lost —
which is exactly the upgrade most worth scanning."""
txt = _fetch(url, _gitea_auth(url))
out, svc = {}, None
in_services = False
for line in txt.splitlines():
if re.match(r"^services:\s*$", line):
in_services = True
continue
if in_services and re.match(r"^\S", line):
in_services = False
if not in_services:
continue
m = re.match(r"^ (\S+):\s*$", line)
if m:
svc = m.group(1)
continue
m = re.match(r"^\s+image:\s*[\"']?([^\"'\s]+)", line)
if m and svc:
ref = m.group(1).split("@", 1)[0]
if "${" in ref or "$(" in ref:
continue
repo, _, tag = ref.rpartition(":")
if repo and tag:
out[svc] = (repo, tag)
return out
def _default_branch_compose(url: str) -> str | None:
"""Same repo as `url`, but its DEFAULT branch — resolved from the API, never assumed.
Several coopcloud recipes keep a stale `main` beside the real default `master` (gitea's `main`
is 1.24.2-rootless while `master` has 1.27.1-rootless), so guessing the branch produces a
confidently wrong baseline."""
m = re.match(r"(https?://[^/]+)/([^/]+)/([^/]+)/(?:raw|src)/branch/[^/]+/(.*)$", url)
if not m:
return None
host, owner, repo, path = m.groups()
try:
meta = json.loads(_fetch(f"{host}/api/v1/repos/{owner}/{repo}", _gitea_auth(host)))
br = meta.get("default_branch")
except Exception: # noqa: BLE001
return None
return f"{host}/{owner}/{repo}/raw/branch/{br}/{path}" if br else None
def windows_from_compose(to_url: str, from_url: str | None = None) -> tuple[list, str | None]:
"""Derive the scan's version windows by DIFFING two compose files.
This is the deterministic alternative to a human (or a model) deciding which `--image` args a
given upgrade needs. Point it at a PR's compose and it reads the windows straight off the diff:
every image whose tag changed becomes a window, every image that did not change is correctly
left out, and nothing depends on anyone remembering that the recipe also bumped its redis.
Returns (windows, note) where windows is [(image-name, from, to)].
"""
if from_url is None:
from_url = _default_branch_compose(to_url)
if not from_url:
raise SystemExit("could not resolve a baseline compose; pass --compose-from explicitly")
new, old = _compose_images(to_url), _compose_images(from_url)
app, others = None, []
for svc, (repo, tag) in sorted(new.items()):
if svc not in old:
continue
prev_repo, prev_tag = old[svc]
if prev_tag == tag and prev_repo == repo:
continue
# The `app` service is the recipe's primary image by coop-cloud convention; its window drives
# --from/--to so the scan's primary advisory source is judged against it. Everything else is
# a sidecar window keyed by its image name.
if svc == "app":
app = (repo.split("/")[-1], prev_tag, tag)
else:
others.append((repo.split("/")[-1], prev_tag, tag))
wins = ([app] if app else []) + others
return wins, f"baseline {from_url}"
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("recipe")
@@ -1231,13 +828,6 @@ def main() -> int:
"fix version published), fetch their full text + references and append a "
"block for the agent to judge. Additive: it never changes the count above.")
ap.add_argument("--registry", default=REGISTRY_DIR)
ap.add_argument("--compose-to", default=None, metavar="URL",
help="derive the windows by DIFFING this compose against its baseline, instead "
"of passing --from/--to/--image by hand. Point it at a PR's compose.yml "
"(e.g. .../raw/branch/<pr-branch>/compose.yml).")
ap.add_argument("--compose-from", default=None, metavar="URL",
help="baseline compose for --compose-to. Default: the same repo's DEFAULT "
"branch, resolved from the API (never assumed to be `main`).")
ap.add_argument("--image", action="append", default=[], metavar="NAME=FROM:TO",
help="a sidecar image and the versions it moved between, e.g. "
"--image redis=7.4:8.10 (repeatable). NAME matches a source repo name; "
@@ -1245,20 +835,6 @@ def main() -> int:
"being left unclassified.")
a = ap.parse_args()
images = []
if a.compose_to:
wins, note = windows_from_compose(a.compose_to, a.compose_from)
if not wins:
print(f"### Advisory scan — {a.recipe}\n\n**No image versions changed between the two "
f"compose files, so this upgrade fixes no CVEs by definition.**\n\n_{note}_")
return 0
print(f"_derived from compose diff ({note}):_", file=sys.stderr)
for n_, f_, t_ in wins:
print(f"_ {n_}: {f_}{t_}_", file=sys.stderr)
# The `app` service (first entry when present) drives --from/--to; the rest are --image
# windows. Passing every window as --image too is harmless: each is matched by name against
# the advisory sources, and an unmatched name is simply ignored.
a.v_from, a.v_to = a.v_from or wins[0][1], a.v_to or wins[0][2]
images = list(wins[1:])
for spec in a.image:
name, _, rng = spec.partition('=')
vf, _, vt = rng.partition(':')
-314
View File
@@ -1,314 +0,0 @@
#!/usr/bin/env python3
"""audit-sources — are we still looking in the right place for each recipe's updates?
A recipe tracks an image repo and a set of registry URLs. Upstreams move: they rename the image,
switch registry, archive the GitHub repo, or split a community edition out of the original. When that
happens nothing errors — the old repo simply stops receiving tags, and the recipe looks "up to date"
forever while real releases happen somewhere else.
plausible is the worked example. It tracked `plausible/analytics` on Docker Hub; upstream moved to
`ghcr.io/plausible/community-edition`. The old repo still exists and still serves v2.0.0, so every
survey said "no upgrades available" while v3 shipped elsewhere.
This reports the signals that catch that, per image and per registry URL:
* IMAGE GONE QUIET — newest tag is older than --quiet-days (default 365). The single strongest
signal that releases moved somewhere else.
* DEPRECATION WORDING — the registry description says deprecated / moved / no longer maintained.
* GITHUB REPO ARCHIVED — upstream archived it.
* GITHUB REPO RENAMED — the API redirects to a different owner/name than we ask for.
* GITHUB REPO GONE — 404.
Everything is a SIGNAL, not a verdict: a genuinely stable image (mumble, custom-html) can be quiet
for good reason. The output is for a human to judge, so each finding says what was measured.
audit-sources.py [recipe ...] [--ssh HOST] [--quiet-days N] [--json]
"""
from __future__ import annotations
import argparse
import importlib.util
import json
import os
import re
import sys
import urllib.error
import urllib.request
from datetime import datetime, timezone
HERE = os.path.dirname(os.path.abspath(__file__))
_spec = importlib.util.spec_from_file_location("resolve_images", os.path.join(HERE, "resolve-images.py"))
RI = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(RI)
# advisory-scan supplies the source-fetching + changelog-attribution used by --security-sources
_aspec = importlib.util.spec_from_file_location("advisory_scan", os.path.join(HERE, "advisory-scan.py"))
A = importlib.util.module_from_spec(_aspec)
_aspec.loader.exec_module(A)
REGISTRY_DIR = os.environ.get("CCCI_UPSTREAM_REGISTRY", os.path.join(HERE, "upstream"))
USED_RECIPES = os.path.join(HERE, "used-recipes.md")
DEPRECATION_RE = re.compile(
r"\b(deprecat|no longer maintain|unmaintained|superseded|moved to|migrated to|"
r"has moved|discontinued|end.of.life|archived)\b", re.I)
def _days_since(iso: str | None) -> int | None:
if not iso:
return None
try:
d = datetime.fromisoformat(iso.replace("Z", "+00:00"))
except ValueError:
return None
return (datetime.now(timezone.utc) - d).days
def hub_repo_meta(repo: str) -> dict:
"""Docker Hub repo metadata: when it was last pushed to, and how it describes itself."""
try:
d = RI._json(f"https://hub.docker.com/v2/repositories/{repo}", RI._hub_auth())
except urllib.error.HTTPError as e:
return {"status": f"HTTP {e.code}"}
except Exception as e: # noqa: BLE001
return {"status": f"{type(e).__name__}"}
text = f"{d.get('description') or ''}\n{d.get('full_description') or ''}"
m = DEPRECATION_RE.search(text)
return {"status": "ok", "last_updated": d.get("last_updated"),
"deprecation_hint": (m.group(0) if m else None),
"archived": bool(d.get("is_archived") or d.get("status") == "inactive")}
def github_repo_meta(owner: str, repo: str) -> dict:
"""GitHub repo state — archived, renamed (the API answers with the CURRENT full_name), or gone."""
hdrs = {"Accept": "application/vnd.github+json"}
tok = RI._gh_token()
if tok:
hdrs["Authorization"] = f"Bearer {tok}"
try:
d = RI._json(f"https://api.github.com/repos/{owner}/{repo}", hdrs)
except urllib.error.HTTPError as e:
return {"status": f"HTTP {e.code}"}
except Exception as e: # noqa: BLE001
return {"status": f"{type(e).__name__}"}
asked, got = f"{owner}/{repo}".lower(), (d.get("full_name") or "").lower()
return {"status": "ok", "archived": bool(d.get("archived")), "pushed_at": d.get("pushed_at"),
"renamed_to": (d.get("full_name") if got and got != asked else None),
"description": d.get("description") or ""}
def newest_tag_date(registry: str, repo: str, tag: str) -> str | None:
"""When was the repo's newest same-shape tag pushed? Docker Hub only (it dates its tags)."""
if registry not in ("docker.io", "registry-1.docker.io"):
return None
try:
d = RI._json(f"https://hub.docker.com/v2/repositories/{repo}/tags"
f"?page_size=100&ordering=last_updated", RI._hub_auth())
except Exception: # noqa: BLE001
return None
want = RI.shape(tag)
for row in d.get("results", []):
if RI.shape(row.get("name") or "") == want:
return row.get("last_updated")
return (d.get("results") or [{}])[0].get("last_updated")
def security_source_audit(recipe: str) -> list[dict]:
"""Per source: are its CVEs USABLE, or merely visible?
The nginx lesson. nginx publishes no GitHub advisories; all its CVEs live in nginx.org/en/CHANGES.
The scan saw them and could do nothing with them, because nothing said which release fixed which
CVE — so every nginx bump in the fleet reported 0. Attribution (advisory-scan §2b) fixed that for
changelogs organised by release, but a page that lists CVEs with NO release structure is still a
blind spot: visible, uncountable. This finds those.
Per source: `advisory-feed` (structured, best), `changelog` (CVEs attributable to a release),
`unattributable` (CVEs present but no release structure — BLIND), or `no-cve-data`.
"""
urls, _ = _registry_urls(recipe)
out = []
# NVD CPE entries are a first-class source: for projects publishing nothing machine-readable
# (mattermost, mumble) they are the ONLY structured source, and omitting them here made two
# recipes look permanently blind after they had been fixed.
for key, cpe in A.registry_cpes(recipe, REGISTRY_DIR):
e = A.nvd_advisories(cpe, key)
n = len(e.get("advisories") or [])
out.append({"source": e["source"] + f" ({cpe.split(':')[4]}/{cpe.split(':')[3]})",
"kind": "advisory-feed" if n else "no-cve-data",
"status": e["status"], "cves": n, "usable": n})
for entry in A.github_advisories(urls):
out.append({"source": entry["source"], "kind": "advisory-feed",
"status": entry["status"], "cves": len(entry.get("advisories") or []),
"usable": len(entry.get("advisories") or [])})
for entry in A.vendor_pages(urls):
if entry["status"].startswith("skipped"):
continue
n = len(entry.get("cves") or [])
attributed = len(entry.get("fixed_in") or {})
kind = ("no-cve-data" if n == 0 else
"changelog" if attributed else "unattributable")
out.append({"source": entry["source"], "kind": kind, "status": entry["status"],
"cves": n, "usable": attributed})
return out
def audit_recipe(recipe: str, ssh: str | None, quiet_days: int) -> dict:
out = {"recipe": recipe, "findings": [], "images": [], "sources": []}
try:
refs = (RI.compose_images_ssh(recipe, ssh, "~/.abra/recipes") if ssh
else RI.compose_images(recipe, RI.RECIPE_DIR))
except Exception as e: # noqa: BLE001
out["findings"].append({"level": "error", "what": f"could not read compose: {e}"})
return out
for ref in refs:
if "${" in ref:
continue
info = RI.parse_ref(ref)
row = {"ref": ref, "registry": info["registry"], "repo": info["repo"], "tag": info["tag"]}
if info["registry"] in ("docker.io", "registry-1.docker.io"):
meta = hub_repo_meta(info["repo"])
row.update(meta)
newest = newest_tag_date(info["registry"], info["repo"], info["tag"])
row["newest_tag_pushed"] = newest
age = _days_since(newest)
row["newest_tag_age_days"] = age
if age is not None and age > quiet_days:
out["findings"].append({
"level": "warn", "what": "image has gone quiet",
"detail": f"{info['repo']}: newest {RI.shape(info['tag'])}-shaped tag pushed "
f"{age} days ago — releases may have moved elsewhere"})
if meta.get("deprecation_hint"):
out["findings"].append({
"level": "warn", "what": "registry text suggests deprecation",
"detail": f"{info['repo']}: says {meta['deprecation_hint']!r}"})
if meta.get("archived"):
out["findings"].append({"level": "warn", "what": "registry repo archived/inactive",
"detail": info["repo"]})
out["images"].append(row)
urls, reg_path = ([], None)
try:
urls, reg_path = _registry_urls(recipe)
except Exception: # noqa: BLE001
pass
if reg_path is None:
out["findings"].append({"level": "warn", "what": "no upstream registry file",
"detail": f"cc-ci-plan/upstream/{recipe}.md is missing — the advisory "
f"scan has nowhere to look"})
seen = set()
for u in urls:
m = re.match(r"https?://github\.com/([^/]+)/([^/#?]+)", u)
if not m:
continue
owner, repo = m.group(1), m.group(2).removesuffix(".git")
if (owner, repo) in seen:
continue
seen.add((owner, repo))
meta = github_repo_meta(owner, repo)
row = {"repo": f"{owner}/{repo}", **meta}
age = _days_since(meta.get("pushed_at"))
row["pushed_age_days"] = age
out["sources"].append(row)
if meta.get("status") != "ok":
out["findings"].append({"level": "warn", "what": "registry source unreachable",
"detail": f"{owner}/{repo}: {meta['status']}"})
continue
if meta.get("renamed_to"):
out["findings"].append({"level": "alert", "what": "GitHub repo has MOVED",
"detail": f"{owner}/{repo} now answers as {meta['renamed_to']}"})
if meta.get("archived"):
out["findings"].append({"level": "alert", "what": "GitHub repo is ARCHIVED",
"detail": f"{owner}/{repo} — upstream development has stopped here"})
if age is not None and age > quiet_days:
out["findings"].append({"level": "warn", "what": "GitHub repo quiet",
"detail": f"{owner}/{repo}: last push {age} days ago"})
return out
def _registry_urls(recipe: str):
path = os.path.join(REGISTRY_DIR, f"{recipe}.md")
if not os.path.exists(path):
return [], None
text = open(path).read()
urls = []
for u in re.findall(r"https?://[^\s)|\]]+", text):
u = u.rstrip("`'\"*.,;:>)")
if u and u not in urls:
urls.append(u)
return urls, path
def all_recipes() -> list[str]:
out = []
for ln in open(USED_RECIPES):
ln = ln.strip()
if not ln or ln.startswith("#") or ln.startswith("`"):
continue
parts = ln.split()
if len(parts) >= 2 and parts[1] in ("weekly", "external"):
out.append(parts[0])
return out
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("recipes", nargs="*")
ap.add_argument("--ssh", default=None)
ap.add_argument("--quiet-days", type=int, default=365)
ap.add_argument("--json", action="store_true")
ap.add_argument("--security-sources", action="store_true",
help="audit whether each recipe's CVE sources are USABLE (structured advisory "
"feed / release-attributable changelog) or merely visible")
a = ap.parse_args()
recipes = a.recipes or all_recipes()
if a.security_sources:
# What matters is whether the RECIPE can see CVEs at all — not whether some individual page
# is unparseable. A page with no release structure is harmless when the same project also
# publishes an advisory feed (redis, gitea, minio, clickhouse all do); it is only a blind
# spot when nothing else covers that project.
blind_recipes, noisy = [], 0
for r in recipes:
rows = security_source_audit(r)
feeds = [x for x in rows if x["kind"] == "advisory-feed" and x["cves"] > 0]
logs = [x for x in rows if x["kind"] == "changelog"]
unattr = [x for x in rows if x["kind"] == "unattributable"]
noisy += len(unattr)
usable = len(feeds) + len(logs)
if usable == 0:
blind_recipes.append(r)
print(f"!! {r}: NO USABLE CVE SOURCE — {len(unattr)} unparseable page(s), "
f"0 advisory feeds, 0 attributable changelogs")
for x in rows:
print(f" {x['kind']:15} {x['source'][:64]} ({x['cves']} CVEs)")
else:
print(f"OK {r}: {len(feeds)} advisory-feed(s), {len(logs)} changelog(s)"
+ (f", {len(unattr)} unparseable page(s) (redundant — covered by a feed)"
if unattr else ""))
for x in logs:
print(f" changelog {x['source'][:62]} ({x['usable']}/{x['cves']})")
print(f"\n{len(recipes)} recipes · {len(blind_recipes)} with NO usable CVE source"
+ (f": {', '.join(blind_recipes)}" if blind_recipes else "")
+ f" · {noisy} unparseable page(s) elsewhere (harmless where a feed covers them)")
return 0
reports = [audit_recipe(r, a.ssh, a.quiet_days) for r in recipes]
if a.json:
print(json.dumps(reports, indent=2))
return 0
alerts = 0
for rep in reports:
fs = rep["findings"]
mark = "OK " if not fs else ("!! " if any(f["level"] == "alert" for f in fs) else " ? ")
print(f"{mark} {rep['recipe']}")
for f in fs:
alerts += f["level"] == "alert"
print(f" [{f['level']}] {f['what']}: {f.get('detail','')}")
print(f"\n{len(reports)} recipes audited · "
f"{sum(len(r['findings']) for r in reports)} findings · {alerts} alerts")
return 0
if __name__ == "__main__":
sys.exit(main())
-231
View File
@@ -1,231 +0,0 @@
#!/usr/bin/env python3
"""pr-survey — deterministic facts about every open recipe PR, for /cc-ci-cleanup to judge.
Open recipe PRs rot in specific, detectable ways. This gathers the evidence; it does NOT decide
anything — closing a PR is a judgement the skill makes, with these facts in hand.
RUN `reconcile-upstream.sh --all` FIRST. Every signal below is measured against the mirror's `main`,
and an unreconciled mirror makes all of them wrong: on 2026-08-11 three PRs (discourse #6 carrying
140 CVEs, keycloak #6 carrying 12, n8n #5) looked pending against a stale mirror while upstream had
already merged them. This tool refuses to guess about that — see `reconciled_recently`.
Per PR:
behind_main commits on main not in the branch — the "out of date" measure
ahead commits on the branch not on main
mergeable gitea's own verdict (false = conflicts, needs a rebase)
diff_files files the PR touches (0 = nothing left to merge)
adds_images the `+ image:` lines it introduces
already_in_main those `+ image:` lines ALREADY present in main -> the bump landed another way
obsolete true when every image it adds is already in main (nothing to contribute)
ci newest `!testme` verdict + build number parsed from the PR comments
branch_kind upgrade / fix / ci-artifact (`ci/*` sweep + probe branches) / other
age_days, stale_days (since last update)
pr-survey.py [recipe ...] [--json]
"""
from __future__ import annotations
import argparse
import base64
import json
import os
import re
import sys
import urllib.error
import urllib.parse
import urllib.request
from datetime import datetime, timezone
HERE = os.path.dirname(os.path.abspath(__file__))
USED_RECIPES = os.path.join(HERE, "used-recipes.md")
TESTENV = os.environ.get("CCCI_TESTENV", "/srv/cc-ci/.testenv")
NS = "recipe-maintainers"
def _env() -> dict:
e = {}
try:
for ln in open(TESTENV):
ln = ln.strip()
if "=" in ln and not ln.startswith("#"):
k, v = ln.split("=", 1)
e[k] = v.strip().strip('"').strip("'")
except OSError:
pass
return e
ENV = _env()
GITEA = os.environ.get("GITEA_URL") or ENV.get("GITEA_URL", "git.autonomic.zone")
_AUTH = base64.b64encode(
f"{os.environ.get('GITEA_USERNAME') or ENV.get('GITEA_USERNAME','')}:"
f"{os.environ.get('GITEA_PASSWORD') or ENV.get('GITEA_PASSWORD','')}".encode()
).decode()
def _get(path: str, raw: bool = False):
req = urllib.request.Request(
f"https://{GITEA}{path}",
headers={"Authorization": f"Basic {_AUTH}", "User-Agent": "cc-ci-pr-survey"},
)
with urllib.request.urlopen(req, timeout=60) as r:
body = r.read()
return body.decode(errors="replace") if raw else json.loads(body)
def _days(iso: str | None) -> int | None:
if not iso:
return None
try:
d = datetime.fromisoformat(iso.replace("Z", "+00:00"))
except ValueError:
return None
return (datetime.now(timezone.utc) - d).days
def _branch_kind(ref: str) -> str:
if ref.startswith("ci/"):
return "ci-artifact" # regall/cfold sweeps + testme probes; never meant to merge
if ref.startswith("upgrade"):
return "upgrade"
if re.match(r"^(fix|feat|chore|revert)", ref):
return "fix"
return "other"
def _main_images(recipe: str) -> set[str]:
"""Image refs pinned on the mirror's main — the baseline a PR is judged against."""
out = set()
for f in ("compose.yml",):
try:
txt = _get(f"/{NS}/{recipe}/raw/branch/main/{f}", raw=True)
except Exception: # noqa: BLE001
continue
for m in re.finditer(r"^\s*image:\s*[\"']?([^\"'\s]+)", txt, re.M):
out.add(m.group(1))
return out
def _ci_verdict(recipe: str, number: int) -> dict:
"""Newest cc-ci !testme outcome recorded on the PR."""
try:
cs = _get(f"/api/v1/repos/{NS}/{recipe}/issues/{number}/comments?limit=100")
except Exception: # noqa: BLE001
return {"verdict": "unknown", "build": None}
for c in reversed(cs):
b = c.get("body") or ""
if "cc-ci:testme" not in b:
continue
m = re.search(r"/cc-ci/(\d+)", b)
if "" in b or "passed" in b:
return {"verdict": "passed", "build": m.group(1) if m else None}
if "" in b or "failure" in b:
return {"verdict": "failed", "build": m.group(1) if m else None}
if "" in b or "in progress" in b:
return {"verdict": "running", "build": m.group(1) if m else None}
return {"verdict": "never-run", "build": None}
def survey_pr(recipe: str, pr: dict, main_images: set[str]) -> dict:
n = pr["number"]
head = pr["head"]["ref"]
row = {
"recipe": recipe, "number": n, "title": pr.get("title", ""), "head": head,
"url": pr.get("html_url"), "branch_kind": _branch_kind(head),
"age_days": _days(pr.get("created_at")), "stale_days": _days(pr.get("updated_at")),
"mergeable": pr.get("mergeable"),
}
try:
row["behind_main"] = _get(
f"/api/v1/repos/{NS}/{recipe}/compare/{urllib.parse.quote(head, safe='')}...main"
).get("total_commits", 0)
row["ahead"] = _get(
f"/api/v1/repos/{NS}/{recipe}/compare/main...{urllib.parse.quote(head, safe='')}"
).get("total_commits", 0)
except Exception: # noqa: BLE001
row["behind_main"], row["ahead"] = None, None
# A FAILED diff fetch must never look like an empty diff: gitea#4 404s on .diff (force-pushed
# branch) and would otherwise be flagged EMPTY-DIFF and closed — while being a verified, green,
# needed fix. Unknown is its own state.
diff = None
try:
body = _get(f"/{NS}/{recipe}/pulls/{n}.diff", raw=True)
if body.lstrip().startswith(("diff --git", "From ")) or not body.strip():
diff = body
except Exception: # noqa: BLE001
diff = None
row["diff_files"] = None if diff is None else len(re.findall(r"^diff --git ", diff, re.M))
adds = re.findall(r"^\+\s*image:\s*[\"']?([^\"'\s]+)", diff or "", re.M)
row["adds_images"] = sorted(set(adds))
row["already_in_main"] = sorted({i for i in set(adds) if i in main_images})
# Nothing left to contribute: it touches files but every image it introduces is already pinned.
# Only claim obsolete when the diff was actually READ. No diff, no verdict.
row["obsolete"] = diff is not None and bool(adds) and set(adds).issubset(main_images)
row["ci"] = _ci_verdict(recipe, n)
return row
def all_recipes() -> list[str]:
out = []
for ln in open(USED_RECIPES):
p = ln.split()
if len(p) >= 2 and not ln.startswith(("#", "`")) and p[1] in ("weekly", "external"):
out.append(p[0])
return out
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("recipes", nargs="*")
ap.add_argument("--json", action="store_true")
a = ap.parse_args()
rows = []
for r in (a.recipes or all_recipes()):
try:
prs = _get(f"/api/v1/repos/{NS}/{r}/pulls?state=open&limit=50")
except urllib.error.HTTPError as e:
rows.append({"recipe": r, "error": f"HTTP {e.code}"})
continue
if not prs:
continue
mi = _main_images(r)
for pr in prs:
rows.append(survey_pr(r, pr, mi))
if a.json:
print(json.dumps(rows, indent=2))
return 0
print(f"{len(rows)} open PR(s)\n")
for x in sorted(rows, key=lambda z: (z.get("recipe", ""), z.get("number", 0))):
if x.get("error"):
print(f" {x['recipe']}: {x['error']}")
continue
flags = []
if x["obsolete"]:
flags.append("OBSOLETE(images already in main)")
if x["branch_kind"] == "ci-artifact":
flags.append("CI-ARTIFACT")
if x["diff_files"] == 0:
flags.append("EMPTY-DIFF")
if x["diff_files"] is None:
flags.append("DIFF-UNREADABLE(do not close on this)")
if x["mergeable"] is False:
flags.append("CONFLICTS")
if (x["behind_main"] or 0) > 0:
flags.append(f"BEHIND-{x['behind_main']}")
print(f" {x['recipe']}#{x['number']:<3} {x['title'][:52]}")
print(f" {x['branch_kind']:12} age={x['age_days']}d idle={x['stale_days']}d "
f"ci={x['ci']['verdict']}({x['ci']['build'] or '-'}) files={x['diff_files'] if x['diff_files'] is not None else '?'}")
if x["adds_images"]:
print(f" adds: {', '.join(i.split('/')[-1] for i in x['adds_images'][:4])}")
if flags:
print(f" >> {' | '.join(flags)}")
return 0
if __name__ == "__main__":
sys.exit(main())
+14 -45
View File
@@ -9,11 +9,7 @@ Subcommands (the /recipe-report agent runs them around its own review/classifica
survey [DATE] JSON of the run + every recipe's open PRs + CI verdict + per-recipe upgrade
notes (breaking-change/CVE analysis), and the /upgrade-all summary.
render SPEC.json OUT.html render the agent's report spec -> a self-contained newspaper HTML page
publish OUT.html DATE [KIND] copy to cc-ci:/var/lib/cc-ci-reports/<KIND>-DATE.html and regen the
archive index. KIND is `week` (default, the weekly /recipe-report) or `cve`
(a /cve-check advisory sweep). BOTH kinds appear in the SAME archive index,
newest first, each row suffixed "full" or "CVE check"; the distinct
filename prefix just stops a sweep overwriting a weekly edition.
publish OUT.html DATE copy to cc-ci:/var/lib/cc-ci-reports/week-DATE.html and regen the archive index
Page order: short lead → the full wire table (priority-sorted, CVEs column) → Addendum → Security
Bulletin → per-recipe "What changed".
@@ -48,10 +44,6 @@ LOGDIR = "/srv/cc-ci/.cc-ci-logs"
TESTENV = "/srv/cc-ci/.testenv"
INFRA = {"cc-ci", "cc-ci-orchestrator", "cc-ci-secrets"}
HOST_REPORTS = "/var/lib/cc-ci-reports"
# Both kinds live in ONE archive, distinguished by a suffix on a common title.
# prefix -> (page title, index label)
KINDS = {"week": ("The Recipe Report", "Week of {d} — full"),
"cve": ("The Recipe Report — CVE check", "{d} — CVE check")}
def _env():
@@ -223,16 +215,8 @@ def _table(rows, repo_url=None):
if repo_url and r.get("recipe") in repo_url:
name = f'<a href="{repo_url[r["recipe"]]}">{name}</a>'
cve = r.get("cve")
# "?" = advisory scan absent or had failed sources → count NOT authoritative. Per the
# /recipe-report guardrail this must NEVER render as "none" (a blank-that-reads-clean is
# exactly how two CVSS-9.8 gitea RCEs were misreported as "none" on 2026-08-07). A positive
# int is the confirmed CVE count; 0/omit is a confirmed-clean scan.
if isinstance(cve, str) and cve.strip() == "?":
cve_cell = '<span class="muted" title="advisory scan incomplete or absent — CVE count unknown">?</span>'
elif isinstance(cve, (int, float)) and cve:
cve_cell = f'<span class="cve">{int(cve)}</span>'
else:
cve_cell = '<span class="muted">none</span>'
cve_cell = (f'<span class="cve">{int(cve)}</span>' if isinstance(cve, (int, float)) and cve
else '<span class="muted">none</span>')
ci = _esc(r.get("ci"))
if r.get("ci_url"):
ci = f'<a href="{_esc(r["ci_url"])}">{ci}</a>'
@@ -286,10 +270,8 @@ def _mast():
def render(spec_path, out_path):
s = json.load(open(spec_path))
kind = s.get("kind", "week")
title = KINDS.get(kind, KINDS["week"])[0]
gen = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC")
sub = s.get("subtitle", ("Week of " if kind == "week" else "CVE check ") + s["date"])
sub = s.get("subtitle", "Week of " + s["date"])
lead = s.get("lead", "") or ""
# Auto-link recipe-name mentions in the lead to their mirror repos.
gitea = _env().get("GITEA_URL", "git.autonomic.zone")
@@ -303,9 +285,7 @@ def render(spec_path, out_path):
f'<span>report.ci.commoninternet.net</span><span>{gen}</span></div>'
f'<div class="lead">{lead}</div>')
# 1) the full wire — every recipe, in the agent's recommended priority order (CVEs first); CVEs column.
wire = ("The full wire — every recipe, in priority order" if kind == "week"
else "Advisory sweep — every recipe, worst first")
body += f'<h2>{wire}</h2>{_table(s.get("table"), repo_url)}'
body += f'<h2>The full wire — every recipe, in priority order</h2>{_table(s.get("table"), repo_url)}'
# 2) addendum — special issues to look into (normal-size header); omitted entirely if there are none.
add = [a for a in (s.get("addendum") or []) if str(a).strip()]
if add:
@@ -318,30 +298,19 @@ def render(spec_path, out_path):
# 4) what changed — a short section per recipe that has a PR
if s.get("changes"):
body += f'<h2>What changed</h2>{_changes(s.get("changes"), repo_url)}'
body += (f'<footer>{title} · generated {gen} · '
body += (f'<footer>The Recipe Report · generated {gen} · '
f'<a href="https://ci.commoninternet.net/">dashboard</a> · <a href="./">archive</a></footer>')
open(out_path, "w").write(_page(f"{title} · " + s["date"], body))
open(out_path, "w").write(_page("The Recipe Report — " + s["date"], body))
print("wrote", out_path)
def publish(html_path, date, kind="week"):
if kind not in KINDS:
print(f"unknown kind {kind!r}; expected one of {', '.join(KINDS)}"); sys.exit(2)
page = f"{kind}-{date}.html"
def publish(html_path, date):
page = f"week-{date}.html"
subprocess.run(["ssh", "cc-ci", f"cat > {HOST_REPORTS}/{page}"], input=open(html_path, "rb").read(), check=True)
# One index over BOTH families, newest first, each row labelled by its kind — an operator looking
# for "the latest security picture" should not have to know which skill produced which page.
entries = []
for k in KINDS:
listing = subprocess.run(["ssh", "cc-ci", f"ls -1 {HOST_REPORTS}/{k}-*.html 2>/dev/null"],
capture_output=True, text=True).stdout.split()
for pth in listing:
d = os.path.basename(pth)[len(k) + 1:-5]
if re.fullmatch(r"\d{4}-\d{2}-\d{2}", d):
entries.append((d, k))
lis = "\n".join(
f'<li><a href="{k}-{d}.html">{KINDS[k][1].format(d=d)}</a><span class="d">{d}</span></li>'
for d, k in sorted(set(entries), reverse=True))
listing = subprocess.run(["ssh", "cc-ci", f"ls -1 {HOST_REPORTS}/week-*.html 2>/dev/null"],
capture_output=True, text=True).stdout.split()
dates = sorted({os.path.basename(p)[5:-5] for p in listing}, reverse=True)
lis = "\n".join(f'<li><a href="week-{d}.html">Week of {d}</a><span class="d">{d}</span></li>' for d in dates)
idx = _page("The Recipe Report — Archive", _mast() +
'<div class="dateline"><span>Weekly review of Co-op Cloud recipe upgrades &amp; CI</span>'
'<span>report.ci.commoninternet.net</span></div>'
@@ -359,7 +328,7 @@ def main():
elif cmd == "render":
render(a[1], a[2])
elif cmd == "publish":
publish(a[1], a[2], a[3] if len(a) > 3 else "week")
publish(a[1], a[2])
else:
print(__doc__); sys.exit(2)
-64
View File
@@ -1,64 +0,0 @@
#!/usr/bin/env bash
# reconcile-upstream — sync recipe mirrors from TRUE upstream. Run this FIRST, always.
# ----------------------------------------------------------------------------------
# Every recipe we maintain is a MIRROR of a coopcloud recipe. Work done against a stale
# mirror is wasted or wrong, in three ways we have actually hit:
#
# 1. A PR whose changes upstream ALREADY MERGED. mailu #6 (2024.06.57 + redis 8.10,
# two internet-facing Roundcube CVEs) sat open and was reported as the fix for
# those CVEs — while upstream had merged and released it as 3.1.3+2024.06.57. The
# work was done; only our mirror was behind.
# 2. A survey that reads the stale mirror and reports "no upgrades available", so a
# recipe silently drops out of the weekly run.
# 3. Reading the WRONG BRANCH. Several coopcloud recipes keep a stale `main` beside
# the real default `master` — gitea's `main` is at 1.24.2-rootless while `master`
# has 1.27.1-rootless plus the merged PRs. Reading `main` there says the recipe is
# three releases behind and missing two CVSS-9.8 RCE fixes, which reads exactly
# like a real finding. open-recipe-pr.sh resolves the default branch itself
# (main OR master) — never hand-pick one.
#
# This is deterministic: it force-syncs each mirror's `main` to upstream's default
# branch and closes any mirror PR whose changes are already upstream. No AI judgement.
#
# reconcile-upstream.sh <recipe>... # specific recipes
# reconcile-upstream.sh --all # every recipe in used-recipes.md
#
# Safe to run repeatedly; a mirror already in sync is a no-op. Recipe work lives in
# BRANCHES, never on mirror `main`, so force-syncing `main` discards nothing.
set -o errexit -o nounset -o pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
ORCH="$(dirname "$HERE")"
SSH="${SSH:-cc-ci}"
TESTENV="${TESTENV:-/srv/cc-ci/.testenv}"
RECONCILE="${RECONCILE:-$ORCH/.claude/skills/recipe-upgrade/open-recipe-pr.sh}"
USED_RECIPES="${USED_RECIPES:-$HERE/used-recipes.md}"
[ -f "$RECONCILE" ] || { echo "ERROR: reconcile helper not found: $RECONCILE" >&2; exit 1; }
set -a; . "$TESTENV"; set +a
: "${GITEA_USERNAME:?}"; : "${GITEA_PASSWORD:?}"; : "${GITEA_URL:?}"
if [ "${1:-}" = "--all" ]; then
mapfile -t RECIPES < <(awk '!/^[[:space:]]*#/ && ($2=="weekly" || $2=="external") {print $1}' "$USED_RECIPES")
else
[ "$#" -gt 0 ] || { echo "usage: reconcile-upstream.sh <recipe>... | --all" >&2; exit 2; }
RECIPES=("$@")
fi
synced=0; closed=0; failed=0
for r in "${RECIPES[@]}"; do
echo "── $r"
if out="$(ssh "$SSH" "GITEA_USERNAME='$GITEA_USERNAME' GITEA_PASSWORD='$GITEA_PASSWORD' GITEA_URL='$GITEA_URL' bash -s $r --reconcile-only" < "$RECONCILE" 2>&1)"; then
printf '%s\n' "$out" | grep -E "Force-syncing|already in sync|closed PR|still open|✓" | sed 's/^/ /' || true
synced=$((synced + 1))
closed=$((closed + $(printf '%s' "$out" | grep -c "closed PR" || true)))
else
printf '%s\n' "$out" | tail -3 | sed 's/^/ /'
echo " ✗ FAILED — do NOT proceed against this mirror until it reconciles"
failed=$((failed + 1))
fi
done
echo
echo "reconcile-upstream: ${synced} mirror(s) synced, ${closed} already-upstream PR(s) closed, ${failed} failed"
[ "$failed" -eq 0 ]
-459
View File
@@ -1,459 +0,0 @@
#!/usr/bin/env python3
"""resolve-images — what version is each of a recipe's images on, and what is newest?
An abra-independent version resolver. `abra recipe upgrade` is the normal path, but it has a hard
failure mode: an image pinned with BOTH a tag and a digest makes it FATA and abandon the WHOLE
recipe — even images it already parsed. immich pins two that way:
ghcr.io/immich-app/postgres:14-vectorchord0.4.3-pgvectors0.2.0@sha256:bcf6…
docker.io/valkey/valkey:9@sha256:3acc…
so immich contributes NO version data at all and silently drops out of every survey. That is
indistinguishable from "up to date" unless a human notices the missing row — which is exactly how it
kept getting skipped, and why a CVE sweep reported it as unknown.
This reads the compose files directly and queries the registries itself, so a digest pin is just a
digest pin. Output is JSON (default) or a table.
resolve-images.py <recipe> [--ssh HOST] [--recipe-dir DIR] [--table] [--only IMAGE]
The cc-ci host has no python3, so `--ssh cc-ci` reads the compose files from that host's checkout
over ssh and does the resolving locally. That keeps the source of truth the SAME tree abra and CI
use, rather than a second copy that can drift.
TAG SHAPES. Registries mix wildly different tag conventions in one repo, so "newest" is meaningless
without a shape. Each tag is reduced to a signature by replacing digit runs with '#':
v3.1.0 -> v#.#.#
1.27.1-rootless -> #.#.#-rootless
8.10-alpine -> #.#-alpine
14-vectorchord0.4.3-pgvectors0.2.0 -> #-vectorchord#.#.#-pgvectors#.#.#
Only tags sharing the CURRENT pin's shape are candidates. That keeps `-alpine` on `-alpine`, and
stops a `latest`/`release`/`sha-…` tag from ever being proposed as an upgrade.
TWO ANSWERS, NOT ONE. It reports `newest_same_shape` AND `newest_within_major` (same leading number).
For a plain app image they usually agree. For a compatibility-pinned sidecar they do not, and taking
the max would be wrong: immich's postgres tag encodes the pg major plus the vectorchord/pgvectors
versions that immich-server is built against, so jumping pg major because a newer tag exists breaks
the deployment. The caller picks; this tool refuses to guess and shows both.
"""
from __future__ import annotations
import argparse
import glob
import shlex
import subprocess
import gzip
import json
import os
import re
import sys
import time
import urllib.error
import urllib.parse
import urllib.request
UA = "cc-ci-resolve-images (+https://git.autonomic.zone/recipe-maintainers/cc-ci)"
TIMEOUT = int(os.environ.get("RESOLVE_IMAGES_TIMEOUT", "45"))
RECIPE_DIR = os.environ.get("ABRA_RECIPE_DIR", os.path.expanduser("~/.abra/recipes"))
MAX_TAG_PAGES = int(os.environ.get("RESOLVE_IMAGES_MAX_PAGES", "40"))
IMAGE_RE = re.compile(r"""^\s*image:\s*["']?([^"'\s]+)["']?\s*$""", re.M)
RETRIES = int(os.environ.get("RESOLVE_IMAGES_RETRIES", "4"))
def _fetch(url: str, headers: dict | None = None) -> bytes:
"""GET with backoff on rate limits.
Docker Hub throttles anonymous clients hard, and a sweep re-reads the same popular repos
(nginx, redis, postgres) for recipe after recipe. A 429 mid-sweep used to surface as
'unresolved', which is indistinguishable from a real lookup failure — so retry, and let the
per-repo cache below remove most of the requests entirely."""
h = {"User-Agent": UA, "Accept-Encoding": "gzip"}
h.update(headers or {})
delay = 2.0
for attempt in range(RETRIES):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=h), timeout=TIMEOUT) as r:
raw = r.read()
if r.headers.get("Content-Encoding") == "gzip":
raw = gzip.decompress(raw)
return raw
except urllib.error.HTTPError as e:
if e.code in (429, 503) and attempt < RETRIES - 1:
time.sleep(delay)
delay *= 2
continue
raise
raise RuntimeError("unreachable")
_HUB_JWT: list = []
def _hub_auth() -> dict:
"""Authenticated Docker Hub calls get a far higher rate limit than anonymous ones.
Credentials come from /srv/cc-ci/.testenv (DOCKERHUB_USERNAME / DOCKERHUB_TOKEN), the same pair
the CI host already uses. Absent creds are fine — the sweep just runs anonymous and slower."""
if _HUB_JWT:
return _HUB_JWT[0]
env = {}
try:
for ln in open(os.environ.get("CCCI_TESTENV", "/srv/cc-ci/.testenv")):
if "=" in ln and not ln.strip().startswith("#"):
k, v = ln.strip().split("=", 1)
env[k] = v.strip().strip("\"'")
except OSError:
pass
u = os.environ.get("DOCKERHUB_USERNAME") or env.get("DOCKERHUB_USERNAME")
t = os.environ.get("DOCKERHUB_TOKEN") or env.get("DOCKERHUB_TOKEN")
hdrs = {}
if u and t:
try:
body = json.dumps({"username": u, "password": t}).encode()
req = urllib.request.Request("https://hub.docker.com/v2/users/login",
data=body, method="POST",
headers={"Content-Type": "application/json", "User-Agent": UA})
with urllib.request.urlopen(req, timeout=TIMEOUT) as r:
tokj = json.load(r).get("token")
if tokj:
hdrs = {"Authorization": f"JWT {tokj}"}
except Exception: # noqa: BLE001 — anonymous is a valid fallback
hdrs = {}
_HUB_JWT.append(hdrs)
return hdrs
def _json(url: str, headers: dict | None = None):
return json.loads(_fetch(url, headers))
def shape(tag: str) -> str:
"""Signature of a tag with every digit run replaced by '#'. See module docstring."""
return re.sub(r"\d+", "#", tag)
def vkey(tag: str) -> tuple:
"""Ordering key: every number in the tag, in order. '1.27.10' > '1.27.9'; text ignored."""
return tuple(int(x) for x in re.findall(r"\d+", tag))
def parse_ref(ref: str) -> dict:
"""Split an image reference into registry / repo / tag / digest."""
digest = None
if "@" in ref:
ref, _, digest = ref.partition("@")
host, repo, tag = "docker.io", ref, "latest"
# A leading component is a REGISTRY only when there is a path after it. Without the slash test,
# a bare `postgres:15.18` looks like host "postgres:15.18" because of the tag's colon — which
# silently sent every library image to a nonexistent registry.
if "/" in ref:
first = ref.split("/")[0]
if "." in first or ":" in first or first == "localhost":
host, _, repo = ref.partition("/")
if ":" in repo.split("/")[-1]:
repo, _, tag = repo.rpartition(":")
if host == "docker.io" and "/" not in repo:
repo = f"library/{repo}" # bare `redis` is really `library/redis`
return {"registry": host, "repo": repo, "tag": tag, "digest": digest}
HUB_RECENT_PAGES = int(os.environ.get("RESOLVE_IMAGES_HUB_PAGES", "10"))
def _hub_tag_exists(repo: str, tag: str) -> bool:
try:
_json(f"https://hub.docker.com/v2/repositories/{repo}/tags/{tag}", _hub_auth())
return True
except Exception: # noqa: BLE001
return False
def _hub_tags(repo: str) -> list[str]:
"""Recently-pushed tags, newest first.
Popular Docker Hub repos carry many thousands of tags, so a full enumeration is impractical —
but it is also unnecessary: a tag NEWER than the one we run must have been pushed AFTER it, so
ordering by last_updated and reading a bounded recent window is sufficient to find any upgrade.
(ghcr offers no ordering, which is why that path needs a different strategy.)
"""
tags, url = [], (f"https://hub.docker.com/v2/repositories/{repo}/tags"
f"?page_size=100&ordering=last_updated")
auth = _hub_auth()
for _ in range(HUB_RECENT_PAGES):
d = _json(url, auth)
tags += [r["name"] for r in d.get("results", [])]
url = d.get("next")
if not url:
break
return tags
def _oci_bearer(host: str, repo: str) -> dict:
"""Token for an OCI registry, discovered from its own auth challenge.
Registries do NOT share a token endpoint. ghcr answers at /token?scope=…&service=ghcr.io, but
lscr.io and dock.mau.dev advertise different realms, and assuming ghcr's shape made both 401 —
which then read as "could not resolve" rather than "asked the wrong URL". The registry tells us
where to go in its WWW-Authenticate header; use that."""
try:
urllib.request.urlopen(
urllib.request.Request(f"https://{host}/v2/{repo}/tags/list?n=1",
headers={"User-Agent": UA}), timeout=TIMEOUT)
return {} # no auth needed
except urllib.error.HTTPError as e:
if e.code != 401:
return {}
chal = e.headers.get("WWW-Authenticate", "") or ""
except Exception: # noqa: BLE001
return {}
if not chal.lower().startswith("bearer"):
return {}
parts = dict(re.findall(r'(\w+)="([^"]*)"', chal))
realm = parts.get("realm")
if not realm:
return {}
q = {"service": parts.get("service", host), "scope": parts.get("scope", f"repository:{repo}:pull")}
url = realm + ("&" if "?" in realm else "?") + urllib.parse.urlencode(q)
try:
tok = (_json(url) or {}).get("token") or (_json(url) or {}).get("access_token")
return {"Authorization": f"Bearer {tok}"} if tok else {}
except Exception: # noqa: BLE001
return {}
def _oci_tags(host: str, repo: str) -> list[str]:
"""Tags from any OCI/v2 registry, with challenge-derived auth and Link pagination.
ghcr paginates hard — immich-server has >40,000 tags — and a truncated listing silently hides
the newest release line, so follow the cursor and let the caller's integrity check catch a read
that never reached the current pin."""
hdrs = _oci_bearer(host, repo)
tags, url = [], f"https://{host}/v2/{repo}/tags/list?n=1000"
for _ in range(MAX_TAG_PAGES):
req = urllib.request.Request(url, headers={"User-Agent": UA, **hdrs})
with urllib.request.urlopen(req, timeout=TIMEOUT) as r:
tags += (json.load(r) or {}).get("tags") or []
link = r.headers.get("Link", "") or ""
m = re.search(r'<([^>]+)>;\s*rel="next"', link)
if not m:
break
nxt = m.group(1)
url = f"https://{host}{nxt}" if nxt.startswith("/") else nxt
return tags
def _gh_token() -> str | None:
tok = os.environ.get("GITHUB_TOKEN")
if tok:
return tok.strip()
try:
return open(os.environ.get("GITHUB_TOKEN_FILE", "/srv/cc-ci/.github-token")).read().strip() or None
except OSError:
return None
def github_release_tags(owner: str, repo: str, max_pages: int = 4) -> list[str]:
"""Release tag names for a GitHub repo, newest first.
FALLBACK for registries whose tag listing cannot be enumerated. ghcr has no ordering and no
server-side filter, and immich-machine-learning carries >40,000 tags — a full read is impractical
and a partial read silently hides the newest release line. The project's RELEASES are ordered,
small, and authoritative: container tags track them. (The GitHub Packages API would answer this
directly but needs a scoped token; this scan's token deliberately has none.)
"""
hdrs = {"Accept": "application/vnd.github+json"}
tok = _gh_token()
if tok:
hdrs["Authorization"] = f"Bearer {tok}"
out = []
for page in range(1, max_pages + 1):
try:
rows = _json(f"https://api.github.com/repos/{owner}/{repo}/releases"
f"?per_page=100&page={page}", hdrs)
except Exception: # noqa: BLE001
break
if not rows:
break
out += [r.get("tag_name") or "" for r in rows]
return [t for t in out if t]
def _release_fallback_repos(registry: str, repo: str) -> list[tuple[str, str]]:
"""Candidate GitHub repos whose releases track this image's tags."""
if "ghcr.io" not in registry:
return []
parts = repo.split("/")
if len(parts) < 2:
return []
owner, name = parts[0], parts[-1]
cands = [(owner, name)]
# ghcr.io/immich-app/immich-machine-learning is built from immich-app/immich.
if name.startswith(owner.split("-")[0]):
cands.append((owner, owner.split("-")[0]))
return cands
_TAG_CACHE: dict[tuple[str, str], tuple[list[str], str | None]] = {}
def list_tags(registry: str, repo: str) -> tuple[list[str], str | None]:
if (registry, repo) in _TAG_CACHE:
return _TAG_CACHE[(registry, repo)]
res = _list_tags_uncached(registry, repo)
_TAG_CACHE[(registry, repo)] = res
return res
def _list_tags_uncached(registry: str, repo: str) -> tuple[list[str], str | None]:
try:
return (_hub_tags(repo) if registry in ("docker.io", "registry-1.docker.io")
else _oci_tags(registry, repo)), None
except urllib.error.HTTPError as e:
return [], f"HTTP {e.code}"
except Exception as e: # noqa: BLE001
return [], f"{type(e).__name__}: {e}"
def resolve(ref: str) -> dict:
"""Current pin -> newest same-shape tag, and newest within the current major."""
if "${" in ref or "$(" in ref:
# The tag is a compose variable (ghost pins `ghost:${IMAGE_VERSION}-alpine`). Its real value
# lives in .env, not here. Report it as skipped, never as a failed lookup.
return {**parse_ref(ref), "ref": ref, "shape": None, "candidates": 0,
"newest_same_shape": None, "newest_within_major": None,
"upgrade_available": False,
"status": "skipped: templated ref (tag comes from a compose variable)"}
info = parse_ref(ref)
out = {**info, "ref": ref, "shape": shape(info["tag"]), "status": "ok",
"newest_same_shape": None, "newest_within_major": None, "candidates": 0,
"upgrade_available": False}
tags, err = list_tags(info["registry"], info["repo"])
if err:
out["status"] = f"error: {err}"
return out
out["tags_seen"] = len(set(tags))
# INTEGRITY CHECK: the tag we are currently running MUST appear in the listing. If it does not,
# the listing is incomplete and any "newest" derived from it is a guess — ghcr paginates to tens
# of thousands of tags and a truncated read silently hides whole release lines. immich's
# machine-learning image is pinned v3.1.0, which EXISTS, yet a short read reported v1.134.0 as
# newest; without this check that becomes a confident, wrong answer.
if info["tag"] not in set(tags):
# Docker Hub: the window is recency-ordered, so the pin being outside it just means the pin
# is old — which is fine, because anything NEWER is necessarily inside the window. Confirm
# the pin genuinely exists (so a typo is still caught) and carry on.
if info["registry"] in ("docker.io", "registry-1.docker.io") and _hub_tag_exists(info["repo"], info["tag"]):
out["source"] = f"docker-hub:recent-{HUB_RECENT_PAGES * 100}"
tags = list(tags) + [info["tag"]]
else:
for owner, name in _release_fallback_repos(info["registry"], info["repo"]):
rel = github_release_tags(owner, name)
if info["tag"] in rel:
tags = rel
out["source"] = f"github-releases:{owner}/{name}"
out["tags_seen"] = len(set(rel))
break
else:
out["status"] = ("error: tag listing incomplete — the current pin "
f"{info['tag']!r} is absent from {len(set(tags))} registry tags "
f"and from the project's GitHub releases")
return out
want, cur = out["shape"], vkey(info["tag"])
same = [t for t in set(tags) if shape(t) == want and vkey(t)]
out["candidates"] = len(same)
if not same:
# Not a failure: digest-only pins and `latest`/`stable` have no comparable siblings.
out["status"] = "no comparable tags (shape has no numeric siblings)"
return out
newest = max(same, key=vkey)
out["newest_same_shape"] = newest
if cur:
within = [t for t in same if vkey(t)[:1] == cur[:1]]
if within:
out["newest_within_major"] = max(within, key=vkey)
out["upgrade_available"] = bool(cur and vkey(newest) > cur)
return out
def compose_images_ssh(recipe: str, host: str, recipe_dir: str) -> list[str]:
"""Same as compose_images, but the recipe tree lives on another host (cc-ci has no python3)."""
# NB: no shell-quoting of the directory — it may legitimately start with ~ or $HOME, and
# quoting it stops the remote shell expanding it, which yields an empty (and silent) result.
d = f"{recipe_dir}/{shlex.quote(recipe)}".replace("~", "$HOME")
cmd = (f'for f in {d}/compose*.yml; do case "$f" in *compose.ccci.yml) continue;; esac; '
f'[ -f "$f" ] && {{ cat "$f"; echo; }}; done; exit 0')
out = subprocess.run(["ssh", host, cmd], capture_output=True, text=True, timeout=120)
if out.returncode != 0:
raise RuntimeError(f"ssh {host}: {(out.stderr.strip() or 'no output')[:200]}")
if not out.stdout.strip():
raise RuntimeError(f"ssh {host}: no compose files found under {d}")
refs = []
for m in IMAGE_RE.finditer(out.stdout):
if m.group(1) not in refs:
refs.append(m.group(1))
return refs
def compose_images(recipe: str, recipe_dir: str) -> list[str]:
"""Every `image:` ref in the recipe's own compose files (the cc-ci overlay is NOT the recipe)."""
refs, base = [], os.path.join(recipe_dir, recipe)
for path in sorted(glob.glob(os.path.join(base, "compose*.yml"))):
if os.path.basename(path) == "compose.ccci.yml":
continue
try:
for m in IMAGE_RE.finditer(open(path).read()):
if m.group(1) not in refs:
refs.append(m.group(1))
except OSError:
continue
return refs
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("recipe")
ap.add_argument("--ssh", default=None, metavar="HOST",
help="read the recipe's compose files from HOST over ssh (e.g. --ssh cc-ci); "
"resolving still happens locally")
ap.add_argument("--recipe-dir", default=RECIPE_DIR)
ap.add_argument("--table", action="store_true", help="human-readable table instead of JSON")
ap.add_argument("--only", default=None, help="resolve just the images whose ref contains this")
a = ap.parse_args()
rdir = a.recipe_dir if a.recipe_dir != RECIPE_DIR or not a.ssh else "~/.abra/recipes"
refs = (compose_images_ssh(a.recipe, a.ssh, rdir) if a.ssh
else compose_images(a.recipe, a.recipe_dir))
if a.only:
refs = [r for r in refs if a.only in r]
results = [resolve(r) for r in refs]
report = {
"recipe": a.recipe,
"images": results,
"upgrades_available": [r["ref"] for r in results if r["upgrade_available"]],
"unresolved": [r["ref"] for r in results if r["status"].startswith("error")],
# The whole point: distinguish "checked, current" from "could not check".
"all_resolved": not any(r["status"].startswith("error") for r in results),
}
if not a.table:
print(json.dumps(report, indent=2))
return 0
print(f"{a.recipe}{len(results)} images")
for r in results:
flag = "UPGRADE" if r["upgrade_available"] else ("ERROR" if r["status"].startswith("error") else "current")
print(f" [{flag:7}] {r['repo']}:{r['tag']}" + (" (digest-pinned)" if r["digest"] else ""))
print(f" shape={r['shape']} candidates={r['candidates']}"
f" newest_same_shape={r['newest_same_shape']}"
f" newest_within_major={r['newest_within_major']}")
if r["status"] != "ok":
print(f" status: {r['status']}")
return 0
if __name__ == "__main__":
sys.exit(main())
+8 -228
View File
@@ -455,38 +455,13 @@ class TestReleaseNoteResolution(unittest.TestCase):
self.assertEqual(rep["indeterminate"], [])
self.assertEqual(rep["resolved_by_release_notes"]["CVE-TBD"], ["7.4.5", "8.0.3"])
def test_naming_releases_all_below_the_window_means_ALREADY_fixed(self):
# Every known fix predates the version we were already on, so this upgrade did not deliver
# it. That is a DECISION, not an unknown — mailu's redis 8.8.0 → 8.10.0 crosses 12 such
# advisories, and calling them "could not judge" overstates the uncertainty.
def test_release_naming_it_only_outside_the_window_stays_indeterminate(self):
rep = run_scan([gh("redis/redis", [adv("CVE-TBD", patched="TBD")])],
v_from="7.4", v_to="8.10", urls=["https://github.com/redis/redis"],
releases={"CVE-TBD": ["6.2.19"]})
self.assertEqual(rep["fixed_by_this_upgrade"], [])
self.assertEqual(rep["indeterminate"], [])
self.assertIn("CVE-TBD", rep["already_fixed_before_upgrade"])
self.assertIn("outside-window", rep["cves"]["CVE-TBD"]["classification"])
def test_naming_releases_only_ABOVE_the_window_stays_indeterminate(self):
# The fix landed after our target, so we are still exposed. Deliberately NOT decided as a
# tidy "not fixed": it is an open vulnerability and must stay visible to the operator.
rep = run_scan([gh("redis/redis", [adv("CVE-TBD", patched="TBD")])],
v_from="7.4", v_to="8.10", urls=["https://github.com/redis/redis"],
releases={"CVE-TBD": ["9.0.0"]})
self.assertEqual(rep["fixed_by_this_upgrade"], [])
self.assertIn("CVE-TBD", rep["indeterminate"])
def test_vendor_page_cve_on_the_same_repo_uses_release_notes(self):
# mailu announces its Roundcube CVEs only on github.com/Mailu/Mailu/releases. Requiring an
# advisory feed sent a deterministic case to pass 2; it is now decided in pass 1.
rep = run_scan([gh("Mailu/Mailu", [])],
[vendor("https://github.com/Mailu/Mailu/releases", ["CVE-2026-54432"])],
v_from="2024.06.55", v_to="2024.06.57",
urls=["https://github.com/Mailu/Mailu"],
releases={"CVE-2026-54432": ["2024.06.56"]})
self.assertIn("CVE-2026-54432", rep["fixed_by_this_upgrade"])
self.assertEqual(rep["cve_count_fixed"], 1)
def test_release_evidence_is_recorded_for_audit(self):
rep = run_scan([gh("redis/redis", [adv("CVE-TBD", patched="TBD")])],
v_from="7.4", v_to="8.10", urls=["https://github.com/redis/redis"],
@@ -503,195 +478,6 @@ class TestReleaseNoteResolution(unittest.TestCase):
self.assertNotIn("CVE-OK", rep.get("resolved_by_release_notes") or {})
class TestReleaseLineSemantics(unittest.TestCase):
"""A fix inside the numeric window is not a fix on the branch you actually land on."""
def test_fix_later_on_the_targets_own_line_is_not_counted(self):
# ClickHouse fixed CVE-2023-48704 in 23.9.6.20 AND 23.10.5.20. Landing on 23.10.4.25 crosses
# the 23.9 fix numerically but is BELOW its own line's fix, so it does not have it.
rep = run_scan([gh("ClickHouse/ClickHouse",
[adv("CVE-2023-48704", patched="v23.10.5.20; v23.9.6.20; v23.8.8.20")])],
v_from="23.4.2.11", v_to="23.10.4.25",
urls=["https://github.com/ClickHouse/ClickHouse"])
self.assertEqual(rep["fixed_by_this_upgrade"], [])
def test_fix_earlier_on_the_targets_own_line_is_counted(self):
rep = run_scan([gh("ClickHouse/ClickHouse",
[adv("CVE-2023-47118", patched="v23.10.2.13; v23.8.6.16")])],
v_from="23.4.2.11", v_to="23.10.4.25",
urls=["https://github.com/ClickHouse/ClickHouse"])
self.assertEqual(rep["fixed_by_this_upgrade"], ["CVE-2023-47118"])
def test_fix_exactly_at_the_target_is_counted(self):
rep = run_scan([gh("ClickHouse/ClickHouse",
[adv("CVE-2023-48298", patched="v23.10.4.25; v23.9.5.29")])],
v_from="23.4.2.11", v_to="23.10.4.25",
urls=["https://github.com/ClickHouse/ClickHouse"])
self.assertEqual(rep["fixed_by_this_upgrade"], ["CVE-2023-48298"])
def test_no_fix_on_the_target_line_falls_back_to_the_window(self):
# redis fixes 7.4.6/8.0.4/8.2.2 with no 8.10.x entry; landing on 8.10 still has them,
# because nothing on the 8.10 line is named as a LATER fix.
rep = run_scan([gh("redis/redis", [adv("CVE-2025-49844", patched="7.4.6; 8.0.4; 8.2.2")])],
v_from="7.4", v_to="8.10", urls=["https://github.com/redis/redis"])
self.assertEqual(rep["fixed_by_this_upgrade"], ["CVE-2025-49844"])
class TestAlreadyFixedOnFromLine(unittest.TestCase):
"""A fix that landed on the line we upgrade FROM was already ours before the upgrade."""
def test_backport_to_our_own_line_is_not_credited(self):
# mattermost patches every maintained line at once. 10.11.22 -> 10.12.4 crosses 10.12.1, but
# 10.11.22 is already past 10.11.4, so the deployment HAD the fix. Counting it credits the
# upgrade with work it did not do.
rep = run_scan([gh("mattermost/mattermost",
[adv("CVE-1", patched="10.11.4; 10.12.1; 10.5.12")])],
v_from="10.11.22", v_to="10.12.4",
urls=["https://github.com/mattermost/mattermost"])
self.assertEqual(rep["fixed_by_this_upgrade"], [])
def test_a_fix_ABOVE_our_position_on_the_same_line_still_counts(self):
rep = run_scan([gh("mattermost/mattermost", [adv("CVE-2", patched="10.11.30; 10.12.1")])],
v_from="10.11.22", v_to="10.12.4",
urls=["https://github.com/mattermost/mattermost"])
self.assertEqual(rep["fixed_by_this_upgrade"], ["CVE-2"])
def test_placeholders_never_feed_this_rule(self):
# "7.4.X" parses to a bare 7.4, which would read as "already fixed at 7.4" and silently drop
# a real fix — this is exactly how redis CVE-2024-46981 was lost when the rule was added.
rep = run_scan([gh("redis/redis", [adv("CVE-3", patched="6.2.X, 7.2.X, 7.4.X")])],
v_from="7.4", v_to="8.10", urls=["https://github.com/redis/redis"])
self.assertIn("CVE-3", rep["indeterminate"])
self.assertEqual(rep["fixed_by_this_upgrade"], [])
class TestChangelogAttribution(unittest.TestCase):
"""Projects that publish no advisory feed still say which release fixed what — in their changelog."""
CHANGES = """
Changes with nginx 1.31.3 11 Aug 2026
*) Security: a flaw ... (CVE-2026-60005)
*) Security: another ... (CVE-2026-56434)
Changes with nginx 1.31.2 04 Aug 2026
*) Security: something ... (CVE-2026-48142)
Changes with nginx 1.31.1 21 Jul 2026
*) Security: older ... (CVE-2026-9256)
Changes with nginx 1.20.0 01 Jan 2021
*) Security: ancient ... (CVE-2013-2028)
"""
def test_each_cve_is_attributed_to_the_release_that_fixed_it(self):
got = A._changelog_versions(self.CHANGES)
self.assertEqual(got["CVE-2026-60005"], "1.31.3")
self.assertEqual(got["CVE-2026-48142"], "1.31.2")
self.assertEqual(got["CVE-2026-9256"], "1.31.1")
self.assertEqual(got["CVE-2013-2028"], "1.20.0")
def _scan(self, wfrom, wto):
# nginx publishes NO GitHub advisories — the feed is empty and the changelog is everything.
return run_scan(
[gh("nginx/nginx", [])],
[{"source": "https://nginx.org/en/CHANGES", "status": "ok",
"cves": sorted(A._changelog_versions(self.CHANGES)),
"context": {}, "fixed_in": A._changelog_versions(self.CHANGES)}],
images=[("nginx", wfrom, wto)], urls=["https://github.com/nginx/nginx"])
def test_window_counts_only_the_releases_it_crosses(self):
rep = self._scan("1.31.1", "1.31.3") # 1.31.1 is the FROM, so its CVE is already fixed
self.assertEqual(set(rep["fixed_by_this_upgrade"]),
{"CVE-2026-48142", "CVE-2026-56434", "CVE-2026-60005"})
def test_a_narrower_window_counts_fewer(self):
rep = self._scan("1.31.2", "1.31.3")
self.assertEqual(set(rep["fixed_by_this_upgrade"]), {"CVE-2026-56434", "CVE-2026-60005"})
def test_ancient_entries_are_not_swept_in(self):
# The changelog lists the project's whole history; only the crossed releases may count.
rep = self._scan("1.31.1", "1.31.3")
self.assertNotIn("CVE-2013-2028", rep["fixed_by_this_upgrade"])
def test_evidence_is_recorded(self):
rep = self._scan("1.31.1", "1.31.3")
self.assertEqual(rep["resolved_by_changelog"]["CVE-2026-60005"], "1.31.3")
class TestComposeDerivedWindows(unittest.TestCase):
"""Windows read off a compose diff, so nobody has to remember which --image args an upgrade needs."""
OLD = """
services:
app:
image: "plausible/analytics:v2.0.0"
db:
image: pgautoupgrade/pgautoupgrade:18-alpine
plausible_events_db:
image: clickhouse/clickhouse-server:23.4.2.11-alpine
volumes:
data:
"""
NEW = """
services:
app:
image: "ghcr.io/plausible/community-edition:v3.2.1"
db:
image: pgautoupgrade/pgautoupgrade:18-alpine
plausible_events_db:
image: clickhouse/clickhouse-server:24.12-alpine
volumes:
data:
"""
def _windows(self, old=None, new=None):
pages = {"to": new if new is not None else self.NEW,
"from": old if old is not None else self.OLD}
with unittest.mock.patch.object(A, "_fetch", lambda u, h=None: pages["to" if "to" in u else "from"]), \
unittest.mock.patch.object(A, "_gitea_auth", lambda u: {}):
return A.windows_from_compose("http://x/to", "http://x/from")[0]
def test_app_service_leads_and_sidecars_follow(self):
w = self._windows()
self.assertEqual(w[0], ("community-edition", "v2.0.0", "v3.2.1"))
self.assertIn(("clickhouse-server", "23.4.2.11-alpine", "24.12-alpine"), w)
def test_unchanged_images_are_not_windows(self):
# pgautoupgrade is identical in both; inventing a window for it would be a false count.
self.assertNotIn("pgautoupgrade", [n for n, _, _ in self._windows()])
def test_a_changed_image_REPO_is_still_the_same_service(self):
# plausible/analytics -> ghcr.io/plausible/community-edition. Keyed by image repo this reads
# as one image vanishing and another appearing, and the app window is lost entirely.
w = self._windows()
self.assertTrue(any(n == "community-edition" and f == "v2.0.0" for n, f, _ in w))
def test_no_change_yields_no_windows(self):
self.assertEqual(self._windows(old=self.NEW, new=self.NEW), [])
def test_templated_tags_are_skipped(self):
new = self.NEW.replace('ghcr.io/plausible/community-edition:v3.2.1', 'ghost:${IMAGE_VERSION}')
self.assertNotIn("ghost", [n for n, _, _ in self._windows(new=new)])
class TestImageNameMatching(unittest.TestCase):
"""An image name and its advisory source rarely spell each other exactly."""
def test_matches_when_the_image_name_is_LONGER_than_the_source(self):
# clickhouse/clickhouse-server vs source ClickHouse/ClickHouse — one-directional matching
# dropped this window silently when the key came from a compose file.
rep = run_scan([gh("ClickHouse/ClickHouse", [adv("CVE-1", patched="23.10.2.13")])],
images=[("clickhouse-server", "23.4.2.11", "24.12")],
urls=["https://github.com/ClickHouse/ClickHouse"])
self.assertIn("github-advisories:ClickHouse/ClickHouse", rep["windows"])
self.assertEqual(rep["cve_count_fixed"], 1)
def test_matches_when_the_image_name_is_SHORTER_than_the_source(self):
rep = run_scan([gh("redis/redis", [adv("CVE-2", patched="7.4.1")])],
images=[("redis", "7.4", "8.10")], urls=["https://github.com/redis/redis"])
self.assertEqual(rep["cve_count_fixed"], 1)
class TestAdjudicationEvidenceAssembly(unittest.TestCase):
"""Pass 2's JUDGEMENT is a model's and not testable; what IS testable is what it gets shown."""
@@ -813,22 +599,16 @@ class TestHistoricReportNumbers(unittest.TestCase):
for cve in delta:
self.assertIn("redis", both["cves"][cve]["sources"][0])
def test_mailu_finds_the_roundcube_pair_without_an_agent(self):
# Published as 2 on 2026-08-07, but only because an agent read the release notes; the scan
# itself contributed 0. It now reaches 2 deterministically: the CVEs appear only on
# github.com/Mailu/Mailu/releases, and release 2024.06.56 (inside the window) names them.
def test_mailu_scan_finds_zero_and_says_so_knowingly(self):
# Published as 2 via the UNION with release-note reading; the scan's own contribution is 0,
# and 0 here must mean "checked, none", not "could not check".
rep = self._count("mailu", "2024.06.55", "2024.06.57", [("redis", "8.8.0", "8.10.0")])
self.assertEqual(rep["cve_count_fixed"], 2)
self.assertEqual(set(rep["fixed_by_this_upgrade"]), {"CVE-2026-54432", "CVE-2026-54433"})
# The redis bump fixes nothing new — every advisory it crosses was fixed at or before 8.6.3.
self.assertEqual(rep["cve_count_indeterminate"], 0)
self.assertEqual(rep["cve_count_fixed"], 0)
self.assertTrue(rep["count_known"])
def test_keycloak_26_7_0_to_26_7_1_is_12(self):
# Was 7 while only the GHSA feed was consulted. keycloak lists five more CVEs in the 26.7.1
# release notes' fixed-issues section that it never filed as advisories — the gitea pattern.
def test_keycloak_26_7_0_to_26_7_1_is_7(self):
rep = self._count("keycloak", "26.7.0", "26.7.1")
self.assertEqual(rep["cve_count_fixed"], 12)
self.assertEqual(len(rep.get("resolved_by_release_notes") or {}), 5)
self.assertEqual(rep["cve_count_fixed"], 7)
def _main():
-9
View File
@@ -75,12 +75,3 @@
recreate DB, reimport dump. `DROP DATABASE WITH (FORCE)` requires PostgreSQL 13+ — safe on postgres:15-alpine.
The previous inline-label approach (no restore hook) was a defect: raw PGDATA restore without a reload
was a silent no-op. Fixed in PR #2 (restore fix cherry-picked from PR #1 ci/pg-restore).
## NVD CPE fallback
This project publishes nothing machine-readable we can reach — no GitHub advisory feed,
no release-attributable changelog — so its CVE count was `?` (nothing measured). NVD is
CPE-indexed and carries structured version ranges, so it can answer where the vendor
cannot. It LAGS the vendor, so it is a fallback, never the primary source.
- nvd-cpe: mattermost-team-edition = cpe:2.3:a:mattermost:mattermost_server:*:*:*:*:*:*:*:*
- nvd-cpe: postgres = cpe:2.3:a:postgresql:postgresql:*:*:*:*:*:*:*:*
-29
View File
@@ -1,29 +0,0 @@
# Upstream sources — mumble
| service | image | source repo | releases / changelog |
|---------|-------|-------------|----------------------|
| app | mumblevoip/mumble-server | https://github.com/mumble-voip/mumble | https://github.com/mumble-voip/mumble/releases |
| web | rankenstein/mumble-web | https://github.com/rankenstein/mumble-web | https://github.com/rankenstein/mumble-web/releases |
## Standing notes
- This file was **missing entirely** until 2026-08-11. Without it the advisory scan had no source to
query, and still printed "0 identified by the deterministic scan" — which was then published as a
clean `0` in the 2026-08-11 CVE check. The scan now refuses to emit a count when it has no usable
source (it reports UNKNOWN), and `audit-sources.py` flags a missing registry file directly.
- `mumblevoip/mumble-server` tracks the upstream server releases and DOES publish GitHub security
advisories, so it is the recipe's primary CVE source.
- `rankenstein/mumble-web` is a **fork** of the original `Johni0702/mumble-web`, which has been
dormant since 2023-05. The fork itself last pushed 2023-07 and its Docker tag `0.5` was last built
well over five years ago. Neither is archived, but treat the web client as effectively unmaintained:
if a CVE lands there, expect no upstream fix and plan a replacement rather than an upgrade.
- The server image tag is `v<version>-<build>` (e.g. `v1.6.870-4`); the trailing number is the image
build, not an app version, and moves independently of upstream releases — `abra recipe upgrade`
reports "no new versions" for it, so use `resolve-images.py` to see those bumps.
## NVD CPE fallback
This project publishes nothing machine-readable we can reach — no GitHub advisory feed,
no release-attributable changelog — so its CVE count was `?` (nothing measured). NVD is
CPE-indexed and carries structured version ranges, so it can answer where the vendor
cannot. It LAGS the vendor, so it is a fallback, never the primary source.
- nvd-cpe: mumble-server = cpe:2.3:a:mumble:mumble:*:*:*:*:*:*:*:*