A recipe tracks an image repo and a set of registry URLs. When upstream moves,
nothing errors — the old repo just stops receiving tags and the recipe looks
'up to date' forever. plausible is the case: it tracked plausible/analytics on
Docker Hub while upstream moved to ghcr.io/plausible/community-edition. Every
survey said 'no upgrades available' while v3 shipped elsewhere.
audit-sources.py reports the signals that catch it, per image and per registry
URL: image gone quiet (newest tag older than --quiet-days), deprecation wording
in the registry description, and GitHub repos that are archived, renamed or
gone. Signals, not verdicts — a stable image can be quiet for good reason — so
each finding says what was measured.
First run over 22 recipes, 11 findings, 4 alerts. It independently re-derived
the plausible case (analytics quiet 1126 days), and found:
- drone: harness/drone now answers as harness/harness (the image is fine)
- lasuite-docs, lasuite-drive: minio/minio is ARCHIVED on GitHub
- lasuite-docs: docspecio/api is ARCHIVED
- matrix-synapse: halfshot/matrix-appservice-discord image quiet 2078 days
- mumble: NO cc-ci-plan/upstream/mumble.md at all
That last one exposed a scanner bug. With no registry file there is no source to
query, yet the scan still printed '0 identified by the deterministic scan' — and
that 0 was published as a clean count in the 2026-08-11 CVE check. A scan with no
usable source has measured nothing and must not report a number, least of all 0.
It now returns UNKNOWN and says the registry file is missing.
upstream/mumble.md added; mumble now scans 6 sources for a genuine 0.
immich pins two images with BOTH a tag and a digest, which makes abra FATA and
abandon the WHOLE recipe. It therefore contributed no version data at all and
silently dropped out of every survey — indistinguishable from 'up to date'. The
standing answer was prose in three skills telling an agent to check registries by
hand. This replaces it with a tool.
resolve-images.py reads the compose files and queries registries itself:
- Docker Hub, ghcr, and any OCI registry via its own auth challenge (lscr.io
and dock.mau.dev advertise different realms; assuming ghcr's shape 401'd).
- tag SHAPES (digits -> '#') so -alpine stays on -alpine and 'latest' is never
proposed as an upgrade.
- reports newest_within_major AND newest_same_shape, and refuses to choose:
immich's postgres tag encodes the pg major plus the vectorchord/pgvectors
build immich-server expects, so taking the newest breaks the deploy.
- integrity check: if the CURRENT pin is absent from the listing, the listing
was truncated and any 'newest' is a guess. ghcr caps out past 40k tags, so
that falls back to the project's GitHub releases.
- per-repo cache + backoff + Docker Hub auth: a fleet sweep re-reads nginx,
redis and postgres many times and was getting 429s reported as 'unresolved'.
21/21 recipes now resolve. It found upgrades abra missed entirely in five:
mumble (abra said 'no new versions'; four patches behind), plausible's
clickhouse, lasuite-drive's collabora, gitea's mariadb, immich's postgres.
plausible's carried four CVEs, three high.
Also fixes a real over-count found while validating that: a fix inside the
numeric window is not a fix on the branch you land on. ClickHouse patched
CVE-2023-48704 in 23.9.6.20 AND 23.10.5.20 — landing on 23.10.4.25 crosses the
23.9 fix but sits below its own line's, so it does NOT have it. A fix named on
the target's own line and above the target is now proof of absence.
70 tests (64 offline + 6 live). keycloak's live expectation moves 7 -> 12 and
mailu's 0 -> 2: both are the release-note source finding real fixes that were
never filed as advisories.
1. Release-note resolution now covers vendor pages on the same repo. It required
a github-advisories: source, so mailu's Roundcube CVEs — announced only on
github.com/Mailu/Mailu/releases — went to pass 2 even though the answer was
sitting in the release notes. mailu now reports 2 deterministically, matching
what previously took an agent reading the notes.
2. 'All known fix versions predate the version we were on' is now a DECISION,
not an unknown. mailu's redis 8.8.0 -> 8.10.0 crosses 12 advisories all fixed
by 8.6.3 or earlier; reporting them as 'could not judge' overstated the
uncertainty. Recorded as outside-window with the naming tags as evidence.
A fix landing ABOVE the window still stays indeterminate on purpose: that is
an open vulnerability and must stay visible.
60 offline tests (was 58). discourse 140 / gitea 2 unchanged.
Adds test-advisory-scan.py (58 offline tests on fixtures + 6 live regressions
against the week-2026-08-07 report) and audit-advisory-scan.py, which re-derives
every count with a SEPARATE semver implementation and its own release fetch and
diffs against the scanner. Both found real defects:
1. Window membership was compared on ragged tuples, so (18,) < (18,0) — a CVE
patched in 18.0 fell OUTSIDE a window ending at 18. Bare major tags are the
norm for sidecars (postgres:18, redis:8-alpine). Now zero-padded, which also
keeps the upper bound conservative (18.5 stays out of a window ending at 18).
2. Advisories with no knowable fix version were silently counted as 'not fixed'.
Twelve redis advisories say patched_versions 'TBD' or '7.4.X' with an
open-ended range — six of them high severity. They are now INDETERMINATE:
not counted, not dismissed, and surfaced in the output.
All twelve turned out to be genuinely fixed: redis names each in the release
notes of every branch that got the fix (CVE-2025-32023 -> 6.2.19, 7.2.10,
7.4.5, 8.0.3, 8.2.0). So a third deterministic method resolves them from
release notes, with the naming tags recorded as the citation. discourse's
redis contribution goes 5 -> 17, and its total 128 -> 140.
Pass 2 (--adjudicate) is the model-judged stage for what arithmetic cannot
settle: it hands over each open case's full evidence, plus every verdict pass 1
reached, and takes FIXED/NOT-FIXED/STILL-UNKNOWN with a reason citing that
evidence. It may only raise a count. Vendor-page-only CVEs — the shape of both
gitea CVSS-9.8 RCEs — now reach it instead of being dropped.
Tests cover pass 1 only, by design; pass 2's judgement is a model's. What is
tested there is deterministic: which cases it selects, and that truncation is
announced rather than silent.
SPEC.md rewritten around the two passes.
Restores the single-value form (operator preference) under the --image name.
Repeat the flag per image, all in one call. Malformed values warn on stderr and
are skipped rather than aborting the scan, since it is an additive pre-step.
Counts unchanged: discourse 128 with redis / 123 without, gitea 2.
The image name was packed into the value, so the flag needed a hand-rolled
KEY=FROM:TO parser with its own malformed-input branch, and 'window' named the
wrong thing — the tool has two kinds of window (version ranges and, on the date
fallback, real date windows) and the flag meant only the first.
Now each part is its own argument: --image redis 7.4 8.10, repeatable, all in
one call. argparse enforces the arity, so the string parsing and its error path
are deleted. 'windows' survives internally as the computed-range concept.
Counts unchanged: discourse 128 with redis / 123 without, gitea 2.
A recipe upgrades several images, each through its own version range. The scan
previously classified only the app repo, so sidecar bumps contributed nothing — the
alternative to the earlier bug where sidecars were judged by the APP's window and
produced a false 133.
Now: --window KEY=FROM:TO (repeatable) gives any other source its own range; each
window is classified independently (one may use patched-version ranges while another
falls back to advisory dates) and the count is the UNION. An image with no window is
still not counted — the scan will not guess a range it was not given. If ANY requested
window cannot be ordered, the total is UNKNOWN rather than a partial number.
/recipe-upgrade now instructs passing a --window per bumped sidecar.
Verified on discourse app 3.5.3->2026.7.1 + redis 7.4->8.10: 128 = 123 (app, by
publish date) + 5 (redis, by version range). The redis five are genuine for that bump
(patched 7.4.1 / 7.4.6 / 8.2.3) and include CVE-2025-49844, CRITICAL — previously
invisible. Regressions clean: gitea still 2, discourse without the sidecar window
still 123.
Answers 'how can the weekly run produce counts like the hand count?' — by doing
exactly what the hand count did, deterministically. Two changes:
1. PAGINATION. The scanner requested per_page=100 and stopped. This endpoint caps at
100 AND ignores ?page= (it re-returns the same rows — which is how a manual count
first produced exact triplicates and a bogus 300). Busy projects were silently
truncated: discourse has 286 advisories, so a single page could not see the window
at all. Now follows the Link rel=next cursor to exhaustion.
2. DATE-BASED FALLBACK. Version strings cannot be ordered across a scheme change
(discourse semver 3.5.3 -> calver 2026.7.1), which is why the scan first reported a
false 133, then correctly refused. Release DATES always order. When the version path
refuses, the scan now resolves both versions to their git tag dates on the primary
repo and counts advisories PUBLISHED in that window, labelling the method in the
output. The version path is still preferred when usable — it is exact rather than
temporal.
Verified: discourse 3.5.3 -> 2026.7.1 now reports 123, matching the hand count
(1 critical, 16 high, 91 medium, 16 low; window 2025-12-30 -> 2026-07-31); gitea
1.27.0 -> 1.27.1 still reports 2 via the version path.
Operator: 'the scanner should not say 0 when it was not able to scan.' Correct — the
previous patch still led with '0 identified' and relegated the caveat to a footnote,
so the headline number was wrong even though the prose was right. A 0 in a security
column is an assertion of safety; it must never be emitted for an undetermined result.
Now: cve_count_fixed is null (not 0) in JSON, a count_known flag distinguishes
'counted zero' from 'could not count', and the markdown headline reads
'CVEs fixed by this upgrade: UNKNOWN — the scan could NOT determine a count' with an
explicit 'This is NOT zero' and instructions to render '?'.
Verified: discourse 3.5.3 -> 2026.7.1 (semver->calver) now reports UNKNOWN; gitea
1.27.0 -> 1.27.1 still reports 2.
Operator disbelieved discourse's '133 CVEs fixed' — correctly. Two defects made it
confidently wrong:
1. ONE WINDOW APPLIED TO EVERY IMAGE. The scan queries all source repos in the
recipe's registry (app + redis/postgres/nginx sidecars) but judged them all with
the APP's version window. 34 of the 133 were redis advisories, including
CVE-2021-21309 — patched in redis 6.0.11 back in 2021 — scored as 'fixed by this
upgrade' purely because 6.0.11 sits numerically inside discourse's 3.5.3 ->
2026.7.1 range. Only the PRIMARY app repo is now classified; other sources are
reported as unclassified so they stay visible without inflating the count.
2. VERSION-SCHEME CHANGES BREAK ORDERING. discourse moved semver -> calver
(3.5.3 -> 2026.7.1), so 2025.12.2 compares 'newer' than 3.5.3 while shipping
earlier. Numeric comparison cannot order that. The scan now detects a leading-
component jump >= 100, refuses to classify, and says so in the block: the count
is '0 by refusal, not by evidence — read the vendor's release notes'.
Refusing to answer beats answering wrongly: a fabricated 133 in a public security
report is worse than an explicit 'cannot determine'.
Verified after the fix: discourse 133 -> 0 (with the refusal caveat), gitea still
exactly 2 (both criticals, patched 1.27.1), keycloak 7 all genuinely from
keycloak/keycloak patched in 26.7.1, plausible 1. No other count changed.
The 2026-08-07 regeneration rendered '?' for 5 of 21 recipes. '?' is meant to be a
rare 'we tried and could not tell'; at that rate it is indistinguishable from noise
and hides the real unknowns. Three causes, none of them genuine uncertainty:
1. URL EXTRACTION BUG (mine). The registry is markdown, so urls appear inside
`backticks` and 'quotes'. The extractor captured the trailing punctuation, so
it fetched https://docs.n8n.io/release-notes/` and https://git.autonomic.zone'`
— both 404 on the malformed url, both 200 when clean. Trailing markdown
punctuation is now stripped. Fixed immich + n8n.
2. STALE REGISTRY URL. mattermost-lts pointed at
docs.mattermost.com/about/mattermost-changelog.html, which 404s; the page moved
to /deploy/. Corrected (same class as the pgautoupgrade fix).
3. WRONG SEMANTICS FOR 'NO UPGRADE'. lasuite-docs and custom-html-tiny were
up-to-date this run, so no scan block existed and the report fell back to '?'.
But a recipe with no upgrade has nothing an upgrade could have fixed — that is
0, not unknown. The report skill now says so explicitly, restricts '?' to scans
that RAN and reported genuinely failed sources, states that benign notes
(no-advisories-published / template URL) never trigger '?', and instructs that
many '?' is itself a bug to raise in the Addendum.
Result across all 16 scanned recipes of that run: 0 failed sources (was 5).
Counts also improved with the classifier fix: discourse 130->133, keycloak ->7.
Exposed by asking whether the scan catches the n8n CVEs (CVE-2026-42231/42232). It
did not — the advisories were fetched correctly but both misclassified as
out-of-window. Two bugs:
1. Only vulnerabilities[0] was read. An advisory carries ONE ENTRY PER PATCHED
RELEASE LINE: n8n patches three (1.123.32, 2.17.4, 2.18.1), so whichever line
the deployment is actually on was silently dropped. gitea passed only because it
patches a single line. Now all entries are kept.
2. patched_versions is a RANGE EXPRESSION ('>= 2.18.1'), not a bare version. Naive
parsing produced (18,1) instead of (2,18,1), so no comparison could ever match.
Version tokens are now extracted with a regex and the advisory counts as
fixed-by-this-upgrade if ANY patched line falls in (from, to].
Verified: n8n 2.17.0 -> 2.18.1 now reports 12 CVEs including both criticals
(CVE-2026-42231 GHSA-q5f4-99jv-pgg5, CVE-2026-42232); gitea 1.27.0 -> 1.27.1 still
reports exactly 2. Note our deployed n8n (2.27.2+) is already past all three patched
lines, so these were never outstanding for us — the bug was in detection, not
exposure.
Two refinements found by running the scan across all 14 recipes of the 2026-08-07 run:
1. A repo with no advisory feed returns HTTP 404 on /security-advisories (e.g. the
pgautoupgrade sidecar image). That is a BENIGN ABSENCE, not a failed check.
Likewise registry entries that are TEMPLATE urls for humans
(…/changelog/v<VERSION>/, …/<vX.Y.Z>/…) are documentation, not fetchable.
Counting either as a failure pushed most recipes to '?', which would make the
unknown-vs-clean distinction meaningless again — the exact signal the ? exists to
preserve. Both are now recorded in sources_benign; only genuine errors (rate
limit, network, 5xx, wrong URL) land in sources_failed.
2. upstream/*.md pointed at github.com/pgautoupgrade/pgautoupgrade, which 404s —
the repo is pgautoupgrade/docker-pgautoupgrade. Corrected in n8n, lasuite-docs,
lasuite-drive, lasuite-meet. A 404ing registry URL means we were not scanning a
source we believed we were.
Effect on the 2026-08-07 data: recipes with genuine failed sources 5 -> 3 (the
remainder are really unreachable vendor pages). CVE counts unchanged where they
were already sound: discourse 130, gitea 2, plausible 1.
Anonymous GitHub API is 60 req/hr — a full weekly sweep across ~20 recipes exhausts
it and the scan then reports sources as failed (visible, but degraded coverage). A
token lifts it to 5000/hr.
_github_token(): GITHUB_TOKEN env wins, else GITHUB_TOKEN_FILE (default
/srv/cc-ci/.github-token, 0600, gitignored). Reading PUBLIC advisories needs NO
scopes — a classic PAT with nothing ticked, or fine-grained limited to 'Public
repositories: read'. The tool only ever GETs advisories; do not grant write scopes.
A missing token is not an error: the scan runs anonymously and surfaces failures.
Also gitignores .github-token and .hcloud-token.
Why: gitea 1.27.1 fixed CVE-2026-60004 + CVE-2026-59774 (both CVSS 9.8). The
2026-08-03 report printed gitea's CVE count as '1', the 2026-08-07 report as
'none'. Cause chain: the upgrade subagent read the GitHub release notes, which
name NEITHER cve (they are announced only in the vendor blog's security section),
so it recorded one unrelated minor item; the report then derived security content
from those notes plus model knowledge, and the model's training predates the CVEs.
Nothing in the pipeline ever queried an advisory source.
cc-ci-plan/advisory-scan.py — deterministic, per recipe, per upgrade window:
1. GitHub Security Advisories API for every source repo in the upstream registry.
PRIMARY: CVE + GHSA + severity + vulnerable/patched ranges, so 'fixed by THIS
upgrade' is computed. Needs no new per-recipe config (134 registry URLs are
already github.com).
2. Vendor release/security pages — every registry URL, fetched + regex-scanned.
This is the source that actually had the gitea CVEs.
3. OSV where a package mapping exists — supplementary.
Each source reports its own status so 'checked, none found' is never confused with
'not checked'. Source selection was measured, not assumed: for these two CVEs OSV
404'd and NVD's API had them by neither CPE, id, nor keyword — advisory DBs lag the
vendor, hence 1+2 lead.
Wiring is strictly ADDITIVE:
- /recipe-upgrade gains step 2a: run the scan, paste the block into the per-recipe
log, and report the UNION of it and the existing release-note reading. The scan
may never lower a count established by reading.
- /recipe-report treats the block as a FURTHER source, prefers its advisory ids /
severities / fixed-in versions for citation, and must render '?' (not 'none')
when a scan is absent or has failed sources — the false-clean 'none' is exactly
what happened on 2026-08-07.
- upstream/gitea.md records blog.gitea.com as the security-announcement URL.
Verified on the real regression: 1.27.0 -> 1.27.1 now yields exactly the 2 missed
criticals with their GHSA ids; the wider 1.26.2 -> 1.27.1 window yields 62.