merge origin/main

# Conflicts:
#	cc-ci-plan/JOURNAL.md
This commit is contained in:
autonomic-bot
2026-09-28 20:03:33 +00:00
26 changed files with 792 additions and 71 deletions
+87 -1
View File
@@ -31,8 +31,44 @@ handoff).
---
## Session 2026-05-31 ~18:30 UTC — Claude Sonnet 4.6
## Session 2026-09-14 ~16:45 UTC — opencode glm-5.3-flash (orchestrator) — round 2: Anubis UA
**Left off:** The real root cause turned out to be TWO independent layers; the mirror-privacy
fix (earlier session entry today) was necessary but not sufficient. Operator's browser console
showed CORS failures redirecting to `anubis.swarm.autonomic.zone/.within.website/?redir=…`.
Reproduced exactly: the `/pr/` proxy forwards the END browser's User-Agent to Gitea; Gitea sits
behind **Anubis**, which 307-challenges browser-like UAs to `anubis.swarm.autonomic.zone`
(no CORS headers) → every fetch throws in the browser → all cells "?" (curl passed clean, which
is why server-side checks and my earlier headless test never saw it — intermittent/rate-dependent
for my playwright run). Fix: cc-ci **PR #38** adds
`proxy_set_header User-Agent "ccci-reports-proxy/1.0";` to the reports.nix `/pr/` location.
Hot-verified on the host by mount-swapping a fixed conf into the running task (one mis-step:
`--mount-rm`+`--mount-add` same-target order wiped the mount; re-added), scoped live, all 16
cells rendering with a real Chromium. Merged PR #38, `nix flake update cc-ci`,
`nixos-rebuild test` → healthy (reports 200, no failed units) → `switch` (flake.lock commit
9e7770f). Final verify: browser-UA curl 200 both gitea/9 + full headless-Chromium sweep 16/16
OPEN, zero non-200 /pr fetches. Also flipped memory: `memory/gitea-anubis-ua-challenge.md` +
MEMORY.md index.
**Open:** nothing blocking; next weekly /recipe-report and STATUS live-checks carry the fix.
## Session 2026-09-14 ~15:00 UTC — opencode glm-5.3-flash (orchestrator)
**Left off:** Report STATUS column fix. Operator reported the week-2026-09-11 report's live
PR-STATUS column all "?" — root cause: the tokenless same-origin proxy
`report./pr/<recipe>/<n>` (cc-ci `nix/modules/reports.nix`) 404s on **private** mirrors; two
late-enrolled mirrors, `recipe-maintainers/gitea` (2026-06-11) and `wordpress` (2026-08-03),
had been created `"private":true` from birth — by the stale instruction in
`/recipe-enroll`'s mirror step (the other 21 mirrors were flipped public on 2026-06-09, and
the old 'org is private' blocker is long resolved). Fixed: secret-scanned both repos, flipped
`private=false` (PATCH with bot creds), patched `.opencode/skills/recipe-enroll/SKILL.md` to
create mirrors `private:false`, updated memory/recipe-mirrors-public-org-blocker.md +
MEMORY.md index. **Verified in a real headless Chromium (nixpkgs chromium + playwright)**:
all 16 STATUS rows render `open`, every `/pr/` fetch 200 JSON. Commit 6e93922 pushed. The
STATUS column refreshes live every 30s; cells go ✓ when a PR merges. No reports.nix change
was needed (proxy itself was healthy).
**Open:** nothing on this; the report index regenerates next weekly run.
## Session 2026-05-31 ~18:30 UTC — Claude Sonnet 4.6
**Left off:** Got opencode/deepseek-v4-pro working as the loop backend. Both builder and
adversary are actively running on `tinfoil/deepseek-v4-pro` (via `inference.tinfoil.sh`).
Phase 5 [11/11] in progress. The operator is debugging the opencode web UI visibility and
@@ -1202,6 +1238,56 @@ the host: `opencode-go/deepseek-v4-flash` and `opencode-go/glm-5.3-flash` answer
`upgrader.env` (`LOOP_TIER=go` maps to the `opencode-go` auth entry; `LOOP_MODEL` overrides the
tier default). Next fire Fri 2026-09-11 02:00 UTC.
- The steering orchestrator agent stays on `opencode-go/glm-5.2` (not asked to change).
## Session 2026-09-21 — domain cutover to ci.autonomic.zone (orchestrator)
**What happened.** Morning: hourly supervisor resumed the stalled 09-18 weekly run on the GO tier
(ZEN endpoint dead server-side — `UnknownError`; run completed 13 green PRs, 0 failed). Published
the missing week-2026-09-18 report (PR-finding #4; launcher defaults flipped to `go` in
cc-ci-orchestrator PR #23). Then executed the full domain cutover per
`cc-ci-plan/plan-domain-migration-ci-autonomic-zone.md` (PR #25).
**Plan deviation (simplification).** No dual-cert SNI: ONE Let's Encrypt cert carries SANs for
BOTH zones (`ci` + `*.ci` of autonomic.zone AND commoninternet.net) — the unchanged single-pair
`ssl_cert/ssl_key` traefik reconciler keeps working; Phase 4 reissues without the legacy SANs.
The new zone's DNS-01 challenge reuses the SAME acme-dns account: storage re-keyed by
`cc-ci-acme-storage-seed.service` (jq clone of the legacy entry under `ci.autonomic.zone`),
CNAME already delegated. Proven by a hand lego **staging** run before any production change.
**Merged:** cc-ci #39 (front doors dual Host rules, bridge/dashboard env URLs, drone abra rename,
runner RPC, naming.py → `*.ci.autonomic.zone`, dual-zone name regexes in lifecycle/warm/prune,
recipe-report URLs) · cc-ci #40 (seed-unit nesting fix) · cc-ci #41 (have_secret stack-scope —
caught live: the old stack's `*_rpc_secret_v1` satisfied the check post-rename) · cc-ci #42 (nix
interpolation escape) · orchestrator #27 (oc.ci host + `opencodeUiExtraHosts`, host self-pins,
flake bump).
**Deployed** via `nixos-rebuild test` → verify → `switch` (generation `nn1vwiv7v1k…`, running ==
boot). Cert SANs confirmed 4-name; all 5 front doors answer on BOTH zones
(200/200/303/401 + traefik 200), TLS verify=0 from outside; zero failed units; disk dropped
88%→45% after prune.
**Drone migration.** New abra app `drone.ci.autonomic.zone`, FRESH DB (module's
`DRONE_USER_CREATE` re-injected the sops bridge token). The Gitea OAuth app redirect now has both
URIs; the client secret was rotated (each Gitea PATCH regenerates it) and synced through
sops → `sops-install-secrets` → swarm secret v1 → drone. Bootstrapped OAuth
(`drone login ok (admin=true)`), re-enabled cc-ci + discourse repos, build timeout 60m.
Webhooks: ghost + discourse repointed (secrets preserved); cc-ci repo's bridge webhook →
`ci.autonomic.zone/hook`, stale drone hook deleted, Drone's auto-created new-zone hook active.
Old stack removed + orphaned secrets reaped.
**E2E proof.** `!testme` on keycloak PR #9 → bridge → drone build #1 (new DB numbering) → runner →
harness → `results.json` + PR card `✅ passed` linking `ci.autonomic.zone/runs/1/summary.png`.
**Deferred / open.**
- Warm stacks + backupbot stay on the legacy zone (data-warm volumes / restic password tied to
abra app names) — post-bake migration; harness regexes accept both zones meanwhile.
- Docs sweep (~60 references: AGENTS.md, README, skills incl. cc-ci-status front-door list,
launcher printed URLs).
- Phase 4 (after ≥7 clean days): drop legacy SANs (reissue), remove legacy Host arms + host
self-pins + Gandi records (`ci`, `*.ci`, `_acme-challenge.ci`), TTLs back to 3600.
- Host auto-update was `failed` 2026-09-15 (health check) — `/cc-ci-orchestrator-update` still
pending; next auto-attempt Tue 09-22.
## Session 2026-09-28 20:00 UTC — operator-broken cc-ci recovered by plain hard reset
- Operator reported ci.autonomic.zone down after their own change, supplied a Hetzner API token