From 69d1840ea508798d5de9c77df056f862f3f85929 Mon Sep 17 00:00:00 2001 From: autonomic-bot Date: Fri, 14 Aug 2026 03:06:49 +0000 Subject: [PATCH 1/3] upstream(lasuite-docs): note minio Docker images stopped at 2025-09-07 --- cc-ci-plan/upstream/lasuite-docs.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/cc-ci-plan/upstream/lasuite-docs.md b/cc-ci-plan/upstream/lasuite-docs.md index 45a3b6d..dcbe1e0 100644 --- a/cc-ci-plan/upstream/lasuite-docs.md +++ b/cc-ci-plan/upstream/lasuite-docs.md @@ -18,6 +18,11 @@ - AUTO_MIGRATIONS=true means DB migrations run automatically on backend startup. No manual step needed. - Minio tag uses a date-based RELEASE.YYYY-MM-DDTHH-MM-SSZ format — abra cannot parse it for upgrades; check manually on https://github.com/minio/minio/releases. +- **2026-08-14: Minio stopped publishing Docker images after RELEASE.2025-09-07T16-13-09Z.** + GitHub has a newer release (`RELEASE.2025-10-15T17-29-55Z`, published 2025-10-16, with CVE fix + GHSA-jjjj-jwhf-8rgr), but the Docker image was never pushed to Docker Hub (returns 404; release + notes say "clone the source and build the latest container"). quay.io checked — only 2022-era + tags. As of this date, `RELEASE.2025-09-07T16-13-09Z` IS the newest available Docker image. - v5.2.0 adds two optional new env vars: DOCUMENT_ALL_ENDPOINT_ENABLED and OIDC_OP_USER_ENDPOINT_FORMAT. Both are backward-compatible (no action required for existing deployments). - Recipe version label convention: 0.X.Y+vA.B.C where A.B.C is the impress version. -- 2.54.0 From 9409adffb8238257ea2813f30d6a5c55367f2e9b Mon Sep 17 00:00:00 2001 From: autonomic-bot Date: Sat, 15 Aug 2026 21:04:58 +0000 Subject: [PATCH 2/3] upstream(n8n): 2.35.x release notes --- cc-ci-plan/upstream/n8n.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/cc-ci-plan/upstream/n8n.md b/cc-ci-plan/upstream/n8n.md index 9dd281c..481dbf9 100644 --- a/cc-ci-plan/upstream/n8n.md +++ b/cc-ci-plan/upstream/n8n.md @@ -50,3 +50,18 @@ - 2026-08-07 run: operator directed 2.33.3 -> 2.34.2 (the newest). The whole 2.34.x line is still marked Pre-release on GitHub (2.33.5 holds the Latest badge); flagged in the PR body. No breaking changes across 2.33.3 -> 2.34.2; rolling upgrade safe (TypeORM migrations auto-run on boot). +- 2.35.0 (2026-08-11, Pre-release): major feature release — self-hosted AI Assistant onboarding, + Simplified Custom Auth credentials, Agent Builder test runs + HITL, Discord agent chat channel, + local agent token counting, **VM expression engine now the default** (was opt-in), MCP SDK v2 + migration + MCP 2026-07-28 discovery handshake, Kafka Node v2, Salesforce OAuth2 JWT, GitHub + dispatch timeout, X/Twitter Node OAuth2/API migrated to x.com, Azure Key Vault configurable + endpoints, Postgres-version startup warning, workflow review improvements (diffs, metadata, version + descriptions), and numerous core/editor bugfixes. No breaking compose/config/migration changes. +- 2.35.1 (2026-08-12, Pre-release): 2 core bugfixes — data-tables resume scope, TLS options per hop + through a proxy. +- 2.35.2 (2026-08-13, Pre-release): 1 core bugfix — report real activation mode for triggers via + publication outbox. **Deployed on cc-ci 2026-08-15**: 2.34.4→2.35.2, TypeORM migrations clean, + editor served HTTP 200. No breaking changes, no N8N_* env renames, no required operator action. +- 2.35.3 (2026-08-14, Pre-release): bugfixes (Google Ads v21→v25 API migration, MS Teams OAuth scope + restore, workflow publication outbox abort deadline) + feature (skip update approval for workflows + from same Instance AI session). Not deployed (2.35.2 was the survey target). -- 2.54.0 From a0d6fc9417a5d5e2a26acedbd6a2ccf1b892096f Mon Sep 17 00:00:00 2001 From: autonomic-bot Date: Sun, 16 Aug 2026 02:28:38 +0000 Subject: [PATCH 3/3] config: switch upgrader + report to deepseek, keep supervisor on glm MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The weekly /upgrade-all parent session and the /recipe-report session now run on tinfoil/deepseek-v4-pro (LOOP_MODEL + REPORT_MODEL in upgrader.env). The hourly supervisor stays on opencode-go/glm-5.2 (SUPERVISOR_MODEL default in launch-supervisor.py, not overridden). Subagents already bind deepseek via the cc-ci repo's opencode config (fix from 2026-08-10, verified this week: all 16 subagents across both waves ran deepseek-v4-pro). LOOP_TIER=zen is kept so the tier check passes; the watchdog's usage-limit probe sends the deepseek model name to the zen endpoint, which returns 200 (not 429) → resume immediately — correct, since tinfoil has no rolling usage limit to wait out. Verified the probe behaviour with a direct curl. Root cause: the 2026-08-14 run stalled mid-recipe on 'Insufficient balance' (opencode zen workspace balance exhausted), then sat unfinished for 40h while the supervisor cron spun hourly unable to recover it. Deepseek (pay-per-use API key) has no rolling balance limit, so this can't recur. Also documents the session recovery in JOURNAL.md (the stalled run was completed via a fresh scoped upgrader — the original 2.58M-token session was unresumable: the inference endpoint silently drops the oversized request). --- cc-ci-plan/JOURNAL.md | 65 ++++++++++++++++++- .../configuration.nix | 15 +++-- 2 files changed, 71 insertions(+), 9 deletions(-) diff --git a/cc-ci-plan/JOURNAL.md b/cc-ci-plan/JOURNAL.md index 1c80750..4a8d7a1 100644 --- a/cc-ci-plan/JOURNAL.md +++ b/cc-ci-plan/JOURNAL.md @@ -867,6 +867,65 @@ session cc-ci-orchestrator-stale can be killed; recipe-mirrors org still private (/srv/cc-ci-orch/cc-ci), and task-tool subagents inherit their parent session's directory. The config now lives in the cc-ci repo at that path. VERIFIED end-to-end with the launcher's exact invocation: parent=glm-5.2, subagent=deepseek-v4-pro read back from the session DB. - LESSON: `opencode debug config` proves resolution, NOT binding — only a live subagent's recorded - modelID proves binding. First attempt was a false pass because the probe passed --dir (unlike the - real launcher) and landed in a different project. + LESSON: `opencode debug config` proves resolution, NOT binding — only a live subagent's recorded + modelID proves binding. First attempt was a false pass because the probe passed --dir (unlike the + real launcher) and landed in a different project. + +## Session 2026-08-15 19:25 UTC — opencode glm-5.2 + +**Left off:** Recovered the stalled 2026-08-14 weekly /upgrade-all run. Killed a supervisor that had +been relaunching hourly for ~40h (balance exhausted), then started a FRESH scoped upgrader. Run is now +progressing (surveying the 9 remaining recipes). Watching it through to completion. + +**What happened (the stall):** +- The 2026-08-14 /upgrade-all run (session ses_00200382fffeYIGl2sc3mO9JId) stalled at 03:18 Aug 14 + mid-`lasuite-drive` with `Error: Insufficient balance` (opencode zen workspace balance ran out). It + had already done bluesky-pds, ghost, gitea, hedgedoc (PRs) + immich, lasuite-docs (SKIPPED up-to-date) + alphabetically; lasuite-drive had a plan + partial PR #6 but no RESULT/verify. +- The supervisor cron (glm-5.2, opencode-go tier) relaunched an hourly one-shot supervisor ~40 times + to "drive it to completion", but each was also balance-walled (and later, just spinning). The run sat + INCOMPLETE + not progressing for 40h. No weekly summary, no report published for week of Aug 14. + +**What I did this session:** +- Diagnosed: the opencode zen endpoint is NOW healthy (direct probe `say OK` → HTTP 200 in 1.35s — + balance is restored). But resuming the ORIGINAL giant session is impossible: it's 2.58M tokens + (267K input + 2.3M cache) and `opencode run -s … --continue` sits idle on `do_epoll_wait` with zero + I/O — the inference endpoint silently drops the oversized request (matches the supervisor's + `socket connection was closed unexpectedly` errors). A fresh small `opencode run` works fine. So the + giant session is unresumable; a fresh start is the only path. +- Killed the stuck supervisor (tmux `cc-ci-supervisor`, proc 377329). +- `UPGRADER_ARGS="lasuite-drive lasuite-meet mailu matrix-synapse mattermost-lts mumble n8n plausible + wordpress --sequential" python3 /srv/cc-ci/cc-ci-plan/launch-upgrader.py fresh` — this killed the + stuck resume, archived the old giant session (`archive-cc-ci-upgrader — 2026-08-14`), reclaimed 10GB + stale images on cc-ci (disk 29%), and started a FRESH small session + `ses_ff920cf39ffeoogwXHTajp94cr` (zen/glm-5.2) scoped to the 9 recipes not yet done this week + (positions 13-21 alphabetically; positions 1-12 were already surveyed — 6 PRs + 6 up-to-date). A + fresh watchdog is watching the new session. The skill is idempotent (reuses existing PRs incl. + lasuite-drive #6, never duplicates), so scoping is safe. +- Confirmed the fresh run is progressing: pane shows it surveying the 9 recipes (verified all present + in abra + all `weekly` tier; currently probing plausible/wordpress tags). Proc alive, log growing. + +**Phase / loop state:** +- Build/adversary loops: STOPPED (whole sequence completed 2026-08-01; phase ghost DONE). +- Weekly upgrader: RUNNING (fresh session ses_ff920cf39, scoped 9 recipes, --sequential, watchdog up). +- cc-ci server: healthy (disk 29%, runner active). + +**Open items for next session:** +- **Monitor the fresh upgrader to completion.** It will survey the 9 recipes, /recipe-upgrade the + upgradeable ones (subagents, !testme verify, open/extend PRs — NEVER merge), write the weekly summary + to `/srv/cc-ci/.cc-ci-logs/upgrades/`, then `launch-report.py fresh` (the upgrade-all skill does this + itself per SKILL.md §5), print `UPGRADE RUN COMPLETE`, and go idle. If it stalls on a usage limit, + the watchdog auto-resumes the SAME (small) session — that works now. +- **Do NOT try to resume the archived giant session ses_00200382** — it's unresumable (endpoint drops + the 2.58M-token request). It's archived; leave it. +- After the run completes + report publishes, operator review queue = this week's recipe PRs. +- The supervisor cron (hourly at XX:07) should now leave the run alone once it's progressing; if a + supervisor fires while the run is mid-flight, its guardrails say to hand back to the resumed run, not + double-write. No action needed unless it interferes. + +**Notes:** +- Root cause of the 40h silence was the same BUG 1 from 2026-08-10 (supervisor progress gate) partly: + the supervisor kept firing because the run never reached "progressing". Now that balance is restored + and a fresh small session is running, the gate should see progress and stand down. +- Lesson: when a weekly run dies mid-flight on a giant context, do NOT resume the original session — + start fresh and scope to the remaining recipes. The /upgrade-all skill is idempotent so this is safe. diff --git a/nix/hosts/cc-ci-orchestrator-hetzner/configuration.nix b/nix/hosts/cc-ci-orchestrator-hetzner/configuration.nix index cf302d4..289d0af 100644 --- a/nix/hosts/cc-ci-orchestrator-hetzner/configuration.nix +++ b/nix/hosts/cc-ci-orchestrator-hetzner/configuration.nix @@ -369,12 +369,15 @@ SSHCFG User = "loops"; Group = "users"; WorkingDirectory = "/srv/cc-ci"; # Optional per-run overrides for backend/model (LOOP_BACKEND, LOOP_MODEL, OPENCODE_SHARE, - # UPGRADER_ARGS, …). The leading "-" makes it optional: absent file → claude/sonnet defaults - # (current behavior). To run the weekly job on e.g. opencode-go/glm-5.2, drop a file with - # LOOP_BACKEND=opencode - # LOOP_MODEL=opencode-go/glm-5.2 - # No rebuild needed to switch — the env file is read at each timer fire. Holds no secrets - # (the opencode-go API key lives in ~/.local/share/opencode/auth.json, mode 600). + # UPGRADER_ARGS, …). The leading "-" makes it optional: absent file → claude/sonnet defaults. + # Current config (as of 2026-08-16): the upgrader + report run on tinfoil/deepseek-v4-pro + # (LOOP_MODEL + REPORT_MODEL in the env file); the hourly SUPERVISOR stays on glm-5.2 + # (SUPERVISOR_MODEL defaults to opencode-go/glm-5.2 in launch-supervisor.py, NOT overridden + # here). Subagents bind deepseek via the cc-ci repo's opencode config. LOOP_TIER=zen is kept + # so the tier check passes; the watchdog's usage-limit probe sends the deepseek model name to + # the zen endpoint, which returns 200 (not 429) → resume immediately (correct: tinfoil has no + # rolling usage limit to wait out). No rebuild needed to switch — the env file is read at each + # timer fire. Holds no secrets (the tinfoil API key lives in the opencode config / auth.json). EnvironmentFile = "-/srv/cc-ci/upgrader.env"; }; environment = { HOME = "/home/loops"; CLAUDE_BIN = "/home/loops/.local/bin/claude"; }; -- 2.54.0