--- name: weekly-upgrader-opencode-go-glm description: weekly upgrade run now defaults to opencode-go/glm-5.3-flash (PR #20, 2026-09-14); zen tier dead on this host; how backend/model + monitoring are wired metadata: node_type: memory type: project originSessionId: 85355980-5e4f-4f90-b1ca-d0e4fe82f04b --- The weekly recipe-upgrade run (`cc-ci-plan/launch-upgrader.py`, fired by the `cc-ci-upgrade-all` systemd timer Thu 22:00 Boston) now **defaults to backend=opencode, model=opencode-go/glm-5.3-flash** when no env is set (operator choice 2026-09-08 via `/srv/cc-ci/upgrader.env`; PR #20 hardcoded the same defaults into the launchers 2026-09-14 — for all three launchers: upgrader, report, supervisor). Model default tracks the backend — an explicit `LOOP_BACKEND=claude` (no model) falls back to `sonnet`. **2026-09-14 — ZEN IS DEAD ON THIS HOST.** The ZEN provider models (`opencode/glm-5.2`) are gone from the provider catalogue AND the host holds no ZEN key, so any launch that bypassed upgrader.env died instantly with `ProviderModelNotFoundError: opencode/glm-5.2` — this is what killed the 2026-09-11 report agent at launch and every manual retry (all valid ZEN credentials mean nothing without the ZEN models existing server-side). If a launcher log shows `ProviderModelNotFoundError`, the model id does not resolve on this host: run `opencode models | grep glm` for what does, then fix the tier/model default. See PR recipe-maintainers/cc-ci-orchestrator#20. **Provider:** "OpenCode **Go**" — a flat $10/mo subscription, DISTINCT from "opencode Zen" (pay-as-you-go). Endpoint `https://opencode.ai/zen/go/v1/` (OpenAI-compatible), model id **`opencode-go/glm-5.2`** (Zen tops out at glm-5.1 — `opencode.ai/zen/v1` does NOT have 5.2; updating the opencode CLI does NOT change this, the catalog is server-side). API key lives in `~/.local/share/opencode/auth.json` (mode 600, entry `opencode-go` type api) — **never in the repo**. opencode CLI installed: 1.15.13. **Override per run without a rebuild:** drop `/srv/cc-ci/upgrader.env` (read via `EnvironmentFile=-` on the unit) with e.g. `LOOP_BACKEND=claude` / `LOOP_MODEL=sonnet`. Absent file → the glm-5.2 default. **Monitoring (opencode backend):** the launcher does `opencode run -m --share --attach http://127.0.0.1:4096`. `opencode-web.service` (=`opencode serve` :4096, nginx → http://oc.commoninternet.net, tailnet-only) must be active. `OPENCODE_SHARE=1` (default) also mints a public `https://opncd.ai/share/` link (printed in the TUI / readable from the server API `GET /session/` → `.share.url`). Note: launch-upgrader.py's `session_busy()` regex is claude-tuned, so it misreports an active opencode run as "idle/finishing" — check the process (`opencode run`) or web UI instead. Gotcha fixed: opencode `-m/--model`/`--attach`/`--share`/`--title` are flags on the `run` SUBCOMMAND — they must come AFTER `run` (a global `opencode --model X run` is silently ignored). `launch-report.py` also defaults to opencode-go/glm-5.2 now (overridable via REPORT_BACKEND/REPORT_MODEL). In the weekly flow the upgrader writes the summary inline AND auto-launches the report (upgrade-all §6). **Usage limit + auto-resume watchdog (added 2026-06-23):** opencode-go enforces a ROLLING usage window (~5h) capped by dollar value. A full glm-5.2 upgrade run is heavy — ~8 recipes (≈2h20m wall) exhausted it mid-run, returning HTTP 429 + `retry-after` (seconds) on the go endpoint. When limited, the `opencode run` agent loop ENDS and does NOT self-resume (process lingers but idle). Fix: `launch-upgrader.py` now has `resume` (continue the SAME session, context preserved, via `opencode run -s --continue` — finds the session from the web server, kills the idle proc via /proc scan NOT `pkill -f` which self-matches) and `watchdog` (polls the session log; on >15min stall, waits out the 429 retry-after then auto-resumes). An opencode `start` auto-spawns the watchdog; logs to `cc-ci-upgrader-watchdog.log`. To resume by hand: `python3 /srv/cc-ci/cc-ci-plan/launch-upgrader.py resume`. (Subcommand `watchdog`, matching agents.py — renamed from `babysit` 2026-06-23.) **Watchdog covers ALL parts of the weekly run (2026-06-23):** (1) it runs in a tmux session (`-watchdog`) under the persistent tmux server, NOT a Popen child — so it survives the `cc-ci-upgrade-all` systemd Type=oneshot whose cgroup is reaped on exit (a Popen child would be killed). (2) The watchdog is generic (env: `UPGRADER_SESSION/_MODEL/_DONE_MARKER/_RESUME_FILE`), so `launch-report.py` reuses it for the REPORT job (`cc-ci-report` session, marker `RECIPE REPORT COMPLETE`) — the report runs on glm-5.2 sharing the same opencode-go budget the upgrade run drains, so it can 429-stall too. `_run_pids()` is scoped to the managed session (title or `-s `) so the two watchdogs never kill each other's process. Weekly timer: Thu 22:00 America/New_York; no `LOOP_BACKEND`/`upgrader.env` override ⇒ glm-5.2 defaults. See [[orchestrator-backend-switch-gotcha]]. **2026-07-04 hardening (after the 2026-07-03 run stalled on a HOST DISK-FULL and never resumed — the log-idle/429 watchdog does NOT cover an environmental wedge).** Three changes (commits 1bd156e, 399e999), host rebuilt: - **Step-0 stale-image pre-reclaim** (`launch-upgrader.py` `prereclaim_cc_ci()`): before each weekly run, `ssh cc-ci docker image prune -af --filter until=168h` — only images unused AND >1 week old, so this week's likely-reused images stay. Best-effort; env `UPGRADER_PRERECLAIM*`. A full run's images pile up week-over-week and had run root to 100%. (The 2026-07-03 pile was 93 GB / 194 images; pruning took it 88%→17%.) - **watchdog proc-death detection**: an `opencode run` proc EXITS when the model ends its turn, so a long /upgrade-all dies+restarts many times before completing (and the log mtime freezes on death, making log-idle useless). `watchdog()` now resumes on "no live run proc for the session + not `_completed()`", not just log-idle. Capped at `UPGRADER_MAX_RESUMES` (20) → hands off if stuck. `resume()` now auto-spawns a watchdog if none is alive (skips when the watchdog itself called it). - **Hourly glm-5.2 SUPERVISOR** (`launch-supervisor.py`, systemd `cc-ci-upgrade-supervisor` .service + .timer, hourly at :07): cheap deterministic gate — no-ops (0 tokens) if the run is complete / progressing / a per-run watchdog is alive; only when a run stalled+died AND the watchdog gave up does it launch a short-lived glm-5.2 agent to diagnose (disk/wedge) + drive to completion. `force`/`status`/ `stop` subcommands. Progress judged by `lu._run_pids(sid)` + log mtime (session_busy() is claude-tuned, misreads a headless opencode run as idle — do NOT use it for opencode). Window guard: won't auto-resurrect a session >96h old. Overrides via `/srv/cc-ci/upgrader.env` (e.g. `SUPERVISOR_MODEL=`). Ownership: **watchdog = prompt recovery (single writer while alive); supervisor = backstop when the watchdog is absent/gave up.** They're mutually exclusive (gate defers while watchdog tmux alive). **2026-07-04 finish + more watchdog fixes (commits 69dd178, ddedea8, b6a95c4).** Finishing the stalled 2026-07-03 run exposed that a glm-5.2 `opencode run` is FLAKY for long autonomous jobs — its turn ENDS (proc exits) repeatedly, `--continue` of a bloated session often doesn't advance, and it can wedge on a hung subagent. Bugs fixed in `launch-upgrader.py`: (1) `_session_idle_min()` now measures staleness across the whole session TREE (parent + descendant subagents) via the opencode server time.updated — NOT the tmux log mtime (freezes on `--continue`) and NOT the parent alone (a 20-40min recipe subagent made the parent look idle → false-resume killed the productive run). (2) watchdog loop is wrapped in try/except so a transient blip can't kill it. (3) **`_completed()` false-positive**: it matched DONE_MARKER inside TOOL parts (a subagent `task` prompt / bash echoing "print UPGRADE RUN COMPLETE") → watchdog exited early, run wedged. Now requires the marker in the LAST assistant TEXT (prose) message. (4) supervisor gate also treats the run as done once the weekly summary file exists (a run can finish work+summary+report yet never print a clean marker). **Operational fallback that WORKED:** when the report agent wedged twice, published the report by hand from the deterministic path — `recipe-report.py survey ` → build the spec JSON from survey.json → `render` → `publish`. This is reliable and bypasses the flaky agent entirely; keep it as the go-to when an agent-driven report stalls. The 2026-07-03 report is live at report.ci.commoninternet.net/week-2026-07-03.html (16 GREEN, plausible + lasuite-drive RED, discourse 5-PR reconcile).