Files
agent-orchestrator/tests/test_unit.py
T
notplantsandClaude Fable 5 e185cec88c tangled_pr: verify against the pulls list, not a response header; and a tools test suite
THE BUG THAT PROMPTED THIS. tangled_pr.py judged success ONLY by an HX-Redirect header on the POST.
A create that SUCCEEDED but answered without that header read as a failure, so the caller retried
and Tangled grew duplicates — that is exactly how #397, #398 and #399 were filed for one branch. A
response header describes what the server meant to say; it is not the artifact.

Now it checks the artifact, in both directions:
  * BEFORE posting, refuse if an open pull already exists for this source branch, naming it. A
    retry cannot duplicate, whatever the response said. (--allow-duplicate to override.)
  * AFTER posting, confirm against the pulls list: a new pull number that did not exist before,
    whose page names this source branch, IS the success — with or without a redirect header.
  * Failure is reported only when no such pull appeared. A false failure is worse than a loud
    error here, because the caller's remedy is to retry.
Verified live: a dry-run against a branch that already has a pull refuses with rc=3, naming #417.

TWO REAL DEFECTS FOUND BY WRITING THE TESTS.

agents.py shelled out to `pgrep -P` and `ps -o comm=`. Neither is on the agent PATH on this host,
and a missing binary under shell=True returns rc=127 with EMPTY stdout — indistinguishable from
"this process has no children" and "no build is running". So _build_running was ALWAYS False and
the stall detector could reboot an agent mid-build. Both now read /proc directly: no PATH
dependency, and it cannot fail silently in that direction.

That shipped because the unit tests MOCKED pgrep and ps. The fakes stood in for the broken
dependency, so the suite passed on a host where neither tool was reachable and never exercised the
real path. The tests now patch _proc_descendants and _comms — the seams this repo owns. A test that
mocks a dependency proves the mock works.

Also fixed a monkeypatch leak those tests had: restoration used a name derivation that silently
matched nothing, so the patch escaped into another test class and failed an unrelated test — only
in a full run, never when that test ran alone. Now addCleanup, which cannot be ordered wrong.

NEW: tests/test_tools.py, 24 tests over tangled_pr, tangled_pr_close and gateway-domain, with every
HTTP boundary injected so they run offline. Mutation-checked: breaking classify(), the pull-number
regex, the branch match, or the scan bound each turns the suite red. Suite is 93 tests, green, and
order-stable across repeated runs.

README: a "PATH on a NixOS host" section. Every one of ps, pgrep, free, cmp, awk, curl, diff,
strings, nm, getent and ping is INSTALLED here and simply not on the agent PATH, so each reports
"command not found" and reads as a missing package. Documents how to check before concluding a tool
is absent, how to add the system profile, `nix shell` for what is genuinely missing, and the rule
that harness code should not shell out for what the kernel already exposes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01V3LdmEL7CvCYTNpoBq1kce
2026-08-21 03:50:11 +00:00

772 lines
35 KiB
Python
Executable File

#!/usr/bin/env python3
"""Unit tests for the agent-orchestrator harness (agents.py).
Pure-logic tests — NO agent CLIs spawned, NO live tmux sessions created. Every test builds a
throwaway config + fixture files in a tempdir and exercises the harness functions directly.
The one function that would spawn sessions (phase_advance_check → start/stop_loops) is tested
with those two hooks monkeypatched to recorders, so the phase-machine *logic* is covered without
launching anything.
Run: python3 -m unittest tests.test_unit (from repo root)
or python3 tests/test_unit.py
"""
import os
import sys
import time
import textwrap
import tempfile
import shutil
import re
import signal
import subprocess
import unittest
from datetime import datetime, timedelta
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(REPO_ROOT))
import agents # noqa: E402
# ── shared fixture config ────────────────────────────────────────────────────────
BASE_TOML = r"""
[watchdog]
signal_interval = 30
heavy_interval = 300
limit_probe_fallback = 300
limit_reset_slack = 45
stall_grace = 180
[defaults]
session_prefix = "aotest-ut-"
log_dir = "state"
backend = "claude"
model = "claude-sonnet-4-6"
watch = "none"
[backend.claude]
bin = "claude"
flags = "--dangerously-skip-permissions"
remote_control = true
supports_resume = true
prompt_delivery = "arg"
process_name = "claude"
submit_key = "Enter"
stall_idle = 300
active_re = "esc to interrupt|Running tool|\\u00b7 \\d+"
limit_re = "spend limit|usage limit|limit reached|reached your .*limit|out of (credits|tokens)"
fatal_re = "redacted_thinking|blocks cannot be modified"
[backend.opencode]
bin = "opencode"
attach = "{bin} attach {server} --dir {dir}"
server = "http://127.0.0.1:4096"
supports_resume = false
prompt_delivery = "ping"
process_name = "opencode"
footer_ui = true
log_grace = 180
connect_delay = 12
submit_key = "C-m"
stall_idle = 900
active_re = "esc interrupt|thinking|inferring|running tool|tool call|preparing patch|reading|searching"
limit_re = "usage limit|limit reached"
[backend.demo]
bin = "echo up; exec sleep 100000"
prompt_delivery = "exec"
[[agent]]
name = "builder"
kind = "loop"
role = "builder"
backend = "demo"
[[agent]]
name = "adversary"
kind = "loop"
role = "adversary"
backend = "demo"
[[agent]]
name = "cl"
kind = "persistent"
backend = "claude"
prompt = "hi"
[[agent]]
name = "oc"
kind = "persistent"
backend = "opencode"
prompt = "hi"
[[agent]]
name = "custom"
kind = "persistent"
session = "explicit-session"
model = "override-model"
dir = "/abs/somewhere"
backend = "demo"
prompt = "x"
[[service]]
name = "svc"
command = "sleep 1"
[loop]
state_file = "phase-idx"
resume_phase = true
auto_advance = true
done_marker = "## DONE"
kickoff_template = "prompts/kickoff.md"
roles_dir = "prompts"
handoff = { repo = ".", claim_pings = "adversary", review_pings = "builder", inboxes = ["ADVERSARY-INBOX.md", "BUILDER-INBOX.md"], state_subdir = "machine-docs" }
phases = [
{ id = "p1", plan = "PLAN1.md", status = "STATUS-p1.md" },
{ id = "p2", plan = "PLAN2.md", status = "STATUS-p2.md", models = { builder = "opus-x" } },
]
"""
KICKOFF_TMPL = "*** PROJECT PHASE: {phase_id} ***\nPLAN: {plan}\nSTATUS: {status}\nROLE: {role}\n---\n"
BUILDER_PROMPT = "You are the **Builder** agent. (builder role body marker)\n"
ADVERSARY_PROMPT = "You are the **Adversary** agent. (adversary role body marker)\n"
def _make_project(tmp, toml=BASE_TOML):
"""Write a self-contained project (config + prompts + machine-docs) into tmp; return cfg path."""
root = Path(tmp)
(root / "prompts").mkdir(parents=True, exist_ok=True)
(root / "machine-docs").mkdir(parents=True, exist_ok=True)
(root / "prompts" / "kickoff.md").write_text(KICKOFF_TMPL)
(root / "prompts" / "builder.md").write_text(BUILDER_PROMPT)
(root / "prompts" / "adversary.md").write_text(ADVERSARY_PROMPT)
cfg_path = root / "agents.toml"
cfg_path.write_text(toml)
return cfg_path
# ── config loading + defaults merge ────────────────────────────────────────────────
class TestConfigLoad(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg_path = _make_project(self.tmp)
self.cfg = agents.load_config(self.cfg_path)
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_defaults_merge_into_agents(self):
b = self.cfg["agents"]["builder"]
self.assertEqual(b["session_prefix"], "aotest-ut-")
self.assertEqual(b["watch"], "none") # from defaults
self.assertEqual(b["kind"], "loop") # explicit
def test_session_name_defaults_to_prefix_plus_name(self):
self.assertEqual(self.cfg["agents"]["builder"]["session"], "aotest-ut-builder")
def test_explicit_session_overrides_prefix(self):
self.assertEqual(self.cfg["agents"]["custom"]["session"], "explicit-session")
def test_per_agent_override_wins_over_default(self):
# default model is claude-sonnet-4-6; custom overrides
self.assertEqual(self.cfg["agents"]["custom"]["model"], "override-model")
self.assertEqual(self.cfg["agents"]["builder"]["model"], "claude-sonnet-4-6")
def test_relative_dir_resolved_against_project_root(self):
# builder has no dir → defaults dir "." → project_dir
self.assertEqual(self.cfg["agents"]["builder"]["dir"], self.cfg["project_dir"])
def test_absolute_dir_kept(self):
self.assertEqual(self.cfg["agents"]["custom"]["dir"], "/abs/somewhere")
def test_log_dir_and_state_dir_resolved(self):
self.assertEqual(self.cfg["log_dir"], str(Path(self.cfg["project_dir"]) / "state"))
self.assertEqual(self.cfg["state_dir"], os.path.join(self.cfg["log_dir"], "state"))
self.assertTrue(Path(self.cfg["state_dir"]).is_dir()) # created on load
def test_service_session_named(self):
self.assertIn("svc", self.cfg["services"])
self.assertEqual(self.cfg["services"]["svc"]["session"], "aotest-ut-svc")
def test_backend_of_resolves(self):
b = agents.backend_of(self.cfg, self.cfg["agents"]["cl"])
self.assertEqual(b["prompt_delivery"], "arg")
self.assertEqual(b["submit_key"], "Enter")
def test_backend_of_unknown_dies(self):
a = dict(self.cfg["agents"]["cl"]); a["backend"] = "nope"
with self.assertRaises(SystemExit):
agents.backend_of(self.cfg, a)
def test_missing_session_prefix_dies(self):
bad = self.tmp + "/bad1"
p = _make_project(bad, toml='[defaults]\nlog_dir = "state"\n')
with self.assertRaises(SystemExit):
agents.load_config(p)
def test_missing_log_dir_dies(self):
bad = self.tmp + "/bad2"
p = _make_project(bad, toml='[defaults]\nsession_prefix = "x-"\n')
with self.assertRaises(SystemExit):
agents.load_config(p)
def test_env_override_model_single_invocation(self):
os.environ["AGENT_MODEL_cl"] = "env-only-model"
try:
cfg2 = agents.load_config(self.cfg_path)
self.assertEqual(cfg2["agents"]["cl"]["model"], "env-only-model")
finally:
del os.environ["AGENT_MODEL_cl"]
# without the env var the file value stands again
cfg3 = agents.load_config(self.cfg_path)
self.assertEqual(cfg3["agents"]["cl"]["model"], "claude-sonnet-4-6")
class TestExampleConfig(unittest.TestCase):
"""The SHIPPED agents.example.toml must parse and define the documented shape."""
def test_example_config_loads(self):
ex = REPO_ROOT / "agents.example.toml"
self.assertTrue(ex.exists(), "agents.example.toml missing from repo")
cfg = agents.load_config(ex)
self.assertIn("builder", cfg["agents"])
self.assertIn("adversary", cfg["agents"])
for be in ("demo", "claude", "opencode"):
self.assertIn(be, cfg["backends"], f"backend {be} missing from example")
self.assertEqual(len(agents.phases(cfg)), 2)
# ── kickoff-template assembly ──────────────────────────────────────────────────────
class TestKickoff(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_kickoff_renders_slots_and_appends_role(self):
out = agents.build_loop_kickoff(self.cfg, self.cfg["agents"]["builder"])
self.assertIn("PROJECT PHASE: p1", out) # phase_id slot filled (phase idx 0)
self.assertIn("PLAN: PLAN1.md", out)
self.assertIn("STATUS: STATUS-p1.md", out)
self.assertIn("ROLE: builder", out)
self.assertIn("builder role body marker", out) # role prompt appended
self.assertNotIn("{phase_id}", out) # no unrendered slot
self.assertNotIn("{role}", out)
def test_kickoff_picks_correct_role_prompt(self):
out = agents.build_loop_kickoff(self.cfg, self.cfg["agents"]["adversary"])
self.assertIn("adversary role body marker", out)
self.assertNotIn("builder role body marker", out)
def test_agent_prompt_loop_returns_kickoff(self):
out = agents.agent_prompt(self.cfg, self.cfg["agents"]["builder"])
self.assertIn("PROJECT PHASE: p1", out)
def test_agent_prompt_persistent_returns_inline_prompt(self):
out = agents.agent_prompt(self.cfg, self.cfg["agents"]["cl"])
self.assertEqual(out, "hi")
def test_role_model_phase_override(self):
# phase p2 overrides builder model to opus-x; advance index to 1
Path(agents.phase_idx_file(self.cfg)).write_text("1")
self.assertEqual(agents.role_model(self.cfg, self.cfg["agents"]["builder"]), "opus-x")
# adversary has no override → its configured/default model
self.assertEqual(agents.role_model(self.cfg, self.cfg["agents"]["adversary"]),
"claude-sonnet-4-6")
# ── phase machine ──────────────────────────────────────────────────────────────────
class TestPhaseMachine(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
self.md = Path(self.cfg["project_dir"]) / "machine-docs"
# monkeypatch the session-spawning hooks so the machine logic runs without tmux
self._orig = (agents.stop_loops, agents.start_loops, agents.handoff_reset)
self.calls = []
agents.stop_loops = lambda cfg: self.calls.append("stop")
agents.start_loops = lambda cfg: self.calls.append("start")
agents.handoff_reset = lambda: self.calls.append("reset")
def tearDown(self):
agents.stop_loops, agents.start_loops, agents.handoff_reset = self._orig
shutil.rmtree(self.tmp, ignore_errors=True)
def _status(self, basename, text):
(self.md / basename).write_text(text)
def test_phase_done_detects_marker(self):
self._status("STATUS-p1.md", "header\n## DONE\nall verified PASS\n")
self.assertTrue(agents.phase_done(self.cfg, "STATUS-p1.md"))
def test_phase_done_rejects_placeholder_body(self):
self._status("STATUS-p1.md", "## DONE\nnot yet — written here only when complete\n")
self.assertFalse(agents.phase_done(self.cfg, "STATUS-p1.md"))
def test_phase_done_false_when_no_marker(self):
self._status("STATUS-p1.md", "## In progress\nworking\n")
self.assertFalse(agents.phase_done(self.cfg, "STATUS-p1.md"))
def test_phase_done_false_when_file_missing(self):
self.assertFalse(agents.phase_done(self.cfg, "STATUS-nope.md"))
def test_cur_idx_reads_state_file(self):
Path(agents.phase_idx_file(self.cfg)).write_text("1")
self.assertEqual(agents.cur_idx(self.cfg), 1)
def test_advance_on_done(self):
Path(agents.phase_idx_file(self.cfg)).write_text("0")
self._status("STATUS-p1.md", "## DONE\nverified\n")
advanced = agents.phase_advance_check(self.cfg)
self.assertTrue(advanced)
self.assertEqual(agents.cur_idx(self.cfg), 1) # moved to p2
self.assertIn("stop", self.calls)
self.assertIn("start", self.calls)
def test_no_advance_when_not_done(self):
Path(agents.phase_idx_file(self.cfg)).write_text("0")
self._status("STATUS-p1.md", "## In progress\n")
self.assertFalse(agents.phase_advance_check(self.cfg))
self.assertEqual(agents.cur_idx(self.cfg), 0)
self.assertEqual(self.calls, [])
def test_sequence_complete_idempotent(self):
Path(agents.phase_idx_file(self.cfg)).write_text("1") # last phase
self._status("STATUS-p2.md", "## DONE\nverified\n")
marker = Path(self.cfg["log_dir"]) / "SEQUENCE-COMPLETE"
# first call: completes the sequence
self.assertTrue(agents.phase_advance_check(self.cfg))
self.assertTrue(marker.exists())
self.assertEqual(self.calls.count("stop"), 1)
# second call: idempotent — no re-stop, returns False
self.assertFalse(agents.phase_advance_check(self.cfg))
self.assertEqual(self.calls.count("stop"), 1)
def test_append_phase_clears_marker_and_resumes(self):
# simulate "sequence already complete", then a 3rd phase appended to the config
Path(agents.phase_idx_file(self.cfg)).write_text("1")
self._status("STATUS-p2.md", "## DONE\nverified\n")
marker = Path(self.cfg["log_dir"]) / "SEQUENCE-COMPLETE"
marker.write_text("stale completion\n")
self.cfg["loop"]["phases"].append(
{"id": "p3", "plan": "PLAN3.md", "status": "STATUS-p3.md"})
advanced = agents.phase_advance_check(self.cfg)
self.assertTrue(advanced)
self.assertEqual(agents.cur_idx(self.cfg), 2) # resumed onto p3
self.assertFalse(marker.exists()) # stale marker cleared
self.assertIn("start", self.calls)
def test_custom_done_marker(self):
self.cfg["loop"]["done_marker"] = "## SHIPPED"
self._status("STATUS-p1.md", "## SHIPPED\nverified\n")
self.assertTrue(agents.phase_done(self.cfg, "STATUS-p1.md"))
self.assertFalse(agents.phase_done(self.cfg, "STATUS-p2.md"))
# ── usage-limit banner reset parsing ───────────────────────────────────────────────
class TestLimitParsing(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_parse_reset_pm(self):
ep = agents._parse_reset_epoch("You've hit your limit · resets at 10pm")
self.assertIsNotNone(ep)
self.assertEqual(datetime.fromtimestamp(ep).hour, 22)
def test_parse_reset_am_with_minutes(self):
ep = agents._parse_reset_epoch("resets 3:30am")
self.assertIsNotNone(ep)
dt = datetime.fromtimestamp(ep)
self.assertEqual((dt.hour, dt.minute), (3, 30))
def test_parse_reset_12am_is_midnight(self):
ep = agents._parse_reset_epoch("resets at 12am")
self.assertEqual(datetime.fromtimestamp(ep).hour, 0)
def test_parse_reset_invalid_hour_none(self):
self.assertIsNone(agents._parse_reset_epoch("resets at 25"))
def test_parse_reset_no_match_none(self):
self.assertIsNone(agents._parse_reset_epoch("everything is fine here"))
def test_parse_reset_picks_last_match(self):
ep = agents._parse_reset_epoch("resets at 9am ... actually resets at 11am")
self.assertEqual(datetime.fromtimestamp(ep).hour, 11)
def test_next_limit_until_unparsable_fallback(self):
now = time.time()
until, parsed = agents._next_limit_until(self.cfg, "limit reached, no time given", now)
self.assertFalse(parsed)
self.assertEqual(int(until), int(now + 300)) # limit_probe_fallback
def test_next_limit_until_within_window_uses_banner(self):
now = time.time()
t = datetime.now() + timedelta(hours=2)
h12 = t.hour % 12 or 12
ampm = "am" if t.hour < 12 else "pm"
banner = f"weekly limit · resets at {h12}:{t.minute:02d}{ampm}"
until, parsed = agents._next_limit_until(self.cfg, banner, now)
self.assertTrue(parsed)
self.assertGreater(until, now)
self.assertLessEqual(until - now, 6 * 3600 + 60) # within 6h window (+slack)
def test_next_limit_until_far_future_falls_back(self):
now = time.time()
t = datetime.now() + timedelta(hours=7) # > 6h window
h12 = t.hour % 12 or 12
ampm = "am" if t.hour < 12 else "pm"
banner = f"limit · resets at {h12}:{t.minute:02d}{ampm}"
until, parsed = agents._next_limit_until(self.cfg, banner, now)
self.assertFalse(parsed)
self.assertEqual(int(until), int(now + 300))
# ── stall / WAITING-UNTIL parsing ──────────────────────────────────────────────────
class TestWaitingUntil(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
self.claude_agent = self.cfg["agents"]["cl"] # non-footer backend
self.oc_agent = self.cfg["agents"]["oc"] # footer_ui backend
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
def test_non_footer_finds_marker_anywhere(self):
pane = "blah blah\nWAITING-UNTIL: 2030-06-13T12:00:00Z\nmore output after\n"
ep = agents._parse_waiting_until(self.cfg, self.claude_agent, pane)
self.assertIsNotNone(ep)
self.assertEqual(ep, datetime.fromisoformat("2030-06-13T12:00:00+00:00").timestamp())
def test_non_footer_none_without_marker(self):
self.assertIsNone(agents._parse_waiting_until(
self.cfg, self.claude_agent, "just working, no marker"))
def test_footer_honors_marker_above_the_footer(self):
# A footer_ui backend renders its input-box/status footer BELOW the agent's message, so the
# marker is never the literal last line. It must still be honored (this is the real-claude case).
pane = "WAITING-UNTIL: 2030-06-13T12:00:00Z\n ▣ Build · GPT · 2m 19s\n"
ep = agents._parse_waiting_until(self.cfg, self.oc_agent, pane)
self.assertIsNotNone(ep)
self.assertEqual(ep, datetime.fromisoformat("2030-06-13T12:00:00+00:00").timestamp())
def test_takes_most_recent_marker(self):
pane = ("WAITING-UNTIL: 2030-01-01T00:00:00Z\nwork\n"
"WAITING-UNTIL: 2031-06-13T12:00:00Z\n footer\n")
ep = agents._parse_waiting_until(self.cfg, self.oc_agent, pane)
self.assertEqual(ep, datetime.fromisoformat("2031-06-13T12:00:00+00:00").timestamp())
def test_bad_timestamp_none(self):
self.assertIsNone(agents._parse_waiting_until(
self.cfg, self.claude_agent, "WAITING-UNTIL: not-a-time"))
# ── backend activity detectors (claude + opencode footers) ──────────────────────────
class TestActivityDetection(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
self.claude_agent = self.cfg["agents"]["cl"]
self.oc_agent = self.cfg["agents"]["oc"]
def tearDown(self):
shutil.rmtree(self.tmp, ignore_errors=True)
# claude: non-footer, active_re matched anywhere in the pane
def test_claude_active_esc_to_interrupt(self):
self.assertTrue(agents.pane_active(
self.cfg, self.claude_agent, "thinking...\n esc to interrupt", use_log=False))
def test_claude_active_running_tool(self):
self.assertTrue(agents.pane_active(
self.cfg, self.claude_agent, "Running tool: Bash", use_log=False))
def test_claude_active_spinner_dot_count(self):
self.assertTrue(agents.pane_active(
self.cfg, self.claude_agent, "Compiling · 137 tokens", use_log=False))
def test_claude_idle_is_not_active(self):
self.assertFalse(agents.pane_active(
self.cfg, self.claude_agent, "Done.\n> ", use_log=False))
# opencode: footer_ui — only the bottom rows count as activity
def test_opencode_active_footer(self):
pane = "~ Preparing patch...\n ⬝⬝■ esc interrupt 137.6K\n"
self.assertTrue(agents.pane_active(self.cfg, self.oc_agent, pane, use_log=False))
def test_opencode_idle_footer_not_active(self):
pane = " ▣ Build · GPT-5.4 · 2m 19s\n 178.4K (17%) ctrl+p commands\n"
self.assertFalse(agents.pane_active(self.cfg, self.oc_agent, pane, use_log=False))
def test_opencode_active_only_at_top_is_ignored(self):
# active marker far above the bottom 10 lines → a footer UI ignores it
pane = "running tool now\n" + "\n".join(f"line {i}" for i in range(20)) + \
"\n ▣ Build · GPT · idle\n"
self.assertFalse(agents.pane_active(self.cfg, self.oc_agent, pane, use_log=False))
def test_opencode_log_grace_fallback(self):
# idle footer, but a freshly-touched session log within the grace window → active
idle = " ▣ Build · GPT · idle\n 178K (17%) ctrl+p\n"
logp = agents._session_log_path(self.cfg, self.oc_agent["session"])
logp.parent.mkdir(parents=True, exist_ok=True)
logp.write_text("recent activity\n") # mtime = now
self.assertTrue(agents.pane_active(self.cfg, self.oc_agent, idle, use_log=True))
# remove the log → no fallback → idle footer reads as not active
logp.unlink()
self.assertFalse(agents.pane_active(self.cfg, self.oc_agent, idle, use_log=True))
# ── build-aware stall detection ─────────────────────────────────────────────────────
# A silent pane whose claude session has a real compile/coverage/test process running is a
# running build, NOT a stall. These cover the three moving parts: the process-name match set,
# the session-scoped descendant walk, _build_running end-to-end, and stall_check_one's defer.
class TestBuildProcRegex(unittest.TestCase):
"""High-signal build/test tool names match; generic interpreters + the harness itself must
NOT (bare python/node/bash would false-positive on the orchestrator's own engine, and the
claude root's args embed the prompt which names cargo/rustc/…)."""
def setUp(self):
self.rx = re.compile(agents.DEFAULT_BUILD_PROCS_RE)
def test_build_tools_match(self):
for name in ["cargo", "cargo-llvm-cov", "cargo-mutants", "cargo-nextest", "nextest",
"rustc", "rustdoc", "cc1", "cc1plus", "collect2", "lld", "llvm-cov",
"llvm-profdata", "lichen-server", "lichen-cms", "lichen-shell",
"lichen-cli", "chromium", "chrome", "playwright"]:
self.assertTrue(self.rx.match(name), f"{name!r} should be treated as a build proc")
def test_generic_procs_do_not_match(self):
for name in ["python", "python3", "bash", "sh", "node", "claude", "tmux",
"sleep", "grep", "git", "ssh", "vim", "pgrep", "ps"]:
self.assertIsNone(self.rx.match(name), f"{name!r} must NOT be treated as a build proc")
def test_anchored_no_substring_false_positives(self):
# ^...$ anchored — a longer name merely containing a tool word must not match
for name in ["cargotool", "xrustc", "mycc1", "playwrightish", "notchromium", "rustcx"]:
self.assertIsNone(self.rx.match(name), f"{name!r} must not match (substring)")
class TestProcDescendants(unittest.TestCase):
"""_proc_descendants returns the real child tree of the given roots and EXCLUDES the roots."""
def test_finds_children_excludes_root(self):
p = subprocess.Popen("sleep 30 & sleep 30 & wait", shell=True, start_new_session=True)
try:
time.sleep(0.4) # let the two children spawn
kids = agents._proc_descendants([str(p.pid)])
self.assertNotIn(str(p.pid), kids) # root excluded
self.assertGreaterEqual(len(kids), 2) # the two sleeps
# Read /proc/<pid>/comm rather than shelling to `ps -o comm=`: this host's ps
# returns nothing for that form, and an empty result from a tool that is absent or
# unsupported is indistinguishable from "the children are not sleeps". Same class of
# bug as the pgrep dependency this test just caught in _proc_descendants.
comms = []
for k in sorted(kids):
try:
comms.append(open(f"/proc/{k}/comm").read().strip())
except OSError:
pass
self.assertIn("sleep", comms, f"expected a sleep among {comms}")
finally:
os.killpg(os.getpgid(p.pid), signal.SIGKILL); p.wait()
def test_leaf_process_has_no_descendants(self):
p = subprocess.Popen(["sleep", "30"], start_new_session=True)
try:
time.sleep(0.2)
self.assertEqual(agents._proc_descendants([str(p.pid)]), set())
finally:
os.killpg(os.getpgid(p.pid), signal.SIGKILL); p.wait()
class TestBuildRunning(unittest.TestCase):
"""_build_running is scoped to the watched session's pane_pid descendants (never the root
claude process) and matches on process comm."""
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
self.agent = self.cfg["agents"]["cl"]
self._orig_run = agents.subprocess.run
self.ps_targets = ""
def tearDown(self):
# Restore EXPLICITLY. A cleverer derivation of these names silently matched nothing, so the
# monkeypatch leaked into the next test class and made an unrelated test fail — visible only
# in a full run, never when that test was run alone.
agents.subprocess.run = self._orig_run
shutil.rmtree(self.tmp, ignore_errors=True)
def _patch(self, descendant_comms):
"""pane_pid=1000; its children are 1001,1002 with the given comms.
Patches the two internal seams (_proc_descendants, _comms) rather than faking
`pgrep`/`ps` subprocess calls. WHY (2026-08-21): the previous version mocked those two
commands, so the suite passed on hosts where NEITHER IS INSTALLED — the fake stood in for
the broken dependency and the real bug (empty output read as "no build running", which
lets the watchdog reboot an agent mid-build) was invisible to every test. Mock the seam
you own, not the tool you depend on, or the test proves the mock works.
"""
outer = self
class R:
def __init__(self, out): self.stdout = out; self.returncode = 0
def fake_run(cmd, *a, **k):
return R("1000\n") if "list-panes" in cmd else R("")
agents.subprocess.run = fake_run
orig_desc, orig_comms = agents._proc_descendants, agents._comms
def _restore():
agents._proc_descendants, agents._comms = orig_desc, orig_comms
self.addCleanup(_restore) # runs even if the test errors; no tearDown ordering to get wrong
agents._proc_descendants = lambda roots: (
{"1001", "1002"} if "1000" in list(roots) else set())
def fake_comms(pids):
outer.ps_targets = ",".join(sorted(pids))
return list(descendant_comms)
agents._comms = fake_comms
def test_detects_running_build(self):
self._patch(["bash", "cargo"])
self.assertTrue(agents._build_running(self.cfg, self.agent))
def test_no_build_when_only_shells(self):
self._patch(["bash", "vim"])
self.assertFalse(agents._build_running(self.cfg, self.agent))
def test_only_inspects_descendants_not_the_claude_root(self):
self._patch(["bash", "cargo"])
agents._build_running(self.cfg, self.agent)
self.assertNotIn("1000", self.ps_targets) # root pane_pid never ps-inspected
self.assertIn("1001", self.ps_targets)
def test_custom_build_procs_re_override(self):
cfg = dict(self.cfg)
cfg["watchdog"] = dict(self.cfg["watchdog"], build_procs_re=r"^mybuild$")
self._patch(["mybuild"])
self.assertTrue(agents._build_running(cfg, self.agent))
self._patch(["cargo"]) # default tool no longer counts
self.assertFalse(agents._build_running(cfg, self.agent))
class TestBuildAwareStall(unittest.TestCase):
"""stall_check_one defers the kill+reboot while a build runs, up to the stall_idle_max cap."""
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix="aotest-ut-")
self.cfg = agents.load_config(_make_project(self.tmp))
self.agent = self.cfg["agents"]["cl"] # stall_idle=300, stall_idle_max default 1800
self.session = self.agent["session"]
self.reboots = []
self._orig = {}
def patch(name, fn):
if name not in self._orig: # capture the TRUE original once, so a
self._orig[name] = getattr(agents, name) # test that re-patches doesn't leak the
setattr(agents, name, fn) # earlier patch into tearDown's restore
patch("session_alive", lambda s: True)
patch("capture_pane", lambda *a, **k: "")
patch("limit_tick", lambda *a, **k: False)
patch("pane_active", lambda *a, **k: False)
patch("_parse_waiting_until", lambda *a, **k: None)
patch("_pane_last_active", lambda s: None)
patch("phases", lambda cfg: []) # skip the DONE-nudge branch
patch("log", lambda *a, **k: None)
patch("start_agent", lambda cfg, a, force=False: self.reboots.append(a["name"]))
self.patch = patch
agents._idle_since.clear(); agents._build_deferred.discard(self.session)
def tearDown(self):
for name, fn in self._orig.items():
setattr(agents, name, fn)
agents._idle_since.clear(); agents._build_deferred.discard(self.session)
shutil.rmtree(self.tmp, ignore_errors=True)
def _set_idle(self, seconds):
agents._idle_since[self.session] = time.time() - seconds
def test_defers_reboot_while_building(self):
self.patch("_build_running", lambda *a, **k: True)
self._set_idle(700) # > stall_idle(300), < max(1800)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, []) # deferred, not rebooted
self.assertIn(self.session, agents._build_deferred)
def test_reboots_when_idle_and_no_build(self):
self.patch("_build_running", lambda *a, **k: False)
self._set_idle(700)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, [self.agent["name"]]) # genuinely idle → rebooted
def test_hard_cap_reboots_even_while_building(self):
self.patch("_build_running", lambda *a, **k: True)
self._set_idle(2000) # > stall_idle_max(1800)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, [self.agent["name"]]) # cap reached → reboot despite build
def test_waiting_until_defers_reboot(self):
# agent signalled a remote run in progress → hold off even with no local build + long idle
self.patch("_build_running", lambda *a, **k: False)
self.patch("_parse_waiting_until", lambda *a, **k: time.time() + 600)
self._set_idle(3000)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, []) # deferred until the stated deadline
def test_waiting_until_capped_reboots_when_no_build(self):
# a runaway that parked itself and is doing NOTHING still reboots at the cap ("max no matter
# what"), however far out its stated deadline is
self.cfg["watchdog"]["waiting_until_max"] = 7200
self.patch("_build_running", lambda *a, **k: False)
self.patch("_parse_waiting_until", lambda *a, **k: time.time() + 100000)
self._set_idle(8000) # idle > cap(7200), nothing running
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, [self.agent["name"]]) # idle past cap, no build → rebooted
def test_live_build_defers_past_the_stated_deadline(self):
# the deadline is only the agent's ESTIMATE, and an agent blocked on a shell cannot re-emit a
# fresh marker — so a still-running build (cargo-mutants overrunning its guess) must not be
# killed at the estimate. Proof of life outranks the deadline.
self.cfg["watchdog"]["waiting_until_max"] = 7200
self.patch("_build_running", lambda *a, **k: True)
self.patch("_parse_waiting_until", lambda *a, **k: time.time() - 1000) # deadline already past
self._set_idle(3000) # but still under the absolute cap
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, []) # build running → deferred
def test_past_deadline_with_no_build_reboots(self):
# no build, deadline blown → the self-wake did not fire → reboot
self.patch("_build_running", lambda *a, **k: False)
self.patch("_parse_waiting_until", lambda *a, **k: time.time() - 1000)
self._set_idle(3000)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, [self.agent["name"]])
def test_cap_is_absolute_and_reboots_a_hung_build(self):
# the cap is the ONE absolute bound: past it we reboot even mid-build — that is precisely how
# a genuinely HUNG build gets caught, and why a runaway can never park forever
self.cfg["watchdog"]["waiting_until_max"] = 7200
self.patch("_build_running", lambda *a, **k: True)
self.patch("_parse_waiting_until", lambda *a, **k: time.time() + 100000)
self._set_idle(8000) # idle > cap, build "running" (hung)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, [self.agent["name"]])
def test_no_build_check_below_base_threshold(self):
def boom(*a, **k):
raise AssertionError("_build_running must not be consulted below stall_idle")
self.patch("_build_running", boom)
self._set_idle(100) # < stall_idle(300)
agents.stall_check_one(self.cfg, self.agent)
self.assertEqual(self.reboots, [])
if __name__ == "__main__":
unittest.main(verbosity=2)