v1.84.0 ·
The build that failed without running
The problem
advancePr in packages/cli/src/autopilot/driver.ts counts every check in bucket fail or
cancel as a failed build. It spends one of the item's two CI tries and relaunches the agent to
fix code that was never tested. The second such failure parks the item.
What the real jobs looked like on 2026-10-10, read with gh api .../runs/<id>/jobs:
- PR #175, run 38019031121,
gate:runner_name: "",steps: [], conclusioncancelledafter queueing 03:00–08:24Z. No runner ever took it. - PR #176, run 38044963454,
gate: ran onfcb214b9358e-aidlc; setup and installsuccess,pnpm buildcancelled, lint and testskipped; job conclusionfailure. The runner went away mid-job. No step reportedfailure. - PR #176, run 38045705008: both jobs
successon every step — the same code, once a runner held.
The roadmap item said the second case's later steps had "no conclusion". The data says
cancelled then skipped. Both readings share one fact, and the rule below rests on it: no
step concluded failure. A real test failure always leaves the test step at failure.
How it could be solved
The first choice was what counts as "tested". A job that no runner ever took has no steps at all. A job whose runner vanished halfway has steps that ended as cancelled or skipped. A real failure is different from both: the step that ran the tests reports failure. So that became the whole rule. A check counts against the item only when one of its steps failed. Asking whether the job had a runner name was rejected, because the job that lost its runner halfway had one. Naming the setup steps was rejected too, because those names belong to each project's workflow and would drift.
The second was how often to retry. A job that never ran is re-run once, and only once. After that the item waits and the person is told the runner is the problem. The count of re-runs is not kept by autopilot at all. GitHub already numbers each attempt of a run, whoever started it, and a new push starts again at one. The risk is the opposite mistake: a test that hangs until a timeout cancels it looks like a lost runner. That costs one extra run and a message, never an agent or a parked item.
The third was what to do when the job details cannot be read. Then the failure counts exactly as it did before. The change can only ever take away a false failure; it cannot hide a real one.
How AIDLC solves it
Autopilot no longer counts a CI job that never ran its tests as a failed build. Before it counts a failure, it reads the failing checks' jobs from GitHub Actions. A check counts only when one of its job's steps failed. A job no runner took, or whose runner went away part way, spends none of the item's two CI tries and launches no agent.
Such a job is re-run once. If it still does not run, the item waits at its pull request, shows
CI untested in aidlc autopilot watch, says the runner is the problem, and sends one
notification. Re-running the job once a runner is back lets autopilot carry on as usual. A job
whose details cannot be read counts as before, and a real test failure is handled exactly as
before.
Changes
packages/cli/src/autopilot/ci-jobs.ts— new; reads a check's run and job, and decides whether the failing checks tested the code.packages/cli/src/autopilot/driver.ts—advancePrasks for each check's link, reads the jobs once per run, and re-runs or waits; the jobs read is allowed in a dry run.packages/cli/src/autopilot/state.ts—CiStategainsuntested.packages/content/docs/roadmap.md— step 4 of the autopilot section explains the new handling. This is the topic page forautopilot; no other page applied.- Spec:
autopilotgains "CI failures that never tested the code" when the instance is closed (fold dry run clean). - Blog post files:
the-build-that-failed-without-running.