Skip to content

v1.84.0 ·

The build that failed without running

The problem

advancePr in packages/cli/src/autopilot/driver.ts counts every check in bucket fail or cancel as a failed build. It spends one of the item's two CI tries and relaunches the agent to fix code that was never tested. The second such failure parks the item.

What the real jobs looked like on 2026-10-10, read with gh api .../runs/<id>/jobs:

  • PR #175, run 38019031121, gate: runner_name: "", steps: [], conclusion cancelled after queueing 03:00–08:24Z. No runner ever took it.
  • PR #176, run 38044963454, gate: ran on fcb214b9358e-aidlc; setup and install success, pnpm build cancelled, lint and test skipped; job conclusion failure. The runner went away mid-job. No step reported failure.
  • PR #176, run 38045705008: both jobs success on every step — the same code, once a runner held.

The roadmap item said the second case's later steps had "no conclusion". The data says cancelled then skipped. Both readings share one fact, and the rule below rests on it: no step concluded failure. A real test failure always leaves the test step at failure.

How it could be solved

The first choice was what counts as "tested". A job that no runner ever took has no steps at all. A job whose runner vanished halfway has steps that ended as cancelled or skipped. A real failure is different from both: the step that ran the tests reports failure. So that became the whole rule. A check counts against the item only when one of its steps failed. Asking whether the job had a runner name was rejected, because the job that lost its runner halfway had one. Naming the setup steps was rejected too, because those names belong to each project's workflow and would drift.

The second was how often to retry. A job that never ran is re-run once, and only once. After that the item waits and the person is told the runner is the problem. The count of re-runs is not kept by autopilot at all. GitHub already numbers each attempt of a run, whoever started it, and a new push starts again at one. The risk is the opposite mistake: a test that hangs until a timeout cancels it looks like a lost runner. That costs one extra run and a message, never an agent or a parked item.

The third was what to do when the job details cannot be read. Then the failure counts exactly as it did before. The change can only ever take away a false failure; it cannot hide a real one.

How AIDLC solves it

Autopilot no longer counts a CI job that never ran its tests as a failed build. Before it counts a failure, it reads the failing checks' jobs from GitHub Actions. A check counts only when one of its job's steps failed. A job no runner took, or whose runner went away part way, spends none of the item's two CI tries and launches no agent.

Such a job is re-run once. If it still does not run, the item waits at its pull request, shows CI untested in aidlc autopilot watch, says the runner is the problem, and sends one notification. Re-running the job once a runner is back lets autopilot carry on as usual. A job whose details cannot be read counts as before, and a real test failure is handled exactly as before.

Changes

  • packages/cli/src/autopilot/ci-jobs.ts — new; reads a check's run and job, and decides whether the failing checks tested the code.
  • packages/cli/src/autopilot/driver.ts — advancePr asks for each check's link, reads the jobs once per run, and re-runs or waits; the jobs read is allowed in a dry run.
  • packages/cli/src/autopilot/state.ts — CiState gains untested.
  • packages/content/docs/roadmap.md — step 4 of the autopilot section explains the new handling. This is the topic page for autopilot; no other page applied.
  • Spec: autopilot gains "CI failures that never tested the code" when the instance is closed (fold dry run clean).
  • Blog post files: the-build-that-failed-without-running.