Skip to content

v1.52.0 ·

The agent that wrote the code also graded it

The problem

Implementation runs every task in tasks.md in one session. By task 8 the context holds tasks 1–7, and the agent that wrote the code also judges it — including the tests that define done. Nothing stops it from loosening a test until it passes.

How it could be solved

Three decisions, each choosing the plainer mechanism and accepting what it gives up.

The first was how the reviewer spots a weakened test. Asking it to notice would make the answer depend on how closely it looked. Instead one fixed git command lists every test or spec file that existed before the task and was changed, deleted or renamed, plus the requirements and the task list, so the definition of done cannot be edited either. Two reviewers of the same task see the same list. Judgment is left for one question: did the task ask for this change. The filter also catches some files that are not tests, which is the safe side to be wrong on.

The second was when work is kept. The agent doing a task does not commit it; the commit waits for the reviewer's pass. So a rejected attempt never reaches the branch, and each review sees one task and nothing else. A rejection goes to a new agent with the reasons. A second rejection ends it: the main session finishes the task or puts it off with a dated reason. Some similar tools keep retrying and debugging on their own; this one stops at two.

The third was how tasks run at the same time. A list of what each task depends on was the richer choice, but it needs a rule for tasks that list nothing and a check for loops. A single number, a wave, shared by tasks next to each other in the list needs neither, and the order is still the list's order. Those tasks share one copy of the code; separate copies were ruled out from the start. So two agents can still collide on a file. That risk was accepted, held back only by the rule that tasks in one wave touch different files and by commits waiting for review.

How AIDLC solves it

One agent per task. On Claude Code, the implementation skill now offers to hand each task in tasks.md to a fresh agent, with a second agent that edits nothing reviewing it. The reviewer gets one fixed git diff listing every test or spec file that existed before the task and was changed, deleted or renamed, and fails each one the task did not ask to change. It is asked once per instance, recorded in implementation-questions.md, and off by default because it roughly doubles the agents.

A wave marker in tasks.md. (wave: <n>) on adjacent tasks says they may run at the same time. A bad marker keeps the task and warns; a wave split by another task warns; aidlc gate --json carries wave on each task entry that has one.

Minor, not patch: a new optional behaviour in the implementation skill and a new task-line form. A tasks.md without waves parses exactly as before; no gate, transition or state format changes.

Changes

  • packages/cli/src/tasks/parser.ts, types.ts, gate-criterion.ts — the wave marker.
  • packages/content/capabilities/claude-code.yaml — subagents entry for the implementation phase; it lands in six Claude Code skill entries (implementation, continue, add-skill, add-action, autopilot, autopilot-stop).
  • packages/content/resources/rules/task-agents.md — the per-task loop (new rule file).
  • packages/content/skills/30-design.md (wave marker), 40-implementation.md (the offer), 00-overview.md (the rule's path; scope-table separator shortened to stay under 12,000 bytes).
  • Regenerated .aidlc/skills/, .aidlc/resources/rules/, .claude/skills/, .agents/skills/. CLAUDE.md/AGENTS.md are deliberately unchanged (see code-complete.md).
  • Tests: tasks-wave.test.ts, task-agents-content.test.ts (new); additions to cli-gate-tasks.test.ts; pins updated in both overview-rules.test.ts files and skill-code-contract.test.ts.
  • Spec delta: ADDED task-breakdown: Wave marker on task lines and ADDED task-agents: One agent per task, checked by a reviewer — two new spec files at completion (spec fold --dry-run is clean).