v1.52.0 ·
The agent that wrote the code also graded it
The problem
Implementation runs every task in tasks.md in one session. By task 8 the context holds tasks
1–7, and the agent that wrote the code also judges it — including the tests that define done.
Nothing stops it from loosening a test until it passes.
How it could be solved
Three decisions, each choosing the plainer mechanism and accepting what it gives up.
The first was how the reviewer spots a weakened test. Asking it to notice would make the answer depend on how closely it looked. Instead one fixed git command lists every test or spec file that existed before the task and was changed, deleted or renamed, plus the requirements and the task list, so the definition of done cannot be edited either. Two reviewers of the same task see the same list. Judgment is left for one question: did the task ask for this change. The filter also catches some files that are not tests, which is the safe side to be wrong on.
The second was when work is kept. The agent doing a task does not commit it; the commit waits for the reviewer's pass. So a rejected attempt never reaches the branch, and each review sees one task and nothing else. A rejection goes to a new agent with the reasons. A second rejection ends it: the main session finishes the task or puts it off with a dated reason. Some similar tools keep retrying and debugging on their own; this one stops at two.
The third was how tasks run at the same time. A list of what each task depends on was the richer choice, but it needs a rule for tasks that list nothing and a check for loops. A single number, a wave, shared by tasks next to each other in the list needs neither, and the order is still the list's order. Those tasks share one copy of the code; separate copies were ruled out from the start. So two agents can still collide on a file. That risk was accepted, held back only by the rule that tasks in one wave touch different files and by commits waiting for review.
How AIDLC solves it
One agent per task. On Claude Code, the implementation skill now offers to hand each task in
tasks.md to a fresh agent, with a second agent that edits nothing reviewing it. The reviewer gets
one fixed git diff listing every test or spec file that existed before the task and was changed,
deleted or renamed, and fails each one the task did not ask to change. It is asked once per
instance, recorded in implementation-questions.md, and off by default because it roughly doubles
the agents.
A wave marker in tasks.md. (wave: <n>) on adjacent tasks says they may run at the same
time. A bad marker keeps the task and warns; a wave split by another task warns; aidlc gate --json carries wave on each task entry that has one.
Minor, not patch: a new optional behaviour in the implementation skill and a new task-line form.
A tasks.md without waves parses exactly as before; no gate, transition or state format changes.
Changes
packages/cli/src/tasks/parser.ts,types.ts,gate-criterion.ts— the wave marker.packages/content/capabilities/claude-code.yaml—subagentsentry for the implementation phase; it lands in six Claude Code skill entries (implementation, continue, add-skill, add-action, autopilot, autopilot-stop).packages/content/resources/rules/task-agents.md— the per-task loop (new rule file).packages/content/skills/30-design.md(wave marker),40-implementation.md(the offer),00-overview.md(the rule's path; scope-table separator shortened to stay under 12,000 bytes).- Regenerated
.aidlc/skills/,.aidlc/resources/rules/,.claude/skills/,.agents/skills/.CLAUDE.md/AGENTS.mdare deliberately unchanged (see code-complete.md). - Tests:
tasks-wave.test.ts,task-agents-content.test.ts(new); additions tocli-gate-tasks.test.ts; pins updated in bothoverview-rules.test.tsfiles andskill-code-contract.test.ts. - Spec delta: ADDED
task-breakdown: Wave marker on task linesand ADDEDtask-agents: One agent per task, checked by a reviewer— two new spec files at completion (spec fold --dry-runis clean).