Diagnose, review, eval and repair AI-generated code.
Work through a repeatable loop: diagnose the diff, choose review methods, evaluate evidence gaps, define defect specs, plan, run and repeat the first three stages on the new diff.
Current fixture behavior still exposes diagnose, plan, validate and report. V1 targets diagnose → review → eval → define → plan → run. Stage 2 selects methods before results are shown, and Stage 7 remains blocked until the unseen assessment change exists.
01
Learning outcomes
What these stages should change.
The stages succeed when practitioners recognize the same risks, describe them in the same words and change what they do before approving work.
Outcome 01
Shared intuition
Engineers recognize common AI failure modes and understand why different risks require different validation methods.
ObservableGiven an AI-generated change, learners identify plausible failure modes before selecting a method.
Outcome 02
Common language
Teams consistently describe failure modes, validation methods, evidence strength, unexamined areas and human decisions.
ObservableLearners explain a validation decision using the same terms for risk, method, evidence and uncertainty.
Outcome 03
Everyday habits
Before approving AI-generated code, engineers diagnose risk, judge evidence, define bounded repairs and repeat the loop after every new diff.
ObservableLearners perform the workflow on an unfamiliar change without a prescribed tool list.
02
The learning path
Seven stages. One decision each.
Each stage adds one decision and produces an artifact the next stage works from. Preparation happens before instructional time.
The target sequence is explicit about what exists now, what the new stages require and where delivery is still blocked.
Current fixture behaviorExisting command surface
The current fixture exposes diagnose, plan, validate and report. validate executes the workflow and report summarizes its evidence. Results may appear before Stage 2 method selection.
V1 targetSix commands, then repeat
diagnose → review → eval → define → plan → run. Review methods are chosen before results are shown; define writes defects/*.md; plan writes tasks.json; run creates the next diff. Stage 7 is blocked until the unseen assessment change exists.
04
Where this sits
Before and after these stages.
Tracks at the same number are peers on the same fixture. The 501 assumes both of them.