Skip to content
arula
301Tooled judgment

Diagnose, review, eval and repair AI-generated code.

Work through a repeatable loop: diagnose the diff, choose review methods, evaluate evidence gaps, define defect specs, plan, run and repeat the first three stages on the new diff.

Stages
7
175 minutes of instructional time
Fixture
Payments
AI-added refund-retry support
Handoffs
6
YAML → defects/*.md → tasks.json → new diff
The practice170 minutes
01
Diagnosethe diff and its plausible risks
Stage 1
02
Reviewthe diagnosis with chosen methods
Stage 2
03
Evalthe evidence and its gaps
Stage 3
04
Defineauditable defect specs
Stage 4
05
Planthe specs into tasks.json
Stage 5
06
Runthe plan and create a new diff
Stage 6
07
RepeatDiagnose → Review → Eval
Stage 7

Current fixture behavior still exposes diagnose, plan, validate and report. V1 targets diagnose → review → eval → define → plan → run. Stage 2 selects methods before results are shown, and Stage 7 remains blocked until the unseen assessment change exists.

Learning outcomes

What these stages
should change.

The stages succeed when practitioners recognize the same risks, describe them in the same words and change what they do before approving work.

Outcome 01

Shared intuition

Engineers recognize common AI failure modes and understand why different risks require different validation methods.

ObservableGiven an AI-generated change, learners identify plausible failure modes before selecting a method.

Outcome 02

Common language

Teams consistently describe failure modes, validation methods, evidence strength, unexamined areas and human decisions.

ObservableLearners explain a validation decision using the same terms for risk, method, evidence and uncertainty.

Outcome 03

Everyday habits

Before approving AI-generated code, engineers diagnose risk, judge evidence, define bounded repairs and repeat the loop after every new diff.

ObservableLearners perform the workflow on an unfamiliar change without a prescribed tool list.

The learning path

Seven stages.
One decision each.

Each stage adds one decision and produces an artifact the next stage works from. Preparation happens before instructional time.

··

Prepare

Sibling repository: ../payments-validation-fixture. Complete this before the first stage.

Before the stagesOutside instructional time
01

Diagnose the diff

Run workbench diagnose, inspect the AI-authored refund-retry diff and decide which failure modes are plausible before selecting a review or security method.

Risk surface YAML20 minutes · Pairs
02

Review the diagnosis

Choose security and code-review methods before seeing their results, then run workbench review to test whether the diagnosed risks are supported by the diff and requirements.

Review YAML25 minutes · Pairs
03

Eval the evidence gaps

Run workbench eval on the review output, find gaps in the tests and evidence, and separate code defects, test gaps and specification gaps.

Eval YAML25 minutes · Pairs
04

Define the defect specs

Run workbench define to turn newly identified code defects and test gaps into bounded, auditable defects/*.md files. Route policy gaps to an accountable human instead of inventing requirements.

defects/1.md and defects/2.md20 minutes · Pairs
05

Plan the spec

Run workbench plan to turn defects/*.md into tasks.json: an executable implementation and validation plan with explicit checks, gates, budget and omissions.

tasks.json25 minutes · Pairs
06

Run the plan

Run workbench run with tasks.json, apply only the approved repair scope and inspect the new implementation diff and run output.

New implementation diff and run output30 minutes · Pairs
07

Repeat Diagnose → Review → Eval

After run creates a new implementation diff, repeat workbench diagnose, workbench review and workbench eval independently before approving, defining another defect or escalating policy.

Loop evaluation record25 minutes · Individual
››

The workbench

Failure classes and the checks that examine them

ReferenceCarried through every stage

Current versus V1

The fixture today is not the V1 target.

The target sequence is explicit about what exists now, what the new stages require and where delivery is still blocked.

Current fixture behaviorExisting command surface

The current fixture exposes diagnose, plan, validate and report. validate executes the workflow and report summarizes its evidence. Results may appear before Stage 2 method selection.

V1 targetSix commands, then repeat

diagnose → review → eval → define → plan → run. Review methods are chosen before results are shown; define writes defects/*.md; plan writes tasks.json; run creates the next diff. Stage 7 is blocked until the unseen assessment change exists.

Where this sits

Before and after
these stages.

Tracks at the same number are peers on the same fixture. The 501 assumes both of them.