Skip to content
arula

Stage 120 minutesPairs301 · Tooled judgment

Diagnose the diff

Run workbench diagnose, inspect the AI-authored refund-retry diff and decide which failure modes are plausible before selecting a review or security method.

You will produce

Risk surface YAML

We will demonstrate

From a change signal to a plausible failure mode

Done when

Every plausible failure mode cites a signal in the requirement or change.

Sibling repository../payments-validation-fixture836a75dThe preparation guide explains how to verify a newer revision before using it.

Orient

Why this stage exists

The concrete change is an AI agent adding refund-retry support to the payments service.

The diff may log the provider request and leak card data, include a weak happy-path test that misses broken retry behavior, silently regress unrelated payment behavior, and leave refund idempotency unspecified.

Start from

What you start from

Work from these inputs only. Anything not on this list is either a later stage’s concern or a decision you do not own.

Product spec
specs/product/payments.md, owned by payments-product.
Technical spec
specs/tech/payments.md, the invariants and money contract.
Change
The diff between main and the seeded round-0 branch.
Failure classes
F1 through F8 in corpus/classes.ts.
Demonstrate

From the refund-retry diff to a risk surface

One worked pass, so the shape of the work is visible before you do it on your own.

  1. Run workbench diagnose to create the first YAML artifact.
  2. Read the changed logging, timeout path, tests and payment scope.
  3. Mark card-data leakage, weak retry tests, silent regression and missing refund idempotency as plausible or ruled out.
  4. Rank each plausible risk by the cost of missing it.

What this turns onDiagnosis says what could be wrong. It does not call a signal a finding or choose the next method.

Practice

Do the work

  1. Inspect the refund-retry diff against main.
  2. Complete the generated risk surface with plausibility, costOfMissing and findableByReading.
  3. Challenge any claim that a green happy-path test rules out broken retries.
  4. Do not run validation yet.
Output

You produce

Artifact

Risk surface YAML

  • Failure mode
  • Observable signal
  • Plausible or ruled out
  • Cost of missing
  • Findable by reading
Pressure-test

Challenge your own result

The refund test is green. Does that rule out a broken timeout retry or a second refund?

Show how this resolves

No. The test covers only the happy path. It does not examine timeout behavior, repeated submission, card-data sinks or whether refund idempotency was ever required.

Gate

Ready to continue when

Readiness gate
  • Every plausible failure mode cites a signal in the requirement or change.
  • Every ruled-out mode has a reason another pair can challenge.

What transfersDiagnosing before selecting a tool applies to any agent-authored change, regardless of stack.