Skip to content
arula

Stage 5 / 20 min

Stage 5: Validate with fresh context

Prove the pair with deterministic evidence and independent judgment.

Your orientation Stage 5 · JudgeIntuition · Common language · Habit

Where you are in the work

  1. Define / Plan
  2. Execute
  3. Judge
  4. Learn

What you are developing

Intuition

Distinguish green local tests from compatibility evidence.

Build this understanding
Common language / Standardization

State findings, dispositions, compatibility, and verification limits precisely.

Use the shared language
Habit / Behavior

Use fresh context and disposition every finding before accepting the work.

Practice this behavior
STAGE 5 · VALIDATE WITH FRESH CONTEXTProve the Pair ★ NEVER CUT20 min

Concept — fresh-context validation and independent judgment

You leave with — evidence you did not produce, and rulings on what it found

Separate creation from judgment.

If Stage 4 runs long, the facilitator will apply a checkpoint and move the room here anyway. Arriving with four of five fixes done and seeing what independent judgment catches is a far better session than finishing the code and never finding out.

Intuition

◆ Predict #4

Both repositories are green and you believe the work is done. What will a fresh validator, or the pair harness, catch that your green builds did not?

One thing, written down, before you run either.

Common language / Standardization

1 · Deterministic evidence first

⌘  python3 scripts/run_pair_verification.py

This answers the only question that matters: are the two services actually in agreement? Each failure names the disagreement in its own message.

flowchart LR
  accTitle: The compatibility matrix
  accDescr: Four distinct consumer-to-producer cases. Previous to previous is the baseline, the world before this change. Previous to current must succeed. Current to previous must be shown to fail. Current to current is the intended final state.
  A["previous consumer"] -->|"baseline. The world before this change."| A1["previous producer"]
  B["previous consumer"] -->|"MUST SUCCEED"| B1["current producer"]
  C["current consumer"] -->|"MUST BE SHOWN TO FAIL"| C1["previous producer"]
  D["current consumer"] -->|"the intended final state"| D1["current producer"]

That third row is the interesting one. The test passes by detecting the incompatibility — it is diagnostic, not permanently red. It is how the rollout order stops being an assertion somebody made and becomes something you can point at.

2 · Then independent judgment

⌘  python3 .claude/scripts/build_validator_brief.py

The brief is assembled mechanically from an allowlist: the specification, both diffs, the ledger, the pair results, the plan, the scope documents.

It contains no chat history and no builder rationale — not because the script is careful about leaving them out, but because it has no way to reach them. That is a structural guarantee rather than an instruction to be discreet, which is exactly why it is a script and not a prompt.

Dispatch the fresh code-to-spec-validator against it. It has read and test tools and no write tools, so it cannot quietly repair what it finds. It never saw your session, so it cannot inherit your confidence in your own work.

Habit / Behavior

3 · Disposition every finding

In docs/finding-dispositions.md:

  │ Finding │ Evidence │ In scope? │ Material? │ Disposition │ Rationale │

⚠ A validator finding does not authorise a code change.

Some findings are correct and out of scope. Some are simply wrong. Some are right but immaterial. Sorting them is the judgment this stage exists to build — and you are graded on the disposition, not on agreeing with the validator.

Reveal: compare against Prediction #4.

⌘  /hand-off

⏸ Q&A pause — 3 min


Carry the work forward

Complete the stage’s instructions and hand-off above before continuing. Keep your evidence and unresolved questions with the work.

If you fall behind or need help