Stage 5 / 20 min
Stage 5: Validate with fresh context
Prove the pair with deterministic evidence and independent judgment.
Your orientation Stage 5 · JudgeIntuition · Common language · Habit
Where you are in the work
- Define / Plan
- Execute
- Judge
- Learn
What you are developing
Distinguish green local tests from compatibility evidence.
Build this understandingState findings, dispositions, compatibility, and verification limits precisely.
Use the shared languageUse fresh context and disposition every finding before accepting the work.
Practice this behaviorConcept — fresh-context validation and independent judgment
You leave with — evidence you did not produce, and rulings on what it found
Separate creation from judgment.
If Stage 4 runs long, the facilitator will apply a checkpoint and move the room here anyway. Arriving with four of five fixes done and seeing what independent judgment catches is a far better session than finishing the code and never finding out.
Intuition
◆ Predict #4
Both repositories are green and you believe the work is done. What will a fresh validator, or the pair harness, catch that your green builds did not?
One thing, written down, before you run either.
Common language / Standardization
1 · Deterministic evidence first
⌘ python3 scripts/run_pair_verification.py
This answers the only question that matters: are the two services actually in agreement? Each failure names the disagreement in its own message.
flowchart LR
accTitle: The compatibility matrix
accDescr: Four distinct consumer-to-producer cases. Previous to previous is the baseline, the world before this change. Previous to current must succeed. Current to previous must be shown to fail. Current to current is the intended final state.
A["previous consumer"] -->|"baseline. The world before this change."| A1["previous producer"]
B["previous consumer"] -->|"MUST SUCCEED"| B1["current producer"]
C["current consumer"] -->|"MUST BE SHOWN TO FAIL"| C1["previous producer"]
D["current consumer"] -->|"the intended final state"| D1["current producer"]
That third row is the interesting one. The test passes by detecting the incompatibility — it is diagnostic, not permanently red. It is how the rollout order stops being an assertion somebody made and becomes something you can point at.
2 · Then independent judgment
⌘ python3 .claude/scripts/build_validator_brief.py
The brief is assembled mechanically from an allowlist: the specification, both diffs, the ledger, the pair results, the plan, the scope documents.
It contains no chat history and no builder rationale — not because the script is careful about leaving them out, but because it has no way to reach them. That is a structural guarantee rather than an instruction to be discreet, which is exactly why it is a script and not a prompt.
Dispatch the fresh code-to-spec-validator against it. It has read and test tools and no write
tools, so it cannot quietly repair what it finds. It never saw your session, so it cannot
inherit your confidence in your own work.
Habit / Behavior
3 · Disposition every finding
In docs/finding-dispositions.md:
│ Finding │ Evidence │ In scope? │ Material? │ Disposition │ Rationale │
⚠ A validator finding does not authorise a code change.
Some findings are correct and out of scope. Some are simply wrong. Some are right but immaterial. Sorting them is the judgment this stage exists to build — and you are graded on the disposition, not on agreeing with the validator.
Reveal: compare against Prediction #4.
⌘ /hand-off
⏸ Q&A pause — 3 min
Carry the work forward
Complete the stage’s instructions and hand-off above before continuing. Keep your evidence and unresolved questions with the work.
If you fall behind or need help