Skip to content
arula

Stage 2 / 18–20 min

Stage 2: Author & validate the spec

Turn the intended change into a buildable specification.

Your orientation Stage 2 · Define / PlanIntuition · Common language · Habit

Where you are in the work

  1. Define / Plan
  2. Execute
  3. Judge
  4. Learn

What you are developing

Intuition

Recognize ambiguity that would force an agent to invent behavior.

Build this understanding
Common language / Standardization

Use testable acceptance criteria, authority, and explicit exclusions.

Use the shared language
Habit / Behavior

Trace requirements to evidence and validate readiness before planning.

Practice this behavior
STAGE 2 · AUTHOR & VALIDATE THE SPECMake the Spec Buildable18–20 min

Concept — spec-as-context and readiness gates

You leave with — a specification an agent can build from without guessing

Whatever stays vague here becomes an invention later. Not might. Does.

The validated specification is the bounded authority every agent in Stage 4 will build from. It is the highest-leverage document in the lab, and right now it is not good enough.

⌘  python3 .claude/scripts/validate_spec.py

It will refuse the specification and name the failing checks. They map to the eight checks in docs/SPEC_COMPLETENESS_BAR.md.

The eight readiness checks

Specification Completeness Bar

Eight checks. A specification clears the bar when all eight hold.

The deterministic gate — python3 .claude/scripts/validate_spec.py — checks the structural half. The rest is a human judgement, and the status file says so rather than implying a machine blessed the content.

# Check What it means in practice
1 Ownership is named Every contested decision has exactly one owning service. A decision with two owners is a decision that will drift.
2 Acceptance criteria are executable Each criterion describes something a test can observe. “Handled correctly” is not an outcome.
3 Out-of-scope survives intact Nothing excluded quietly reappears as in-scope because it was convenient.
4 Unknowns remain explicit Open questions are still recorded as open. Closing one by picking a plausible answer is the failure this check exists for.
5 Error semantics are named where source-backed Which condition produces which status, where the source says so.
6 No invented field, default, service or endpoint Every named thing traces to source material or is labelled a lab representation.
7 Source silence is surfaced, not guessed Where the material does not say, the specification says that it does not say.
8 Compatibility impact is identified Which combinations of old and new must keep working, and which must be shown to fail.

Why 4 and 7 are the ones that matter most here

The other six are the sort of thing a careful reviewer catches. These two are the ones an agent will quietly close for you, because a specification with no open questions reads finished. In payments, an invented refund threshold, retry count or expiry window is a business decision made by something with no authority to make it — and it will read as perfectly reasonable right up until it moves the wrong amount of money.

A specification that says “we do not know this yet” is more finished than one that guessed.


Download the unchanged original source

Intuition

The facilitator demonstrates — one weak requirement becoming testable

   ✗  WEAK
      "Handle duplicate refunds correctly."

   ✓  BUILDABLE
      "When Payment Processor identifies a duplicate logical refund, the TTA
       boundary preserves the duplicate-conflict semantics, and the downstream
       repository contains no second refund record."

Look at what actually changed. The second version names who decides, what the caller observes, and what must be true of stored state afterwards. Three things a test can check.

The first names none of them. The word “correctly” was carrying the entire requirement, and “correctly” is exactly where an agent inserts its own judgement.

Common language / Standardization

▶ Your turn — harden the rest

The specification carries several more weaknesses of the same shape. Work through them with your agent and re-run the gate until it reports READY.

Three rules while you do:

1 · Do not close an open question by answering it.

OQ-1 asks how the idempotency key is derived in production. The source material states the duplicate rule and the status code, and never states the derivation.

⚠ Trap — the strongest one in this lab. Your agent will offer you a reasonable-sounding derivation. It will look like diligence. In payments, an invented key derivation is a business decision made by something with no authority to make it — and it will read as perfectly sensible right up until it moves the wrong amount of money.

It stays open. Refusing to answer it is the single most important thing you do today.

2 · Do not weaken the out-of-scope list to make something fit. It is write-protected, so the gate will stop you. The instinct is the thing worth noticing in yourself.

3 · Everything you add traces to docs/PGS_DECISIONS.md — as a Meridian fact, or as an explicitly labelled lab representation. Nothing gets invented into existence.

Habit / Behavior

Human gate before you move on

Read your hardened specification once more and ask one question:

Did we invent any Meridian behaviour to get here?

If yes, take it out and put the question back.

The gate is structural. It tells you the specification is well-formed, never that it is right — which is why the status file records "semantic_authority": "human-reviewed" rather than quietly implying a machine approved the content.

⌘  /hand-off

⏸ Q&A pause — 3 min

Domain, spec, or gate questions. Environment problems go to the parking lot instead of into the room.


Carry the work forward

Complete the stage’s instructions and hand-off above before continuing. Keep your evidence and unresolved questions with the work.

If you fall behind or need help