Skip to content
arula

Meridian Engineering / Lab 2

Prepare for the lab

Start setup, understand the seam, and learn the shared language before the session.

Your orientation Prepare · Define / PlanIntuition · Common language · Habit

Where you are in the work

  1. Define / Plan
  2. Execute
  3. Judge
  4. Learn

What you are developing

Intuition

Use the refund examples to explain how two passing repositories can still disagree at a service boundary.

Build this understanding
Common language / Standardization

Read the grounding and term card here: distinguish facts, lab representations, and seeded failures; learn the ten shared terms.

Use the shared language
Habit / Behavior

Start setup, read while Maven downloads, and confirm “Setup complete”. Bring uncertainties to the session.

Practice this behavior

Start setup first. Read while Maven downloads, then confirm “Setup complete” before the session.

Define / Plan · Habit / Behavior

Get the workspace ready.

Prerequisites

  • JDK 17 or newer. A 21 JDK compiling to the Java 17 target is the documented Meridian pattern and is what this lab is built and tested against.
  • Maven 3.9+
  • Python 3.9+, with PyYAML for the grader (python3 -m pip install pyyaml)
  • Git
  • The workbench plugin, for /lab, /journey, /hand-off, the planner and the code-to-spec-validator
  • A warm ~/.m2. The lab makes no network calls at runtime, but a first Maven build on a cold cache resolves dependencies like any other.

Setup — before session day, not during it

python3 scripts/verify_setup.py

This checks the toolchain, creates the two service repositories from starter/ (each a real Git repository with one committed starter commit), warms the Maven cache, and confirms the starting state. It must end with “Setup complete”.

--check verifies without changing anything. --reset is destructive: it discards the working copies and restores them from starter/, and asks for confirmation naming exactly what it will delete.


Define / Plan · Intuition, shared language & habits

Understand the work before the session.

Pre-read — Lab 2

10–15 minutes, before session day. Setup takes longer than reading this, so start with the setup step and read while Maven downloads.


1. Setup first

python3 scripts/verify_setup.py

This checks your toolchain, creates the two service repositories, warms the Maven cache and confirms the starting state. It must end with “Setup complete”. If it does not, bring the output to the facilitator before the session rather than during it.

A cold Maven cache is the single most common way a room loses fifteen minutes. Run this at your desk, on the machine you will use.


2. The situation, in plain terms

A merchant took a payment. Now they need to refund it.

The refund request arrives in a WSAPI-facing shape, is translated by TTA, and is sent to Payment Processor, which decides it and records it.

flowchart TB
  accTitle: The merchant refund journey
  accDescr: A merchant refund request passes through WSAPI to TTA, the translation boundary, then Payment Processor, the refund decision. The wider Meridian flow continues through CPC, Injection, LCS, and DCF; these downstream services are not implemented in this lab.
  R["Merchant refund request"] --> W["WSAPI"]
  W --> T["TTA · translation boundary"]:::focus
  T --> P["Payment Processor · refund decision"]:::focus
  P --> C
  subgraph downstream["Wider Meridian flow continues · not implemented in this lab"]
    C["CPC"] --> I["Injection"] --> L["LCS"] --> D["DCF"]
  end

You work on the TTA → Payment Processor portion only.

Your task: bring that portion into agreement — across contract, implementation, tests and technical requirements — without expanding into the rest of the flow.


3. The thing that makes it interesting

Both repositories build. Both have passing tests. Neither is obviously broken.

They still do not agree with each other.

That is the whole point of the exercise. Correctness across a service boundary lives in the relationship between two implementations, not inside either one, and no amount of testing one repository in isolation will show you a disagreement with the other.

Two examples of the shape of the problem, neither of which is a hint about where to look:

  • A service answers a duplicate refund with “this is a duplicate”. The service in front of it turns that into “the system failed”. The caller retries a request that was correctly refused.
  • Both sides implement retry protection. They protect against different retry identities. Each looks correct alone; together, the operation is not idempotent.

4. Lab versus real Meridian

The lab uses real Meridian terminology and real refund behaviour, simplified deliberately so the exercise fits in two hours.

Read SCENARIO_GROUNDING.md — it separates grounded Meridian behaviour from lab simplification from deliberately planted defect. The planted defects are teaching fixtures and do not imply anything about real Meridian systems.

The one rule worth carrying into the room: inventing lab code is fine; inventing Meridian platform behaviour is not.


Understand what is fact, representation, and fixture

Scenario grounding — what is real, what is simplified, what is planted

This lab sits in the Meridian refund domain and uses real Meridian terminology. That makes it worth being exact about which parts are grounded behaviour, which are simplifications made so the exercise fits in two hours, and which are deliberate teaching fixtures.

The rule the lab authors worked to: inventing lab code is acceptable; inventing Meridian platform behaviour is not. A failure mechanism may be simulated. The engineering principle it demonstrates may not be made up.


The three layers

Layer 1 — Meridian fact

Behaviour supported by the supplied Meridian material, recorded decision by decision in PGS_DECISIONS.md with its citation.

Examples: the wider refund flow and where TTA and Payment Processor sit in it; the request field mapping; the two refund endpoints and the identifiers that select between them; the status codes, including 409 for a duplicate; refunds not exceeding the captured amount, per order and per capture; idempotency being required on money-moving paths; correlation IDs propagating end to end; Void being out of scope; settlement being downstream.

Layer 2 — Lab representation

Deliberate simplifications, each labelled as such in PGS_DECISIONS.md.

Examples: two editable repositories; only the TTA → Payment Processor seam being executable; online/offline supplied as an input rather than modelling CPC’s decision; in-memory stores instead of Oracle; a deterministic authorization stub with no network; the contract version numbers; the specific transport field carrying the retry identity; the correlation header name; a local compatibility harness standing in for a deployment pipeline.

Layer 3 — Seeded failure

Defects introduced deliberately by the lab authors to create the exercise.

Their presence does not imply that the same defect exists, or ever existed, in a Meridian production system. They are teaching fixtures. What is grounded is the principle each one demonstrates and the behaviour that correcting it restores.

This document does not say where they are. Finding them is the work.


What the lab deliberately does not model

CPC behaviour beyond the state supplied at the seam · refund injection · LCS API · DCF generation and the settlement lifecycle · ISO 8583 and DE48 mapping · A3RS, BECS, BPSS and surrounding platform components · production routing, regional and blue-green topology · the production CI/CD, integration-environment and release-governance path · the real team-ownership and merge-approval model across repositories · merchant-privilege enforcement.

These are scope reductions for the exercise. They are not architectural claims. Nothing here says the real system lacks them.


What “success” means, precisely

The represented TTA → Payment Processor slice is internally consistent, contract-compatible, tested, and independently validated against the supplied technical authority.

That is narrower than “the refund capability works”. The lab proves the represented seam locally. It does not claim the wider Meridian refund capability is production-ready, and pair verification does not replace Meridian integration testing or release governance.


Why the grounding discipline is the point, not the paperwork

The habit this lab is trying to build is the one that matters when the AI is confident and the source is silent. An agent asked to finish a refund path will produce something plausible for a threshold nobody specified, an endpoint nobody published, or a dependency nobody built. It will read well. In payments it will also be a business decision made by something with no authority to make it.

Separating fact from representation from fixture — here, and in your own work — is what makes that difference visible before it ships.


Download the unchanged original source

5. Vocabulary

Skim TERM_CARD.md. Ten terms. You will use seam, context ledger and agent brief constantly.


Use the same language

Term Card

Ten terms. Everything else is ordinary engineering vocabulary.

Term Meaning
Context boundary What a given agent is allowed to see. Chosen deliberately, not by default.
Seam The place two services meet — here, TTA → Payment Processor. Correctness at a seam lives in the relationship between two implementations, not inside either one.
Repo brief The instruction you write for an agent auditing one repository: what to inspect, what evidence to return, and what it must not infer.
Context ledger Your reconciliation of what the agents reported: claim, where it was asserted, evidence, what contradicts it, your ruling, and its status.
Scoped sub-agent An agent deliberately given one repository and read-only tools, so its findings stay local and checkable.
Agent brief The contract for an implementation agent: outcome, authoritative inputs, repository scope, allowed and excluded areas, tools, the acceptance criteria it owns, expected return shape, and stop conditions.
Spec readiness Whether a specification is well-formed enough to build from. Structural readiness is machine-checkable; whether it is correct is not.
Compatibility matrix The combinations of old and new that must work, and the one that must be shown to fail. It is what turns rollout order from an assertion into evidence.
Fresh-context validator A judging agent that never saw the work being done, so it cannot inherit the builder’s confidence.
Hand-off The written checkpoint at a stage boundary: what we concluded, what evidence exists, what remains open, what the next context needs to know.

One deliberate naming choice

The human running the session is the coordinating engineer, not the “orchestrator”. Two reasons: “orchestration” already means something specific in payments platform work, and the human role here is not to run more agents but to decide what each is allowed to know, and to rule on what they come back with.

Journey and hand-off are different things

Journey is automatic and machine-observable — what actually happened. Hand-off is written by you at a stage boundary — what we concluded and what remains open.

Neither replaces the other. A journey with no hand-off is a trail nobody can act on; a hand-off with no journey is a claim with nothing behind it.


Download the unchanged original source

6. What is graded

Read ESSENTIAL_OUTCOMES.md.

The short version: engineering outcomes and evidence, never whether your screen matches the facilitator’s.


Know what the evidence must show

Essential Outcomes Card

Write-protected. Keep this open during the session.

What is graded

Engineering outcomes and evidence. Not whether your screen matches the facilitator’s.

Model wording, ordering and formatting will differ between people, and between two runs by the same person. That is normal and expected. We grade what you concluded, what you can show for it, and what you refused to do — never whether your response looks like anyone else’s.

If your agent phrases something differently, finds the same problem by another route, or returns its findings in a different order, you are not behind.

The eight outcomes

By the end of the session your workspace should show:

  1. The context ledger identifies the contract drift between the two repositories, with evidence.
  2. The context ledger identifies the divergent business rule, and names which service owns it.
  3. At least one unknown remains unresolved rather than filled in with a plausible default.
  4. The specification preserves its out-of-scope list — nothing excluded got built.
  5. The plan defines repository boundaries and a compatibility rationale, not just a task list.
  6. The retry identity is stable across the seam — what one side sends is what the other deduplicates against.
  7. Duplicate semantics survive the seam — a duplicate reaches the caller as a duplicate.
  8. Pair verification is green, and you can say what that does and does not prove.

Two things that are outcomes, not failures

A “no diff” result, with evidence, is a real outcome. Deciding that a repository is correct and proving it is engineering work. So is refusing to change something because it is out of scope, and recording why.

An unresolved question left visible beats a plausible answer invented to close it. In payments, guessing a threshold or a default is a business decision you were not authorised to make. Surfacing it is the correct move, and it is scored as such.

What this lab does not claim

The lab proves the represented seam locally. It does not claim the wider Meridian refund capability is production-ready, and pair verification does not replace integration testing or release governance.


Download the unchanged original source

7. Model output will vary — plan for it

Your agent will word things differently from the person next to you, find the same problem by a different route, and return findings in a different order. Running the same prompt twice will not give you the same text.

None of that means you are behind. If you find yourself trying to make your output look like the demonstration, stop and ask what the demonstration was showing you instead.


8. What this lab assumes you already have

From Lab 1: fresh-context review, sub-agents, human gates, deterministic checks, journey and hand-off. These are not re-taught. If any of them is hazy, say so early — Stage 0 is the moment for it, not Stage 4.

You do not need to know anything new about payments. You need to be willing to say “the source does not tell us that” and leave it unanswered.


Download the unchanged original source

Define / Plan · Bring this into Stage 0

Ready for the session.

Your setup must end with “Setup complete”. If it does not, bring the output to the facilitator before the session. If a Lab 1 concept is hazy, raise it at Stage 0.

Intuition

Revisit the two examples of how services can each look correct and still disagree.

The thing that makes it interesting
Common language / Standardization

Keep the ten terms and the distinction between fact, representation, and seeded failure available.

Revisit the term card
Habit / Behavior

Check the setup result and bring unresolved questions into the session.

Return to setup

In Stage 0 you will start the lab, establish authority, and seal your first prediction before investigating.

Start Stage 0

Available throughout the lab

Workspace & architecture guide.

The repository layout, commands, architecture, and scope notes to consult as you work.

Lab 2 — One Refund Across a Service Boundary

A 120-minute hands-on lab on using AI safely when a money-moving change spans more than one repository, and the decisions that matter live at the seam between them.

Lab 1 taught governing one AI-assisted change inside one repository. This lab is about orchestrating several scoped AI contexts across a service boundary while the engineer keeps authority over the seam.

Participants start with LAB_ACTION_GUIDE.md. This file covers architecture and setup.


The problem

A merchant refund crosses the TTA → Payment Processor boundary. Two repositories. Both build. Both have passing tests. Neither is obviously broken.

They still do not agree with each other.

flowchart LR
  accTitle: The complete service flow
  accDescr: WSAPI to TTA to Payment Processor to CPC to Injection to LCS API to DCF. TTA and Payment Processor: represented here. Downstream services: not implemented in this lab.
  W("WSAPI") --> T
  subgraph lab["represented here"]
    T("TTA"):::focus --> P("Payment Processor"):::focus
  end
  P --> C
  subgraph downstream["not implemented in this lab"]
    C("CPC") --> I("Injection") --> L("LCS API") --> D("DCF")
  end

Correctness across a service boundary lives in the relationship between two implementations, not inside either one — which is why the central fact of the lab is:

Repo A green + Repo B green ≠ pair correct.


Layout

specs/          the specification, plus the write-protected authority documents
docs/           pre-read, term card, outcomes, grounding, decisions, ledger, tracker
starter/        pristine source for the two services; setup copies it out
lab-harness/    pair-verification harness -- readable, not writable
.claude/        lab config, write gate, auditor agent, validators, rubric, grader
scripts/        setup and pair verification

After setup, pgs-tta/ and pgs-payment-processor/ appear at the root as independent Git repositories. They are gitignored here on purpose: nesting one repository’s history inside another’s is how solution history leaks.

The seven stages

# Stage Focus Min
0 Ground the Work Frame the Boundary 10
1 Audit Context Map the Seam 18–20
2 Author & Validate the Spec Make the Spec Buildable 18–20
3 Plan Across Repositories Design the Orchestration 14–15
4 Build & Validate Build the Bounded Slice 28–30
5 Validate with Fresh Context Prove the Pair 20
6 Review, Handoff & Close Transfer the Learning 5–6

Commands

python3 scripts/verify_setup.py             # set up and verify the workspace
python3 scripts/run_pair_verification.py    # prove the seam
python3 scripts/run_pair_verification.py --explain   # what the compatibility matrix asks

python3 .claude/scripts/validate_spec.py    # Stage 2 readiness gate
python3 .claude/scripts/validate_plan.py    # Stage 3 plan gate
python3 .claude/scripts/build_validator_brief.py     # Stage 5 brief, assembled deterministically
python3 .claude/scripts/grade_repo.py       # deterministic grading

Architecture notes

The write gate. .claude/hooks/gate_guard.py blocks writes to the harness, the out-of-scope and non-negotiables documents, the outcomes card and the facilitator package. Reading them is expected. A lab whose grading harness can be edited by the thing being graded is not measuring anything. Run --self-test to confirm all four bypass classes are still covered.

The generated client. pgs-tta/src/main/java/.../client/contract/ is generated from the contract that service is pinned to:

cd pgs-tta && mvn -Pgenerate-client generate-sources

Generation runs at construction and remediation time only, never during an ordinary build, so the lab needs no network. The same contract and configuration reproduce the committed output exactly — regenerate rather than hand-editing.

Determinism. grade_repo.py produces the same score for the same workspace state every time. Nothing samples, calls a model, or depends on the clock. Two people who reach the same outcome by different routes score identically.

A limitation, stated rather than papered over. A journey event records that a tool ran, not what it returned, so no rubric check can prove a build went green from the journey alone. Build outcomes are graded from recorded results and repository state, and the facilitator’s live spot-check covers the rest.

Grounding

docs/SCENARIO_GROUNDING.md separates grounded Meridian behaviour from lab simplification from deliberately seeded defect, and docs/PGS_DECISIONS.md records every decision with its source and layer.

Seeded defects are teaching fixtures. Their presence does not imply the same defect exists, or ever existed, in a Meridian production system.

The lab proves the represented seam locally. It does not claim the wider Meridian refund capability is production-ready, and pair verification does not replace integration testing or release governance.


Download the unchanged original source