Meridian Engineering / Lab 2
Prepare for the lab
Start setup, understand the seam, and learn the shared language before the session.
Your orientation Prepare · Define / PlanIntuition · Common language · Habit
Where you are in the work
- Define / Plan
- Execute
- Judge
- Learn
What you are developing
Use the refund examples to explain how two passing repositories can still disagree at a service boundary.
Build this understandingRead the grounding and term card here: distinguish facts, lab representations, and seeded failures; learn the ten shared terms.
Use the shared languageStart setup, read while Maven downloads, and confirm “Setup complete”. Bring uncertainties to the session.
Practice this behaviorStart setup first. Read while Maven downloads, then confirm “Setup complete” before the session.
Define / Plan · Habit / Behavior
Get the workspace ready.
Prerequisites
- JDK 17 or newer. A 21 JDK compiling to the Java 17 target is the documented Meridian pattern and is what this lab is built and tested against.
- Maven 3.9+
- Python 3.9+, with PyYAML for the grader (
python3 -m pip install pyyaml) - Git
- The
workbenchplugin, for/lab,/journey,/hand-off, the planner and thecode-to-spec-validator - A warm
~/.m2. The lab makes no network calls at runtime, but a first Maven build on a cold cache resolves dependencies like any other.
Setup — before session day, not during it
python3 scripts/verify_setup.py
This checks the toolchain, creates the two service repositories from starter/ (each a real Git
repository with one committed starter commit), warms the Maven cache, and confirms the starting
state. It must end with “Setup complete”.
--check verifies without changing anything. --reset is destructive: it discards the working
copies and restores them from starter/, and asks for confirmation naming exactly what it will
delete.
Define / Plan · Intuition, shared language & habits
Understand the work before the session.
Pre-read — Lab 2
10–15 minutes, before session day. Setup takes longer than reading this, so start with the setup step and read while Maven downloads.
1. Setup first
python3 scripts/verify_setup.py
This checks your toolchain, creates the two service repositories, warms the Maven cache and confirms the starting state. It must end with “Setup complete”. If it does not, bring the output to the facilitator before the session rather than during it.
A cold Maven cache is the single most common way a room loses fifteen minutes. Run this at your desk, on the machine you will use.
2. The situation, in plain terms
A merchant took a payment. Now they need to refund it.
The refund request arrives in a WSAPI-facing shape, is translated by TTA, and is sent to Payment Processor, which decides it and records it.
flowchart TB
accTitle: The merchant refund journey
accDescr: A merchant refund request passes through WSAPI to TTA, the translation boundary, then Payment Processor, the refund decision. The wider Meridian flow continues through CPC, Injection, LCS, and DCF; these downstream services are not implemented in this lab.
R["Merchant refund request"] --> W["WSAPI"]
W --> T["TTA · translation boundary"]:::focus
T --> P["Payment Processor · refund decision"]:::focus
P --> C
subgraph downstream["Wider Meridian flow continues · not implemented in this lab"]
C["CPC"] --> I["Injection"] --> L["LCS"] --> D["DCF"]
end
You work on the TTA → Payment Processor portion only.
Your task: bring that portion into agreement — across contract, implementation, tests and technical requirements — without expanding into the rest of the flow.
3. The thing that makes it interesting
Both repositories build. Both have passing tests. Neither is obviously broken.
They still do not agree with each other.
That is the whole point of the exercise. Correctness across a service boundary lives in the relationship between two implementations, not inside either one, and no amount of testing one repository in isolation will show you a disagreement with the other.
Two examples of the shape of the problem, neither of which is a hint about where to look:
- A service answers a duplicate refund with “this is a duplicate”. The service in front of it turns that into “the system failed”. The caller retries a request that was correctly refused.
- Both sides implement retry protection. They protect against different retry identities. Each looks correct alone; together, the operation is not idempotent.
4. Lab versus real Meridian
The lab uses real Meridian terminology and real refund behaviour, simplified deliberately so the exercise fits in two hours.
Read SCENARIO_GROUNDING.md — it separates grounded Meridian behaviour from
lab simplification from deliberately planted defect. The planted defects are teaching fixtures and
do not imply anything about real Meridian systems.
The one rule worth carrying into the room: inventing lab code is fine; inventing Meridian platform behaviour is not.
5. Vocabulary
Skim TERM_CARD.md. Ten terms. You will use seam, context ledger and
agent brief constantly.
6. What is graded
Read ESSENTIAL_OUTCOMES.md.
The short version: engineering outcomes and evidence, never whether your screen matches the facilitator’s.
7. Model output will vary — plan for it
Your agent will word things differently from the person next to you, find the same problem by a different route, and return findings in a different order. Running the same prompt twice will not give you the same text.
None of that means you are behind. If you find yourself trying to make your output look like the demonstration, stop and ask what the demonstration was showing you instead.
8. What this lab assumes you already have
From Lab 1: fresh-context review, sub-agents, human gates, deterministic checks, journey and hand-off. These are not re-taught. If any of them is hazy, say so early — Stage 0 is the moment for it, not Stage 4.
You do not need to know anything new about payments. You need to be willing to say “the source does not tell us that” and leave it unanswered.
Define / Plan · Bring this into Stage 0
Ready for the session.
Your setup must end with “Setup complete”. If it does not, bring the output to the facilitator before the session. If a Lab 1 concept is hazy, raise it at Stage 0.
Revisit the two examples of how services can each look correct and still disagree.
The thing that makes it interestingKeep the ten terms and the distinction between fact, representation, and seeded failure available.
Revisit the term cardCheck the setup result and bring unresolved questions into the session.
Return to setupIn Stage 0 you will start the lab, establish authority, and seal your first prediction before investigating.
Start Stage 0Available throughout the lab
Workspace & architecture guide.
The repository layout, commands, architecture, and scope notes to consult as you work.
Lab 2 — One Refund Across a Service Boundary
A 120-minute hands-on lab on using AI safely when a money-moving change spans more than one repository, and the decisions that matter live at the seam between them.
Lab 1 taught governing one AI-assisted change inside one repository. This lab is about orchestrating several scoped AI contexts across a service boundary while the engineer keeps authority over the seam.
Participants start with LAB_ACTION_GUIDE.md. This file covers
architecture and setup.
The problem
A merchant refund crosses the TTA → Payment Processor boundary. Two repositories. Both build. Both have passing tests. Neither is obviously broken.
They still do not agree with each other.
flowchart LR
accTitle: The complete service flow
accDescr: WSAPI to TTA to Payment Processor to CPC to Injection to LCS API to DCF. TTA and Payment Processor: represented here. Downstream services: not implemented in this lab.
W("WSAPI") --> T
subgraph lab["represented here"]
T("TTA"):::focus --> P("Payment Processor"):::focus
end
P --> C
subgraph downstream["not implemented in this lab"]
C("CPC") --> I("Injection") --> L("LCS API") --> D("DCF")
end
Correctness across a service boundary lives in the relationship between two implementations, not inside either one — which is why the central fact of the lab is:
Repo A green + Repo B green ≠ pair correct.
Layout
specs/ the specification, plus the write-protected authority documents
docs/ pre-read, term card, outcomes, grounding, decisions, ledger, tracker
starter/ pristine source for the two services; setup copies it out
lab-harness/ pair-verification harness -- readable, not writable
.claude/ lab config, write gate, auditor agent, validators, rubric, grader
scripts/ setup and pair verification
After setup, pgs-tta/ and pgs-payment-processor/ appear at the root as independent Git
repositories. They are gitignored here on purpose: nesting one repository’s history inside
another’s is how solution history leaks.
The seven stages
| # | Stage | Focus | Min |
|---|---|---|---|
| 0 | Ground the Work | Frame the Boundary | 10 |
| 1 | Audit Context | Map the Seam | 18–20 |
| 2 | Author & Validate the Spec | Make the Spec Buildable | 18–20 |
| 3 | Plan Across Repositories | Design the Orchestration | 14–15 |
| 4 | Build & Validate | Build the Bounded Slice | 28–30 |
| 5 | Validate with Fresh Context | Prove the Pair | 20 |
| 6 | Review, Handoff & Close | Transfer the Learning | 5–6 |
Commands
python3 scripts/verify_setup.py # set up and verify the workspace
python3 scripts/run_pair_verification.py # prove the seam
python3 scripts/run_pair_verification.py --explain # what the compatibility matrix asks
python3 .claude/scripts/validate_spec.py # Stage 2 readiness gate
python3 .claude/scripts/validate_plan.py # Stage 3 plan gate
python3 .claude/scripts/build_validator_brief.py # Stage 5 brief, assembled deterministically
python3 .claude/scripts/grade_repo.py # deterministic grading
Architecture notes
The write gate. .claude/hooks/gate_guard.py blocks writes to the harness, the out-of-scope
and non-negotiables documents, the outcomes card and the facilitator package. Reading them is
expected. A lab whose grading harness can be edited by the thing being graded is not measuring
anything. Run --self-test to confirm all four bypass classes are still covered.
The generated client. pgs-tta/src/main/java/.../client/contract/ is generated from the
contract that service is pinned to:
cd pgs-tta && mvn -Pgenerate-client generate-sources
Generation runs at construction and remediation time only, never during an ordinary build, so the lab needs no network. The same contract and configuration reproduce the committed output exactly — regenerate rather than hand-editing.
Determinism. grade_repo.py produces the same score for the same workspace state every time.
Nothing samples, calls a model, or depends on the clock. Two people who reach the same outcome by
different routes score identically.
A limitation, stated rather than papered over. A journey event records that a tool ran, not what it returned, so no rubric check can prove a build went green from the journey alone. Build outcomes are graded from recorded results and repository state, and the facilitator’s live spot-check covers the rest.
Grounding
docs/SCENARIO_GROUNDING.md separates grounded Meridian behaviour from lab simplification from
deliberately seeded defect, and docs/PGS_DECISIONS.md records every decision with its source and
layer.
Seeded defects are teaching fixtures. Their presence does not imply the same defect exists, or ever existed, in a Meridian production system.
The lab proves the represented seam locally. It does not claim the wider Meridian refund capability is production-ready, and pair verification does not replace integration testing or release governance.