Meridian Engineering / Lab 2
One refund across a service boundary.
Two repositories. Both build. Both have passing tests. They still do not agree with each other.
Prepare for the labYour orientation The complete journeyIntuition · Common language · Habit
Where you are in the work
- Define / Plan
- Execute
- Judge
- Learn
What you are developing
Develop a mental model of the seam, then revise it against evidence.
Use shared terms for context, authority, specifications, compatibility, and findings.
Practice bounded delegation, independent verification, and carrying learning into the next task.
The learning design
Understand the course.
The scenario
One change. A shared boundary.
A merchant refund crosses TTA and the Payment Processor. Your task is to coordinate scoped AI contexts, reconcile what they find, and prove that the two services work together.
Engineers with Lab 1 experience governing a change within one repository.
120 minutes · Seven hands-on stages
JDK 17+, Maven 3.9+, Python 3.9+ with PyYAML, Git, and the Workbench plugin.
What success means
Essential Outcomes Card
Write-protected. Keep this open during the session.
What is graded
Engineering outcomes and evidence. Not whether your screen matches the facilitator’s.
Model wording, ordering and formatting will differ between people, and between two runs by the same person. That is normal and expected. We grade what you concluded, what you can show for it, and what you refused to do — never whether your response looks like anyone else’s.
If your agent phrases something differently, finds the same problem by another route, or returns its findings in a different order, you are not behind.
The eight outcomes
By the end of the session your workspace should show:
- The context ledger identifies the contract drift between the two repositories, with evidence.
- The context ledger identifies the divergent business rule, and names which service owns it.
- At least one unknown remains unresolved rather than filled in with a plausible default.
- The specification preserves its out-of-scope list — nothing excluded got built.
- The plan defines repository boundaries and a compatibility rationale, not just a task list.
- The retry identity is stable across the seam — what one side sends is what the other deduplicates against.
- Duplicate semantics survive the seam — a duplicate reaches the caller as a duplicate.
- Pair verification is green, and you can say what that does and does not prove.
Two things that are outcomes, not failures
A “no diff” result, with evidence, is a real outcome. Deciding that a repository is correct and proving it is engineering work. So is refusing to change something because it is out of scope, and recording why.
An unresolved question left visible beats a plausible answer invented to close it. In payments, guessing a threshold or a default is a business decision you were not authorised to make. Surfacing it is the correct move, and it is scored as such.
What this lab does not claim
The lab proves the represented seam locally. It does not claim the wider Meridian refund capability is production-ready, and pair verification does not replace integration testing or release governance.
The complete journey
Seven stages of practice.
- Stage 0 · Define / Plan · 10 min
Ground the work
Frame the boundary before an agent acts.
- Stage 1 · Define / Plan · 18–20 min
Audit context
Map the seam through scoped investigation and human reconciliation.
- Stage 2 · Define / Plan · 18–20 min
Author & validate the spec
Turn the intended change into a buildable specification.
- Stage 3 · Define / Plan · 14–15 min
Plan across repositories
Define bounded work and a defensible compatibility plan.
- Stage 4 · Execute · 28–30 min
Build & validate
Implement the bounded slice while preserving the seam.
- Stage 5 · Judge · 20 min
Validate with fresh context
Prove the pair with deterministic evidence and independent judgment.
- Stage 6 · Learn · 5–6 min
Review, handoff & close
Transfer the evidence and make the learning reusable.
Before the session
Session briefing & conventions.
╔════════════════════════════════════════════════════════════════════════════╗
║ ║
║ L A B 2 ║
║ ║
║ ONE REFUND ACROSS A SERVICE BOUNDARY ║
║ ───────────────────────────────────── ║
║ ║
║ Two repositories. Both green. They still disagree. ║
║ ║
║ 120 minutes · 7 stages · 2 services · 1 seam ║
║ ║
╚════════════════════════════════════════════════════════════════════════════╝
The 90-second version
A merchant took a payment. Now they want it refunded.
That refund crosses a boundary: TTA translates it, Payment Processor decides it. Two services, two repositories, two teams, two release cadences.
Here is the situation you are walking into:
flowchart TB
accTitle: Both green; the two do not agree
accDescr: pgs-tta and pgs-payment-processor each run mvn verify and are GREEN. Both lead to the same unknown at their boundary: the two do not agree.
A["pgs-tta · mvn verify · GREEN"] --> Q["????? · the two do not agree"]
B["pgs-payment-processor · mvn verify · GREEN"] --> Q
Nobody’s tests are failing. Nobody’s build is red. No linter is complaining. And the refund still does the wrong thing, because correctness across a service boundary does not live inside either service — it lives in the relationship between them, and neither repository can see it.
The one line to take away
Repo A green + Repo B green ≠ pair correct.
Lab 1 taught you to govern one AI-assisted change inside one repository. This lab is about what happens when the change spans two, and the decisions that actually matter live in the space between them — where no single agent, and no single test suite, is looking.
Your job, stated plainly
You are not here to fix bugs. You are here to establish what is true across a boundary where neither witness can see the whole picture — and then to direct AI safely inside that picture.
By minute 120 you will have:
▸ audited both repositories with scoped agents that cannot see each other
▸ reconciled their conflicting claims yourself, with evidence
▸ turned a vague specification into one you can build from
▸ written the contracts that bound what each agent may do
▸ let AI implement only inside those bounds
▸ proved the pair with evidence a fresh context produced
The map
The whole lab on one screen. Each stage teaches one agentic-engineering concept, and each hands the next stage something concrete.
flowchart TB
accTitle: Stage, concept, and what you leave with
accDescr: Seven ordered stages, with Q&A after stages 2 and 5. Each stage retains its concept and deliverable.
S0["0 · Ground the Work / context boundary & authority / a sealed prediction"] --> S1["1 · Audit Context / scoped agents, parallel delegation / the context ledger"]
S1 --> S2["2 · Author & Validate the Spec / spec-as-context, readiness gates / a buildable spec"]
S2 --> Q1["Q&A"] --> S3["3 · Plan Across Repositories / orchestration & context isolation / plan + 2 agent briefs"]
S3 --> S4["4 · Build & Validate / bounded execution, deterministic guardrails / a remediated seam"]
S4 --> S5["5 · Validate with Fresh Context / independent judgment / evidence you didn't write"]
S5 --> Q2["Q&A"] --> S6["6 · Review, Handoff & Close / context handoff, learning loop / a practice worth reusing"]
Where the time goes
S0 ██████ 10 min Ground the Work
S1 ███████████ 18 min Audit Context
S2 ███████████ 18 min Author & Validate the Spec
Q&A ██ 3 min ⏸
S3 ████████ 14 min Plan Across Repositories
S4 █████████████████ 28 min Build & Validate
S5 ████████████ 20 min Validate with Fresh Context
Q&A ██ 3 min ⏸
S6 ███ 5 min Review, Handoff & Close
─────────────────────────────────────────
119 min at the low end of every stage
Stages carry ranges (Stage 1 is 18–20, Stage 4 is 28–30). Those ranges are where the facilitator trades time between stages — not where extra time comes from. Something always runs long. When it does, the trade comes out of Stage 4, never Stage 5.
How to read this guide
This guide shows you one worked example per technique and then asks you to write the next one yourself. It deliberately does not hand you prompts to paste. You will not have this document at your desk next month, and a prompt you copied teaches you nothing about how to write the one you will actually need.
The markers
◆ Predict Write your answer down BEFORE you find out. Including — especially —
when you turn out to be wrong. That is the part that sticks.
▶ Your turn The facilitator demonstrated one. You write the next one.
⚠ Trap A place rooms reliably lose time or reach for the wrong instinct.
⏸ Q&A pause Scheduled, so questions land somewhere instead of derailing the room.
⌘ Run A command to actually execute.
Before anything else
Your agent’s output will not match your neighbour’s. Different wording, different ordering, the same problem found by a different route. Running the same prompt twice will not give you the same text either.
None of that means you are behind. If you catch yourself trying to make your output look like the demonstration, stop and ask what the demonstration was actually showing you.
See
docs/ESSENTIAL_OUTCOMES.md— we grade what you concluded and what you can show for it, never whether your screen matches anyone else’s.
Preflight
python3 scripts/verify_setup.py
Must end with “Setup complete”. If it does not, flag it now — not at minute forty.