Skip to content
arula

Meridian Engineering / Lab 2

One refund across a service boundary.

Two repositories. Both build. Both have passing tests. They still do not agree with each other.

Prepare for the lab
Your orientation The complete journeyIntuition · Common language · Habit

Where you are in the work

  1. Define / Plan
  2. Execute
  3. Judge
  4. Learn

What you are developing

Intuition

Develop a mental model of the seam, then revise it against evidence.

Common language / Standardization

Use shared terms for context, authority, specifications, compatibility, and findings.

Habit / Behavior

Practice bounded delegation, independent verification, and carrying learning into the next task.

The learning design

Understand the course.

The scenario

One change. A shared boundary.

A merchant refund crosses TTA and the Payment Processor. Your task is to coordinate scoped AI contexts, reconcile what they find, and prove that the two services work together.

For

Engineers with Lab 1 experience governing a change within one repository.

Format

120 minutes · Seven hands-on stages

Before starting

JDK 17+, Maven 3.9+, Python 3.9+ with PyYAML, Git, and the Workbench plugin.

What success means

Essential Outcomes Card

Write-protected. Keep this open during the session.

What is graded

Engineering outcomes and evidence. Not whether your screen matches the facilitator’s.

Model wording, ordering and formatting will differ between people, and between two runs by the same person. That is normal and expected. We grade what you concluded, what you can show for it, and what you refused to do — never whether your response looks like anyone else’s.

If your agent phrases something differently, finds the same problem by another route, or returns its findings in a different order, you are not behind.

The eight outcomes

By the end of the session your workspace should show:

  1. The context ledger identifies the contract drift between the two repositories, with evidence.
  2. The context ledger identifies the divergent business rule, and names which service owns it.
  3. At least one unknown remains unresolved rather than filled in with a plausible default.
  4. The specification preserves its out-of-scope list — nothing excluded got built.
  5. The plan defines repository boundaries and a compatibility rationale, not just a task list.
  6. The retry identity is stable across the seam — what one side sends is what the other deduplicates against.
  7. Duplicate semantics survive the seam — a duplicate reaches the caller as a duplicate.
  8. Pair verification is green, and you can say what that does and does not prove.

Two things that are outcomes, not failures

A “no diff” result, with evidence, is a real outcome. Deciding that a repository is correct and proving it is engineering work. So is refusing to change something because it is out of scope, and recording why.

An unresolved question left visible beats a plausible answer invented to close it. In payments, guessing a threshold or a default is a business decision you were not authorised to make. Surfacing it is the correct move, and it is scored as such.

What this lab does not claim

The lab proves the represented seam locally. It does not claim the wider Meridian refund capability is production-ready, and pair verification does not replace integration testing or release governance.


Download the unchanged original source

The complete journey

Seven stages of practice.

  1. Stage 0 · Define / Plan · 10 min

    Ground the work

    Frame the boundary before an agent acts.

  2. Stage 1 · Define / Plan · 18–20 min

    Audit context

    Map the seam through scoped investigation and human reconciliation.

  3. Stage 2 · Define / Plan · 18–20 min

    Author & validate the spec

    Turn the intended change into a buildable specification.

  4. Stage 3 · Define / Plan · 14–15 min

    Plan across repositories

    Define bounded work and a defensible compatibility plan.

  5. Stage 4 · Execute · 28–30 min

    Build & validate

    Implement the bounded slice while preserving the seam.

  6. Stage 5 · Judge · 20 min

    Validate with fresh context

    Prove the pair with deterministic evidence and independent judgment.

  7. Stage 6 · Learn · 5–6 min

    Review, handoff & close

    Transfer the evidence and make the learning reusable.

Before the session

Session briefing & conventions.

╔════════════════════════════════════════════════════════════════════════════╗
║                                                                            ║
║   L A B   2                                                                ║
║                                                                            ║
║   ONE REFUND ACROSS A SERVICE BOUNDARY                                     ║
║   ─────────────────────────────────────                                    ║
║                                                                            ║
║   Two repositories. Both green. They still disagree.                       ║
║                                                                            ║
║   120 minutes · 7 stages · 2 services · 1 seam                             ║
║                                                                            ║
╚════════════════════════════════════════════════════════════════════════════╝

The 90-second version

A merchant took a payment. Now they want it refunded.

That refund crosses a boundary: TTA translates it, Payment Processor decides it. Two services, two repositories, two teams, two release cadences.

Here is the situation you are walking into:

flowchart TB
  accTitle: Both green; the two do not agree
  accDescr: pgs-tta and pgs-payment-processor each run mvn verify and are GREEN. Both lead to the same unknown at their boundary: the two do not agree.
  A["pgs-tta · mvn verify · GREEN"] --> Q["????? · the two do not agree"]
  B["pgs-payment-processor · mvn verify · GREEN"] --> Q

Nobody’s tests are failing. Nobody’s build is red. No linter is complaining. And the refund still does the wrong thing, because correctness across a service boundary does not live inside either service — it lives in the relationship between them, and neither repository can see it.

The one line to take away

Repo A green + Repo B green ≠ pair correct.

Lab 1 taught you to govern one AI-assisted change inside one repository. This lab is about what happens when the change spans two, and the decisions that actually matter live in the space between them — where no single agent, and no single test suite, is looking.


Your job, stated plainly

You are not here to fix bugs. You are here to establish what is true across a boundary where neither witness can see the whole picture — and then to direct AI safely inside that picture.

By minute 120 you will have:

  ▸ audited both repositories with scoped agents that cannot see each other
  ▸ reconciled their conflicting claims yourself, with evidence
  ▸ turned a vague specification into one you can build from
  ▸ written the contracts that bound what each agent may do
  ▸ let AI implement only inside those bounds
  ▸ proved the pair with evidence a fresh context produced

The map

The whole lab on one screen. Each stage teaches one agentic-engineering concept, and each hands the next stage something concrete.

flowchart TB
  accTitle: Stage, concept, and what you leave with
  accDescr: Seven ordered stages, with Q&A after stages 2 and 5. Each stage retains its concept and deliverable.
  S0["0 · Ground the Work / context boundary & authority / a sealed prediction"] --> S1["1 · Audit Context / scoped agents, parallel delegation / the context ledger"]
  S1 --> S2["2 · Author & Validate the Spec / spec-as-context, readiness gates / a buildable spec"]
  S2 --> Q1["Q&A"] --> S3["3 · Plan Across Repositories / orchestration & context isolation / plan + 2 agent briefs"]
  S3 --> S4["4 · Build & Validate / bounded execution, deterministic guardrails / a remediated seam"]
  S4 --> S5["5 · Validate with Fresh Context / independent judgment / evidence you didn't write"]
  S5 --> Q2["Q&A"] --> S6["6 · Review, Handoff & Close / context handoff, learning loop / a practice worth reusing"]

Where the time goes

  S0  ██████                             10 min   Ground the Work
  S1  ███████████                        18 min   Audit Context
  S2  ███████████                        18 min   Author & Validate the Spec
  Q&A ██                                  3 min   ⏸
  S3  ████████                           14 min   Plan Across Repositories
  S4  █████████████████                  28 min   Build & Validate
  S5  ████████████                       20 min   Validate with Fresh Context
  Q&A ██                                  3 min   ⏸
  S6  ███                                 5 min   Review, Handoff & Close
      ─────────────────────────────────────────
                                        119 min   at the low end of every stage

Stages carry ranges (Stage 1 is 18–20, Stage 4 is 28–30). Those ranges are where the facilitator trades time between stages — not where extra time comes from. Something always runs long. When it does, the trade comes out of Stage 4, never Stage 5.


How to read this guide

This guide shows you one worked example per technique and then asks you to write the next one yourself. It deliberately does not hand you prompts to paste. You will not have this document at your desk next month, and a prompt you copied teaches you nothing about how to write the one you will actually need.

The markers

  ◆ Predict     Write your answer down BEFORE you find out. Including — especially —
                when you turn out to be wrong. That is the part that sticks.

  ▶ Your turn   The facilitator demonstrated one. You write the next one.

  ⚠ Trap        A place rooms reliably lose time or reach for the wrong instinct.

  ⏸ Q&A pause   Scheduled, so questions land somewhere instead of derailing the room.

  ⌘ Run         A command to actually execute.

Before anything else

Your agent’s output will not match your neighbour’s. Different wording, different ordering, the same problem found by a different route. Running the same prompt twice will not give you the same text either.

None of that means you are behind. If you catch yourself trying to make your output look like the demonstration, stop and ask what the demonstration was actually showing you.

See docs/ESSENTIAL_OUTCOMES.md — we grade what you concluded and what you can show for it, never whether your screen matches anyone else’s.

Preflight

python3 scripts/verify_setup.py

Must end with “Setup complete”. If it does not, flag it now — not at minute forty.