# Lab 3 speaker script

## Opening (0–8 min)

### Timing

| Minutes | What happens |
|---|---|
| 0–2½ | Welcome, the problem, and how the afternoon works |
| 2½–3½ | The business background and the change |
| 3½–7 | Walk through the four files together |
| 7–8 | Where the specs are, then bridge to Diagnose |

### Say

> Hi, I'm [name]. [Coach names] are here to help you through the afternoon.
>
> Quick show of hands. Who's reviewed a pull request in the last month that was mostly
> written by an AI?
>
> Keep your hand up if you read every line.
>
> Yeah. Me neither.
>
> That's why we're here. The agents write code faster than we can read it. If our plan is
> to read harder, we lose.
>
> So today you'll use tools that do a lot of the reading for you. But tools get things
> wrong. They flag things that are fine. They miss things.
>
> And some decisions aren't theirs to make. If nobody has written down how something
> should work, a tool can't decide that for you. Neither can you. You find the person who
> owns that decision, and you ask them.
>
> By the end of today, you'll be able to take a change an AI wrote and work out what could
> be wrong. You'll check it with the tools. Then you'll say what you'd approve, what you'd
> send back, and who has to decide the rest.
>
> Here's how the afternoon works. It's just under three hours, with a break at about the
> seventy-minute mark. You'll sit in pods of four and work in pairs. Swap who's typing
> after every step. If you get stuck, ask your partner, then your pod, then grab a coach.
> At the end, you'll do one round on your own. That's how you and I will both know it
> stuck.
>
> Some background first. This is a card payments service. A merchant reserves money on a
> customer's card, then takes the money later. When goods come back, the merchant refunds
> the customer. Every refund is tied to the payment it came from.
>
> Here's the change. Sometimes a merchant calls support and says, "I refunded my customer,
> and the money never arrived." The agent was asked to give support a way to send that
> refund again. If sending fails, it tries again, up to three times in total.
>
> Why should you care about this one? It's payments code. It moves real money, and it
> handles card numbers. A mistake here costs someone money, or leaks card data.
>
> Let's look at what the agent produced.

### Do

Show the agent's commit. Participants run the same commands on their own machines:

```sh
git show --stat 965029f
git show 965029f
```

The first command lists the four changed files. The second shows the changes line by line.

Don't diff the whole `round-0` branch against `main`. The branch also holds the course
setup (SPEED config, docs, specs and test renames), which buries the agent's change.

### Say

Walk through the files in this order. Scroll to each one as you name it.

> Four files. Let's go through them.
>
> `retry.ts` is new. It's the retry itself. It sends the refund, and if that fails, it
> tries again, up to three times.
>
> `service.ts` is the main payments service. The agent changed two lines here. One is in
> how a failed card check gets logged. The other is in how the fee is worked out.
>
> `cards.ts` is test data. The agent added one card number for testing.
>
> `refund-retry.test.ts` is the new test. Its comment says it confirms that a retried
> refund is accepted.

### Do

Put the two spec paths on screen and leave them there:

- Product spec: `specs/product/payments.md` (owner: payments-product)
- Tech spec: `specs/tech/payments.md` (owner: payments-engineering)

### Say

> You also have what the agent was working from. The product spec says what the payments
> service must do. The tech spec says how. Both are in the `specs` folder. Each one names
> the team that owns it. That's who you'd ask.
>
> So that's the change. Before anyone decides whether it's good, let's see what the tools
> point at.

### Notes

- **Describe, don't judge.** Say what each file does. Don't say whether anything is wrong.
  If someone asks "is that a bug?", answer: "Hold that thought. We'll check it."
- **The two lines in `service.ts` are both findings.** One logs the card number
  unredacted. The other rounds the fee down instead of half-up. Name where they are, not
  what's wrong with them.
- **The test is a finding too.** It never calls the retry helper. Don't point that out.
- **Don't hand out the test spec yet.** `specs/tests/payments.md` names the refund
  idempotency question in its RETRY-01 row. Hold it back until Eval.
- **Say "AI coding agent".** The agent wrote the code. The code itself has no AI in it.

## Diagnose (8–23 min)

### Three teaching goals

- **Shared intuition:** explain why Diagnose selected the added logging line and what
  that match can establish. Build this by following the line through the diff, matcher,
  and output, including why a redacted logging call would also match.
- **Shared language:** describe the result using category, rule, match, and signal;
  distinguish an observation from a claim and a confirmed defect. Introduce each term
  beside its example, and distinguish matching-line, unique-file, and signal counts.
- **Shared behavior:** run Diagnose for the selected task, locate the matching change,
  explain what the rule measured, and carry an unresolved question into Review. Model
  this with the worked report while leaving the judgment open.

### Timing

| Minutes | What happens |
|---|---|
| 8–10 | Why Diagnose helps, and what a task record is |
| 10–12 | Introduce the first F2 rule and run the command |
| 12–15 | Follow the task record into the branch diff |
| 15–19 | Follow the added logging line through the matcher |
| 19–21 | Read the resulting signal and its limits |
| 21–23 | Connect what Diagnose notices, how we describe it, and what we do next |

### Say

> We've seen the refund retry change. We've also seen that the agent changed logging
> and fee calculation in the existing service. That gives us several things to check,
> even though the request was about refunds.
>
> Diagnose helps us choose where to start looking. It checks the change against rules
> this project has supplied and records what matched. A match gives us a concrete place
> to investigate. It doesn't tell us whether the behaviour is wrong.
>
> We'll follow one rule all the way through: the first rule in the cardholder data
> leakage category. It looks for added lines that mention certain logging or reporting
> names. We'll use the logging change we just saw to understand exactly how it works.
>
> We're building three things together. Shared intuition: why Diagnose selected that
> line and what it tells us. Shared language: how we describe the result. And shared
> behavior: what we check and carry into Review. We'll build all three from this one
> example.
>
> First, how does SPEED know which change we're talking about? It keeps a task record:
> a JSON file with information about a piece of work. That includes its branch, the
> recorded agent model, and a list of files. The feature groups those records together.
> Here, the feature is payments, and we're using Task 1.

### Do

Keep the payments fixture open. Open SPEED's implementation in a separate repository
window. All paths below are relative to the root of the named repository: either the
payments fixture or SPEED.

Show fixture file `.speed/features/payments/tasks/1.json`, **lines 2–11**. Point to
`id`, `branch`, `agent_model`, and `files_touched`. Stay on these top-level fields.

### Say

> This record points to the round-0 branch and lists the four files from the opening.
> That's how we identify the work. The task number isn't a commit number, and that file
> list isn't a filter on everything Diagnose will scan. We'll see the actual input in
> a moment.
>
> Now we need the project's rule. SPEED supplies the machinery for matching text.
> The project supplies what to look for and how to describe it. That matters here:
> a payments project needs to pay attention to card data reaching logs.

### Do

Show fixture file `speed.toml`, **lines 8–10**. Point to `classes_file` at **line 9**,
then open the named fixture file `.speed/classes.yaml`, **lines 17–22**:

```yaml
  - id: F2
    title: Cardholder data leakage
    rules:
      - look: added-lines
        match: '\b(log|logger|webhook|telemetry)\s*\.'
        say: "{n} file(s) with an observability sink call in the added lines"
```

Keep the second F2 rule, **lines 23–25**, folded. It remains enabled; it is outside
this walkthrough. F2 names the category, not the individual rule.

### Say

> F2 is the category. This first rule has three parts. “Look” chooses added lines as
> the input. “Match” is the text pattern. “Say” is the message to write if it finds
> something, with the count inserted where the braces are.
>
> The pattern looks for log, logger, webhook, or telemetry, starting at a word boundary,
> followed by optional whitespace and a dot. Our changed line contains log dot error,
> so it fits that pattern.
>
> Notice how little that pattern asks. It doesn't ask whether the argument contains a
> card number. It doesn't even check that this is a function call. The category tells
> us why the project cares about a match; the pattern tells us what it can actually
> notice.
>
> Let's run it. The feature selects the payments records. The task selects the record
> we've just opened. The command writes a risk surface: a list of the project's failure
> categories, the signals it found, and fields for judgments that are still open.
> A signal is the recorded observation from a rule that matched.

### Do

From the fixture root, run:

```sh
speed diagnose --feature payments --task 1
```

Show the console's `Change: main...round-0` and the F2 summary, `2 signal(s)`.
Explain its meaning immediately; do not ask learners to interpret it first.

### Say

> F2 has two signals because both of its rules found something. That isn't two defects,
> and it isn't the count for our selected rule. We're going to follow the first signal.
>
> To understand where it came from, we'll follow four steps: read the task record,
> make the branch diff, match the added line, and record the signal.
>
> First, the shell command reads the task record we opened earlier. It takes the branch
> name and uses it to select the change. It also reads the recorded model and declared
> file list. It deliberately leaves out the description, acceptance criteria, and
> review feedback. The comment explains why: those fields would supply a verdict that
> Diagnose must not make.

### Do

In the SPEED repository, show `lib/cmd/diagnose.sh`, **lines 32–54**:

- **Lines 34–35:** resolve the configured rules file relative to the fixture root.
- **Lines 44–54:** read the task and extract `branch`, `agent_model`, and `files_touched`.
  The comment explains why description, acceptance criteria, and review feedback are
  excluded: this command must not supply a verdict from those fields.

Show `lib/tasks.sh`, **lines 66–76**. `task_get` reads the selected task's JSON file.

### Say

> Now we have the branch name from the record. The next step uses that name and the base
> branch to make a Git diff. The result is a temporary text file that the engine can read.

### Do

Return to `lib/cmd/diagnose.sh`, **lines 61–81**, focusing on **line 74**, which writes
the branch diff into a temporary file.

Show `lib/git.sh`, **lines 326–330**. The helper runs `git diff` with the base and
task branch separated by three dots. In this checked run, the base is `main`.

Return to the fixture terminal and show the service's part of that same diff:

```sh
git diff main...round-0 -- src/payments/service.ts
```

Locate the added logging call in fixture file `src/payments/service.ts`, **line 67**:

```ts
      log.error(serialiseFailure(error, req));
```

### Say

> There's our line, with a plus sign in the diff. The previous logging line is removed;
> this one is added. Even though most of the call stayed the same, Git presents the
> replacement as a whole added line. That's the text this rule receives.
>
> There's also a difference from the opening. We opened the agent's four-file commit
> there. Diagnose compares the branch with its common ancestor with main. That includes
> the course setup as well as the agent's change. It doesn't narrow the diff to the four
> files listed in the task record. Our selected rule still matches just this one line
> in the current branch diff.
>
> That gives us the input for the third step: matching the added line. The shell command
> passes the diff and the project's rules to the Python engine.

### Do

In SPEED's `lib/cmd/diagnose.sh`, show **lines 93–100**. The command invokes
`lib/diagnose_engine.py` with the rules path, temporary diff file, declared file list,
and, when configured, a spec path. The selected added-lines rule does not use that
file list or spec text to match.

Open SPEED's `lib/diagnose_engine.py`, **lines 505–513**. `diagnose` loads the classes,
extracts added lines, and builds the inputs used by the rules. Then show
**lines 307–320**, `_added_lines_by_file`.

### Say

> The engine turns the diff into pairs: a file path and the text of an added line.
> It remembers the path from the file header. It skips that header itself, then removes
> the leading plus sign from each added line.
>
> So our pair contains the service's path and the logging call. The removed call isn't
> in this list. Neither are the unchanged lines around it. At this point, we're handling
> strings from a diff. We haven't parsed TypeScript or followed any values through the
> payments service.

### Do

In SPEED's `lib/diagnose_engine.py`, show **lines 516–520**, where each rule is run.
Follow `_run_rule` to **lines 482–486**: `added-lines` selects `_look_added_lines`.
Show that function at **lines 360–372**.

### Say

> The rule's “look” value chooses this matcher. It compiles the configured pattern,
> searches each added line, and adds one to the count when a line matches.
>
> On our logging line, the matching text is log followed by a dot. That makes the count
> one. The matcher also records the service's path so we know where to look.
>
> These are two different measurements. The count is matching lines. The location list
> contains each matching file once. Two matching lines in the same file would give a
> count of two and one file in that list. Several matches on one line would still count
> as one matching line.
>
> That exposes a wording problem in our project's rule. Its message says “files,” but
> this matcher counts lines. In this example, one line in one file makes the numbers
> agree. We still need to understand what the implementation measured.
>
> We now have a count of one and the service's file path. The fourth step turns those
> into a recorded signal, using the message supplied by the project rule.

### Do

Show SPEED's `lib/diagnose_engine.py`, **lines 498–502** and **516–526**. A positive
count produces one signal; `_format_say` substitutes the count into the configured
message, and `where` receives the file list.

Return to SPEED's `lib/cmd/diagnose.sh`:

- **Lines 109–116:** the console counts signal entries in each category.
- **Lines 135–139:** write the risk-surface YAML.
- **Lines 158–184:** render category fields, set judgments to `undecided`, and copy each
  signal's message and locations into the YAML.

Open fixture output `.speed/features/payments/risk-surface.yaml`, **lines 16–24**.
These locations were verified after generating Task 1's output on 2026-09-29:

```yaml
  - id: F2
    failureMode: "Cardholder data leakage"
    plausible: undecided
    costOfMissing: undecided
    findableByReading: undecided
    signals:
      - observed: "1 file(s) with an observability sink call in the added lines"
        where:
          - "src/payments/service.ts"
```

### Say

> Here's the result of that whole path. The matcher counted one added line. The engine
> inserted that count into the project's message and attached the file path. The shell
> command wrote them here, under F2.
>
> The three undecided fields are separate judgments: whether this failure is plausible,
> what missing it would cost, and whether reading could find it. The writer puts
> undecided in all three. A match doesn't fill them in.
>
> Diagnose hasn't established what reaches the log when the card check fails. It hasn't
> inspected what serialiseFailure returns, traced the request, or exercised the failure
> path. It would match a properly redacted
> logging call too. And it could miss logging through a name the pattern doesn't cover.
> The result gives us a place to investigate, rather than a decision to approve or
> reject the change.

> We've followed the logging line from the branch diff to this output. The engine kept
> it because it was added, matched log dot against the project's pattern, and recorded
> the count and file path. That's how we can explain the result. Because we also saw
> what the matcher reads, we know why this result can't tell us what data was logged.
>
> Each part we've seen has a name. F2 is the category. Its first rule supplies the
> pattern. Log dot is the matching text. The signal is the observation written here.
> If we say “this change leaks card data,” we've moved from that observation to a claim.
> We'd need evidence about what actually reaches the log before calling it a confirmed
> defect. Using those words consistently lets someone else understand what we've checked.
>
> Here's how I'd report this result: “The first F2 rule matched the added log.error
> line in the payments service. It found the logging name; it didn't inspect the data.
> We still need to establish what reaches that log when the card check fails.”
>
> That gives the next person the location, the observation, and a specific question
> to investigate. That's how we'll use Diagnose: run it for the task we're checking,
> locate the matching change, explain what the rule noticed, and carry the unresolved
> question forward. Leave the judgment open until we have evidence to answer it.
>
> That's the question we'll carry into Review: what reaches the log when the card check
> fails? We'll keep the matching line and that question beside us as we examine the change.
>
> We've built shared intuition: we can explain why this line matched and what the match
> tells us.
>
> We've built shared language: the rule produces a signal. We need evidence about what
> reaches the log before we can call the suspected leak a confirmed defect.
>
> And we've shown the shared behavior: locate the change, explain the observation, and
> carry the unanswered question into the next check.

### Notes

- **Scope:** a presenter-led worked example of the first F2 rule only. No new learner
  exercises, quizzes, discussions, or verdict requests are introduced.
- **Preserve the opening's starting point:** learners have seen the business request,
  four changed files, and two spec paths. They have not yet learned SPEED's task records,
  rules, signals, or command internals; introduce those here before using them.
- **Use the checked SPEED implementation.** The walkthrough
  follows selected shell and Python blocks, not the other matchers, YAML parser,
  or evidence-storage implementation.
- **Task provenance:** use `.speed/features/payments/tasks/1.json`, lines 2–11, and the
  generated risk surface's task marker at line 2. The record's title at line 3 is a
  regression-fixture label; its nested `workbench_source_task` is retained metadata
  from other work. Neither supplies the payments story or this rule's input.
- **Output locations can move:** regenerate Task 1's output and check its line numbers
  before teaching from a changed fixture. The F2 excerpt above omits the second signal;
  the console's two signals are expected because both F2 rules remain enabled.
- **Count accurately:** this selected rule produces one signal summarizing one matching
  added line in one file. Do not call the category's signal count a defect count.
- **Do not change the fixture to improve the demo:** its misleading `file(s)` label and
  broad branch diff are part of the checked behaviour. Explain them; leave them intact.
- **Do not follow the console's verdict instruction here:** the command prints
  “Every verdict is undecided. Fill them in before running any check.” This lesson does
  not ask for a verdict during Diagnose. Leave those judgments open while explaining
  what matched and what still needs investigation.
- **Keep the open question open:** do not inspect logging helper internals or demonstrate
  a leak during this section. Those would go beyond the selected Diagnose rule.
- **Verification:** source locations and Task 1 output were checked on 2026-09-29.
  The Diagnose command was run with Homebrew Bash against this workspace. The script
  has been reviewed as text; it has not been rehearsed aloud.

## Review (23–45 min)

### Three teaching goals

- **Shared intuition:** explain how Review reasons about a change, and why its input
  limits what its findings can establish. Build the review message from the task record
  and diff, connect it to the reviewer instructions, then check the F2 claim against
  the current implementation.
- **Shared language:** distinguish a signal, a review finding, a requirement, and evidence.
  Use the logging example to show why a finding still needs checking.
- **Shared behavior:** run task Review, inspect a finding's support, and carry a precise
  question into Eval. Keep Task 1's evidence separate from other task artifacts.

### Say

> Diagnose noticed the logging name. Our question is still what reaches the log when
> the card check fails. Review gives us another way to examine the change: an agent
> reads the diff and reports problems it can support from that input.
>
> We're building shared intuition about how that reading works, shared language for
> its findings, and shared behavior for checking them. We'll start with the logging
> line, then examine the fee change from the same fixture.

### Do

From the payments fixture root, run:

```sh
speed review --feature payments --task 1
```

In the SPEED repository, show `lib/cmd/review.sh`, **lines 194–212**. Point to the
`--task` route into `_cmd_review_task`. Then show **lines 95–104** and **119–125**.
Keep fixture `.speed/features/payments/tasks/1.json`, **lines 2–11**, beside them.

### Say

> SPEED gets the branch name from Task 1 and builds the diff we saw in Diagnose.
> It also reads the recorded author model and declared files. Here, the record says
> sonnet and lists the four files from the opening. Those values and the diff become
> the message the reviewer receives. Let's see how it's put together.

### Do

Show SPEED `lib/cmd/review.sh`, **lines 129–140**. Point to the text between
`cat <<EOF` and `EOF`: this is the message template. Follow each substitution back
to its source:

- `${author_model}` comes from the task record at **line 98**.
- `${declared_files}` comes from the file-list object built at **lines 99–104**.
  In this fixture, `files_touched` supplies the `modified` list; `created` and
  `deleted` are empty. These are recorded declarations, not Git's classification.
- `${main_branch}`, `${branch}`, and `${diff}` identify and contain the comparison
  made at **lines 119–125**.

Show this shortened view of the assembled message. The logging hunk is an excerpt;
the code inserts the full branch diff at `${diff}`.

````text
Recorded author model: sonnet
Declared file lists:
{"created":[],"modified":["src/payments/retry.ts","src/payments/service.ts","test/refund-retry.test.ts","test/fixtures/cards.ts"],"deleted":[]}

Diff (main...round-0):
```diff
-      log.error(serialiseFailure(error, redact(req)));
+      log.error(serialiseFailure(error, req));
```
````

### Say

> This block builds a string. The shell fills in the model name, file lists, and Git
> diff. There isn't an agent selecting or summarizing the change at this step.
>
> In that message, the reviewer can see the old and new logging calls together. It
> can see that redact(req) became req. It also receives the fee change and the rest of
> the branch diff. The declared files don't trim that diff to four files.
>
> This is the change to read. The instructions for how to review it live in a second
> place: the reviewer role file.

### Do

Show SPEED `agents/clean-context-reviewer.md`, **lines 3–12** and **70–71**.
Then return to `lib/cmd/review.sh`, **lines 145–150**. Point to the role-file path
and `$review_prompt` as the two inputs passed to `provider_run`.

### Say

> The role tells the reviewer to report actionable problems supported by the diff and
> to say when it can't establish something. It treats the diff and metadata as material
> to examine; comments in the changed code aren't instructions for the reviewer to obey.
>
> This call starts the reviewer through SPEED's configured agent provider. It passes
> both pieces: these instructions and the message we just built. It also selects the
> reviewer model and read-only tool permissions. The sonnet label inside the message
> records who wrote the change; it doesn't select the reviewer model.
>
> The role also tells the reviewer not to run commands. It can identify a concern about
> the logging call, but a suggested reproduction isn't a test it has executed.

### Do

Show SPEED `lib/cmd/review.sh`, **lines 83–85**, and
`agents/clean-context-reviewer.md`, **lines 3–12**. Then show fixture
`.speed/features/payments/tasks/1.json`, **lines 50–54**. Return to the assembled
message at SPEED `lib/cmd/review.sh`, **lines 129–140**, to show that the criterion
isn't inserted. No separate product spec or Diagnose risk surface is inserted either.

In the branch diff, locate the calculation from fixture
`src/payments/service.ts`, **lines 118–120**, and the assertions from
`test/service.test.ts`, **lines 107–114**. Keep them beside the reviewer instruction
to report problems supported by the diff.

### Say

> The refund criterion checks that two refund records exist. But the agent also
> changed the fee calculation and logging. We want those changes examined too.
>
> This reviewer gets the whole branch diff and instructions to find problems in it.
> The acceptance criterion isn't part of its prompt. Look at the fee calculation
> beside the added test to see a problem it can identify from this input.
>
> For a capture of two hundred, the new calculation produces a fee of two. The test
> expects three. Both are visible in the diff. The reviewer can point out that they
> disagree and tell us where to check. It hasn't run the test.
>
> That's how we'll use a finding: read the code it points to, check the expected
> behavior, and run the proposed reproduction. Let's start with its logging claim
> and follow the request into the serializer.

### Do

Show SPEED `agents/clean-context-reviewer.md`, **lines 23–31** and **37–64**.
Then show `lib/cmd/review.sh`, **lines 160–184**.

### Say

> The output is an issues array. Each item needs a message and severity. A location,
> observed behavior, expected behavior, and reproduction are included when the input
> supports them. These are review findings. They aren't approval decisions.
>
> A scenario identifier names a case in the test catalog. The reviewer may copy one
> only when it appears in the diff and belongs to the behavior being discussed. It
> mustn't invent an identifier to make a finding look testable.
>
> The command archives the evidence and writes a normalized task review. That keeps
> the finding available for later steps without turning its claims into established facts.

### Do

Open fixture `.speed/features/payments/reviews/task-1.review`. Its checked saved
version has the logging claim at **lines 4–11** and the fee claim at **lines 13–20**.
Treat these as recorded claims; a fresh agent run may change both wording and locations.
Recheck line numbers after running Review.

For the logging claim, show fixture:

- `src/payments/service.ts`, **lines 62–68**: the request reaches `serialiseFailure`.
- `src/obs/index.ts`, **lines 49–64**: the serializer calls `redact` on its context.
- `src/obs/index.ts`, **lines 27–31**: the logger stores the resulting string.

### Say

> The saved finding says the raw request exposes card data. Let's follow the actual
> call. The service passes the request into the serializer. But the current serializer
> redacts its context before making the string. The logger receives that string.
>
> That changes our conclusion. Removing redaction at the call site is visible in the
> diff, but it doesn't establish a leak when the helper still redacts. Diagnose's match
> was accurate. This saved Review claim isn't supported by the current implementation.
> We need to check the code behind the claim, including changes outside the displayed diff.

### Do

Show fixture `src/vault.ts`, **lines 21–26**, to establish the controlled failure.
Then run this in the fixture root. It imports `src/payments/service.ts`,
`src/obs/index.ts`, and `test/fixtures/cards.ts`; the test card is defined in
`test/fixtures/cards.ts`, **lines 8–16**.

```sh
node --experimental-strip-types --input-type=module <<'JS'
import { PaymentsService } from './src/payments/service.ts';
import { sinks, resetSinks } from './src/obs/index.ts';
import { TEST_CARDS, EXPIRY } from './test/fixtures/cards.ts';
resetSinks();
try {
  new PaymentsService().authorise({
    pan: Number(TEST_CARDS.visa), expiry: EXPIRY, amount: 1000,
  });
} catch {}
console.log({
  records: sinks.logs.length,
  redacted: sinks.logs.every(r => JSON.parse(r.message).context.pan === '[redacted]'),
  containsTestPan: sinks.logs.some(r => r.message.includes(TEST_CARDS.visa)),
});
JS
```

Checked result: one log record, `redacted: true`, `containsTestPan: false`.

### Say

> For this demonstration, we deliberately supplied a number where the TypeScript
> contract requires a string. Calling replace on that number makes tokenisation throw,
> so we can inspect the failure path. The stored record has a redacted pan field.
>
> This is evidence for this controlled failure. It doesn't prove every logging path
> is safe. It does answer the claim we just checked. Bad expiry isn't a reproduction
> here: this tokenise function doesn't validate expiry.

### Do

Turn to the fee finding. Show fixture `src/payments/service.ts`, **lines 118–120**,
alongside `specs/tech/payments.md`, **lines 40–43**, and
`test/service.test.ts`, **lines 107–114**.

### Say

> Review also points to the fee change from the opening. The requirement says to round
> half up. The service floors the result. The existing test expects three minor units
> for a capture of two hundred. That gives us a concrete claim and a check we can run.
>
> Our shared intuition is that Review reasons from supplied code; its claims still
> need checking against the current implementation.
>
> Our shared language separates a review finding from the requirement and evidence
> that support or contradict it.
>
> Our shared behavior is to follow the claim, check its support, and carry the specific
> question into Eval: what did the tests actually examine?

### Notes

- Use `--task 1`. `--task-id 1` selects a different Review path in this build.
- Source references checked against SPEED revision `6a3637a`.
- The saved task review was inspected, not regenerated for this edit. Its output is
  historical evidence, not a promised result from the next agent invocation.
- The controlled logging reproduction was run and checked. Preserve its limited claim.
- Presenter-led demonstration only; no new learner interaction is introduced.

## Eval (45–70 min)

### Three teaching goals

- **Shared intuition:** explain how a scenario becomes a selected test and why acceptance
  applies to that scope. Trace Task 1's RETRY-01 result through selection and execution.
- **Shared language:** distinguish a scenario, a selector, an assertion, execution
  evidence, and acceptance. Separate a passing task result from an unexamined finding.
- **Shared behavior:** inspect selection before interpreting a result, then run the
  specific check needed for the remaining fee claim.

### Say

> We have a fee claim to check. Before reading a green or red result, we need to know
> what Eval selected. A task can pass its selected tests while another finding remains
> unanswered.
>
> Eval joins scenarios in a test catalog to runnable tests, executes them, and records
> evidence. We'll build intuition about that join, language for the result, and the
> habit of checking its scope before using it to make a decision.

### Do

Show fixture `.speed/features/payments/tasks/1.json`, **lines 49–54**;
`specs/tests/payments.md`, **lines 88–92**; and
`test/refund-retry.test.ts`, **lines 9–15**.
Then show fixture `speed.toml`, **lines 1–6**.

### Say

> RETRY-01 is the scenario identifier in this task's test criterion. The catalog says
> what that scenario means: the existing test records two direct refunds. The test's
> title carries the same identifier.
>
> Those identifiers connect the task, catalog, and test. The configuration supplies
> the test command and tells Eval which declared files are candidate test files.
> A selector identifies the file and individual test to execute.
>
> This scenario characterizes the current behavior. It doesn't establish a refund
> idempotency policy, and its passing assertion won't prove the retry helper works.

### Do

From the fixture root, run:

```sh
speed eval --feature payments --task 1
```

In SPEED, show `lib/cmd/eval.sh`, **lines 247–259** and **274–295**.
Then show `lib/eval_runtime.py`, **lines 202–236**.

### Say

> The shell locates the test spec and asks Python to prepare a run. Python reads the
> tasks, requires the selected task to be done, reads the catalog, and chooses a plan.
>
> An explicit test plan takes priority. Otherwise, task Eval tries the persisted task
> review. If that supplies no executable cases, it can select from the task criteria.
> Review is an additional input, not a prerequisite.
>
> That's an actual Review-to-Eval connection in this build. It depends on declared
> scenario identifiers and test ownership. It doesn't turn every review sentence
> into a test automatically.

### Do

Show SPEED `lib/eval_review_adapter.py`, **lines 34–50**, **112–147**, and
**230–249**. Show `lib/eval_runtime.py`, **lines 171–187**.
For the criteria fallback, show `lib/eval_selection.py`, **lines 28–39** and **57–100**.

### Say

> The adapter starts with findings that name a scenario and aren't explicitly marked
> untestable. It looks for tagged tests in the task's candidate files, then checks that
> the task owns them. The criteria path uses the task's scenario identifiers to find
> tagged tests in those candidate files too.
>
> Our saved Review names RISK-02, but Task 1's declared test file is the refund retry
> test. That file doesn't contain RISK-02. The adapter can select RETRY-01; it cannot
> borrow a different task's test and present it as Task 1 evidence.

### Do

Show SPEED `lib/eval_runtime.py`, **lines 284–297** and **333–363**;
`lib/eval_execution.py`, **lines 350–383** and **388–405**;
`lib/eval_report.py`, **lines 511–526** and **776–783**; and
`lib/eval_runtime.py`, **lines 422–462**.

### Say

> Preparation saves the inputs and selected plan under a new attempt directory. The
> executor reads that plan and runs each selected case. For a named Node test, it makes
> an exact name pattern and records machine-readable results and command output.
>
> The report combines the evidence with the required criteria and gates. Acceptance
> requires at least one applicable result and all applicable results to pass. A missing
> required check isn't a pass.
>
> Finally, the runtime publishes the latest report and summary while retaining the
> attempt. That's how we can inspect both the outcome and what produced it.

### Do

Open fixture `.speed/features/payments/eval/task-1/summary.md`, **lines 3–13**,
**19–25**, and **33–39**, and `.speed/features/payments/eval/task-1/report.json`,
**lines 74–95**. These locations were checked for attempt
`7f8356983ce54c539dee2c77a33694f3`; recheck after a new run.

The checked attempt's plan is
`.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/test-plan.json`,
**lines 2–10** and **25–27**. Its command evidence is
`commands/8852ef476b5b4c909893d3e4ea51d6dd/result.json` relative to that attempt.

### Say

> One scenario passed: RETRY-01. One task criterion reused that evidence; it wasn't
> another test execution. The review notes explain why the logging and fee scenarios
> weren't selected for Task 1.
>
> Read accepted with its scope attached: accepted for the declared scenarios, criteria,
> and recorded gates. It doesn't answer the fee claim. And the retry test calls refund
> directly twice, so it doesn't exercise the helper's retry loop either.

### Do

Run the fee check separately in the same fixture:

```sh
node --experimental-strip-types --test --test-name-pattern '^\[RISK-02\]' test/service.test.ts
```

Show fixture `test/service.test.ts`, **lines 102–114**, and
`src/payments/service.ts`, **lines 118–136**. The checked test fails at
`test/service.test.ts`, **line 112**, with actual `2`, expected `3`.

### Say

> This test captures two hundred minor units and checks the ledger's fee account. It
> expects three; the implementation posts two. We now have a reproduced fee failure.
> This separate command isn't part of Task 1's Eval acceptance result.
>
> Our shared intuition connects selection to evidence: Eval can establish only what
> its selected assertions examine.
>
> Our shared language distinguishes a passing scenario from acceptance of the whole
> change, and a missing check from a failing one.
>
> Our shared behavior is to inspect scope, keep the remaining questions visible, and
> run the check that can answer them. Now we'll describe this fee defect for someone
> who has to decide what to do about it.

### Notes

- Task 1 Eval and the separate RISK-02 command were run for this edit. The first passed;
  the second failed. The report identifies an uncommitted working tree, not a clean commit.
- Source references checked against SPEED revision `6a3637a`. Older course notes saying
  Eval never consumes Review are stale for this implementation.
- Do not run Task 2 or widen this lesson to feature Eval to make the fee failure appear
  in Task 1's result. Explain the boundary and preserve the separate evidence.
- Do not edit criteria, the catalog, or test mappings to improve the demonstration.
- Pause for the break at 70 minutes; resume at 78 minutes.

## Think through one defect (78–88 min)

### Three teaching goals

- **Shared intuition:** connect the fee failure to the fee/merchant split and the
  existing requirement, without inventing an incident or its scale.
- **Shared language:** separate observed behavior, expected behavior, impact, evidence,
  and a decision request.
- **Shared behavior:** write a reproducible report that helps a product manager decide
  priority and ownership, while keeping the repair tied to the agreed requirement.

### Say

> Let's stay with the fee defect. A product manager shouldn't have to reverse-engineer
> Math.floor to understand why this matters. Our report needs to explain the behavior,
> show how we checked it, and say which decisions remain.
>
> For the reproduced capture, the fee is two minor units and the merchant net is one
> hundred and ninety-eight. The requirement calls for a fee of three and net of one
> hundred and ninety-seven. The total still balances. The split is wrong.
>
> That means a balanced ledger alone wouldn't answer this question. We need evidence
> about which accounts receive the amounts.

### Do

Show fixture `specs/tech/payments.md`, **lines 42–43**;
`specs/tests/payments.md`, **line 54**;
`src/payments/service.ts`, **lines 118–136**; and
`test/service.test.ts`, **lines 102–114**.

Use this worked report aloud:

> “Captures can allocate one minor unit too little to the scheme fee and one too much
> to the merchant. In our reproduced case, a capture of 200 posts fee 2/net 198, where
> TR4 requires fee 3/net 197. The existing RISK-02 test fails with 2 versus 3.
> We need to restore the required rounding and pass that check. Payments-product needs
> to decide priority and whether deployed exposure needs investigation. This fixture
> doesn't establish any production exposure.”

### Say

> The report starts with the effect, then gives a small example and its evidence. It
> doesn't claim all captures are wrong. It doesn't claim customers have lost money.
>
> The rounding rule is already decided in the spec. We don't need a product manager
> to choose between floor and half-up again. We do need a decision about priority and
> whether to investigate exposure outside this fixture.
>
> Refund idempotency is a separate unanswered policy question. Keeping it out of this
> repair gives us a defect we can act on without inventing a refund policy.
>
> Our shared intuition connects the code failure to a concrete business effect.
>
> Our shared language separates the observed split, required split, evidence, and
> decisions that remain.
>
> Our shared behavior is to make the report reproducible and useful to its decision
> owner. Let's put that report into Define.

## Define (88–108 min)

### Three teaching goals

- **Shared intuition:** explain how Define assembles evidence into a finding and how
  an editable draft becomes a filed defect. Follow the CLI, inbox, preview, and filing path.
- **Shared language:** distinguish source evidence, a finding, a draft, a disposition,
  and a filed defect. Explain each using the fee example and the UI fields.
- **Shared behavior:** inspect the source, fill unsupported draft fields with checked
  facts, explain impact and decisions, and check existing defects before filing.

### Say

> We have a checked fee failure and a report that explains its effect. Define gives us
> a place to organize the evidence and record what we intend to do with it.
>
> The command shows the feature's finding inbox. The UI lets us inspect evidence,
> choose what to do with a finding, and review a defect draft before filing it.
> A draft is editable text. A filed defect is a saved report with tracking state.

### Do

From the fixture root, run:

```sh
speed define feature payments
```

Open `http://localhost:3000/define/payments/findings`. Show the fee finding's source
and **Create defect from finding** form, as in the supplied screenshots. Keep the
source producer and task visible: use `clean_review`, Task 1.

In SPEED, show `lib/cmd/define.sh`, **lines 11–29**, and
`lib/define_cli.py`, **lines 44–59** and **79–90**.

### Say

> The shell passes the feature name to Python. Python reads the findings and prints
> their status and sources. In an interactive terminal, the shell may open this route
> if the dashboard frontend is already running. It doesn't start the dashboard itself.
>
> The inbox combines source evidence with our recorded decisions. A finding is the
> item we are deciding what to do with. Its evidence keeps the producer, task, and
> source artifact attached, so a summary doesn't lose where it came from.

### Do

Show SPEED `dashboard/frontend/app/define/[feature]/findings/page.tsx`,
**lines 569–591**; `dashboard/frontend/lib/graphql/queries/feature-defects.ts`,
**lines 3–13**; `dashboard/backend/schema.py`, **lines 781–799**; and
`dashboard/backend/resolvers/feature_defects.py`, **lines 42–43** and **59–60**.
Then show `lib/defect_findings.py`, **lines 467–491** and **549–569**.

### Say

> The page requests the feature's findings through GraphQL. That's the request layer
> connecting the browser to the backend. The backend uses the same read-findings
> function as the CLI.
>
> That function gathers the evidence and decision history, applies recorded groups,
> and builds each finding with its evidence, resolution, and permitted actions. The
> UI shows that view; it isn't independently deciding which claims are defects.

### Do

Show SPEED `dashboard/frontend/app/define/[feature]/findings/page.tsx`,
**lines 379–405**, and `lib/defect_intake.py`, **lines 125–161**.
Return to the UI's draft. Point to the source, blank fields, and **NO WRITE** indicator.

### Say

> Opening the draft makes a query, not a filing request. The preview copies explicit
> observed, expected, and reproduction fields from the evidence. It doesn't construct
> missing facts from the finding's message.
>
> That's why this draft has gaps. The title can be blank when the source title is too
> long. Severity starts unset. Reproducibility, last known working, environment, and
> error output are left for us to establish. Missing fields are information we still
> need to supply, not evidence that the form is broken.

### Do

Show the form fields in SPEED
`dashboard/frontend/app/define/[feature]/findings/page.tsx`, **lines 488–515**
and **521–538**. Fill a demonstration draft using the checked fee example:

| Field | Worked example |
|---|---|
| Title | Capture fee truncates instead of rounding half up |
| Severity | Review P1 as a proposed priority: the wrong fee/net split affects money. Explain the priority choice; don't claim a production incident. |
| Related features | payments |
| Reproducibility | always for the reproduced capture of 200 |
| Last known working | Parent of 965029f uses applyRate; runtime behavior at that revision has not been checked in this walkthrough. |
| Observed | Capture 200 posts fee 2/net 198. The existing RISK-02 assertion expects fee 3 and fails. |
| Expected | TR4 requires half-up rounding: 200 gives fee 3/net 197; 9999 gives fee 149/net 9850. Fee plus net must equal the capture. |
| Reproduction | Numbered steps: run the exact RISK-02 command from Eval; observe 2 versus 3; inspect service.ts line 119. |
| Environment | Current round-0 payments fixture, uncommitted working tree, macOS, Node v26.9.0, node:test, published fixture test card. Refresh these facts on the teaching machine. |
| Error output | Paste the captured AssertionError with actual 2 and expected 3. |
| Additional context | Link the selected Task 1 evidence and TR4/RISK-02. Explain the fee/merchant split, unknown production exposure, and requested product priority/exposure decisions. |
| Filing rationale | Track the reproduced rounding regression separately from refund retry policy and coverage. Restore the existing requirement using the existing helper and test. |

The reproduction field should contain actual numbered lines, not a paragraph or
literal `\n` characters. Confirm severity only after explaining its rationale.

For the last-known-working field, inspect the historical fixture file with
`git show 965029f^:src/payments/service.ts`, **lines 116–119**. Its parent resolves
to `d08a680` in this checkout; **line 118** uses `applyRate`. This checks the previous
code, not a successful runtime test at that revision.

### Say

> Observed says what we reproduced. Expected says what the requirement calls for.
> Reproduction lets the next person check our work. Environment and error output say
> where we ran it and what happened.
>
> Additional context explains why the mismatch matters and what decisions remain.
> Filing rationale explains why we're tracking it as a defect. Those fields help the
> product manager understand the problem without treating the agent's severity as
> their decision.
>
> We're also being precise about history. The previous call used applyRate, but that
> alone isn't a tested last-known-good release. We say what we've checked.

### Do

Show SPEED `dashboard/frontend/app/define/[feature]/findings/page.tsx`,
**lines 406–427**; `dashboard/frontend/lib/graphql/queries/feature-defects.ts`,
**lines 71–75**; `dashboard/backend/schema.py`, **lines 1347–1353**; and
`dashboard/backend/resolvers/feature_defects.py`, **lines 111–128**.
Follow into `lib/defect_intake.py`, **lines 182–205**, **422–452**, **462–516**,
and **304–348**.

### Say

> File defect sends the edited report, rationale, and identifiers for the evidence
> we reviewed. The backend checks required fields, evidence revisions, and duplicates.
> If the evidence changed while the draft was open, we need to review it again.
>
> Filing renders the report, prepares its tracking state and decision, and publishes
> them from a staged receipt. That keeps the report and its filing decision connected.
> The source evidence stays attached; the act of filing doesn't strengthen it.

### Do

This fixture already has the fee defect filed. Inspect the duplicate information,
then cancel the demonstration draft and open the existing report:

- `specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md`,
  **lines 12–24**, **30–41**.
- `.speed/defects/capture-fee-truncates-instead-of-rounding-half-up/report.md`,
  **lines 12–24**.
- `.speed/defects/capture-fee-truncates-instead-of-rounding-half-up/state.json`,
  **lines 7–24**.

The existing report's environment is historical, and its named parent commit does
not match this checkout. Use the freshly checked evidence when explaining the current
reproduction and history. Don't create another defect merely to repeat the demo.

### Say

> Here's the saved report we can pass to Plan. Its planning scope already limits the
> implementation change to the capture calculation and keeps the existing regression
> test as evidence. The tracking state is separate from the report's explanation.
>
> Our shared intuition follows evidence into an editable draft and then a saved defect.
>
> Our shared language distinguishes the finding, its source evidence, the draft, and
> the filing decision.
>
> Our shared behavior is to check the facts, explain the impact and needed decisions,
> and review duplicates before filing. Now we'll turn this report into bounded work.

### Notes

- Teach the actual UI fields shown in the supplied screenshots. This is a presenter
  demonstration, not a newly designed learner exercise.
- The CLI is read-only; the UI filing mutation writes. Don't describe the CLI invocation
  itself as creating a defect.
- UI behavior was checked in source and against the supplied screenshots. A live filing
  was not performed for this edit. The existing fee report is already filed.
- Use the current finding and evidence identifiers from the inbox. Don't reuse a stale
  identifier from a screenshot or replace a filed report to make the demo repeatable.

## Plan (108–128 min)

### Three teaching goals

- **Shared intuition:** explain how a defect report becomes task records, and why a
  generated plan still needs review. Follow the report through the planner input and writer.
- **Shared language:** distinguish defect scope, task scope, acceptance criteria,
  dependencies, and provenance. Keep source Task 1 distinct from new repair task numbers.
- **Shared behavior:** check that the plan addresses the reproduced fee defect, uses
  the existing helper and regression test, and introduces no refund-policy decision.

### Say

> Define gave us a report describing the problem, evidence, and repair boundary.
> Plan reads that report and asks an architect agent to turn it into tasks.
>
> A task record is now familiar: it identifies work, files, criteria, and dependencies.
> Plan creates new records for the repair. It hasn't performed the repair or proved
> that its proposed tasks are sufficient.
>
> We'll build intuition about that transformation, use the same terms for its parts,
> and check that the resulting work stays inside this defect's scope.

### Do

Show fixture `specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md`,
**lines 14–24** and **39–41**. Show fixture `src/domain/money.ts`, **lines 38–45**,
and `test/service.test.ts`, **lines 107–114**.

### Say

> The report says to use the existing half-up helper in the capture calculation.
> That helper already exists, and the regression test already checks the two amounts.
> Those are inputs and verification evidence. They don't both need to become edit tasks.
>
> This distinction matters for scope. We need the service change and a passing existing
> test. We don't need to redesign money arithmetic or decide refund idempotency.

### Do

Use a dedicated repair feature in the same fixture. From its root, the command is:

```sh
speed plan --feature payments-fee-repair \
  --defects specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md
```

Before a live demonstration, check that `payments-fee-repair` is available for this
repair. The command creates task state; use a prepared teaching copy for repeat runs.
Keep the original payments Task 1 records available for reference.

In SPEED, show `lib/cmd/plan.sh`, **lines 479–527**,
**388–403**, **634–655**, and **668–685**.

### Say

> The defects option selects defect-driven planning. The command resolves the named
> report and reads its contents. The feature name identifies the repair's task set;
> it doesn't change which source defect we're addressing.
>
> This mode has no primary tech spec from which to derive a test catalog, so the normal
> spec-mode catalog loading doesn't run. It also skips the new-spec vision check. The
> code comment explains that a defect repair isn't a new-spec proposal.
>
> That's why the report must carry the requirement, evidence, and scope clearly. We
> shouldn't assume the planner silently receives every document we opened in class.

### Do

Show SPEED `lib/cmd/plan.sh`, **lines 883–918**, **100–111**, **126–151**,
and **1125–1133**. Show `agents/architect.md`, **lines 1–20**.

### Say

> The command can reuse a cached architect result when the inputs haven't changed.
> Otherwise, it builds a message containing each defect's text and exact path, plus
> the code context gathered for planning. The role describes the architect's job:
> turn the input into a task plan.
>
> The message requires each task to reference its source defect. The provider runs
> the architect and returns structured output, which the shell parses. That reference
> gives us a way to trace proposed work back to the problem it is supposed to solve.

### Do

Show SPEED `lib/cmd/plan.sh`, **lines 1283–1308**, **1323–1372**;
`lib/tasks.sh`, **lines 14–58**; and
`lib/cmd/plan.sh`, **lines 346–380**.

### Say

> Plan writes the architect's tasks as JSON records. It preserves the criteria,
> dependencies, model, and declared files, then adds the current base commit and spec
> references. For defect-driven tasks, matching the source reference attaches the
> source defect path and any declared failure-class or evidence fields.
>
> These are new pending tasks. A task numbered one under the repair feature is different
> from the source Task 1 under payments.
>
> Plan replaces the selected feature's task set. Its guard blocks a defect plan from
> silently deleting an unrelated task set. Our separate repair feature also keeps the
> original Task 1 evidence available through the lesson.

### Do

After the live plan, open its generated files under
`.speed/features/payments-fee-repair/tasks/`. Name each actual JSON file on screen
and verify its line numbers before delivery. No repair task files were generated
for this script edit; don't present an invented plan as checked output.

Compare the actual records with this review checklist:

- The implementation change addresses `src/payments/service.ts`'s capture fee.
- Verification uses the existing RISK-02 assertions in `test/service.test.ts`.
- The plan uses `applyRate` without changing `src/domain/money.ts`.
- Each repair task references the named fee defect, and its criteria state observable
  outcomes. Files read or executed aren't automatically files to modify.
- Dependencies describe real prerequisites; no task decides refund idempotency.

### Say

> Read the task criteria against the report. Can we tell what must be true when the
> work is done? For this defect, the capture uses the required rounding and the existing
> fee/net checks pass. A criterion saying “fix rounding” would leave too much unstated.
>
> A generated plan is proposed work. We review its scope and evidence requirements
> before execution. Planning alone doesn't change the fee result or close the defect.
>
> Our shared intuition follows a defect report into proposed tasks, while keeping
> planning separate from implementation and verification.
>
> Our shared language distinguishes the source defect, repair tasks, criteria,
> dependencies, and the references connecting them.
>
> Our shared behavior is to inspect that connection, keep the repair bounded, and
> require evidence for completion. That's the behavior we'll carry into execution.

### Notes

- Source references checked against SPEED revision `6a3637a`. Plan was inspected, not
  executed for this edit. Generated task counts, IDs, wording, and line numbers remain
  a delivery preflight check.
- This section demonstrates defect-driven Plan. Don't substitute spec-mode behavior
  or claim that it automatically injects the payments test catalog.
- Check generated criteria and file scope before teaching from the output. Do not
  broaden the fee repair to cover the separate refund-policy or retry-coverage findings.
- Keep all displayed paths repository-relative. Use actual current output identifiers.
