Lab 3: the payments change, from Diagnose to Plan
Read the main paragraphs aloud; italic text gives sources and presenter instructions.
Fixture paths are relative to 301-config.
SPEED paths are relative to /Users/sanjay/Documents/code/301.
Keep both repositories open so the payments code stays beside the implementation.
This lesson replaces the continuation in docs/lab-3-speaker-script.md.
Source references were checked on 2026-09-29 against SPEED revision 1de583a.
Saved outputs below identify their task or attempt.
Fresh Review output and generated repair Plan tasks still need verification before
delivery.
The continuation uses presenter demonstrations and adds no learner interactions.
Opening (0–8 min)
| Minutes | What happens |
|---|---|
| 0–2½ | Welcome, the problem, and how the afternoon works |
| 2½–3½ | The business background and the change |
| 3½–7 | Walk through the four files together |
| 7–8 | Where the specs are, then bridge to Diagnose |
Hi, I’m [name].
[Coach names] are here to help you through the afternoon.
Quick show of hands.
Who’s reviewed a pull request in the last month that was mostly written by an AI?
Keep your hand up if you read every line.
Yeah.
Me neither.
That’s why we’re here.
The agents write code faster than we can read it.
If our plan is to read harder, we lose.
So today you’ll use tools that do a lot of the reading for you.
But tools get things wrong.
They flag things that are fine.
They miss things.
And some decisions aren’t theirs to make.
If nobody has written down how something should work, a tool can’t decide that for you.
Neither can you.
You find the person who owns that decision, and you ask them.
By the end of today, you’ll be able to take a change an AI wrote and work out what could
be wrong.
You’ll check it with the tools.
Then you’ll say what you’d approve, what you’d send back, and who has to decide the
rest.
Here’s how the afternoon works.
It’s just under three hours, with a break at about the seventy-minute mark.
You’ll sit in pods of four and work in pairs.
Swap who’s typing after every step.
If you get stuck, ask your partner, then your pod, then grab a coach.
At the end, you’ll do one round on your own.
That’s how you and I will both know it stuck.
Some background first.
This is a card payments service.
A merchant reserves money on a customer’s card, then takes the money later.
When goods come back, the merchant refunds the customer.
Every refund is tied to the payment it came from.
Here’s the change.
Sometimes a merchant calls support and says, “I issued a refund, but the money hasn’t
reached my customer’s account.”
The agent was asked to give support a way to send that refund again.
If sending fails, it tries again, up to three times in total.
Why should you care about this one?
It’s payments code.
It moves real money, and it handles card numbers.
A mistake here costs someone money, or leaks card data.
Let’s look at what the agent produced.
Show the agent’s commit; participants run these commands on their machines.
git show --stat 965029f
git show 965029f
The first command lists the four files; the second shows their changes.
Use this commit for the opening.
The whole branch diff also includes course configuration, docs, specs, and test renames.
Four files.
Let’s go through them.
retry.ts is new.
It’s the retry itself.
It sends the refund, and if that fails, it tries again, up to three times.
service.ts is the main payments service.
The agent changed two lines here.
One is in how a failed card check gets logged.
The other is in how the fee is worked out.
cards.ts is test data.
The agent added one card number for testing.
refund-retry.test.ts is the new test.
Its comment says it confirms that a retried refund is accepted.
Scroll through the fixture files in that order: src/payments/retry.ts, lines 10–28;
src/payments/service.ts, lines 62–68 and 118–120;
test/fixtures/cards.ts, lines 18–19;
test/refund-retry.test.ts, lines 6–16.
You also have what the agent was working from.
The product spec says what the payments service must do.
The tech spec says how.
Both are in the specs folder.
Each one names the team that owns it.
That’s who you’d ask.
So that’s the change.
Before anyone decides whether it’s good, let’s see what the tools point at.
Leave both fixture paths visible: specs/product/payments.md, line 3, names
payments-product;
specs/tech/payments.md, line 3, names payments-engineering.
Presenter notes: Keep the approved opening descriptive.
If someone asks whether something is a bug, say, “Hold that thought.
We’ll check it.”
Name where logging and fees changed without deciding either claim.
The old presenter note claimed an unredacted log; the current serializer redacts its
context.
That correction is supported by fixture src/obs/index.ts, lines 49–64.
The retry test never calls the helper; save that explanation for the later checks.
Hold the test spec until Eval; its RETRY-01 row characterizes behavior without approving
a refund policy.
That row is in fixture specs/tests/payments.md, lines 88–92.
Say “AI coding agent”: the agent wrote the code, and the code itself has no AI in it.
Diagnose (8–23 min)
| Minutes | What happens |
|---|---|
| 8–10 | Why Diagnose helps, and what a task record is |
| 10–12 | Introduce the first F2 rule and run the command |
| 12–15 | Follow the task record into the branch diff |
| 15–19 | Follow the added logging line through the matcher |
| 19–21 | Read the signal and its limits |
| 21–23 | Explain the observation and carry the question into Review |
The agent was asked to add refund retries, but it also changed logging and fees.
Diagnose gives us a way to pick out changes that deserve a closer look.
It searches the diff using rules written for this project.
We’ll follow one of those rules to the logging line, then look at how SPEED records the
match.
We first need to tell SPEED which work we’re checking.
It stores each task in a JSON file called a task record.
This one belongs to the payments feature and has the ID 1.
It names the round-0 branch, records the agent model, and lists the four files from the
opening.
We’ll use its branch to find the changes.
The task ID identifies this record; it isn’t a commit number.
Fixture: .speed/features/payments/tasks/1.json, lines 2–11.
Point to id, branch, agent_model, and files_touched.
The project also needs to tell Diagnose what to look for.
In speed.toml, classes_file points to the YAML file beside it.
F2 is the cardholder-data leakage category in that file.
It has two rules, and we’re going to read the first one.
The second stays enabled, so it can still appear in the output.
Fixture: speed.toml, lines 8–10, especially classes_file at line 9;
.speed/classes.yaml, lines 17–22.
Keep the second rule at lines 23–25 folded.
- id: F2
title: Cardholder data leakage
rules:
- look: added-lines
match: '\b(log|logger|webhook|telemetry)\s*\.'
say: "{n} file(s) with an observability sink call in the added lines"
Read the three fields together.
The look field selects added lines, match gives us the pattern, and say gives us the
message to write when something matches.
Here the pattern looks for log, logger, webhook, or telemetry at a word boundary.
After that name, it allows whitespace and requires a dot.
So log dot error fits the pattern.
The name “Cardholder data leakage” tells us why the project includes this rule.
The pattern itself only looks for those names and a dot; it doesn’t check the arguments
or even require a function call.
When a rule matches, Diagnose records an observation called a signal.
It writes these signals into a file called the risk surface, under the project’s failure
categories.
That file also has fields for judgments we’ll make later.
Let’s run Diagnose for the payments task we just opened.
speed diagnose --feature payments --task 1
The console identifies the comparison as main to round-0 and reports two signals under
F2.
Each of F2’s rules found something, which accounts for the two entries.
We’ll inspect the first entry.
To explain where it came from, start with the command that read our task record.
SPEED: lib/cmd/diagnose.sh, lines 109–116, counts signal entries for the console.
Fixture: saved .speed/features/payments/risk-surface.yaml, lines 16–28, contains both
F2 signals.
The shell joins the configured rules path to the fixture root, then reads Task 1’s JSON.
It takes the branch, agent model, and declared files from that record.
The comment calls out three fields it leaves behind: description, acceptance criteria,
and review feedback.
Those fields describe or judge the work; the comment says Diagnose mustn’t take its
verdict from them.
The shell then writes the branch diff into a temporary file for the Python engine to
read.
SPEED: lib/cmd/diagnose.sh, lines 32–54;
lib/tasks.sh, lines 66–76, reads the task JSON;
lib/cmd/diagnose.sh, lines 61–81, especially line 74, creates the diff file;
lib/git.sh, lines 326–330, runs the three-dot comparison.
The Git helper uses a three-dot comparison.
Here that means changes on round-0 since its common ancestor with main.
The opening showed one commit, but this comparison includes the course setup as well.
The task’s four declared files don’t restrict this diff.
This command shows just the service’s portion so we can find the logging change.
git diff main...round-0 -- src/payments/service.ts
log.error(serialiseFailure(error, req));
Fixture: src/payments/service.ts, line 67, is the added logging call.
Git shows this replacement as a whole added line, even though most of the call stayed
the same.
That’s why an added-lines rule will examine it.
The shell sends the diff and rules to Python, along with the declared files and any
configured spec path.
Those last two inputs serve other rules; this rule searches the added text.
SPEED: lib/cmd/diagnose.sh, lines 93–100, invokes lib/diagnose_engine.py;
lib/diagnose_engine.py, lines 598–605, loads the rules and prepares the inputs.
The Python code walks through the diff, remembering the filename from each file header.
The hunk header supplies the line position.
When it reaches an added line, it removes the leading plus sign and keeps the filename,
position, and text together.
For our example, that’s the service file, line 67, and the logging call.
Headers, removed lines, and unchanged lines aren’t entries in this list.
No TypeScript parsing is involved.
SPEED: lib/diagnose_engine.py, lines 315–347, extracts the records.
This replaces the older script’s reference to the previous extraction helper.
The rule’s look field sends that list to the added-lines matcher.
Look at pattern.search: it searches each line using the regular expression from our YAML
file.
On the logging line it finds log dot, so count increases by one.
The matcher keeps the filename once and records the matching line and snippet as
evidence.
That distinction matters if two lines in one file match: the count is two, but the file
list still has one entry.
Several matches on the same line also increase the count only once.
SPEED: lib/diagnose_engine.py, lines 609–612, runs the rules; lines 575–579 select the
matcher; lines 411–432 implement it.
Now read the rule’s message beside that count.
It says “files,” although the code counts matching lines.
We happen to have one line in one file, so the number looks right here.
The engine substitutes that count for the braces in the message and makes one signal for
the rule.
It includes the file list and evidence, then the shell writes them to the risk surface.
SPEED: lib/diagnose_engine.py, lines 591–618, formats and assembles the signal;
lib/cmd/diagnose.sh, lines 135–139 and 158–199, writes the output.
- id: F2
failureMode: "Cardholder data leakage"
plausible: undecided
costOfMissing: undecided
findableByReading: undecided
signals:
- observed: "1 file(s) with an observability sink call in the added lines"
where:
- "src/payments/service.ts"
Fixture: saved .speed/features/payments/risk-surface.yaml, line 2, identifies Task 1.
Lines 16–24 contain this excerpt; it omits the second signal.
The saved file predates the current writer’s detailed evidence fields.
After a fresh run, verify the output and its line numbers again.
There’s the message we just built, with the service path below it.
The three judgment fields remain undecided.
They ask whether the failure is plausible, what missing it would cost, and whether
reading could find it.
The writer leaves them open because the match hasn’t answered those questions.
It hasn’t followed the request into serialiseFailure or checked what the logger
receives.
A redacted logging call would match too, while a logging name outside this pattern could
be missed.
I would describe this result like this.
“The first F2 rule matched the added log.error line in the payments service.
We need to follow serialiseFailure to see what reaches the log.”
That gives another engineer the location and the question to check.
We’ll keep both beside us when we run Review; its input doesn’t include this
risk-surface file.
Shared intuition: we can explain the match by pointing to log dot in the added line.
Shared language: F2 is the category, the YAML supplies the rule, and the recorded
observation is a signal.
A leak is a claim we’d need to check against what actually gets logged.
Shared behavior: open the named code and follow the data before reporting a defect.
Presenter notes: Stay with Task 1 and F2’s first rule.
The audience has seen the business request, four files, and spec paths, but hasn’t
learned SPEED vocabulary before this section.
Use selected shell and Python blocks; the other matchers, YAML parser, and evidence
storage aren’t part of the close reading.
Task 1’s title is a regression-fixture label, and its nested workbench_source_task is
retained metadata from other work.
Neither supplies the payments story or this rule’s input.
Those fields are in fixture .speed/features/payments/tasks/1.json, lines 3 and 12–48.
Presenter notes: Preserve the broad branch diff and the rule’s misleading “files” label.
Leave judgments open despite the console instruction to fill verdicts before running
checks.
That instruction is in SPEED lib/cmd/diagnose.sh, line 143.
Save the serializer walkthrough for Review; don’t stage a leak during Diagnose.
The original script records a Homebrew Bash Diagnose run on 2026-09-29.
This port inspected the saved output and current code without regenerating it.
The script has been reviewed as text, not rehearsed aloud.
Review (23–45 min)
Review asks an agent to examine the diff for problems it can support from the code
shown.
That can give us a more specific claim to check than Diagnose’s text match.
We’ll start with the logging call, then use the fee change to see what else the reviewer
can notice.
Run Review for Task 1, and let’s look at exactly what the agent receives.
speed review --feature payments --task 1
SPEED: lib/cmd/review.sh, lines 194–212, routes --task into _cmd_review_task.
Use --task 1;
--task-id 1 selects a different path in this build.
Task Review reads the same branch from Task 1 and gets its diff against main.
It also reads the author model and declared file list.
Here those are sonnet and the four files we’ve already seen.
The full diff goes into the review message; the file declarations don’t filter it.
Fixture: .speed/features/payments/tasks/1.json, lines 2–11.
SPEED: lib/cmd/review.sh, lines 95–104 and 119–125.
Look at the message template in the shell script.
It inserts the author model, a JSON object containing the file declarations, and the
diff.
Our task’s files_touched list becomes the modified list; created and deleted are empty.
Those labels come from the task record, rather than Git classifying each change.
Here’s the assembled message with just the logging hunk shown.
SPEED: lib/cmd/review.sh, lines 129–140, builds the message.
Line 98 supplies the author model; lines 99–104 build the declarations; lines 119–125
supply the diff.
Recorded author model: sonnet
Declared file lists:
{
"created": [],
"modified": [
"src/payments/retry.ts",
"src/payments/service.ts",
"test/refund-retry.test.ts",
"test/fixtures/cards.ts"
],
"deleted": []
}
Diff (main…round-0):
- log.error(serialiseFailure(error, redact(req)));
+ log.error(serialiseFailure(error, req));
Fixture: src/payments/service.ts, line 67.
These are message excerpts; the file-list JSON is expanded for readability.
The actual message contains the full branch diff.
The reviewer can see both calls: redact(req) has been replaced by req.
It also receives the fee change and the rest of the branch diff.
A separate file supplies its instructions.
That file asks for actionable problems supported by the diff, with uncertainty stated
where information is missing.
It also tells the agent to treat comments and metadata as review material, rather than
follow instructions embedded in them.
SPEED: agents/clean-context-reviewer.md, lines 3–12.
The provider call passes those instructions and the message to the reviewer.
It chooses the review model from SPEED’s configuration and grants read-only tool
permissions.
The sonnet value inside the message is the recorded author model; it doesn’t choose who
reviews the code.
The role file also says not to run commands.
When this reviewer suggests a reproduction, we still have to run it.
SPEED: lib/cmd/review.sh, lines 145–150;
agents/clean-context-reviewer.md, lines 70–71.
Compare that message with Task 1’s acceptance criterion.
The criterion asks whether two refund records exist, but the reviewer doesn’t receive
it.
It receives no separately loaded spec or Diagnose output either.
This lets us examine the whole diff for problems beyond satisfying that one refund
criterion.
Any requirement the reviewer uses must be supported by what is actually in the diff.
Fixture: .speed/features/payments/tasks/1.json, lines 50–54.
SPEED: lib/cmd/review.sh, lines 83–85 and 129–140;
agents/clean-context-reviewer.md, lines 3–12.
Task criterion, omitted from the Review message:
{
"acceptance_criteria": [
{
"criterion": "[RETRY-01] a retried refund against a capture records both refunds",
"verify_by": "test"
}
]
}
The fee change is a useful example.
The rate is 1.49 percent, so a capture of two hundred gives 2.98 before rounding.
The new calculation floors that to two; the test expects three.
Both the calculation and the added test are in this branch diff.
The reviewer can point out the disagreement and suggest running the test.
Fixture: src/payments/service.ts, line 43.
const SCHEME_FEE_RATE = 0.0149;
Fee-change excerpt from main...round-0; current fixture location:
src/payments/service.ts, lines 118–120.
- const fee = applyRate(requested, SCHEME_FEE_RATE);
+ // Fees are never rounded up against the merchant.
+ const fee = minor(Math.floor(requested * SCHEME_FEE_RATE));
const net = sub(requested, fee);
Fixture: test/service.test.ts, lines 107–115.
test('[RISK-02] the scheme fee rounds half up at the capture boundary', () => {
for (const [amount, fee, net] of [[200, 3, 197], [9_999, 149, 9_850]]) {
const s = svc();
const p = auth(s, amount);
s.capture(p.id, amount);
assert.equal(sumAccount(s, p.id, 'scheme_fees'), fee);
assert.equal(sumAccount(s, p.id, 'merchant_settled'), net);
}
});
The reply is JSON containing an issues array.
Each issue needs a message and severity, with a location and reproduction details where
the input supports them.
It can also include observed and expected behavior.
A scenario identifier, such as RISK-02, names a case in the test catalog.
The reviewer may copy it only if it appears in the diff and applies to this finding.
That rule prevents the agent from inventing an identifier just to make its finding look
testable.
SPEED saves the review and archives its evidence so later commands can use it.
Do: Show SPEED agents/clean-context-reviewer.md, lines 23–31 and 37–64;
and lib/cmd/review.sh, lines 160–184.
In the fixture, open .speed/features/payments/reviews/task-1.review:
lines 4–11 contain the logging claim;
lines 13–20 contain the fee claim;
lines 22–29 contain the retry-coverage claim.
Recheck the wording and locations after live Review.
These saved claims aren’t a promised new result.
Selected fields from three saved findings:
{
"issues": [
{
"message": "The authorise() failure path now serialises the raw request instead of the redacted one, so a full PAN and CVV can reach the log sink on every tokenise failure.",
"severity": "critical",
"file": "src/payments/service.ts",
"line": 67,
"scenario_id": "RISK-01"
},
{
"message": "capture() now floors the scheme fee instead of rounding half up, contradicting TR4/RISK-02 and the RISK-02 test added in this same diff.",
"severity": "major",
"file": "src/payments/service.ts",
"line": 119,
"scenario_id": "RISK-02"
},
{
"message": "reissueRefund(), the retry helper this diff introduces, is never called from any source or test file, so its retry loop, attempt limit, and error-rethrow behaviour are entirely unexercised.",
"severity": "major",
"file": "src/payments/retry.ts",
"line": 12,
"scenario_id": "RETRY-01"
}
]
}
Our saved review claims that passing the raw request exposes card data.
Open serialiseFailure and follow what it does with that request.
It calls redact on the context before JSON.stringify builds the string.
The logger then stores that string.
So removing redact at the call site hasn’t established a leak: the helper still does the
redaction.
The saved claim doesn’t hold up against this implementation.
This is why we need to follow a call beyond the lines changed in the diff.
Fixture: src/payments/service.ts, line 67, passes the request to the serializer.
log.error(serialiseFailure(error, req));
Fixture: src/obs/index.ts, lines 49–64, redacts the context and builds the string.
const PAN_FIELDS = new Set(['pan', 'cardNumber', 'primaryAccountNumber', 'cvv', 'cvc']);
export const redact = (value: unknown): unknown => {
if (Array.isArray(value)) return value.map(redact);
if (value && typeof value === 'object') {
return Object.fromEntries(
Object.entries(value as Record<string, unknown>).map(([k, v]) =>
PAN_FIELDS.has(k) ? [k, '[redacted]'] : [k, redact(v)],
),
);
}
return value;
};
export const serialiseFailure = (error: unknown, context: unknown): string =>
JSON.stringify({ error: error instanceof Error ? error.message : String(error), context: redact(context) });
Fixture: src/obs/index.ts, line 30, stores the string in the log sink.
error: (message: string) => sinks.logs.push({ level: 'error', message, at: Date.now() }),
Let’s inspect the stored log from a failure.
The request expects a string for the card number; we’ll deliberately pass a number
instead.
Tokenisation calls replace on it, which throws and takes us into the logging path.
We can then read the captured log using the fixture’s published test card.
Fixture: src/payments/service.ts, lines 35–41 and 62–68;
src/vault.ts, lines 21–26;
test/fixtures/cards.ts, lines 8–16;
src/obs/index.ts, lines 21–31 and 66–70, exposes and resets the captured logs.
Run from the fixture root; the imports name every module used by this demonstration.
node --experimental-strip-types --input-type=module <<'JS'
import { PaymentsService } from './src/payments/service.ts';
import { sinks, resetSinks } from './src/obs/index.ts';
import { TEST_CARDS, EXPIRY } from './test/fixtures/cards.ts';
resetSinks();
try {
new PaymentsService().authorise({
pan: Number(TEST_CARDS.visa), expiry: EXPIRY, amount: 1000,
});
} catch {}
console.log({
records: sinks.logs.length,
redacted: sinks.logs.every(r => JSON.parse(r.message).context.pan === '[redacted]'),
containsTestPan: sinks.logs.some(r => r.message.includes(TEST_CARDS.visa)),
});
JS
The recorded run produced one log record.
Its card-number field was redacted, and the stored message didn’t contain the test card
number.
That answers the logging claim for this failure path.
It doesn’t test every other path through the service.
Also notice that the saved review suggested bad expiry as a reproduction.
This tokenise function never validates expiry, so that suggestion wouldn’t cause the
failure we need.
The original script records running and checking this reproduction.
This port rechecked the source; it hasn’t repeated that execution or generated fresh
Review output.
Recorded result from that reproduction:
{ records: 1, redacted: true, containsTestPan: false }
The saved review also raises a concern about retry coverage.
Let’s compare the helper with the test.
Do: Show fixture src/payments/retry.ts, lines 19–27;
and test/refund-retry.test.ts, lines 9–16, side by side.
The helper catches a failed refund call and tries again, up to its attempt limit.
The test calls refund twice directly, then checks that two refund records exist.
It never calls the helper or makes a refund attempt fail.
So this test doesn’t check the retry loop, despite having retry in its name.
The fee finding also has an existing test we can run.
The spec requires half-up rounding, the service uses floor, and the existing test
expects a fee of three for a capture of two hundred.
Do: Show fixture src/payments/service.ts, lines 118–120;
specs/tech/payments.md, lines 40–43;
and test/service.test.ts, lines 107–115.
We’ve checked the logging claim for the failure path we demonstrated.
The retry test leaves the helper’s behavior unchecked.
The fee calculation disagrees with the expected amount in its test.
Next, we’ll run Eval for Task 1 and examine what evidence it gives us for those
remaining concerns.
Shared intuition: the review gave us claims whose support we could inspect.
Shared language: the finding is what the reviewer says; the requirement and checked
code help us assess it.
Shared behavior: follow the cited code and run a suitable check.
Keep all Review evidence tied to Task 1; a saved task review is historical evidence.
Eval (45–70 min)
Review left us with two concerns: the retry helper isn’t exercised by its test,
and the fee calculation disagrees with the required rounding.
Eval finds existing tests, runs them, and records whether their results satisfy the
selected task’s requirements.
We’ll follow that process for Task 1, then compare the evidence with both findings.
speed eval --feature payments --task 1
The task option matters: we’re asking about Task 1, not every change on the branch.
Start with its task record.
The status is done, which means the work is recorded as completed.
Eval can still find that the completed work fails its checks.
The acceptance criterion tells us what this task is required to demonstrate.
Fixture: .speed/features/payments/tasks/1.json, lines 49–55.
Show these fields together:
{
"status": "done",
"acceptance_criteria": [
{
"criterion": "[RETRY-01] a retried refund against a capture records both refunds",
"verify_by": "test"
}
]
}
Here, the criterion asks for two refund records, and verify_by says to check it with a
test.
That’s narrower than checking whether the helper retries after a failure.
Keep that difference in mind when we read the result.
RETRY-01 is the label connecting this criterion to a case in the test spec.
The test spec is a catalog of cases and their expected results; SPEED calls each case a
scenario.
Open its RETRY-01 entry beside the task record.
It describes the existing two-refund assertion and explicitly leaves refund policy
undecided.
Neither this entry nor the task criterion requires the test to exercise the retry helper.
Fixture: specs/tests/payments.md, lines 88–92.
Introduce this test spec here; the opening introduced only the product and tech specs.
The command hands the task ID and test-spec path to the Python code that prepares the
evaluation.
That code reads the task records and checks the selected task’s status.
If Task 1 isn’t marked done, it stops before running tests.
With that check passed, it works out which tests to run.
SPEED: lib/cmd/eval.sh, lines 247–259 and 274–295;
lib/eval_runtime.py, lines 202–218.
The status check and error are at lines 213–215.
For this execution, test selection starts from Task 1’s saved Review file.
When that file supplies a usable selection, Eval uses it before selecting from task
criteria.
The retry-coverage finding we just read carries the label RETRY-01.
That links it to the same scenario named in Task 1’s criterion.
Eval takes the labels from findings marked testable, including findings where that field
is omitted.
It then looks for existing tests with those labels.
A reviewer’s explanation supplies a question to investigate; Eval doesn’t turn that
explanation into a new test.
Fixture: .speed/features/payments/reviews/task-1.review, lines 22–29.
SPEED: lib/eval_runtime.py, lines 171–187 and 219–236;
lib/eval_review_adapter.py, lines 34–50.
The implementation comment at lib/eval_runtime.py, lines 225–228, explains why a
usable Review takes priority over task criteria.
To find those tests, Eval reads the four files declared in the task record.
The project configuration supplies a pattern identifying which of them are test files.
Compare those two inputs: only the refund test matches.
That’s the file Eval searches for labelled tests belonging to Task 1.
Fixture: .speed/features/payments/tasks/1.json, lines 6–11;
speed.toml, line 6.
The following shows the declared files and the configured pattern:
Task 1 files_touched:
src/payments/retry.ts
src/payments/service.ts
test/refund-retry.test.ts
test/fixtures/cards.ts
test_file_patterns: ["test/*.test.ts"]
Matching file: test/refund-retry.test.ts
The file filter reads files_touched and keeps names matching the configured pattern.
The Review adapter then reads the test names and finds RETRY-01.
It checks that the scenario belongs to this task; here, Task 1’s criterion declares
RETRY-01.
Together, the file list, label, and ownership check select this particular test.
SPEED: lib/eval_execution.py, lines 73–76;
lib/eval_review_adapter.py, lines 112–147 and 230–270.
Fixture: test/refund-retry.test.ts, lines 9–16:
test('[RETRY-01] a retried refund is accepted', () => {
const s = new PaymentsService();
const p = s.authorise({ pan: TEST_CARDS.visa, expiry: EXPIRY, amount: 10_000 });
const c = s.capture(p.id, 10_000);
s.refund(p.id, c.id, 2_500);
s.refund(p.id, c.id, 2_500);
assert.equal(s.payments.get(p.id)!.refunds.length, 2);
});
This is the assertion we read during Review: two direct refund calls produce two records.
It matches Task 1’s criterion.
A pass can satisfy that criterion while leaving the helper’s retry behavior unchecked.
Eval writes the selection into a test plan: the list of tests it intends to execute.
Here are the fields connecting our scenario to its file and test name.
Fixture:
.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/test-plan.json,
lines 2–10.
Selected fields from that plan:
{
"scenario_id": "RETRY-01",
"task_id": "1",
"selector": "test/refund-retry.test.ts",
"runner": "node",
"source": "review_findings",
"test_name": "[RETRY-01] a retried refund is accepted"
}
The selector names the file; test_name identifies the individual test inside it.
The source field records that the selection came from Review.
Next, Eval takes the Node command configured in the project and adds a filter for this
exact test name.
It escapes the brackets as literal characters and puts anchors around the name so a
partial match won’t select another test.
Fixture: speed.toml, lines 1–6.
SPEED: lib/eval_execution.py, lines 350–383, especially lines 375–377 below:
pattern = "^" + re.escape(test_name) + "$"
options += (["--test-name-pattern=" + pattern] if kind == "node"
else ["--testNamePattern=" + pattern])
Node executes the selected test and writes a JUnit report, a structured file containing
the test results.
Eval needs that report; a successful command exit without it isn’t accepted as evidence.
It keeps the command output and result with copies of the task, test spec, and test
plan.
Those files let us trace a reported pass back to the test that actually ran.
Fixture: speed.toml, lines 2–5.
SPEED: lib/eval_runtime.py, lines 284–297 and 333–363;
lib/eval_execution.py, lines 388–405.
Now open the evaluation summary.
It reports one passing scenario and one satisfied criterion.
That’s one test execution: RETRY-01 passed, and its result also satisfied the task’s
criterion.
The report uses the word discharged for a satisfied criterion.
Fixture: .speed/features/payments/eval/task-1/summary.md, lines 3 and 11–13.
Exact summary lines:
**Outcome:** ACCEPTED for the declared scenarios, criteria, and recorded gates.
**Scenarios:** 1 scenarios, 1 pass
**Criteria:** 1 of 1 criteria discharged
**Gates:** 0 total, 0 pass, 0 not examined
Gates are additional checks that can prevent acceptance.
For example, a required scenario missing from the catalog creates a blocking gate.
There are none in this result.
Eval reports accepted because it has an applicable result and every applicable result
passed.
The latest report and summary are published, and the files for this execution are
retained.
SPEED: lib/eval_report.py, lines 442–458, 511–526, and 776–783;
lib/eval_runtime.py, lines 422–462.
Fixture: .speed/features/payments/eval/task-1/summary.md, lines 19–25 and 33–39, links
the result to the selected test.
Now compare the result with the two concerns we carried out of Review.
The two-record assertion passed, satisfying Task 1’s criterion.
As we saw earlier, that test never calls the retry helper, so the coverage concern
remains.
The fee test wasn’t selected, so this run supplies no result for the fee finding.
Fixture: test/refund-retry.test.ts, lines 9–16;
.speed/features/payments/reviews/task-1.review, lines 22–29.
The selection notes explain why: the fee test sits outside Task 1’s declared test files
and belongs to Task 2.
So Task 1’s accepted result leaves the fee finding unanswered.
We’ll run that existing test directly, keeping its result separate from Task 1 Eval.
Fixture: .speed/features/payments/eval/task-1/report.json, lines 74–95, especially
line 88;
.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/test-plan.json,
lines 25–27.
These notes also explain why the logging scenario, RISK-01, wasn’t selected.
The fee test is labelled RISK-02.
Its first case captures 200 minor units and expects 3 in the fee account and 197 in the
merchant account.
Its second case checks the required split for a capture of 9999.
The assertions compare the amounts in those accounts with the expected amounts.
Fixture: test/service.test.ts, lines 102–115.
Lines 102–105 define sumAccount; lines 107–115 contain the test:
test('[RISK-02] the scheme fee rounds half up at the capture boundary', () => {
for (const [amount, fee, net] of [[200, 3, 197], [9_999, 149, 9_850]]) {
const s = svc();
const p = auth(s, amount);
s.capture(p.id, amount);
assert.equal(sumAccount(s, p.id, 'scheme_fees'), fee);
assert.equal(sumAccount(s, p.id, 'merchant_settled'), net);
}
});
Run only the test whose name starts with RISK-02.
node --experimental-strip-types --test --test-name-pattern '^\[RISK-02\]' test/service.test.ts
The first fee assertion fails: actual 2, expected 3.
Execution stops there, before the merchant assertion and the second case.
That result supports Review’s fee finding: the calculation disagrees with the required
rounding.
We now have a specific failure to explain and repair.
Fixture: test/service.test.ts, line 112, is the failing assertion;
src/payments/service.ts, lines 118–139, calculates and posts the fee and merchant net;
specs/tech/payments.md, lines 42–43, states the required split and rounding rule.
Shared intuition: a pass is useful when we know which test ran and what it checked.
Shared language: Task 1 Eval passed RETRY-01; the separate RISK-02 fee check failed.
Shared behavior: read the selected test and its assertion before using the result to
answer a finding.
After the break, we’ll explain the fee error’s impact in a defect report.
Presenter notes: Keep the main walkthrough on this selection path.
Without an explicit test plan, Eval tries a usable Task Review, then tagged task
criteria, then the catalog’s execution mapping.
An explicit plan takes priority over those sources; Review isn’t required to run Eval.
SPEED: lib/eval_runtime.py, lines 219–236;
lib/eval_selection.py, lines 28–39 and 57–100, implements the criteria path.
Leave the refund-policy discussion aside here.
Preparation evidence: The verified Task 1 attempt is
7f8356983ce54c539dee2c77a33694f3, from an uncommitted working tree.
Its command evidence is
.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/commands/8852ef476b5b4c909893d3e4ea51d6dd/result.json,
lines 2–15.
RETRY-01 passed; the separately run RISK-02 failed with actual 2, expected 3.
This rewrite checked the existing evidence and source; it didn’t rerun either test.
Fresh Review output still needs verification before delivery and may change which source
supplies the plan.
Recheck the plan, summary, and evidence paths after that preparation run.
Keep those preparation details out of the spoken explanation.
Don’t substitute Task 2 or feature Eval, or alter the criteria, catalog, or mappings to
put the fee failure into Task 1.
Older notes claiming Eval never consumes Review are incorrect for this implementation.
Break at 70 minutes; resume at 78 minutes.
Think through one defect (78–88 min)
We have a passing Task 1 Eval and a separately reproduced fee failure.
Let’s take that fee failure into a report a product manager can act on.
We’ll use one defect throughout: “Capture fee truncates instead of rounding half up.”
The request was to add refund retries.
The fee failure happens during capture, before any refund is attempted.
That matters when we describe the problem: a merchant can encounter this calculation
without using the new retry helper.
Do: Open fixture test/service.test.ts, lines 107–115.
Point to the capture at line 111 and the assertion at line 112.
Open the captured run in docs/assets/lab-3-define/risk-02-output.txt, lines 13–28.
The test captures 200 minor units and checks the fee recorded in the ledger.
It expects three and gets two.
That’s the failure we reproduced.
The test stops at this assertion, so this run doesn’t reach the merchant assertion or
the second capture case.
A failed test gives us a disagreement to explain.
We still need to establish why the expected value is correct.
Do: Open fixture specs/tech/payments.md, lines 42–43,
and specs/tests/payments.md, line 54.
The tech spec requires half-up rounding and explicitly rules out truncation.
The test spec applies that rule to this capture: fee three, merchant amount 197.
So the test’s expectation has a requirement behind it.
We don’t need the product manager to choose a rounding rule.
Do: Show fixture src/payments/service.ts, line 43 and lines 118–139.
Follow the fee calculation at line 119 into the subtraction at line 120,
then the merchant and fee postings at lines 126–139.
Keep all the amounts in minor units, as the test does.
The capture is 200 and the fee rate is 1.49 percent.
Multiplying those gives 2.98.
The required rounding gives three; Math.floor discards the fraction and gives two.
The service subtracts the fee from the capture to calculate the merchant amount.
With a fee of two, that leaves 198.
With the required fee of three, it should leave 197.
The ledger therefore assigns one minor unit too little to the fee account and one too
much to the merchant.
Both splits total 200, so a check of the total alone would miss this error.
That gives us the opening of our report.
For Observed, we’ll write:
Capturing 200 minor units records a scheme fee of 2 and a merchant amount of 198. The RISK-02 fee assertion fails: actual 2, expected 3.
Then Expected explains the rule and the required amounts:
TR4 requires half-up rounding. For a capture of 200 minor units, record a fee of 3 and a merchant amount of 197.
Those sentences give the reader something concrete to compare.
The first reports the behavior we checked; the second gives its required replacement.
The merchant amount follows from the calculation and posting we just read, rather
than from an assertion that never ran.
Now explain why someone should care.
Write that the fee account receives too little and the merchant account too much.
The one-unit example demonstrates the wrong allocation.
To understand its wider impact, we would need to find which deployed versions contain
this calculation and how many payments it affects.
We haven’t established either fact in this fixture.
Do: Show fixture specs/product/payments.md, line 3, naming payments-product.
Keep the reproduction output beside the report wording.
Our request to that owner will be:
Set repair priority and agree whether to investigate deployed versions and affected payments. Production exposure and frequency are unknown.
The engineering repair can follow the existing rounding requirement.
The product decision concerns urgency and the scope of follow-up.
Keep those separate so the report asks for a decision the reader actually needs to make.
We also have to make the failure repeatable.
We’ll include the exact fee-test command, its failure output, and the environment that
produced it.
That evidence came from our separate RISK-02 check.
Task 1’s passing RETRY-01 result doesn’t supply it.
Shared intuition: a defect report connects a failure to a requirement and a
consequence someone can understand.
Shared language: observed describes what happened; expected names the required
behavior; impact explains who receives the wrong amount.
Shared behavior: attach a reproduction, identify what remains unknown, and ask for a
specific decision.
Now let’s put those exact facts into Define.
Define (88–108 min)
Define brings the tool finding, our explanation, and the decision about what happens
next into one place.
We’ll follow the fee finding from Task 1 into an editable defect draft.
The fixture already contains a report for this fee defect.
We’ll prepare the draft to show what we need to add, then inspect that existing report
instead of filing another copy.
Do: From the fixture root, run the command below.
Open http://localhost:3000/define/payments/findings.
Filter Source to Clean review and Task ID to 1.
Expand the finding beginning “capture() now floors the scheme fee instead of rounding
half up”. Use this wording and source to distinguish it from the older fee findings.
speed define feature payments

Actual Define UI, captured on 2026-09-29. This selected finding is Unresolved; the existing report shown later comes from an older finding for the same fee defect.
The top of the card identifies Clean review and Task 1.
Inside it are the reviewer’s claim, the service location, and a suggested reproduction.
We checked that reproduction ourselves.
The card’s text is still the review evidence; our separate test execution hasn’t been
automatically added to it.
Notice the two visible actions: File new defect and Send back to task.
The recommended action is to send the finding back to the feature task.
For this walkthrough we’re examining the defect-report route, because we need to
explain and track the fee regression separately from the retry work.
A recommendation doesn’t make that choice for us.
Do: In SPEED, show lib/cmd/define.sh, lines 11–29,
and lib/define_cli.py, lines 79–89.
Then show lib/defect_findings.py, lines 467–491 and 549–569.
The command reads the feature’s findings.
The reader combines tool evidence with saved decisions and returns each finding’s
source, status, and available actions.
That’s why this view can show more than the last review message.
The CLI can open an already-running dashboard; it doesn’t start one.
Do: Click File new defect on this fee finding. Inspect the initial form before entering anything.

The form says NO WRITE.
Opening it has prepared a draft; it hasn’t created a defect.
The related feature is already payments, and Reproduction contains the reviewer’s
suggestion.
But Title, Severity, Observed, and Expected are empty.
Do: In SPEED, show lib/defect_intake.py, lines 125–161.
In the fixture, show
.speed/features/payments/evidence/clean_review/1/467d7ec1-ef20-43ac-80c0-00ece22c69b3.json,
lines 43–58: expected and observed are null; reproduction and summary contain text.
observed=_first_fact(evidence, "observed"),
expected=_first_fact(evidence, "expected"),
reproduction=_first_fact(evidence, "reproduction"),
These assignments copy facts from matching evidence fields.
They don’t rewrite the review message into a complete report.
The title is also blank because this finding’s message exceeds the preview’s
hundred-character limit.
This is where the explanation we just worked through becomes useful: we supply the
observed behavior, required behavior, and evidence ourselves.
Do: Enter the following values. The full copyable field values used in every draft screenshot are in fee-defect-draft.json. That file is lesson material, not a saved SPEED defect.
| Field | Text to enter |
|---|---|
| Title | Capture fee truncates instead of rounding half up |
| Severity | P1 |
| Related features | payments |
| Reproducibility | always |
| Last known working | Unknown at runtime; parent d08a680 uses applyRate, but was not tested. |
For this demonstration, P1 carries forward the reported severity on the existing
defect.
Review’s “major” label didn’t automatically become P1.
Confirming the form’s severity records our review of that choice; the one-unit example
by itself doesn’t establish production frequency or repair priority.
Always refers to the reproduced capture of 200.
The old calculation is visible in the parent commit, but we haven’t run that revision.
That’s why Last known working says unknown at runtime.
Do: Check I reviewed and confirm this severity for the demonstration draft.
Verify the existing report’s severity in fixture
specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md, line 3.
For the history, run git show 965029f^:src/payments/service.ts:
line 118 in that historical file uses applyRate; the parent resolves to d08a680.
For Observed, enter the sentence we prepared:
Capturing 200 minor units records a scheme fee of 2 and a merchant amount of 198. The RISK-02 fee assertion fails: actual 2, expected 3.
For Expected, include the requirement and both cases already in the regression test:
TR4 requires half-up rounding. For a capture of 200 minor units, record a fee of 3 and a merchant amount of 197. RISK-02 also requires fee 149 and merchant amount 9850 for a capture of 9999. Fee plus merchant amount must equal the capture.
Do: Keep fixture specs/tech/payments.md, lines 42–43,
and test/service.test.ts, lines 107–115, beside these fields.

Now a reader can compare actual and required behavior without reconstructing the fee
calculation.
The second case is an expected result from the test; the failing run stopped at the
first case.
We haven’t described it as a second observed test failure.
Do: Replace the generated Reproduction text with these numbered steps. Use real line breaks in the form.
1. From the 301-config fixture root, run:
node --experimental-strip-types --test --test-name-pattern '^\[RISK-02\]' test/service.test.ts
2. Observe test/service.test.ts:112 fail: actual 2, expected 3.
3. Inspect src/payments/service.ts:119-120: Math.floor gives fee 2; subtracting it from 200 gives merchant amount 198.
The reviewer’s suggestion described a case.
These steps tell another engineer exactly how to run it and where to look.
Environment records the checkout and runtime that produced the error.
Error Output keeps the actual failure available beside those steps.
Do: Enter Environment and Error Output from
fee-defect-draft.json.
The environment is the preparation run on 2026-09-29: round-0 at aa7cebc with an
uncommitted working tree, macOS on arm64, Node v26.9.0, and node:test.
The data comes from fixture test/fixtures/cards.ts, lines 8–16.
The screenshot contains an error excerpt; the complete captured output is
risk-02-output.txt, lines 1–30.
Refresh environment and output together if demonstrating a later run.

Additional Context is where we explain the consequence and ask for a decision.
Keep the source reference that Define already supplied, then add our checked impact,
what remains unknown, and the proposed repair scope.
Do: Enter Additional Context and Filing rationale from the field-values file. The added context is reproduced below; preserve the existing Selected evidence text after it, as shown in the screenshot.
Impact: for the reproduced capture of 200, the fee account receives one minor unit too little and the merchant account one too much. The total remains 200.
Evidence: the fee failure was reproduced separately from Task 1 Eval, which passed RETRY-01. Reproducibility 'always' refers to this capture case. Production exposure and frequency are unknown.
Decision for payments-product: set repair priority and agree whether to investigate deployed versions and affected payments.
Repair scope: use the existing applyRate helper in src/payments/service.ts; verify with the existing RISK-02 test. Keep retry coverage and refund policy separate.
For Filing rationale, enter:
The existing rounding requirement and failing RISK-02 test establish a fee regression. Track this bounded repair separately from retry-helper coverage and the unresolved refund-idempotency policy.
Do: For the proposed repair, verify fixture src/domain/money.ts, lines 38–45;
src/payments/service.ts, lines 118–120;
and test/service.test.ts, lines 107–115.
For the decision owner, verify specs/product/payments.md, line 3.

The report now tells the product manager which account receives the wrong amount and
what we need decided.
It tells an engineer how to reproduce the problem and which existing rule governs the
repair.
The source reference keeps both readers connected to the original Review finding.
The File defect button is now enabled.
Before using it, look at what that action sends and what the backend checks.
Do: In SPEED, show
dashboard/frontend/app/define/[feature]/findings/page.tsx, lines 406–427,
and lib/defect_intake.py, lines 423–451.
The browser sends our edited fields together with the finding and evidence identifiers.
It also sends the versions of the evidence and decisions that we read.
The backend checks that those versions still match, validates the required fields,
and rejects a title already used by a defect.
So an enabled button means the browser has enough information to submit; it doesn’t
establish that this is a new issue.
In this fixture the fee defect already exists.
We won’t submit another report with the same title.
The initial preview didn’t show a duplicate warning because its generated title was
blank.
We checked the existing report ourselves.
Do: Click Cancel. Open the existing report at
http://localhost:3000/editor?spec=specs%2Fdefects%2Fcapture-fee-truncates-instead-of-rounding-half-up.md.
Keep the editor in preview mode.

This is the existing filed report, not the draft we just cancelled.
It describes the same fee failure and already limits the repair to the service
calculation, using the existing helper and test.
Its environment and source finding belong to an earlier run.
Use the current reproduction we just prepared when discussing today’s evidence.
Do: Read fixture
specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md, lines 12–24 and
39–48. Its source is the older finding finding-9fec9618519e39f4.
The screenshot draft used finding-5c661cea21154dd6, the current saved Review’s fee claim.
The report’s line 7 incorrectly names 34a4e40 as the parent; this checkout resolves
965029f^ to d08a680. Line 28 records an older environment.
These historical details were not changed to stage the demonstration.
To understand what filing persists, compare the report with its tracking record.
The report explains the problem.
The state records that it was filed and identifies its source finding and evidence.
Neither says the fee calculation has been repaired.
Do: Open fixture
.speed/defects/capture-fee-truncates-instead-of-rounding-half-up/state.json, lines
7–24, and
.speed/defects/capture-fee-truncates-instead-of-rounding-half-up/report.md, lines
12–24. In SPEED, show lib/defect_intake.py, lines 462–495 and 513–516,
then lines 326–348.
The filing code prepares the report, tracking state, and decision together, then
publishes them.
The decision retains which finding and evidence led to this defect.
That gives the next person a route back from the work we propose to the problem we
actually checked.
We’ll give Plan the existing report’s path next.
Its repair scope calls for the existing rounding helper and regression test.
That lets us judge the proposed work against a specific failure and a bounded repair.
Shared intuition: Define preserves the evidence and our decision in a report
someone else can act on.
Shared language: the finding is the tool’s claim; the draft is our editable report;
filed means the report and its tracking decision have been saved.
Shared behavior: fill the report with checked facts, retain its source, and inspect
existing defects before filing another.
Presenter preparation: All six images are actual local UI captures from 2026-09-29, not mockups. The draft fields were entered and checked in the browser, then cancelled; no filing or other GraphQL mutation was sent. The current selected finding is unresolved, while the older fee report exists. The inbox also warns about incomplete intake transactions on other findings; this lesson neither repairs them nor presents the inbox as clean. Use the current source and wording if finding identifiers change. Capture provenance is in capture-manifest.json. The fee check was rerun and failed as shown. No fresh Review or full Eval was run for these captures. The source and screenshot checks do not constitute an aloud rehearsal. This remains a presenter demonstration, with no added learner exercise.
Plan (108–128 min)
We have a defect report that explains the fee error and limits the repair.
Plan reads that report and asks an architect agent to propose tasks for the repair.
For this defect, the intended work is already clear: use the existing rounding helper in
capture and run the existing fee test.
We’ll check whether the proposed tasks actually describe that work.
Fixture: specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md, lines
14–24 and 39–41.
Look at applyRate before we send the report to the planner.
It multiplies the amount by the rate, rounds the absolute result, and restores the sign.
Its comment states the half-up rule for positive and negative amounts.
The helper already exists, and RISK-02 already checks the two required fee/net splits.
We need the service to use this helper; reading the helper and running the test doesn’t
require editing either file.
Fixture: src/domain/money.ts, lines 38–45;
test/service.test.ts, lines 107–115;
src/payments/service.ts, lines 118–120.
export const applyRate = (amount: Minor, rate: number): Minor => {
const raw = amount * rate;
const sign = raw < 0 ? -1 : 1;
return minor(sign * Math.round(Math.abs(raw)));
};
We’ll put the new repair tasks under payments-fee-repair.
That keeps the original payments Task 1 and its evidence available while we plan the
repair.
The defects option names the report to read; the feature option names the task set to
write.
Both matter because Plan replaces the task files under the selected feature.
speed plan --feature payments-fee-repair \
--defects specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md
Before the live demo, check that payments-fee-repair is available for this repair.
Use a prepared teaching copy for repeat runs.
SPEED: lib/cmd/plan.sh, lines 479–527 and 1283–1308.
The overwrite guard blocks an existing task set with no defect-sourced tasks unless
force is requested.
It isn’t a general guarantee that every existing task set is protected.
The command resolves the defect path and reads the report’s contents.
It doesn’t take the normal spec-planning route that loads the product, tech, design, and
test specs.
Normally it uses the primary tech spec’s name to find the test catalog.
With a defect report as input, that step is skipped.
That’s why our report states the rounding requirement and names the test to run.
The new-spec vision check is skipped too: the comment explains that a defect repair
isn’t a proposal for a new spec.
SPEED: lib/cmd/plan.sh, lines 388–403, 479–527, 634–655, and 668–685.
Plan puts the defect text and its exact path into the architect’s message, along with
the gathered code context.
The role file tells the architect how to turn that input into a task plan.
The message requires each proposed task to reference its source defect file.
That gives us a concrete check later: every task in this repair should point back to the
fee report.
If the checked inputs match a cached architect result, Plan can reuse it; otherwise it
calls the provider and parses the returned JSON.
SPEED: lib/cmd/plan.sh, lines 883–918, assembles the message or reads the cache;
lines 100–111 and 126–151 prepare the provider call and parse the result;
lines 1125–1133 run the architect phase.
agents/architect.md, lines 1–20, defines the role and its planning contract.
The writer turns those proposed tasks into JSON records.
Acceptance criteria describe what must be true when a task is complete; dependencies
name tasks that must finish first.
The writer preserves those fields, the model, and declared files, then adds the current
base commit and source references.
When a reference matches the fee report, it copies the source defect path and any
declared failure-class or evidence fields.
Those fields retain their source values; planning doesn’t strengthen the evidence they
refer to.
The new tasks start as pending work.
SPEED: lib/cmd/plan.sh, lines 1323–1372;
lib/tasks.sh, lines 14–58;
lib/cmd/plan.sh, lines 346–380, matches the defect reference and attaches its
metadata.
Open the generated task records and start with the files they propose to change.
The implementation change should be in the service, using applyRate without modifying
the shared money helper.
Verification should run the existing RISK-02 test without rewriting its assertions.
Then read the criteria against the report: the capture must use the required rounding
and produce the expected fee/net splits.
“Fix rounding” alone wouldn’t say enough to check completion.
Any dependencies should represent actual prerequisites, and no task should decide refund
idempotency as part of this repair.
Fixture sources for that comparison:
specs/defects/capture-fee-truncates-instead-of-rounding-half-up.md, line 41;
src/payments/service.ts, lines 118–120;
src/domain/money.ts, lines 38–45;
test/service.test.ts, lines 107–115.
Finally, check the source reference on each task.
It should name the fee defect we supplied, so we can explain why that work belongs in
the repair.
A task numbered one under payments-fee-repair is a new record, separate from source Task
1 under payments.
The proposed plan still needs our review before execution.
Creating the task records hasn’t changed the fee calculation or closed the defect.
Shared intuition: the report explains the problem; the plan proposes the work and
checks needed to repair it.
Shared language: criteria describe completion, dependencies describe order, and
source references preserve where the work came from.
Those source references are the plan’s provenance.
Shared behavior: compare the proposed files and criteria with the defect, then
require test evidence before calling the repair complete.
Presenter notes: Plan has been inspected in source, not executed for this lesson
preparation.
Generated output under .speed/features/payments-fee-repair/tasks/ still needs
verification.
After the live preparation run, name every actual JSON file and its line range beside
the comparison above.
No task count, identifier, wording, or generated criterion is asserted here as checked
output.
Use the actual current output; don’t present a suggested task as a generated one.
Keep defect-driven behavior distinct from spec-mode planning: the payments test catalog
isn’t automatically injected in this mode.
Check scope and source references before delivery, and keep the fee repair separate from
retry coverage and refund policy.