Chapter 04 / 07
Eval
Trace which test ran, what its pass establishes, and why the separate fee check is still needed.
What you should leave with
- Shared intuition
- a pass is useful when we know which test ran and what it checked.
- Shared language
- Task 1 Eval passed RETRY-01; the separate RISK-02 fee check failed.
- Shared behavior
- read the selected test and its assertion before using the result to answer a finding. After the break, we'll explain the fee error's impact in a defect report.
The task and its criterion
Review left us with two concerns: the retry helper isn’t exercised by its test,
and the fee calculation disagrees with the required rounding.
Eval finds existing tests, runs them, and records whether their results satisfy the
selected task’s requirements.
We’ll follow that process for Task 1, then compare the evidence with both findings.
speed eval --feature payments --task 1The task option matters: we’re asking about Task 1, not every change on the branch.
Start with its task record.
The status is done, which means the work is recorded as completed.
Eval can still find that the completed work fails its checks.
The acceptance criterion tells us what this task is required to demonstrate.
Source references
Fixture: .speed/features/payments/tasks/1.json, lines 49–55.
Show these fields together:
{
"status": "done",
"acceptance_criteria": [
{
"criterion": "[RETRY-01] a retried refund against a capture records both refunds",
"verify_by": "test"
}
]
}Here, the criterion asks for two refund records, and verify_by says to check it with a
test.
That’s narrower than checking whether the helper retries after a failure.
Keep that difference in mind when we read the result.
RETRY-01 is the label connecting this criterion to a case in the test spec.
The test spec is a catalog of cases and their expected results; SPEED calls each case a
scenario.
Open its RETRY-01 entry beside the task record.
It describes the existing two-refund assertion and explicitly leaves refund policy
undecided.
Neither this entry nor the task criterion requires the test to exercise the retry helper.
Source references
Fixture: specs/tests/payments.md, lines 88–92.
Introduce this test spec here; the opening introduced only the product and tech specs.
The command hands the task ID and test-spec path to the Python code that prepares the
evaluation.
That code reads the task records and checks the selected task’s status.
If Task 1 isn’t marked done, it stops before running tests.
With that check passed, it works out which tests to run.
Source references
SPEED: lib/cmd/eval.sh, lines 247–259 and 274–295;
lib/eval_runtime.py, lines 202–218.
The status check and error are at lines 213–215.
Trace test selection
For this execution, test selection starts from Task 1’s saved Review file.
When that file supplies a usable selection, Eval uses it before selecting from task
criteria.
The retry-coverage finding we just read carries the label RETRY-01.
That links it to the same scenario named in Task 1’s criterion.
Eval takes the labels from findings marked testable, including findings where that field
is omitted.
It then looks for existing tests with those labels.
A reviewer’s explanation supplies a question to investigate; Eval doesn’t turn that
explanation into a new test.
Source references
Fixture: .speed/features/payments/reviews/task-1.review, lines 22–29.
SPEED: lib/eval_runtime.py, lines 171–187 and 219–236;
lib/eval_review_adapter.py, lines 34–50.
The implementation comment at lib/eval_runtime.py, lines 225–228, explains why a
usable Review takes priority over task criteria.
To find those tests, Eval reads the four files declared in the task record.
The project configuration supplies a pattern identifying which of them are test files.
Compare those two inputs: only the refund test matches.
That’s the file Eval searches for labelled tests belonging to Task 1.
Source references
Fixture: .speed/features/payments/tasks/1.json, lines 6–11;
speed.toml, line 6.
The following shows the declared files and the configured pattern:
Task 1 files_touched:
src/payments/retry.ts
src/payments/service.ts
test/refund-retry.test.ts
test/fixtures/cards.ts
test_file_patterns: ["test/*.test.ts"]
Matching file: test/refund-retry.test.tsThe file filter reads files_touched and keeps names matching the configured pattern.
The Review adapter then reads the test names and finds RETRY-01.
It checks that the scenario belongs to this task; here, Task 1’s criterion declares
RETRY-01.
Together, the file list, label, and ownership check select this particular test.
Source references
SPEED: lib/eval_execution.py, lines 73–76;
lib/eval_review_adapter.py, lines 112–147 and 230–270.
Fixture: test/refund-retry.test.ts, lines 9–16:
test('[RETRY-01] a retried refund is accepted', () => {
const s = new PaymentsService();
const p = s.authorise({ pan: TEST_CARDS.visa, expiry: EXPIRY, amount: 10_000 });
const c = s.capture(p.id, 10_000);
s.refund(p.id, c.id, 2_500);
s.refund(p.id, c.id, 2_500);
assert.equal(s.payments.get(p.id)!.refunds.length, 2);
});Read the selected test
This is the assertion we read during Review: two direct refund calls produce two records.
It matches Task 1’s criterion.
A pass can satisfy that criterion while leaving the helper’s retry behavior unchecked.
Eval writes the selection into a test plan: the list of tests it intends to execute.
Here are the fields connecting our scenario to its file and test name.
Source references
Fixture:
.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/test-plan.json,
lines 2–10.
Selected fields from that plan:
{
"scenario_id": "RETRY-01",
"task_id": "1",
"selector": "test/refund-retry.test.ts",
"runner": "node",
"source": "review_findings",
"test_name": "[RETRY-01] a retried refund is accepted"
}The selector names the file; test_name identifies the individual test inside it.
The source field records that the selection came from Review.
Next, Eval takes the Node command configured in the project and adds a filter for this
exact test name.
It escapes the brackets as literal characters and puts anchors around the name so a
partial match won’t select another test.
Source references
Fixture: speed.toml, lines 1–6.
SPEED: lib/eval_execution.py, lines 350–383, especially lines 375–377 below:
pattern = "^" + re.escape(test_name) + "$"
options += (["--test-name-pattern=" + pattern] if kind == "node"
else ["--testNamePattern=" + pattern])Node executes the selected test and writes a JUnit report, a structured file containing
the test results.
Eval needs that report; a successful command exit without it isn’t accepted as evidence.
It keeps the command output and result with copies of the task, test spec, and test
plan.
Those files let us trace a reported pass back to the test that actually ran.
Source references
Fixture: speed.toml, lines 2–5.
SPEED: lib/eval_runtime.py, lines 284–297 and 333–363;
lib/eval_execution.py, lines 388–405.
Now open the evaluation summary.
It reports one passing scenario and one satisfied criterion.
That’s one test execution: RETRY-01 passed, and its result also satisfied the task’s
criterion.
The report uses the word discharged for a satisfied criterion.
Source references
Fixture: .speed/features/payments/eval/task-1/summary.md, lines 3 and 11–13.
Exact summary lines:
**Outcome:** ACCEPTED for the declared scenarios, criteria, and recorded gates.
**Scenarios:** 1 scenarios, 1 pass
**Criteria:** 1 of 1 criteria discharged
**Gates:** 0 total, 0 pass, 0 not examinedGates are additional checks that can prevent acceptance.
For example, a required scenario missing from the catalog creates a blocking gate.
There are none in this result.
Eval reports accepted because it has an applicable result and every applicable result
passed.
The latest report and summary are published, and the files for this execution are
retained.
Source references
SPEED: lib/eval_report.py, lines 442–458, 511–526, and 776–783;
lib/eval_runtime.py, lines 422–462.
Fixture: .speed/features/payments/eval/task-1/summary.md, lines 19–25 and 33–39, links
the result to the selected test.
Bound the accepted result
Now compare the result with the two concerns we carried out of Review.
The two-record assertion passed, satisfying Task 1’s criterion.
As we saw earlier, that test never calls the retry helper, so the coverage concern
remains.
The fee test wasn’t selected, so this run supplies no result for the fee finding.
Source references
Fixture: test/refund-retry.test.ts, lines 9–16;
.speed/features/payments/reviews/task-1.review, lines 22–29.
The selection notes explain why: the fee test sits outside Task 1’s declared test files
and belongs to Task 2.
So Task 1’s accepted result leaves the fee finding unanswered.
We’ll run that existing test directly, keeping its result separate from Task 1 Eval.
Source references
Fixture: .speed/features/payments/eval/task-1/report.json, lines 74–95, especially
line 88;
.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/test-plan.json,
lines 25–27.
These notes also explain why the logging scenario, RISK-01, wasn’t selected.
Run the fee check separately
The fee test is labelled RISK-02.
Its first case captures 200 minor units and expects 3 in the fee account and 197 in the
merchant account.
Its second case checks the required split for a capture of 9999.
The assertions compare the amounts in those accounts with the expected amounts.
Source references
Fixture: test/service.test.ts, lines 102–115.
Lines 102–105 define sumAccount; lines 107–115 contain the test:
test('[RISK-02] the scheme fee rounds half up at the capture boundary', () => {
for (const [amount, fee, net] of [[200, 3, 197], [9_999, 149, 9_850]]) {
const s = svc();
const p = auth(s, amount);
s.capture(p.id, amount);
assert.equal(sumAccount(s, p.id, 'scheme_fees'), fee);
assert.equal(sumAccount(s, p.id, 'merchant_settled'), net);
}
});Run only the test whose name starts with RISK-02.
node --experimental-strip-types --test --test-name-pattern '^\[RISK-02\]' test/service.test.tsThe first fee assertion fails: actual 2, expected 3.
Execution stops there, before the merchant assertion and the second case.
That result supports Review’s fee finding: the calculation disagrees with the required
rounding.
We now have a specific failure to explain and repair.
Source references
Fixture: test/service.test.ts, line 112, is the failing assertion;
src/payments/service.ts, lines 118–139, calculates and posts the fee and merchant net;
specs/tech/payments.md, lines 42–43, states the required split and rounding rule.
Shared intuition: a pass is useful when we know which test ran and what it checked.
Shared language: Task 1 Eval passed RETRY-01; the separate RISK-02 fee check failed.
Shared behavior: read the selected test and its assertion before using the result to
answer a finding.
After the break, we’ll explain the fee error’s impact in a defect report.
Preparation and evidence notes
Presenter notes: Keep the main walkthrough on this selection path.
Without an explicit test plan, Eval tries a usable Task Review, then tagged task
criteria, then the catalog’s execution mapping.
An explicit plan takes priority over those sources; Review isn’t required to run Eval.
SPEED: lib/eval_runtime.py, lines 219–236;
lib/eval_selection.py, lines 28–39 and 57–100, implements the criteria path.
Leave the refund-policy discussion aside here.
Preparation and evidence notes
Preparation evidence: The verified Task 1 attempt is
7f8356983ce54c539dee2c77a33694f3, from an uncommitted working tree.
Its command evidence is
.speed/features/payments/eval/task-1/runs/7f8356983ce54c539dee2c77a33694f3/commands/8852ef476b5b4c909893d3e4ea51d6dd/result.json,
lines 2–15.
RETRY-01 passed; the separately run RISK-02 failed with actual 2, expected 3.
This rewrite checked the existing evidence and source; it didn’t rerun either test.
Fresh Review output still needs verification before delivery and may change which source
supplies the plan.
Recheck the plan, summary, and evidence paths after that preparation run.
Keep those preparation details out of the spoken explanation.
Don’t substitute Task 2 or feature Eval, or alter the criteria, catalog, or mappings to
put the fee failure into Task 1.
Older notes claiming Eval never consumes Review are incorrect for this implementation.
Break at 70 minutes; resume at 78 minutes.
Carry forward
Task 1 Eval passed RETRY-01. The separate RISK-02 check failed. Use that fee evidence to explain one bounded defect.
Help me reason through this
Read the selected test’s assertion. Task 1 acceptance and the separate fee reproduction answer different questions.
Your explanation is saved in this browser.