test(studio): print PASS for edit accuracy cases that pass every metric - #4925
Conversation
Edit accuracy: accurate 1216 (base branch 1216), smooth 1058 of thoseThe gate passes. Quarantined, measured but not gated (1)
|
terencecho
left a comment
There was a problem hiding this comment.
Approved at be369b870ef809039aab5bd1c2c1f6e4377510ec.
The case-level pass is false on errors and otherwise requires every metric, including smoothness, after the unsettled overrides. That matches the existing full-pass summary while leaving the headline accuracy count's intentional smoothness exclusion unchanged. The run log and evidence-save guard now use that full case result, so a smoothness-only failure reports FAIL and retains evidence; a clean case reports PASS without failure evidence. The added test covers clean, metric-only, smoothness-only and error cases, and its clean-case assertion would fail on the old result shape.
The exact-head Studio test run included report.test.mjs (6 tests) and passed; the edit-accuracy gate passed with accuracy 1216 vs 1216 on base and smoothness 1058. All 11 currently required checks are terminal and passing. I found no blocking issue in this delta.
— Review by tai (pr-review)
What
The edit accuracy bench printed
FAILfor every case in its run log, including cases that passed every metric.score()built each case'schecksbut never setpass, and the run log line (run.mjsverdict) and the failing-case evidence save both readpass. So a bareFAIL <case> <seconds>with no metric after it was really a full pass, and every case saved "failing" evidence.score()now setspassto every metric passing, smoothness included, andfalsefor a case that errored. The log line then readsPASSfor a clean case andFAIL <case> <s> <metrics>only when a metric failed, and evidence is saved for failing cases only, as intended.Scope
Only the run log line and the evidence save read
pass. The summary,table.md, the ratchet gate,baseline.jsonand the sticky comment count all read the per-metricchecks, so their numbers do not change. The per-casepassmatches the summary's "pass everything" count (smoothness included); the headline "accurate" count leaves smoothness out on purpose, so a case can printFAIL <case> <s> smoothand still count as accurate.Tests
report.test.mjs: a case with every metric in range passes; one with a metric out of range fails; an errored case fails. On the oldscore()the first assertion readsundefined. The bench's unit tests pass (8 files, 57 tests).No visible change
The bench's run log and evidence folders only; nothing in Studio or the player changes.
Size
Two lines in
score()and one test: the bug is a missing field. It ships alone because nothing open can carry it: the other bench PRs are drafts that rebase over this one, and the misleading run log affects every bench reader today.