Skip to content

test(studio): print PASS for edit accuracy cases that pass every metric - #4925

Merged
miguel-heygen merged 2 commits into
mainfrom
test/edit-accuracy-case-verdict
Oct 3, 2026
Merged

miguel-heygen merged 2 commits into
mainfrom
test/edit-accuracy-case-verdict

Conversation

@miguel-heygen

@miguel-heygen miguel-heygen commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

What

The edit accuracy bench printed FAIL for every case in its run log, including cases that passed every metric. score() built each case's checks but never set pass, and the run log line (run.mjs verdict) and the failing-case evidence save both read pass. So a bare FAIL <case> <seconds> with no metric after it was really a full pass, and every case saved "failing" evidence.

score() now sets pass to every metric passing, smoothness included, and false for a case that errored. The log line then reads PASS for a clean case and FAIL <case> <s> <metrics> only when a metric failed, and evidence is saved for failing cases only, as intended.

Scope

Only the run log line and the evidence save read pass. The summary, table.md, the ratchet gate, baseline.json and the sticky comment count all read the per-metric checks, so their numbers do not change. The per-case pass matches the summary's "pass everything" count (smoothness included); the headline "accurate" count leaves smoothness out on purpose, so a case can print FAIL <case> <s> smooth and still count as accurate.

Tests

report.test.mjs: a case with every metric in range passes; one with a metric out of range fails; an errored case fails. On the old score() the first assertion reads undefined. The bench's unit tests pass (8 files, 57 tests).

No visible change

The bench's run log and evidence folders only; nothing in Studio or the player changes.

Size

Two lines in score() and one test: the bug is a missing field. It ships alone because nothing open can carry it: the other bench PRs are drafts that rebase over this one, and the misleading run log affects every bench reader today.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Edit accuracy: accurate 1216 (base branch 1216), smooth 1058 of those

The gate passes.
Smoothness is reported in the artifact, not gated. A case fails only if it fails 2 of 3 runs.

Quarantined, measured but not gated (1)

@miguel-heygen
miguel-heygen marked this pull request as ready for review October 3, 2026 01:10

@terencecho terencecho left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved at be369b870ef809039aab5bd1c2c1f6e4377510ec.

The case-level pass is false on errors and otherwise requires every metric, including smoothness, after the unsettled overrides. That matches the existing full-pass summary while leaving the headline accuracy count's intentional smoothness exclusion unchanged. The run log and evidence-save guard now use that full case result, so a smoothness-only failure reports FAIL and retains evidence; a clean case reports PASS without failure evidence. The added test covers clean, metric-only, smoothness-only and error cases, and its clean-case assertion would fail on the old result shape.

The exact-head Studio test run included report.test.mjs (6 tests) and passed; the edit-accuracy gate passed with accuracy 1216 vs 1216 on base and smoothness 1058. All 11 currently required checks are terminal and passing. I found no blocking issue in this delta.

— Review by tai (pr-review)

@miguel-heygen
miguel-heygen added this pull request to the merge queue Oct 3, 2026
Merged via the queue into main with commit 703d0e6 Oct 3, 2026
169 checks passed
@miguel-heygen
miguel-heygen deleted the test/edit-accuracy-case-verdict branch October 3, 2026 01:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants