port: add check-all; verify-screen ignored its own statistic; 'six expected DIFFERS' was wrong
Eleven tools and nothing ran them together -- the ninth instance of correct, documented and unexercised, one level up. check-all runs the four that assert, reports the oracle table, and gives verify-screen an allowance that EXPIRES when the pin lands rather than standing forever. All eleven exercised first; none had rotted. verify-screen computed over3 because 'a single max cannot tell 2 pixels from 25 444' and then decided the verdict on max alone: main_menu (max 4, over3 0) read DIFFERS while extras (max 3, over3 0) read OK. The bar is unchanged; a frame with no pixel over it now gets its own ROUNDING verdict. And corrects a claim I have given the Decoder more than once. The real count was ten, now eight: six forced-backdrop, two rounding, and TWO UNEXPLAINED -- title at 790 px and title_jp at 20498, neither carrying a forced element. My leaf hypothesis is refuted: emptying draw_leaf_for changes the numbers not at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
This commit is contained in:
@@ -602,3 +602,31 @@ tree. So the candidate fixes are about the *message*, not the order:
|
||||
|
||||
⚠️ Deliberately not chosen here: both change failure semantics across every tool,
|
||||
and neither is measured against anything. It goes to whoever owns that call.
|
||||
|
||||
---
|
||||
|
||||
## `title` differs from `sylpheed-cli` on 790 pixels, and the leaf theory is refuted
|
||||
|
||||
*Derived from HANDOFF `9ca1eb5`. Raised 2026-08-30 by the port. Not urgent — the
|
||||
port agrees with the **oracle** on this screen at 0.21 %, and that is the check
|
||||
that counts.*
|
||||
|
||||
`verify-screen` has `title` at **790 pixels** over the bar and `title_jp` at
|
||||
**20 498**, and neither carries a forced-backdrop element — so they are not
|
||||
covered by the pinned-tag allowance the other six sit under. I had been reporting
|
||||
"six expected DIFFERS"; the real count was ten.
|
||||
|
||||
🔴 **The obvious explanation is wrong.** `authored/rendering.json` records that the
|
||||
consistency harness compares against a renderer drawing no `.rat` leaves, so the
|
||||
`ptloop` sweeps were the candidate. Emptying `draw_leaf_for` and
|
||||
`loop_leaf_on_screens` changes the figures **not at all** — the harness poses at
|
||||
`rest`, where the leaves do not draw.
|
||||
|
||||
What would help: **which elements `sylpheed-cli` draws on build 4 at `rest`**, as a
|
||||
list. The port's own list is in any `--screen=title` log (`drew N: …`). A set
|
||||
difference answers it immediately, and it is a question about our own tool rather
|
||||
than about the game — no oracle run, no emulator.
|
||||
|
||||
⚠️ Do not read this as the port being wrong. Against the **capture**, `title` is at
|
||||
0.21 % and `title_plate` at 0.00093 %. This is two of our renderers disagreeing,
|
||||
and the one with an oracle behind it is not the one under suspicion.
|
||||
|
||||
@@ -6338,3 +6338,67 @@ claim about a tree, and a manifest for a tree that was never finished would be
|
||||
worse. So this is **filed rather than fixed**: the behaviour is defensible and the
|
||||
message is not, since "is that an export tree?" describes the symptom and hides
|
||||
the cause. What a stranger needs to be told is *the last export failed; re-run it*.
|
||||
|
||||
## `check-all`, a verdict that ignored its own statistic, and a claim of mine that was wrong
|
||||
|
||||
Eleven tools under `tools/port/` and **nothing ran them together**, so each had to
|
||||
be remembered individually. That is the ninth instance of this port's recurring
|
||||
shape — correct, documented, unexercised — one level up: the checks were the thing
|
||||
nobody was running.
|
||||
|
||||
`tools/port/check-all` runs the four that assert (`check`, `check-modding`,
|
||||
`check-capture-controls`, `verify-menu-audio`), prints the oracle table, and
|
||||
handles `verify-screen` specially. All eleven were exercised first and **none had
|
||||
rotted**; `which-focus` independently picks NEW_GAME at a **93.8× margin**, which
|
||||
is a second instrument agreeing with the capture fit's 10×.
|
||||
|
||||
Two things it is careful about:
|
||||
|
||||
* the six exploratory tools are **not** listed as passes. They produce artifacts
|
||||
for a person to look at and have no verdict; counting them would invent six.
|
||||
* `verify-capture` is **reported, not asserted** — it always exits 0. Its header
|
||||
is right that the numbers are not a target, but *not a target* is not *not a
|
||||
regression detector*, and nothing would notice `title_plate` moving off 0.00 %.
|
||||
Named as a gap rather than papered over; a real fix needs stored baselines, and
|
||||
what a baseline means when the pose is fitted is a decision, not a chore.
|
||||
* the `verify-screen` allowance **expires on its own condition**. It is allowed to
|
||||
fail only while `formats-pin-2026-08-29d` is not an ancestor of `origin/main`;
|
||||
the day it lands, `check-all` fails instead. A suppression with no expiry is
|
||||
just a hidden failure.
|
||||
|
||||
### 🔴 The verdict ignored the statistic added to inform it
|
||||
|
||||
`verify-screen` computes `over3` — how many pixels exceed the bar — because *"a
|
||||
single `max` cannot tell 2 pixels from 25 444"*, its own words. **The verdict was
|
||||
then decided on `max` alone.** So `main_menu` (max 4, `over3` **0**) read DIFFERS
|
||||
while `extras` (max 3, `over3` 0) read OK: one unit on one pixel separating two
|
||||
frames that are equivalent at the bar.
|
||||
|
||||
⚠️ Not fixed by raising the bar, which this file rightly forbids. The bar is still
|
||||
3. A frame with **no** pixel over it now gets its own verdict, `ROUNDING`, instead
|
||||
of being lumped in with a real disagreement. Tenth instance: the fix was
|
||||
implemented, documented, and never wired to the thing it was for.
|
||||
|
||||
### 🔴 And "six expected DIFFERS" — which I have told the Decoder more than once — was wrong
|
||||
|
||||
The true count was **ten**, now **eight** after the rounding fix:
|
||||
|
||||
| screens | count | explained |
|
||||
|---|---|---|
|
||||
| the forced-backdrop six | 6 | ✅ the pin: two decoder eras |
|
||||
| `main_menu`, `main_menu_jp` | 2 | ✅ now `ROUNDING`, not a disagreement |
|
||||
| **`title`, `title_jp`** | **2** | 🔴 **not explained** |
|
||||
|
||||
`title` differs on **790** pixels and `title_jp` on **20 498**, and neither is the
|
||||
forced-backdrop rule — those screens have no forced element. I had a blanket
|
||||
allowance covering two disagreements I had never accounted for.
|
||||
|
||||
**My hypothesis for them is refuted.** `authored/rendering.json` notes that the
|
||||
consistency harness compares against a renderer that draws no `.rat` leaves, so
|
||||
the port's `ptloop` sweeps looked like the obvious cause. Emptying `draw_leaf_for`
|
||||
and `loop_leaf_on_screens` changes the numbers **not at all** — 790 and 20 498
|
||||
either way. `verify-screen` poses at `rest`, where the leaves evidently do not
|
||||
draw. Filed as open.
|
||||
|
||||
⚠️ `title_jp` is a localisation screen and out of scope (MISSION §7). `title` is on
|
||||
the boot path and is not.
|
||||
|
||||
Reference in New Issue
Block a user