port: sweep instructions above descriptions -- the silent class is clean, two loud hits

Their sharpening: a stale instruction manufactures a false confirmation, strictly
worse than a stale description that merely misleads. Applied to my instruction
surface, the documented invocations in tool and script headers.

All fifteen distinct flags across those examples are parsed, so nothing in my
headers can produce their failure mode by being inert. But 'parsed' is a proxy and
its gap is known -- --shots parses and does nothing on the --boot path -- so I ran
two documented examples end to end rather than trusting the grep, and both produce
a 1280x720 frame.

Two hits, both loud rather than silent: 11 references to tools/verify-capture and
tools/verify-screen, paths that do not exist since the tools are under tools/port/
(fixed in 4 files); and check-all claiming eleven tools where there are fourteen
(now states both so the sentence dates itself).

The distinction worth recording: mine fail loudly, theirs failed silently. A wrong
path announces itself; an inert environment variable returns a clean wrong result.
Both are stale instructions and only one manufactures evidence.

Honest limit: I tested the flag surface plus two examples end to end, not all
thirteen documented invocations -- the --boot ones take 156 s each. That is a
judgement about cost, not a claim of coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
This commit is contained in:
Sylpheed port agent
2026-08-30 18:05:57 +00:00
parent e63f20468f
commit 66f0adce02
5 changed files with 56 additions and 13 deletions

View File

@@ -9,7 +9,7 @@ dies, which is what this file is for.
<!-- INDEX: generated by tools/port/index-decisions -- do not hand-edit -->
212 sections. Search this before re-deriving anything.
213 sections. Search this before re-deriving anything.
* [P0 — the exporter, 2026-08-28](#p0--the-exporter-2026-08-28)
* [P1 — Godot draws the screen, 2026-08-28](#p1--godot-draws-the-screen-2026-08-28)
@@ -223,6 +223,7 @@ dies, which is what this file is for.
* [Their sharpened tell, applied to my tree: two descriptions the code below had already refuted](#their-sharpened-tell-applied-to-my-tree-two-descriptions-the-code-below-had-already-refuted)
* [The grep found two more — and the reason is my correction *habit*, not my attention](#the-grep-found-two-more--and-the-reason-is-my-correction-habit-not-my-attention)
* [Auditing headings — and my own index was amplifying the withdrawn ones](#auditing-headings--and-my-own-index-was-amplifying-the-withdrawn-ones)
* [Ranking instructions above descriptions — swept, and the worst class is clean](#ranking-instructions-above-descriptions--swept-and-the-worst-class-is-clean)
<!-- /INDEX -->
## P0 — the exporter, 2026-08-28
@@ -397,7 +398,7 @@ human sees the screen — but it is not what the numbers come from.
## P1 gate — the diff, and what it found
`tools/verify-screen` renders every screen in the manifest both ways and reports
`tools/port/verify-screen` renders every screen in the manifest both ways and reports
the largest per-channel difference anywhere in the frame. Both renderers are held
to the same inputs: the reference CLI built by `build-reference-cli` from the
revision the exporter is **pinned** to (not `/reborn/target/`, which is a live
@@ -984,7 +985,7 @@ The correction comes from the human, via the RE agent, in their words: Reborn
various files. It may very well be wrong." **The oracle is the Xenia Canary
capture and the game.**
So `tools/verify-screen` is a **consistency check between two decoders that
So `tools/port/verify-screen` is a **consistency check between two decoders that
share their assumptions**, and a regression detector. It is not a correctness
check, and agreement in it is not evidence of correctness.
@@ -1008,7 +1009,7 @@ capture and catchable by nothing else:
### What changes
* `tools/verify-screen` says all of this in its own header, calls the CLI the
* `tools/port/verify-screen` says all of this in its own header, calls the CLI the
**comparison** renderer, and a `DIFFERS` row now means "we moved apart, find
out which of us moved" rather than "the port is wrong".
* The correctness question moves to the captures. The RE agent has committed
@@ -3889,7 +3890,7 @@ the experiment are the same operation rather than two implementations that agree
## The correctness harness the docs promised for eight milestones did not exist
`tools/port/verify-screen`, line 20, since P1: *"Use `tools/verify-capture` for
`tools/port/verify-screen`, line 20, since P1: *"Use `tools/port/verify-capture` for
the correctness question."* **There was no such file.** The port has had a harness
comparing itself to `sylpheed-cli` — two renderers sharing its assumptions — and
none comparing it to the game, while its own documentation said otherwise.
@@ -8756,7 +8757,7 @@ the wrong field name, and a fair metric pointed at the wrong frame.
### Wired so it cannot recur
`tools/verify-capture` takes a fifth per-row field, a capture crop, because this
`tools/port/verify-capture` takes a fifth per-row field, a capture crop, because this
capture is a full 1280×720 display frame with the surface at +0+45 while every
other capture in that directory is pre-cropped to 1279×675 — comparing it whole
would score the port against a 45 px shift. With it, `title_jp` reads **RMSE
@@ -11233,3 +11234,43 @@ was making three withdrawn claims findable first.
own framing: **"record" and "statement" want opposite orders, and a single block
cannot be both without deciding which one leads.** That is more precise than
calling the habit wrong — it isn't wrong, it is under-specified about ordering.
## Ranking instructions above descriptions — swept, and the worst class is clean
Their sharpening: **a stale instruction manufactures a false confirmation**, which
is strictly worse than a stale description that merely misleads. Their example is
a doc naming an environment variable removed with the record-layout fix — a reader
sets something inert, gets default behaviour, and concludes the two readings
agree. So: rank instructions above descriptions when sweeping.
Applied to my tree, the instruction surface is the documented invocations in the
tool and script headers. Fifteen distinct flags appear across them.
✅ **All fifteen are parsed** — no silently ignored flag, so nothing in my headers
can produce their failure mode by being inert.
⚠️ **But "parsed" is a proxy and I know its gap**: `--shots` parses and does
**nothing** on the `--boot` path, which I found two iterations ago. Parsing is not
working. So I ran two documented examples end to end rather than trusting the
grep — `--screen=main_menu --pose=rest --capture` and
`--screen=title --overlay=press_start --time=4` — and both produce a 1280×720
frame. (`--boot --shots` is not a documented combination, which is why the gap has
not bitten a reader.)
### Two hits, both of the *loud* kind
| | |
|---|---|
| **11 references** to `tools/verify-capture` / `tools/verify-screen` | those paths do not exist; the tools are under `tools/port/`. Fixed in 4 files. |
| `check-all`: *"There are **eleven** tools under `tools/port/`"* | there are **fourteen**. Now states both, so the sentence dates itself. |
📌 **The distinction worth recording: mine fail loudly, theirs failed silently.** A
wrong path errors out and announces itself; an inert environment variable returns
a clean, wrong result. **Both are stale instructions and only one manufactures
evidence.** That is the ranking their sharpening earns, and it means my two hits —
while real — are the cheap kind.
⚠️ And the honest limit on this sweep: I tested the **flag surface**, plus two
examples end to end. I did not run all thirteen documented invocations. The `--boot`
ones take 156 s each and I judged the flag-parse check plus two spot runs
sufficient; that is a judgement about cost, not a claim of coverage.