port: the game decodes all three voice streams at once, and two baseline rows were comparing blank frames

TWO FINDINGS, one mine and one handed to me, and the second retires a premise I
built on twice.

THE P1 BASELINE HAD ROWS THAT PROVED NOTHING. `build_12` and `build_15` render
pure black in BOTH renderers -- mean 0, max 0 -- so the difference is zero and
`verify-screen` scored them `max 0  over3 0  OK`, the strongest verdict it has.
Two of sixteen rows were comparing nothing against nothing. Worse than a missing
test, because a missing test is visible in the count.

Cause isolated by a control, not by reading: `build_00`/`build_01` are the same
loading screen minus three elements and render fine (mean 1.913, max 214.5). The
dressed variants add `pgloading_eff00`, a 1280x720 primitive resting OPAQUE BLACK
at t=38 inside its own opening black hold, with no layer key so paint order puts
it last.

The rule I was about to write -- "rest.t before the last timed keyframe is the
pathology" -- was killed by running the census first: 152 of 212 elements in this
export have rest.t earlier than their last timed keyframe. It is the norm. What
is actually unusual is the CONTENT, and its reach is one: `pgloading_eff00` is
the only element in the export whose resting pose is a fully opaque full-frame
quad. One instance is not a rule, so the renderer is unchanged and the HARNESS is
fixed: a blank pair now reports BLANK -- both renderers drew nothing; this row
proves nothing. `status` is untouched, so an unrelated DIFFERS still fails.

THE VOICE EXPORT IS KNOWN INCOMPLETE. The Decoder booted Canary with
--xma_param_probe and the game decodes ALL THREE streams CONCURRENTLY, in three
XMA contexts whose byte sizes match the disc payloads exactly. So "three
presentations of one take, pick one" is refuted by the running game and the
question I had been arguing -- WHICH presentation -- has no answer.

This one no census could have caught. Every measurement was right: the streams
are equal-duration, one is silence, one is 0.60x another with the residual 26.8
dB down. The frame around them was wrong, and the file says ChannelMask 0x0002 on
all three. It took the running game -- which is the mission's own sentence
arriving in practice.

BEHAVIOUR HELD DELIBERATELY. An equal-gain 1/n sum of channel pairs is not a
downmix either -- MISSION section 6 pins an explicit matrix for exactly that
reason -- and summing cost S00A 6.02 dB when one stream was silence. Swapping one
guess for another on a message is what produced this entry twice. What changed is
that the wrongness is now LOUD, because this failure sounds like success: one
stream is clean audible dialogue. A top-level manifest warning per movie, the
console line, and the authored entry all say `1 of 3 streams`.

"They are 5.1" is recorded as the Decoder's HYPOTHESIS with its own
counter-evidence attached, and nothing builds on it. What settles it is asked: a
recording of the game's own output over ADV through the null sink, which turns
channel roles into a fit against an oracle.

Refutation attempt, survived: the Decoder's loading-screen variant map. Entries
0/1 carry 7 elements and 12/15 carry those seven plus baseeff, eff00 and loop5 --
exact in count and identity, and it is what made build_00 a control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
This commit is contained in:
Sylpheed port agent
2026-08-29 15:50:02 +00:00
parent c43d44f57e
commit 8fba7944d4
6 changed files with 276 additions and 71 deletions

View File

@@ -133,11 +133,33 @@ print(json.load(open("export/"+f))["source"]["build"])' "$name")
-compose difference -composite -colorspace Gray -threshold $((3*65535/255)) \
-format "%[fx:int(mean*w*h)]" info:)
# 3/255 is what integer-truncating compositing in the CLI and float rounding
# in a GPU differ by. Anything above that is a placement, order or colour
# disagreement and needs a reason, not a threshold.
# BOTH FRAMES BLANK IS NOT AGREEMENT, AND THIS SCRIPT USED TO SAY IT WAS.
#
# `build_12` and `build_15` -- the two dressed loading screens -- render as
# pure black in BOTH renderers, mean 0 and max 0, so the difference is 0 and
# the row read `max 0 over3 0 OK`. Two of the sixteen rows in the committed
# baseline were comparing nothing against nothing and reporting the strongest
# verdict this script has.
#
# That is worse than a missing test: it is a test that reports a pass. The
# screens are black because `pgloading_eff00` is a full-frame opaque black
# quad whose `rest.t` (38) sits inside its own opening black hold, and
# `--pose=rest` freezes it there -- see docs/port/DECISIONS.md. Whether that
# is the port's bug or the decoders' reading of `rest` is open; what is not
# open is that a blank pair may not be scored.
#
# So blankness is checked FIRST and reported as its own verdict. It is not a
# failure -- the port may legitimately have nothing to draw -- but it is not a
# pass either, and `status` is left alone so an unrelated screen's DIFFERS is
# still what fails the run.
ink=$(convert "$OUT/$name.godot.png" "$OUT/$name.ref.png" \
-evaluate-sequence max -colorspace Gray -format "%[fx:maxima*255]" info:)
verdict=OK
awk "BEGIN{exit !($max > 3)}" && { verdict=DIFFERS; status=1; }
if awk "BEGIN{exit !($ink <= 0)}"; then
verdict="BLANK -- both renderers drew nothing; this row proves nothing"
else
awk "BEGIN{exit !($max > 3)}" && { verdict=DIFFERS; status=1; }
fi
printf '%-17s build %-3s max %-5s mean %-8s over3 %-7s %s\n' \
"$name" "$build" "$max" "${mean:0:6}" "$over" "$verdict"
done