port: withdraw the dual-mono generalisation -- the measurement stands, the rule does not

I argued `highest_rate` had no case because ADV's higher-rate presentation is
dual-mono while its louder one is mono-in-stereo, so the extra bytes buy a
duplicated channel rather than fidelity. The Decoder tested that disc-wide over
the 28 three-stream cues: the stream-3/stream-2 size ratio runs min 0.0778,
median 1.2565, max 2.9163, sd 0.5057, with only 12 of 28 within 15% of 1.0, and
declared rates scatter with them. A 37x spread is not a duplicated channel.

The CHANNEL MEASUREMENT STANDS -- ADV chunk 1 is mono-in-stereo and chunk 2 is
dual-mono at -8.318574, this port's own decode, which the Decoder could not
re-run and did not dispute. What fails is the step from one asset to the format.

NOTHING IN THE EXPORT CHANGES. `loudest` is a per-asset content rule -- it reads
the peak of the streams in front of it -- so a scattering structural ratio cannot
undermine it. What changes is the REASON, in four places: authored/audio.json's
presentation_why, the selector comment in audio.rs, BLOCKED.md's row, and
DECISIONS.md. The honest statement is narrower: `highest_rate` was never refuted,
it was never argued for, and neither is `loudest`. That is why the entry is
marked CHOSEN rather than measured, and why one capture deletes it.

Recorded on the pattern rather than just the instance: this is the third claim of
mine in two iterations that generalised a single-asset observation, after "the
chunks are two stems" and "everything the sequencer paces off rest.t is late".
All three were true of the thing I looked at. The failure is reaching for the
rule a measurement would imply if it held everywhere and writing that down in the
same breath as the measurement.

Also noted, not mine and not affecting export_voice: S12B's three streams are
byte-size identical, and BIRD_224 is three-stream while not being a movie cue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
This commit is contained in:
Sylpheed port agent
2026-08-29 15:39:22 +00:00
parent 81ea5cb324
commit 92f1436836
4 changed files with 80 additions and 18 deletions

View File

@@ -135,19 +135,26 @@
"different streams: ADV chunk 1 is 1118268 B at 0.0 dBFS, chunk 2 is",
"1171516 B at -8.3 dBFS.",
"",
"The reason for switching is a measurement of this port's, not a preference.",
"The case for `highest_rate` was that more bytes per second means a better",
"encode. It does not, here: ADV chunk 1 is MONO-IN-STEREO (channel 2",
"digitally silent) and chunk 2 is DUAL-MONO (both channels identical at",
"-8.318574). So chunk 2's extra bytes go into encoding a duplicate of its own",
"channel, not into fidelity, and the byte-rate difference is explained without",
"any appeal to quality. That removes the only argument for it.",
"WHY `loudest` AND NOT `highest_rate`: this is a PER-ASSET CONTENT choice,",
"and that is the whole of its justification. The disc masters its other audio",
"near full scale -- the SE cues decode to +0.18 dBFS -- and S00A's only",
"surviving full-length stream is its louder one at -4.2 dBFS, so `loudest`",
"makes the two cutscenes' dialogue sit at comparable levels instead of 4.4 dB",
"apart. Neither half is strong alone; together they are what there is.",
"",
"What is left points the other way, and both parts are weak on their own:",
"the disc masters its other audio near full scale (the SE cues decode to",
"+0.18 dBFS), and S00A's only surviving full-length stream is its louder one",
"at -4.2 dBFS -- so `loudest` makes the two cutscenes' dialogue sit at",
"comparable levels instead of 4.4 dB apart.",
"⚠️ A STRUCTURAL ARGUMENT WAS OFFERED HERE AND IS WITHDRAWN. It said ADV",
"chunk 1 is mono-in-stereo (channel 2 digitally silent) while chunk 2 is",
"dual-mono (channels identical at -8.318574), therefore the extra bytes",
"encode a duplicated channel rather than fidelity, therefore the rate",
"argument collapses. The CHANNEL MEASUREMENT stands -- it is ADV's, and it is",
"this port's own. THE GENERALISATION DOES NOT. The Decoder tested it",
"disc-wide over the 28 three-stream cues: the stream-3/stream-2 size ratio",
"runs min 0.0778, median 1.2565, max 2.9163, sd 0.5057, with only 12 of 28",
"within 15% of 1.0, and declared rates scatter with them (S06A: 5661 against",
"16513 B/s). A 37x spread is not a duplicated channel.",
"",
"So `highest_rate` was not refuted as a rule; it was simply never argued for,",
"and neither was this. That is why the entry is CHOSEN and why it says so.",
"",
"STILL A CHOICE. One capture of the intro with dialogue audible settles it,",
"and it is the last unforced decision in the voice pipeline."

View File

@@ -623,11 +623,17 @@ pub fn export_voice<S: DiscSource + ?Sized>(
// WHICH of the equal-duration survivors is a CHOICE, and it lives in
// `authored/audio.json` rather than here -- see [`Presentation`]. It was
// `highest_rate` on the Decoder's recommendation until that was withdrawn as
// self-contradictory, and the reason it is now `loudest` is a measurement:
// `ADV`'s louder presentation is mono-in-stereo while its higher-rate one is
// DUAL-MONO, so the extra bytes encode a duplicate channel rather than
// fidelity, and the rate difference is explained without appealing to
// quality at all.
// self-contradictory. `loudest` is a PER-ASSET CONTENT choice and nothing
// more: the disc masters its other audio near full scale, and it puts the
// two cutscenes' dialogue at comparable levels.
//
// ⚠️ A structural argument for it — `ADV`'s higher-rate stream is dual-mono,
// so its extra bytes encode a duplicated channel rather than fidelity — was
// offered here and **does not generalise**. The channel measurement is
// `ADV`'s and stands; the inference was tested disc-wide over the 28
// three-stream cues and the size ratio runs 0.0778 to 2.9163. Neither rule
// has a structural argument behind it, which is exactly why the choice is
// authored rather than derived.
let tied: Vec<usize> = (0..all.len())
.filter(|&i| !silent.contains(&i) && (longest - lengths[i]).abs() < 0.001)
.collect();

View File

@@ -114,7 +114,7 @@ HANDOFF.
| P6 BGM — the sub-wave count | **is a music bank's LEADING REGION a stem, or a decoder artefact?** | Q10 | 🔴 **HANDOFF and the decoders disagree, and P6 ships the disagreement.** `media::sound_bank_riffs("BGM_103.slb")` returns **three** sub-waves; HANDOFF Q10's census says a music bank is *"exactly two waves of identical duration (32/32 banks on the disc)"*. The third comes from `slb.rs:380` `to_xma_riffs`, whose hybrid branch emits a leading headerless packet region ahead of the `RIFF` waves — and `docs/re/REFUTED.md` already records that region as what makes `BGM_106``BGM_109` *"break the two-wave rule"*. Derived at HANDOFF `9ca1eb5`. **The exporter sums all three and writes a manifest warning**, because choosing which sub-wave to drop is a decoding question and MISSION §2 forbids this exporter answering one. So the menu currently plays a sum of three things where the census predicts two. What settles it: whether that leading region carries music. Raised with the Decoder 2026-08-29. |
| ~~P3 — the plate's ONSET~~ | ~~visible 2.13 s after settle, or group starts then?~~ | Q2 | ✅ **resolved 2026-08-29, and the answer is AUTHOR NOTHING.** The port's refutation held and produced a better answer than either option it offered. Correction at `5b0a6e6` on `auto/no-disc-and-menu-captures`: **both builds run on one clock, started together**, and the plate arrives at its own declared `t=238`. Checked against this export rather than taken on trust — build 4's visible build-in ends at `t=118` (`pteff01`, `pteff02`, `ptlogoall_eff` finish together), `ptbtn00` reaches alpha 255 at `t=238`, difference **120 units = 2.000 s**, against a measured 2.138 / 2.132 s at an emulator presenting 28.1 fps rather than 30. The 2.13 s constant is **deleted**. |
| ~~P3/P5 — `settle_time()`~~ | ~~`rest.t` is not when a screen settles, and the port's sequencer uses it~~ | — | ✅ **MEASURED 2026-08-29 and the row was HALF WRONG — mine.** The Decoder took it on a cold profile with no shader cache (`auto/no-disc-and-menu-captures` at `4bd4779`, `docs/re/boot-settle-times-measured.md`). The principle holds: the title's `rest.t` is 251 units = **4.183 s** where its art finishes at ~2 s. **But "everything the sequencer paces off that landmark is therefore late" does not.** Measured the port the way the game was measured — visible span, `--film` at 4 fps — the publisher wordmark runs **4.25 s** against the game's 4.297/4.604/4.370 and the developer logos **3.50 s** against 3.508/3.503/3.366. Dead on. My earlier reading compared the port's *arrival-to-arrival* timestamps against the game's *visible spans*, which differ by the exit ramp plus the black hold — the whole of the discrepancy I was about to chase. `rest.t` is still the wrong landmark; its blast radius is `_script_settled` waiting longer than it needs to, which is a slow test and not a wrong frame. `dwell_seconds` stays `null`, now for a measured reason. 🔴 **Do not author an Ⓐ→menu dwell**: it measures 3.763 s and contains a 1.53 s guest load stall, third independent reproduction. 🟡 Menu build-in 0.531 s and Ⓑ→title 0.482 s rest on one run and are not authored; the port is within ~0.1 s of both from the disc. |
| P4/P7 — which voice presentation | **which of a region's two full-length streams does the game play?** | — | 🟡 **the last unforced decision in the voice pipeline, and it is now unambiguously the port's.** The Decoder's "highest byte rate" was **withdrawn as self-contradictory** — its sentence read *"the highest-rate, highest-gain one is chunk 1"*, and those select different streams (`ADV` chunk 1: 1 118 268 B at 0.0 dBFS; chunk 2: 1 171 516 B at 8.3). Nothing on the disc ranks them: `wEncodeOptions`, channel count and channel mask are byte-identical. Moved to `authored/audio.json` `voice.presentation` per MISSION §3, set to **`loudest`**, and the reason is a measurement of mine rather than a preference: chunk 1 is **mono-in-stereo** and chunk 2 is **dual-mono**, so chunk 2's extra bytes encode a duplicate channel rather than fidelity — which explains the rate difference and removes the only argument for it. What settles it: **one capture of the intro with dialogue audible.** That did not ride the settle-time boot, which drove the title path and never played the movie with audio. |
| P4/P7 — which voice presentation | **which of a region's two full-length streams does the game play?** | — | 🟡 **the last unforced decision in the voice pipeline, and it is now unambiguously the port's.** The Decoder's "highest byte rate" was **withdrawn as self-contradictory** — its sentence read *"the highest-rate, highest-gain one is chunk 1"*, and those select different streams (`ADV` chunk 1: 1 118 268 B at 0.0 dBFS; chunk 2: 1 171 516 B at 8.3). Nothing on the disc ranks them: `wEncodeOptions`, channel count and channel mask are byte-identical. Moved to `authored/audio.json` `voice.presentation` per MISSION §3, set to **`loudest`** — a **per-asset content** choice and nothing more: the disc masters its other audio near full scale (the SE cues decode to +0.18 dBFS), and it puts the two cutscenes' dialogue at comparable levels instead of 4.4 dB apart. ⚠️ **A structural argument for it was offered and is withdrawn.** I said `ADV`'s higher-rate stream is dual-mono where the louder is mono-in-stereo, so its extra bytes encode a duplicated channel rather than fidelity. The **channel measurement stands** — it is `ADV`'s and it is mine — but the Decoder tested the *inference* disc-wide over the 28 three-stream cues and the stream-3/stream-2 size ratio runs **min 0.0778, median 1.2565, max 2.9163, sd 0.5057**, only 12 of 28 within 15 % of 1.0. A 37× spread is not a duplicated channel. So `highest_rate` was never *refuted*, it was merely never argued for — and neither is `loudest`. That is precisely why the entry is marked CHOSEN. What settles it: **one capture of the intro with dialogue audible.** That did not ride the settle-time boot, which drove the title path and never played the movie with audio. |
| ~~P3/P5 — the title plate~~ | ~~does the idle title show `PRESS Ⓐ`~~ | Q2 | ✅ **answered and TAKEN at this iteration.** `auto/no-disc-and-menu-captures` at `fb536df`, `docs/re/title-plate-delay-measured.md`, traces in `docs/re/data/plate-timing-run{1,2}.tsv`. It is the third case: build 4 alone, then the plate composited over it. ⚠️ The delay is timed from where build 4 **stops animating**, not from where it first appears — measured the other way the two runs differ by 0.48 s against 6 ms. `ScreenView` now draws two builds at once, as a second `ScreenView` in the same `SubViewport` rather than a subordinate screen inside one. The onset question above is what is left. |
| P3 — the plate's PULSE | **does the plate's focus record loop, and with what period?** | Q2 | ❔ **open, and the port's earlier reading of it was wrong.** The port had looked for the pulse in `ptbtn00`'s own group; `5b0a6e6` identifies it as the plate's **focus record** `ptbtn00f` — a glow ramping alpha `0x00``0x50` and back, t=6…105. Measured on the running game at 2.12 / 2.19 / 2.34 / 2.31 s, mean **2.24 s**. 🟡 **The port has not taken it.** Looping that record needs a period, and its group is 105 timed units plus the **authored** 24-unit exit ramp = 129 units = 2.15 s — composing an authored constant with a loop assumption to land on a measured number is tuning, not measuring. Separately: the port draws no focus record on `press_start` at all, because the screen has no `buttons` and nothing is focused, so *whether the game always draws it* is its own question. |

View File

@@ -2922,3 +2922,52 @@ not to be trusted on these.** Its "16 channels / 4310 Hz / 2-bit" is
offsets — its XMA1 reader is misaligned. That is a tool in this repository
reporting confident nonsense, and it is the second time a renderer or reader of
ours has been believed before it was checked.
## Refutation of my dual-mono inference — the measurement stands, the generalisation does not
I argued that `highest_rate` had no case because `ADV`'s higher-rate presentation
is **dual-mono** while its louder one is mono-in-stereo, so the extra bytes buy a
duplicated channel rather than fidelity. The Decoder tested that disc-wide, as a
refutation attempt, and **it fails**.
Over the 28 three-stream cues, the stream-3 / stream-2 size ratio runs:
| min | median | max | sd | within 15 % of 1.0 |
|---|---|---|---|---|
| 0.0778 | 1.2565 | 2.9163 | 0.5057 | **12 of 28** |
Declared rates scatter with them — `S06A` is 5 661 against 16 513 B/s. **A 37×
spread is not a duplicated channel.**
**The channel measurement itself stands**: `ADV` chunk 1 really is mono-in-stereo
and chunk 2 really is dual-mono at 8.318574, and that is this port's own decode,
which the Decoder could not re-run and did not dispute. What fails is the step
from *one asset* to *the format*.
### What this changes, and what it does not
Nothing in the export changes. `loudest` is a **per-asset content** rule — it
reads the peak of the actual streams in front of it — so a scattering structural
ratio cannot undermine it, and `ADV`'s dialogue at +0.3 dBFS instead of 8.7 is
plainly the better outcome either way.
What changes is the *reason*, in four places: `authored/audio.json`'s
`presentation_why`, the selector comment in `audio.rs`, `BLOCKED.md`'s row, and
this page. The honest statement is narrower and slightly less satisfying:
**`highest_rate` was never refuted — it was never argued for, and neither is
`loudest`.** Which is exactly why the entry is marked *chosen* rather than
*measured*, and why one capture deletes it.
⚠️ **This is the third claim of mine in two iterations that generalised a
single-asset observation** — after "the chunks are two stems" and "everything the
sequencer paces off `rest.t` is late". All three were true of the thing I looked
at. The pattern is not carelessness about the measurement; it is reaching for the
rule the measurement would imply if it held everywhere, and writing that down in
the same breath. The corpus catches it because someone else runs the census.
### Two things in that data that are not mine, recorded so they are not lost
* **`S12B`'s three streams are byte-size identical** (14 396 each).
* **`BIRD_224` is three-stream and is not a movie cue** — so the three-stream
shape is not exclusive to cutscenes, which narrows how it was described to this
port earlier. Neither affects `export_voice`, which only resolves movies.