Files
Sylpheed/docs/re/structures/intro-audio-decomposed.md
sylph-decoder 3e8235ddbd re: resolve_movie_voice_region starts INSIDE the first stream, 8 of 10 multichannel regions
Found because the port refused to apply my stream assignment and did the
arithmetic instead: the running decoder's three ADV contexts sum to 3 584 000 B
against a resolved region of 3 114 352 -- 15 % too small to hold them. Two spans,
one wrong, and it was the disc side.

The gap is 238 packets exactly (487 424 B), which is what a start offset looks
like; ctx0 declares 632 packets and the resolver's leading chunk has 394.

Verified against the decoder's own byte_sizes, which cannot be fitted to: at
-238 packets to_xma_riffs yields [1294336, 1118208, 1171456], all three exactly.
It is a real boundary and not the end of a sweep -- at -300 the previous asset's
chunks appear while the three ADV sizes stay stable.

Disc-wide: 24 of 24 single-chunk regions start at a boundary; 8 of 10 three-chunk
regions start mid-stream. The defect is specific to the multichannel case.

The audit's per-movie number is an UPPER BOUND, not the clip -- its stopping rule
is the chunk count changing, and to_xma_riffs absorbs a few packets of the
previous asset first (243 reported for ADV against a true 238). Only ADV has
external ground truth.

Consequence: in those 8 movies the leading chunk is a truncated first stream, not
a spurious artefact, and anything measured on it was measured on a fragment --
including this corpus's own chunk-0 level, though the assignment survives because
its ratio test was chosen to be immune to the clipping.

The resolver is NOT patched. Why the predecessor cue's trailer lands 238 packets
into the next asset is unanswered, and a fix guessed from one movie would be
worse than a documented defect.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
2026-08-30 08:41:22 +00:00

164 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ✅ The boot intro's audio: the movie's own 5.1 bed at 0.600, **plus** three streams in 5.1
**Classification: measured.** Xenia Canary, 2026-08-30, one boot, no input. This
settles the port's ask #4 and **confirms** the corpus's leading hypothesis from the
output side, where it had been recorded as "not established".
## 🔴 First: `ADV.wmv` is not three XMA streams. It is one WMA Pro 5.1 track.
```
ffprobe /disc/dat/movie/ADV.wmv
Stream #0:0(jpn): Audio: wmapro, 48000 Hz, 5.1, fltp, 384 kb/s
Stream #0:1(jpn): Video: wmv3, 1280x720, 30 fps
```
One audio stream, **5.1**, decoded by Xenia's WMA path — not the XMA path the three
probe contexts come from. Any framing of the intro's audio as *only* "which of three
voice streams to ship" was missing the bed entirely.
## The decomposition
Aligning the [148 s capture](intro-audio-output-census.md) against that track and
solving `capture = g × movie + residual`
([`../data/intro-audio-decomposition.txt`](../data/intro-audio-decomposition.txt)):
| ch | gain | r | capture rms | residual rms | residual/capture |
|---|---|---|---|---|---|
| FL | 0.600 | +0.905 | −21.86 | −29.28 | −7.43 dB |
| FR | 0.600 | +0.935 | −20.28 | −29.29 | −9.02 dB |
| **FC** | 0.597 | +0.146 | −25.18 | −25.27 | **−0.09 dB** |
| **LFE** | 0.600 | **+1.000** | −43.66 | **−115.73** | **−72.06 dB** |
| BL | 0.600 | +0.956 | −24.89 | −35.46 | −10.58 dB |
| BR | 0.600 | +0.964 | −24.02 | −35.49 | −11.47 dB |
**The gain is 0.600 on every channel** — a uniform −4.44 dB, which is a mixer
setting, not a fit artefact. **LFE is reproduced to −115.73 dBFS**, 72 dB below the
signal: at that residual the two decoders agree essentially exactly, which is what
rules out "the leftovers are just codec differences".
🔴 **And `FC` is the exception that carries the answer.** The movie explains
**nothing** of the capture's centre channel — the residual is the whole signal
(−0.09 dB). The movie's own FC is 91.6 % silent; the capture's is not.
## The residual is three signals, not one
| | FL | FR | FC | LFE | BL | BR |
|---|---|---|---|---|---|---|
| FL | 1.000 | **0.918** | 0.017 | 0.001 | 0.009 | 0.004 |
| FR | **0.918** | 1.000 | 0.034 | 0.001 | 0.015 | 0.014 |
| FC | 0.017 | 0.034 | 1.000 | 0.000 | 0.385 | 0.378 |
| LFE | 0.001 | 0.001 | 0.000 | 1.000 | 0.000 | −0.000 |
| BL | 0.009 | 0.015 | 0.385 | 0.000 | 1.000 | **0.929** |
| BR | 0.004 | 0.014 | 0.378 | −0.000 | **0.929** | 1.000 |
Three coherent groups: a **front pair** (0.918), a **rear pair** (0.929), and a
**centre** whose partner LFE is empty. The FC residual's 100 ms frame levels span
**34 dB** (median −53.9, p90 −19.9) — bursty, not steady noise.
## ✅ This confirms the 5.1 hypothesis, and predicts the silent channel correctly
[`voice-three-streams-are-concurrent.md`](voice-three-streams-are-concurrent.md)
recorded "three concurrent stereo streams is six channels" as the obvious reading
and marked it **not established**, with a specific piece of supporting detail: that
`ADV` stream 2 is **mono-in-stereo**, *"a centre paired with a silent LFE looks
exactly like that"*.
That is exactly what the residual shows — a live centre whose paired channel is
empty to −115 dB. Measured from the output, with no access to the stream contents:
| XMA stream | lands in |
|---|---|
| one | **FL, FR** |
| one | **FC**, LFE silent |
| one | **BL, BR** |
**So both things are true and the port needs both**: the movie's own 5.1 WMA Pro
track *and* the three streams mixed over it in 5.1.
## 🔴 Correction to the census page, and to the recipe page's channel order
[`intro-audio-output-census.md`](intro-audio-output-census.md) labelled its channels
using the permutation `[0,1,4,5,2,3]` that
[`audio-capture-alsa-file-tee.md`](../audio-capture-alsa-file-tee.md) records for
ALSA. **That permutation does not apply to this capture.** The 6×6 correlation
matrix was computed without assuming any order, every row's maximum falls on a
distinct movie channel, and the result is the **identity**.
So the census's *"BR is 82 % silent"* was really **LFE** — which also reconciles it
with the movie, whose LFE is 80.64 % silent. ⚠️ The recipe page's permutation was
measured on a different chain and is not wrong there; what is wrong is assuming it
travels. **Measure the channel order per capture; a 6×6 matrix that comes out a
clean permutation is its own control.**
## Reach
⚠️ **One boot, one movie.** `ADV` only.
✅ **The assignment is now determined** — see the section below. It was open when
this page was first written.
⚠️ `--gpu=null`, so no video cross-check.
✅ The 0.600 gain is measured on this run; whether it is a fixed mix constant or a
volume setting is not established.
## ✅ Which stream is which (2026-08-30, later)
The three chunks were dumped from the resolved voice region
(`examples/adv_voice_dump.rs`) and decoded:
[`../data/adv-stream-assignment.txt`](../data/adv-stream-assignment.txt).
| chunk | `byte_size` | probe ctx | L rms | R rms | R silent |
|---|---|---|---|---|---|
| 0 | 806 912 | **ctx0, clipped tail** (full 1 294 336) | −24.79 | −24.81 | 53.1 % |
| 1 | 1 118 208 | ctx1 | −20.33 | **−inf** | **100 %** |
| 2 | 1 171 456 | ctx2 | −30.67 | −30.68 | 53.6 % |
### 🔴 Two instruments failed first, and both look convincing
* **Envelope correlation cannot discriminate.** A per-pair lag search returns
**0.86–0.95 for every chunk against every channel**, because all six residual
channels share the dialogue's activity timing. A number that high reads as a
result; it is the instrument having no resolving power. Recorded so nobody
reports it as one.
* **Sample-level correlation returns ≈ 0.** The chunks do not start with the movie
and the XMA decode's framing offset is unknown.
### ✅ Level settles it, under the same 0.600 gain
| chunk | level | × 0.600 | nearest residuals (error, dB) |
|---|---|---|---|
| 0L | −24.79 | −29.23 | **FL 0.05** · FR 0.06 · FC 3.96 |
| 1L | −20.33 | −24.77 | **FC 0.50** · FL 4.51 |
| 2L | −30.67 | −35.11 | **BL 0.35** · BR 0.38 · FR 5.82 |
Each stream lands within **0.5 dB** of exactly one residual pair and misses the
others by ~4–6 dB. **The same 0.600 that scales the movie bed also scales the
voice** — which is itself worth having: it is one mixer gain, not two.
✅ **Ratio test, immune to chunk 0 being clipped:** chunk0 − chunk2 = **+5.88 dB**
against FL − BL = **+6.18 dB**, agreeing to **0.30 dB**; swapped, the ratio would be
wrong by **11.76 dB**.
✅ **Structural confirmation.** Chunk 1 is the *only* chunk with a digitally silent
channel, and LFE is the *only* output channel with an empty residual (−115.73 dBFS).
One to one. And the internal L/R correlations track: chunk 0 **+0.932** against the
FL/FR residual's **+0.918**, chunk 2 **+0.962** against BL/BR's **+0.929**.
| stream | → |
|---|---|
| ctx0 · 1 294 336 | **FL, FR** |
| ctx1 · 1 118 208 | **FC** (LFE silent) |
| ctx2 · 1 171 456 | **BL, BR** |
🔴 **And the identifier this is indexed by does not resolve against the disc.** The
port refused to apply this assignment because the three contexts sum to 3 584 000 B
against a resolved region of 3 114 352. It was right to: the **region start is
wrong**, by 238 packets for `ADV` and on 8 of 10 multichannel regions disc-wide —
[`voice-region-starts-late.md`](voice-region-starts-late.md). The assignment above
still stands (the ratio test was chosen to be immune to the clipping), but chunk 0's
absolute level was measured over 62 % of its stream.
⚠️ **Reach.** Levels, not waveforms — this is an argument from three numbers
agreeing to 0.5 dB and a 1:1 structural match, not from a matched waveform. One
movie, one boot. And chunk 0 is a clipped tail, which is why the ratio test is
quoted alongside the absolute match.