re: which ADV stream sits where -- settled by level, not by waveform

Completes ask #4. The three chunks were dumped from the resolved voice region
and decoded; the assignment is ctx0 -> FL/FR, ctx1 -> FC with LFE silent,
ctx2 -> BL/BR.

Two instruments failed first and both look like results, so both are recorded.
Envelope correlation with a per-pair lag search returns 0.86-0.95 for EVERY
chunk against EVERY channel, because all six residual channels share the
dialogue's activity timing -- that is an instrument with no resolving power, not
a finding. Sample-level correlation returns about zero, because the chunks do
not start with the movie and the XMA decode's framing offset is unknown.

Level settles it under the same 0.600 gain the bed uses: each stream lands
within 0.5 dB of exactly one residual pair and misses the others by 4-6 dB. The
ratio test is immune to chunk 0 being a clipped tail of ctx0 -- chunk0 - chunk2
is +5.88 dB against FL - BL at +6.18 dB, agreeing to 0.30 dB, where a swap would
be wrong by 11.76 dB.

Structural confirmation: chunk 1 is the only chunk with a digitally silent
channel and LFE is the only output channel with an empty residual (-115.73
dBFS), one to one; and the internal L/R correlations track the residual pairs'
(0.932 vs 0.918, 0.962 vs 0.929).

Worth having on its own: the same 0.600 scales both the movie bed and the voice,
so it is one mixer gain rather than two.

Reach: levels, not waveforms; one boot, one movie; and whether 0.600 is a fixed
constant or a volume setting is still unknown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
This commit is contained in:
sylph-decoder
2026-08-30 08:32:14 +00:00
parent bcb822d4b9
commit edab16d7ca
4 changed files with 169 additions and 6 deletions

View File

@@ -93,9 +93,63 @@ clean permutation is its own control.**
## Reach
⚠️ **One boot, one movie.** `ADV` only.
⚠️ **The stream→channel assignment is by position, not by content.** Which of the
three XMA contexts is the front, centre or rear stream is *not* determined here —
that needs the streams decoded and correlated individually, which was not done.
**The assignment is now determined** — see the section below. It was open when
this page was first written.
⚠️ `--gpu=null`, so no video cross-check.
✅ The 0.600 gain is measured on this run; whether it is a fixed mix constant or a
volume setting is not established.
## ✅ Which stream is which (2026-08-30, later)
The three chunks were dumped from the resolved voice region
(`examples/adv_voice_dump.rs`) and decoded:
[`../data/adv-stream-assignment.txt`](../data/adv-stream-assignment.txt).
| chunk | `byte_size` | probe ctx | L rms | R rms | R silent |
|---|---|---|---|---|---|
| 0 | 806 912 | **ctx0, clipped tail** (full 1 294 336) | 24.79 | 24.81 | 53.1 % |
| 1 | 1 118 208 | ctx1 | 20.33 | **inf** | **100 %** |
| 2 | 1 171 456 | ctx2 | 30.67 | 30.68 | 53.6 % |
### 🔴 Two instruments failed first, and both look convincing
* **Envelope correlation cannot discriminate.** A per-pair lag search returns
**0.860.95 for every chunk against every channel**, because all six residual
channels share the dialogue's activity timing. A number that high reads as a
result; it is the instrument having no resolving power. Recorded so nobody
reports it as one.
* **Sample-level correlation returns ≈ 0.** The chunks do not start with the movie
and the XMA decode's framing offset is unknown.
### ✅ Level settles it, under the same 0.600 gain
| chunk | level | × 0.600 | nearest residuals (error, dB) |
|---|---|---|---|
| 0L | 24.79 | 29.23 | **FL 0.05** · FR 0.06 · FC 3.96 |
| 1L | 20.33 | 24.77 | **FC 0.50** · FL 4.51 |
| 2L | 30.67 | 35.11 | **BL 0.35** · BR 0.38 · FR 5.82 |
Each stream lands within **0.5 dB** of exactly one residual pair and misses the
others by ~46 dB. **The same 0.600 that scales the movie bed also scales the
voice** — which is itself worth having: it is one mixer gain, not two.
**Ratio test, immune to chunk 0 being clipped:** chunk0 chunk2 = **+5.88 dB**
against FL BL = **+6.18 dB**, agreeing to **0.30 dB**; swapped, the ratio would be
wrong by **11.76 dB**.
**Structural confirmation.** Chunk 1 is the *only* chunk with a digitally silent
channel, and LFE is the *only* output channel with an empty residual (115.73 dBFS).
One to one. And the internal L/R correlations track: chunk 0 **+0.932** against the
FL/FR residual's **+0.918**, chunk 2 **+0.962** against BL/BR's **+0.929**.
| stream | → |
|---|---|
| ctx0 · 1 294 336 | **FL, FR** |
| ctx1 · 1 118 208 | **FC** (LFE silent) |
| ctx2 · 1 171 456 | **BL, BR** |
⚠️ **Reach.** Levels, not waveforms — this is an argument from three numbers
agreeing to 0.5 dB and a 1:1 structural match, not from a matched waveform. One
movie, one boot. And chunk 0 is a clipped tail, which is why the ratio test is
quoted alongside the absolute match.