Completes ask #4. The three chunks were dumped from the resolved voice region and decoded; the assignment is ctx0 -> FL/FR, ctx1 -> FC with LFE silent, ctx2 -> BL/BR. Two instruments failed first and both look like results, so both are recorded. Envelope correlation with a per-pair lag search returns 0.86-0.95 for EVERY chunk against EVERY channel, because all six residual channels share the dialogue's activity timing -- that is an instrument with no resolving power, not a finding. Sample-level correlation returns about zero, because the chunks do not start with the movie and the XMA decode's framing offset is unknown. Level settles it under the same 0.600 gain the bed uses: each stream lands within 0.5 dB of exactly one residual pair and misses the others by 4-6 dB. The ratio test is immune to chunk 0 being a clipped tail of ctx0 -- chunk0 - chunk2 is +5.88 dB against FL - BL at +6.18 dB, agreeing to 0.30 dB, where a swap would be wrong by 11.76 dB. Structural confirmation: chunk 1 is the only chunk with a digitally silent channel and LFE is the only output channel with an empty residual (-115.73 dBFS), one to one; and the internal L/R correlations track the residual pairs' (0.932 vs 0.918, 0.962 vs 0.929). Worth having on its own: the same 0.600 scales both the movie bed and the voice, so it is one mixer gain rather than two. Reach: levels, not waveforms; one boot, one movie; and whether 0.600 is a fixed constant or a volume setting is still unknown. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
47 lines
2.1 KiB
Plaintext
47 lines
2.1 KiB
Plaintext
# Which ADV voice stream sits where -- the assignment, and how it was settled.
|
|
#
|
|
# 2026-08-30. Chunks dumped by examples/adv_voice_dump.rs from the resolved
|
|
# movie voice region (dat/sound 433930240..437044592), decoded with ffmpeg's
|
|
# xma decoder to f32le 48 kHz stereo.
|
|
#
|
|
# chunk bytes byte_size probe ctx dur L rms R rms L/R r R silent
|
|
# 0 806972 806912 ctx0 TAIL 84.55s -24.79 -24.81 +0.932 53.1%
|
|
# 1 1118268 1118208 ctx1 137.32s -20.33 -inf +0.000 100.0%
|
|
# 2 1171516 1171456 ctx2 137.32s -30.67 -30.68 +0.962 53.6%
|
|
#
|
|
# ctx0's full byte_size is 1294336; the region resolver starts at the
|
|
# predecessor cue's trailer, so chunk 0 is its clipped tail (62 %).
|
|
#
|
|
# 🔴 WHAT DID NOT WORK: envelope correlation cannot discriminate.
|
|
# Every residual channel shares the dialogue's activity timing, so a
|
|
# per-pair lag search returns 0.86-0.95 for EVERY chunk against EVERY
|
|
# channel. Recorded because it looks like a strong result and is not.
|
|
#
|
|
# 🔴 Sample-level correlation also fails: the chunks do not start with the
|
|
# movie and the XMA decode's framing offset is unknown, so r ~ 0.
|
|
#
|
|
# ✅ WHAT SETTLES IT: level, under the SAME 0.600 gain the bed uses.
|
|
#
|
|
# chunk level x0.600 nearest residuals (|error| dB)
|
|
# 0L -24.79 -29.23 FL 0.05 FR 0.06 FC 3.96
|
|
# 0R -24.81 -29.25 FL 0.03 FR 0.04 FC 3.98
|
|
# 1L -20.33 -24.77 FC 0.50 FL 4.51 FR 4.52
|
|
# 2L -30.67 -35.11 BL 0.35 BR 0.38 FR 5.82
|
|
# 2R -30.68 -35.12 BL 0.34 BR 0.37 FR 5.83
|
|
#
|
|
# Ratio test, immune to any worry about chunk 0 being clipped:
|
|
# chunk0L - chunk2L = +5.88 dB ; FL - BL = +6.18 dB -> agree to 0.30 dB
|
|
# swapped, the ratio would be wrong by 11.76 dB
|
|
#
|
|
# Structural confirmation: chunk 1 is the ONLY chunk with a digitally silent
|
|
# channel (R, 100 %), and LFE is the ONLY output channel with an empty
|
|
# residual (-115.73 dBFS). One-to-one.
|
|
#
|
|
# Internal L/R correlation also tracks:
|
|
# chunk0 +0.932 <-> FL/FR residual +0.918
|
|
# chunk2 +0.962 <-> BL/BR residual +0.929
|
|
#
|
|
# ==> ctx0 (1294336) -> FL, FR
|
|
# ==> ctx1 (1118208) -> FC, LFE silent
|
|
# ==> ctx2 (1171456) -> BL, BR
|