# 🔴 A voice cue's three streams are DECODED TOGETHER — they are not alternative presentations **Classification: measured.** Xenia Canary, `--xma_param_probe=true`, one boot into the intro movie, 2026-08-29. Log excerpt committed at [`data/voice-three-streams-runtime.txt`](../data/voice-three-streams-runtime.txt). **This refutes a framing of mine that two documents and the port's exporter were built on**, so it is written as its own page rather than as an edit. ## What was believed [`voice-region-leading-chunk.md`](voice-region-leading-chunk.md) established that a long voice cue's byte span holds **three** streams, and read them as *three presentations of one take* — a defensible reading of the evidence then available (they are the same duration, and their content correlates strongly). From it came the instruction **"take one stream, do not sum"**, which the port implemented. The open question was *which* presentation the game plays. ## What the game does The question has no answer, because the premise is wrong. Canary's `xma_param_probe` — a cvar whose own comment says it is keyed so as to reveal "**WHICH sub-wave of a movie's `.slb` the game actually decodes**" — shows the guest opening **three XMA contexts and decoding all three concurrently**: | context | packets | `byte_size` | channels | rate | disc stream | |---|---|---|---|---|---| | **0** | 632 | **1 294 336** | 2 | 48 000 | `ADV` stream 1 (RIFF 1 294 396 − 60) | | **1** | 546 | **1 118 208** | 2 | 48 000 | `ADV` stream 2 (1 118 268 − 60) | | **2** | 572 | **1 171 456** | 2 | 48 000 | `ADV` stream 3 (1 171 516 − 60) | **Three-way, byte-exact**, against sizes taken independently off the disc. Only these three contexts appear in the run. So the three streams are **simultaneous**, not alternative. A consumer that picks one is discarding two thirds of what the game mixes. ## 🟡 The obvious reading is 5.1, and it is NOT established Three concurrent **stereo** streams is six channels, and N stereo streams is exactly how XMA carries multichannel audio on the 360. It would explain a lot at once: * the differing declared byte rates — different channel pairs, different content, same encoder quality; * the port's content measurements, which become measurements *of channels*: `ADV` stream 2 is **mono-in-stereo** (one channel digitally silent — a centre paired with a silent LFE looks exactly like that), stream 3 is **dual-mono** (a centre-panned line in an L/R pair is L = R exactly), and `S00A`'s third stream is **digital silence** (surrounds, on a dialogue-only cue); * stream 3 measuring **0.60 ×** stream 2 with the residual 26.8 dB down — which is what a correlated channel pair at a lower level looks like, and *not* only what a duplicate looks like; * ✅ and the census dichotomy already on record: inter-descriptor spans hold **1 stream or 3, never 2** (258 and 28). If 3-stream is 5.1 and 1-stream is mono/stereo, the missing 2 is the missing 4-channel configuration. ⚠️ **Against it**, and the reason this stays 🟡: all three `fmt ` chunks declare `ChannelMask = 0x0002` **identically**, which is not what distinct channel roles should look like. Either the mask is unset on this disc or the offset is mis-taken; it is weak evidence either way, and no channel-role assignment is claimed here. ❔ **Which stream is which channel pair is unknown.** Nothing measured here assigns roles, and the port must not infer them from stream order. ## 🔴 What this withdraws * **"Three presentations of one take"** — withdrawn. Three concurrent streams of one take. * **"Take one stream, do not sum"** — withdrawn. It was my instruction, the port implemented it, and it discards two of three decoded streams. ⚠️ **This does not make the previous behaviour right either.** An equal-gain `1/n` sum of channel pairs is not a downmix, and the port's measured 6.02 dB loss from summing a silent stream was real. The correct handling needs the channel roles, which are open. **Neither "pick one" nor "sum at 1/n" is established; a consumer is authoring, and should say so.** * **"Which presentation does the game play?"** — dissolved rather than answered. ✅ **What survives untouched:** every content measurement the port made, and the byte-level structure in [`voice-region-leading-chunk.md`](voice-region-leading-chunk.md) — the leading chunk being stream 1 clipped by our own 1.5 MB guard, the 70 + 8 + 17 decomposition, the bank-header discriminator. Those are about bytes and they did not depend on the framing. ## Reach * **One cue, one boot.** `ADV` only. That 28 cues on the disc are 3-stream is decoded from the bytes, but that all 28 decode concurrently is measured on **one** of them. * The probe fires on **first decode** per `(buffer ptr, packet count)`, so this shows all three being *started*; it does not by itself prove they play to the end in lockstep. Their equal durations and the port's sample-synchronous correlation both point that way. * Nothing here identifies the **mix** the guest applies downstream of the three decoders.