Files
Sylpheed/docs/re/structures/voice-region-leading-chunk.md
sylph-decoder 7a0e18889f re: the dual-mono explanation for the extra stream bytes does not generalise
The port chose a voice presentation on the argument that ADV chunk 1 is
mono-in-stereo and chunk 2 is dual-mono, so chunk 2 s extra bytes encode a
duplicated channel rather than fidelity -- which would explain its higher
declared PsuedoBytesPerSec without appealing to encode quality.

Their ADV channel measurement stands. The generalisation does not. If stream 3
were systematically the same take with its channel duplicated, its size ratio to
stream 2 would be tight across the 28 three-stream cues. Measured:

  min 0.0778 (S00A, the silent one)
  median 1.2565
  max 2.9163 (S06A)
  sd 0.5057
  within 15 percent of 1.0: 12 of 28

A 37x spread is not a duplicated channel, and the declared rates scatter with
them -- S06A is 5661 against 16513 B/s. Whatever distinguishes the three streams
varies per cue rather than being a fixed channel-configuration triple.

This does not touch the port s decision, which is to take the loudest
presentation: that is a per-asset content measurement, not a structural rule, so
a scattering ratio cannot undermine it. It touches the explanation, which should
not harden into a fact about the format.

Two curiosities recorded: S12B s three streams are byte-size identical at 14396
each, and BIRD_224 is 3-stream while being a non-movie cue, so the shape is not
exclusive to cutscenes.

Also narrows the settle-time page s own generalisation. The port measured its
boot the way this corpus measured the game and found the sequencer NOT late --
its 0.6 s discrepancy was arrival-to-arrival timestamps compared against visible
spans, the plate-delay trap in a second place. So what is supported is that
rest.t is the wrong landmark for the TITLE, not that everything paced off it is
late. And the offered re-take of the one-run menu figures is recorded as
declined, with the reason, rather than left looking unfinished.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QsEPXWVaEpyfudtR6re1Pd
2026-08-29 15:37:46 +00:00

17 KiB
Raw Blame History

🟡 A movie-voice region's THIRD chunk is not the bank-header case — and nothing else claims it

Status: the structure is decoded, disc-wide, 95/95 regions. 🟡 what the leading chunk contains is open, and this page states the reach of that negative rather than guessing. 🔴 One claim this page carried — that chunks 1 and 2 are two stems of one performance — is withdrawn; see the bottom of the page.

Raised by the port: media::sound_bank_riffs("BGM_103.slb") used to return three sub-waves where bgm-two-stems says two, and slb-bank-header-not-a-wave attributed the extra to the bank header. The port then hit the same 2+1 signature on a different asset kind — a resolved movie-voice region also decoding to three chunks — and asked whether one explanation covers both.

It does not. They are two different structures, and the corpus's own code already tells them apart; what it does not do is say which one it is looking at.

Tool: cargo run -p sylpheed-formats --example voice_region_chunks -- $SYLPHEED_DISC. Census committed at data/voice-region-chunk-census.txt.

Disc-wide: a voice region never begins at a RIFF

All 95 English movie-voice regions the manifest binds:

how the region opens regions chunks it yields
a bank headerbank_header_len fires, 10 240 B = 5 packets exactly, every time 78 1 (×70) or 3 (×8)
a leading headerless streambank_header_len is None 17 3, every time
directly at a RIFF 0

And the leading streams are not ragged. All 17 have a length ≡ 1392 (mod 2048) — no other residue occurs — which is exactly HEADERLESS_DATA_OFFSET, the <lang>\etc\ data offset that slb-data-offset derives. So a leading stream is 1392 B of preamble followed by a whole number of 2048-byte XMA1 packets: 394 of them on ADV, 646 on S00A, 900 on S12C.

That is the discriminator the port needed, and it is mechanical:

bank_header_len(region) == Some(n)   ->  n is 10240, a header, already consumed
bank_header_len(region) == None      ->  first_riff % 2048 == 1392, a real stream

🔴 So the BGM explanation does not transfer

slb.rs's own doc comment predicted this and disagrees with "drop it": the header signature fires on 28 sound.pak entries, all music banks, with "zero false positives on the 7 993 mid-bank windows, where the leading region IS real". A movie-voice region is a mid-bank window by construction — resolve_movie_voice_region anchors its start at the predecessor cue's trailer, deliberately, because the cue may sit either side of its own .slb chunk.

⚠️ The 3-chunk count is not evidence of the leading region at all. Eight regions open with a bank header and still yield three chunks (S11A, S12A, S12B, S13B, S15B, …). Counting chunks cannot distinguish the two cases; only bank_header_len can.

🔴 "It is the previous cue's audio, so dropping it is right" — TESTED, and it fails

The obvious defence of dropping the leading chunk is that the region starts at the predecessor's trailer, so those bytes are the previous line of dialogue. That is checkable without decoding anything: take each leading span [start, start + first_riff) and ask whether any other resolved region covers it.

leading spans lying wholly or partly inside another movie-voice region 0 of 17
…expressed as covered fraction 0.0 % on every one

For contrast, the regions themselves are not disjoint — 16 overlapping pairs, 60 exactly-adjacent boundaries, 18 gaps — so the test is capable of finding an overlap, and it finds none here. 73 of 78 bank-header regions start exactly where another region ends; 0 of 17 leading-stream regions do.

So the leading chunk is not another movie's voice.

RESOLVED 2026-08-29 — the leading chunk is the MOVIE'S OWN cue, and the mechanism is a guard

The section that stood here left this open and named an in-mission VOICE_D_* line as the leading hypothesis. That hypothesis is refuted. The port pointed out that the byte-span test already built settles it without anyone listening, if the enumeration is widened past the 95 manifest-bound movies — and it does.

Rather than resolving cues one at a time, scan the stream for every trailer descriptor: the (id: u32be, 0x11, …) pair whose id repeats at +0x800, which movie_voice documents as the end of a cue's audio, with a false-match probability of ~2⁻⁶⁴. The full descriptor list is the stream's complete cue partition, movie and mission alike. Over a 116.2 MB window covering every region: 287 descriptors, and all 287 carry an id the registry names (4 280 cue names).

Tool: cargo run -p sylpheed-formats --example voice_stream_cue_map -- $SYLPHEED_DISC, output at data/voice-stream-cue-map.txt.

The leading span belongs to the movie itself — 17 of 17

Each leading span is bracketed by desc(N-1) .. desc(N), and in every case desc(N) is that movie's own cue id:

movie leading span ends at descriptor
ADV id 1600 = VOICE_ADV movie cue
S00A 1501 = VOICE_S00A movie cue
S14A 1524 = VOICE_S14A movie cue
…all 17 movie cue, 0 mission lines

By the stream's own rule — cue N's audio is [desc(N-1) .. desc(N)] — those bytes are this movie's dialogue. 🔴 So "it is an in-mission VOICE_D_* line" is dead, and so is any reading in which the leading chunk is foreign audio.

And the mechanism is resolve_movie_voice_region's own guard

resolve_movie_voice_region takes the predecessor trailer as the region start, but guards it with end - start < 1_500_000 and falls back to the .slb TOC anchor when that fails. If the guard is the cause, the stream-opening regions should be exactly the cues whose true span exceeds it:

cues of which stream-opening
true cue span ≥ 1.5 MB 17 17
true cue span < 1.5 MB 78 0

Perfect discrimination, both ways. A long cue's region does not start at its cue boundary; it starts at the anchor, mid-cue, and everything from the anchor to the next .slb RIFF becomes the leading chunk. That is the whole phenomenon.

⚠️ The anchor sits a constant 504 464 B after the true predecessor trailer on all 17 — not an approximate constant, the same number every time. That regularity is unexplained and is worth someone's attention; it says the .slb chunk boundary is placed at a fixed distance from a trailer.

RESOLVED 2026-08-29 (later) — a long cue is stored as THREE presentations of one take

The section that stood here declined to convert bytes into seconds and left open why one cue's byte span decodes to ~2.6× the movie. That is now answered, by two independent routes that agree.

The port's route — envelope cross-correlation, with controls. Sliding with overhang at both ends and normalised over the overlap only:

best r at lag overlap
ADV chunk 0 → chunk 1 0.998 +52.8 s 84.5 s
S00A chunk 0 → chunk 1 0.932 +25.6 s 68.0 s
control — ADV chunk 0 → itself 1.000 0.0 s
control — ADV chunk 0 → S00A chunk 1 0.289 28.2 s

Both lags put chunk 0 flush against the end of chunk 1 (52.8 + 84.55 = 137.35 against 137.324; 25.6 + 68.07 = 93.67 against 93.694). In the sample domain, after refining the lag and best-fitting a scalar, the residual is 16.70 dB below the target over 84.5 s (ADV) and 23.15 dB over 68.1 s (S00A) — 9899.5 % of the energy is a scaled copy.

⚠️ Their earlier 0.768 is withdrawn by them: that search only tried lags where the shorter chunk fitted wholly inside the longer, and scored best on the boundary of its own range — where a statistic lands when it has found nothing.

My route — byte rates, from the disc, with no decoder. If the leading chunk is the tail of a full-length first stream, then the whole leading stream (the part before the region's anchor plus the part after) should be one complete take of chunk 1's duration. Using only the port's durations and the disc's byte counts:

full leading stream, ADV 504 464 + 808 304 = 1 312 768 B
its byte rate, from chunk 0 808 304 B / 84.553 s = 9 559.7 B/s
implied duration of the whole leading stream 137.323 s
chunk 1's measured duration 137.324 s

Agreement to 1 ms over 137 s, from a quantity (byte rate) independent of the one the port measured (envelope correlation). And it confirms their point that bytes-per-second is not a constant: the three streams run at 9 560 / 8 143 / 8 531 B/s for the same 137.324 s.

And the byte structure says three, disc-wide

The reading above predicts a fixed number of stream starts inside a cue's true span [desc(N-1) .. desc(N)]. Counted directly over every inter-descriptor span in the 116.2 MB window — RIFFs, plus a leading headerless run when it is not a bank header:

streams in the span spans
1 258
3 28
2, or anything else 0

Bimodal, with nothing in between, and all 20 spans ≥ 1.5 MB are 3-stream. So a cue is stored as one stream or as three, never two.

That closes the account, and the three routes agree on the arithmetic:

the 95 movie regions cue shape what the region yields
70 1-stream 1 chunk, bank header
8 3-stream, span < 1.5 MB 3 chunks, bank header
17 3-stream, span ≥ 1.5 MB — the guard fires 3 chunks, headerless leading one

70 + 8 + 17 = 95, and the 8 here are independently the same 8 the first census found as "bank header, 3 chunks". 359 s = 84.55 + 137.32 + 137.32 — the region catches the tail of stream 1 and the whole of streams 2 and 3. The 2.6× is three presentations of one take, one of them clipped by our own guard.

⚠️ The 504 464 B constant is structural, not proportional. It is the same number on all 17 despite their differing durations. For ADV a proportional prediction lands within 8 bytes of it, which is a coincidence and should not be built on — the same prediction for S00A is 4 305 B out.

What a consumer should do meanwhile

Drop the leading chunk. It is a DUPLICATE, not a truncation. It is the cutscene's own dialogue — but it is the tail of the stream that follows it, so removing it loses nothing any other chunk does not carry. An earlier version of this page said "I would not change the exporter yet"; that hedge is lifted, and the reason the exporter's existing behaviour is right is now on the record.

🔴 But DO NOT sum chunk 1 and chunk 2 — they are the same take, not two stems. ADV chunk 2 is 0.60 × chunk 1 with the residual 26.8 dB down; S00A chunk 2 is digital silence. Summing a take with a scaled copy of itself adds ~4 dB and colours it; summing it with silence at 1/n costs 6.02 dB. Take one stream.

🔴 Which stream — my recommendation was self-contradictory, and is withdrawn

This page said "the highest-rate, highest-gain one is chunk 1". Those two criteria do not select the same stream, and the sentence should never have joined them. On ADV, chunk 1 is 1 118 268 B at 0.0 dBFS and chunk 2 is 1 171 516 B at 8.3 dBFS: chunk 2 has the higher rate and the lower level. The port implemented "highest rate", correctly, and thereby selected the quieter presentation — the opposite of what the parenthetical intended. My error, and the port caught it by checking the consequence rather than the instruction.

What the header does decode. The fmt chunk is a 32-byte XMAWAVEFORMAT, little-endian, and +0x20 is PsuedoBytesPerSec — a declared field, not an inference:

ADV chunk 1 ADV chunk 2 S00A chunk 1
+0x20 declared bytes/sec 8 142 8 530 13 485
computed from size ÷ duration 8 143.3 8 531.0 13 487.3
+0x24 sample rate 48 000 48 000 48 000
+0x18 wEncodeOptions 0x10d6 0x10d6 0x10d6
+0x31/+0x32 channels / mask 2 / 0x0002 2 / 0x0002 2 / 0x0002

So the rate is decoded and agrees with the measured one to ~0.02 %. But nothing in the header ranks the presentations. wEncodeOptions, channel count and channel mask are byte-identical across them. The header says how fast each stream is, and says nothing about which one the game plays.

🟡 So stream selection is an authored choice, and the port must know it is authoring. It is flagged in the port's manifest with the level consequence stated, which is the right handling. This is settleable in one emulator run — a capture of the intro with the dialogue audible tells you which level the game plays — and it has not been done.

Why the disc stores three presentations — quality tiers, a mix the engine selects between, an authoring artefact — is not answered here.

🔴 Refutation attempt, 2026-08-29 — "the extra bytes are a duplicated channel" does NOT generalise

The port selected a presentation on this argument: ADV chunk 1 is mono-in-stereo (channel 2 digitally silent) and chunk 2 is dual-mono (both channels identical at 8.318574), so chunk 2's extra bytes encode a duplicate of its own channel rather than fidelity — which would explain its higher declared PsuedoBytesPerSec without appealing to encode quality.

The ADV measurement is theirs and stands. The generalisation does not. If stream 3 were systematically "the same take with its channel duplicated", its size would sit in a tight ratio to stream 2 on every 3-stream cue. Measured over all 28 — data/voice-three-stream-sizes.txt, --example voice_three_stream_sizes:

stream3 / stream2
min 0.0778 (S00A, the silent one)
median 1.2565
max 2.9163 (S06A)
sd 0.5057
within 15 % of 1.0 12 of 28

A 37× spread is not a duplicated channel. The declared rates scatter with them — S06A is 5 661 against 16 513 B/s, S00A 13 485 against 1 049 — so whatever distinguishes the three streams varies per cue rather than being a fixed channel-configuration triple.

⚠️ Two curiosities worth someone's time: S12B's three streams are byte-size identical (14 396 each), and S11A's first two are (81 980). And BIRD_224 is 3-stream while being a non-movie cue, so the 3-stream shape is not exclusive to cutscenes.

What this does and does not touch. It does not touch the port's decision, which is to take the loudest presentation — that is a per-asset content measurement, not a structural rule, so a scattering ratio cannot undermine it. What it touches is the explanation: "more bytes means a duplicated channel, not better fidelity" is true of ADV and is not a fact about the format. It should not harden into one.

⚠️ Note for anyone reading our own tooling: sylpheed-cli audio info reports these chunks as 16 channels, 4310 Hz, 2-bit. Those are the wBitsPerSample (16), wEncodeOptions (0x10d6 = 4310) and channel fields read at the wrong offsets. The header above is the correct layout; the CLI's reader is misaligned for XMA1 and should not be used on these.

⚠️ Do not "fix" it by concatenating. The port measured a concatenated region at 359 s against a 137 s movie.

🔴 But the REASON this page gave was wrong, and is withdrawn (2026-08-29). It said chunks 1 and 2 are "the two-stem pattern bgm-two-stems documents for music — equal duration, played together". That claim originated with the port, I adopted it here on the strength of equal duration, and the port then refuted its own claim by decoding the content:

  • S00A chunk 2 is digital silence — 4 497 300 samples, peak −∞. Not a quiet stem. Nothing at all.
  • ADV chunk 2 is 0.60 × chunk 1 — best-fit scalar, residual 26.8 dB below the target. ~95 % of its energy is a 4.4 dB copy of chunk 1, not an independent performance.

Equal duration was a shape match and Q10's music census should not have been carried across to voice on it. ⚠️ This is how a wrong belief hardens: it was asserted in one place, adopted in a second, and the second citing the first would have made it look corroborated. It was caught because the port measured its own claim rather than the other agent's.

What ADV's near-duplicate chunk 2 is remains open — a decoding question, not a port one. What is not open is that summing a digitally silent chunk at 1/n costs 6.02 dB for nothing; the port drops silent chunks before summing, which is arithmetic rather than a content judgement.