re: the .slb leading region is 1392+n*2048 — and my fix for it is withdrawn

The structure is exact. In all five resupply banks the first RIFF sits at
HEADERLESS_DATA_OFFSET + n*2048, where 1392 is a constant this crate already
had and 2048 is the XMA1 packet size: n = 8, 1, 7, 22, 29. No free parameter
to tune, and the raw bytes agree -- high entropy from offset 0, then a zero
run immediately before the RIFF. VOICE_D_451 is the control, its single
packet being all zeros.

So I made the obvious fix, emitting that region as a sub-wave, and then
withdrew it on two measurements:

* It does not recover audio. Coverage went 5.4% -> 89.9% for VOICE_D_453, but
  the emitted stream decodes through FFmpeg to 1792 PCM bytes -- silence --
  while the RIFF sub-waves from the same banks decode to 150-270 KB. Byte
  coverage was the wrong success metric and it looked like progress.
* It is not narrow. The rule matches 1524 of the 8021 RIFF-bearing entries in
  sound.pak, including RT* movie banks that decode correctly today. Landing
  it would have risked a wide regression in order to not-fix five banks.

to_xma_riffs is back to its previous behaviour, verified by re-measuring:
coverage is 5.4% / 9.7% again. The refuted attempt is recorded in the code
beside the branch it would have changed, so the next person does not
re-derive the arithmetic and re-make the change.

XMA1_PACKET is kept as a named constant because the blast-radius scan uses
it. Artifacts: examples/voice_bank_shape.rs (structure), voice_bank_dump.rs
(sub-waves for decoding), slb_hybrid_scan.rs (the 1524 count).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PMRJjbxLqZtsb5Vb7KunPE
This commit is contained in:
Sylpheed RE agent
2026-08-25 23:30:42 +00:00
parent fedb31a5f9
commit 6fa6564be8
6 changed files with 133 additions and 6 deletions

View File

@@ -23,6 +23,10 @@
/// Fixed offset of the raw XMA1 stream in a headerless `.slb` (no `RIFF`).
pub const HEADERLESS_DATA_OFFSET: usize = 1392;
/// XMA1 packet size. A headerless stream is always a whole number of these, which
/// is how a leading stream is told apart from arbitrary bytes before a `RIFF`.
pub const XMA1_PACKET: usize = 2048;
/// Voice language for cutscene audio. Only English and Japanese voice exist on
/// the disc (subtitles cover more languages, voice does not).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
@@ -138,6 +142,22 @@ pub fn to_xma_riffs(slb: &[u8]) -> Vec<Vec<u8>> {
}
return out;
}
// ❌ A "leading headerless stream" rule was tried here and WITHDRAWN.
//
// The structure is real: in all five resupply banks the first `RIFF` sits at
// exactly `HEADERLESS_DATA_OFFSET + n*XMA1_PACKET` (n = 8, 1, 7, 22, 29), and
// 87 % of `VOICE_D_453` lies in front of it. Emitting that region as a
// sub-wave raised byte coverage from 5.4 % to 89.9 %.
//
// But byte coverage was the wrong success metric. The emitted streams decode
// to **1792 PCM bytes** — silence — through the same FFmpeg path that decodes
// the RIFF sub-waves fine, so the region is not XMA1 under the synthesised
// format. And the rule is not narrow: it matches **1524 of the 8021**
// RIFF-bearing entries in `sound.pak`, including `RT*` banks that decode
// correctly today. Landing it would have risked a large regression to fix
// five banks it does not actually fix.
//
// See docs/re/voice-bank-leading-region.md.
let mut pos = 0usize;
while let Some(ri) = find(slb, b"RIFF", pos) {
// Parse this sub-wave's fmt + data (declared size is honest per sub-wave).