slb: derive the leading-stream data offset instead of assuming 1392

HEADERLESS_DATA_OFFSET is the value the offset takes in <lang>\etc\, not a
property of the format. The leading stream is a whole number of 2048-byte XMA1
packets ending at the first RIFF, so its start is first_riff % XMA1_PACKET.
Disc-wide that takes four values -- 1392, 1468, 1600, 1728 -- varying by
language and subdirectory.

Verified by decoding, not by arithmetic: on a random 140-bank sample with a
non-empty leading region, the derived offset yields more audio in 85, identical
in 54 (the eng\etc controls, where it must and does reproduce the old
behaviour) and less in 1. Median gain among the improved is 70x --
eng\Voice\VOICE_TCAF_592 goes 1506 -> 97152 bytes, jpn 2910 -> 127178.

This withdraws my own claim from earlier today that the Japanese banks were a
different undecoded layout. They are the same format with a different offset;
I had treated a constant derived from one subdirectory as a property of the
format. The same error was hiding the identical defect in 1873 eng\Voice banks.
This commit is contained in:
Sylpheed RE agent
2026-08-26 03:56:50 +00:00
parent a6fb570677
commit d15b3d8d85
3 changed files with 178 additions and 3 deletions

View File

@@ -84,3 +84,72 @@ fn all_zero_leading_region_is_skipped() {
// Two sub-waves, both from the RIFF section — no synthesised third.
assert_eq!(slb::to_xma_riffs(&b).len(), 2);
}
fn bank_named(root: &Path, path: &str) -> Vec<u8> {
let snd = PakArchive::open(root.join("dat/sound.pak")).expect("sound.pak");
let entry = snd.find_by_name(path).unwrap_or_else(|| panic!("{path} present"));
snd.read(entry).expect("read")
}
/// `HEADERLESS_DATA_OFFSET` is the `<lang>\etc\` case, not the format.
///
/// The leading stream is a whole number of packets ending at the first `RIFF`,
/// so its start is `first_riff % XMA1_PACKET`. Disc-wide that takes four values
/// and only 1392 matches the old constant — assuming it elsewhere starts the
/// decode mid-packet. See docs/re/structures/slb-data-offset.md.
#[test]
fn leading_data_offset_is_derived_not_assumed() {
skip_without_disc!(root);
// (bank, expected derived offset). The `etc` banks must still land on the
// old constant — that is the no-regression half of the test.
for (path, want) in [
("eng\\etc\\VOICE_D_452.slb", 1392usize),
("eng\\etc\\VOICE_D_453.slb", 1392),
("eng\\Voice\\VOICE_TCAF_592.slb", 1468),
("jpn\\Voice\\VOICE_TCAF_592.slb", 1728),
("jpn\\etc\\VOICE_D_452.slb", 1600),
] {
let b = bank_named(&root, path);
let ri = b.windows(4).position(|w| w == b"RIFF").expect("has a RIFF");
let got = slb::leading_data_offset(ri);
assert_eq!(got, want, "{path}: derived offset");
assert_eq!(
(ri - got) % slb::XMA1_PACKET,
0,
"{path}: leading stream is not a whole packet count"
);
assert!(
got == slb::HEADERLESS_DATA_OFFSET || got > slb::HEADERLESS_DATA_OFFSET,
"{path}: offsets below the old constant are unexplained"
);
}
}
/// The banks the old constant mis-decoded now carry a leading sub-wave, and the
/// ones it decoded correctly are untouched.
#[test]
fn derived_offset_recovers_voice_banks_without_regressing_etc() {
skip_without_disc!(root);
for path in ["eng\\Voice\\VOICE_TCAF_592.slb", "jpn\\Voice\\VOICE_TCAF_592.slb"] {
let b = bank_named(&root, path);
let ri = b.windows(4).position(|w| w == b"RIFF").expect("has a RIFF");
// Under the old constant this leading region was not a whole packet
// count, so `to_xma_riffs` emitted no leading sub-wave at all.
assert_ne!(
(ri - slb::HEADERLESS_DATA_OFFSET) % slb::XMA1_PACKET,
0,
"{path}: expected the OLD constant to mis-align here"
);
let riffs = slb::to_xma_riffs(&b);
assert!(
riffs.len() >= 2,
"{path}: expected a leading sub-wave plus at least one RIFF, got {}",
riffs.len()
);
}
// Control: an `etc` bank still produces what it did before.
let b = bank_named(&root, "eng\\etc\\VOICE_D_452.slb");
assert_eq!(slb::leading_data_offset(
b.windows(4).position(|w| w == b"RIFF").unwrap()),
slb::HEADERLESS_DATA_OFFSET);
}