Of the 7586 banks with a RIFF and a data chunk after it, 5296 declare a data size larger than the pak entry holds; 2290 declare less (the ordinary multi-sub-wave case); NONE declare exactly what they hold. This contradicts the decoder comment claiming the declared size 'is honest per sub-wave'. The code clamps, so it is a documentation defect, not a crash. It also closes the loose end from the offset work: eng\Voice\VOICE_TCAF_608, the single bank where neither offset decoded, is 99% short -- there is nothing there to decode. Method note recorded: my first pass searched for 'data' from offset 0, which can match by chance inside the leading audio region. Anchoring the search after the first RIFF moved the count 5038 -> 5296. Separately, the 55 'early RIFF' English banks are not an anomaly: all 55 sit at exactly 1392 behind a zero-filled header -- a zero-length leading region, which both the old and new code already handle correctly.
7.6 KiB
.slb leading-stream data offset — 1392 was never a constant
✅ Settled 2026-08-26, verified by decoding. A bank's leading headerless
packet stream does not start at a fixed offset. It starts at
first_riff % 2048. HEADERLESS_DATA_OFFSET = 1392 is the value that
offset happens to take in <lang>\etc\, and assuming it everywhere starts the
decode mid-packet and throws away almost all of the audio.
The rule
XMA1 packets are 2048 bytes and the leading stream is a whole number of them
ending at the first RIFF. So its start is forced:
start = first_riff % 2048
Disc-wide that lands on exactly four values — 1392, 1468, 1600, 1728 — all
of the form 1392 + 4k. Across the 3 965 Japanese and 3 393 English banks with
a non-empty leading region, no other value occurs:
| 1392 | 1468 | 1600 | 1728 | |
|---|---|---|---|---|
eng\etc, eng\Movie, eng\Briefing |
1 520 | — | — | — |
eng\Voice |
8 | 1 873 | — | — |
jpn\etc |
— | 1 402 | 303 | — |
jpn\Briefing |
— | 71 | — | — |
jpn\Movie |
— | — | 61 | — |
jpn\Voice |
— | — | 2 033 | 95 |
It varies by language and subdirectory, which is why a constant derived
from eng\etc\ looked right for years' worth of the banks anyone had reason to
open.
Verified by decoding, not by arithmetic
The alignment argument alone proves nothing — any offset can be made to "align"
by definition. The test is whether more audio comes out. Decoded through
FFmpeg's xma1 at mono/48 kHz, on a random sample of 140 banks that have a
non-empty leading region:
| outcome | banks |
|---|---|
more audio at ri % 2048 |
85 |
| byte-identical | 54 |
| less audio | 1 |
Median gain among the improved: 70×. The 54 identical ones are the control —
they are the eng\etc-style banks where ri % 2048 is 1392, so the rule
must and does reproduce the old behaviour exactly. Individual cases:
eng\Voice\VOICE_TCAF_592.slb 1 506 -> 97 152 bytes (65x)
jpn\Voice\VOICE_TCAF_592.slb 2 910 -> 127 178 bytes (44x)
eng\etc\VOICE_D_452.slb 30 154 -> 30 154 bytes (unchanged, control)
The one counterexample — ✅ explained
eng\Voice\VOICE_TCAF_608.slb decodes 2 840 bytes at 1392 and 896 at 1468.
It is not a bank where the old constant works and the derived offset fails: both offsets yield well under a tenth of a second from a 38 988-byte region, i.e. both fail, and 1392 merely produces marginally more garbage.
The reason is that the bank is truncated. Its data chunk declares 759 808
bytes and the pak entry holds 8 864 — 99 % short. There is almost nothing
there to decode at any offset. See the section below.
❌ This withdraws my own claim from earlier the same day
sound-pak-contents.md reported that the leading
region rule holds for "0 of 5 100 Japanese banks" and filed a backlog item
saying the Japanese banks were a different, undecoded layout. That was wrong.
The Japanese banks are the same format; only the offset differs. The measurement
behind it was correct — zero of them satisfy (riff − 1392) % 2048 == 0 — but
the conclusion drawn from it was not, and the reason is instructive: I treated
HEADERLESS_DATA_OFFSET as a property of the format when it was a property of
the sample the format was derived from.
The same error was hiding a defect in the English set too: 1 873 eng\Voice
banks sit at 1468 and were being decoded mid-packet just as badly.
The RIFF-less banks had the same bug, plus a worse one
✅ Settled 2026-08-26. 1 495 banks (799 jpn, 696 eng) carry no RIFF at
all and take a separate code path. That path was wrong twice over:
- it used the constant offset, with no
RIFFto derive from; and - it built a stereo
fmtchunk.
Decoded across a random 48-bank sample:
| banks where the old stereo-at-1392 beat the best mono offset | 0 of 48 |
| median gain | 184× |
| range | 25× – 489 344× |
Stereo is the same failure signature recorded for the leading segment: it stops after one frame. Individual banks went from 0–4 816 bytes to 180 000–380 000.
The winning offsets fall out by directory, and they reproduce the
distribution measured independently from the RIFF-bearing banks — which is the
cross-check that makes this more than curve-fitting:
eng\etc 1392 (11/11) eng\Voice 1468 (9/9) eng\Briefing 1392 (2/2)
jpn\Voice 1600 (12/13) jpn\etc 1468 (8/12), 1600 (4)
Note jpn\etc splits, so the path alone is not enough to pick the offset.
Picking the offset without a decoder
An XMA1 packet opens with a big-endian header — 6 bits frame count, 15 bits frame-offset-in-bits, 3 bits metadata, 8 bits packet-skip. At the true offset those fields stay in range packet after packet; one byte off and they do not. Scoring the first 24 packets and taking the best candidate:
7 330 of 7 358 (99.62 %) on the labelled set — every bank that has a
RIFF, where the answer is forced and therefore known. All 28 misses are
ties on the top score; there is not a single case where the scan picks wrongly
with a unique winner. scan_data_offset therefore falls back to 1392 on a tie.
This is used only for the RIFF-less banks. Where a RIFF exists the offset is
derived from it exactly, never scanned.
What this does not settle
-
Why the offset takes those four values, and what the bytes before it are. There is no length field in the first 64 bytes — banks open on high-entropy data — so the offset is derived, not read.
-
The 28 ties. The scan cannot separate them and falls back to 1392, which is right for roughly a third of that population and wrong for the rest.
-
Why the offset takes exactly these four values by directory is still unexplained — see above.
-
Nothing here was run in the game — this is a decoder-side result measured with FFmpeg as the oracle.
🟡 Most banks declare more data than they store
Measured 2026-08-26. Of the 7 586 banks that carry both a RIFF and a
data chunk after it, 5 296 (69.8 %) declare a data size larger than the
bytes actually present in the pak entry. The remaining 2 290 declare less,
which is the ordinary multi-sub-wave case. Not one declares exactly what it
holds.
Worst cases run to 99 % short:
eng\Movie\VOICE_RT16C.slb declared 1 810 432 available 489 392 -73 %
jpn\etc\VOICE_D_589.slb declared 1 177 600 available 6 708 -99 %
eng\Voice\VOICE_TCAF_608.slb declared 759 808 available 8 864 -99 %
This contradicts a claim in the decoder's own comment, which says the
declared size "is honest per sub-wave". It is not, for about seven banks in ten.
The code is nonetheless safe — it clamps the range with .min(slb.len()) — so
this is a documentation defect and an integrity observation, not a crash.
⚠️ Method note on this measurement. My first pass searched for data from
offset 0, which can hit those four bytes by chance inside the leading audio
region and read a garbage length. Re-running it anchored after the first
RIFF changed the count from 5 038 to 5 296 — the flaw was slightly
under-counting, but it could as easily have gone the other way, and an
unanchored chunk search over binary audio is not a safe way to ask this
question.
❔ Why the declared sizes are too large is not settled. Plausible readings — an authoring-time allocation that was never trimmed, or deliberate truncation of unused tails — are guesses; nothing here distinguishes them, and the game has not been observed reading one of these banks.