The labelled set has a RIFF and the scan is unbounded, so it reads past the RIFF there -- the headline number could have been borrowing discrimination that a RIFF-less bank cannot offer. Confining the scan to the leading region gives 69.98%, which looks like exactly that problem. It is not. Split by how much leading audio there is: on the 989 banks with >=24 packets of it, the scan is 100% correct with ZERO ties, whether or not the RIFF is in range. The 69.98% is an artifact of short leading regions, where two or three packets are not enough to separate candidates. A RIFF-less bank is a whole pak entry, so 24 packets are always available. The 99.62% is conservative for the population the scan serves, not optimistic.
199 lines
9.2 KiB
Markdown
199 lines
9.2 KiB
Markdown
# `.slb` leading-stream data offset — 1392 was never a constant
|
||
|
||
**✅ Settled 2026-08-26, verified by decoding.** A bank's leading headerless
|
||
packet stream does not start at a fixed offset. It starts at
|
||
**`first_riff % 2048`**. `HEADERLESS_DATA_OFFSET = 1392` is the value that
|
||
offset happens to take in `<lang>\etc\`, and assuming it everywhere starts the
|
||
decode mid-packet and throws away almost all of the audio.
|
||
|
||
## The rule
|
||
|
||
XMA1 packets are 2048 bytes and the leading stream is a whole number of them
|
||
ending at the first `RIFF`. So its start is forced:
|
||
|
||
start = first_riff % 2048
|
||
|
||
Disc-wide that lands on exactly **four** values — 1392, 1468, 1600, 1728 — all
|
||
of the form `1392 + 4k`. Across the 3 965 Japanese and 3 393 English banks with
|
||
a non-empty leading region, no other value occurs:
|
||
|
||
| | 1392 | 1468 | 1600 | 1728 |
|
||
|---|---|---|---|---|
|
||
| `eng\etc`, `eng\Movie`, `eng\Briefing` | 1 520 | — | — | — |
|
||
| `eng\Voice` | 8 | 1 873 | — | — |
|
||
| `jpn\etc` | — | 1 402 | 303 | — |
|
||
| `jpn\Briefing` | — | 71 | — | — |
|
||
| `jpn\Movie` | — | — | 61 | — |
|
||
| `jpn\Voice` | — | — | 2 033 | 95 |
|
||
|
||
It varies by **language and subdirectory**, which is why a constant derived
|
||
from `eng\etc\` looked right for years' worth of the banks anyone had reason to
|
||
open.
|
||
|
||
## Verified by decoding, not by arithmetic
|
||
|
||
The alignment argument alone proves nothing — any offset can be made to "align"
|
||
by definition. The test is whether more audio comes out. Decoded through
|
||
FFmpeg's `xma1` at mono/48 kHz, on a random sample of **140** banks that have a
|
||
non-empty leading region:
|
||
|
||
| outcome | banks |
|
||
|---|---|
|
||
| more audio at `ri % 2048` | **85** |
|
||
| byte-identical | 54 |
|
||
| less audio | **1** |
|
||
|
||
Median gain among the improved: **70×**. The 54 identical ones are the control —
|
||
they are the `eng\etc`-style banks where `ri % 2048` *is* 1392, so the rule
|
||
must and does reproduce the old behaviour exactly. Individual cases:
|
||
|
||
eng\Voice\VOICE_TCAF_592.slb 1 506 -> 97 152 bytes (65x)
|
||
jpn\Voice\VOICE_TCAF_592.slb 2 910 -> 127 178 bytes (44x)
|
||
eng\etc\VOICE_D_452.slb 30 154 -> 30 154 bytes (unchanged, control)
|
||
|
||
## The one counterexample — ✅ explained
|
||
|
||
`eng\Voice\VOICE_TCAF_608.slb` decodes 2 840 bytes at 1392 and 896 at 1468.
|
||
|
||
It is **not** a bank where the old constant works and the derived offset fails:
|
||
both offsets yield well under a tenth of a second from a 38 988-byte region,
|
||
i.e. both fail, and 1392 merely produces marginally more garbage.
|
||
|
||
**The reason is that the bank is truncated.** Its `data` chunk declares 759 808
|
||
bytes and the pak entry holds 8 864 — **99 % short**. There is almost nothing
|
||
there to decode at any offset. See the section below.
|
||
|
||
## ❌ This withdraws my own claim from earlier the same day
|
||
|
||
[`sound-pak-contents.md`](sound-pak-contents.md) reported that the leading
|
||
region rule holds for "0 of 5 100 Japanese banks" and filed a backlog item
|
||
saying the Japanese banks were a different, undecoded layout. **That was wrong.**
|
||
The Japanese banks are the same format; only the offset differs. The measurement
|
||
behind it was correct — zero of them satisfy `(riff − 1392) % 2048 == 0` — but
|
||
the conclusion drawn from it was not, and the reason is instructive: I treated
|
||
`HEADERLESS_DATA_OFFSET` as a property of the format when it was a property of
|
||
the sample the format was derived from.
|
||
|
||
The same error was hiding a defect in the **English** set too: 1 873 `eng\Voice`
|
||
banks sit at 1468 and were being decoded mid-packet just as badly.
|
||
|
||
## The `RIFF`-less banks had the same bug, plus a worse one
|
||
|
||
**✅ Settled 2026-08-26.** 1 495 banks (799 `jpn`, 696 `eng`) carry no `RIFF` at
|
||
all and take a separate code path. That path was wrong twice over:
|
||
|
||
1. it used the constant offset, with no `RIFF` to derive from; and
|
||
2. it built a **stereo** `fmt` chunk.
|
||
|
||
Decoded across a random 48-bank sample:
|
||
|
||
| | |
|
||
|---|---|
|
||
| banks where the old stereo-at-1392 beat the best mono offset | **0 of 48** |
|
||
| median gain | **184×** |
|
||
| range | 25× – 489 344× |
|
||
|
||
Stereo is the same failure signature recorded for the leading segment: it stops
|
||
after one frame. Individual banks went from 0–4 816 bytes to 180 000–380 000.
|
||
|
||
The winning offsets fall out **by directory**, and they reproduce the
|
||
distribution measured independently from the `RIFF`-bearing banks — which is the
|
||
cross-check that makes this more than curve-fitting:
|
||
|
||
eng\etc 1392 (11/11) eng\Voice 1468 (9/9) eng\Briefing 1392 (2/2)
|
||
jpn\Voice 1600 (12/13) jpn\etc 1468 (8/12), 1600 (4)
|
||
|
||
Note `jpn\etc` splits, so the **path alone is not enough** to pick the offset.
|
||
|
||
### Picking the offset without a decoder
|
||
|
||
An XMA1 packet opens with a big-endian header — 6 bits frame count, 15 bits
|
||
frame-offset-in-bits, 3 bits metadata, 8 bits packet-skip. At the true offset
|
||
those fields stay in range packet after packet; one byte off and they do not.
|
||
Scoring the first 24 packets and taking the best candidate:
|
||
|
||
**7 330 of 7 358 (99.62 %)** on the labelled set — every bank that *has* a
|
||
`RIFF`, where the answer is forced and therefore known. All **28** misses are
|
||
ties on the top score; there is not a single case where the scan picks wrongly
|
||
with a unique winner. `scan_data_offset` therefore falls back to 1392 on a tie.
|
||
|
||
This is used only for the `RIFF`-less banks. Where a `RIFF` exists the offset is
|
||
derived from it exactly, never scanned.
|
||
|
||
### Is that 99.62 % transferable? — checked, and it is conservative
|
||
|
||
The labelled set has a `RIFF`; the population the scan actually serves does not.
|
||
Since the scan is unbounded it reads *past* the `RIFF` on labelled banks, so the
|
||
99.62 % could have been borrowing discriminating power that a `RIFF`-less bank
|
||
cannot offer. That would make the headline number optimistic for the only case
|
||
it is used in — worth checking before trusting it.
|
||
|
||
Confining the scan to the leading region drops it to **69.98 %** with 1 910
|
||
ties, which at first looks like exactly that problem. It is not. Splitting by
|
||
how much leading audio there is separates the two explanations:
|
||
|
||
| | correct | ties |
|
||
|---|---|---|
|
||
| unbounded, all 7 358 labelled banks | 99.62 % | 28 |
|
||
| confined to the leading region, all 7 358 | 69.98 % | 1 910 |
|
||
| **≥24 packets of leading audio (989 banks), unbounded** | **100 %** | **0** |
|
||
| **≥24 packets of leading audio (989 banks), confined** | **100 %** | **0** |
|
||
|
||
The last two rows settle it. Where there is enough audio to score, the
|
||
discriminator is perfect **whether or not the `RIFF` is in range** — so it is
|
||
not leaning on the `RIFF`. The 69.98 % is an artifact of *short* leading
|
||
regions: with only two or three packets to judge, candidates tie and the
|
||
tie-break decides. Unboundedness helps those banks by giving the scan more bytes,
|
||
which is why the two columns differ at all.
|
||
|
||
A `RIFF`-less bank is a whole pak entry, tens of kilobytes, so 24 packets are
|
||
always available — it is always in the 100 % regime. **The 99.62 % figure is
|
||
therefore conservative for the population the scan is used on**, not optimistic.
|
||
|
||
|
||
## What this does not settle
|
||
|
||
* **Why the offset takes those four values**, and what the bytes before it are.
|
||
There is no length field in the first 64 bytes — banks open on high-entropy
|
||
data — so the offset is derived, not read.
|
||
|
||
* **The 28 ties.** The scan cannot separate them and falls back to 1392, which
|
||
is right for roughly a third of that population and wrong for the rest.
|
||
* **Why the offset takes exactly these four values by directory** is still
|
||
unexplained — see above.
|
||
* Nothing here was run **in the game** — this is a decoder-side result measured
|
||
with FFmpeg as the oracle.
|
||
|
||
|
||
## 🟡 Most banks declare more `data` than they store
|
||
|
||
**Measured 2026-08-26.** Of the 7 586 banks that carry both a `RIFF` and a
|
||
`data` chunk after it, **5 296 (69.8 %)** declare a `data` size larger than the
|
||
bytes actually present in the pak entry. The remaining 2 290 declare *less*,
|
||
which is the ordinary multi-sub-wave case. **Not one declares exactly what it
|
||
holds.**
|
||
|
||
Worst cases run to 99 % short:
|
||
|
||
eng\Movie\VOICE_RT16C.slb declared 1 810 432 available 489 392 -73 %
|
||
jpn\etc\VOICE_D_589.slb declared 1 177 600 available 6 708 -99 %
|
||
eng\Voice\VOICE_TCAF_608.slb declared 759 808 available 8 864 -99 %
|
||
|
||
This **contradicts a claim in the decoder's own comment**, which says the
|
||
declared size "is honest per sub-wave". It is not, for about seven banks in ten.
|
||
The code is nonetheless safe — it clamps the range with `.min(slb.len())` — so
|
||
this is a documentation defect and an integrity observation, not a crash.
|
||
|
||
⚠️ **Method note on this measurement.** My first pass searched for `data` from
|
||
offset 0, which can hit those four bytes by chance inside the leading audio
|
||
region and read a garbage length. Re-running it anchored *after* the first
|
||
`RIFF` changed the count from 5 038 to 5 296 — the flaw was slightly
|
||
*under*-counting, but it could as easily have gone the other way, and an
|
||
unanchored chunk search over binary audio is not a safe way to ask this
|
||
question.
|
||
|
||
❔ **Why** the declared sizes are too large is **not settled**. Plausible
|
||
readings — an authoring-time allocation that was never trimmed, or deliberate
|
||
truncation of unused tails — are guesses; nothing here distinguishes them, and
|
||
the game has not been observed reading one of these banks.
|