Files
Sylpheed/docs/re/structures/slb-data-offset.md
Sylpheed RE agent 59c68b2f61 re: the four .slb offsets are a segment-packing phase, not a header size
Reproduced independently: X = (cumulative start of the .pNN segment) mod 2048.
The running sums of the five sound.pNN file sizes are 0, 1392, 1468, 1600, 1728
-- exactly the four values measured, plus zero. The XMA grid is 2048-aligned
inside each segment file, the segments are not multiples of 2048 long, so each
join shifts the phase.

Two of my explanations on this page were wrong even though the measurements
stand. 'Varies by language and subdirectory' was correlation, not cause --
directories cluster into segments. And there is no header at all: those bytes
are the PREVIOUS bank's audio, which is why they had no length field and looked
like data.

Also records a disagreement I could not resolve. auto/slb-loader withdraws my
'69.8% over-declare' finding as declared-sizes-are-exact, but VOICE_TCAF_608
declares 759808 bytes of data with 55296 bytes to the next TOC entry. Both
readings agree the declared size exceeds the window; mine said 'truncated',
which I withdraw as an over-claim, but 'exact' requires a wave to span ~14 named
entries. Recorded with the experiment that would decide it.
2026-08-26 05:03:56 +00:00

16 KiB
Raw Blame History

SUPERSEDED IN ITS CAUSE (2026-08-26) — see branch auto/slb-loader. The four offsets are not a header size. They are a packing phase:

X = (cumulative start of the .pNN segment holding the bank) mod 2048
segment size cumulative start start mod 2048
sound.p00 267 930 992 0 0
sound.p01 268 404 812 267 930 992 1392
sound.p02 268 404 868 536 335 804 1468
sound.p03 268 384 384 804 740 672 1600
sound.p04 14 903 296 1 073 125 056 1728

Independently reproduced here: the four values I measured are exactly the running sums of the five segment file sizes, mod 2048. The XMA grid is 2048-aligned inside each .pNN file; the segments are not multiples of 2048 long; so every join shifts the phase, and the flat concatenation the TOC addresses inherits the shift.

Two things I wrote on this page are therefore wrong in their explanation, even though the measurements stand:

  • "It varies by language and subdirectory" — that was a correlation, not a cause. Directories cluster into segments, so the per-directory table is real but explains nothing.
  • "The header is high-entropy content of a size the loader must know a priori" — there is no header. Those bytes are the previous bank's audio, which is why they looked like data and had no length field: they are data.

The heuristics below (99.62 % packet scan, 99.97 % seek residue, 99.95 % combined) are all superseded by an exact rule, verified 8 783/8 783 on the other branch. slb.rs still uses the heuristics; the exact fix needs PakArchive to expose the segment phase, which is an API change.

.slb leading-stream data offset — 1392 was never a constant

Settled 2026-08-26, verified by decoding. A bank's leading headerless packet stream does not start at a fixed offset. It starts at first_riff % 2048. HEADERLESS_DATA_OFFSET = 1392 is the value that offset happens to take in <lang>\etc\, and assuming it everywhere starts the decode mid-packet and throws away almost all of the audio.

The rule

XMA1 packets are 2048 bytes and the leading stream is a whole number of them ending at the first RIFF. So its start is forced:

start = first_riff % 2048

Disc-wide that lands on exactly four values — 1392, 1468, 1600, 1728 — all of the form 1392 + 4k. Across the 3 965 Japanese and 3 393 English banks with a non-empty leading region, no other value occurs:

1392 1468 1600 1728
eng\etc, eng\Movie, eng\Briefing 1 520
eng\Voice 8 1 873
jpn\etc 1 402 303
jpn\Briefing 71
jpn\Movie 61
jpn\Voice 2 033 95

It varies by language and subdirectory, which is why a constant derived from eng\etc\ looked right for years' worth of the banks anyone had reason to open.

Verified by decoding, not by arithmetic

The alignment argument alone proves nothing — any offset can be made to "align" by definition. The test is whether more audio comes out. Decoded through FFmpeg's xma1 at mono/48 kHz, on a random sample of 140 banks that have a non-empty leading region:

outcome banks
more audio at ri % 2048 85
byte-identical 54
less audio 1

Median gain among the improved: 70×. The 54 identical ones are the control — they are the eng\etc-style banks where ri % 2048 is 1392, so the rule must and does reproduce the old behaviour exactly. Individual cases:

eng\Voice\VOICE_TCAF_592.slb    1 506 ->  97 152 bytes   (65x)
jpn\Voice\VOICE_TCAF_592.slb    2 910 -> 127 178 bytes   (44x)
eng\etc\VOICE_D_452.slb        30 154 ->  30 154 bytes   (unchanged, control)

The one counterexample — explained

eng\Voice\VOICE_TCAF_608.slb decodes 2 840 bytes at 1392 and 896 at 1468.

It is not a bank where the old constant works and the derived offset fails: both offsets yield well under a tenth of a second from a 38 988-byte region, i.e. both fail, and 1392 merely produces marginally more garbage.

The reason is that the bank is truncated. Its data chunk declares 759 808 bytes and the pak entry holds 8 864 — 99 % short. There is almost nothing there to decode at any offset. See the section below.

This withdraws my own claim from earlier the same day

sound-pak-contents.md reported that the leading region rule holds for "0 of 5 100 Japanese banks" and filed a backlog item saying the Japanese banks were a different, undecoded layout. That was wrong. The Japanese banks are the same format; only the offset differs. The measurement behind it was correct — zero of them satisfy (riff 1392) % 2048 == 0 — but the conclusion drawn from it was not, and the reason is instructive: I treated HEADERLESS_DATA_OFFSET as a property of the format when it was a property of the sample the format was derived from.

The same error was hiding a defect in the English set too: 1 873 eng\Voice banks sit at 1468 and were being decoded mid-packet just as badly.

The RIFF-less banks had the same bug, plus a worse one

Settled 2026-08-26. 1 495 banks (799 jpn, 696 eng) carry no RIFF at all and take a separate code path. That path was wrong twice over:

  1. it used the constant offset, with no RIFF to derive from; and
  2. it built a stereo fmt chunk.

Decoded across a random 48-bank sample:

banks where the old stereo-at-1392 beat the best mono offset 0 of 48
median gain 184×
range 25× 489 344×

Stereo is the same failure signature recorded for the leading segment: it stops after one frame. Individual banks went from 04 816 bytes to 180 000380 000.

The winning offsets fall out by directory, and they reproduce the distribution measured independently from the RIFF-bearing banks — which is the cross-check that makes this more than curve-fitting:

eng\etc 1392 (11/11)   eng\Voice 1468 (9/9)   eng\Briefing 1392 (2/2)
jpn\Voice 1600 (12/13) jpn\etc 1468 (8/12), 1600 (4)

Note jpn\etc splits, so the path alone is not enough to pick the offset.

Picking the offset without a decoder

An XMA1 packet opens with a big-endian header — 6 bits frame count, 15 bits frame-offset-in-bits, 3 bits metadata, 8 bits packet-skip. At the true offset those fields stay in range packet after packet; one byte off and they do not. Scoring the first 24 packets and taking the best candidate:

7 330 of 7 358 (99.62 %) on the labelled set — every bank that has a RIFF, where the answer is forced and therefore known. All 28 misses are ties on the top score; there is not a single case where the scan picks wrongly with a unique winner. scan_data_offset therefore falls back to 1392 on a tie.

This is used only for the RIFF-less banks. Where a RIFF exists the offset is derived from it exactly, never scanned.

A second, independent signal — and it breaks the ties

Settled 2026-08-26. The 28 ties needed a different signal, not more of the same one, and the banks carry one: a seek chunk sitting on a packet boundary. Its position modulo 2048 therefore is the data offset.

seek at 3 516 / 5 564 / 7 612 / 9 660 / 13 756 / 19 900  —  all ≡ 1468 (mod 2048)

On the 6 033 labelled banks that have a seek before their first RIFF, 6 031 agree (99.97 %) and 2 disagree. That is better than the packet scan and, more importantly, structural rather than statistical — which is why it is now tried first.

Applied to the packet scan's 28 ties: 26 resolved correctly, 0 wrongly, and 2 with no usable seek. The combined rule — seek residue, else packet plausibility, else 1392 — scores 7 354 / 7 358 = 99.95 % on the labelled set, up from 99.62 %.

762 of the 1 495 RIFF-less banks carry a seek, and its residue lands on the four known offsets there too (1468 ×343, 1600 ×255, 1392 ×148, 1728 ×16), so the signal is available in the population that needs it.

The header is not audio being discarded

Worth ruling out, since a wrong data offset was the whole subject of this page: if the bytes before the offset were audio, we would be throwing away the start of every clip. Adding 0 to the candidate set and re-running the scan, it wins 6 of 7 358 — noise. The header is genuinely not part of the packet stream. (1 482 banks have an all-zero header; 5 876 have content in it, which is what prompted the check.)

Is that 99.62 % transferable? — checked, and it is conservative

The labelled set has a RIFF; the population the scan actually serves does not. Since the scan is unbounded it reads past the RIFF on labelled banks, so the 99.62 % could have been borrowing discriminating power that a RIFF-less bank cannot offer. That would make the headline number optimistic for the only case it is used in — worth checking before trusting it.

Confining the scan to the leading region drops it to 69.98 % with 1 910 ties, which at first looks like exactly that problem. It is not. Splitting by how much leading audio there is separates the two explanations:

correct ties
unbounded, all 7 358 labelled banks 99.62 % 28
confined to the leading region, all 7 358 69.98 % 1 910
≥24 packets of leading audio (989 banks), unbounded 100 % 0
≥24 packets of leading audio (989 banks), confined 100 % 0

The last two rows settle it. Where there is enough audio to score, the discriminator is perfect whether or not the RIFF is in range — so it is not leaning on the RIFF. The 69.98 % is an artifact of short leading regions: with only two or three packets to judge, candidates tie and the tie-break decides. Unboundedness helps those banks by giving the scan more bytes, which is why the two columns differ at all.

A RIFF-less bank is a whole pak entry, tens of kilobytes, so 24 packets are always available — it is always in the 100 % regime. The 99.62 % figure is therefore conservative for the population the scan is used on, not optimistic.

What this does not settle

  • Why the offset takes those four values, and what the bytes before it are. This was probed and remains open; what is now ruled out is recorded below.

  • The 28 ties. The scan cannot separate them and falls back to 1392, which is right for roughly a third of that population and wrong for the rest.

  • Why the offset takes exactly these four values by directory is still unexplained — see above.

  • Nothing here was run in the game — this is a decoder-side result measured with FFmpeg as the oracle.

🟡 Most banks declare more data than they store

Measured 2026-08-26. Of the 7 586 banks that carry both a RIFF and a data chunk after it, 5 296 (69.8 %) declare a data size larger than the bytes actually present in the pak entry. The remaining 2 290 declare less, which is the ordinary multi-sub-wave case. Not one declares exactly what it holds.

Worst cases run to 99 % short:

eng\Movie\VOICE_RT16C.slb    declared 1 810 432   available 489 392   -73 %
jpn\etc\VOICE_D_589.slb      declared 1 177 600   available   6 708   -99 %
eng\Voice\VOICE_TCAF_608.slb declared   759 808   available   8 864   -99 %

This contradicts a claim in the decoder's own comment, which says the declared size "is honest per sub-wave". It is not, for about seven banks in ten. The code is nonetheless safe — it clamps the range with .min(slb.len()) — so this is a documentation defect and an integrity observation, not a crash.

⚠️ Method note on this measurement. My first pass searched for data from offset 0, which can hit those four bytes by chance inside the leading audio region and read a garbage length. Re-running it anchored after the first RIFF changed the count from 5 038 to 5 296 — the flaw was slightly under-counting, but it could as easily have gone the other way, and an unanchored chunk search over binary audio is not a safe way to ask this question.

Why the declared sizes are too large is not settled. Plausible readings — an authoring-time allocation that was never trimmed, or deliberate truncation of unused tails — are guesses; nothing here distinguishes them, and the game has not been observed reading one of these banks.

What the header is — four things it is not

The bytes before the data offset are still unexplained, but the field has been narrowed. Probing the header of banks at each of the four offsets:

  • Not a length field. There is no word in the first 64 bytes equal to the offset, the offset minus 1392, the RIFF position or the entry size, in either endianness. The offset has to be derived; it is not read.
  • Not a seek table or any ascending index. Treated as big-endian words, only about half of consecutive pairs are non-decreasing — which is what random data gives. Every word is distinct and none is zero, across all four offsets.
  • Not zero padding, at least not usually: 1 482 of 7 358 banks have an all-zero header, but 5 876 have content in it.
  • Not audio being discarded. Adding 0 to the offset candidates, it wins 6 of 7 358 — noise. (Recorded above.)

So it is high-entropy content of a size that is constant per language and subdirectory, carrying no field that names its own length. That combination suggests something the loader knows the size of a priori rather than something self-describing.

First step if this is picked up again: find the loader. SETTINGS.PATH is game:\dat\sound.pak+ and SETTINGS.PARAM is Pj_Silph.xgs, so there is code that opens a bank by name and seeks to its data; the constant, or the table it indexes, should be visible there. That is static PE work (/work/*.pe, offset = VA 0x82000000), not another pass over the archive — this page has taken the byte-level evidence about as far as it goes.

⚠️ Disagreement on the declared-data question — not resolved

auto/slb-loader withdraws the 🟡 finding above that 69.8 % of banks declare more data than they store, reporting instead that declared sizes are exact (260/260 checked) and that the extra bytes live outside the TOC window but still in the .pNN stream — and specifically that VOICE_TCAF_608 is not truncated.

I could not reproduce that, and the arithmetic is against it. Walking that bank's chunks gives a clean, internally consistent structure:

window = [662 403 072, 662 456 412)   size 53 340
RIFF at +40 380, its size field 761 360
  fmt   32
  Dmmy  4 028
  data  759 808          <-- declared

The next TOC entry begins at 662 458 368, i.e. 55 296 bytes after this one starts. 759 808 bytes of audio cannot fit there. They would have to span roughly fourteen further TOC windows.

Both readings agree on the underlying fact — the declared size exceeds the TOC window — and differ on what follows from it. Mine said "truncated", which was an over-claim I withdraw: the 1 928 non-zero bytes in the 1 956-byte gap after the window, and the audio-looking bytes at the next entry, are consistent with a bank's data simply continuing past its window. But "declared sizes are exact" requires a wave to span many named entries, which is a much stronger claim than "the bytes are outside the window".

Unresolved. The other branch's own open question — "which bank in a window belongs to the entry's name (some windows hold two, ids drift)" — is the same question from the other side, and settling it is what would decide this. First step: take VOICE_TCAF_608, read 759 808 bytes from its data chunk straight out of the flat stream ignoring window boundaries, and decode. If it yields ~8 s of continuous speech, the data really does span windows; if it turns to noise at the window edge, it does not.