Files
Sylpheed/docs/re/structures/slb-data-offset.md
Sylpheed RE agent aa84aaf53f slb: read Channels instead of assuming mono -- and retract the TCAF_608 conclusion
I read the seek chunk's packet count big-endian; it is little-endian at seek+12,
with size == 8 + 4*count. And a seek sits immediately AFTER its own data, so an
entry's first seek usually belongs to the PREVIOUS bank (implied start -25232
for D_452, -145988 for TCAF_608). I was comparing an entry's first seek against
its first data -- different waves by construction, which is why no reading lined
up.

With that fixed, the declared sizes are honest: every RIFF-bearing entry on the
disc has seek magic at exactly data_at + declared_size with count*2048 ==
declared. 7620/7620, zero failures.

VOICE_TCAF_608 is not truncated. Its Channels is 2 and I decoded it as mono;
read as stereo it gives 6520176 bytes = 33.96 s, agreeing with both length
signals in the bank (33.88 s from cumulative samples, 33.97 s from
PsuedoBytesPerSec). 170 of 8021 banks (2.12%) are stereo -- exactly the rate of
my 1-in-60 outlier.

This is the mono/stereo trap already documented on this very page, met from the
other direction: I had written 'at two channels every bank yields one frame' and
then spent several passes blaming missing data for a one-frame decode.

Code fix: to_xma_riffs built the leading segment with a hard-wired mono fmt. It
now reads Channels from the bank's first RIFF. 7 disc tests pass.
2026-08-26 05:27:42 +00:00

506 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
> ✅✅ **SUPERSEDED IN ITS CAUSE (2026-08-26) — see branch `auto/slb-loader`.**
> The four offsets are **not a header size**. They are a *packing phase*:
>
> X = (cumulative start of the .pNN segment holding the bank) mod 2048
>
> | segment | size | cumulative start | start mod 2048 |
> |---|---|---|---|
> | `sound.p00` | 267 930 992 | 0 | **0** |
> | `sound.p01` | 268 404 812 | 267 930 992 | **1392** |
> | `sound.p02` | 268 404 868 | 536 335 804 | **1468** |
> | `sound.p03` | 268 384 384 | 804 740 672 | **1600** |
> | `sound.p04` | 14 903 296 | 1 073 125 056 | **1728** |
>
> Independently reproduced here: the four values I measured are exactly the
> running sums of the five segment file sizes, mod 2048. The XMA grid is
> 2048-aligned *inside each `.pNN` file*; the segments are not multiples of 2048
> long; so every join shifts the phase, and the flat concatenation the TOC
> addresses inherits the shift.
>
> Two things I wrote on this page are therefore wrong in their explanation, even
> though the measurements stand:
>
> * **"It varies by language and subdirectory"** — that was a *correlation*, not
> a cause. Directories cluster into segments, so the per-directory table is
> real but explains nothing.
> * **"The header is high-entropy content of a size the loader must know a
> priori"** — there is no header. Those bytes are the **previous bank's
> audio**, which is why they looked like data and had no length field: they
> are data.
>
> The heuristics below (99.62 % packet scan, 99.97 % `seek` residue, 99.95 %
> combined) are all superseded by an exact rule, verified 8 783/8 783 on the
> other branch. `slb.rs` still uses the heuristics; the exact fix needs
> `PakArchive` to expose the segment phase, which is an API change.
# `.slb` leading-stream data offset — 1392 was never a constant
**✅ Settled 2026-08-26, verified by decoding.** A bank's leading headerless
packet stream does not start at a fixed offset. It starts at
**`first_riff % 2048`**. `HEADERLESS_DATA_OFFSET = 1392` is the value that
offset happens to take in `<lang>\etc\`, and assuming it everywhere starts the
decode mid-packet and throws away almost all of the audio.
## The rule
XMA1 packets are 2048 bytes and the leading stream is a whole number of them
ending at the first `RIFF`. So its start is forced:
start = first_riff % 2048
Disc-wide that lands on exactly **four** values — 1392, 1468, 1600, 1728 — all
of the form `1392 + 4k`. Across the 3 965 Japanese and 3 393 English banks with
a non-empty leading region, no other value occurs:
| | 1392 | 1468 | 1600 | 1728 |
|---|---|---|---|---|
| `eng\etc`, `eng\Movie`, `eng\Briefing` | 1 520 | — | — | — |
| `eng\Voice` | 8 | 1 873 | — | — |
| `jpn\etc` | — | 1 402 | 303 | — |
| `jpn\Briefing` | — | 71 | — | — |
| `jpn\Movie` | — | — | 61 | — |
| `jpn\Voice` | — | — | 2 033 | 95 |
It varies by **language and subdirectory**, which is why a constant derived
from `eng\etc\` looked right for years' worth of the banks anyone had reason to
open.
## Verified by decoding, not by arithmetic
The alignment argument alone proves nothing — any offset can be made to "align"
by definition. The test is whether more audio comes out. Decoded through
FFmpeg's `xma1` at mono/48 kHz, on a random sample of **140** banks that have a
non-empty leading region:
| outcome | banks |
|---|---|
| more audio at `ri % 2048` | **85** |
| byte-identical | 54 |
| less audio | **1** |
Median gain among the improved: **70×**. The 54 identical ones are the control —
they are the `eng\etc`-style banks where `ri % 2048` *is* 1392, so the rule
must and does reproduce the old behaviour exactly. Individual cases:
eng\Voice\VOICE_TCAF_592.slb 1 506 -> 97 152 bytes (65x)
jpn\Voice\VOICE_TCAF_592.slb 2 910 -> 127 178 bytes (44x)
eng\etc\VOICE_D_452.slb 30 154 -> 30 154 bytes (unchanged, control)
## The one counterexample — ✅ explained
`eng\Voice\VOICE_TCAF_608.slb` decodes 2 840 bytes at 1392 and 896 at 1468.
It is **not** a bank where the old constant works and the derived offset fails:
both offsets yield well under a tenth of a second from a 38 988-byte region,
i.e. both fail, and 1392 merely produces marginally more garbage.
**The reason is that the bank is truncated.** Its `data` chunk declares 759 808
bytes and the pak entry holds 8 864 — **99 % short**. There is almost nothing
there to decode at any offset. See the section below.
## ❌ This withdraws my own claim from earlier the same day
[`sound-pak-contents.md`](sound-pak-contents.md) reported that the leading
region rule holds for "0 of 5 100 Japanese banks" and filed a backlog item
saying the Japanese banks were a different, undecoded layout. **That was wrong.**
The Japanese banks are the same format; only the offset differs. The measurement
behind it was correct — zero of them satisfy `(riff 1392) % 2048 == 0` — but
the conclusion drawn from it was not, and the reason is instructive: I treated
`HEADERLESS_DATA_OFFSET` as a property of the format when it was a property of
the sample the format was derived from.
The same error was hiding a defect in the **English** set too: 1 873 `eng\Voice`
banks sit at 1468 and were being decoded mid-packet just as badly.
## The `RIFF`-less banks had the same bug, plus a worse one
**✅ Settled 2026-08-26.** 1 495 banks (799 `jpn`, 696 `eng`) carry no `RIFF` at
all and take a separate code path. That path was wrong twice over:
1. it used the constant offset, with no `RIFF` to derive from; and
2. it built a **stereo** `fmt` chunk.
Decoded across a random 48-bank sample:
| | |
|---|---|
| banks where the old stereo-at-1392 beat the best mono offset | **0 of 48** |
| median gain | **184×** |
| range | 25× 489 344× |
Stereo is the same failure signature recorded for the leading segment: it stops
after one frame. Individual banks went from 04 816 bytes to 180 000380 000.
The winning offsets fall out **by directory**, and they reproduce the
distribution measured independently from the `RIFF`-bearing banks — which is the
cross-check that makes this more than curve-fitting:
eng\etc 1392 (11/11) eng\Voice 1468 (9/9) eng\Briefing 1392 (2/2)
jpn\Voice 1600 (12/13) jpn\etc 1468 (8/12), 1600 (4)
Note `jpn\etc` splits, so the **path alone is not enough** to pick the offset.
### Picking the offset without a decoder
An XMA1 packet opens with a big-endian header — 6 bits frame count, 15 bits
frame-offset-in-bits, 3 bits metadata, 8 bits packet-skip. At the true offset
those fields stay in range packet after packet; one byte off and they do not.
Scoring the first 24 packets and taking the best candidate:
**7 330 of 7 358 (99.62 %)** on the labelled set — every bank that *has* a
`RIFF`, where the answer is forced and therefore known. All **28** misses are
ties on the top score; there is not a single case where the scan picks wrongly
with a unique winner. `scan_data_offset` therefore falls back to 1392 on a tie.
This is used only for the `RIFF`-less banks. Where a `RIFF` exists the offset is
derived from it exactly, never scanned.
### ✅ A second, independent signal — and it breaks the ties
**Settled 2026-08-26.** The 28 ties needed a different signal, not more of the
same one, and the banks carry one: a **`seek` chunk sitting on a packet
boundary**. Its position modulo 2048 therefore *is* the data offset.
seek at 3 516 / 5 564 / 7 612 / 9 660 / 13 756 / 19 900 — all ≡ 1468 (mod 2048)
On the 6 033 labelled banks that have a `seek` before their first `RIFF`,
**6 031 agree (99.97 %)** and 2 disagree. That is better than the packet scan
and, more importantly, *structural* rather than statistical — which is why it is
now tried first.
Applied to the packet scan's 28 ties: **26 resolved correctly, 0 wrongly**, and
2 with no usable `seek`. The combined rule — `seek` residue, else packet
plausibility, else 1392 — scores **7 354 / 7 358 = 99.95 %** on the labelled set,
up from 99.62 %.
762 of the 1 495 `RIFF`-less banks carry a `seek`, and its residue lands on the
four known offsets there too (1468 ×343, 1600 ×255, 1392 ×148, 1728 ×16), so the
signal is available in the population that needs it.
### ❌ The header is not audio being discarded
Worth ruling out, since a wrong data offset was the whole subject of this page:
if the bytes *before* the offset were audio, we would be throwing away the start
of every clip. Adding **0** to the candidate set and re-running the scan, it wins
**6 of 7 358** — noise. The header is genuinely not part of the packet stream.
(1 482 banks have an all-zero header; 5 876 have content in it, which is what
prompted the check.)
### Is that 99.62 % transferable? — checked, and it is conservative
The labelled set has a `RIFF`; the population the scan actually serves does not.
Since the scan is unbounded it reads *past* the `RIFF` on labelled banks, so the
99.62 % could have been borrowing discriminating power that a `RIFF`-less bank
cannot offer. That would make the headline number optimistic for the only case
it is used in — worth checking before trusting it.
Confining the scan to the leading region drops it to **69.98 %** with 1 910
ties, which at first looks like exactly that problem. It is not. Splitting by
how much leading audio there is separates the two explanations:
| | correct | ties |
|---|---|---|
| unbounded, all 7 358 labelled banks | 99.62 % | 28 |
| confined to the leading region, all 7 358 | 69.98 % | 1 910 |
| **≥24 packets of leading audio (989 banks), unbounded** | **100 %** | **0** |
| **≥24 packets of leading audio (989 banks), confined** | **100 %** | **0** |
The last two rows settle it. Where there is enough audio to score, the
discriminator is perfect **whether or not the `RIFF` is in range** — so it is
not leaning on the `RIFF`. The 69.98 % is an artifact of *short* leading
regions: with only two or three packets to judge, candidates tie and the
tie-break decides. Unboundedness helps those banks by giving the scan more bytes,
which is why the two columns differ at all.
A `RIFF`-less bank is a whole pak entry, tens of kilobytes, so 24 packets are
always available — it is always in the 100 % regime. **The 99.62 % figure is
therefore conservative for the population the scan is used on**, not optimistic.
## What this does not settle
***Why the offset takes those four values**, and what the bytes before it
are. This was probed and remains open; what is now ruled out is recorded
below.
* **The 28 ties.** The scan cannot separate them and falls back to 1392, which
is right for roughly a third of that population and wrong for the rest.
* **Why the offset takes exactly these four values by directory** is still
unexplained — see above.
* Nothing here was run **in the game** — this is a decoder-side result measured
with FFmpeg as the oracle.
## 🟡 Most banks declare more `data` than they store
**Measured 2026-08-26.** Of the 7 586 banks that carry both a `RIFF` and a
`data` chunk after it, **5 296 (69.8 %)** declare a `data` size larger than the
bytes actually present in the pak entry. The remaining 2 290 declare *less*,
which is the ordinary multi-sub-wave case. **Not one declares exactly what it
holds.**
Worst cases run to 99 % short:
eng\Movie\VOICE_RT16C.slb declared 1 810 432 available 489 392 -73 %
jpn\etc\VOICE_D_589.slb declared 1 177 600 available 6 708 -99 %
eng\Voice\VOICE_TCAF_608.slb declared 759 808 available 8 864 -99 %
This **contradicts a claim in the decoder's own comment**, which says the
declared size "is honest per sub-wave". It is not, for about seven banks in ten.
The code is nonetheless safe — it clamps the range with `.min(slb.len())` — so
this is a documentation defect and an integrity observation, not a crash.
⚠️ **Method note on this measurement.** My first pass searched for `data` from
offset 0, which can hit those four bytes by chance inside the leading audio
region and read a garbage length. Re-running it anchored *after* the first
`RIFF` changed the count from 5 038 to 5 296 — the flaw was slightly
*under*-counting, but it could as easily have gone the other way, and an
unanchored chunk search over binary audio is not a safe way to ask this
question.
**Why** the declared sizes are too large is **not settled**. Plausible
readings — an authoring-time allocation that was never trimmed, or deliberate
truncation of unused tails — are guesses; nothing here distinguishes them, and
the game has not been observed reading one of these banks.
## ❔ What the header is — four things it is *not*
The bytes before the data offset are still unexplained, but the field has been
narrowed. Probing the header of banks at each of the four offsets:
* **Not a length field.** There is no word in the first 64 bytes equal to the
offset, the offset minus 1392, the `RIFF` position or the entry size, in
either endianness. The offset has to be derived; it is not read.
* **Not a seek table or any ascending index.** Treated as big-endian words, only
about half of consecutive pairs are non-decreasing — which is what random data
gives. Every word is distinct and none is zero, across all four offsets.
* **Not zero padding**, at least not usually: 1 482 of 7 358 banks have an
all-zero header, but **5 876 have content** in it.
* **Not audio being discarded.** Adding 0 to the offset candidates, it wins 6 of
7 358 — noise. (Recorded above.)
So it is high-entropy content of a size that is constant per language and
subdirectory, carrying no field that names its own length. That combination
suggests something the *loader* knows the size of a priori rather than something
self-describing.
**First step if this is picked up again**: find the loader. `SETTINGS.PATH` is
`game:\dat\sound.pak+` and `SETTINGS.PARAM` is `Pj_Silph.xgs`, so there is code
that opens a bank by name and seeks to its data; the constant, or the table it
indexes, should be visible there. That is static PE work
(`/work/*.pe`, offset = VA 0x82000000), not another pass over the archive —
this page has taken the byte-level evidence about as far as it goes.
## ⚠️ Disagreement on the declared-`data` question — not resolved
`auto/slb-loader` **withdraws** the 🟡 finding above that 69.8 % of banks declare
more `data` than they store, reporting instead that declared sizes are exact
(260/260 checked) and that the extra bytes live outside the TOC window but still
in the `.pNN` stream — and specifically that `VOICE_TCAF_608` is not truncated.
**I could not reproduce that, and the arithmetic is against it.** Walking that
bank's chunks gives a clean, internally consistent structure:
window = [662 403 072, 662 456 412) size 53 340
RIFF at +40 380, its size field 761 360
fmt 32
Dmmy 4 028
data 759 808 <-- declared
The next TOC entry begins at 662 458 368, i.e. **55 296 bytes** after this one
starts. 759 808 bytes of audio cannot fit there. They would have to span roughly
fourteen further TOC windows.
Both readings agree on the underlying fact — **the declared size exceeds the
TOC window** — and differ on what follows from it. Mine said "truncated", which
was an over-claim I withdraw: the 1 928 non-zero bytes in the 1 956-byte gap
after the window, and the audio-looking bytes at the next entry, are consistent
with a bank's data simply continuing past its window. But "declared sizes are
exact" requires a wave to span many named entries, which is a much stronger
claim than "the bytes are outside the window".
### ✅ The experiment was run — the data does not span windows
Read `VOICE_TCAF_608`'s declared 759 808 bytes straight out of the flat stream,
ignoring the window boundary, and decoded through FFmpeg's `xma1`:
| read | bytes in | decoded |
|---|---|---|
| `data` → window end | 8 864 | 896 bytes = **0.01 s** |
| `data` → full declared size | **759 808** | 896 bytes = **0.01 s** |
| control `VOICE_D_452`, declared fits | 26 624 | 208 970 bytes = **2.18 s** |
**Reading 86× more bytes yields not one extra sample.** The bytes past the
window are not this bank's audio, so "the bytes are outside the TOC window but
still in the `.pNN` stream" does not hold here — and this is the very bank the
other branch named as *not* truncated. (No segment join is crossed: the read
sits inside `sound.p02`, so the packing phase is not a confound.)
So for `VOICE_TCAF_608` the audio really is not present. I still withdraw the
word **"truncated"** as an over-claim about the other 5 295 banks — I measured
that their declared size exceeds their window, not what is in the bytes beyond
it, and I have now tested exactly one of them. What is established is narrower
and worth stating exactly:
* ✅ the declared `data` size exceeds the TOC window for 5 296 of 7 586 banks;
* ✅ for `VOICE_TCAF_608` the missing bytes cannot be recovered from the stream;
* ❔ whether that generalises is **untested** — the same decode would have to be
run across the population, which is the obvious next step and was not done.
## ❌ Decoded length is not a valid test of where a bank ends — my own test withdrawn
I set out to generalise the `VOICE_TCAF_608` result across the 5 296
over-declaring banks, and the first pass looked like a clean reversal: on a
random 60, reading the **full declared size** instead of stopping at the TOC
window gained audio in **59**, median **2.10×**, up to 59×. That reads as
"the data really does continue past the window", i.e. the other branch is right
and my truncation reading was wrong.
**Then I checked whether the declared size is an honest boundary at all**, by
reading *twice* it. If the stream ends where the header says, doubling the input
should add little:
| | |
|---|---|
| reading 2× the declared size yields >1.5× the audio | **33 of 40** |
| ratio of decoded bytes, 2× input vs 1× | median **1.64**, p90 1.75 |
It keeps producing audio indefinitely. **XMA1 packets are self-contained**, so
feeding the decoder the *next* bank's packets yields perfectly good audio that
simply is not this bank's. The decoder cannot tell the difference, and neither
can a byte count.
**So the 59-of-60 result is withdrawn as evidence.** It does not show that the
declared size is honest; it shows only that *something* decodes after the window,
which was never in doubt — the bytes there are audio, just possibly someone
else's. This is the same error in a new costume as the unanchored `data` search
earlier on this page: a measurement that returns a plausible number for a
question it cannot actually answer.
What survives:
*`VOICE_TCAF_608` is still special, and now more clearly so: for it, reading
86× more bytes gained **nothing**, where the typical over-declaring bank keeps
yielding audio without limit. Whatever is at its offset does not decode at all.
***Whether the declared sizes are honest is unresolved by this method** and
cannot be resolved by it. A valid test has to identify the bank boundary
independently of the decoder — the `seek`-chunk packet-count chain used on
`auto/slb-loader` is exactly such a signal, and is the right next step.
* ❌ Both my "69.8 % are truncated" and my attempted reversal of it are off the
table. The measured fact is unchanged and narrow: **the declared `data` size
exceeds the TOC window for 5 296 of 7 586 banks.**
## ❔ The `seek` chunk's layout — identified, but it does not yield a packet count
The decoder-independent boundary signal this page called for is the `seek`
chunk. Its **shape** is now readable; its **arithmetic** is not.
Immediately after the `seek` tag sits a little-endian size, then a short header,
then a strictly ascending table:
eng\etc\VOICE_D_452 seek at +3 440 size 64 -> 16 words
eng\Voice\VOICE_TCAF_592 seek at +9 660 size 108 -> 27 words
eng\Voice\VOICE_TCAF_608 seek at +28 092 size 348 -> 87 words
words after seek+8 (big-endian):
0x01000000, <varies>, 0, then ascending: 0, 1572864, 3932160, 6160384, ...
Word 0 is `0x01000000` in every bank examined — a version or entry-size marker.
Word 1 varies and its top byte is 14 / 25 / 85 for the three above. Word 2 is 0,
and the ascending run begins after it.
**Two readings tried, both fail:**
| reading | D_452 | TCAF_592 | TCAF_608 |
|---|---|---|---|
| declared `data` size ÷ 2048 (packets) | 13 | 17 | 371 |
| `seek` size ÷ 4 (entries) | 16 | 27 | 87 |
| entries 3 header words | **13 ✅** | 24 ✗ | 84 ✗ |
The third row is the near-miss that would be easy to adopt: it is exactly right
for `VOICE_D_452` and wrong for the other two. That is a one-of-three fit, and
this page has already recorded two measurements today that returned plausible
numbers for questions they could not answer — so it is recorded as **failed**,
not as a rule with exceptions.
The ascending values are not packet indices: the steps (≈1.52.2 million) are far
too large for a bank of a few hundred packets, so they are sample or
fixed-point offsets on some other scale that has not been identified.
**What the next attempt should know**: the chunk is real, its size field is
little-endian, three header words precede the table, and the entry count is
*not* the packet count. `auto/slb-loader` reports chaining `seek` packet counts
successfully across consecutive entries — whatever field it used is not one of
the two tried here, and reconciling the two readings is the cheapest way in.
## ❌❌ `VOICE_TCAF_608` is NOT truncated — I was decoding it as mono
**2026-08-26, resolving the disagreement above in the other branch's favour.**
Everything I concluded about this bank was an artefact of a wrong `fmt ` chunk,
and the declared sizes are honest after all.
**The `seek` chunk, read correctly.** I read the packet count big-endian. It is
**little-endian**, and the layout is:
+0 'seek'
+4 u32 LE chunk size (always 8 + 4*packets)
+8 u32 LE stream count (always 1) <-- my "0x01000000" was LE 1 here
+12 u32 LE PACKET COUNT <-- my "<varies>" read big-endian
+16 packets x u32 LE cumulative decoded sample totals
Verified: `size == 8 + 4*count` on every bank checked. My "entries 3" reading
matched `VOICE_D_452` by coincidence; the real relation is `size/4 2`.
**And a `seek` sits immediately *after* its own data**, so the *first* `seek` in
an entry usually belongs to the *previous* bank — its implied start is negative
(25 232 for D_452, 145 988 for TCAF_608). I was comparing an entry's first
`seek` against its first `data`, which are different waves by construction. That
is why no reading could line up.
**The declared sizes are honest — 7 620 / 7 620.** For every `RIFF`-bearing entry
on the disc there is `seek` magic at exactly `data_at + declared_size`, and its
packet count × 2048 equals the declared size. **Zero failures.** For TCAF_608:
probe at 663 207 356 → `seek`, count **371**, 371 × 2048 = **759 808** = declared.
**Why it decoded to 896 bytes:** its `Channels` is **2**. I decoded it as mono.
Reading it as stereo gives **6 520 176 bytes = 33.96 s** — and two independent
length signals in the bank agree: the last cumulative sample 1 626 112 / 48 000 =
33.88 s, and 759 808 / `PsuedoBytesPerSec` = 33.97 s. The audio was there the
whole time.
**170 of 8 021 banks (2.12 %) are stereo**`Channels` is the byte at
`RIFF + 49`. That is exactly the 1-in-60 rate of my "gains nothing" outlier.
### What this retracts
* ❌ "`VOICE_TCAF_608`'s missing bytes cannot be recovered from the stream" —
**wrong**, nothing was missing.
* ❌ "The declared `data` size exceeds the TOC window for 5 296 banks" stands as
a fact about the *window*, but my framing of it as a problem is withdrawn: the
window is simply not the wave boundary, and `data_at + declared_size` is.
* ❌ My `seek`-layout write-up above (entry count, "failed readings") was wrong
in its endianness and in its pairing assumption. Left in place as a record.
This is the mono/stereo trap **already documented on this page** — "at two
channels every bank yields exactly 1792 bytes, one frame" — met from the other
direction. Having written that down, I then spent several passes attributing a
one-frame decode to missing data instead of checking the channel count.
### Code fixed
`to_xma_riffs` built the leading segment with a hard-wired mono `fmt `. It now
reads `Channels` from the bank's own first `RIFF` (`riff_channels`), falling
back to mono only when there is no `RIFF` to read. 7 disc tests pass.
### The decoder-independent boundary, for the record
bytes = u32 LE at seek+12 x 2048 (== the `data` chunk size)
validate = 'seek' magic at data_at + declared_size (7 620/7 620)
samples = the LAST u32 LE entry in the seek table
channels = byte at RIFF + 49 <-- read it, never assume