This repository has been archived on 2026-09-16. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
Syplheed-Reborn/docs/re/structures/sound-pak-contents.md
Sylpheed RE agent 2320c56867 re: the seek table's sample total is not the wave length -- durations corrected
Chasing w0 I decoded leading waves and the output matched neither trailer word.
Following that into RIFF waves, where the extent is certain, the decoded sample
count exceeds the seek table's last cumulative sample by a median 9.7%.

The tiebreak is the bank's own PsuedoBytesPerSec: over 14 banks it agrees with
an actual FFmpeg decode to a mean of 0.007 s and with the seek-table duration
only to 0.287 s. Two independent quantities agree with each other and both
disagree with the seek total, so that last entry is the last SEEK POINT, not the
sample count.

Both artifacts regenerated with data_bytes / PsuedoBytesPerSec. Total audio is
408.3 minutes, not the 390.9 I published -- 4.3% in aggregate. Dialogue timings
all rise, e.g. 'They got Leader!' from 1.78/2.46 s to 2.25/2.93 s.

What made this hard to catch is worth recording: the seek total gave BGM of 2.4
minutes, chatter of 2.8 seconds and cutscenes of 11 minutes, and I cited that as
the chain validating itself. It validated the structure, not the scale -- a
uniform 10% error preserves every ratio I checked.
2026-08-26 06:17:07 +00:00

152 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# What is actually in `sound.pak` — and the language substitution
**✅ Settled 2026-08-26.** `sound.pak`'s 9 519 entries are almost entirely
accounted for: the Japanese voice set, the English voice set, and 35 music /
jingle banks. The `FILES` record in `tables.pak` names the Japanese half, and
the game reaches the English half by **substituting the leading directory**.
Until now the corpus described `sound.pak` as "~9 519 headerless entries, params
in an unknown table". The table is
[`sound-cue-table.md`](sound-cue-table.md), and this page is the census.
## The name is the key
Every path in `FILES` hashes straight into the `sound.pak` TOC with
`name_hash` — **5 135 of 5 135**, using the path exactly as stored
(backslashes, case-insensitive):
jpn\etc\VOICE_D_257.slb -> name_hash -> TOC entry (offset, size)
No transformation, no prefix, no stem. Taking the basename instead resolves only
35 — precisely the 35 root-level `BGM_*` / `JNGL_*` banks whose basename *is*
their path, which is the control that shows the full path is what matters.
This is not luck: the hash's low 24 bits are a residue mod `0x00FFF9D7` and the
top byte is a byte-sum, so for 5 135 probes against 9 519 occupied slots the
expected number of accidental hits is about **0.01**.
## The census
| what | entries |
|---|---|
| `jpn\…` — named directly by `FILES` | 5 100 |
| `eng\…` — the same names with the directory substituted | 4 382 |
| `BGM_*.slb`, `JNGL_*.slb` at the root | 35 |
| `static.slb` and `Pj_Silph.xgs` — named by `BANK_SE` / `SETTINGS` | 2 |
| **total accounted** | **9 519** |
| TOC total | **9 519** |
**The archive is now fully accounted for.** The last two took a correction:
I first reported them as unidentified, having looked past the fact that the
same IDXD object names them. `BANK_SE`'s single field is `Static.slb`
(8 970 240 bytes — the shared sound-effect bank) and `SETTINGS.PARAM` is
`Pj_Silph.xgs` (533 bytes — the XACT global settings). Both hash into the TOC,
and both were sitting in the table I had already printed at the top of this
page. For 18 candidate names against 9 519 occupied slots the expected number of
accidental matches is about 4 × 10⁻⁵, and these two are corroborated by being
named in the table rather than guessed.
Only `eng` and `jpn` resolve. `deu`, `fra`, `ita`, `esp` and the alternative
spellings `fre`/`ger`/`spa` produce **zero** hits, which is the measurement
behind the claim that only two voice sets are shipped — the other four languages
are subtitle-only.
**718 of the 5 100 Japanese banks have no English counterpart.** That is
consistent with the previously recorded finding that a set of lines exists only
undubbed, and it puts a number on it from the archive side.
> ❌ **CORRECTED the same day — see [slb-data-offset.md](slb-data-offset.md).**
> The section below measures correctly but concludes wrongly. The Japanese banks
> are **not** a different layout: the leading-stream offset is
> `first_riff % 2048`, which is 1392 only in `<lang>\etc\`. Every "misaligned"
> row here is a bank being decoded mid-packet, in English as well as Japanese.
> The numbers stand as a measurement of the *old constant's* reach; the
> interpretation does not.
## ⚠️ The leading-region rule is English-only
This is the part that matters most, and it bounds an existing result.
[`voice-bank-leading-region.md`](../voice-bank-leading-region.md) established
the hybrid layout: a bank may carry a leading headerless XMA1 region before its
first `RIFF`, recognised by the first `RIFF` sitting exactly at
`1392 + n·2048`. That was derived from `VOICE_D_450`–`454`, and it reproduces
**exactly** — all five, at 16 384 / 2 048 / 14 336 / 45 056 / 59 392 bytes,
matching the recorded figures to the byte. The pipeline is sound.
But those five are `eng\etc\` banks, and the rule does not generalise:
| set | aligned | misaligned | no `RIFF` | first `RIFF` ≤ 1392 |
|---|---|---|---|---|
| `eng\…` (4 382) | **1 571** | 2 060 | 696 | 55 |
| `jpn\…` (5 100) | **0** | 4 301 | 799 | — |
**Not one of the 5 100 Japanese banks satisfies the alignment condition.** The
rule describes a third of the English set and none of the Japanese set.
Being careful about what that does and does not mean: "misaligned" only means
the leading-segment branch does not fire, and `to_xma_riffs` still walks the
bank's sub-waves normally — so this is not evidence that those banks are
broken. What it does mean is that **the leading-region finding is a property of
some English banks, not of `.slb` in general**, and any code or documentation
that treats it as the bank format is overreaching. The Japanese banks are
laid out differently and that layout is **not decoded**.
## How I nearly got this wrong
I first resolved the five documented banks through their `FILES` paths, which
are `jpn\etc\…`, and got leading regions of 59 600 / — / 6 352 / 14 544 /
28 880 — none a multiple of 2 048, none matching the recorded figures. The
tempting reading was that my extraction was broken. The alignment condition is
what said otherwise: a broken reader would not produce *exactly* the recorded
numbers for `eng\etc\`, which it then did, for all five. The extraction was
right and the assumption that `FILES`' path was the one the corpus had measured
was wrong.
## What this does not settle
* **The 2 unaccounted TOC entries.**
* **The Japanese bank layout.** 4 301 banks have a `RIFF` at an unaligned
offset and 799 have none at all; neither shape has been decoded, and no
Japanese bank has been decoded to audio here.
* **Why 55 English banks have their first `RIFF` at or before 1392**, inside
what the headerless reading treats as the header.
* **Nothing here was run.** This is a static census of the archive.
## ✅ A full bank manifest — and the durations validate the whole chain
**2026-08-26.** With the wave boundary exact and `Channels` read rather than
assumed, every field needed to describe a bank can be *read*:
path | cue | sound id | channels | rate | data bytes | packets | samples | seconds
Committed as [`../data/voice-bank-manifest.txt`](../data/voice-bank-manifest.txt)
— **4 114 banks**, every one carrying a full `RIFF`/`fmt `/`data`/`seek`
structure. (1 021 of the 5 135 `FILES` paths have no `RIFF` and are excluded;
those are the headerless banks.) Duration comes from the **last cumulative
sample** in the `seek` table divided by the sample rate — arithmetic, no
decoding.
**Total: 408.3 minutes — about 6¾ hours of audio.** ⚠️ An earlier version of this page said 390.9 minutes, computed from the `seek` table's sample total; that understates the true length by a median 9.7 % per bank. See [slb-data-offset.md](slb-data-offset.md).
The reason this is worth more than a listing is that nothing in it was fitted to
an expectation, yet every category lands where its content says it should:
| directory | banks | total | mean | longest |
|---|---|---|---|---|
| root (`BGM_*`, `JNGL_*`) | 34 | 82.8 min | **146.0 s** | 277.5 s |
| `Movie` | 62 | 81.4 min | **78.8 s** | **672.7 s** |
| `Briefing` | 72 | 12.9 min | **10.8 s** | 26.8 s |
| `etc` | 1 776 | 112.9 min | **3.8 s** | 136.8 s |
| `Voice` | 2 170 | 100.9 min | **2.79 s** | 7.3 s |
Music tracks average two and a half minutes; a cutscene track runs up to eleven
minutes; briefing lines are ten seconds; and in-mission radio chatter averages
**2.8 seconds and never exceeds 7.3**. If the boundary rule, the channel field
or the seek table were being misread, these numbers would not sort themselves
into those five shapes.
Two incidental facts fall out: **4 banks run at 44 100 Hz** where everything else
is 48 000, and the `BGM_*` tracks are the stereo ones.