Files
Sylpheed/docs/re/structures/idxd-container.md
Sylpheed RE agent 2498922e9a re: the '504 unnamed keys' is 504 entries, not 504 names -- it is 42 keys
Independently reproduced across all 33 paks: 7750 IDXD objects, 1485577 unnamed
field entries, 7094 distinct never-named keys splitting cleanly into 7052 in an
ordinal band (<=0x2198, 94.6% equal to their own field index) and 42 hash-shaped
(>=0x2677C), with ZERO keys in the gap between. The 42 carry exactly 504
entries -- six language copies of one object times two records.

So the preimage target was 42, not 504, and my earlier wording invited the
misreading. Cross-referenced to the idxd-unnamed-keys write-up, which shows the
42 belong to <lang>\script\ID.tbl and cannot be recovered from a 24-bit hash.
2026-08-26 03:40:55 +00:00

170 lines
8.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The IDXD container — record/field table
Status: ✅ **CONFIRMED**, decoded in full and verified over **the whole disc**
(not just `dat/``hidden/DefTables.pak` holds another 1425 objects, which the
first version of this sweep missed): **7750 / 7750**
objects parse, **190,782 / 190,782** records reproduce their stored name hash,
**1,271,462 / 1,271,462** named fields reproduce their key from their stored name.
Zero failures of any kind.
This closes the **"binary node/index region — not yet decoded"** note that stood
in `crates/sylpheed-formats/src/idxd.rs` for the whole life of the parser, and it
**demotes two beliefs the corpus was built on** (see *Two corrections* below).
Tests: `crates/sylpheed-formats/tests/idxd_records_disc.rs`
`records_roundtrip_disc`, `first_header_word_is_record0_hash`,
`field_names_are_stored_disc`, `movie_table_ids_are_literal_field_keys`.
## ✅ The layout
All fields big-endian.
```text
Offset Size Field
0x00 4 "IDXD"
0x04 4 record_count n
0x08 16*n records { u32 name_hash, u32 name_off, u32 field_begin, u32 field_end }
.. 4 field_count m (equals max(field_end))
.. 12*m fields { u32 key, u32 name_off, u32 value_off }
.. 4 pool_size (equals file_len - pool_base)
.. .. string pool (every *_off above is a byte offset from here)
```
* **Records are sorted ascending by `name_hash`** and the guest binary-searches
them with a 16-byte stride (`sub_82448AA0`). Verified: sorted in every object.
* `name_hash` is [`tag_hash`](idxd-tag-hash.md) of the record's **own name** —
not the pak TOC hash (different modulus, and not lowercased).
* A record owns the half-open field range `[field_begin, field_end)`. Ranges are
*not* contiguous in record order — records are in hash order while their fields
sit in name order, so the field array is shared, not partitioned by position.
* **Fields are sorted ascending by `key`** and lower-bounded (`sub_8244E338`).
* `field_count` and `pool_size` are what make the layout self-checking: a wrong
record stride lands on a `pool_size` that does not equal the remaining bytes.
That identity is what caught the first wrong version of this decode.
## ✅ Field names are on the disc
A field's middle word is **its own name's pool offset**, and then
`key == tag_hash(name)`. Field names never have to be recovered from their hash.
`name_off == 0xFFFF_FFFF` means the field has **no name**; its `key` is then a
literal positional integer — a line-slot index (`0,1,2,3`), a movie id (`105`,
`1303`). So a field is positional *exactly* when it stores no name; you do not
have to guess from the key's magnitude.
Disc-wide census of all **2,757,039** fields:
| kind | count |
|---|---|
| named, `key == tag_hash(name)` | 1,271,462 |
| unnamed, literal key (`< 0x10000`) | 1,485,073 |
| unnamed, hash-shaped key | 504 |
Those **504** are the entire remaining preimage problem on the disc — but the
count is easy to misread, so state it precisely: 504 is a count of **field
entries**, not of distinct names. They are **42 distinct keys**, each appearing
in 12 places (six byte-identical language copies of one object × two records).
The preimage target was never 504 names; it is 42.
❔ Their names are unrecovered, and
[idxd-unnamed-keys.md](idxd-unnamed-keys.md) shows why that is not a matter of
more compute: the 42 belong to `<lang>\script\ID.tbl`, and at 7 characters one
of them already has 1 176 preimages under a 24-bit modulus. 30 of the 42 have
their trailing digits pinned algebraically; the prefixes cannot be recovered
from the hash alone.
## Two corrections
### ❌ WITHDRAWN — "the word at `0x08` is a schema hash"
The corpus (and `IdxdObject::schema_hash`, and every `pak list` line printing
`schema 067025b9`) read the third header word as an object-type id whose preimage
was unknown. **It is not a schema id.** It is simply **record 0's `name_hash`**:
the record array is uniform 16-byte entries starting at `0x08`, and the header has
no type field at all.
Evidence: `tag_hash(records[0].name) == word@0x08` for **all 7750** objects on the
disc, zero exceptions (`first_header_word_is_record0_hash`).
How it was caught, which is the useful part: a test asserted that every movie id in
`BASE_INFO` names a real record, and one — `1005 -> "STAGE10_PHASE01"` — did not
resolve. The reason was that `tag_hash("STAGE10_PHASE01")` **is** `0x067025B9`, the
movie table's supposed schema id. A "coincidence" at 1-in-2^32 is not a
coincidence; the record was being eaten by the header.
It still *works* as a type discriminator, because tables of one kind share their
lowest-hashed record name — which is exactly why it went unquestioned for so long.
`schema_hash` is therefore kept under its established name, with its docs corrected,
rather than renamed across 33 call sites. **Nothing on disc names an object's type.**
### ❌ WITHDRAWN — "the field's middle word is an `aux`/flags word, always `0xFFFFFFFF`"
It is `0xFFFFFFFF` for 54% of fields, which is enough to look constant in a small
sample. It is a name offset (above). The tell was that the word is constant *per
key across records* (`key=0x6c43a78d` always carried `1743`) — a per-record flag
cannot do that, but a per-name pointer must.
## Worked example — the movie table
`dat/tables.pak`, the object whose record 0 is `STAGE10_PHASE01` (105 records,
433 fields). Its `BASE_INFO` record mixes both field kinds:
```text
named: PATH = "dat\movie\" VERSION = "0x060329" SUBTITLE_FONT …
positional: 105 -> "STAGE01_PHASE01" 205 -> "STAGE02_PHASE01" 1005 -> "STAGE10_PHASE01"
```
Each positional value names another record in the same object, which carries the
real files:
```text
STAGE02_PHASE01: MOVIE = "RT02A.wmv" VOICETRACK = "VOICE_RT02A"
SUBTITLE = "…+SUBTITLE_RT02A.tbl" TELOP = "…+pwrt02.prt"
```
All 104 positional ids resolve to a real record — verified, no dangling entries.
This is the id space the mission script's cutscene request uses; see
[movie-subtitle-link](../movie-subtitle-link.md).
## ✅ `IXUD` is the same container
The wide-string sibling has an identical shape, with three substitutions: the hash
is `ixud_hash`, strings are UTF-16BE, and **every offset — record name, field name
and value alike — is in 16-bit chars**, so `pool_base + 2*off`. `pool_size` is
likewise a char count, which is the same identity as the already-known
`STR + 2*strsize == filesize` ([ixud-localised-text](ixud-localised-text.md)).
Verified independently over all **1104** IXUD objects on the disc: the header word
at `0x08` is record 0's `ixud_hash` (1104/1104), and `key == ixud_hash(field name)`
for **628,165 / 628,165** named fields, zero mismatches. Only 48 fields are
unnamed — 6 objects × 8, all in a `CATEGORY_DESC` record with keys `0..7`, a
positional array of weapon-category descriptions.
One extra rule shows up in IXUD's comparator (`sub_82447F38`) and is worth assuming
for IDXD too: `name_off == 0xFFFF_FFFF` is the **primary** sort key, so a record's
field slice is *unnamed fields first, then named fields*, each key-ascending
(1476/1476 slices). That is why there are two getters — lookup-by-integer searches
only the unnamed run, lookup-by-name only the named run.
🟡 The loader `sub_82448D00` also accepts legacy magics `IIDX`, `IDX2`, `IDX3`,
`IDXC` and rejects `IDXD` on that path with `"Old virsion binary table. Not
supported."` [sic]. **None of them occur on this disc** (magic census over every
pak entry: `IDXD` 7750, `T8aD` 4525, `RATC` 2985, `IXUD` 1104, `LSTA` 64, and zero
of the four legacy magics), so that path is read-only knowledge, untestable here.
## What this does *not* settle
* ❔ The names of the 42 unnamed hash-keyed fields (504 entries — see
[idxd-unnamed-keys.md](idxd-unnamed-keys.md); shown to be unrecoverable from
the hash alone).
* ❔ How much of the existing corpus the old string-pool reader got wrong. The
legacy `get_f32`/`get_raw` path infers a field's value from *pool adjacency*
(`<value>\0<key>\0`). That adjacency is a consequence of the field table, not a
rule of the format, and it cannot represent a field whose value string is shared
or reordered. Every number in `docs/re/` that came from it is now re-checkable
against the true table, and has not yet been re-checked.
* ❔ Whether `IXUD`, the wide-string sibling, uses the same shape. Its offsets are
in 16-bit chars and its hash is different ([ixud-localised-text](ixud-localised-text.md)).
* The type-identification question is now open rather than closed: with no schema
field, an object's kind is known only from the caller that loads it.