Files
Sylpheed/docs/re/structures/idxd-container.md
Sylpheed RE agent af32540190 re: decode the IDXD/IXUD record table — and there is no schema hash
The binary region in front of the string pool was the parser's oldest open
note ("Not yet decoded"). It is a uniform 16-byte record array sorted by
name hash, a field count, a 12-byte field array sorted by key, a pool size,
and the pool. The trailing `pool_size == file_len - pool_base` identity makes
the layout self-checking, which is what caught the first wrong version.

Verified over the WHOLE disc with zero failures: 7750/7750 IDXD objects,
190782/190782 records reproducing their stored tag_hash, 1271462/1271462
named fields reproducing their key. IXUD is the same container with
ixud_hash, UTF-16BE and every offset in chars — 1104/1104 objects,
628165/628165 fields, checked with an independent parser.

Field names are stored on disc, so no preimage search is needed: a field's
middle word points at its own name. Only 504 fields disc-wide are hash-keyed
with no name; the other 1485073 nameless fields are positional, keyed by a
literal integer (line slots, movie ids).

Two long-held beliefs are WITHDRAWN:

* The word at 0x08 is not a schema hash. It is record 0's name_hash — the
  format has no type field at all, and an object's kind is known only from
  the caller that loads it. It survived as "schema" because tables of one
  kind share their lowest-hashed record name. Caught by a test asserting
  every movie id names a real record: 1005 -> STAGE10_PHASE01 failed because
  tag_hash("STAGE10_PHASE01") IS 0x067025B9, that table's supposed schema id.
* The field's middle word is not an always-0xFFFFFFFF flags word. It is
  0xFFFFFFFF for 54% of fields, enough to look constant in a small sample;
  the tell was that it is constant per key ACROSS records, which a per-record
  flag cannot be but a per-name pointer must.

`schema_hash` keeps its name rather than churn 33 call sites, with corrected
docs. The first sweep globbed dat/** and missed hidden/DefTables.pak (1425
objects); the test now walks the whole disc root.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PMRJjbxLqZtsb5Vb7KunPE
2026-08-25 21:40:23 +00:00

7.8 KiB
Raw Blame History

The IDXD container — record/field table

Status: CONFIRMED, decoded in full and verified over the whole disc (not just dat/hidden/DefTables.pak holds another 1425 objects, which the first version of this sweep missed): 7750 / 7750 objects parse, 190,782 / 190,782 records reproduce their stored name hash, 1,271,462 / 1,271,462 named fields reproduce their key from their stored name. Zero failures of any kind.

This closes the "binary node/index region — not yet decoded" note that stood in crates/sylpheed-formats/src/idxd.rs for the whole life of the parser, and it demotes two beliefs the corpus was built on (see Two corrections below).

Tests: crates/sylpheed-formats/tests/idxd_records_disc.rsrecords_roundtrip_disc, first_header_word_is_record0_hash, field_names_are_stored_disc, movie_table_ids_are_literal_field_keys.

The layout

All fields big-endian.

Offset   Size    Field
0x00     4       "IDXD"
0x04     4       record_count  n
0x08     16*n    records      { u32 name_hash, u32 name_off, u32 field_begin, u32 field_end }
..       4       field_count  m       (equals max(field_end))
..       12*m    fields       { u32 key, u32 name_off, u32 value_off }
..       4       pool_size            (equals file_len - pool_base)
..       ..      string pool          (every *_off above is a byte offset from here)
  • Records are sorted ascending by name_hash and the guest binary-searches them with a 16-byte stride (sub_82448AA0). Verified: sorted in every object.
  • name_hash is tag_hash of the record's own name — not the pak TOC hash (different modulus, and not lowercased).
  • A record owns the half-open field range [field_begin, field_end). Ranges are not contiguous in record order — records are in hash order while their fields sit in name order, so the field array is shared, not partitioned by position.
  • Fields are sorted ascending by key and lower-bounded (sub_8244E338).
  • field_count and pool_size are what make the layout self-checking: a wrong record stride lands on a pool_size that does not equal the remaining bytes. That identity is what caught the first wrong version of this decode.

Field names are on the disc

A field's middle word is its own name's pool offset, and then key == tag_hash(name). Field names never have to be recovered from their hash.

name_off == 0xFFFF_FFFF means the field has no name; its key is then a literal positional integer — a line-slot index (0,1,2,3), a movie id (105, 1303). So a field is positional exactly when it stores no name; you do not have to guess from the key's magnitude.

Disc-wide census of all 2,757,039 fields:

kind count
named, key == tag_hash(name) 1,271,462
unnamed, literal key (< 0x10000) 1,485,073
unnamed, hash-shaped key 504

Those 504 are the entire remaining preimage problem on the disc. Their names are unrecovered.

Two corrections

WITHDRAWN — "the word at 0x08 is a schema hash"

The corpus (and IdxdObject::schema_hash, and every pak list line printing schema 067025b9) read the third header word as an object-type id whose preimage was unknown. It is not a schema id. It is simply record 0's name_hash: the record array is uniform 16-byte entries starting at 0x08, and the header has no type field at all.

Evidence: tag_hash(records[0].name) == word@0x08 for all 7750 objects on the disc, zero exceptions (first_header_word_is_record0_hash).

How it was caught, which is the useful part: a test asserted that every movie id in BASE_INFO names a real record, and one — 1005 -> "STAGE10_PHASE01" — did not resolve. The reason was that tag_hash("STAGE10_PHASE01") is 0x067025B9, the movie table's supposed schema id. A "coincidence" at 1-in-2^32 is not a coincidence; the record was being eaten by the header.

It still works as a type discriminator, because tables of one kind share their lowest-hashed record name — which is exactly why it went unquestioned for so long. schema_hash is therefore kept under its established name, with its docs corrected, rather than renamed across 33 call sites. Nothing on disc names an object's type.

WITHDRAWN — "the field's middle word is an aux/flags word, always 0xFFFFFFFF"

It is 0xFFFFFFFF for 54% of fields, which is enough to look constant in a small sample. It is a name offset (above). The tell was that the word is constant per key across records (key=0x6c43a78d always carried 1743) — a per-record flag cannot do that, but a per-name pointer must.

Worked example — the movie table

dat/tables.pak, the object whose record 0 is STAGE10_PHASE01 (105 records, 433 fields). Its BASE_INFO record mixes both field kinds:

named:       PATH = "dat\movie\"   VERSION = "0x060329"   SUBTITLE_FONT …
positional:  105 -> "STAGE01_PHASE01"   205 -> "STAGE02_PHASE01"   1005 -> "STAGE10_PHASE01"

Each positional value names another record in the same object, which carries the real files:

STAGE02_PHASE01:  MOVIE = "RT02A.wmv"   VOICETRACK = "VOICE_RT02A"
                  SUBTITLE = "…+SUBTITLE_RT02A.tbl"   TELOP = "…+pwrt02.prt"

All 104 positional ids resolve to a real record — verified, no dangling entries. This is the id space the mission script's cutscene request uses; see movie-subtitle-link.

IXUD is the same container

The wide-string sibling has an identical shape, with three substitutions: the hash is ixud_hash, strings are UTF-16BE, and every offset — record name, field name and value alike — is in 16-bit chars, so pool_base + 2*off. pool_size is likewise a char count, which is the same identity as the already-known STR + 2*strsize == filesize (ixud-localised-text).

Verified independently over all 1104 IXUD objects on the disc: the header word at 0x08 is record 0's ixud_hash (1104/1104), and key == ixud_hash(field name) for 628,165 / 628,165 named fields, zero mismatches. Only 48 fields are unnamed — 6 objects × 8, all in a CATEGORY_DESC record with keys 0..7, a positional array of weapon-category descriptions.

One extra rule shows up in IXUD's comparator (sub_82447F38) and is worth assuming for IDXD too: name_off == 0xFFFF_FFFF is the primary sort key, so a record's field slice is unnamed fields first, then named fields, each key-ascending (1476/1476 slices). That is why there are two getters — lookup-by-integer searches only the unnamed run, lookup-by-name only the named run.

🟡 The loader sub_82448D00 also accepts legacy magics IIDX, IDX2, IDX3, IDXC and rejects IDXD on that path with "Old virsion binary table. Not supported." [sic]. None of them occur on this disc (magic census over every pak entry: IDXD 7750, T8aD 4525, RATC 2985, IXUD 1104, LSTA 64, and zero of the four legacy magics), so that path is read-only knowledge, untestable here.

What this does not settle

  • The names of the 504 unnamed hash-keyed fields.
  • How much of the existing corpus the old string-pool reader got wrong. The legacy get_f32/get_raw path infers a field's value from pool adjacency (<value>\0<key>\0). That adjacency is a consequence of the field table, not a rule of the format, and it cannot represent a field whose value string is shared or reordered. Every number in docs/re/ that came from it is now re-checkable against the true table, and has not yet been re-checked.
  • Whether IXUD, the wide-string sibling, uses the same shape. Its offsets are in 16-bit chars and its hash is different (ixud-localised-text).
  • The type-identification question is now open rather than closed: with no schema field, an object's kind is known only from the caller that loads it.