This repository has been archived on 2026-09-16. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
Syplheed-Reborn/docs/re/structures/mcol-collision.md
Sylpheed RE agent 4f0d21f50d re: MCOL's 0x5C block is bounding spheres at stride 16 -- the 0.75 was 12/16
The unexplained ~0.75 ratio left at the end of the last iteration was my own
stride.  I had read the block as 12-byte points because REGN's vertex section
is 12 bytes, and never checked it: len(0x5C) is not a multiple of 12 in 5 of
the 11 objects, so that stride was never arithmetically possible.

At stride 16 the relation is exact in 11/11 -- max u16 == len(0x5C)/16 - 1 --
and the record reads as {centre f32[3], radius f32}.  Powered test, since a
u16 is reached through a specific grid cell: the sphere it names reaches that
cell in 18 559/18 577 = 99.90%, against a 12.02% random-sphere control.  Both
fields carry signal (centre alone 26.75%, radius shuffled 70.19%).

The converse -- is the list *exactly* the intersecting set? -- is 0.38%, which
is the expected direction: a bounding sphere is conservative, so membership
implies overlap but not the reverse.  The tighter geometry is in 0x54/0x58,
still undecoded.  18 entries (0.10%) go the wrong way and are recorded as open.

tools/re-capture/regn_decode.py is copied unchanged from auto/regn-reader so
the probe's POF0 reader is the known-good one rather than a second copy.
2026-08-26 09:04:30 +00:00

378 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `MCOL` — the same container as `REGN`, and the same map parameters
**🟡 Opened 2026-08-26.** `MCOL` sits beside `REGN` in `hidden/MiscBin.pak`, 11
of each, and has never been decoded. This page establishes what it shares with
`REGN` — which is a lot, and gives the next attempt a large head start.
## ✅ Same container
Over all **11** objects:
| | |
|---|---|
| `POF0` fixup table at `data_size@+4` + 16 | **11 / 11** |
| bbox pad words are `1.0` / `1.0` / `0.0` at `+0x1C`, `+0x2C`, `+0x3C` | **11 / 11** |
| `extent == max − min` for the `0x30` block | **11 / 11** |
So the header prefix is byte-for-byte the same shape as `REGN`'s: magic, data
size, then bbox min / bbox max / extent as `f32[4]`, then a triple at `0x40`.
The `POF0` mechanism applies, which means **the `chunk + 0x10` base and the
loader's own pointer list are available here too** — the two things that cracked
`REGN`.
## ✅ And the same map parameters, exactly
The bounding boxes and the `0x40` triple are not merely similar — the
distributions are **identical**:
| | `MCOL` | `REGN` |
|---|---|---|
| bbox ±250 000 | 2 | 2 |
| bbox ±50 000 | 6 | 6 |
| bbox ±25 000 | 3 | 3 |
| `0x40` = 50 000 | 2 | 2 |
| `0x40` = 10 000 | 9 | 9 |
Eleven maps, and for each one an `MCOL` and a `REGN` describing the same volume
at the same cell size. `0x40` is the **cell size** in `REGN`; the same values in
the same multiplicities here is strong evidence it is the cell size in `MCOL`
too — though note this is a match of *distributions*, not a demonstrated
object-to-object pairing, which would need the two linked by name or by a stage's
tables.
## ❔ What is not yet known
* **Everything past `0x40`.** `MCOL`'s words at `0x50`–`0x84` do **not** look
like `REGN`'s (`REGN` has grid dims at `0x50`, six `u16` counts at `0x60` and
six section pointers at `0x70`; `MCOL` has a large value, two mid-range values
and `112` at `0x50`, mostly zeros at `0x60`, and `0x05050501` at `0x70`). The
headers agree on the spatial prefix and diverge after it.
* **Everything past `0x40`** — but see below; the pointer layout is now known.
## ✅ The pointer layout, from `POF0`
**2026-08-26.** Running the known-good decoder (`regn_decode.py` on
`auto/regn-reader`) rather than my own broken one. Sanity check first: on `REGN`
it returns header slots `0x70`–`0x84` exactly — the six section pointers — so the
tool and my use of it are right.
On `MCOL`, over all **11** objects:
| | |
|---|---|
| header-region relocated slots are **exactly `0x54`, `0x58`, `0x5C`, `0x74`** | **11 / 11** |
| slot `0x5C` resolves to **`0x80`** — the first byte after the header | **11 / 11** |
| slot `0x74` resolves to **(first array pointer − 8)** | **10 / 11** |
So `MCOL` has **four** top-level pointers where `REGN` has six, and one of them
(`0x5C`) always addresses the data immediately following the header.
**The bulk of the relocations form record arrays.** 92.7 % of the gaps between
consecutive relocated words are **32 bytes**, arranged in 7–127 contiguous runs
per object. Combined with `0x74` landing 8 bytes before the first of them, the
reading is an array of **32-byte records each carrying one pointer at `+8`**.
What the four targets look like:
0x5C -> 0x80 c685620b 4596789d 4694b3b5 44a9a634 floats
0x54 -> … c6826964 456b1aa4 469ab065 c685cfe8 floats
0x58 -> … 00000001 00020003 00040005 00050004 small ints / u16 pairs
0x74 -> … 00000001 00000001 00007710 00000000 counts, then the array
Two float blocks, an index block and a record array is the shape of a mesh —
which is what a name like `MCOL` beside a navigation mesh would suggest. 🟡 That
is a reading of the shape; none of the four blocks has been decoded.
❔ Still open: the record layout, what the index block indexes, the **one object
in eleven** where `0x74` does not land 8 before the array, and the 7.3 % of gaps
that are not 32 (they are the boundaries between runs, but that has not been
checked).
## 🟡 The 32-byte record — a cell entry, and there are two interleaved arrays
**2026-08-26.** Reading each record as 8 big-endian words (record start =
pointer slot − 8), the first entries of the smallest object are:
@0x6780 00000001 00000001 00007710 00000000 00000000 00000000 00000000 47295092
@0x67A0 01000001 00000001 00007730 00000000 …
@0x67C0 02000001 00000001 00007750 00000000 46023555 C6023555 C6023555 471FA1A7
@0x6820 00010001 00000001 000077B0 00000000 …
Word 0 read as **four bytes** is `(x, y, z, 1)` — a **3-D cell index**. That
object's grid is 5×5×5 (bbox ±25 000, cell 10 000), and the values run 0–4 in
the first byte and step the second byte at the right point. Words 4–6 are a
position and word 7 a positive scalar — a bounding sphere. Word 1 is a count and
word 2 the relocated pointer.
**Every record pointer lands in the same region — 8 976 / 8 976 (100 %)** — and
each points 0xFA0 further on with the same stride, so there are **two parallel
arrays**, not one: array A at `0x6780` and array B at `0x7720`.
### The 50 % is the tell, not a failure
Testing the cell-index reading over *all* relocated records gives almost exactly
half:
byte 3 == 1 4 491 / 8 976 (50.03 %)
bytes 0..2 a valid cell index 4 488 (50.00 %)
the record's sphere reaches that cell 4 485 (49.97 %)
Three independent criteria all landing on 50.0 % is not a partial fit — it says
**half the records are not this type**. The `POF0` slot list interleaves both
arrays, and I was testing array B's records against array A's layout. Reported as
a rate it would read like a half-working hypothesis; split by array it is two
clean populations.
🟡 So array A is a **per-cell record** — cell index, count, pointer into array B,
bounding sphere — the same role `REGN`'s section 3 plays. ❔ Array B's layout is
unread, and the split has not yet been re-run per array to confirm 100 % on A.
### ✅ Split by array, and it goes to 100 %
Done. Separating the records by address and re-running the same three criteria:
| | array A (2 509) | array B (6 467) |
|---|---|---|
| byte 3 of word 0 == 1 | **100.00 %** | 30.65 % |
| bytes 0–2 a valid cell index | **100.00 %** | 30.60 % |
| the record's sphere reaches that cell | **99.92 %** | — |
**Array A is the per-cell record**, exactly as read: cell index `(x, y, z)`, a
count, a pointer into array B, and a bounding sphere — the same role `REGN`'s
section 3 plays. Array B is a different record type; its ~30 % is incidental,
and it is the control that shows A's 100 % is not something any 32-byte block
would score.
Worth noting the split was crude — "first half by address", giving 2 509 vs
6 467 rather than an even cut — and A still came out clean at 100 %. A rough
partition that isolates a perfect population is stronger evidence than a careful
one that isolates a good-ish population.
### ✅ A → B is one-to-one, and every count is 1
Filtering array A by the cell-index criteria (so only genuine A records) and
following each one's pointer, over all **11** objects:
| | |
|---|---|
| every A record points at a **distinct** B record | **11 / 11** |
| every A record's count field is **exactly 1** | **11 / 11** |
Per object the A-record count is the number of **occupied cells** — 110, 118,
117, 488, 488, 575, 514, 468, 492, 546, 572 — and the total of the count fields
equals it exactly.
That is the *same* design `REGN` uses: this corpus already records for `REGN`
that "every occupied cell has count exactly 1 — total items equals occupied
cells". Two sibling formats, same cell-index convention. It is a further
independent confirmation of the A-record reading, since the filter and the
cardinality are unrelated criteria.
### ❌ Array B resisted a first pass, and the tests I reached for were bad ones
I could not read B's record layout this iteration, and both attempts failed in
ways worth recording rather than retrying:
* **The boundary was off by 8 again.** Dumping from `pointer − 8` (as A's layout
needed) produced records that begin with what is plainly the *tail* of the
previous structure — a zero and `0x47295092`, the same radius value A's first
record carries. Same mistake as the `0x74` check two iterations ago.
* **The `u16`-index test had no power.** The `0x5C` block holds ~1 232 points, so
"is this `u16` below the point count" is satisfied by almost any small value —
and duly reported 100 % at seven different offsets. That is the *fourth* time
in this pair of formats that a bound-check against a large collection has
produced a meaningless 100 %.
❔ So B's layout is open. **What would have power**: B records are 1:1 with
occupied cells and A already carries the cell's bounding sphere, so a candidate
field in B can be tested by whether it is spatially consistent with *that
specific cell* — the same design that worked for `REGN`'s faces, and the same
design that these two bound-checks lack.
## ✅ Array B decoded — `{count, u16 index array}`, and the chain closes
**2026-08-26.** Two corrections got there.
**Where the other relocations live.** `MCOL` has 9 020 relocated words and only
4 488 are A-record pointers. I assumed the rest sat at `B+8`, mirroring A —
**refuted, 0 / 4 488**. Measuring their offset from the nearest preceding B
record instead:
at B + 4 4 488 (99.0 %)
before the first B 44 (1.0 %)
and those 44 are exactly the four header pointers × 11 objects. **Nothing
unaccounted for.**
**So B is `{u32 count, pointer}`** — pointer at `+4`, not `+8`. Read that way:
B@0x7720 count 0x2B ptr 0x84D0
B@0x7740 count 0x1C ptr 0x8526 0x8526 − 0x84D0 = 0x56 = 2 × 0x2B
B@0x7760 count 0x02 ptr 0x855E 0x855E − 0x8526 = 0x38 = 2 × 0x1C
B@0x7780 count 0x02 ptr 0x8562 = 2 × 2
The pointers advance by exactly twice the count — a **packed `u16` array**, no
padding. Over all 11 objects:
| | |
|---|---|
| consecutive B pointers differ by **exactly `2 × count`** | **4 477 / 4 477 (100.00 %)** |
| the `u16` entries are valid `0x5C` point indices | 18 379 / 18 379 |
The first row is the load-bearing one: an exact arithmetic identity over 4 477
consecutive pairs. (The second is the same weak bound-check flagged above — the
point block is large, so almost any `u16` passes. It is consistent, not
evidence.)
### The chain
position → cell → A record {cell index, count 1, →B, bounding sphere}
→ B record {count n, →u16[n]}
→ n indices into the 0x5C point block ← wrong, see below
which is the same shape as `REGN`'s `cell → item → refs → geometry`, as the two
formats' shared header and shared count-1 convention already suggested.
❔ Still open: the `0x54` and `0x58` blocks (neither is reached by this chain).
What the `u16` entries index is **not** the point block — see immediately below.
## ❌ The `u16` entries do NOT index the point block — and the weak test said they did
**2026-08-26.** Two tests, and the contrast between them is the point of this
section.
**Not a triangle list.** If the `u16` array held triangle corners, every count
would be divisible by 3. Counts modulo 3 across all objects:
n % 3 == 0 639 n % 3 == 1 2 063 n % 3 == 2 1 786
Spread across all three residues. Refuted.
**Not cell-local either.** A B record is reached *through* a specific cell, so a
point it references should lie in that cell. With a random-point control:
| | |
|---|---|
| referenced point inside the cell that reached it | 142 / 17 871 = **0.79 %** |
| a random point inside that cell (control) | 85 / 17 871 = **0.48 %** |
Chance. So whatever the `u16`s address, it is not the `0x5C` point block in any
spatially meaningful way.
### ⚠️ This is the bound-check hazard caught in the act
One section above, the same `u16` entries scored **18 379 / 18 379 (100 %)** on
"are these valid `0x5C` point indices". I flagged that at the time as the weak
bound-check rather than evidence, and recorded it as *consistent* rather than as
a finding. **That caution was correct**: the powered version of the same question
now returns chance.
This is the fifth appearance of the pattern across `REGN` and `MCOL`, and the
first time both halves have been run side by side on the same field, so it is
worth stating exactly:
> A bound-check asks *"could this be an index?"*. Almost always, yes — the answer
> is set by how large the target collection is, not by whether the field means
> anything. The powered version asks *"does the thing it points at make sense
> where it was reached from?"*, and only that version can be wrong.
Had the 100 % been written up as the decode, this page would now carry a
confident and false statement about `MCOL`'s geometry.
❔ What the `u16`s index is open. A datum for the next attempt: the maximum
`u16` is consistently ≈ **0.75 ×** the point count (923/1 232, 1 019/1 360,
1 163/1 552, 59/80) — too consistent to be coincidence, and not explained.
> **Resolved in the next section.** 0.75 is `12 / 16`: "the point count" was
> computed with an assumed 12-byte stride that the block lengths refute. The
> `u16`s *do* index this block — at stride 16. Both this test and the 100 %
> bound-check above were reading the block wrongly; only the powered one could
> say so.
## ✅ The `0x5C` block is **bounding spheres at stride 16** — and that explains the 0.75
**2026-08-26.** The unexplained ≈0.75 ratio left at the end of the section above
was **12 / 16**: my own stride. I had been reading the `0x5C` block as 12-byte
points because `REGN`'s vertex section is 12 bytes, and never checked the
assumption.
It does not survive the cheapest possible check — **`len(0x5C)` is not a
multiple of 12** in 5 of the 11 objects, so a 12-byte stride was never
arithmetically possible:
| object | `len(0x5C)` | ÷12 | ÷16 | max `u16` |
|---|---|---|---|---|
| `2cf7eb47` | 960 | 80.00 | 60 | 59 |
| `cbb99d34` | 192 | 16.00 | 12 | 11 |
| `d84a95fb` | 3 488 | 290.67 ❌ | 218 | 217 |
| `db066592` | 12 640 | 1 053.33 ❌ | 790 | 789 |
| `db61c506` | 2 144 | 178.67 ❌ | 134 | 133 |
| `dc44fe0c` | 2 784 | 232.00 | 174 | 173 |
| `dc4b0896` | 14 784 | 1 232.00 | 924 | 923 |
| `dd89a110` | 18 624 | 1 552.00 | 1 164 | 1 163 |
| `df4628c2` | 4 160 | 346.67 ❌ | 260 | 259 |
| `e084c13c` | 192 | 16.00 | 12 | 11 |
| `e16460cf` | 16 320 | 1 360.00 | 1 020 | 1 019 |
At stride 16 the relation is not "≈0.75×" but **exact, in 11 / 11 objects**:
max u16 == len(0x5C) / 16 − 1
The `u16` array indexes the `0x5C` block at stride 16, and *covers it fully* —
the largest index is always the last element.
### The record is `{ centre f32[3], radius f32 }`
Read at stride 16, the first three floats lie inside the object's own bounding
box in **every record of every object**, and are never unit-length, so this is a
position and not a plane normal. The fourth float is a positive scalar which is
**not** `|centre|`.
The powered test is the one the previous section said was needed: a `u16` is
reached *through a specific cell*, so the thing it names should be present in
that cell. Treating the record as a sphere and the cell as its grid box:
| | |
|---|---|
| **referenced sphere intersects the cell that reached it** | **18 559 / 18 577 = 99.90 %** |
| a random sphere from the same object (control) | 2 233 / 18 577 = 12.02 % |
And both halves of the record are load-bearing — ablating either one costs most
of the signal:
| | |
|---|---|
| centre + radius | **99.90 %** |
| centre alone, radius treated as 0 | 26.75 % |
| centre kept, radius shuffled within the object | 70.19 % |
| radius kept, centre shuffled within the object | 23.74 % |
So the `0x5C` block is a **broad-phase bounding-sphere array**, and each grid
cell's `u16` list names the primitives that reach into that cell — the standard
shape for a collision mesh, and the sibling of `REGN`'s `cell → tetrahedra`.
### The list is a *subset* of what the spheres allow, which is the expected direction
Testing the converse — is the `u16` set **exactly** the set of spheres that
intersect the cell? — gives **17 / 4 488 cells (0.38 %)**, with 46 525 spheres
intersecting a cell but absent from its list. That is the right direction and
not a problem: a bounding sphere is a conservative bound on the primitive inside
it, so "sphere overlaps cell" must be implied by membership but cannot imply it.
The tighter true geometry lives in the `0x54` / `0x58` blocks, still undecoded.
❔ **18 exceptions (0.10 %)** go the wrong way — listed, but the sphere misses
the cell, by `dist / radius` of 1.004 to 1.129. They are spread over 7 of the 11
objects with no object dominating, so this looks like a small build-time margin
rather than a decode error, but it is **not explained** and is recorded as open.
### The chain, corrected
position → cell → A record {cell index (x,y,z,1), count 1, →B at +8, bounding sphere}
→ B record {u32 count, →u16[n] at +4}
→ n indices into the 0x5C array of 16-byte bounding spheres
❔ Still open: the `0x54` and `0x58` blocks — the actual collision geometry that
these spheres bound. Their lengths are **not** a constant multiple of the sphere
count (`0x58 / n` is ≈6.0 for the large objects but 6.13 and 6.67 for the two
smallest), so at least one of them is variable-stride or has its own count.