This repository has been archived on 2026-09-16. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
Syplheed-Reborn/docs/re/structures/mcol-collision.md
Sylpheed RE agent 25ccb6723a re: MCOL A->B is 1:1 with count 1, matching REGN's cell-index convention
Filtering array A by the cell-index criteria and following each pointer: every A
record points at a distinct B record, 11/11, and every A count field is exactly
1, 11/11. Per object the A-record count is the number of occupied cells (110 to
575) and the counts sum to it exactly.

That is the same design REGN uses -- the corpus already records 'every occupied
cell has count exactly 1' there. Two sibling formats, one convention, and a
further independent confirmation of the A reading since the filter and the
cardinality are unrelated criteria.

Array B resisted, and both attempts failed in ways worth recording. The record
boundary was off by 8 again, producing records that start with the tail of the
previous structure -- the same mistake as the 0x74 check two iterations ago. And
the u16-index test had no power: the 0x5C block holds ~1232 points, so 'is this
u16 below the point count' passes for almost any small value, and duly reported
100% at seven offsets. Fourth time in these two formats that a bound-check
against a large collection produced a meaningless 100%.

Recorded what would have power instead: B records are 1:1 with occupied cells
and A carries that cell's bounding sphere, so a B field can be tested for
spatial consistency with that specific cell.
2026-08-26 08:38:46 +00:00

191 lines
8.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `MCOL` — the same container as `REGN`, and the same map parameters
**🟡 Opened 2026-08-26.** `MCOL` sits beside `REGN` in `hidden/MiscBin.pak`, 11
of each, and has never been decoded. This page establishes what it shares with
`REGN` — which is a lot, and gives the next attempt a large head start.
## ✅ Same container
Over all **11** objects:
| | |
|---|---|
| `POF0` fixup table at `data_size@+4` + 16 | **11 / 11** |
| bbox pad words are `1.0` / `1.0` / `0.0` at `+0x1C`, `+0x2C`, `+0x3C` | **11 / 11** |
| `extent == max − min` for the `0x30` block | **11 / 11** |
So the header prefix is byte-for-byte the same shape as `REGN`'s: magic, data
size, then bbox min / bbox max / extent as `f32[4]`, then a triple at `0x40`.
The `POF0` mechanism applies, which means **the `chunk + 0x10` base and the
loader's own pointer list are available here too** — the two things that cracked
`REGN`.
## ✅ And the same map parameters, exactly
The bounding boxes and the `0x40` triple are not merely similar — the
distributions are **identical**:
| | `MCOL` | `REGN` |
|---|---|---|
| bbox ±250 000 | 2 | 2 |
| bbox ±50 000 | 6 | 6 |
| bbox ±25 000 | 3 | 3 |
| `0x40` = 50 000 | 2 | 2 |
| `0x40` = 10 000 | 9 | 9 |
Eleven maps, and for each one an `MCOL` and a `REGN` describing the same volume
at the same cell size. `0x40` is the **cell size** in `REGN`; the same values in
the same multiplicities here is strong evidence it is the cell size in `MCOL`
too — though note this is a match of *distributions*, not a demonstrated
object-to-object pairing, which would need the two linked by name or by a stage's
tables.
## ❔ What is not yet known
* **Everything past `0x40`.** `MCOL`'s words at `0x50`–`0x84` do **not** look
like `REGN`'s (`REGN` has grid dims at `0x50`, six `u16` counts at `0x60` and
six section pointers at `0x70`; `MCOL` has a large value, two mid-range values
and `112` at `0x50`, mostly zeros at `0x60`, and `0x05050501` at `0x70`). The
headers agree on the spatial prefix and diverge after it.
* **Everything past `0x40`** — but see below; the pointer layout is now known.
## ✅ The pointer layout, from `POF0`
**2026-08-26.** Running the known-good decoder (`regn_decode.py` on
`auto/regn-reader`) rather than my own broken one. Sanity check first: on `REGN`
it returns header slots `0x70`–`0x84` exactly — the six section pointers — so the
tool and my use of it are right.
On `MCOL`, over all **11** objects:
| | |
|---|---|
| header-region relocated slots are **exactly `0x54`, `0x58`, `0x5C`, `0x74`** | **11 / 11** |
| slot `0x5C` resolves to **`0x80`** — the first byte after the header | **11 / 11** |
| slot `0x74` resolves to **(first array pointer − 8)** | **10 / 11** |
So `MCOL` has **four** top-level pointers where `REGN` has six, and one of them
(`0x5C`) always addresses the data immediately following the header.
**The bulk of the relocations form record arrays.** 92.7 % of the gaps between
consecutive relocated words are **32 bytes**, arranged in 7–127 contiguous runs
per object. Combined with `0x74` landing 8 bytes before the first of them, the
reading is an array of **32-byte records each carrying one pointer at `+8`**.
What the four targets look like:
0x5C -> 0x80 c685620b 4596789d 4694b3b5 44a9a634 floats
0x54 -> … c6826964 456b1aa4 469ab065 c685cfe8 floats
0x58 -> … 00000001 00020003 00040005 00050004 small ints / u16 pairs
0x74 -> … 00000001 00000001 00007710 00000000 counts, then the array
Two float blocks, an index block and a record array is the shape of a mesh —
which is what a name like `MCOL` beside a navigation mesh would suggest. 🟡 That
is a reading of the shape; none of the four blocks has been decoded.
❔ Still open: the record layout, what the index block indexes, the **one object
in eleven** where `0x74` does not land 8 before the array, and the 7.3 % of gaps
that are not 32 (they are the boundaries between runs, but that has not been
checked).
## 🟡 The 32-byte record — a cell entry, and there are two interleaved arrays
**2026-08-26.** Reading each record as 8 big-endian words (record start =
pointer slot − 8), the first entries of the smallest object are:
@0x6780 00000001 00000001 00007710 00000000 00000000 00000000 00000000 47295092
@0x67A0 01000001 00000001 00007730 00000000 …
@0x67C0 02000001 00000001 00007750 00000000 46023555 C6023555 C6023555 471FA1A7
@0x6820 00010001 00000001 000077B0 00000000 …
Word 0 read as **four bytes** is `(x, y, z, 1)` — a **3-D cell index**. That
object's grid is 5×5×5 (bbox ±25 000, cell 10 000), and the values run 0–4 in
the first byte and step the second byte at the right point. Words 4–6 are a
position and word 7 a positive scalar — a bounding sphere. Word 1 is a count and
word 2 the relocated pointer.
**Every record pointer lands in the same region — 8 976 / 8 976 (100 %)** — and
each points 0xFA0 further on with the same stride, so there are **two parallel
arrays**, not one: array A at `0x6780` and array B at `0x7720`.
### The 50 % is the tell, not a failure
Testing the cell-index reading over *all* relocated records gives almost exactly
half:
byte 3 == 1 4 491 / 8 976 (50.03 %)
bytes 0..2 a valid cell index 4 488 (50.00 %)
the record's sphere reaches that cell 4 485 (49.97 %)
Three independent criteria all landing on 50.0 % is not a partial fit — it says
**half the records are not this type**. The `POF0` slot list interleaves both
arrays, and I was testing array B's records against array A's layout. Reported as
a rate it would read like a half-working hypothesis; split by array it is two
clean populations.
🟡 So array A is a **per-cell record** — cell index, count, pointer into array B,
bounding sphere — the same role `REGN`'s section 3 plays. ❔ Array B's layout is
unread, and the split has not yet been re-run per array to confirm 100 % on A.
### ✅ Split by array, and it goes to 100 %
Done. Separating the records by address and re-running the same three criteria:
| | array A (2 509) | array B (6 467) |
|---|---|---|
| byte 3 of word 0 == 1 | **100.00 %** | 30.65 % |
| bytes 0–2 a valid cell index | **100.00 %** | 30.60 % |
| the record's sphere reaches that cell | **99.92 %** | — |
**Array A is the per-cell record**, exactly as read: cell index `(x, y, z)`, a
count, a pointer into array B, and a bounding sphere — the same role `REGN`'s
section 3 plays. Array B is a different record type; its ~30 % is incidental,
and it is the control that shows A's 100 % is not something any 32-byte block
would score.
Worth noting the split was crude — "first half by address", giving 2 509 vs
6 467 rather than an even cut — and A still came out clean at 100 %. A rough
partition that isolates a perfect population is stronger evidence than a careful
one that isolates a good-ish population.
### ✅ A → B is one-to-one, and every count is 1
Filtering array A by the cell-index criteria (so only genuine A records) and
following each one's pointer, over all **11** objects:
| | |
|---|---|
| every A record points at a **distinct** B record | **11 / 11** |
| every A record's count field is **exactly 1** | **11 / 11** |
Per object the A-record count is the number of **occupied cells** — 110, 118,
117, 488, 488, 575, 514, 468, 492, 546, 572 — and the total of the count fields
equals it exactly.
That is the *same* design `REGN` uses: this corpus already records for `REGN`
that "every occupied cell has count exactly 1 — total items equals occupied
cells". Two sibling formats, same cell-index convention. It is a further
independent confirmation of the A-record reading, since the filter and the
cardinality are unrelated criteria.
### ❌ Array B resisted a first pass, and the tests I reached for were bad ones
I could not read B's record layout this iteration, and both attempts failed in
ways worth recording rather than retrying:
* **The boundary was off by 8 again.** Dumping from `pointer − 8` (as A's layout
needed) produced records that begin with what is plainly the *tail* of the
previous structure — a zero and `0x47295092`, the same radius value A's first
record carries. Same mistake as the `0x74` check two iterations ago.
* **The `u16`-index test had no power.** The `0x5C` block holds ~1 232 points, so
"is this `u16` below the point count" is satisfied by almost any small value —
and duly reported 100 % at seven different offsets. That is the *fourth* time
in this pair of formats that a bound-check against a large collection has
produced a meaningless 100 %.
❔ So B's layout is open. **What would have power**: B records are 1:1 with
occupied cells and A already carries the cell's bounding sphere, so a candidate
field in B can be tested by whether it is spatially consistent with *that
specific cell* — the same design that worked for `REGN`'s faces, and the same
design that these two bound-checks lack.