Pure static analysis, no run spent. XObject::Wait's prologue does mov %rdi,%rbx at 8fbc9c, so the this pointer lives in a callee-saved register rather than a stack slot. And the binary carries full unwind information: .eh_frame with 127231 FDEs, which survives in Release builds because C++ exceptions need it, including an FDE covering 8fbc90 to 8fbde2 that tracks rbx explicitly. Together those mean that from a thread frozen deep in pthread_cond_wait, moving to the XObject::Wait frame and reading rbx yields the XObject being waited on -- gdb reconstructs callee-saved registers during the unwind from .eh_frame alone, with no DWARF involved. Reading the first quadword at that pointer gives the vtable, and vtable symbols are in the symtab, so the object's concrete type is identifiable too. This revises the previous entry, which listed route 1 as per-frame archaeology that must be redone whenever the binary changes, and route 2, a RelWithDebInfo rebuild, as what would make the question easy. Route 1 is neither expensive nor fragile: two gdb commands per thread, no rebuild, and the oracle stays byte-identical to the binary every other measurement in this corpus was taken against. Not yet executed on a frozen run, which is the next step and is now a small one.
Reverse-engineering knowledge base
This directory is the spec-side of the clean-room: it records what the original Project Sylpheed binary does (behaviour) and how its data is laid out, so that the Rust port can be implemented from these specs without re-deriving anything and without ever copying original code.
It exists to answer one question fast: "do we already know how X works, and how sure are we?"
The one rule that matters
Never document a claim more confidently than the evidence supports, and never paste original code here.
A wrong-but-confident note is worse than no note: someone builds on it and the bug hides for weeks. Every entry therefore carries an explicit confidence and its evidence. This mirrors the project method — measure the oracle, never infer; refute before believing.
Clean-room firewall
- ✅ Allowed: behaviour descriptions, field offsets/types, formulas, state machines,
observed input→output pairs, and references to the original by address
(
sub_821B68C0) or toxenia-rs/sylpheed.db. - ❌ Forbidden here and in
crates/: pasted decompiled C/C++ or verbatim disassembled function bodies presented as the thing to reimplement. Cite the address; describe the behaviour in your own words. Disassembly is a tool for understanding, not a source to copy.
Confidence levels
| Level | Meaning | Bar to reach it |
|---|---|---|
CONFIRMED |
Behaviour verified against ground truth. | ≥2 independent observations or one observation cross-checked against an oracle (canary framebuffer, a known-correct value, a second code path). |
PROBABLE |
Strong single-source inference. | One clean observation, or an unambiguous static read of the disassembly. |
HYPOTHESIS |
Educated guess, not yet tested. | Anything else. Must say what would confirm/refute it. |
Promotion requires new evidence, not re-reading the old evidence. A HYPOTHESIS that
"looks right again" is still a HYPOTHESIS. Only an independent check promotes it.
If evidence later contradicts an entry, demote it and record the contradiction — do
not silently edit the conclusion.
When to document
- Right after a function/structure crosses from
HYPOTHESISto at leastPROBABLE— before moving to the next code path, so the knowledge isn't lost or re-derived. - Whenever confidence changes (up or down) — append to the Evidence log, don't overwrite.
- Not while it's still a pure guess with no evidence — a one-liner in the relevant backlog/plan is enough until there's something to stand on.
What to document
- Functions/code paths →
docs/re/functions/<name>.md(one file per function or tight cluster). - Data structures / formats →
docs/re/structures/<name>.md. - Keep the index in
INDEX.md(one line each: name · confidence · one-line summary).
Use the templates: _TEMPLATE.function.md,
_TEMPLATE.structure.md.
How we find and confirm code paths (the toolchain)
Everything joins on the guest virtual address (PC) — code addresses are fixed by the XEX load, identical across our emulator and canary.
- Static (cheap, try first):
xenia-rs/sylpheed.db(DuckDB: 25 481 functions, xrefs, strings, vtables, imports). Query withxenia-rs/zq.py—zq.py grep <str>,zq.py xref <addr>,zq.py dis <lo> <hi>,zq.py fn <pc>. Entry points are usually a string (zq.py grep MSG_DEMO) or an import (movie/XMA API) xref'd back to the loader. - Dynamic (when static is ambiguous): run
xenia-rswith its probe suite —--pc-probe/--audit-pc-probe-hex(fires at block entry),--mem-watch(mid-block reads/writes of a VA),--lr-trace(call/return chains),--trace-instructions,--dump-addr(read guest memory). These already exist; prefer them over hacking canary. - Oracle (correctness ground truth): canary — the Wine cross-build
xenia-canary/build-cross/bin/Windows/Debug/xenia_canary.exe(the native Linux ELF crashes / does not run — do not use it). This is the only emulator that reaches the in-game menu; ourxenia-rsnever got past the intro video. Use canary to observe output (capture its framebuffer for texture colours), not usually to instrument code — though itsbuild-crosstoolchain does compile, so small C++ probes + rebuild are possible when needed. Run muted, one emulator process at a time, point it at the real ISO (not the symlink).
⚠️ VA-equality caveat: join code by PC (fixed), but never assume a data VA holds the same bytes across emulators — allocators differ. Compare data by content/layout.