Files
Sylpheed/docs/re
Sylpheed RE agent 749567e835 re: the +184 writer is blocked - and xrefs.ind_call is a CROSS PRODUCT, not a call graph
Chasing what supplies the EX_ mode word. Two routes, both measured to have no
power, plus one structural fact that did come out.

The offset route is dead. stw ..., 184(rN) occurs 301 times in the image, lwz
from +184 occurs 351 times, and 115 functions touch both +144 and +184. +184 is
an ordinary small offset shared by many unrelated classes - the same shape the
corpus already recorded as failing three times. Nothing narrows to a writer.

The bigger result is an instrument refutation with corpus-wide reach. Asking
xrefs for the callers of the two screens returns 633 sources for EACH, and the
two lists are IDENTICAL, which cannot be right. Measuring the relation itself:

  ind_call rows                       1827297
  distinct targets                       1710
  distinct sources                       6992
  targets with EXACTLY 633 sources        236

236 different functions sharing an identical source count is the signature of an
unresolved-indirect-call cross product, not of a call graph. Any reading that
treats an ind_call edge as "X calls Y" is void, here and anywhere else in the
corpus it may have been used.

The control shows the other kinds are sound: sub_82286BC8 has exactly one
caller, kind call. And both EX_ screens have ZERO non-ind_call edges - they are
reached only through function pointers, which is why the direct graph is empty
for them.

What did come out: scanning the entries of all 1150 catalogued vtables in the
flat .pe, both EX_ screens are slot 1 of their own class - sub_822814D8 in
ANON_Class_271D5F25 and sub_8227A3A0 in ANON_Class_CAA8AD62 - while sub_82286BC8
is in no catalogued vtable at all. So the 2/0/0/1 partition from the previous
commit reflects a structural difference rather than a coincidence: the two
screens that select on the mode word are vtable methods of their own classes and
the one that does not select is not a vtable method.

Still not settled: what writes +184. Both the offset sweep and the call graph are
exhausted for it. A route with actual power would be the RTTI behind those two
anonymous classes, or a runtime watch on the field - not another static offset
search.

All seventeen artefacts byte-identical.
2026-08-28 11:08:18 +00:00
..

Reverse-engineering knowledge base

This directory is the spec-side of the clean-room: it records what the original Project Sylpheed binary does (behaviour) and how its data is laid out, so that the Rust port can be implemented from these specs without re-deriving anything and without ever copying original code.

It exists to answer one question fast: "do we already know how X works, and how sure are we?"


The one rule that matters

Never document a claim more confidently than the evidence supports, and never paste original code here.

A wrong-but-confident note is worse than no note: someone builds on it and the bug hides for weeks. Every entry therefore carries an explicit confidence and its evidence. This mirrors the project method — measure the oracle, never infer; refute before believing.

Clean-room firewall

  • Allowed: behaviour descriptions, field offsets/types, formulas, state machines, observed input→output pairs, and references to the original by address (sub_821B68C0) or to xenia-rs/sylpheed.db.
  • Forbidden here and in crates/: pasted decompiled C/C++ or verbatim disassembled function bodies presented as the thing to reimplement. Cite the address; describe the behaviour in your own words. Disassembly is a tool for understanding, not a source to copy.

Confidence levels

Level Meaning Bar to reach it
CONFIRMED Behaviour verified against ground truth. ≥2 independent observations or one observation cross-checked against an oracle (canary framebuffer, a known-correct value, a second code path).
PROBABLE Strong single-source inference. One clean observation, or an unambiguous static read of the disassembly.
HYPOTHESIS Educated guess, not yet tested. Anything else. Must say what would confirm/refute it.

The status markers

The table above is the confidence scale. The markers that appear in BACKLOG.md and the structures/ pages are a separate, and until now undefined, vocabulary. They mean:

Marker Meaning
Confirmed — verified against ground truth.
🟡 Partial: true as far as it goes, or true under a stated assumption.
Open question. Nobody has answered it yet.
🔴 Refuted — shown false — or blocked by something the container cannot do.
A specific claim that was tried and failed. Prefer 🔴.
🚧 Work started and not finished.

🔴 never means "we have not run it yet." That is or 🚧. Reserve 🔴's "blocked" sense for a real limit of the box — no push credentials, no hardware Vulkan (lavapipe only), or a decision only the user can make. The box can run the emulator, script input, screenshot, read guest memory, and build and test Rust, so "needs a run" is never a blocker. This paragraph exists because the marker was undefined for 98 uses and three of them were mislabelled that way.

Promotion requires new evidence, not re-reading the old evidence. A HYPOTHESIS that "looks right again" is still a HYPOTHESIS. Only an independent check promotes it. If evidence later contradicts an entry, demote it and record the contradiction — do not silently edit the conclusion.


When to document

  • Right after a function/structure crosses from HYPOTHESIS to at least PROBABLE — before moving to the next code path, so the knowledge isn't lost or re-derived.
  • Whenever confidence changes (up or down) — append to the Evidence log, don't overwrite.
  • Not while it's still a pure guess with no evidence — a one-liner in the relevant backlog/plan is enough until there's something to stand on.

What to document

  • Functions/code pathsdocs/re/functions/<name>.md (one file per function or tight cluster).
  • Data structures / formatsdocs/re/structures/<name>.md.
  • Keep the index in INDEX.md (one line each: name · confidence · one-line summary).

Use the templates: _TEMPLATE.function.md, _TEMPLATE.structure.md.


How we find and confirm code paths (the toolchain)

Everything joins on the guest virtual address (PC) — code addresses are fixed by the XEX load, identical across our emulator and canary.

  • Static (cheap, try first): xenia-rs/sylpheed.db (DuckDB: 25 481 functions, xrefs, strings, vtables, imports). Query with xenia-rs/zq.pyzq.py grep <str>, zq.py xref <addr>, zq.py dis <lo> <hi>, zq.py fn <pc>. Entry points are usually a string (zq.py grep MSG_DEMO) or an import (movie/XMA API) xref'd back to the loader.
  • Dynamic (when static is ambiguous): run xenia-rs with its probe suite — --pc-probe / --audit-pc-probe-hex (fires at block entry), --mem-watch (mid-block reads/writes of a VA), --lr-trace (call/return chains), --trace-instructions, --dump-addr (read guest memory). These already exist; prefer them over hacking canary.
  • Oracle (correctness ground truth): canary — the Wine cross-build xenia-canary/build-cross/bin/Windows/Debug/xenia_canary.exe (the native Linux ELF crashes / does not run — do not use it). This is the only emulator that reaches the in-game menu; our xenia-rs never got past the intro video. Use canary to observe output (capture its framebuffer for texture colours), not usually to instrument code — though its build-cross toolchain does compile, so small C++ probes + rebuild are possible when needed. Run muted, one emulator process at a time, point it at the real ISO (not the symlink).

⚠️ VA-equality caveat: join code by PC (fixed), but never assume a data VA holds the same bytes across emulators — allocators differ. Compare data by content/layout.