The database mount below was one of three ways the decoder could not reach its own oracle. The other two are here. `run-canary` never looked in `Checked/`. It tried `Release/` then `Debug/`, and both of those exist on this box -- an Aug 28 binary and a Jul 19 one. They boot the game perfectly well and carry NO `audit_61` branch probe, so a probe run against either returns zero hits that read as a finding about the game rather than as a stale binary. Configuration is now the outer loop and location the inner one, so a `Checked` build anywhere beats a `Release` build anywhere; `$XENIA_BIN` still overrides everything. Measured here: `Checked` has `audit_61_branch_probe_pcs`, `Release` and `Debug` do not. The launcher also now says which instrumentation is missing BEFORE the run, because the alternative is reading an empty log afterwards and guessing. `build-canary` built `$PROJECT_DIR/xenia-canary`, which does not exist in this container -- the source is bind-mounted at `/canary` and the launcher already exports `XENIA_SRC=/canary`. CONTAINER-NOTES has carried that defect since 2026-08-29 with a symlink workaround and a warning to remember to delete the symlink afterwards. It now reads `$XENIA_SRC` first, so there is nothing to remember. Its default configuration moves Release -> Checked to match what `run-canary` picks; the old default spent a full build on a binary nothing ran. Two documented blockers are refuted rather than deleted, since the sequence of wrong readings is what makes the right one checkable: the CONTAINER-NOTES symlink dance (the warm build volume it was configured against is gone too, removed in the 2026-09-18 cleanup, so the next build configures cleanly against `/canary`), and `upstream-baseline.md`'s "`version.h` is never generated" -- `CMakeLists.txt` generates it at configure time now, with a stub fallback. `decoder-loop.md` claimed the oracle was at `Linux/Release/` and that the probe was on two side branches; both were true when written and neither is now. Verified: five pick_bin cases against the extracted function body, and `strings` on all three real binaries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
300 lines
16 KiB
Markdown
300 lines
16 KiB
Markdown
# Notes for an agent working inside this container
|
||
|
||
Read this before starting a dynamic-RE run. Everything here is something that
|
||
already went wrong once.
|
||
|
||
## The container fixes three old traps for you
|
||
|
||
* **The display outlives the turn.** Xvfb and openbox are children of PID 1, not
|
||
of your shell. The old "Xvfb dies on its own every few minutes" note is gone —
|
||
you no longer have to wrap a whole session in one blocking foreground call to
|
||
keep it alive.
|
||
* **The toolchain is real.** `tools/re-capture/rebuild_canary.sh` exists because
|
||
the old box had no cmake/ninja/clang and only runtime sonames, so it hand-
|
||
relinked object files. **Do not use it here.** Use `build-canary`.
|
||
✅ **FIXED 2026-09-21 — `build-canary` now reads `$XENIA_SRC` first**, which the
|
||
launcher already sets to `/canary`, so there is no symlink to make and none to
|
||
forget. The two paragraphs below are kept because they are how the defect was
|
||
found and refuted, but **neither describes the script any more.**
|
||
|
||
🔴 ~~**`build-canary` does not work in this container as it stands
|
||
(2026-08-29).** It builds `${PROJECT_DIR:-/work}/xenia-canary`, which **does
|
||
not exist here** — the Canary source is at **`/canary`** (`$XENIA_SRC`). The
|
||
warm 235 MB tree at `/sylph-home/re/canary-build` is configured with
|
||
`CMAKE_HOME_DIRECTORY=/work/xenia-canary`, also missing, and its
|
||
`build-Release.ninja` carries **no per-file rules** — it wants to re-run CMake
|
||
first, which would fail on the absent source root.~~ The warm tree is gone too:
|
||
every `sylph-*` volume was removed in the 2026-09-18 cleanup, so the next build
|
||
configures from scratch against `/canary` and caches the right source root.
|
||
|
||
✅ **The conclusion drawn from all that is REFUTED (2026-08-31).** The note went
|
||
on to say *"any Canary change is a full reconfigure against `/canary` plus a full
|
||
compile, not an incremental one"*, and budgeted a session for it. That is wrong,
|
||
and the fix is one line:
|
||
|
||
```bash
|
||
ln -sfn /canary /work/xenia-canary # the path the warm tree was configured with
|
||
cmake --build /sylph-home/re/canary-build --config Release --parallel 4 \
|
||
--target xenia_canary
|
||
rm /work/xenia-canary # it is UNTRACKED inside the repo — remove it
|
||
```
|
||
|
||
Measured: a one-file edit to `src/xenia/gpu/command_processor.cc` rebuilt and
|
||
relinked `bin/Linux/Release/xenia_canary` in **under 10 minutes at `-j4`**, exit
|
||
0, no reconfigure, no OOM. The missing source root was the only defect; the
|
||
warm 235 MB tree is otherwise intact and the ninja re-run resolves its rules
|
||
from the symlink.
|
||
|
||
⚠️ ~~**Remove the symlink when you are done.**~~ No longer needed — but if you
|
||
ever make one by hand, note that `/work` is the repository and `xenia-canary`
|
||
is not in `.gitignore`, so it shows up as untracked and can be swept into a
|
||
`git add -A`.
|
||
* **numpy and Pillow are installed.** `entities2.py`, `flight_probe.py` and the
|
||
image oracles work. Their absence used to look like a logic bug.
|
||
|
||
## Method (the part that matters more than the tooling)
|
||
|
||
* **Measure the oracle; never infer it.** A session with zero Canary runs is a
|
||
red flag.
|
||
* **Trace upstream to where data first goes wrong**, rather than patching the
|
||
symptom you can see.
|
||
* **Try to refute before believing.** Record demotions rather than editing them
|
||
away — `docs/re/README.md` has the ✅/🟡/❔ convention, and a withdrawn result
|
||
is more useful than a quietly deleted one.
|
||
* **A probe that never performs the action will "prove" the action does not
|
||
exist.** The "targeting is automatic" conclusion came from a sweep that only
|
||
ever tapped once; target select is Ⓐ pressed *twice*.
|
||
* **Do not poll faster than the guest updates** — it manufactures a clean curve
|
||
out of noise. `rate-curve-aliased-BAD.csv` is committed as the bad example.
|
||
|
||
## Running the emulator
|
||
|
||
```bash
|
||
run-canary # correct audio/pad/display flags baked in
|
||
pad.py tap A ; pad.py dpad down # scripted input (--hid=file, no uinput)
|
||
screenshot ~/shots/now.png # cropped to the GAME surface, not the window
|
||
python3 tools/re-capture/gmem.py find hex:820af844 400
|
||
```
|
||
|
||
* **One emulator at a time.** `run-canary` enforces it with a lockfile.
|
||
* 🔴 **`run-canary` is SILENT TWICE OVER, and that defeats `audio-capture`.**
|
||
Line 82 is `export SDL_AUDIODRIVER="${SDL_AUDIODRIVER:-dummy}"`, and its
|
||
header explains why: `--apu=nop` stalls the guest in the intro movie, so the
|
||
SDL driver against a *dummy* device is what lets the title advance. The
|
||
comment's premise — "there is no PulseAudio here" — **stopped being true when
|
||
`tools/audio-capture` landed**, and it starts a daemon on demand.
|
||
So a capture through the null sink records **pure silence**, at the right
|
||
length, with a perfectly healthy-looking run behind it. To actually record the
|
||
game:
|
||
|
||
⚠️ **And that is only the first of TWO layers.** `run-canary` also passes
|
||
**`--mute=true`** on its own command line (line 98). With the driver fixed and
|
||
the mute left alone, Canary attaches a healthy 6-channel stream to the sink,
|
||
holds it at 100 % volume for the whole run — and emits silence. Both have to go:
|
||
|
||
```bash
|
||
audio-capture start # or load a null sink yourself
|
||
PULSE_SINK=cap SDL_AUDIODRIVER=pulseaudio \
|
||
run-canary --mute=false … # `"$@"` is last, so this wins
|
||
```
|
||
|
||
Record at the monitor's real format, too — `parec` defaults to stereo/44.1 kHz
|
||
and will silently resample a 6-channel monitor:
|
||
`parec -d cap.monitor --channels=6 --rate=48000 --format=s16le`.
|
||
|
||
🔴 **And even with both mutes off, a PulseAudio-monitor capture is not
|
||
faithful — use the ALSA tee instead.** A null sink's *monitor* is sampled on a
|
||
wall clock and **invents silence** whenever the client is late, so a capture
|
||
through it is 39 % holes that the game never emitted. `PULSE_LATENCY_MSEC`
|
||
only trades gap count against gap size and never wins.
|
||
✅ **The working route is `--apu=alsa` with an ALSA `file` tee in front of a
|
||
paced slave** — full recipe, controls and three configuration traps in
|
||
[`audio-capture-alsa-file-tee.md`](../re/audio-capture-alsa-file-tee.md).
|
||
* 🔴 **A BARE ALSA `file` tee WILL FILL THE DISK — always run a size guard.**
|
||
Xenia's ALSA writer thread pads silence whenever its ring is empty
|
||
(`alsa_audio_driver.cc:359`), so against a device that never blocks it
|
||
free-runs: measured at **~250× real time, 7.34 GB in 50 seconds**. The slave
|
||
must pace — `slave.pcm { type pulse }` — and the capture loop should abort
|
||
above ~3× real time. The next person to try a bare tee hits this in the first
|
||
minute.
|
||
* ✅ **And use `--gpu=null` for an audio capture.** It is what takes the guest
|
||
from 0.70× to **0.96×** real time, which stops Xenia padding at all: 0.31 %
|
||
silence and 0.01 gaps/s, against 9.98 % / 8.37 rendered. ⚠️ No video, so
|
||
screen-based provenance is unavailable (use the XMA probe), and `--gpu=null`
|
||
runs here die at ~70 s with `PM4_DRAW_INDX: Failed in backend`.
|
||
🔴 **REFUTED 2026-08-30 — that lifetime does not hold.** A `--gpu=null` capture
|
||
ran **148.02 s** and ended on its own probe's timer with the emulator still
|
||
alive, having decoded the whole `ADV` movie
|
||
([`intro-audio-output-census.md`](../re/structures/intro-audio-output-census.md)).
|
||
More than twice the quoted figure. Whatever produced the ~70 s was fixed or was
|
||
never general, and this note had been the reason not to use `--gpu=null` for
|
||
anything long — which is exactly the configuration a clean audio capture needs.
|
||
|
||
⚠️ **Do not read Canary's 6-channel PulseAudio stream as evidence the GAME is
|
||
5.1.** `pactl` will show `float32le 6ch 48000Hz`, channel-mapped to a full 5.1
|
||
layout, on any title. That is `AudioDriver::kFrameChannelsDefault = 6`, a
|
||
hardcoded constant — the code path actually used
|
||
(`SDLAudioSystem::CreateDriver(index, semaphore, &driver)`) constructs
|
||
`SDLAudioDriver(semaphore)` and takes every default. The *format* is Xenia's;
|
||
only the *content* of those six channels is the guest's.
|
||
|
||
⚠️ **Check `pactl list sink-inputs` before trusting a recording.** If it is
|
||
empty, Canary never attached and you are recording zeroes; the sink also sits
|
||
at `IDLE`. `audio-capture run` warns on a `-inf` peak afterwards, which is the
|
||
backstop — but a live check fails in seconds instead of after the whole run.
|
||
* Boot is slow cold, ~25 s once the shader/code caches are warm — so a
|
||
launch-and-dump fits in a single call.
|
||
* **Screens: classify by whole-image statistics** (`screen_id.py`), not named
|
||
pixels. Named-pixel oracles are only valid while the game image sits at a
|
||
known place, and nothing errors when it moves.
|
||
|
||
## Verifying your own work
|
||
|
||
* Reborn's disc-gated tests **self-skip** without `SYLPHEED_DISC`. A green run
|
||
with it unset means almost nothing.
|
||
🔴 **But `build-reborn` does not work in this container (2026-08-29).** Line 15
|
||
is `SRC="${PROJECT_DIR:-/work}/Syplheed-Reborn"` — note the transposed letters —
|
||
and no such directory exists; the workspace is at **`/work`** itself. It fails
|
||
immediately with `cd: /work/Syplheed-Reborn: No such file or directory`, so the
|
||
documented way to run the disc-gated tests is broken.
|
||
✅ **Run them directly instead**, setting the variable yourself:
|
||
```bash
|
||
SYLPHEED_DISC=/disc cargo test -p sylpheed-formats --test <name>
|
||
```
|
||
⚠️ This is the **second** wrapper in this container pointing at a source root
|
||
that does not exist — `build-canary` has the same defect. Check a wrapper's
|
||
`SRC` before trusting that a green or a failure came from your code.
|
||
* Prefer a headless self-verify over "it compiles": `sylpheed-cli mesh render`,
|
||
`screen render`, `save info` all produce checkable artifacts.
|
||
* A Bevy system-parameter conflict is invisible to the type checker and panics
|
||
at startup. If you touch viewer systems, *run the binary*, don't just build it.
|
||
|
||
## Reporting
|
||
|
||
State what you measured, what you assumed, and what you could not settle. If a
|
||
result is withdrawn, say so and keep the reasoning — that is the corpus's whole
|
||
convention, and the reason its numbers can be trusted.
|
||
|
||
## The static-analysis corpus — mounted, not reproducible
|
||
|
||
Two things arrive read-only from the host because **nothing in this repository
|
||
can produce them yet**:
|
||
|
||
| path | what | env |
|
||
|---|---|---|
|
||
| `/xenia-rs/sylpheed.db` | the disassembly database, 586 MB | `SYLPHEED_DB` |
|
||
| `/image/sylpheed.pe` | the decompressed image, flat VA dump | `SYLPHEED_PE` |
|
||
|
||
**The `.pe` is a flat VA dump**: file offset = `VA - 0x82000000`
|
||
(`SYLPHEED_IMAGE_BASE`). So reading `0x820A1630` is `seek(0xA1630)` — no XEX
|
||
decrypt, no LZX, and **no booted emulator**. An earlier belief that this file was
|
||
stale was tested and **refuted**; it is current.
|
||
|
||
Recovering the image by dumping `/dev/shm/xenia_memory_*` also works and
|
||
self-validates, but it needs a running emulator — a poor dependency for
|
||
something the whole static corpus rests on. Use the file.
|
||
|
||
The database is far richer than the four scripts that read it use:
|
||
|
||
```
|
||
functions 25 481 address, name, end_address, frame_size,
|
||
saved_gprs, is_leaf, pdata_validated, has_eh
|
||
classes 851 name, vtable_address, rtti_present, base_classes_json
|
||
imports 398 library, ordinal, name, address
|
||
eh_funcinfo / eh_try_blocks 2 588 / 315
|
||
function_pointer_arrays 1 526 + 8 568 entries
|
||
indirect_dispatch_candidates 1 827 297 dispatch_pc, vtable_address, method_address
|
||
```
|
||
|
||
`instructions.raw` is an **INT, not hex** — a trap this corpus has already paid
|
||
for. Query with `python3 -c 'import duckdb'`; it is not SQLite.
|
||
|
||
### ⚠️ It is derived data, and it is wrong in places
|
||
|
||
The image is **primary** — those are the bytes the console executed. The database
|
||
is **an analysis of them**, produced by a disassembler that had to guess, and it
|
||
fails the way disassemblers fail:
|
||
|
||
* **misdecoded mnemonics** — data read as code, or a decoder-table gap, produces
|
||
a plausible instruction that was never executed as one;
|
||
* **wrong function boundaries** — `end_address` short or long, neighbours merged,
|
||
one function split in two;
|
||
* **incomplete coverage** — code reached only through indirect dispatch may be
|
||
absent entirely; the 1.8 M `indirect_dispatch_candidates` are *candidates*;
|
||
* **invented names** — largely derived rather than symbols, so a name is a
|
||
hypothesis wearing a label.
|
||
|
||
**A finding resting on a database row is not established until the bytes agree.**
|
||
Read the same address out of the `.pe` and check. Where they disagree the image
|
||
wins, and the disagreement is worth recording — it tells the next reader which
|
||
parts of the database to distrust.
|
||
|
||
It is a fast index into 9.2 MB of machine code. It is not a source of truth.
|
||
|
||
### These mounts are reference material, not a deliverable
|
||
|
||
They are read-only and they come from outside the repository, which means a
|
||
fresh checkout on another machine has neither. **Reimplementing the producer —
|
||
XEX decrypt + LZX decompress, and the disassembly-to-database step — belongs in
|
||
`crates/sylpheed-formats`.** Until then, every static finding rests on an
|
||
artefact this project cannot rebuild, and that is a real gap in the corpus rather
|
||
than a convenience.
|
||
|
||
## 🔴 Pressing Ⓐ on the title faults the guest — and the fault fills the disk
|
||
|
||
Three attempts to capture the main menu on 2026-08-29 ended the same way. Every
|
||
run that tapped Ⓐ **on the title** faulted; every run that tapped nothing there
|
||
completed and produced its capture.
|
||
|
||
| run | input on the title | outcome |
|
||
|---|---|---|
|
||
| 1 | Ⓐ, then Ⓐ again on the transition | guest fault, **519 MB** of register dump |
|
||
| 2 | one Ⓐ | drifted to a `flight` classification, 97 MB |
|
||
| 3 | one Ⓐ | guest fault, **223 MB** of register dump |
|
||
| 4–6 | none (`NOTAP=1`) | all completed normally |
|
||
|
||
This is the crash `ui_draw_capture.sh`'s own header records from 2026-08-18 — "a
|
||
stray A there sends the guest into the save-data probe". ⚠️ The corpus's existing
|
||
menu measurements (Q4, Q5, the focus ring) were taken by some route that survived
|
||
this; what differs has not been found. **Menu-side dynamic RE is blocked until it
|
||
is.**
|
||
|
||
⚠️ **A guest fault writes an UNBOUNDED register dump to stdout.** Xenia runs with
|
||
`break_on_unimplemented_instructions = true`, and the dump is `vN = [...]` / `rN =
|
||
...` lines at roughly 100 MB per 30 s. The filesystem here sits at **91 %**. Any
|
||
scripted run that presses a button must watch `canary.stdout` and kill on growth —
|
||
`ls -la` on it before trusting a long run.
|
||
|
||
📌 Two knobs added to `ui_draw_capture.sh` for boot-side work: `GRACE=1` (the fixed
|
||
8 s wait before arming means an `ARM=early` capture otherwise misses both splashes,
|
||
which run at ~1.2–9.5 s of guest time) and `NOTAP=1` (no input at all — the movie
|
||
tap fires on "the screen changed a lot", which is also true of a fading splash).
|
||
|
||
### ⚠️ `ARM=early` loses its F10 about 40 % of the time
|
||
|
||
Five `ui_draw_capture.sh ARM=early` runs on 2026-08-29: **two logged `ARMED EARLY`
|
||
and produced no `xenia_re_ui_draws_NN.log` at all.** The keypress goes to the
|
||
window and is silently lost — nothing in the session log distinguishes a run that
|
||
armed from one that did not, so **check the log file exists before spending the
|
||
run**, and treat a repeat measurement as needing more attempts than samples.
|
||
|
||
* 🔴 **`sylpheed-cli` in `$CARGO_TARGET_DIR` can be STALE, and `screen info` lies
|
||
quietly when it is.** The copy here was built **2026-08-29 12:38**, before the
|
||
keyframe-record-layout fix. The old parser shifted every keyframe time by one slot
|
||
and could not time a group's final pose, printing a trailing `-`:
|
||
|
||
```
|
||
stale pteff00.prm 4 kf rest t=70 [12:0,0 70:0,0 80:0,0 -:0,0]
|
||
fresh pteff00.prm 4 kf rest t=12 [ 0:0,0 12:0,0 70:0,0 80:0,0]
|
||
```
|
||
|
||
Both outputs are well-formed and neither announces its age. A whole page of this
|
||
corpus (`screen-transitions.md`) argued from *"there is exactly one untimed
|
||
keyframe"*, which was the stale parser's artefact.
|
||
|
||
⚠️ **`cargo build -p sylpheed-cli` before trusting `screen info`** — it takes 8 s
|
||
against a warm cache. ✅ Renders are **byte-identical** across the two binaries
|
||
(checked on `GP_TUTORIAL` build 0, max per-channel difference **0**), so
|
||
`screen render` output and anything derived from element identity, pivots or
|
||
keyframe *counts* is unaffected. It is the *times* that move.
|