Recorded as a blocker rather than worked around, because it changes what the
next session can plan.
The faithful-capture route is an internal tap at SDLAudioDriver::SubmitFrame,
which receives exactly frame_size_ bytes of the guest s own frame in guest order
with no wall clock in the loop. A cvar-gated WAV writer there would record what
the guest PRODUCED rather than what a device CONSUMED, so it would be gap-free
however slowly the emulator runs -- which is precisely the defect that made both
ADV captures unusable.
The change is small. The build is not. build-canary builds
${PROJECT_DIR:-/work}/xenia-canary, which does not exist in this container; the
source is at /canary. The warm 235 MB tree at /sylph-home/re/canary-build is
configured with CMAKE_HOME_DIRECTORY=/work/xenia-canary, also missing, and its
build-Release.ninja carries no per-file rules -- it re-runs CMake first, and that
reconfigure fails on the absent root. So any Canary change is a full reconfigure
against /canary plus a full compile, at SYLPH_JOBS=4 on a box sitting at about
700 MB free with a documented history of full-parallel builds OOM-killing the
host.
Not attempted: that is a whole session s risk for one probe, and the next session
should decide with the cost in front of it rather than discover it halfway
through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QsEPXWVaEpyfudtR6re1Pd
9.1 KiB
Notes for an agent working inside this container
Read this before starting a dynamic-RE run. Everything here is something that already went wrong once.
The container fixes three old traps for you
- The display outlives the turn. Xvfb and openbox are children of PID 1, not of your shell. The old "Xvfb dies on its own every few minutes" note is gone — you no longer have to wrap a whole session in one blocking foreground call to keep it alive.
- The toolchain is real.
tools/re-capture/rebuild_canary.shexists because the old box had no cmake/ninja/clang and only runtime sonames, so it hand- relinked object files. Do not use it here. Usebuild-canary. 🔴 Butbuild-canarydoes not work in this container as it stands (2026-08-29). It builds${PROJECT_DIR:-/work}/xenia-canary, which does not exist here — the Canary source is at/canary($XENIA_SRC). The warm 235 MB tree at/sylph-home/re/canary-buildis configured withCMAKE_HOME_DIRECTORY=/work/xenia-canary, also missing, and itsbuild-Release.ninjacarries no per-file rules — it wants to re-run CMake first, which would fail on the absent source root. So any Canary change is a full reconfigure against/canaryplus a full compile, not an incremental one. Budget for that before starting:SYLPH_JOBS=4, and this box has been sitting at ~700 MB free with a documented history of full-parallel builds OOM-killing the host. - numpy and Pillow are installed.
entities2.py,flight_probe.pyand the image oracles work. Their absence used to look like a logic bug.
Method (the part that matters more than the tooling)
- Measure the oracle; never infer it. A session with zero Canary runs is a red flag.
- Trace upstream to where data first goes wrong, rather than patching the symptom you can see.
- Try to refute before believing. Record demotions rather than editing them
away —
docs/re/README.mdhas the ✅/🟡/❔ convention, and a withdrawn result is more useful than a quietly deleted one. - A probe that never performs the action will "prove" the action does not exist. The "targeting is automatic" conclusion came from a sweep that only ever tapped once; target select is Ⓐ pressed twice.
- Do not poll faster than the guest updates — it manufactures a clean curve
out of noise.
rate-curve-aliased-BAD.csvis committed as the bad example.
Running the emulator
run-canary # correct audio/pad/display flags baked in
pad.py tap A ; pad.py dpad down # scripted input (--hid=file, no uinput)
screenshot ~/shots/now.png # cropped to the GAME surface, not the window
python3 tools/re-capture/gmem.py find hex:820af844 400
-
One emulator at a time.
run-canaryenforces it with a lockfile. -
🔴
run-canaryis SILENT TWICE OVER, and that defeatsaudio-capture. Line 82 isexport SDL_AUDIODRIVER="${SDL_AUDIODRIVER:-dummy}", and its header explains why:--apu=nopstalls the guest in the intro movie, so the SDL driver against a dummy device is what lets the title advance. The comment's premise — "there is no PulseAudio here" — stopped being true whentools/audio-capturelanded, and it starts a daemon on demand. So a capture through the null sink records pure silence, at the right length, with a perfectly healthy-looking run behind it. To actually record the game:⚠️ And that is only the first of TWO layers.
run-canaryalso passes--mute=trueon its own command line (line 98). With the driver fixed and the mute left alone, Canary attaches a healthy 6-channel stream to the sink, holds it at 100 % volume for the whole run — and emits silence. Both have to go:audio-capture start # or load a null sink yourself PULSE_SINK=cap SDL_AUDIODRIVER=pulseaudio \ run-canary --mute=false … # `"$@"` is last, so this winsRecord at the monitor's real format, too —
parecdefaults to stereo/44.1 kHz and will silently resample a 6-channel monitor:parec -d cap.monitor --channels=6 --rate=48000 --format=s16le.⚠️ Do not read Canary's 6-channel PulseAudio stream as evidence the GAME is 5.1.
pactlwill showfloat32le 6ch 48000Hz, channel-mapped to a full 5.1 layout, on any title. That isAudioDriver::kFrameChannelsDefault = 6, a hardcoded constant — the code path actually used (SDLAudioSystem::CreateDriver(index, semaphore, &driver)) constructsSDLAudioDriver(semaphore)and takes every default. The format is Xenia's; only the content of those six channels is the guest's.⚠️ Check
pactl list sink-inputsbefore trusting a recording. If it is empty, Canary never attached and you are recording zeroes; the sink also sits atIDLE.audio-capture runwarns on a-infpeak afterwards, which is the backstop — but a live check fails in seconds instead of after the whole run. -
Boot is slow cold, ~25 s once the shader/code caches are warm — so a launch-and-dump fits in a single call.
-
Screens: classify by whole-image statistics (
screen_id.py), not named pixels. Named-pixel oracles are only valid while the game image sits at a known place, and nothing errors when it moves.
Verifying your own work
- Reborn's disc-gated tests self-skip without
SYLPHEED_DISC. A green run with it unset means almost nothing.build-reborn testwires it up for you. - Prefer a headless self-verify over "it compiles":
sylpheed-cli mesh render,screen render,save infoall produce checkable artifacts. - A Bevy system-parameter conflict is invisible to the type checker and panics at startup. If you touch viewer systems, run the binary, don't just build it.
Reporting
State what you measured, what you assumed, and what you could not settle. If a result is withdrawn, say so and keep the reasoning — that is the corpus's whole convention, and the reason its numbers can be trusted.
The static-analysis corpus — mounted, not reproducible
Two things arrive read-only from the host because nothing in this repository can produce them yet:
| path | what | env |
|---|---|---|
/xenia-rs/sylpheed.db |
the disassembly database, 586 MB | SYLPHEED_DB |
/image/sylpheed.pe |
the decompressed image, flat VA dump | SYLPHEED_PE |
The .pe is a flat VA dump: file offset = VA - 0x82000000
(SYLPHEED_IMAGE_BASE). So reading 0x820A1630 is seek(0xA1630) — no XEX
decrypt, no LZX, and no booted emulator. An earlier belief that this file was
stale was tested and refuted; it is current.
Recovering the image by dumping /dev/shm/xenia_memory_* also works and
self-validates, but it needs a running emulator — a poor dependency for
something the whole static corpus rests on. Use the file.
The database is far richer than the four scripts that read it use:
functions 25 481 address, name, end_address, frame_size,
saved_gprs, is_leaf, pdata_validated, has_eh
classes 851 name, vtable_address, rtti_present, base_classes_json
imports 398 library, ordinal, name, address
eh_funcinfo / eh_try_blocks 2 588 / 315
function_pointer_arrays 1 526 + 8 568 entries
indirect_dispatch_candidates 1 827 297 dispatch_pc, vtable_address, method_address
instructions.raw is an INT, not hex — a trap this corpus has already paid
for. Query with python3 -c 'import duckdb'; it is not SQLite.
⚠️ It is derived data, and it is wrong in places
The image is primary — those are the bytes the console executed. The database is an analysis of them, produced by a disassembler that had to guess, and it fails the way disassemblers fail:
- misdecoded mnemonics — data read as code, or a decoder-table gap, produces a plausible instruction that was never executed as one;
- wrong function boundaries —
end_addressshort or long, neighbours merged, one function split in two; - incomplete coverage — code reached only through indirect dispatch may be
absent entirely; the 1.8 M
indirect_dispatch_candidatesare candidates; - invented names — largely derived rather than symbols, so a name is a hypothesis wearing a label.
A finding resting on a database row is not established until the bytes agree.
Read the same address out of the .pe and check. Where they disagree the image
wins, and the disagreement is worth recording — it tells the next reader which
parts of the database to distrust.
It is a fast index into 9.2 MB of machine code. It is not a source of truth.
These mounts are reference material, not a deliverable
They are read-only and they come from outside the repository, which means a
fresh checkout on another machine has neither. Reimplementing the producer —
XEX decrypt + LZX decompress, and the disassembly-to-database step — belongs in
crates/sylpheed-formats. Until then, every static finding rests on an
artefact this project cannot rebuild, and that is a real gap in the corpus rather
than a convenience.