Snapshot for handoff. Contains the mission-audio + threading fixes and the
crash-investigation instrumentation (all diagnostic cvars default-OFF).
Fixes (behavioral):
- threading_posix.cc: reap-once guard on PosixCondition<Thread>::post_execution
(double pthread_join at mission teardown -> fault loop -> audio death + freeze).
- xma_decoder.cc: work_event_->Set() in Pause() so the idle XMA worker observes
paused_ and signals pause_fence_ (Pause() deadlock -> permanent audio death).
- audio_system / xma_context_master / xboxkrnl_audio / apu_flags / alsa: mission
audio keepalive + guest_audio_flags + watchdogs.
Instrumentation (additive, default-off): xboxkrnl_debug cache-throw diag +
guest-catch dispatcher, xex_module PE/PDATA/EH scans, kernel_state mem_watch
(NOTE: mem_watch DEFAULTS TRUE -- an always-on host poll thread; prime crash suspect).
NEW ORACLE (see HANDOFF-crash-oracle-2026-07-16.md): stock 6e5b8324f built with
our toolchain + zero custom code = NO crash, NO sound-stop, plays the Ready Room.
=> the Ready-Room out_of_range crash is introduced by THESE changes, not the game
and not the (LTO-broken) build chain. Bisection plan + suspect ranking in the note.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
6.0 KiB
HANDOFF — Ready-Room crash: NEW ORACLE says it's OUR code, not the game (2026-07-16)
TL;DR
The mid-game crash (guest std::out_of_range from the cache-manager flush, seen in the
Ready Room after a mission) is introduced by our own changes, not by the game and not by the
build toolchain. A clean bisection of our 19 commits + uncommitted working tree will find it.
The new oracle (this is the key result)
We built stock upstream 6e5b8324f [APU] Pace audio subsystem with our own toolchain and
zero custom code (0 custom cvars in the binary — verified). The user ran it:
Stock
6e5b8324f→ NO crash, NO sound-stop, played through the Ready Room fine.
This kills two earlier hypotheses and re-frames the whole problem:
| Hypothesis | Verdict |
|---|---|
| "Genuine game data race, unfixable without masking" | WRONG. Stock game code is stable. |
| "Our build chain miscompiles (broken LTO) → crash" | WRONG. Same toolchain, stock source, no crash. |
| "The crash is in OUR 19 commits + 16 working-tree files" | This is where it is. Bisect it. |
The out_of_range race is real, but it's our timing changes that push a concurrent cache-add into the flush's unlocked iteration window. Stock timing never lands there.
Reproduce the oracle
cd "/home/fabi/RE - Project Sylpheed/xenia-canary-native"
git stash push -u
git checkout 6e5b8324f4101464de0f8c2334edb03cac8826c4
git submodule update --init third_party/DirectXShaderCompiler third_party/snappy # sync to stock pins
cmake --build build --config Release --target xenia_canary -j16 # reuses object cache
# copy bin/Linux/Release/xenia_canary aside, then restore:
git checkout <this-branch>
git -C third_party/DirectXShaderCompiler checkout dc3e6c48d
git -C third_party/snappy checkout 77c78fad
git stash pop
A prebuilt stock binary from this session is at ../stock-oracle/xenia_canary_STOCK_6e5b832
(see "Artifacts" below). Run any binary through the launcher with CANARY_BIN=<path> ./run-canary-native.sh.
Build-chain note (real, but NOT the crash cause)
Our LTO is silently disabled: clang++ is LLVM 18.1.3 but the linker's gold LTO plugin is
LLVM 17.0.6, so the plugin rejects clang-18 bitcode —
failed to create LTO module: Unknown attribute kind (91) ×420 (≈ every TU). ThinLTO is not
applied; our codegen differs from the fully-LTO'd official AppImage. Harmless for the crash (proven by
the oracle) but worth fixing for release parity: install/point at a matching LLVMgold.so (llvm-18),
or build ld.lld+-fuse-ld=lld which handles LTO in-tree, or disable LTO explicitly.
Bisection plan (suspects ranked by how much they perturb guest thread timing)
Each test = one build + one Ready-Room run. Do them in order; stop at the first that removes the crash.
mem_watchpolling thread — FREE, no rebuild.kernel_state.ccadds a detached host thread (default--mem_watch=true) that wakes every 3 s and reads the guest cache-manager memory — the exact struct the crash is about. Test: run our binary with--mem_watch=false. → no crash = default it off / remove it; keep every real fix. Best case.- Threading edits (uncommitted, load-bearing).
threading_posix.ccreap-once double-join fix +xma_decoder.cc work_event_->Set()inPause(). These cure the mission-teardown freeze + permanent audio-death (seeproject_audio_stall_freeze_diagnostics_2026_07_14). Rebuild with them reverted to stock; test. → If these are the trigger, it's a real trade-off (crash vs freeze/audio-death) and needs a smarter fix, e.g. make the guest cache-flush loop (PC0x8245A154..0x8245A1CC) atomic w.r.t. other guest threads instead of reverting the reap fix. - Other audio-fix files:
audio_system.cc(+218),xma_context_master.cc(+44),xboxkrnl_audio.cc,apu_flags.*,alsa_audio_driver.cc. - Committed audio fix
f10484834 [APU] Fix mission-audio silence. - event_log Phase-A trampoline per-export overhead (
shim_utils.h, committed).
Faster alternative to 2–5: git stash the uncommitted tree and test the committed HEAD alone — splits
"uncommitted working changes" from "the 19 commits" in one run (caveat: without the threading fixes the
mission-teardown freeze may hit before you reach the Ready Room; if it freezes instead of crashing,
that itself localizes the reap fix as load-bearing).
Crash mechanism (already characterized — for context)
Guest cache manager (VA 0x828F4838). Flush sub_8245A098 snapshots the block deque under a critsec,
releases the lock, then iterates the ~654-entry deque unlocked doing map::at (0x823070B0,
throws at 0x8230711C). A concurrent locked cache-add lands in the µs window → new deque key absent
from the frozen snapshot → out_of_range. Cascade count = 1 (a clean single-add TOCTOU). xenia never
dispatches the guest C++ throw, so it just crashes. Diagnostics for this live behind --cache_throw_diag
(default off) in xboxkrnl_debug.cc.
Current repo state
- Branch
phase-a-args-fileread(fork remotegit.mc02.dev/fabi/Xenia-Canary.git). - This commit = WIP snapshot of all audio/threading fixes + investigation instrumentation (all diag cvars default-OFF). 1139 insertions / 16 src files + 2 new headers + DXC submodule bump.
build/currently holds stock objects (we built the oracle there). To get OUR binary back:cmake --build build --config Release --target xenia_canary -j16(~148 TU).- Cleanup owed before release: strip/gate the diagnostic instrumentation
(
LogGuestThrow/DumpCacheManagerOnThrow/DispatchGuestCatchinxboxkrnl_debug.cc, thexex_module.ccPE/PDATA/EH scans,kernel_state.ccmem_watch). The broken tree-walk node layout inDumpCacheManagerOnThrow(finds 1–17 of 654 nodes) is only used by the moot orphan diag — low prio.
Artifacts (this session, scratchpad — EPHEMERAL, copy out if needed)
xenia_canary_OURS— our full build (the crashing one)xenia_canary_STOCK_6e5b832— stock oracle (the stable one)