# HANDOFF — Ready-Room crash: NEW ORACLE says it's OUR code, not the game (2026-07-16) ## TL;DR The mid-game **crash** (guest `std::out_of_range` from the cache-manager flush, seen in the **Ready Room** after a mission) is **introduced by our own changes**, not by the game and not by the build toolchain. A clean bisection of our 19 commits + uncommitted working tree will find it. ## The new oracle (this is the key result) We built **stock upstream `6e5b8324f` `[APU] Pace audio subsystem`** with our own toolchain and **zero custom code** (0 custom cvars in the binary — verified). The user ran it: > **Stock `6e5b8324f` → NO crash, NO sound-stop, played through the Ready Room fine.** This kills two earlier hypotheses and re-frames the whole problem: | Hypothesis | Verdict | |---|---| | "Genuine game data race, unfixable without masking" | **WRONG.** Stock game code is stable. | | "Our build chain miscompiles (broken LTO) → crash" | **WRONG.** Same toolchain, stock source, no crash. | | "The crash is in OUR 19 commits + 16 working-tree files" | **This is where it is.** Bisect it. | The out_of_range race is real, but it's **our timing changes** that push a concurrent cache-add into the flush's unlocked iteration window. Stock timing never lands there. ### Reproduce the oracle ```bash cd "/home/fabi/RE - Project Sylpheed/xenia-canary-native" git stash push -u git checkout 6e5b8324f4101464de0f8c2334edb03cac8826c4 git submodule update --init third_party/DirectXShaderCompiler third_party/snappy # sync to stock pins cmake --build build --config Release --target xenia_canary -j16 # reuses object cache # copy bin/Linux/Release/xenia_canary aside, then restore: git checkout git -C third_party/DirectXShaderCompiler checkout dc3e6c48d git -C third_party/snappy checkout 77c78fad git stash pop ``` A prebuilt stock binary from this session is at `../stock-oracle/xenia_canary_STOCK_6e5b832` (see "Artifacts" below). Run any binary through the launcher with `CANARY_BIN= ./run-canary-native.sh`. ## Build-chain note (real, but NOT the crash cause) Our LTO is silently disabled: `clang++` is **LLVM 18.1.3** but the linker's gold LTO plugin is **LLVM 17.0.6**, so the plugin rejects clang-18 bitcode — `failed to create LTO module: Unknown attribute kind (91)` **×420** (≈ every TU). ThinLTO is not applied; our codegen differs from the fully-LTO'd official AppImage. Harmless for the crash (proven by the oracle) but worth fixing for release parity: install/point at a matching `LLVMgold.so` (llvm-18), or build `ld.lld`+`-fuse-ld=lld` which handles LTO in-tree, or disable LTO explicitly. ## Bisection plan (suspects ranked by how much they perturb guest thread timing) Each test = one build + one Ready-Room run. Do them in order; stop at the first that removes the crash. 1. **`mem_watch` polling thread — FREE, no rebuild.** `kernel_state.cc` adds a detached host thread (default **`--mem_watch=true`**) that wakes every 3 s and *reads the guest cache-manager memory* — the exact struct the crash is about. Test: run our binary with `--mem_watch=false`. → no crash = default it off / remove it; keep every real fix. **Best case.** 2. **Threading edits (uncommitted, load-bearing).** `threading_posix.cc` reap-once double-join fix + `xma_decoder.cc work_event_->Set()` in `Pause()`. These *cure* the mission-teardown freeze + permanent audio-death (see `project_audio_stall_freeze_diagnostics_2026_07_14`). Rebuild with them reverted to stock; test. → **If these are the trigger, it's a real trade-off** (crash vs freeze/audio-death) and needs a smarter fix, e.g. make the guest cache-flush loop (PC `0x8245A154..0x8245A1CC`) atomic w.r.t. other guest threads instead of reverting the reap fix. 3. Other audio-fix files: `audio_system.cc` (+218), `xma_context_master.cc` (+44), `xboxkrnl_audio.cc`, `apu_flags.*`, `alsa_audio_driver.cc`. 4. Committed audio fix `f10484834 [APU] Fix mission-audio silence`. 5. event_log Phase-A trampoline per-export overhead (`shim_utils.h`, committed). Faster alternative to 2–5: `git stash` the uncommitted tree and test the committed HEAD alone — splits "uncommitted working changes" from "the 19 commits" in one run (caveat: without the threading fixes the mission-teardown freeze may hit before you reach the Ready Room; if it freezes instead of crashing, that itself localizes the reap fix as load-bearing). ## Crash mechanism (already characterized — for context) Guest cache manager (VA `0x828F4838`). Flush `sub_8245A098` snapshots the block deque under a critsec, releases the lock, then iterates the ~654-entry deque **unlocked** doing `map::at` (`0x823070B0`, throws at `0x8230711C`). A concurrent locked cache-add lands in the µs window → new deque key absent from the frozen snapshot → `out_of_range`. Cascade count = 1 (a clean single-add TOCTOU). xenia never dispatches the guest C++ throw, so it just crashes. Diagnostics for this live behind `--cache_throw_diag` (default off) in `xboxkrnl_debug.cc`. ## Current repo state - Branch **`phase-a-args-fileread`** (fork remote `git.mc02.dev/fabi/Xenia-Canary.git`). - This commit = WIP snapshot of all audio/threading fixes + investigation instrumentation (all diag cvars default-OFF). 1139 insertions / 16 src files + 2 new headers + DXC submodule bump. - `build/` currently holds **stock** objects (we built the oracle there). To get OUR binary back: `cmake --build build --config Release --target xenia_canary -j16` (~148 TU). - Cleanup owed before release: strip/gate the diagnostic instrumentation (`LogGuestThrow`/`DumpCacheManagerOnThrow`/`DispatchGuestCatch` in `xboxkrnl_debug.cc`, the `xex_module.cc` PE/PDATA/EH scans, `kernel_state.cc` mem_watch). The broken tree-walk node layout in `DumpCacheManagerOnThrow` (finds 1–17 of 654 nodes) is only used by the moot orphan diag — low prio. ## Artifacts (this session, scratchpad — EPHEMERAL, copy out if needed) - `xenia_canary_OURS` — our full build (the crashing one) - `xenia_canary_STOCK_6e5b832` — stock oracle (the stable one)