re: the A-press fault is SOLVED -- Xenia swallows input, the guest pump is unbounded

The 326 MB log from the failing run was still on disk, so this needed no
emulator time at all.

Mechanism: Xenia's XamInputGetKeystrokeEx returns X_ERROR_SUCCESS with a zeroed
keystroke on every call while a XAM dialog is up (xam_input.cc:197, upstream
Canary). The game's keystroke pump -- sub_82457038, read out of the image -- is
an unbounded 'while (GetKeystrokeEx() == SUCCESS) queue.push_back()'. It queued
8 388 608 empty keystrokes, grew its vector to 64 MB, asked for 128 MB, got a
failed allocation back unchecked, and copied off the top of the guest stack.

Two independent instruments agree to within 7: the Canary counter's last report
before the crash says 8 388 601 swallowed calls; the crash dump's r29 says the
vector held 8 388 608. The reporting granularity is 600.

Retracts this page's own 'r9 is a wild pointer above 4 GB'. Xenia prints
si_addr, a host address; the guest is mapped at 0x100000000, so the fault
address is guest 0x701D0000 -- which is exactly r9 in the register dump.

Also refutes nothing of the port's, but answers its ask #3: the two press-a
captures are different frames (40.84 % of the band's pixels differ at the
best alignment, which has a sharp minimum), so its 0.301 % is not an
instrument floor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
This commit is contained in:
sylph-decoder
2026-08-30 06:44:49 +00:00
parent 6d7bc87b0f
commit e94e203a71
5 changed files with 417 additions and 73 deletions

View File

@@ -1,91 +1,205 @@
# Pressing Ⓐ on the title faults the guest — diagnosed, and it is a third failure
# Pressing Ⓐ on the title faults the guest — SOLVED, and it is the emulator swallowing input
**Classification: measured.** Xenia Canary, 2026-08-29, four runs. This is the
blocker that gates every menu-side dynamic question in this container.
**Classification: measured** (the mechanism, from the run's own retained log) on a
**decoded** code path (the three functions, read out of the image). Xenia Canary,
2026-08-29. This is the blocker that gated every menu-side dynamic question in this
container, and it is not a mystery any more.
## What happens
## The one-line answer
A single Ⓐ press on the title screen produces a Xenia **`==== CRASH DUMP ====`**:
Xenia's `XamInputGetKeystrokeEx` returns **`X_ERROR_SUCCESS` with a zeroed
keystroke, on every call, for as long as a XAM dialog is up**. The game's
keystroke pump is `while (GetKeystrokeEx(...) == SUCCESS) queue.push_back(ks);`
with **no bound**. Something raised a XAM dialog immediately after the third Ⓐ was
delivered, and the pump then queued **8 388 608** empty keystrokes, grew its vector
to 64 MB, asked for 128 MB, got a failed allocation back **unchecked**, and copied
off the end of the guest thread stack.
So the fault is a *symptom two levels down* from an emulator-side input blackout.
Nothing is wrong with the disc, the title screen, or the Ⓐ button.
## 🔴 Retraction — "`r9` is a wild pointer, above 4 GB, never a guest address"
That is this page's own claim, written 2026-08-29 at `72e45a7`, and it is **wrong**.
`Access Violation: write at 0x00000001701D0000` prints `ex->fault_address()`, which
`exception_handler_posix.cc:154` fills from **`signal_info->si_addr`** — a *host*
address. Xenia maps the guest at `mapping_base_`, chosen in `memory.cc:193` as the
first `1ull << n` from n=32 that maps, i.e. **`0x100000000`**.
The register file proves the translation rather than assuming it: the faulting
instruction is `sth r6, 0(r9)` and the dump shows
```
PC: 0x824578A0
Access Violation: write at 0x00000001701D0000
r9 = 00000000701D0000 Access Violation: write at 0x00000001701D0000
```
It repeats **32 356 times** in one run, writing **326 MB** of register dump to
stdout in roughly ten seconds. Four runs that pressed Ⓐ on the title all faulted;
four runs in the same sessions that pressed nothing all completed normally.
`0x1701D0000 0x100000000 = 0x701D0000 = r9`. So `r9` **is** a guest address, in
the `v40000000` heap (`0x40000000 … 0x7EFFFFFF`), and the page is simply not
committed. The distinction matters: "garbage pointer" pointed the next probe at
memory corruption; the truth points it at an allocation that failed.
## 🔴 Refuted: my own hypothesis, that it was an unimplemented instruction
⚠️ **Generalise this.** Every `Access Violation: … at 0x1________` in a Canary log
from this container is a guest address plus `0x100000000`. Subtract before reading.
The config dump carries `break_on_unimplemented_instructions = true`, and Xenia's
own message for that path reads *"report the game to Xenia developers; to skip,
disable break_on_unimplemented_instructions"* — so the flag looked like the fix.
## The code path, read out of the image (0 mismatches against `sylpheed.db`)
**It is not.** Booting with `--break_on_unimplemented_instructions=false` faulted
identically, and **no `Unimplemented instr` line is ever logged**, on stdout or
stderr, in any run. That path emits its `XELOGE` *before* the guarded
`DebugBreak()`, so its absence rules the mechanism out rather than leaving it open.
All three functions were disassembled from `/image/sylpheed.pe` and cross-checked
word-for-word against the database: **466 + 120 instructions, zero disagreements**
across `sub_82457038`, `sub_82457780` and their callees.
The dump comes from `Emulator::ExceptionCallback` (`emulator.cc:1468`), which fires
on a **genuine guest exception** — an access violation or illegal instruction
inside guest code — not on a translation failure.
## The faulting instruction, read from the image
`sylpheed.db` puts the PC inside `sub_82457780` (0x82457780…0x82457958, non-leaf,
frame 144), +0x120 in. The image is primary and gives the instruction itself:
| address | word | instruction |
| | what it is | how that is known |
|---|---|---|
| **0x824578A0** | `b0c90000` | **`sth r6, 0(r9)`** ← faults |
| 0x824578A4 | `b0a90002` | `sth r5, 2(r9)` |
| 0x824578A8 | `b0890004` | `sth r4, 4(r9)` |
| 0x824578AC | `b1490006` | `sth r10, 6(r9)` |
| 0x824578B0 | `409affcc` | `bne-` — loops back |
| `sub_824574C0` | lazy singleton getter for the **input manager** at guest `0x828F3888`, guarded by a bit-0 "constructed" flag at `0x828F3A70` | `lis r11,0x828F; addi r30,r11,14472` = `0x828F3888`; classic guard-variable shape |
| `sub_82457038` | the **keystroke pump**: drains `XamInputGetKeystrokeEx` into a vector at `this+68` = `0x828F38CC` | calls `sub_824AA870`, which is `b 0x8284DBDC` = the **`XamInputGetKeystrokeEx`** import thunk (`imports`, ordinal 408) |
| `sub_82457780` | that vector's **insert-with-grow** | `{ptr@+0, size@+4, capacity@+8}`; doubles capacity, clamps at `0x1FFFFFFF`, `slwi r3,r27,3` for the byte count |
Four consecutive halfword stores at offsets 0/2/4/6 through **`r9`**, inside a
loop: the code is filling an array of 8-byte records with four `u16` fields each.
The element is **8 bytes copied as four halfwords** at offsets 0/2/4/6 — which is
exactly `X_INPUT_KEYSTROKE` `{u16 VirtualKey; u16 Unicode; u16 Flags; u8 UserIndex;
u8 HidCode}`. That is what makes the vector identifiable as a keystroke queue and
not some other 8-byte record.
**`r9` is a wild pointer.** The faulting address `0x1701D0000` is above 4 GB and so
outside the guest's 32-bit address space entirely — not a null dereference and not
a small overrun, but a base that was never a guest address.
The pump, in C:
## It is a *third* failure mode, not either known one
```c
// sub_82457038, 0x82457174 … 0x824571C8
while (XamInputGetKeystrokeEx(&user, 3, &ks) == X_ERROR_SUCCESS) {
if (v->size < v->capacity) v->data[v->size++] = ks; // 0x8245718C
else insert_slow(v, end, &ks); // 0x824571B0 → sub_82457780
}
```
| | PC | crash dumps | this |
There is no iteration cap and no check on the allocator's return.
## The emulator half — `xam_input.cc:197`
```cpp
if (kernel_state()->xam_state()->IsUIActive()) {
...
return X_ERROR_SUCCESS; // keystroke was zeroed above
}
```
`IsUIActive()` is `is_xam_dialog_present_`, set to true by every non-headless
`XamShow*UI` path in `xam_ui.cc` and cleared only by a dialog's close handler. While
it is set, the guest's `== SUCCESS` loop can never terminate.
⚠️ This is **upstream Canary behaviour**, not one of this container's RE patches.
The RE patch is only the `[RE-INPUT]` logging around it — and that logging is what
made the diagnosis possible, so it earned its keep.
## The number that closes it
The instrumentation reports one line per 600 swallowed calls. Immediately before the
first crash dump:
```
[RE-INPUT] XamInputGetKeystrokeEx swallowed by IsUIActive (ui_active=true, 8388601 so far)
```
and the crash dump's own registers say how many records the vector held:
```
r29 = 0000000000800000 = 8 388 608 elements to copy
r26 = 0000000000800001 = new size
r27 = 0000000001000000 = new capacity (doubled)
r30 = FFFFFFFF828F38CC = the vector object — the pump's queue
r31 = 00000000A7AC0000 r7 = 00000000A3AC0000 → 0x04000000 = 64 MB of live data
```
**8 388 601 swallowed calls against 8 388 608 queued records — a gap of 7, inside
the 600-call reporting granularity.** One push per swallowed poll. The two numbers
are independent instruments (a Canary log counter and a guest register file) and
they agree; that is the whole argument, and it needs no further run.
## Why `r3` looked like a stack pointer
`slwi r3, r27, 3` = `0x8000000` = **128 MB** requested from `sub_824F7240`
(`b 0x82150000`, the game's `heap_alloc(*0x828E2B14, size, &out)` wrapper). It came
back as `0x701CF5F0`**below** the pump thread's own `r1 = 0x701CF7B0`, i.e. a
pointer into a stack frame that had already been popped. The copy then walked
`+0xA18` and hit the top of the thread's 64 KB stack at `0x701D0000`.
So: the allocation failed, the failure path left a stale `&local` in `r3`, and the
caller never checked. A 128 MB request on top of a live 64 MB one, in a guest with
512 MB total, is not a surprising failure.
## The timeline, from the log
| log line | event |
|---|---|
| 1149 | first `XamInputGetKeystrokeEx` reaches a driver |
| 11851252 | **three** Ⓐ press/release pairs delivered — `vk=5800`, flags `0001` down / `0002` up |
| 1253 | the third Ⓐ **up** is handed to the guest |
| **1254** | `swallowed by IsUIActive (ui_active=true, 1 so far)` — the blackout starts |
| 125415242 | 13 982 swallow reports = ~8.39 M swallowed calls |
| 15243 | first `==== CRASH DUMP ====`, `PC 0x824578A0` |
Evidence: [`../data/a-press-fault-log-extract.txt`](../data/a-press-fault-log-extract.txt).
⚠️ Note the feedback loop that produces **32 356** dumps rather than one: a guest
crash makes Xenia call `ImGuiDialog::ShowMessageBox` (`emulator.cc:1487`), which is
itself a UI — so the swallow can only get worse after the first fault.
## 🟡 What is still open: *which* dialog
`is_xam_dialog_present_` is a single bool with no logging on the setting side, so the
log says a XAM dialog exists and not which one. The static reach:
* `XamShowDeviceSelectorUI`**ruled out**. The run's own config dump carries
`storage_selection_dialog = false`, and `xam_ui.cc:550` takes the headless path in
that case, which never sets the flag.
* `XamShowSigninUI` and `XamShowMessageBoxUIEx` — both set it, both are imported, and
both reach the guest. `sub_821D03A0` calls the game's `XamShowSigninUI` **and**
`XamShowDeviceSelectorUI` wrappers, which is the shape of a "sign in, then pick
storage" flow — exactly what a title screen's Ⓐ would start.
* `XamShowKeyboardUI`, `XamShowDirtyDiscErrorUI` — imported; not excluded, but neither
fits the moment.
**The experiment that settles it** is one line of Canary, not another blind boot: log
the function name at each `is_xam_dialog_present_.store(true)` site in `xam_ui.cc`.
Recorded rather than done, because it is an emulator change and this iteration's
budget went to the diagnosis.
## What this unblocks, and how
The blocked list — main-menu sweeps, whether a `.tbm` draws pixels, `pbafc.prm`'s
blend — needs a screen behind an Ⓐ press. Three routes now exist where before there
were none, in cost order:
1. **Dismiss the dialog.** It is an ImGui window on the emulator surface; the run had
a display. If it can be clicked or key-dismissed, the flag clears and the pump
drains normally. Cheapest, and testable in one boot.
2. **`--headless`.** Both `xeXamShowSigninUI` and `XamShowMessageBoxUIEx` take a
dispatch-headless path that never sets the flag. ⚠️ It also removes the window the
capture harness grabs, so this trades the fault for a different blocker.
3. **Patch the swallow.** Returning `X_ERROR_EMPTY` instead of `X_ERROR_SUCCESS` at
`xam_input.cc:217` terminates the pump immediately and is closer to hardware
(a real Xbox does not hand a game an infinite run of empty keystrokes). This is
an emulator change and needs to be recorded as one wherever it is used.
⚠️ **Whichever route is taken, keep `tools/re-capture/frame_clock.sh`'s size guard.**
It killed this run at its 300 MB cap and worked exactly as designed; without it the
next fault fills a filesystem that was already at 91 %.
## It was never the same failure as the other two
| | PC | crash dumps | cause |
|---|---|---|---|
| cache-flush crash ([`title-crash-stl-tree.md`](../title-crash-stl-tree.md)) | `0x82307128` | yes | different PC |
| loader stall ([`canary-scripted-input-traps.md`](../canary-scripted-input-traps.md)) | — | **zero** | ❌ this has 32 356 |
| **this** | **`0x824578A0`** | 32 356 | |
| cache-flush crash ([`../title-crash-stl-tree.md`](../title-crash-stl-tree.md)) | `0x82307128` | yes | different |
| loader stall ([`../canary-scripted-input-traps.md`](../canary-scripted-input-traps.md)) | — | **zero** | different |
| **this** | `0x824578A0` | 32 356 | **emulator input blackout → unbounded guest queue** |
⚠️ So the advice in the input-traps page — *"retry whole boots"*, because that
failure is not deterministic — does not obviously apply: this one reproduced on
**4 of 4** attempts.
And it explains the thing the old page could not: **why Q4 and Q5 pressed Ⓐ
successfully and these runs did not.** Nothing about the game differs. What differs
is whether a XAM dialog happened to be up, which is emulator state, not guest state
— so "it reproduced 4/4" and "it worked before" are both true and always were.
⚠️ **And it does not explain the corpus's earlier successes.** Q4 and Q5 measured
all five main-menu buttons, and the focus ring was timed on the menu, so Ⓐ worked
then. What differs between those runs and these has not been found.
## 🔴 Also refuted: the earlier "unimplemented instruction" hypothesis
## The disk hazard, quantified
A guest fault writes an **unbounded** register dump: 32 register lines, 32 float
lines and 128 vector lines per fault, at ~30 MB/s, on a filesystem at 91 %.
`tools/re-capture/frame_clock.sh`'s size guard killed this run at its 300 MB cap
and worked exactly as designed — the session log's `EMULATOR GONE at 56s` is the
guard, not the crash. **Any scripted run that presses a button needs it.**
## What this blocks
The main-menu sweeps; whether a `.tbm` draws pixels (needs `GP_SAVE_LOAD`,
`GP_BUNK` or `GP_DEBRIEFING_PILOTLOG`); `pbafc.prm`'s blend (needs
`GP_READY_ROOM`); and any other menu-side question.
## Where the next probe goes
At **`sub_82457780`** and what sets `r9`. The function is non-leaf with a 144-byte
frame; the loop writes 8-byte records. Whether `r9` comes from an allocation that
failed, a table the guest expects the loader to have filled, or a pointer read back
from a structure, is the question — and the corpus's note that the *loader thread*
does no file I/O in a failed run is suggestive but is a different failure.
Kept from the previous version of this page because the negative still stands.
`break_on_unimplemented_instructions = true` looked like a one-flag fix; booting with
it false faults identically, and **no `Unimplemented instr` line is ever logged**.
That path emits its `XELOGE` *before* the guarded break, so its absence rules the
mechanism out. The dump comes from `Emulator::ExceptionCallback`, which fires on a
genuine guest exception.