[iterate-2V] VdSwap: stop bumping primary CP_RB_WPTR out-of-band (canary-faithful)

Ours' `vd_swap` wrote its 64-dword XE_SWAP block at the guest's reserved `buffer_ptr` slot AND then bumped the primary ring `CP_RB_WPTR` out-of-band via `state.gpu.extend_write_ptr_by(64)`. That bump was a bug: `buffer_ptr` (~0x4add6efc) is NOT inside the primary ring (base ~0x4adcd000, 8192 dwords) — it lives ~10k dwords past it, in the renderer indirect-buffer region. The bogus WPTR bump pushed the GPU read-pointer PAST the guest's real write-pointer; the drain treated the overshoot as a circular wrap and re-executed the splash's draw indirect-buffers ~2×, inflating draws to 78 (the real splash geometry is ~28 draws; 12 INDIRECT_BUFFERs vs the real 6). Canary's `VdSwap_entry` (xenia-canary xboxkrnl_video.cc:518-548) writes the fetch-constant patch + PM4_XE_SWAP + NOP pad into the reserved slot and returns — it NEVER touches CP_RB_WPTR. The guest advances the primary ring write-pointer itself via its own doorbell once it has populated the slot; swap-complete CP interrupts come only from the game's in-stream PM4_INTERRUPT packets, never from VdSwap. This fix removes only the out-of-band `extend_write_ptr_by(64)` call, keeping the buffer_ptr block write intact and byte-faithful to canary. Effect at `--gpu-inline -n 50M`: draws 78→28, INDIRECT_BUFFER 12→6 (re-execution artifact gone), swaps 4→2. The run now halts at ~19.27M instructions (worker threads exit) instead of spinning to 50M, because removing the corruption unmasks the real per-present-interrupt deadlock — the title loop needs a per-present PM4_INTERRUPT that the stalled game never submits. That deadlock is a SEPARATE, known gate tracked/addressed elsewhere; it is intentionally NOT papered over here. Re-baselined golden crates/xenia-app/tests/golden/sylpheed_n50m.json to the new honest values (regenerated twice, byte-identical). sylpheed_n2m.json is unaffected (draws=0 at 2M). cargo test --workspace: 675 passed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
[iterate-2U] VdGlobalDevice: allocate a real device cell so the swap counter (clock B) can advance
2026-06-14 19:58:05 +02:00 · 2026-06-14 16:20:08 +02:00
3 changed files with 40 additions and 25 deletions
--- a/crates/xenia-app/src/main.rs
+++ b/crates/xenia-app/src/main.rs
@@ -1540,8 +1540,19 @@ fn cmd_exec_inner(
                    mem.write_u32(addr, block);
                }
                ("xboxkrnl.exe", 0x01BE) => {
-                    // VdGlobalDevice — passed through to Vd* shims. Write 0.
-                    mem.write_u32(addr, 0);
+                    // VdGlobalDevice — a *pointer to* a global D3D-device cell.
+                    // Mirror xenia-canary RegisterVideoExports (xboxkrnl_video.cc:
+                    // 557-564): allocate a 4-byte cell, point the import slot at
+                    // it, and zero the cell. The guest's graphics init then stores
+                    // its device object INTO the cell (e.g. sub_824C6DC0 @
+                    // 0x824C6F18 `stw r31, 0([0x82000750])`), and the swap-complete
+                    // callback sub_824CE2B8 reads it back via the two-level
+                    // `[[VdGlobalDevice]+0]+15160` to bump the swap counter (clock
+                    // B). Writing 0 directly here (the old behaviour) made that
+                    // store land at address 0 and the swap counter never advance —
+                    // freezing the title-loop's per-frame manager update.
+                    let cell = alloc_zero(0x4, &mut mem, &mut kernel);
+                    mem.write_u32(addr, cell);
                }
                ("xboxkrnl.exe", 0x01C0) => {
                    // VdGpuClockInMHz
--- a/crates/xenia-app/tests/golden/sylpheed_n50m.json
+++ b/crates/xenia-app/tests/golden/sylpheed_n50m.json
@@ -1,9 +1,9 @@
 {
-  "instructions": 50000001,
-  "imports": 451500,
+  "instructions": 19274336,
+  "imports": 72513,
  "unimpl": 0,
-  "draws": 78,
-  "swaps": 4,
+  "draws": 28,
+  "swaps": 2,
  "unique_render_targets": 2,
  "shader_blobs_live": 3,
  "texture_cache_entries": 0
--- a/crates/xenia-kernel/src/exports.rs
+++ b/crates/xenia-kernel/src/exports.rs
@@ -2999,24 +2999,25 @@ fn vd_swap(ctx: &mut PpcContext, mem: &GuestMemory, state: &mut KernelState) {
    // xboxkrnl_video.cc:479. Currently skipped (see below).
    let _ = fetch_dwords; // silence unused — will be live again under the deferred path

-    // iterate-2T: mirror xenia-canary `VdSwap_entry` (xboxkrnl_video.cc:518-548)
+    // iterate-2V: mirror xenia-canary `VdSwap_entry` (xboxkrnl_video.cc:518-548)
    // FAITHFULLY. The game reserves 64 dwords (256 bytes) in the primary ring
    // at `buffer_ptr`; canary writes a `PM4_TYPE0(SHADER_CONSTANT_FETCH_00_0)`
    // fetch-constant patch followed by `PM4_TYPE3(PM4_XE_SWAP)`, then pads with
-    // NOPs. We do the same, then bump WPTR by 64 so the drain consumes the
-    // PM4_XE_SWAP **in command-stream order** — i.e. AFTER any in-stream
-    // callback-arming Type-0 writes the game already queued.
+    // NOPs — and **NEVER touches `CP_RB_WPTR`**. The game advances the primary
+    // ring write-pointer itself via its own doorbell once it has finished
+    // populating the reserved slot, so VdSwap only fills the bytes.
    //
-    // Why this matters (the iterate-2T root): the previous M2b short-circuit
-    // called `notify_xe_swap` directly from the HLE, which synthesized a CP
-    // swap-complete interrupt OUT OF BAND. When that interrupt reached the
-    // graphics ISR (`sub_824BE9A0`) before D3D had armed its swap-callback
-    // slot (`[gfx+10772]+16` still the `0xBADF00D` placeholder), the ISR hit
-    // its "ERR[D3D]: Unanticipated CPU_INTERRUPT. Sign of a corrupt command
-    // buffer?" assert (`twi` at 0x824BE9DC). Routing the swap through the ring
-    // packet keeps the interrupt naturally ordered after arming, matching
-    // canary (whose VdSwap raises NO interrupt itself; swap-complete CP
-    // interrupts come only from in-stream `PM4_INTERRUPT` packets).
+    // iterate-2V FIX (the bug this removes): a prior revision bumped the
+    // primary ring `CP_RB_WPTR` out-of-band here (`extend_write_ptr_by(64)`).
+    // But `buffer_ptr` (~0x4add6efc) is NOT inside the primary ring (base
+    // ~0x4adcd000, 8192 dwords) — it lives ~10k dwords past it, in the
+    // renderer indirect-buffer region. The bogus WPTR bump pushed the GPU
+    // read-pointer PAST the guest's real write-pointer, the drain treated the
+    // overshoot as a circular wrap, and **re-executed the splash's draw
+    // indirect-buffers ~2×** — inflating draws to 78 (real splash ≈ 28; 12
+    // INDIRECT_BUFFERs vs the real 6). Canary's `VdSwap_entry` writes the
+    // block and returns; the swap-complete CP interrupt comes only from the
+    // game's own in-stream `PM4_INTERRUPT` packets, never from VdSwap.
    if buffer_ptr != 0 {
        let mut off = 0u32;
        let mut put = |i: &mut u32, v: u32| {
@@ -3052,12 +3053,15 @@ fn vd_swap(ctx: &mut PpcContext, mem: &GuestMemory, state: &mut KernelState) {
            put(&mut off, xenia_gpu::pm4::make_packet_type2());
        }
    }
-    state.gpu.extend_write_ptr_by(64);
+    // NOTE: We deliberately do NOT bump `CP_RB_WPTR` here (see the iterate-2V
+    // comment above). The drain below consumes only the packets the game has
+    // legitimately advanced the write-pointer over.

-    // Drain the ring; the PM4_XE_SWAP we just queued (and any in-stream
-    // PM4_INTERRUPT) executes in order. The PM4_XE_SWAP handler calls
-    // `notify_xe_swap` for host swap bookkeeping; no synthetic interrupt is
-    // raised (see `notify_xe_swap`).
+    // Drain the ring up to whatever the game has actually submitted; any
+    // in-stream `PM4_INTERRUPT` / draw packets execute in order. The
+    // reserved-slot PM4_XE_SWAP is consumed by the GPU only once the game
+    // advances its own doorbell over it. The swap-counter safety net below
+    // keeps host swap bookkeeping live in the meantime.
    let drained = state.gpu.drain_to_current_wptr(mem);
    tracing::debug!(drained, "VdSwap: drained PM4 packets");