[iterate-4A] jit M0: recompiler seam + in-process differential harness
Foundation for a staged PPC block-recompiler (plan: closure-threaded -> Cranelift -> block-linking). No codegen yet — proves the integration seam and the correctness harness before any lowering exists. New crates/xenia-cpu/src/recompiler.rs: - run_block: executes a DecodedBlock; M0 falls back to the interpreter's execute() for every opcode, so it is bit-identical to step_block. Later stages dispatch lowered ops here and fall back only for uncompiled opcodes. - gates jit_enabled()/diff_enabled() (XENIA_JIT / XENIA_JIT_DIFF, cached). - In-process differential harness (diff_step): the interpreter is AUTHORITATIVE (drives real ctx+mem, so a JIT bug can never corrupt a run); the JIT runs SPECULATIVELY on a ctx clone against OverlayMemory (writes buffered in a byte HashMap, reads fall through to real pre-block memory), then registers are compared. Blocks touching MMIO or sync_sensitive (reservation/barrier) are skipped — a device callback can't be run twice and reservation state is shared cross-thread. report_diff_summary() prints checked/skipped/mismatch. Why in-process: the guest is only COARSELY deterministic — coord_idle_advance ticks vsync from wall-clock when idle, so two separate runs are not bit-exact and a cross-run signature compare would measure jitter, not JIT divergence. Supporting changes: PpcContext #[derive(Clone)] (harness clears the speculative clone's reservation_table Arc); is_mmio() on the MemoryAccess trait (default false) + GuestMemory impl via find_mmio; execute() made pub(crate); run_superblock routes the block body (DIFF->diff_step, JIT->run_block, else step_block). Validated: full boot+movie XENIA_JIT_DIFF=1 run = checked 218.4M blocks, skipped(mmio/sync) 511.6K (0.23%), MISMATCHES=0 CLEAN; movie plays (ADVreads=30, tid25 resumes); diff-mode ~2x slower (test-only path). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -48,6 +48,16 @@ pub trait MemoryAccess {
|
||||
}
|
||||
}
|
||||
|
||||
/// True if `addr` falls in an MMIO region (a load/store there invokes a
|
||||
/// device callback with side effects rather than touching backing RAM).
|
||||
///
|
||||
/// Default `false` (mock memories have no MMIO). `GuestMemory` overrides it.
|
||||
/// Used by the recompiler's differential harness to *skip* comparing blocks
|
||||
/// that touch MMIO — a device callback cannot be safely executed twice.
|
||||
fn is_mmio(&self, _addr: u32) -> bool {
|
||||
false
|
||||
}
|
||||
|
||||
/// Get a direct host pointer for the given guest address.
|
||||
/// Returns None if the address is invalid or in an MMIO region.
|
||||
fn translate(&self, addr: u32) -> Option<*const u8>;
|
||||
|
||||
@@ -504,6 +504,11 @@ impl GuestMemory {
|
||||
}
|
||||
|
||||
impl MemoryAccess for GuestMemory {
|
||||
#[inline]
|
||||
fn is_mmio(&self, addr: u32) -> bool {
|
||||
self.find_mmio(addr).is_some()
|
||||
}
|
||||
|
||||
// Tier-3 perf: `#[inline]` on the hot read/write paths lets LLVM
|
||||
// fold the MMIO + mapping checks into the interpreter's load/store
|
||||
// handlers, hoisting the "not-MMIO, mapped" branch out of the loop
|
||||
|
||||
Reference in New Issue
Block a user