[iterate-4C] JIT Phase B incr.2: region-granularity compilation (the real dispatch win)

Compile a straight-line/same-page REGION rooted at a block into one
CompiledBlock (XENIA_JIT_CHAIN): the set of same-page blocks reachable by
compile-time-known direct branches (BFS, capped by XENIA_JIT_REGION_MAX=32
+ 512-instr budget). Internal transitions between member blocks become plain
native jmps to dynasm labels — no per-block prologue/epilogue, no
jit_chain_next lookup, no stack churn — while the lockstep yield guards
(chain-enabled / budget / mmio) are evaluated inline at each block boundary
and a sync-sensitive block ends the region. Any successor not inlined
(indirect target, cross-page) falls to the increment-1 boundary chain:
jit_chain_next tail-jumps to the next region on a hit, or signals Rust to
build+compile it on a miss.

Byte-identical schedule to run_superblock: the SAME blocks run in the SAME
order with the SAME per-block budget/mmio/sync checks — only the dispatch
MECHANISM differs. Inlining is purely an optimization; the boundary chain is
the always-correct fallback, and inlined successors are same-page (⇒ mapped,
non-thunk) with the LR_HALT sentinel filtered, so next_pc_breaks_chain's
runtime arms are statically satisfied.

THROUGHPUT (best-of-5, n=200M --gpu-inline): interp 3.76s · plain JIT 3.98s ·
increment-1 chain 3.97s · REGION 3.49s. Region is ~1.08-1.14x FASTER than the
interpreter (first clear JIT win on the boot bench) and ~1.14x faster than
increment-1's throughput-neutral block chaining — confirming the thesis that
the real dispatch cost was the per-block prologue/epilogue increment 1 kept.
Region-size sweep: gains saturate by max=32 (max=64 identical).

GOLDEN n200m BYTE-IDENTICAL across 6 configs: interp / JIT / JIT+CHAIN(region)
/ JIT+CHAIN+REGCACHE / JIT+CHAIN+REGION_MAX=1 / JIT+CHAIN+budget=1 ==
interp+budget=1. 21 xenia-jit tests (added region_multiblock_matches_interp_
superblock + region_sync_block_ends_superblock).

Bug found+fixed during bring-up: a region chained PAST a sync-sensitive block
via the boundary chain instead of ending the superblock (compile_block gives
sync blocks the plain epilogue; the region must too) — this diverged the whole
200M schedule (draws 3201 vs 3165). Fix: sync block tail = unconditional
jmp l_region_ret; BFS never expands past a sync block. The sync regression
test asserts the region stops after A+B (cycle==4) instead of looping to budget.

New: xenia_cpu::block_cache::decode_block (pub standalone decoder for the
region builder) + DecodedBlock: Clone. Diagnostic knobs XENIA_JIT_REGION_MAX
(cap, default 32; =1 reproduces increment-1 chaining) and XENIA_JIT_NOREGION
(route chain compiles through compile_block — the A/B knob that isolated the
sync bug). Default XENIA_JIT/XENIA_JIT_CHAIN both still OFF; shipping path
unchanged.
This commit is contained in:
MechaCat02
2026-07-04 23:18:15 +02:00
parent d132fb6d8d
commit f562e3ec2b
4 changed files with 605 additions and 78 deletions

View File

@@ -3525,7 +3525,7 @@ fn run_superblock_jit_chained(
// Compile the first block; get its entry func.
let mut cur_func = {
let block = unsafe { &*first_block_ptr };
wc.jit_cache.as_mut().expect("jit active").ensure_compiled(block)
wc.jit_cache.as_mut().expect("jit active").ensure_compiled(block, mem)
};
// Diagnostics ptr for worker_epilogue (only SYSCALL/Trap read block.instrs;
// scheduling-affecting handling uses `result`, not this). Tracks the last
@@ -3551,7 +3551,7 @@ fn run_superblock_jit_chained(
last_pc_before = next_pc;
let block = wc.block_cache.lookup_or_build(next_pc, mem);
last_block_ptr = block as *const DecodedBlock;
cur_func = wc.jit_cache.as_mut().expect("jit active").ensure_compiled(block);
cur_func = wc.jit_cache.as_mut().expect("jit active").ensure_compiled(block, mem);
}
// Yield (budget/mmio/sync — handled inline) or any non-Continue
// result: end the superblock. `result` drives worker_epilogue.

View File

@@ -63,7 +63,7 @@ const GUEST_PAGE_MASK: u32 = !(GUEST_PAGE_SIZE - 1);
/// One cached basic block. Owned by [`BlockCache`]; a `&DecodedBlock`
/// is handed to the interpreter via [`BlockCache::lookup_or_build`] and
/// stays valid until the next `lookup_or_build` on the same slot.
#[derive(Debug)]
#[derive(Debug, Clone)]
pub struct DecodedBlock {
/// Guest PC at which this block starts. Used as the slot tag.
pub start_pc: u32,
@@ -185,6 +185,18 @@ impl BlockCache {
}
}
/// Decode a standalone `DecodedBlock` at `start_pc` against `mem`, computing
/// its `page_version` from memory (the same `(start_pc, page_version)` key the
/// `BlockCache` uses). Unlike [`BlockCache::lookup_or_build`] this does not
/// touch any cache — it is used by the JIT region compiler to decode the
/// straight-line/same-page successor blocks it stitches into one compiled
/// region. `build_block` stops at the 4 KiB page boundary, so the returned
/// block is fully contained in the page whose version this records.
pub fn decode_block(start_pc: u32, mem: &dyn MemoryAccess) -> DecodedBlock {
let page_version = mem.page_version(start_pc);
build_block(start_pc, mem, page_version)
}
/// Walk forward from `pc`, decoding instructions and collecting them
/// into a `DecodedBlock`. The walk stops on the first of:
/// - a [`PpcOpcode::terminates_block`] true (the terminator IS

View File

@@ -363,6 +363,95 @@ struct CompiledBlock {
// it holds are self-owned. It is never shared across threads.
unsafe impl Send for CompiledBlock {}
/// Emit one basic block's instruction stream into `ops`, using the
/// counter/pc-deferral scheme. Native ops accumulate `pending`; pc and the
/// cycle/timebase counters are materialized only at observability points
/// (fallbacks, native-branch edges, block end). On a taken branch or a
/// fall-through end the code reaches `l_cont`; a non-`Continue` fallback result
/// jumps to `l_exit` with the discriminant in `eax`. `r15`=ctx, `rbx`=env must
/// already be set by the prologue; `cache` may be disabled (regions use no
/// register cache). Shared by `compile_block` (single block → one `l_cont`) and
/// `compile_region` (each member block → its own tail `l_cont`).
fn emit_block_body(
ops: &mut dynasmrt::x64::Assembler,
off: &emit::Offsets,
mem_helpers: &MemHelpers,
cache: &emit::RegCache,
instrs: &[DecodedInstr],
l_exit: dynasmrt::DynamicLabel,
l_cont: dynasmrt::DynamicLabel,
) {
let helper = jit_interpret_one as usize as i64;
let mut state = emit::EmitState::new();
let mut tail_pc: Option<u32> = None;
for instr in instrs.iter() {
match emit::try_emit_native(ops, off, mem_helpers, &mut state, cache, instr) {
// Native non-branch: computation only; pc/counters deferred.
emit::Emit::Native => {
tail_pc = Some(instr.addr.wrapping_add(4));
continue;
}
// Native branch: it set pc absolutely and flushed the counters.
// Append the pc-discontinuity check (taken -> exit via l_cont;
// fall-through -> continue), matching step_block.
emit::Emit::Branch => {
let expected_next = instr.addr.wrapping_add(4) as i32;
dynasm!(ops
; .arch x64
; cmp DWORD [r15 + off.pc], expected_next
; jne =>l_cont
);
tail_pc = None;
continue;
}
// Un-ported opcode: emit the interpreter fallback below.
emit::Emit::Fallback => {}
}
// Interpreter fallback for un-ported opcodes. Make counters + pc current
// first (the callee may read timebase via mftb and reads/writes pc), then
// run it and account its own retirement (interpreter order: after execute).
state.flush_counters(ops, off);
let addr = instr.addr as i32;
let instr_ptr = instr as *const DecodedInstr as usize as i64;
let expected_next = instr.addr.wrapping_add(4) as i32;
// The interpreter reads/writes ctx.gpr directly, so flush the cached
// regs to memory before the call and reload them after.
emit::emit_cache_flush(ops, cache, off);
dynasm!(ops
; .arch x64
; mov DWORD [r15 + off.pc], addr
// fallback: eax = jit_interpret_one(env, &instr); env.last_result set
; mov rdi, rbx
; mov rsi, QWORD instr_ptr
; mov rax, QWORD helper
; call rax
);
emit::emit_cache_load(ops, cache, off); // reload possibly-changed gprs
dynasm!(ops
; .arch x64
// determinism postlude: this instruction retired.
; inc QWORD [r15 + off.cycle]
; inc QWORD [r15 + off.timebase]
// non-Continue result -> exit returning the discriminant in eax
; test eax, eax
; jnz =>l_exit
// taken-branch (pc discontinuity) -> stop the block, return Continue
; cmp DWORD [r15 + off.pc], expected_next
; jne =>l_cont
);
tail_pc = None;
}
// Block end (fall-through): flush the deferred counters and materialize the
// final pc if the tail was straight-line native. Emitted BEFORE the caller's
// l_cont so a taken branch/fallback (which jumps to l_cont) skips it — its
// pc/counters are already current.
state.flush_counters(ops, off);
if let Some(p) = tail_pc {
dynasm!(ops ; .arch x64 ; mov DWORD [r15 + off.pc], p as i32);
}
}
/// Compile `block` into a `CompiledBlock`. Phase 0: every instruction is a
/// `call jit_interpret_one` + the mandatory counter/exit postlude.
fn compile_block(block: &DecodedBlock) -> CompiledBlock {
@@ -373,7 +462,6 @@ fn compile_block(block: &DecodedBlock) -> CompiledBlock {
// Field offsets + helper addresses resolved at compile time.
let off = emit::Offsets::resolve();
let mem_helpers = MemHelpers::resolve();
let helper = jit_interpret_one as usize as i64;
// Per-block register cache (a few hot GPRs pinned in r12/r13/r14). Empty
// unless enabled, in which case the accessors read/write the host regs and
@@ -440,77 +528,9 @@ fn compile_block(block: &DecodedBlock) -> CompiledBlock {
// Counter/pc deferral: native ops just accumulate `state.pending`; pc and the
// counters are materialized only at observability points (fallbacks, branch
// edges, block end). `tail_pc` = Some(next_pc) when the last emitted op was
// straight-line native (its pc write is deferred to the fall-through end);
// None when a branch/fallback already set pc.
let mut state = emit::EmitState::new();
let mut tail_pc: Option<u32> = None;
for instr in instrs.iter() {
match emit::try_emit_native(&mut ops, &off, &mem_helpers, &mut state, &cache, instr) {
// Native non-branch: computation only; pc/counters deferred.
emit::Emit::Native => {
tail_pc = Some(instr.addr.wrapping_add(4));
continue;
}
// Native branch: it set pc absolutely and flushed the counters.
// Append the pc-discontinuity check (taken -> exit via l_cont;
// fall-through -> continue), matching step_block.
emit::Emit::Branch => {
let expected_next = instr.addr.wrapping_add(4) as i32;
dynasm!(ops
; .arch x64
; cmp DWORD [r15 + off.pc], expected_next
; jne =>l_cont
);
tail_pc = None;
continue;
}
// Un-ported opcode: emit the interpreter fallback below.
emit::Emit::Fallback => {}
}
// Interpreter fallback for un-ported opcodes. Make counters + pc current
// first (the callee may read timebase via mftb and reads/writes pc), then
// run it and account its own retirement (interpreter order: after execute).
state.flush_counters(&mut ops, &off);
let addr = instr.addr as i32;
let instr_ptr = instr as *const DecodedInstr as usize as i64;
let expected_next = instr.addr.wrapping_add(4) as i32;
// The interpreter reads/writes ctx.gpr directly, so flush the cached
// regs to memory before the call and reload them after.
emit::emit_cache_flush(&mut ops, &cache, &off);
dynasm!(ops
; .arch x64
; mov DWORD [r15 + off.pc], addr
// fallback: eax = jit_interpret_one(env, &instr); env.last_result set
; mov rdi, rbx
; mov rsi, QWORD instr_ptr
; mov rax, QWORD helper
; call rax
);
emit::emit_cache_load(&mut ops, &cache, &off); // reload possibly-changed gprs
dynasm!(ops
; .arch x64
// determinism postlude: this instruction retired.
; inc QWORD [r15 + off.cycle]
; inc QWORD [r15 + off.timebase]
// non-Continue result -> exit returning the discriminant in eax
; test eax, eax
; jnz =>l_exit
// taken-branch (pc discontinuity) -> stop the block, return Continue
; cmp DWORD [r15 + off.pc], expected_next
; jne =>l_cont
);
tail_pc = None;
}
// Block end (fall-through): flush the deferred counters and materialize the
// final pc if the tail was straight-line native. Emitted BEFORE l_cont so a
// taken branch/fallback (which jumps to l_cont) skips it — its pc/counters
// are already current.
state.flush_counters(&mut ops, &off);
if let Some(p) = tail_pc {
dynasm!(ops ; .arch x64 ; mov DWORD [r15 + off.pc], p as i32);
}
// edges, block end). The single block routes both its taken-branch and
// fall-through ends to the one shared `l_cont`.
emit_block_body(&mut ops, &off, &mem_helpers, &cache, &instrs, l_exit, l_cont);
if chaining {
// --- Native-chaining epilogue (XENIA_JIT_CHAIN, non-sync blocks) ---
@@ -616,6 +636,344 @@ fn compile_block(block: &DecodedBlock) -> CompiledBlock {
}
}
/// Guest page granule — mirrors `block_cache`'s page-boundary stop. Region
/// membership is confined to one page so a single `page_version` key covers the
/// whole region (SMC coherence) and every inlined successor is trivially mapped
/// and outside the import-thunk band (game code pages are disjoint from it).
const JIT_GUEST_PAGE_MASK: u32 = !4095u32;
/// Compile-time-known, same-page, direct successor PCs of `block` — the only
/// successors a region may inline as an internal native jump. Returns at most
/// two. Everything else (indirect `bclr`/`bcctr` targets, cross-page branches,
/// syscalls/traps/`Invalid`) is omitted; those exit the region to the Rust
/// chain driver, which dispatches them correctly (inlining is purely an
/// optimization — the boundary chain is the always-correct fallback).
///
/// Every returned PC is on `block`'s own page, so it is necessarily mapped and
/// (game pages being disjoint from the thunk band) never a thunk — the two
/// runtime break arms of `next_pc_breaks_chain` are statically satisfied. The
/// only remaining break condition, the `LR_HALT` sentinel, is filtered here.
fn block_static_successors(block: &DecodedBlock) -> Vec<u32> {
use xenia_cpu::opcode::PpcOpcode;
const LR_HALT: u32 = xenia_cpu::context::LR_HALT_SENTINEL as u32;
let last = match block.instrs.last() {
Some(i) => i,
None => return Vec::new(),
};
let end_pc = last.addr.wrapping_add(4); // fall-through PC
// Candidate successor PCs, matching emit.rs's exact target arithmetic.
let mut cands: Vec<u32> = Vec::new();
match last.opcode {
// Unconditional direct branch: single target, no fall-through.
PpcOpcode::bx => {
let t = if last.aa() {
last.li() as u32
} else {
last.addr.wrapping_add(last.li() as u32)
};
cands.push(t);
}
// Conditional direct branch: taken target + fall-through.
PpcOpcode::bcx => {
let t = if last.aa() {
last.bd() as u32
} else {
last.addr.wrapping_add(last.bd() as u32)
};
cands.push(t);
cands.push(end_pc);
}
// Indirect branch: taken target unknown; only a (possible) conditional
// fall-through is a static successor.
PpcOpcode::bclrx | PpcOpcode::bcctrx => {
cands.push(end_pc);
}
// Non-Continue / trap terminators: never chain past them.
PpcOpcode::sc
| PpcOpcode::td
| PpcOpcode::tdi
| PpcOpcode::tw
| PpcOpcode::twi
| PpcOpcode::Invalid => {}
// Non-branch terminator (block hit the instruction cap or a page edge):
// straight-line fall-through.
_ => {
cands.push(end_pc);
}
}
let page_base = block.start_pc & JIT_GUEST_PAGE_MASK;
let mut out: Vec<u32> = Vec::with_capacity(2);
for pc in cands {
if pc != LR_HALT && (pc & JIT_GUEST_PAGE_MASK) == page_base && !out.contains(&pc) {
out.push(pc);
}
}
out
}
/// Diagnostic: when `XENIA_JIT_NOREGION` is set, chain compiles route through
/// the increment-1 `compile_block` tail-chaining epilogue instead of
/// `compile_region`. Isolates a region bug from a driver/refactor bug.
fn region_disabled() -> bool {
use std::sync::OnceLock;
static D: OnceLock<bool> = OnceLock::new();
*D.get_or_init(|| std::env::var("XENIA_JIT_NOREGION").is_ok())
}
/// Region block cap (`XENIA_JIT_REGION_MAX`, default 32, min 1). Setting it to
/// 1 makes every region a single block — reproducing the increment-1 native
/// tail-chaining schedule exactly (a useful differential/bisection knob).
fn region_max_blocks() -> usize {
use std::sync::OnceLock;
static N: OnceLock<usize> = OnceLock::new();
*N.get_or_init(|| {
std::env::var("XENIA_JIT_REGION_MAX")
.ok()
.and_then(|v| v.parse::<usize>().ok())
.map(|v| v.max(1))
.unwrap_or(32)
})
}
/// Compile a straight-line/same-page **region** rooted at `first` into one
/// `CompiledBlock` (increment 2 — the real dispatch win). The region is the set
/// of same-page blocks reachable from `first` by compile-time-known direct
/// branches (BFS, capped by [`region_max_blocks`] and a 512-instr budget).
/// Internal transitions between member blocks become plain native `jmp`s to
/// dynasm labels — no per-block prologue/epilogue, no `jit_chain_next` lookup,
/// no stack churn — while the lockstep yield guards (chain-enabled / budget /
/// mmio) are still evaluated inline at every block boundary, and a sync block
/// ends the region. Any successor not inlined (indirect target, cross-page)
/// falls to the increment-1 boundary chain: `jit_chain_next` tail-jumps to the
/// next region on a hit, or signals Rust to build+compile it on a miss.
///
/// The schedule is byte-identical to `run_superblock`: the SAME blocks run in
/// the SAME order with the SAME per-block budget/mmio/sync checks; only the
/// dispatch MECHANISM differs. Used only when `chain_active()` (the chained
/// driver path); the plain-JIT path keeps `compile_block`.
fn compile_region(first: &DecodedBlock, mem: &dyn MemoryAccess) -> CompiledBlock {
const REGION_MAX_INSTRS: usize = 512;
let max_blocks = region_max_blocks();
// --- 1. Discover the region (BFS over same-page direct successors). ---
let entry_page = first.start_pc & JIT_GUEST_PAGE_MASK;
let mut blocks: Vec<DecodedBlock> = Vec::new();
let mut index_of: std::collections::HashMap<u32, usize> = std::collections::HashMap::new();
let mut total_instrs = first.instrs.len();
blocks.push(first.clone());
index_of.insert(first.start_pc, 0);
let mut i = 0;
while i < blocks.len() {
// A sync-sensitive block ends the superblock — never expand past it, so
// its successors are not admitted to the region (they'd be dead code).
if blocks[i].sync_sensitive {
i += 1;
continue;
}
// Snapshot successors first (no borrow held across the push below).
let succs = block_static_successors(&blocks[i]);
for s in succs {
if index_of.contains_key(&s) || blocks.len() >= max_blocks {
continue;
}
if (s & JIT_GUEST_PAGE_MASK) != entry_page {
continue; // different page → cannot share the region's page_version
}
let b = xenia_cpu::block_cache::decode_block(s, mem);
// Same page ⇒ same page_version as the entry (single-threaded); guard
// defensively so a mismatch just declines to inline rather than
// producing a region keyed on a version it doesn't fully cover.
if b.page_version != first.page_version {
continue;
}
if total_instrs + b.instrs.len() > REGION_MAX_INSTRS {
continue;
}
total_instrs += b.instrs.len();
index_of.insert(s, blocks.len());
blocks.push(b);
}
i += 1;
}
// --- 2. Own every member block's instrs in one boxed slice. The emitted
// fallback `call`s bake pointers into THIS copy, so it must outlive `func`. ---
let mut all_instrs: Vec<DecodedInstr> = Vec::with_capacity(total_instrs);
let mut ranges: Vec<(usize, usize)> = Vec::with_capacity(blocks.len());
for b in &blocks {
let start = all_instrs.len();
all_instrs.extend_from_slice(&b.instrs);
ranges.push((start, b.instrs.len()));
}
let instrs: Box<[DecodedInstr]> = all_instrs.into_boxed_slice();
// --- 3. Emit: one prologue, member blocks with internal dispatch, one
// epilogue. Regions never use the register cache (default-off + measured
// neutral); GPRs live in ctx.gpr throughout, so no inter-block flush. ---
let off = emit::Offsets::resolve();
let mem_helpers = MemHelpers::resolve();
let cache = emit::RegCache::disabled();
let chain_helper = jit_chain_next as usize as i64;
let mut ops = dynasmrt::x64::Assembler::new().expect("dynasm assembler");
let entry = ops.offset();
let l_exit = ops.new_dynamic_label();
let l_region_ret = ops.new_dynamic_label();
let l_region_miss = ops.new_dynamic_label();
let block_entry: Vec<_> = (0..blocks.len()).map(|_| ops.new_dynamic_label()).collect();
let block_cont: Vec<_> = (0..blocks.len()).map(|_| ops.new_dynamic_label()).collect();
// Prologue: pin env=rbx, ctx=r15; keep rsp 16-aligned (entry rsp%16==8 →
// push/push/sub8 → %16==0) for the helper `call`s.
dynasm!(ops
; .arch x64
; push rbx
; push r15
; sub rsp, 8
; mov rbx, rdi
; mov r15, [rbx + off.env_ctx]
);
// MMIO watermark snapshot for the inline guard, guarded by chain_enabled: a
// region is ALSO run via `run_jit_block` (budget==1 / Phase A / probes),
// where `chain_mmio_ptr` is null — skip the snapshot there or it derefs null.
{
let l_skip = ops.new_dynamic_label();
dynasm!(ops
; .arch x64
; mov cl, BYTE [rbx + off.env_chain_enabled]
; test cl, cl
; jz =>l_skip
; mov rax, QWORD [rbx + off.env_chain_mmio_ptr]
; mov rax, QWORD [rax]
; mov QWORD [rbx + off.env_chain_mmio_before], rax
; =>l_skip
);
}
for (bi, block) in blocks.iter().enumerate() {
let (start, len) = ranges[bi];
let body = &instrs[start..start + len];
// Block entry (internal jumps land here) then its instruction stream.
dynasm!(ops ; .arch x64 ; =>block_entry[bi]);
emit_block_body(&mut ops, &off, &mem_helpers, &cache, body, l_exit, block_cont[bi]);
// --- Block tail: reached (via block_cont[bi]) on a Continue boundary. ---
dynasm!(ops ; .arch x64 ; =>block_cont[bi]);
// Sync-sensitive block: `run_superblock` ends the superblock AFTER it
// (never chains out of it — compile_block achieves this by giving sync
// blocks the plain, non-chaining epilogue). Return Continue → the driver
// yields. This unconditional exit also covers the chain_enabled==0 path.
if block.sync_sensitive {
dynasm!(ops ; .arch x64 ; jmp =>l_region_ret);
continue;
}
// Chaining disabled this run (budget==1 / probe path via run_jit_block):
// return after this block, exactly like a single compiled block.
dynasm!(ops
; .arch x64
; mov cl, BYTE [rbx + off.env_chain_enabled]
; test cl, cl
; jz =>l_region_ret
);
// Budget guard: cycle_count >= deadline → yield (== interp total>=budget).
dynasm!(ops
; .arch x64
; mov rax, QWORD [r15 + off.cycle]
; cmp rax, QWORD [rbx + off.env_chain_deadline]
; jae =>l_region_ret
);
// MMIO guard: watermark moved since region entry → yield.
dynasm!(ops
; .arch x64
; mov rax, QWORD [rbx + off.env_chain_mmio_ptr]
; mov rax, QWORD [rax]
; cmp rax, QWORD [rbx + off.env_chain_mmio_before]
; jne =>l_region_ret
);
// Fast internal dispatch to inlined successors. Only successors actually
// admitted to the region are inlined; the rest fall to the boundary chain.
let succs: Vec<u32> = block_static_successors(block)
.into_iter()
.filter(|s| index_of.contains_key(s))
.collect();
if !succs.is_empty() {
dynasm!(ops ; .arch x64 ; mov edx, DWORD [r15 + off.pc]);
for s in &succs {
let tgt = index_of[s];
dynasm!(ops
; .arch x64
; cmp edx, *s as i32
; je =>block_entry[tgt]
);
}
}
// Not an inlined successor → boundary-chain to the next region (canary-
// style indirection): jit_chain_next(env, pc) → host entry or null.
dynasm!(ops
; .arch x64
; mov rdi, rbx
; mov esi, DWORD [r15 + off.pc]
; mov rax, QWORD chain_helper
; call rax
; test rax, rax
; jz =>l_region_miss
// HIT: env→rdi, unwind THIS frame to entry rsp, tail-jump.
; mov rdi, rbx
; add rsp, 8
; pop r15
; pop rbx
; jmp rax
);
}
// Miss: next_pc is chainable-shaped but not JIT-compiled — signal Rust to
// build+compile it and re-enter (increment-1 protocol). Falls into the
// Continue-return epilogue.
dynasm!(ops
; .arch x64
; =>l_region_miss
; mov BYTE [rbx + off.env_chain_stop], 1
; =>l_region_ret
; xor eax, eax
; add rsp, 8
; pop r15
; pop rbx
; ret
);
// Non-Continue exit (fallback `jnz =>l_exit`): eax already holds the
// StepResult discriminant.
dynasm!(ops
; .arch x64
; =>l_exit
; add rsp, 8
; pop r15
; pop rbx
; ret
);
let buf = ops.finalize().expect("dynasm finalize");
// SAFETY: `entry` is a valid offset into `buf`; the emitted code matches the
// `JitBlockFn` ABI (System V, first arg rdi, return eax).
let func: JitBlockFn = unsafe { std::mem::transmute::<*const u8, JitBlockFn>(buf.ptr(entry)) };
CompiledBlock {
start_pc: first.start_pc,
page_version: first.page_version,
_instrs: instrs,
_buf: buf,
func,
// The chained driver never reads this (the region applies all break
// conditions internally); the non-chained/budget==1 path runs only the
// FIRST block via run_jit_block, so the first block's flag is the correct
// one for run_superblock_jit's post-run sync check.
sync_sensitive: blocks[0].sync_sensitive,
}
}
/// Run a compiled block against `ctx`/`mem`, returning the same `StepResult`
/// the interpreter's `step_block` would. The block bumps `cycle_count`/
/// `timebase` and updates `ctx.pc` in place, exactly like the interpreter.
@@ -757,7 +1115,14 @@ impl JitCache {
self.hits += 1;
} else {
self.compiles += 1;
self.slots[idx] = Some(compile_block(block));
// With chaining on, compile a multi-block REGION rooted here
// (internal transitions become free native jumps); otherwise a
// single block. Both are keyed identically on (start_pc, page_version).
self.slots[idx] = Some(if chain_active() && !region_disabled() {
compile_region(block, mem)
} else {
compile_block(block)
});
}
let cb = self.slots[idx].as_ref().expect("just populated");
run_jit_block(cb, ctx, mem)
@@ -803,7 +1168,7 @@ impl JitCache {
/// keyed on `(start_pc, page_version)` like `run_or_compile`. Used by the
/// chaining integration to (re-)enter a native chain at a freshly-built
/// block.
pub fn ensure_compiled(&mut self, block: &DecodedBlock) -> JitBlockFn {
pub fn ensure_compiled(&mut self, block: &DecodedBlock, mem: &dyn MemoryAccess) -> JitBlockFn {
let idx = ((block.start_pc >> 2) & JIT_CACHE_MASK) as usize;
let fresh = matches!(
&self.slots[idx],
@@ -813,7 +1178,13 @@ impl JitCache {
self.hits += 1;
} else {
self.compiles += 1;
self.slots[idx] = Some(compile_block(block));
// Chaining is always on when this is called (the chained driver), so
// compile a region; keep the `compile_block` fallback for symmetry.
self.slots[idx] = Some(if chain_active() && !region_disabled() {
compile_region(block, mem)
} else {
compile_block(block)
});
}
self.slots[idx].as_ref().expect("just populated").func
}

View File

@@ -991,3 +991,147 @@ fn run_fresh_hit_miss() {
let mut c3 = ctx_from_gpr([0u64; 32], other);
assert!(cache.run_fresh(other, &mut c3, &mem).is_none(), "uncompiled pc must miss");
}
/// Region compilation (increment 2): a multi-block, same-page loop stitched into
/// ONE compiled region with internal native jumps must yield at EXACTLY the
/// interpreter superblock's per-block budget boundary, with byte-identical GPRs,
/// CTR, pc, and cycle/timebase. This is the load-bearing determinism property:
/// the region only changes the dispatch MECHANISM, never the schedule.
#[test]
fn region_multiblock_matches_interp_superblock() {
use super::{compile_region, run_jit_block, run_jit_chain, ChainStop};
use xenia_cpu::block_cache::decode_block;
use xenia_cpu::interpreter::{step_block, StepResult};
let base = 0x8200_1000u32;
// A@1000: addi r3,r3,5 ; b +4 (-> B@1008)
// B@1008: addi r4,r4,1 ; bdnz -> A@1000 (BO=16: dec CTR, branch if CTR!=0)
// C@1010: sc (region member via B's fall-through; never reached).
let prog: [(u32, u32); 5] = [
(base, enc_d(14, 3, 3, 5)), // addi r3, r3, 5
(base + 0x4, enc_bx(4, 0, 0)), // b +4 -> 0x1008
(base + 0x8, enc_d(14, 4, 4, 1)), // addi r4, r4, 1
(base + 0xC, enc_bcx(16, 0, -12, 0, 0)), // bdnz -> 0x1000
(base + 0x10, (17 << 26) | 2), // sc
];
let mem = VecMem::seeded();
for (addr, raw) in prog {
mem.write_u32(addr, raw);
}
let gpr = {
let mut s = 0x5151u64;
fuzz_gpr(&mut s)
};
const CTR0: u64 = 10_000;
// 4 instrs per A->B iteration; 400 = exactly 100 iterations, well under CTR0
// (so the loop is always taken — never falls through to `sc`).
const BUDGET: u64 = 400;
// Reference: the real interpreter superblock loop (same yield structure as
// run_superblock; no mmio/thunk in this mock).
let mut a = ctx_from_gpr(gpr, base);
a.ctr = CTR0;
let r_ref = loop {
let blk = decode_block(a.pc, &mem);
let r = step_block(&mut a, &mem, &blk);
if a.cycle_count >= BUDGET || !matches!(r, StepResult::Continue) || blk.sync_sensitive {
break r;
}
};
// JIT region driven through the native chain.
let mut b = ctx_from_gpr(gpr, base);
b.ctr = CTR0;
let first = decode_block(base, &mem);
let cb = compile_region(&first, &mem);
let mut cache = crate::JitCache::new();
let cp = cache.as_ptr();
let mmio: u64 = 0;
let (r_jit, stop) = run_jit_chain(cb.func, &mut b, &mem, cp, BUDGET, &mmio as *const u64, true);
assert!(matches!(stop, ChainStop::Yield), "region must yield on budget, not miss");
assert_eq!(a.gpr, b.gpr, "region gpr mismatch");
assert_eq!(a.ctr, b.ctr, "region ctr mismatch");
assert_eq!(a.pc, b.pc, "region pc mismatch");
assert_eq!(a.cycle_count, b.cycle_count, "region cycle mismatch");
assert_eq!(a.timebase, b.timebase, "region timebase mismatch");
assert_eq!(a.cycle_count, BUDGET, "region must run exactly to the budget");
assert_eq!(r_ref, r_jit, "region StepResult mismatch");
// Sanity: the loop actually ran (100 iterations of +5 / +1).
assert_eq!(b.gpr[3], gpr[3].wrapping_add(5 * 100), "r3 accumulation wrong");
assert_eq!(b.gpr[4], gpr[4].wrapping_add(100), "r4 accumulation wrong");
// chain_enabled=false (the budget==1 / probe / Phase-A path via run_jit_block):
// the SAME region must run ONLY its first block and stop — identical to
// step_block over block A alone.
let mut c = ctx_from_gpr(gpr, base);
c.ctr = CTR0;
let _ = run_jit_block(&cb, &mut c, &mem);
let mut d = ctx_from_gpr(gpr, base);
d.ctr = CTR0;
let blk_a = decode_block(base, &mem);
let _ = step_block(&mut d, &mem, &blk_a);
assert_eq!(c.gpr, d.gpr, "single-block(region) gpr mismatch");
assert_eq!(c.pc, d.pc, "single-block(region) pc mismatch");
assert_eq!(c.cycle_count, d.cycle_count, "single-block(region) cycle mismatch");
assert_eq!(c.pc, base + 0x8, "block A must exit to B's start (b +4 taken)");
}
/// A sync-sensitive block inside a region must END the superblock after it runs
/// (never chain out — exactly like `run_superblock` and `compile_block`, which
/// gives sync blocks the plain epilogue). Regression guard for the increment-2
/// bug where a region chained PAST a sync point (fewer scheduler yields → whole
/// 200M schedule diverged). Without the fix, the region loops A->B->A... to the
/// budget instead of stopping at B.
#[test]
fn region_sync_block_ends_superblock() {
use super::{compile_region, run_jit_chain, ChainStop};
use xenia_cpu::block_cache::decode_block;
use xenia_cpu::interpreter::{step_block, StepResult};
let base = 0x8200_1000u32;
// A@1000: addi r3,r3,5 ; b +4 (-> B@1008)
// B@1008: sync ; b -0xC (-> A@1000) [B is sync-sensitive]
let prog: [(u32, u32); 4] = [
(base, enc_d(14, 3, 3, 5)), // addi r3, r3, 5
(base + 0x4, enc_bx(4, 0, 0)), // b +4 -> 0x1008
(base + 0x8, 0x7C00_04AC), // sync
(base + 0xC, enc_bx(-0xC, 0, 0)), // b -0xC -> 0x1000
];
let mem = VecMem::seeded();
for (addr, raw) in prog {
mem.write_u32(addr, raw);
}
let gpr = {
let mut s = 0x7373u64;
fuzz_gpr(&mut s)
};
// Large budget so the run stops on SYNC, not budget (the whole point).
const BUDGET: u64 = 100;
// Reference interpreter superblock: breaks AFTER a sync-sensitive block.
let mut a = ctx_from_gpr(gpr, base);
let r_ref = loop {
let blk = decode_block(a.pc, &mem);
let r = step_block(&mut a, &mem, &blk);
if a.cycle_count >= BUDGET || !matches!(r, StepResult::Continue) || blk.sync_sensitive {
break r;
}
};
let mut b = ctx_from_gpr(gpr, base);
let first = decode_block(base, &mem);
let cb = compile_region(&first, &mem);
let mut cache = crate::JitCache::new();
let cp = cache.as_ptr();
let mmio: u64 = 0;
let (r_jit, _stop) = run_jit_chain(cb.func, &mut b, &mem, cp, BUDGET, &mmio as *const u64, true);
assert_eq!(a.cycle_count, b.cycle_count, "sync-region cycle mismatch");
assert_eq!(a.cycle_count, 4, "must stop after A+B (4 instrs), not loop to budget");
assert_eq!(a.gpr, b.gpr, "sync-region gpr mismatch (must run block A exactly once)");
assert_eq!(a.pc, b.pc, "sync-region pc mismatch");
assert_eq!(r_ref, r_jit, "sync-region StepResult mismatch");
let _ = ChainStop::Yield;
}