run_superblock re-resolved the running thread's PpcContext on every chained block via two bounds-checked slot lookups (ctx_mut_ref for the step + ctx(hw_id) for next_pc) — ~20M double-lookups per 100M-instr window. The running thread is FIXED for the whole chain (an import thunk or any sync-sensitive/MMIO op breaks the chain before any kernel mutation could restructure the runqueue; step_block is pure guest interpretation; the probe closure is read-only), so its context heap slot is stable. Resolve the raw *mut PpcContext once before the loop and reuse it — same raw-pointer discipline as block_ptr. ~3% faster clean wall (1.91s -> 1.85s, n=100M, best of 5). Golden sylpheed_n200m BYTE-IDENTICAL under interp, XENIA_JIT=1, and XENIA_JIT=1+XENIA_JIT_REGCACHE=1. Found via exclusive-attribution profiling: the run is overhead-bound with the non-step cost spread thin across the per-block chaining loop (no single dominant lever); step_block itself is ~46% and only the JIT attacks it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>