[iterate-4A] perf: superblock budget 128->192 (movie-validated ~7% wall) + profiler attribution

Speed frontier, increment 1 (measure-first, then bank the cheap win).

Profiler (crates/xenia-gpu/src/prof.rs): extend XENIA_PROFILE with hierarchical
attribution of the single-thread lockstep loop — ROUND (per-round tax) /
PROLOGUE (worker_prologue) / RUNSB (run_superblock) top level, plus the
STEP/EPILOGUE/KERNEL/BUILD subsets. Gated zero-cost via is_on() (off-check emits
0 lines). Movie-time split: interp step_block 42%, run_superblock chain-loop
body ~30%, worker_prologue+epilogue ~22%, kernel HLE 6%, per-round tax 0.6%,
texture 0.4%. => overhead-bound; the ceiling is a JIT (attacks the ~72% interp
+ dispatch). The per-round tax is a non-lever (0.6%).

Budget (crates/xenia-app/src/main.rs SUPERBLOCK_INSTR_BUDGET 128->192): the
per-slot-visit tax (prologue+epilogue ~22%) scales with slot-visit count, and
chains are NOT break-limited here — 128->192 cuts visits 27.4M->18.9M (-31%),
128->256 -44%. Banked 192 (not 256): the movie decode pipeline is more
timing-sensitive than boot. At 256 the decode worker tid25 (0x82506588)
intermittently fails to resume (1-of-2 runs) with a weakened feeder loop
(source-read 27->5-11) — a scheduling race the coarser interleaving exposes.
192 is deterministic and safe: 3/3 runs byte-identical (source-read=21,
tid25-resume=1, ADVreads=30) at both env-override and compiled-default, ~-7%
movie wall, well below the 384 boot cliff. XENIA_SUPERBLOCK_BUDGET still
overrides for A/B.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-07-06 19:45:23 +02:00
parent cc9ebbb5e7
commit e07f93ed0a
2 changed files with 87 additions and 51 deletions

View File

@@ -3001,10 +3001,20 @@ fn worker_epilogue(
/// swaps / packets track the one-block baseline), then a sharp cliff at
/// ~384 collapses the present loop (a producer/consumer boot handoff
/// starves when one slot runs too long without returning to the round).
/// 128 sits 3× below that cliff with ~1.65× boot-to-splash speedup — a
/// deliberately conservative pick (correctness over the last few %). The
/// `XENIA_SUPERBLOCK_BUDGET` env var overrides it for further tuning.
const SUPERBLOCK_INSTR_BUDGET: u64 = 128;
///
/// Re-tuned against the INTRO-VIDEO workload (2026-07-06, XENIA_PROFILE
/// attribution): the per-slot-visit tax (worker_prologue + worker_epilogue)
/// was ~22% of movie wall, and raising the budget cuts slot-visits ~in
/// proportion (chains are NOT break-limited here — 128→192 dropped visits
/// 27.4M→18.9M, 31%; 128→256 44%). But the movie's decode pipeline is more
/// timing-sensitive than boot: at 256 the decode worker tid25 (0x82506588)
/// INTERMITTENTLY fails to resume (1-of-2 runs) and the feeder loop weakens
/// (source-read 27→5-11) — a scheduling race the coarser interleaving exposes.
/// 192 is the sweet spot: 3/3 runs byte-identical (source-read=21,
/// tid25-resume=1, ADVreads=30) with ~7% movie wall, and it stays well below
/// the 384 boot cliff. Do NOT raise past 192 without re-validating tid25
/// resume across several runs. `XENIA_SUPERBLOCK_BUDGET` overrides for A/B.
const SUPERBLOCK_INSTR_BUDGET: u64 = 192;
/// Effective superblock budget. Defaults to [`SUPERBLOCK_INSTR_BUDGET`];
/// `XENIA_SUPERBLOCK_BUDGET` overrides it (A/B tuning without a rebuild).
@@ -3194,6 +3204,12 @@ fn run_superblock(
}
};
let _ep_pt = xenia_gpu::prof::is_on().then(|| {
xenia_gpu::prof::ScopeTimer::new(
&xenia_gpu::prof::EPILOGUE_NS,
&xenia_gpu::prof::EPILOGUE_CALLS,
)
});
worker_epilogue(
wc,
kernel,
@@ -3302,6 +3318,8 @@ fn run_execution(
// then drain any pending auto-signals whose deadline has passed.
// Both calls are no-ops when `XENIA_SILPH_UI_AUTOSIGNAL_DELAY`
// is unset (the pending queue stays empty).
let _round_pt = xenia_gpu::prof::is_on()
.then(|| xenia_gpu::prof::ScopeTimer::new(&xenia_gpu::prof::ROUND_NS, &xenia_gpu::prof::ROUND_CALLS));
kernel.set_now_cycle_hint(stats.instruction_count);
// Drive the coherent monotonic "now" the kernel deadline-arithmetic
// reads (`KernelState::now_basis_at` -> `Scheduler::global_clock`)
@@ -3336,6 +3354,7 @@ fn run_execution(
// a reusable stack array instead of allocating a fresh Vec per round.
kernel.scheduler.begin_round();
let order_n = kernel.scheduler.round_schedule_into(&mut order_buf);
drop(_round_pt); // end per-round tax measurement at the schedule boundary
let order = &order_buf[..order_n];
if order.is_empty() {
@@ -3360,7 +3379,13 @@ fn run_execution(
for &hw_id in order {
let wc = &mut workers[hw_id as usize];
match worker_prologue(
let _pro_pt = xenia_gpu::prof::is_on().then(|| {
xenia_gpu::prof::ScopeTimer::new(
&xenia_gpu::prof::PROLOGUE_NS,
&xenia_gpu::prof::PROLOGUE_CALLS,
)
});
let prologue_outcome = worker_prologue(
wc,
kernel,
mem,
@@ -3368,7 +3393,9 @@ fn run_execution(
&mut db_writer,
thunk_map,
&mut stats,
) {
);
drop(_pro_pt);
match prologue_outcome {
PrologueOutcome::Continue => continue,
PrologueOutcome::BreakOuter => break 'outer,
PrologueOutcome::StepBlock {
@@ -3385,7 +3412,13 @@ fn run_execution(
// the per-round (timebase / coord / round_schedule)
// and per-slot (prologue) tax over hundreds of
// instructions instead of ~6. See `run_superblock`.
match run_superblock(
let _sb_pt = xenia_gpu::prof::is_on().then(|| {
xenia_gpu::prof::ScopeTimer::new(
&xenia_gpu::prof::RUNSB_NS,
&xenia_gpu::prof::RUNSB_CALLS,
)
});
let sb_outcome = run_superblock(
wc,
kernel,
mem,
@@ -3396,7 +3429,9 @@ fn run_execution(
thread_ref,
block_ptr,
pc_before,
) {
);
drop(_sb_pt);
match sb_outcome {
SlotOutcome::Continue => continue,
SlotOutcome::BreakOuter => break 'outer,
}