Removes the per-round phaser barrier that capped the coarse parallel path at ~parity. Workers now FREE-RUN their HW slots (pick slot's thread, run a parallel-safe region on the extracted ctx, writeback+epilogue under the lock, repeat — no waiting on peers); a coordinator thread runs the same housekeeping (coord_pre_round tickers/timers, dispatch_graphics_interrupts, inline-GPU drain, coord_idle_advance) only on a wall-clock cadence, quiescing the workers at a 7-party phaser (via a global `quiesce` AtomicBool, robust vs the earlier epoch-diff which desynced a late-starting worker) so ctx-borrowing housekeeping stays race-free. New: run_execution_parallel_freerun, parallel_region_budget() (default 2048, XENIA_PARALLEL_BUDGET; decoupled from lockstep's 128 so the golden is untouched), scheduler slot_runnable()/any_runnable(), a per-thread `retired=` field in the XENIA_DUMP_SLOTS diagnostic, and dropping the global mmio_access_count region break in the parallel driver (that shared counter, bumped by ANY worker, collapsed every region to one block). STATUS — correct but NOT yet a win. Measured (n=2B, --gpu-inline, JIT): - runs the full 2B and plays the video (4939 draws / 1361 swaps). - BIMODAL: good runs ~21s (edges out lockstep-JIT's ~24s) but a kernel-mutex-contention / guest-spin-wait pathology intermittently latches and makes a run 2-3x slower (~57-87s). Root cause: the single Arc<Mutex<KernelState>> serializes the 6 workers, so the ~4.3x thread-parallelism the workload exposes (measured: work spread across ~5-8 balanced guest threads, top only ~7%) collapses to ~parity. Region-tuning levers (barrier granularity, MMIO break, GPU cadence, sync break) were each measured and none crack it — the real win requires FINE-GRAINED kernel locking. Gated OPT-IN behind XENIA_PARALLEL_FREERUN=1; default --parallel stays the per-round barrier executor. Lockstep untouched (6-config golden byte-identical); parallel_stress_short 20/20 ok under FREERUN=1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>