Extends B1's forward-linear superblock into a full same-page CFG: chains
through conditional bcx (BOTH directions) and loop back-edges (a successor
already in the chain jumps to its existing IR block), not just unconditional
bx + fall-through. This roughly doubles the block-merge ratio (run_block calls
145.7M interp -> 116.2M B1 (1.25x) -> 74.7M B2 (1.95x)) and halves the
run_superblock dispatch bucket (30.4% interp -> 26.2% B1 -> 16.2% B2).
Mechanism:
- enumerate_chain is now a two-pass CFG builder: BFS the reachable, chainable
(same-page / covered / non-thunk) blocks deduped by start PC (so a back-edge
becomes a loop, not a new node), capped at MAX_CHAIN_NODES=64; then resolve
each terminator to Succ edges (node indices or Exit). bx/fall -> One; bcx ->
Two{taken,fall} (each side Node or Exit); bclrx/bcctrx/uncovered -> Exit.
- emit_bcx extracted from emit_op's bcx arm: computes the taken predicate ONCE
(decrementing CTR at most once) and returns it; emit_node_body captures it so
the two-way boundary branches on the SAME value (recomputing would double the
CTR decrement). Single-block compile is unchanged (ignores the return).
- No Cranelift Variables/phi needed for the loop CFG: the budget/MMIO state
lives in memory (cycle_count, *mmio_count), so emit_stop reloads it and
compares against entry-sampled constants (cycle_entry/mmio_entry), which the
entry block dominates across back-edges. Budget = (cycle_count - cycle_entry)
>= remaining, correct across loop iterations; a native loop spins until the
budget is spent then hands back — same bound as the runner. MMIO check still
skipped for load/store-free blocks. seal_all_blocks handles the arbitrary CFG.
- Degenerate guard fixed: compile single-block only when the ENTRY has no
chainable successor (Succ::Exit) — a one-node self-loop is NOT degenerate.
Validation: 3 new unit tests (bcx always-taken chains to target; bcx not-taken
falls through; self-loop runs natively until budget then stops — all bit-equal
to the interpreter over the same PCs) -> 7/7 jit tests pass. Diff MISMATCHES=0
over 47.95M blocks (the emit_bcx refactor didn't perturb op bodies).
XENIA_JIT_CHAIN=1 movie plays: tid25 resumes (decode handoff / interleaving
preserved under the deeper chaining).
Perf finding (honest): despite ~2x merge and half the dispatch, wall is ~parity
with B1 (both ~-6% vs interp same-batch). step_block MIPS dropped (~82 vs B1
~102) because the deeper/bigger native superblocks cost more per boundary (B1's
compile-time-constant budget became a runtime cycle_count reload, needed for
loops) and Cranelift's default regalloc spills more in larger functions. The
dispatch savings are real but offset by native-code overhead that grows with
chain length. Converting the extra chaining into wall time now needs
host-register allocation (tighter code) — the clear next lever. Env gate,
diff-off-under-chain, and probe/mem-watch exclusivity unchanged from B1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>