[iterate-4A] intro-video: fix decode-timeout clock, feeder starvation, k_8 texture decode
Three layered root causes kept ADV.wmv from playing. This lands the first three fixes: the intro video now decodes end-to-end and uploads its YUV planes. The on-screen composite still needs a multi-texture render path (root #3, tracked separately — the shader interpreter binds one texture slot but the YUV->RGB pass samples three). 1. Clock scale (xenia-kernel/state.rs): INSTRUCTIONS_PER_MS 10_000 -> 1_000_000. KeTimeStampBundle tick_count = global_clock / INSTRUCTIONS_PER_MS. At 10_000 (~10 MIPS; global_clock further inflated ~58x by the serialized scheduler summing busy-spin) the movie handler's 2000 ms software-decode deadline (sub_821B4968 @0x821b68c0) tripped on a legitimate ~20M-instruction 720p-YUV420 decode and aborted playback (canary never enters that wait loop). 1_000_000 sits in the validated [~600k, ~2.77M] window that fits both the movie (decode <=2000 ms) and boot (the worker-hub +66 ms gate still elapses before the movie, ~66M instr). XENIA_INSTR_PER_MS overrides. 2. Scheduler fairness (xenia-cpu/scheduler.rs): pick_runnable's equal-priority tiebreak now prefers the incumbent (running_idx) so decrement_quantum's quantum rotation sticks. Previously each round re-picked the lowest index, so co-located equal-priority threads never alternated and the movie's demux feeder (tid24, co-located on hw=1) starved until the STARVE_LIMIT=4096 backstop -- the decode ring never refilled. Priority preemption unaffected. XENIA_INCUMBENT_PICK=0 rolls back. 3. k_8 texture decode (xenia-gpu/texture_cache.rs + xenia-ui/texture_cache_host.rs): the video uploads its YUV420 planes as linear k_8 textures (Y 1280x720, U/V 640x360); with no k_8 decoder ensure_cached rejected them and the frames never reached the GPU. Adds decode_k8 (1 byte/texel expanded to Rgba8Unorm) + the host Rgba8Unorm mapping. Read-only diagnostic probe knobs used to find these remain uncommitted in the working tree. Boot goldens re-baselined in a follow-up commit (the clock change intentionally moves the digest; see tests/golden/README.md). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -980,7 +980,36 @@ impl KernelState {
|
||||
if block == 0 {
|
||||
return;
|
||||
}
|
||||
const INSTRUCTIONS_PER_MS: u64 = 10_000;
|
||||
// tick_count(ms) scale = retired global_clock units per guest ms.
|
||||
//
|
||||
// 1_000_000 models ~1000 MIPS. The prior 10_000 (~10 MIPS, 1 unit≈100ns)
|
||||
// ran the guest clock so fast that the intro-movie's software video
|
||||
// decoder — a legitimate ~20M-instruction / ~720p-YUV420-frame job that
|
||||
// is ~6 ms on the real ~3.2 GHz console — was perceived by the guest as
|
||||
// >2000 ms and tripped the movie handler's 2000 ms decode deadline
|
||||
// (`sub_821B4968` @0x821b68c0), which sets an abort bit and never plays
|
||||
// the video (canary never even enters that wait loop). The all-thread-
|
||||
// SUM `global_clock` is further inflated ~58× by the serialized
|
||||
// scheduler counting busy-spin, so a HW-faithful per-thread rate isn't
|
||||
// usable directly; 1_000_000 sits in the empirically-validated window
|
||||
// [~600k, ~2.77M] that fits BOTH the movie (decode ≤2000 ms) AND boot
|
||||
// (the worker-hub `tick_count + 66 ms` gate still elapses before the
|
||||
// movie at ~66M instr; boot render goldens stay healthy — see
|
||||
// tests/golden/sylpheed_n50m.json). `XENIA_INSTR_PER_MS=<n>` overrides
|
||||
// for diagnostics. See project memory STEP 73-83 for the full trace.
|
||||
let instructions_per_ms: u64 = {
|
||||
use std::sync::OnceLock;
|
||||
static IPM: OnceLock<u64> = OnceLock::new();
|
||||
*IPM.get_or_init(|| {
|
||||
std::env::var("XENIA_INSTR_PER_MS")
|
||||
.ok()
|
||||
.and_then(|s| s.trim().parse::<u64>().ok())
|
||||
.filter(|&n| n > 0)
|
||||
.unwrap_or(1_000_000)
|
||||
})
|
||||
};
|
||||
#[allow(non_snake_case)]
|
||||
let INSTRUCTIONS_PER_MS: u64 = instructions_per_ms;
|
||||
// Perf (Tier-B #5): the bundle is updated once per scheduler round
|
||||
// (~every 7 retired instructions), but the four guest BE memory
|
||||
// writes are ~8.6% of boot-to-splash. `clock` is the retired-
|
||||
@@ -994,13 +1023,13 @@ impl KernelState {
|
||||
// fade-in (3AH-proven vsync-counter driven, NOT this bundle) is
|
||||
// untouched. Throttle threshold is well below 1 ms so no guest-
|
||||
// visible ms boundary is ever skipped.
|
||||
const BUNDLE_QUANTUM: u64 = INSTRUCTIONS_PER_MS / 4; // 2500 units = 0.25 ms
|
||||
let bundle_quantum: u64 = INSTRUCTIONS_PER_MS / 4; // 0.25 ms in units
|
||||
{
|
||||
use std::sync::atomic::Ordering;
|
||||
let last = self.timestamp_bundle_last_clock.load(Ordering::Relaxed);
|
||||
// Always allow the first write (last == u64::MAX sentinel) and any
|
||||
// write that crosses the quantum. Never go backwards.
|
||||
if last != u64::MAX && clock < last.saturating_add(BUNDLE_QUANTUM) {
|
||||
if last != u64::MAX && clock < last.saturating_add(bundle_quantum) {
|
||||
return;
|
||||
}
|
||||
self.timestamp_bundle_last_clock
|
||||
|
||||
Reference in New Issue
Block a user