[iterate-4A] intro-video: fix decode-timeout clock, feeder starvation, k_8 texture decode

Three layered root causes kept ADV.wmv from playing. This lands the first
three fixes: the intro video now decodes end-to-end and uploads its YUV
planes. The on-screen composite still needs a multi-texture render path
(root #3, tracked separately — the shader interpreter binds one texture slot
but the YUV->RGB pass samples three).

1. Clock scale (xenia-kernel/state.rs): INSTRUCTIONS_PER_MS 10_000 -> 1_000_000.
   KeTimeStampBundle tick_count = global_clock / INSTRUCTIONS_PER_MS. At 10_000
   (~10 MIPS; global_clock further inflated ~58x by the serialized scheduler
   summing busy-spin) the movie handler's 2000 ms software-decode deadline
   (sub_821B4968 @0x821b68c0) tripped on a legitimate ~20M-instruction
   720p-YUV420 decode and aborted playback (canary never enters that wait
   loop). 1_000_000 sits in the validated [~600k, ~2.77M] window that fits
   both the movie (decode <=2000 ms) and boot (the worker-hub +66 ms gate
   still elapses before the movie, ~66M instr). XENIA_INSTR_PER_MS overrides.

2. Scheduler fairness (xenia-cpu/scheduler.rs): pick_runnable's equal-priority
   tiebreak now prefers the incumbent (running_idx) so decrement_quantum's
   quantum rotation sticks. Previously each round re-picked the lowest index,
   so co-located equal-priority threads never alternated and the movie's demux
   feeder (tid24, co-located on hw=1) starved until the STARVE_LIMIT=4096
   backstop -- the decode ring never refilled. Priority preemption unaffected.
   XENIA_INCUMBENT_PICK=0 rolls back.

3. k_8 texture decode (xenia-gpu/texture_cache.rs + xenia-ui/texture_cache_host.rs):
   the video uploads its YUV420 planes as linear k_8 textures (Y 1280x720,
   U/V 640x360); with no k_8 decoder ensure_cached rejected them and the frames
   never reached the GPU. Adds decode_k8 (1 byte/texel expanded to Rgba8Unorm)
   + the host Rgba8Unorm mapping.

Read-only diagnostic probe knobs used to find these remain uncommitted in the
working tree. Boot goldens re-baselined in a follow-up commit (the clock change
intentionally moves the digest; see tests/golden/README.md).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-07-01 20:34:01 +02:00
parent 3e17d37b4e
commit 645feb8f5b
4 changed files with 117 additions and 5 deletions

View File

@@ -119,7 +119,8 @@ impl TextureFormat {
pub fn is_host_supported(self) -> bool {
matches!(
self,
TextureFormat::K8888
TextureFormat::K8
| TextureFormat::K8888
| TextureFormat::K565
| TextureFormat::Dxt1
| TextureFormat::Dxt2_3
@@ -394,6 +395,55 @@ pub fn decode_k8888_tiled(
Ok(linear)
}
/// Decode a `k_8` (single 8-bit channel) texture into `Rgba8Unorm` bytes,
/// replicating the lone component into R, G, B and A so the guest sampler
/// reads the value back on whatever channel its swizzle selects (we do not
/// apply the fetch-constant swizzle downstream). Used by the intro-video
/// YUV420 planes — Y at full resolution and U/V at half — which the guest
/// uploads as **linear** `k_8` textures and its pixel shader converts
/// YUV→RGB. `k_8` is one byte per texel, so the 32-bit endian swap that the
/// packed formats need does not apply here.
pub fn decode_k8(
key: &TextureKey,
mem: &dyn xenia_memory::MemoryAccess,
) -> Result<Vec<u8>, DecodeError> {
if key.width == 0 || key.height == 0 {
return Err(DecodeError::ZeroSize);
}
let w = key.width as u32;
let h = key.height as u32;
let pitch_aligned = tiled_address::align_pitch_to_macro_tile(key.pitch_texels as u32);
let total_bytes = (pitch_aligned * h) as usize; // 1 byte per texel
let raw = read_guest_bytes(mem, key.base_address, total_bytes);
if raw.len() < total_bytes {
return Err(DecodeError::OutOfBounds);
}
// Gather one byte per texel into a tightly-packed w*h plane, honoring the
// row pitch (linear) or the Xenos tiling (tiled — bytes_per_element = 1).
let mut plane = vec![0u8; (w * h) as usize];
if key.tiled {
if tiled_address::detile_2d(&raw, &mut plane, w, h, pitch_aligned, 1).is_err() {
return Err(DecodeError::OutOfBounds);
}
} else {
for y in 0..h as usize {
let src = y * (pitch_aligned as usize);
let dst = y * (w as usize);
plane[dst..dst + w as usize].copy_from_slice(&raw[src..src + w as usize]);
}
}
// Expand to Rgba8Unorm, replicating the component to every channel.
let mut rgba = vec![0u8; (w * h * 4) as usize];
for (i, &v) in plane.iter().enumerate() {
let o = i * 4;
rgba[o] = v;
rgba[o + 1] = v;
rgba[o + 2] = v;
rgba[o + 3] = v;
}
Ok(rgba)
}
/// Decode a DXT-compressed texture to raw block bytes (no format
/// conversion — wgpu understands `Bc{1,2,3}RgbaUnorm` natively so the
/// GPU does the actual decompression on upload).
@@ -603,6 +653,7 @@ impl TextureCache {
self.restale_total += 1;
}
let bytes = match key.format {
TextureFormat::K8 => decode_k8(&key, mem)?,
TextureFormat::K8888 => decode_k8888_tiled(&key, mem)?,
TextureFormat::K565 => decode_k565_tiled(&key, mem)?,
TextureFormat::Dxt1 => decode_dxt1_tiled(&key, mem)?,