[iterate-4A] intro-video: fix decode-timeout clock, feeder starvation, k_8 texture decode
Three layered root causes kept ADV.wmv from playing. This lands the first three fixes: the intro video now decodes end-to-end and uploads its YUV planes. The on-screen composite still needs a multi-texture render path (root #3, tracked separately — the shader interpreter binds one texture slot but the YUV->RGB pass samples three). 1. Clock scale (xenia-kernel/state.rs): INSTRUCTIONS_PER_MS 10_000 -> 1_000_000. KeTimeStampBundle tick_count = global_clock / INSTRUCTIONS_PER_MS. At 10_000 (~10 MIPS; global_clock further inflated ~58x by the serialized scheduler summing busy-spin) the movie handler's 2000 ms software-decode deadline (sub_821B4968 @0x821b68c0) tripped on a legitimate ~20M-instruction 720p-YUV420 decode and aborted playback (canary never enters that wait loop). 1_000_000 sits in the validated [~600k, ~2.77M] window that fits both the movie (decode <=2000 ms) and boot (the worker-hub +66 ms gate still elapses before the movie, ~66M instr). XENIA_INSTR_PER_MS overrides. 2. Scheduler fairness (xenia-cpu/scheduler.rs): pick_runnable's equal-priority tiebreak now prefers the incumbent (running_idx) so decrement_quantum's quantum rotation sticks. Previously each round re-picked the lowest index, so co-located equal-priority threads never alternated and the movie's demux feeder (tid24, co-located on hw=1) starved until the STARVE_LIMIT=4096 backstop -- the decode ring never refilled. Priority preemption unaffected. XENIA_INCUMBENT_PICK=0 rolls back. 3. k_8 texture decode (xenia-gpu/texture_cache.rs + xenia-ui/texture_cache_host.rs): the video uploads its YUV420 planes as linear k_8 textures (Y 1280x720, U/V 640x360); with no k_8 decoder ensure_cached rejected them and the frames never reached the GPU. Adds decode_k8 (1 byte/texel expanded to Rgba8Unorm) + the host Rgba8Unorm mapping. Read-only diagnostic probe knobs used to find these remain uncommitted in the working tree. Boot goldens re-baselined in a follow-up commit (the clock change intentionally moves the digest; see tests/golden/README.md). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -119,7 +119,8 @@ impl TextureFormat {
|
||||
pub fn is_host_supported(self) -> bool {
|
||||
matches!(
|
||||
self,
|
||||
TextureFormat::K8888
|
||||
TextureFormat::K8
|
||||
| TextureFormat::K8888
|
||||
| TextureFormat::K565
|
||||
| TextureFormat::Dxt1
|
||||
| TextureFormat::Dxt2_3
|
||||
@@ -394,6 +395,55 @@ pub fn decode_k8888_tiled(
|
||||
Ok(linear)
|
||||
}
|
||||
|
||||
/// Decode a `k_8` (single 8-bit channel) texture into `Rgba8Unorm` bytes,
|
||||
/// replicating the lone component into R, G, B and A so the guest sampler
|
||||
/// reads the value back on whatever channel its swizzle selects (we do not
|
||||
/// apply the fetch-constant swizzle downstream). Used by the intro-video
|
||||
/// YUV420 planes — Y at full resolution and U/V at half — which the guest
|
||||
/// uploads as **linear** `k_8` textures and its pixel shader converts
|
||||
/// YUV→RGB. `k_8` is one byte per texel, so the 32-bit endian swap that the
|
||||
/// packed formats need does not apply here.
|
||||
pub fn decode_k8(
|
||||
key: &TextureKey,
|
||||
mem: &dyn xenia_memory::MemoryAccess,
|
||||
) -> Result<Vec<u8>, DecodeError> {
|
||||
if key.width == 0 || key.height == 0 {
|
||||
return Err(DecodeError::ZeroSize);
|
||||
}
|
||||
let w = key.width as u32;
|
||||
let h = key.height as u32;
|
||||
let pitch_aligned = tiled_address::align_pitch_to_macro_tile(key.pitch_texels as u32);
|
||||
let total_bytes = (pitch_aligned * h) as usize; // 1 byte per texel
|
||||
let raw = read_guest_bytes(mem, key.base_address, total_bytes);
|
||||
if raw.len() < total_bytes {
|
||||
return Err(DecodeError::OutOfBounds);
|
||||
}
|
||||
// Gather one byte per texel into a tightly-packed w*h plane, honoring the
|
||||
// row pitch (linear) or the Xenos tiling (tiled — bytes_per_element = 1).
|
||||
let mut plane = vec![0u8; (w * h) as usize];
|
||||
if key.tiled {
|
||||
if tiled_address::detile_2d(&raw, &mut plane, w, h, pitch_aligned, 1).is_err() {
|
||||
return Err(DecodeError::OutOfBounds);
|
||||
}
|
||||
} else {
|
||||
for y in 0..h as usize {
|
||||
let src = y * (pitch_aligned as usize);
|
||||
let dst = y * (w as usize);
|
||||
plane[dst..dst + w as usize].copy_from_slice(&raw[src..src + w as usize]);
|
||||
}
|
||||
}
|
||||
// Expand to Rgba8Unorm, replicating the component to every channel.
|
||||
let mut rgba = vec![0u8; (w * h * 4) as usize];
|
||||
for (i, &v) in plane.iter().enumerate() {
|
||||
let o = i * 4;
|
||||
rgba[o] = v;
|
||||
rgba[o + 1] = v;
|
||||
rgba[o + 2] = v;
|
||||
rgba[o + 3] = v;
|
||||
}
|
||||
Ok(rgba)
|
||||
}
|
||||
|
||||
/// Decode a DXT-compressed texture to raw block bytes (no format
|
||||
/// conversion — wgpu understands `Bc{1,2,3}RgbaUnorm` natively so the
|
||||
/// GPU does the actual decompression on upload).
|
||||
@@ -603,6 +653,7 @@ impl TextureCache {
|
||||
self.restale_total += 1;
|
||||
}
|
||||
let bytes = match key.format {
|
||||
TextureFormat::K8 => decode_k8(&key, mem)?,
|
||||
TextureFormat::K8888 => decode_k8888_tiled(&key, mem)?,
|
||||
TextureFormat::K565 => decode_k565_tiled(&key, mem)?,
|
||||
TextureFormat::Dxt1 => decode_dxt1_tiled(&key, mem)?,
|
||||
|
||||
Reference in New Issue
Block a user