Cracked how the game maps a movie to its voice track and reimplemented it,
fixing wrong/short voices for RT and hokyu cutscenes.
## Mechanism (RE'd from a Canary file-I/O trace, user-confirmed by ear)
The movie voices are ONE continuous XMA stream, physically chunked into
`<lang>\Movie\VOICE_*.slb` (and `\etc\VOICE_D_*.slb`) sound.pak TOC entries
whose boundaries do NOT align with the cutscene cues. A single cue routinely
spans two `.slb` chunks, so a `.slb` does not necessarily hold the track its
name claims (they are off-by-one vs the cues). Each cue ends at an inline
trailer descriptor `(sound_id:u32be, 0x11, …)` whose id repeats at +0x800.
Resolution chain (all derivable on-disc, no guest-code RE):
1. movie -> cue token (ADVERTISE_MOVIE manifest VOICETRACK)
2. cue token -> sound-id (master sound registry, tables.pak
a2c8c185=eng / b04238e0=jpn, schema 0x13cb84ba)
3. sound-id -> [start,end) scan the stream for trailer(id) and its
predecessor; cue audio = the bytes between,
read continuously across chunk boundaries.
Validated: decoded voice length matches each movie within ~0.4s, including
cross-boundary cues (ADV/RT01A/RT01B/RT01C/S00A/S01A + the 5 bound hokyu).
## Changes
- sylpheed-formats/src/movie_voice.rs (new): registry_voice_ids(),
find_descriptor(), find_descriptor_before() (gap-tolerant predecessor).
Unit-tested.
- viewer iso_loader.rs: resolve_movie_voice_region() + decode_voice_region()
(continuous read via existing read_segment_range + to_xma_riffs). Voice/
audio requests now region-first; a `\Movie\` token that fails region stays
SILENT (never plays the wrong off-by-one chunk). Removed the old
movie_is_demo_based suppress hack.
- Anchor tries subdirs Movie/etc/Voice so hokyu `\etc\VOICE_D_*` resolve too.
- Examples resolve_all/validate_cues/hokyu_final document + validate the pipeline.
## Status
- RT / S / ADV cutscene voices: CORRECT (user-confirmed).
- 5 manifest-bound hokyu (VOICE_D_450-454): CORRECT (user-confirmed).
## OPEN (handoff)
13 of 18 hokyu movies have NO manifest VOICETRACK; the game reuses the 5
recordings across stages via a code rule. `hokyu_fallback_token()` currently
guesses by category (ship LS/DS x source carrier-A/tanker-H: LS/A->450,
LS/H->453, DS/A->452, DS/H->454). USER REPORTS THIS IS SOMETIMES WRONG.
- LS/carrier has two takes (450 from s02A, 451 from s09A); the early-vs-late
split is unknown — unbound LS/A currently all default to 450.
- Definitive fix = Canary `--phase_a_fileio_only` file.read trace of an unbound
hokyu (e.g. hokyu_LS_s03A) to see which VOICE_D the game actually streams,
then encode the real rule. Same method that cracked the RT mapping.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
143 lines
5.5 KiB
Rust
143 lines
5.5 KiB
Rust
//! Cutscene-voice resolution: movie → sound-id → continuous-stream byte region.
|
|
//!
|
|
//! Reverse-engineered 2026-07-22 (Canary file-I/O trace, user-confirmed by ear).
|
|
//! The movie voices form **one continuous XMA stream**, physically chunked into
|
|
//! `<lang>\Movie\VOICE_*.slb` `sound.pak` TOC entries whose boundaries do **not**
|
|
//! align with the cutscene cues — a single cue routinely spans two `.slb` chunks,
|
|
//! so a `.slb` does not necessarily hold the track its name claims. Each cue's
|
|
//! audio ends at an inline **trailer descriptor** `(sound_id: u32be, 0x11, …)`
|
|
//! whose id repeats at +0x800. Cue N's audio = the stream bytes
|
|
//! `[descriptor(N-1) .. descriptor(N)]`, read continuously across chunk
|
|
//! boundaries. Verified: decoded voice length matches each movie within ~0.3 s,
|
|
//! including cross-boundary cues (RT01B/RT01C span `VOICE_ADV`→`VOICE_RT01A`→…).
|
|
//!
|
|
//! Resolution chain:
|
|
//! 1. movie → cue name (`ADVERTISE_MOVIE` manifest `VOICETRACK`, e.g. `VOICE_RT01A`)
|
|
//! 2. cue name → sound-id (master sound registry, [`registry_voice_ids`], → 1601)
|
|
//! 3. sound-id → `[start, end]` global offsets via [`find_descriptor`] over the
|
|
//! `sound.pNN` stream (the caller supplies the bytes + the physical anchor).
|
|
|
|
use std::collections::HashMap;
|
|
|
|
/// Parse the master sound registry — the large per-language `tables.pak` IDXD
|
|
/// entry (schema `0x13cb84ba`) that carries the `<lang>\Movie\VOICE_*.slb` paths —
|
|
/// into `cue-name → sound-id` from its `NAME\0 ID\0` string pool
|
|
/// (e.g. `VOICE_ADV`→1600, `VOICE_RT01A`→1601).
|
|
pub fn registry_voice_ids(registry: &[u8]) -> HashMap<String, u32> {
|
|
let mut toks: Vec<&[u8]> = Vec::new();
|
|
let mut start = 0usize;
|
|
for i in 0..registry.len() {
|
|
if registry[i] == 0 {
|
|
if i > start {
|
|
toks.push(®istry[start..i]);
|
|
}
|
|
start = i + 1;
|
|
}
|
|
}
|
|
let mut out = HashMap::new();
|
|
for w in toks.windows(2) {
|
|
let (name, num) = (w[0], w[1]);
|
|
if name.starts_with(b"VOICE_")
|
|
&& !num.is_empty()
|
|
&& num.iter().all(u8::is_ascii_digit)
|
|
{
|
|
if let (Ok(n), Ok(id)) = (
|
|
std::str::from_utf8(name),
|
|
std::str::from_utf8(num).unwrap().parse::<u32>(),
|
|
) {
|
|
out.entry(n.to_string()).or_insert(id);
|
|
}
|
|
}
|
|
}
|
|
out
|
|
}
|
|
|
|
/// The constant marker following a cue's sound-id in its inline trailer.
|
|
const DESC_MARK: u32 = 0x11;
|
|
/// The trailer repeats its id this many bytes later; used to reject a coincidental
|
|
/// `(id, 0x11)` pair inside XMA audio (probability of a false match is ~2⁻⁶⁴).
|
|
const DESC_REPEAT: usize = 0x800;
|
|
|
|
/// Byte offset (within `buf`, a slice of the continuous voice stream) of cue
|
|
/// `id`'s trailer descriptor: the first big-endian `(id, 0x11)` pair whose id
|
|
/// repeats at `+0x800`. `None` if absent in the slice.
|
|
pub fn find_descriptor(buf: &[u8], id: u32) -> Option<usize> {
|
|
if buf.len() < DESC_REPEAT + 4 {
|
|
return None;
|
|
}
|
|
let be = |o: usize| u32::from_be_bytes([buf[o], buf[o + 1], buf[o + 2], buf[o + 3]]);
|
|
let end = buf.len() - (DESC_REPEAT + 4);
|
|
let mut o = 0;
|
|
while o <= end {
|
|
if be(o) == id && be(o + 4) == DESC_MARK && be(o + DESC_REPEAT) == id {
|
|
return Some(o);
|
|
}
|
|
o += 4;
|
|
}
|
|
None
|
|
}
|
|
|
|
/// Plausible sound-id range for a trailer (all real cue ids are well under this).
|
|
const ID_MAX: u32 = 0x1_0000;
|
|
|
|
/// Offset of the nearest cue trailer strictly **before** `before` — any plausible
|
|
/// id. Used to bound a cue's START when its `id-1` predecessor doesn't exist
|
|
/// (the sound-id sequence has gaps, e.g. `VOICE_D_453`=7142 → `VOICE_D_454`=7188).
|
|
/// Scans downward and returns the first (largest-offset) valid `(id, 0x11)` with
|
|
/// the `+0x800` id-repeat. `None` if there is no trailer below `before`.
|
|
pub fn find_descriptor_before(buf: &[u8], before: usize) -> Option<usize> {
|
|
if before < 4 {
|
|
return None;
|
|
}
|
|
let cap = buf.len().checked_sub(DESC_REPEAT + 4)?;
|
|
let be = |o: usize| u32::from_be_bytes([buf[o], buf[o + 1], buf[o + 2], buf[o + 3]]);
|
|
// Start strictly below `before` (4-aligned) so the target's own trailer is
|
|
// never returned as its predecessor.
|
|
let mut o = (before - 1).min(cap) & !3;
|
|
loop {
|
|
let id = be(o);
|
|
if id >= 1 && id < ID_MAX && be(o + 4) == DESC_MARK && be(o + DESC_REPEAT) == id {
|
|
return Some(o);
|
|
}
|
|
if o < 4 {
|
|
return None;
|
|
}
|
|
o -= 4;
|
|
}
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
#[test]
|
|
fn registry_parses_name_id_pairs() {
|
|
let mut b = Vec::new();
|
|
for tok in [
|
|
b"VOICE_ADV".as_slice(),
|
|
b"1600",
|
|
b"VOICE_RT01A",
|
|
b"1601",
|
|
b"NOT_A_VOICE",
|
|
b"9",
|
|
] {
|
|
b.extend_from_slice(tok);
|
|
b.push(0);
|
|
}
|
|
let ids = registry_voice_ids(&b);
|
|
assert_eq!(ids.get("VOICE_ADV"), Some(&1600));
|
|
assert_eq!(ids.get("VOICE_RT01A"), Some(&1601));
|
|
assert_eq!(ids.get("NOT_A_VOICE"), None);
|
|
}
|
|
|
|
#[test]
|
|
fn find_descriptor_matches_repeat() {
|
|
let mut b = vec![0u8; 0x900];
|
|
b[0..4].copy_from_slice(&1601u32.to_be_bytes());
|
|
b[4..8].copy_from_slice(&0x11u32.to_be_bytes());
|
|
b[DESC_REPEAT..DESC_REPEAT + 4].copy_from_slice(&1601u32.to_be_bytes());
|
|
assert_eq!(find_descriptor(&b, 1601), Some(0));
|
|
assert_eq!(find_descriptor(&b, 1600), None);
|
|
}
|
|
}
|